{"thread":{"id":"52631","subject":"[RFC] Extending git-replace","startedAt":"2020-01-14T05:33:34Z","lastAt":"2020-01-16T03:30:24Z","messageCount":9,"participants":["Kaushik Srenevasan","Elijah Newren","David Turner","Jonathan Tan"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"389740","messageId":"CAN_uzpK1m42J19Xi8oc3Dwmhk9qF1n4hFrBVqD+0RHZuZ_15Jw@mail.gmail.com","threadId":"52631","inReplyTo":null,"subject":"[RFC] Extending git-replace","fromName":"Kaushik Srenevasan","fromEmail":"kaushik@twitter.com","sentAt":"2020-01-14T05:33:19Z","receivedAt":"2020-01-14T05:33:34Z","isPatch":false,"sender":{"key":"kaushik@twitter.com","avatar":null},"body":"We’ve been trying to get rid of objects larger than a certain size\nfrom one of our repositories that contains tens of thousands of\nbranches and hundreds of thousands of commits. While we’re able to\naccomplish this using BFG[0] , it results in ~ 90% of the repository’s\nhistory being rewritten. This presents the following problems\n1. There are various systems (Phabricator for one) that use the commit\nhash as a key in various databases. Rewriting history will require\nthat we update all of these systems.\n2. We’ll have to force everyone to reclone a copy of this repository.\n\nI was looking through the git code base to see if there is a way\naround it when I chanced upon `git-replace`. While the basic idea of\n`git-replace` is what I am looking for, it doesn’t quite fit the bill\ndue to the `--no-replace-objects` switch, the `GIT_NO_REPLACE_OBJECTS`\nenvironment variable, and `--no-replace-objects` being the default for\ncertain git commands. Namely fsck, upload-pack, pack/unpack-objects,\nprune and index-pack. That Git may still try to load a replaced object\nwhen a git command is run with the `--no-replace-objects` option\nprevents me from removing it from the ODB permanently. Not being able\nto run prune and fsck on a repository where we’ve deleted the object\nthat’s been replaced with git-replace effectively rules this option\nout for us.\n\nA feature that allowed such permanent replacement (say a\n`git-blacklist` or a `git-replace --blacklist`) might work as follows:\n1. Blacklisted objects are stored as references under a new namespace\n-- `refs/blacklist`.\n2. The object loader unconditionally translates a blacklisted OID into\nthe OID it’s been replaced with.\n3. The `+refs/blacklist/*:refs/blacklist/*` refspec is implicitly\nalways a part of fetch and push transactions.\n\nThis essentially turns the blacklist references namespace into an\nadditional piece of metadata that gets transmitted to a client when a\nrepository is cloned and is kept updated automatically.\n\nI’ve been playing around with a prototype I wrote and haven’t observed\nany breakage yet. I’m writing to seek advice on this approach and to\nunderstand if this is something (if not in its current form, some\nversion of it) that has a chance of making it into the product if we\nwere to implement it. Happy to write up a more detailed design and\nshare my prototype as a starting point for discussion.\n\n                                           -- Kaushik\n\n[0] https://rtyley.github.io/bfg-repo-cleaner/\n"},{"id":"389741","messageId":"CABPp-BGy3qu_Rd4epore0wLyoh1fg0UH5EAV27shKJ=kLWX4FA@mail.gmail.com","threadId":"52631","inReplyTo":"CAN_uzpK1m42J19Xi8oc3Dwmhk9qF1n4hFrBVqD+0RHZuZ_15Jw@mail.gmail.com","subject":"Re: [RFC] Extending git-replace","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2020-01-14T06:55:15Z","receivedAt":"2020-01-14T06:55:27Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"Hi Kaushik,\n\nOn Mon, Jan 13, 2020 at 9:39 PM Kaushik Srenevasan <kaushik@twitter.com> wrote:\n>\n> We’ve been trying to get rid of objects larger than a certain size\n> from one of our repositories that contains tens of thousands of\n> branches and hundreds of thousands of commits. While we’re able to\n> accomplish this using BFG[0] , it results in ~ 90% of the repository’s\n> history being rewritten. This presents the following problems\n> 1. There are various systems (Phabricator for one) that use the commit\n> hash as a key in various databases. Rewriting history will require\n> that we update all of these systems.\n\nNot necessarily...\n\n> 2. We’ll have to force everyone to reclone a copy of this repository.\n\nTrue.\n\n> I was looking through the git code base to see if there is a way\n> around it when I chanced upon `git-replace`. While the basic idea of\n> `git-replace` is what I am looking for, it doesn’t quite fit the bill\n> due to the `--no-replace-objects` switch, the `GIT_NO_REPLACE_OBJECTS`\n> environment variable, and `--no-replace-objects` being the default for\n> certain git commands. Namely fsck, upload-pack, pack/unpack-objects,\n> prune and index-pack. That Git may still try to load a replaced object\n> when a git command is run with the `--no-replace-objects` option\n> prevents me from removing it from the ODB permanently. Not being able\n> to run prune and fsck on a repository where we’ve deleted the object\n> that’s been replaced with git-replace effectively rules this option\n> out for us.\n>\n> A feature that allowed such permanent replacement (say a\n> `git-blacklist` or a `git-replace --blacklist`) might work as follows:\n> 1. Blacklisted objects are stored as references under a new namespace\n> -- `refs/blacklist`.\n> 2. The object loader unconditionally translates a blacklisted OID into\n> the OID it’s been replaced with.\n> 3. The `+refs/blacklist/*:refs/blacklist/*` refspec is implicitly\n> always a part of fetch and push transactions.\n>\n> This essentially turns the blacklist references namespace into an\n> additional piece of metadata that gets transmitted to a client when a\n> repository is cloned and is kept updated automatically.\n>\n> I’ve been playing around with a prototype I wrote and haven’t observed\n> any breakage yet. I’m writing to seek advice on this approach and to\n> understand if this is something (if not in its current form, some\n> version of it) that has a chance of making it into the product if we\n> were to implement it. Happy to write up a more detailed design and\n> share my prototype as a starting point for discussion.\n\nI'll get back to this in a minute, but wanted to point out a couple\nother ideas for consideration:\n\n1) You can rewrite history, and then use replace references to map old\ncommit IDs to new commit IDs.  This allows anyone to continue using\nold commit IDs (which aren't even part of the new repository anymore)\nin git commands and git automatically uses and shows the new commit\nIDs.  No problems with fsck or prune or fetch either.  Creating these\nreplace refs is fairly simple if your repository rewriting program\n(e.g. git-filter-repo or BFG Repo Cleaner) provides a mapping of old\nIDs to new IDs, and if you are using git-filter-repo it even creates\nthe replace refs for you.  (The one downside is that you can't use\nabbreviated refs to refer to replace refs, thus you can't use\nabbreviated old commit IDs in this scheme.)\n\nThe downside is that various repository hosting tools ignore replace\nrefs.  Thus if you try to browse to a commit in the web UI of Gerrit\nor GitHub using the old commit IDs, it'll just show you a commit not\nfound page.  Phabricator and GitLab may well be the same (haven't\ntried yet).  However, teaching these tools to pay attention to replace\nrefs would make this simple mechanism for rewriting feel close to\nseamless other than asking people to reclone.  It's possible that\nteaching the Webby tools to pay attention to replace refs might not be\ntoo difficult, at least for the open source systems, though I admit I\nhaven't dug into it myself.\n\n2) Some folks might be okay with a clone that won't pass fsck or\nprune, at least in special circumstances.  We're actually doing that\non purpose to deal with one of our large repositories.  We don't\nprovide that to normal developers, but we do use \"cheap, fake clones\"\nin our CI systems.  These slim clones have 99% of all objects, but\nhappen to be missing the really big ones, resulting in only needing\n1/7 of the time to download.  (And no, don't try to point out shallow\nclones to me.  I hate those things, they're an awful hack, *and* they\ndon't work for us.  It's nice getting all commit history, all trees,\nand most blobs including all for at least the last two years while\nstill saving lots of space.)\n\n[For the curious, I did make a simple script to create these \"cheap,\nfake clones\" for repositories of interest.  See\nhttps://github.com/newren/sequester-old-big-blobs.  But they are\ndefinitely a hack with some sharp corners, with failing fsck and\nprunes only being part of the story.]\n\n\n3) Back to your idea...\n\nWhat you're proposing actually sounds very similar to partial clones,\nwhose idea is to make it okay to download a subset of history.  The\nprimary problems with partial clones are (a) they are still under\ndevelopment and are just experimental, (b) they are currently\nimplemented with a \"promisor\" mode, meaning that if a command tries to\nrun over any piece of missing data then the command pauses while the\nobjects are downloaded from the server.  I want an offline mode (even\nif I'm online) where only explicit downloading from the server (clone,\nfetch, etc.) occurs.\n\nInstead of inventing yet another partial-clone-like thing, it'd be\nnice if your new mechanism could just be implemented in terms of\npartial clones, extending them as you need.  I don't like the idea of\nsupporting multiple competing implementations of partial clones\nwithing git.git, but if it's just some extensions of the existing\ncapability then it sounds great.  But you may want to talk with\nJonathan Tan if you want to go this route (cc'd), since he's the\npartial clone expert.\n"},{"id":"389750","messageId":"d4361b6d34513a3fdefa564d10269f60d4732208.camel@novalis.org","threadId":"52631","inReplyTo":"CAN_uzpK1m42J19Xi8oc3Dwmhk9qF1n4hFrBVqD+0RHZuZ_15Jw@mail.gmail.com","subject":"Re: [RFC] Extending git-replace","fromName":"David Turner","fromEmail":"novalis@novalis.org","sentAt":"2020-01-14T18:19:50Z","receivedAt":"2020-01-14T18:20:01Z","isPatch":false,"sender":{"key":"novalis@novalis.org","avatar":"https://avatars.githubusercontent.com/u/77003?v=4"},"body":"On Mon, 2020-01-13 at 21:33 -0800, Kaushik Srenevasan wrote:\n> A feature that allowed such permanent replacement (say a\n> `git-blacklist` or a `git-replace --blacklist`) might work as\n> follows:\n> 1. Blacklisted objects are stored as references under a new namespace\n> -- `refs/blacklist`.\n> 2. The object loader unconditionally translates a blacklisted OID\n> into\n> the OID it’s been replaced with.\n> 3. The `+refs/blacklist/*:refs/blacklist/*` refspec is implicitly\n> always a part of fetch and push transactions.\n\nThere are definitely some security implications here. I assume that\nthere's a config on the client to trust the server's refs/blacklist/*,\nand that the documentation for this explains that it allows your repo\nto be messed with in quite dangerous ways.  And on the server, I would\nexpect that only privileged users could push to refs/blacklist/*\n\nTo Elijah's point that this is related to partial clones and promisors,\nI think Kaushik's idea is subtly different in that it involves\nreplacements, while promisors try to offer a seamless experience.  I\nwonder whether Kaushik actually needs the replacement functionality?  \n\nThat is, would it be sufficient if every replaced file were replaced\nwith the exact text \"me caga en la leche\" instead of a custom hand-\ncrafted replacement?  I guess it's a bit complicated because while\nthat's a reasonable blob, it's not a valid commit.  So maybe this\nmechanism would be limited to blobs.  I thought about whether we could\na different flavor of replacement for commits, but those generally have\nto be custom because they each have different parents. \n\nAnd if that would be sufficient, could promisors be used for this?  I\ndon't know how those interact with fsck and the other commands that\nyou're worried about.  Basically, the idea would be to use most of the\nexisting promisor code, and then have a mode where instead of visiting\nthe promisor, we just always return \"me caga en la leche\" (and this\ndoes not have its SHA checked, of course).\n\nThis could work together with some sort refs/blacklist mechanism to\nenable the server to choose which objects the client replaces.\n\n\n"},{"id":"389762","messageId":"20200114190339.120452-1-jonathantanmy@google.com","threadId":"52631","inReplyTo":"d4361b6d34513a3fdefa564d10269f60d4732208.camel@novalis.org","subject":"Re: [RFC] Extending git-replace","fromName":"Jonathan Tan","fromEmail":"jonathantanmy@google.com","sentAt":"2020-01-14T19:03:39Z","receivedAt":"2020-01-14T19:03:45Z","isPatch":false,"sender":{"key":"jonathantanmy@fastmail.com","avatar":null},"body":"> That is, would it be sufficient if every replaced file were replaced\n> with the exact text \"me caga en la leche\" instead of a custom hand-\n> crafted replacement?  I guess it's a bit complicated because while\n> that's a reasonable blob, it's not a valid commit.  So maybe this\n> mechanism would be limited to blobs.  I thought about whether we could\n> a different flavor of replacement for commits, but those generally have\n> to be custom because they each have different parents.\n\nSince the original email just discussed blobs, I'll confine myself to\ndiscussing blobs. (Commits are trickier, as you said.)\n\n> And if that would be sufficient, could promisors be used for this?  I\n> don't know how those interact with fsck and the other commands that\n> you're worried about.  Basically, the idea would be to use most of the\n> existing promisor code, and then have a mode where instead of visiting\n> the promisor, we just always return \"me caga en la leche\" (and this\n> does not have its SHA checked, of course).\n\nMissing promisor objects do not prevent fsck from passing - this is part\nof the original design (any packfiles we download from the specifically\ndesignated promisor remote are marked as such, and any objects that the\nobjects in the packfile refer to are considered OK to be missing).\n\nCurrently, when a missing object is read, it is first fetched (there are\nsome more details that I can go over if you have any specific\nquestions). What you're suggesting here is to return a fake blob with\nwrong hash - I haven't looked at all the callers of read-object\nfunctions in detail, but I don't think all of them are ready for such a\nbehavioral change. Maybe it would be sufficient to just make this work\nin a more limited scope (e.g. checkout only - and if we need different\nreplacement blobs for different object IDs, maybe we could have\nsomething similar to the clean/smudge filters).\n\n> This could work together with some sort refs/blacklist mechanism to\n> enable the server to choose which objects the client replaces.\n\nIn the original email, Kaushik mentioned objects larger than a certain\nsize - we already have support for that (--filter=blob:limit=1000000,\nfor example). Having said that, Git is already able to tolerate any\nexclusion (of tree or blob) from the server - we already need this in\norder to support changing of filters, for example.\n"},{"id":"389763","messageId":"20200114191156.122421-1-jonathantanmy@google.com","threadId":"52631","inReplyTo":"CABPp-BGy3qu_Rd4epore0wLyoh1fg0UH5EAV27shKJ=kLWX4FA@mail.gmail.com","subject":"Re: [RFC] Extending git-replace","fromName":"Jonathan Tan","fromEmail":"jonathantanmy@google.com","sentAt":"2020-01-14T19:11:56Z","receivedAt":"2020-01-14T19:12:02Z","isPatch":false,"sender":{"key":"jonathantanmy@fastmail.com","avatar":null},"body":"> 2) Some folks might be okay with a clone that won't pass fsck or\n> prune, at least in special circumstances.  We're actually doing that\n> on purpose to deal with one of our large repositories.  We don't\n> provide that to normal developers, but we do use \"cheap, fake clones\"\n> in our CI systems.  These slim clones have 99% of all objects, but\n> happen to be missing the really big ones, resulting in only needing\n> 1/7 of the time to download.  (And no, don't try to point out shallow\n> clones to me.  I hate those things, they're an awful hack, *and* they\n> don't work for us.  It's nice getting all commit history, all trees,\n> and most blobs including all for at least the last two years while\n> still saving lots of space.)\n> \n> [For the curious, I did make a simple script to create these \"cheap,\n> fake clones\" for repositories of interest.  See\n> https://github.com/newren/sequester-old-big-blobs.  But they are\n> definitely a hack with some sharp corners, with failing fsck and\n> prunes only being part of the story.]\n\nIf you want to reduce the sharpness of the corners, it might be possible\nto designate the pack with missing blobs as a promisor pack (add a\n.promisor file - which is just like the .keep file except\ns/keep/promisor/) and a fake promisor remote. That will make fsck and\nrepack (GC) work.\n\n> 3) Back to your idea...\n> \n> What you're proposing actually sounds very similar to partial clones,\n> whose idea is to make it okay to download a subset of history.  The\n> primary problems with partial clones are (a) they are still under\n> development and are just experimental, (b) they are currently\n> implemented with a \"promisor\" mode, meaning that if a command tries to\n> run over any piece of missing data then the command pauses while the\n> objects are downloaded from the server.  I want an offline mode (even\n> if I'm online) where only explicit downloading from the server (clone,\n> fetch, etc.) occurs.\n\nDavid Turner had an idea of what could be done (instead of fetching) in\nsuch an offline mode [1], so I replied there.\n\n[1] https://lore.kernel.org/git/d4361b6d34513a3fdefa564d10269f60d4732208.camel@novalis.org/\n\n> Instead of inventing yet another partial-clone-like thing, it'd be\n> nice if your new mechanism could just be implemented in terms of\n> partial clones, extending them as you need.  I don't like the idea of\n> supporting multiple competing implementations of partial clones\n> withing git.git, but if it's just some extensions of the existing\n> capability then it sounds great.  But you may want to talk with\n> Jonathan Tan if you want to go this route (cc'd), since he's the\n> partial clone expert.\n\nAh, thanks for your kind words.\n"},{"id":"389776","messageId":"CABPp-BF8OHoHo73doekKzf0CmO09_PyAfe4q__DvoftQ+BeY2w@mail.gmail.com","threadId":"52631","inReplyTo":"20200114190339.120452-1-jonathantanmy@google.com","subject":"Re: [RFC] Extending git-replace","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2020-01-14T20:39:49Z","receivedAt":"2020-01-14T20:40:02Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Tue, Jan 14, 2020 at 11:05 AM Jonathan Tan <jonathantanmy@google.com> wrote:\n>\n> > That is, would it be sufficient if every replaced file were replaced\n> > with the exact text \"me caga en la leche\" instead of a custom hand-\n> > crafted replacement?  I guess it's a bit complicated because while\n> > that's a reasonable blob, it's not a valid commit.  So maybe this\n> > mechanism would be limited to blobs.  I thought about whether we could\n> > a different flavor of replacement for commits, but those generally have\n> > to be custom because they each have different parents.\n>\n> Since the original email just discussed blobs, I'll confine myself to\n> discussing blobs. (Commits are trickier, as you said.)\n>\n> > And if that would be sufficient, could promisors be used for this?  I\n> > don't know how those interact with fsck and the other commands that\n> > you're worried about.  Basically, the idea would be to use most of the\n> > existing promisor code, and then have a mode where instead of visiting\n> > the promisor, we just always return \"me caga en la leche\" (and this\n> > does not have its SHA checked, of course).\n\nMaybe; it doesn't necessarily need to be the same object returned, and\nthese replacements could be user-specified via replace refs...\n\n> Missing promisor objects do not prevent fsck from passing - this is part\n> of the original design (any packfiles we download from the specifically\n> designated promisor remote are marked as such, and any objects that the\n> objects in the packfile refer to are considered OK to be missing).\n\nIs there ever a risk that objects in the downloaded packfile come\nacross as deltas against other objects that are missing/excluded, or\ndoes the partial clone machinery ensure that doesn't happen?  (Because\nthis was certainly the biggest pain-point with my \"fake cheap clone\"\nhacks.)\n\n> Currently, when a missing object is read, it is first fetched (there are\n> some more details that I can go over if you have any specific\n> questions). What you're suggesting here is to return a fake blob with\n> wrong hash - I haven't looked at all the callers of read-object\n> functions in detail, but I don't think all of them are ready for such a\n> behavioral change.\n\ngit-replace already took care of that for you and provides that\nguarantee, modulo the --no-replace-objects & fsck & prune & fetch &\nwhatnot cases that ignore replace objects as Kaushik mentioned.  I\ntook advantage of this to great effect with my \"fake cheap clone\"\nhacks.  Based in part on your other email where you made a suggestion\nabout promisors, I'm starting to think a pretty good first cut\nsolution might look like the following:\n\n  * user manually adds a bunch of replace refs to map the unwanted big\nblobs to something else (e.g. a README about how the files were\nstripped, or something similar to this)\n  * a partial clone specification that says \"exclude objects that are\nreferenced by replace refs\"\n  * add a fake promisor to the downloaded promisor pack so that if\nanyone runs with --no-replace-objects or similar then they get an\nerror saying the specified objects don't exist and can't be\ndownloaded.\n\nAnyone see any obvious problems with this?\n\n>  Maybe it would be sufficient to just make this work\n> in a more limited scope (e.g. checkout only - and if we need different\n> replacement blobs for different object IDs, maybe we could have\n> something similar to the clean/smudge filters).\n\n>\n> > This could work together with some sort refs/blacklist mechanism to\n> > enable the server to choose which objects the client replaces.\n>\n> In the original email, Kaushik mentioned objects larger than a certain\n> size - we already have support for that (--filter=blob:limit=1000000,\n> for example). Having said that, Git is already able to tolerate any\n> exclusion (of tree or blob) from the server - we already need this in\n> order to support changing of filters, for example.\n"},{"id":"389786","messageId":"20200114215730.154601-1-jonathantanmy@google.com","threadId":"52631","inReplyTo":"CABPp-BF8OHoHo73doekKzf0CmO09_PyAfe4q__DvoftQ+BeY2w@mail.gmail.com","subject":"Re: [RFC] Extending git-replace","fromName":"Jonathan Tan","fromEmail":"jonathantanmy@google.com","sentAt":"2020-01-14T21:57:30Z","receivedAt":"2020-01-14T21:57:36Z","isPatch":false,"sender":{"key":"jonathantanmy@fastmail.com","avatar":null},"body":"> > Missing promisor objects do not prevent fsck from passing - this is part\n> > of the original design (any packfiles we download from the specifically\n> > designated promisor remote are marked as such, and any objects that the\n> > objects in the packfile refer to are considered OK to be missing).\n> \n> Is there ever a risk that objects in the downloaded packfile come\n> across as deltas against other objects that are missing/excluded, or\n> does the partial clone machinery ensure that doesn't happen?  (Because\n> this was certainly the biggest pain-point with my \"fake cheap clone\"\n> hacks.)\n\nThe server may send thin packs during a fetch or clone, but because the\nclient runs index-pack (which calculates the hash of every object\ndownloaded, necessitating having the full object, which in turn triggers\nfetches of any delta bases), this should not happen.\n\nBut if you create the packfile in some other way and then manually set a\nfake promisor remote (as I perhaps too naively suggested) then the\nmechanism will attempt to fetch missing delta bases, which (I think) is\nnot what you want.\n\n> > Currently, when a missing object is read, it is first fetched (there are\n> > some more details that I can go over if you have any specific\n> > questions). What you're suggesting here is to return a fake blob with\n> > wrong hash - I haven't looked at all the callers of read-object\n> > functions in detail, but I don't think all of them are ready for such a\n> > behavioral change.\n> \n> git-replace already took care of that for you and provides that\n> guarantee, modulo the --no-replace-objects & fsck & prune & fetch &\n> whatnot cases that ignore replace objects as Kaushik mentioned.  I\n> took advantage of this to great effect with my \"fake cheap clone\"\n> hacks.  Based in part on your other email where you made a suggestion\n> about promisors, I'm starting to think a pretty good first cut\n> solution might look like the following:\n> \n>   * user manually adds a bunch of replace refs to map the unwanted big\n> blobs to something else (e.g. a README about how the files were\n> stripped, or something similar to this)\n>   * a partial clone specification that says \"exclude objects that are\n> referenced by replace refs\"\n>   * add a fake promisor to the downloaded promisor pack so that if\n> anyone runs with --no-replace-objects or similar then they get an\n> error saying the specified objects don't exist and can't be\n> downloaded.\n> \n> Anyone see any obvious problems with this?\n\nLooking at the list of commands given in the original email (fsck,\nupload-pack, pack/unpack-objects, prune and index-pack), if we use a\nfilter by blob size (instead of the partial clone specification\nsuggested), this would satisfy the purposes of fsck and prune only.\n\nIf we had a partial clone specification that excludes object referenced\nby replace refs, then upload-pack from this partial repository (and\npack-objects) would work too.\n\nBut there might be non-obvious problems that I haven't thought of.\n"},{"id":"389793","messageId":"CABPp-BHBtwxk7B9LLZdCFBTOdTUCnxXyeMXG3XPavp96N+LKHg@mail.gmail.com","threadId":"52631","inReplyTo":"20200114215730.154601-1-jonathantanmy@google.com","subject":"Re: [RFC] Extending git-replace","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2020-01-14T22:46:17Z","receivedAt":"2020-01-14T22:46:32Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Tue, Jan 14, 2020 at 1:57 PM Jonathan Tan <jonathantanmy@google.com> wrote:\n>\n> > > Missing promisor objects do not prevent fsck from passing - this is part\n> > > of the original design (any packfiles we download from the specifically\n> > > designated promisor remote are marked as such, and any objects that the\n> > > objects in the packfile refer to are considered OK to be missing).\n> >\n> > Is there ever a risk that objects in the downloaded packfile come\n> > across as deltas against other objects that are missing/excluded, or\n> > does the partial clone machinery ensure that doesn't happen?  (Because\n> > this was certainly the biggest pain-point with my \"fake cheap clone\"\n> > hacks.)\n>\n> The server may send thin packs during a fetch or clone, but because the\n> client runs index-pack (which calculates the hash of every object\n> downloaded, necessitating having the full object, which in turn triggers\n> fetches of any delta bases), this should not happen.\n\nSo if a user does a partial clone, filtering by blob size >= 1M, and\nif they have several blobs of size just above and just below that\nlimit, then the partial clone will work but probably cause them to\nstill download several blobs above the limit size anyway?  (Which, if\nI'm understanding correctly, happens because the blobs just smaller\nthan 1M likely will delta well against the blobs just larger than 1M.)\n\n> But if you create the packfile in some other way and then manually set a\n> fake promisor remote (as I perhaps too naively suggested) then the\n> mechanism will attempt to fetch missing delta bases, which (I think) is\n> not what you want.\n\nWell, it's not optimal, but we're currently just dying with cryptic\nerrors whenever we have missing delta bases, and this happens whenever\nwe have an accidental fetch of older branches (although this does have\nthe nice side effect of notifying us of stray fetches in our CI\nscripts).  Your promisor suggestion would at least permit gc's &\nprunes if we use it in more places, so should be an improvement.  I\njust wanted to verify whether this problem with delta bases would\nremain.\n\n> > > Currently, when a missing object is read, it is first fetched (there are\n> > > some more details that I can go over if you have any specific\n> > > questions). What you're suggesting here is to return a fake blob with\n> > > wrong hash - I haven't looked at all the callers of read-object\n> > > functions in detail, but I don't think all of them are ready for such a\n> > > behavioral change.\n> >\n> > git-replace already took care of that for you and provides that\n> > guarantee, modulo the --no-replace-objects & fsck & prune & fetch &\n> > whatnot cases that ignore replace objects as Kaushik mentioned.  I\n> > took advantage of this to great effect with my \"fake cheap clone\"\n> > hacks.  Based in part on your other email where you made a suggestion\n> > about promisors, I'm starting to think a pretty good first cut\n> > solution might look like the following:\n> >\n> >   * user manually adds a bunch of replace refs to map the unwanted big\n> > blobs to something else (e.g. a README about how the files were\n> > stripped, or something similar to this)\n> >   * a partial clone specification that says \"exclude objects that are\n> > referenced by replace refs\"\n> >   * add a fake promisor to the downloaded promisor pack so that if\n> > anyone runs with --no-replace-objects or similar then they get an\n> > error saying the specified objects don't exist and can't be\n> > downloaded.\n> >\n> > Anyone see any obvious problems with this?\n>\n> Looking at the list of commands given in the original email (fsck,\n> upload-pack, pack/unpack-objects, prune and index-pack), if we use a\n> filter by blob size (instead of the partial clone specification\n> suggested), this would satisfy the purposes of fsck and prune only.\n>\n> If we had a partial clone specification that excludes object referenced\n> by replace refs, then upload-pack from this partial repository (and\n> pack-objects) would work too.\n>\n> But there might be non-obvious problems that I haven't thought of.\n\nCool, sounds like it's at least worth investigating.  Maybe Kaushik is\ninterested, or maybe I consider throwing it on my backlog and coming\nback to it in a year or two.  :-)\n"},{"id":"389857","messageId":"CAN_uzp+NrFZivfFgAgHf0Phpgax3jLwfdP4fLhTODM4By5OZTQ@mail.gmail.com","threadId":"52631","inReplyTo":"CABPp-BGy3qu_Rd4epore0wLyoh1fg0UH5EAV27shKJ=kLWX4FA@mail.gmail.com","subject":"Re: [RFC] Extending git-replace","fromName":"Kaushik Srenevasan","fromEmail":"kaushik@twitter.com","sentAt":"2020-01-16T03:30:10Z","receivedAt":"2020-01-16T03:30:24Z","isPatch":false,"sender":{"key":"kaushik@twitter.com","avatar":null},"body":"Hi Elijah,\n\nOn Mon, Jan 13, 2020 at 10:55 PM Elijah Newren <newren@gmail.com> wrote:\n> 1) You can rewrite history, and then use replace references to map old\n> commit IDs to new commit IDs.  This allows anyone to continue using\n> old commit IDs (which aren't even part of the new repository anymore)\n> in git commands and git automatically uses and shows the new commit\n> IDs.  No problems with fsck or prune or fetch either.  Creating these\n> replace refs is fairly simple if your repository rewriting program\n> (e.g. git-filter-repo or BFG Repo Cleaner) provides a mapping of old\n> IDs to new IDs, and if you are using git-filter-repo it even creates\n> the replace refs for you.  (The one downside is that you can't use\n> abbreviated refs to refer to replace refs, thus you can't use\n> abbreviated old commit IDs in this scheme.)\n>\n\nThis is the path we're considering taking unless something easier\ncomes out of this (or other) proposal(s). We're working on determining\ncompatibility with tools. Thanks for the pointer to git-filter-repo.\nIt looks great!\n\n> Instead of inventing yet another partial-clone-like thing, it'd be\n> nice if your new mechanism could just be implemented in terms of\n> partial clones, extending them as you need.  I don't like the idea of\n> supporting multiple competing implementations of partial clones\n> withing git.git, but if it's just some extensions of the existing\n> capability then it sounds great.  But you may want to talk with\n> Jonathan Tan if you want to go this route (cc'd), since he's the\n> partial clone expert.\n\nI agree that it isn't worth inventing another partial clone like\nfeature. It sounds however, like something based on partial clone will\nnot solve the problem on the \"server\"? or perhaps I'm missing\nsomething (as I've not had a chance to check out the implementation\nyet). While I'm not at all insisting that `git-blacklist` be the way\nto achieve it, we'd (Twitter) like to be able to permanently get rid\nof the objects in question while retaining the ability to run GC and\nFSCK on all copies of the repository, preferably without having to\nrewrite history.\n\nEven merely making `--no-replace-objects` be FALSE by default for GC\nand FSCK (and printing a warning instead), while retaining existing\nbehavior when it is explicitly requested, would significantly improve\n`git-replace`'s usability (for this purpose). The bits related to ref\ntransfer in my proposal are optional. Git users can either be required\nto explicitly fetch the refs/replacement namespace (as they do today),\nor we could print a message (at the end of clone), letting the user\nknow that there are replacements available on the server. I'd only\nproposed a new command as changing `git-reaplce` thus, would break\nbackward compatibility.\n\n                                -- Kaushik\n"}]}