{"thread":{"id":"48504","subject":"worktrees vs. alternates","startedAt":"2018-05-16T08:14:09Z","lastAt":"2018-05-19T05:46:30Z","messageCount":34,"participants":["Lars Schneider","Ævar Arnfjörð Bjarmason","Robert P. J. Day","Derrick Stolee","Konstantin Ryabitsev","Martin Fick","Jeff King","Stefan Beller","Sitaram Chamarty","Duy Nguyen"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"347780","messageId":"A33442B1-B37D-42E1-9C58-8AB583A43BC9@gmail.com","threadId":"48504","inReplyTo":null,"subject":"worktrees vs. alternates","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2018-05-16T08:13:59Z","receivedAt":"2018-05-16T08:14:09Z","isPatch":false,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"Hi,\n\nI am looking into different options to cache Git repositories on build\nmachines. The two most promising ways seem to be git-worktree [1] and\ngit-alternates [2].\n\nI wonder if you see an advantage of one over the other? \n\nMy impression is that git-worktree supersedes git-alternates. Would\nthat be a fair statement? If yes, would it makes sense to deprecate\nalternates for simplification?\n\nThanks,\nLars\n\n\n[1] https://git-scm.com/docs/git-worktree\n[2] https://git-scm.com/docs/gitrepository-layout#gitrepository-layout-objectsinfoalternates"},{"id":"347781","messageId":"87po1waqyc.fsf@evledraar.gmail.com","threadId":"48504","inReplyTo":"A33442B1-B37D-42E1-9C58-8AB583A43BC9@gmail.com","subject":"Re: worktrees vs. alternates","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-05-16T09:29:47Z","receivedAt":"2018-05-16T09:29:53Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, May 16 2018, Lars Schneider wrote:\n\n> I am looking into different options to cache Git repositories on build\n> machines. The two most promising ways seem to be git-worktree [1] and\n> git-alternates [2].\n>\n> I wonder if you see an advantage of one over the other?\n>\n> My impression is that git-worktree supersedes git-alternates. Would\n> that be a fair statement? If yes, would it makes sense to deprecate\n> alternates for simplification?\n>\n> [1] https://git-scm.com/docs/git-worktree\n> [2] https://git-scm.com/docs/gitrepository-layout#gitrepository-layout-objectsinfoalternates\n\nIt's not correct that worktrees supersede alternates, or the other way\naround, they're orthagonal features.\n\ngit-worktree allows you to create a new working directory connected to\nthe same local object store.\n\nAlternates allow you to declare in any given local object store, that\nyour set of objects isn't complete, and you can find the rest at some\nother location, those object stores may or may not have more than one\nworktree connected to them.\n"},{"id":"347782","messageId":"alpine.LFD.2.21.1805160540100.7243@localhost.localdomain","threadId":"48504","inReplyTo":"87po1waqyc.fsf@evledraar.gmail.com","subject":"Re: worktrees vs. alternates","fromName":"Robert P. J. Day","fromEmail":"rpjday@crashcourse.ca","sentAt":"2018-05-16T09:42:02Z","receivedAt":"2018-05-16T09:43:29Z","isPatch":false,"sender":{"key":"rpjday@crashcourse.ca","avatar":"https://avatars.githubusercontent.com/u/226084077?v=4"},"body":"On Wed, 16 May 2018, Ævar Arnfjörð Bjarmason wrote:\n\n>\n> On Wed, May 16 2018, Lars Schneider wrote:\n>\n> > I am looking into different options to cache Git repositories on build\n> > machines. The two most promising ways seem to be git-worktree [1] and\n> > git-alternates [2].\n> >\n> > I wonder if you see an advantage of one over the other?\n> >\n> > My impression is that git-worktree supersedes git-alternates. Would\n> > that be a fair statement? If yes, would it makes sense to deprecate\n> > alternates for simplification?\n> >\n> > [1] https://git-scm.com/docs/git-worktree\n> > [2] https://git-scm.com/docs/gitrepository-layout#gitrepository-layout-objectsinfoalternates\n>\n> It's not correct that worktrees supersede alternates, or the other\n> way around, they're orthagonal features.\n>\n> git-worktree allows you to create a new working directory connected\n> to the same local object store.\n>\n> Alternates allow you to declare in any given local object store,\n> that your set of objects isn't complete, and you can find the rest\n> at some other location, those object stores may or may not have more\n> than one worktree connected to them.\n\n  just to be clear here, there should be nothing about how alternates\nare set up for a repository that should affect the normal behaviour of\nworking trees for that repository, correct? i never thought there was,\ni just thought i'd make absolutely sure.\n\nrday"},{"id":"347783","messageId":"81B00B00-00F4-487A-9D3E-6B7514098B29@gmail.com","threadId":"48504","inReplyTo":"87po1waqyc.fsf@evledraar.gmail.com","subject":"Re: worktrees vs. alternates","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2018-05-16T09:51:26Z","receivedAt":"2018-05-16T09:51:42Z","isPatch":false,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 16 May 2018, at 11:29, Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n> \n> \n> On Wed, May 16 2018, Lars Schneider wrote:\n> \n>> I am looking into different options to cache Git repositories on build\n>> machines. The two most promising ways seem to be git-worktree [1] and\n>> git-alternates [2].\n>> \n>> I wonder if you see an advantage of one over the other?\n>> \n>> My impression is that git-worktree supersedes git-alternates. Would\n>> that be a fair statement? If yes, would it makes sense to deprecate\n>> alternates for simplification?\n>> \n>> [1] https://git-scm.com/docs/git-worktree\n>> [2] https://git-scm.com/docs/gitrepository-layout#gitrepository-layout-objectsinfoalternates\n> \n> It's not correct that worktrees supersede alternates, or the other way\n> around, they're orthagonal features.\n> \n> git-worktree allows you to create a new working directory connected to\n> the same local object store.\n> \n> Alternates allow you to declare in any given local object store, that\n> your set of objects isn't complete, and you can find the rest at some\n> other location, those object stores may or may not have more than one\n> worktree connected to them.\n\nOK. I just wonder in what situation I would work with an incomplete\nobject store. The only use case I could imagine is that two repos share\na common set of objects (most likely blobs). However, in that situation\nI would keep the two independent lines of development in a single repo\nwith two root commits.\n\nWould it be fair to say that \"git alternates\" are a good mechanism to \ncache objects across different repos? However, I would consider a cache \nhit  between different repos unlikely. In that line of thinking\n\"git worktree\" would be a good (maybe better?) mechanism to cache objects\nfor a single repo?\n\nThanks,\nLars"},{"id":"347784","messageId":"87muwzc2kv.fsf@evledraar.gmail.com","threadId":"48504","inReplyTo":"81B00B00-00F4-487A-9D3E-6B7514098B29@gmail.com","subject":"Re: worktrees vs. alternates","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-05-16T10:33:20Z","receivedAt":"2018-05-16T10:33:26Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, May 16 2018, Lars Schneider wrote:\n\n>> On 16 May 2018, at 11:29, Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>>\n>>\n>> On Wed, May 16 2018, Lars Schneider wrote:\n>>\n>>> I am looking into different options to cache Git repositories on build\n>>> machines. The two most promising ways seem to be git-worktree [1] and\n>>> git-alternates [2].\n>>>\n>>> I wonder if you see an advantage of one over the other?\n>>>\n>>> My impression is that git-worktree supersedes git-alternates. Would\n>>> that be a fair statement? If yes, would it makes sense to deprecate\n>>> alternates for simplification?\n>>>\n>>> [1] https://git-scm.com/docs/git-worktree\n>>> [2] https://git-scm.com/docs/gitrepository-layout#gitrepository-layout-objectsinfoalternates\n>>\n>> It's not correct that worktrees supersede alternates, or the other way\n>> around, they're orthagonal features.\n>>\n>> git-worktree allows you to create a new working directory connected to\n>> the same local object store.\n>>\n>> Alternates allow you to declare in any given local object store, that\n>> your set of objects isn't complete, and you can find the rest at some\n>> other location, those object stores may or may not have more than one\n>> worktree connected to them.\n>\n> OK. I just wonder in what situation I would work with an incomplete\n> object store. The only use case I could imagine is that two repos share\n> a common set of objects (most likely blobs). However, in that situation\n> I would keep the two independent lines of development in a single repo\n> with two root commits.\n>\n> Would it be fair to say that \"git alternates\" are a good mechanism to\n> cache objects across different repos? However, I would consider a cache\n> hit  between different repos unlikely. In that line of thinking\n> \"git worktree\" would be a good (maybe better?) mechanism to cache objects\n> for a single repo?\n\nThe use case is cloning with e.g. --shared or --reference.\n\nConsider the following scenario:\n\n * You have 100 developers with *nix accounts on a single machine.\n\n * These 100 all need access to the same repo, but .git/objects is 1G\n\n * This would then naïvely require 100G of space + working tree. If the\n   machine has 92G of RAM you'll be swapping the fscache in & out and\n   performance will be horrible.\n\nInstead, you have a single repository maintained on the system designed\nto have all the alternates point to it, cloned as:\n\n    git clone --reference /usr/share/git_tree/bigrepo ssh://....bigrepo.git ~/bigrepo\n\nNow you're using just a bit over 1GB of space in total, but any new\nobjects the devs create will be written to their local .git dir, since\nyou're spending 1GB for those 100 repos instead of 100GB the data is\nalways in the FS cache.\n\nAnd here's where this isn't at all like \"worktree\", each of those 100\nwill have their own \"master\" branch, and they can all create 100\ndifferent branches called \"topic\" that can be different.\n\nWith worktree the references are all shared across the same worktrees,\nso it's designed for one dev working on different topic branches in\ndifferent checkouts.\n\nThe --reference feature is also commonly used in CI-like\nenvironments. Imagine the above example, but except with 100 devs you\nhave CI jobs on the same machine being spun up all the time, although\nhere you get some overlap, if you're OK with the main branch name being\ndifferent you can also do this with worktrees instead of alternates.\n"},{"id":"347785","messageId":"87lgcjc0zq.fsf@evledraar.gmail.com","threadId":"48504","inReplyTo":"alpine.LFD.2.21.1805160540100.7243@localhost.localdomain","subject":"Re: worktrees vs. alternates","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-05-16T11:07:37Z","receivedAt":"2018-05-16T11:07:44Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, May 16 2018, Robert P. J. Day wrote:\n\n> On Wed, 16 May 2018, Ævar Arnfjörð Bjarmason wrote:\n>\n>>\n>> On Wed, May 16 2018, Lars Schneider wrote:\n>>\n>> > I am looking into different options to cache Git repositories on build\n>> > machines. The two most promising ways seem to be git-worktree [1] and\n>> > git-alternates [2].\n>> >\n>> > I wonder if you see an advantage of one over the other?\n>> >\n>> > My impression is that git-worktree supersedes git-alternates. Would\n>> > that be a fair statement? If yes, would it makes sense to deprecate\n>> > alternates for simplification?\n>> >\n>> > [1] https://git-scm.com/docs/git-worktree\n>> > [2] https://git-scm.com/docs/gitrepository-layout#gitrepository-layout-objectsinfoalternates\n>>\n>> It's not correct that worktrees supersede alternates, or the other\n>> way around, they're orthagonal features.\n>>\n>> git-worktree allows you to create a new working directory connected\n>> to the same local object store.\n>>\n>> Alternates allow you to declare in any given local object store,\n>> that your set of objects isn't complete, and you can find the rest\n>> at some other location, those object stores may or may not have more\n>> than one worktree connected to them.\n>\n>   just to be clear here, there should be nothing about how alternates\n> are set up for a repository that should affect the normal behaviour of\n> working trees for that repository, correct? i never thought there was,\n> i just thought i'd make absolutely sure.\n\nThat's correct. The worktree(s) are logically composed of the\nindex/cache, checked-out files, and the local reference store (and some\nauxiliary things, like per-worktree refs like HEAD, and config...).\n\nWhether you have one worktree or many, eventually git needs to look up\nobjects somewhere. The alternates mechanism is just one more way to\nspecify where to look, along with some special logic in pack-objects and\nthe like where we need to be aware of them for the purposes of\nmaintaining objects in the repository.\n"},{"id":"347787","messageId":"fc2f1fdf-222f-aaee-9d58-aae8692920f5@gmail.com","threadId":"48504","inReplyTo":"87muwzc2kv.fsf@evledraar.gmail.com","subject":"Re: worktrees vs. alternates","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2018-05-16T13:02:13Z","receivedAt":"2018-05-16T13:02:19Z","isPatch":false,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 5/16/2018 6:33 AM, Ævar Arnfjörð Bjarmason wrote:\n[big snip]\n>\n> And here's where this isn't at all like \"worktree\", each of those 100\n> will have their own \"master\" branch, and they can all create 100\n> different branches called \"topic\" that can be different.\n\nThis is the biggest difference. You cannot have the same ref checked out \nin multiple worktrees, as they both may edit that ref. The alternates \nallow you to share data in a \"read only\" fashion. If you have one repo \nthat is the \"base\" repo that manages that objects dir, then that is \nprobably a good way to reduce the duplication. I'm not familiar with \nwhat happens when a \"child\" repo does 'git gc' or 'git repack', will it \ndelete the local objects that is sees exist in the alternate?\n\nGVFS uses alternates in this same way: we create a drive-wide \"shared \nobject cache\" that GVFS manages. We put our prefetch packs filled with \ncommits and trees in there, and any loose objects that are downloaded \nvia the object virtualization are placed as loose objects in the \nalternate. We also store the multi-pack-index and commit-graph in that \nalternate. This means that the only objects in each src dir are those \ncreated by the developer doing their normal work.\n\nThanks,\n-Stolee\n\n"},{"id":"347790","messageId":"0f19f9f8-d215-622e-5090-1341c013babc@linuxfoundation.org","threadId":"48504","inReplyTo":"fc2f1fdf-222f-aaee-9d58-aae8692920f5@gmail.com","subject":"Re: worktrees vs. alternates","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2018-05-16T14:58:19Z","receivedAt":"2018-05-16T14:58:26Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On 05/16/18 09:02, Derrick Stolee wrote:\n> This is the biggest difference. You cannot have the same ref checked out\n> in multiple worktrees, as they both may edit that ref. The alternates\n> allow you to share data in a \"read only\" fashion. If you have one repo\n> that is the \"base\" repo that manages that objects dir, then that is\n> probably a good way to reduce the duplication. I'm not familiar with\n> what happens when a \"child\" repo does 'git gc' or 'git repack', will it\n> delete the local objects that is sees exist in the alternate?\n\nThe parent repo is not keeping track of any other repositories that may\nbe using it for alternates, which is why you basically:\n\n1. never run auto-gc in the parent repo\n2. repack it manually using -Ad to keep loose objects that other repos\nmay be borrowing (but we don't know if they are)\n3. never prune the parent repo, because this may delete objects other\nrepos are borrowing\n\nVery infrequently you may consider this extra set of maintenance steps:\n\n1. Find every repo mentioning the parent repository in their alternates\n2. Repack them without the -l switch (which copies all the borrowed\nobjects into those repos)\n3. Once all child repos have been repacked this way, prune the parent\nrepo (it's safe now)\n4. Repack child repos again, this time with the -l flag, to get your\nsavings back.\n\nI would heartily love a way to teach git-repack to recognize when an\nobject it's borrowing from the parent repo is in danger of being pruned.\nThe cheapest way of doing this would probably be to hardlink loose\nobjects into its own objects directory and only consider \"safe\" objects\nthose that are part of the parent repository's pack. This should make\nalternates a lot safer, just in case git-prune happens to run by accident.\n\n> GVFS uses alternates in this same way: we create a drive-wide \"shared\n> object cache\" that GVFS manages. We put our prefetch packs filled with\n> commits and trees in there, and any loose objects that are downloaded\n> via the object virtualization are placed as loose objects in the\n> alternate. We also store the multi-pack-index and commit-graph in that\n> alternate. This means that the only objects in each src dir are those\n> created by the developer doing their normal work.\n\nI'm very interested in GVFS, because it would certainly make my life\neasier maintaining source.codeaurora.org, which is many thousands of\nrepos that are mostly forks of the same stuff. However, GVFS appears to\nonly exist for Windows (hint-hint, nudge-nudge). :)\n\nBest,\n-- \nKonstantin Ryabitsev\nDirector, IT Infrastructure Security\nThe Linux Foundation\n\n"},{"id":"347792","messageId":"87k1s3bomt.fsf@evledraar.gmail.com","threadId":"48504","inReplyTo":"0f19f9f8-d215-622e-5090-1341c013babc@linuxfoundation.org","subject":"Re: worktrees vs. alternates","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-05-16T15:34:34Z","receivedAt":"2018-05-16T15:34:41Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, May 16 2018, Konstantin Ryabitsev wrote:\n\n> On 05/16/18 09:02, Derrick Stolee wrote:\n>> This is the biggest difference. You cannot have the same ref checked out\n>> in multiple worktrees, as they both may edit that ref. The alternates\n>> allow you to share data in a \"read only\" fashion. If you have one repo\n>> that is the \"base\" repo that manages that objects dir, then that is\n>> probably a good way to reduce the duplication. I'm not familiar with\n>> what happens when a \"child\" repo does 'git gc' or 'git repack', will it\n>> delete the local objects that is sees exist in the alternate?\n>\n> The parent repo is not keeping track of any other repositories that may\n> be using it for alternates, which is why you basically:\n>\n> 1. never run auto-gc in the parent repo\n> 2. repack it manually using -Ad to keep loose objects that other repos\n> may be borrowing (but we don't know if they are)\n> 3. never prune the parent repo, because this may delete objects other\n> repos are borrowing\n>\n> Very infrequently you may consider this extra set of maintenance steps:\n>\n> 1. Find every repo mentioning the parent repository in their alternates\n> 2. Repack them without the -l switch (which copies all the borrowed\n> objects into those repos)\n> 3. Once all child repos have been repacked this way, prune the parent\n> repo (it's safe now)\n> 4. Repack child repos again, this time with the -l flag, to get your\n> savings back.\n>\n> I would heartily love a way to teach git-repack to recognize when an\n> object it's borrowing from the parent repo is in danger of being pruned.\n> The cheapest way of doing this would probably be to hardlink loose\n> objects into its own objects directory and only consider \"safe\" objects\n> those that are part of the parent repository's pack. This should make\n> alternates a lot safer, just in case git-prune happens to run by accident.\n\nI may have missed some edge case, but I believe this entire workaround\nisn't needed if you guarantee that the parent repo doesn't contain any\nobjects that will get un-referenced.\n\nYou'd do that in the common case by cloning with --single-branch, and\ndepending on your setup --no-tags (if you delete tags). This is assuming\nthat your HEAD branch points to something like a \"master\" that doesn't\nget rewound.\n\nThe problem you're describing happens if say you clone git.git and have\nthe \"pu\" branch in there in the parent, and as a result you get child\nrepos referencing those objects, but when the parent GCs after \"pu\" is\nrewound the child repos break. Thus your elaborate work-around.\n\nBut that situation isn't possible in the first place if you only ever\nimport the \"master\" branch, or other references guaranteed not to\nchange.\n\nOf course that has the trade-off that every child repo needs to get its\nown objects for the \"next\" branch, \"pu\", etc. But those are\ncomparatively tiny.\n\nI wasn't aware of -l (--local), or had forgotten about it. I thought\nthat we didn't have that and the \"child\" repos would just keep growing\nover time, i.e. not get rid of the objects we're fetching into the\nparent (which the parent might get later due to the child, say if it's\nfetched in a daily cronjob). Good to know that's not the case.\n\nWith that --local flag the trade-off of not fetching \"next\" and \"pu\"\netc. should become irrelevant over time, as they migrate to \"master\"\nthey'll get de-duplicated, or alternatively GC'd by the child repos if\nthey don't make it.\n\n>> GVFS uses alternates in this same way: we create a drive-wide \"shared\n>> object cache\" that GVFS manages. We put our prefetch packs filled with\n>> commits and trees in there, and any loose objects that are downloaded\n>> via the object virtualization are placed as loose objects in the\n>> alternate. We also store the multi-pack-index and commit-graph in that\n>> alternate. This means that the only objects in each src dir are those\n>> created by the developer doing their normal work.\n>\n> I'm very interested in GVFS, because it would certainly make my life\n> easier maintaining source.codeaurora.org, which is many thousands of\n> repos that are mostly forks of the same stuff. However, GVFS appears to\n> only exist for Windows (hint-hint, nudge-nudge). :)\n\nThis should make you happy:\n\nhttps://arstechnica.com/gadgets/2017/11/microsoft-and-github-team-up-to-take-git-virtual-file-system-to-macos-linux/\n\nBut I don't know what the current status is or where it can be followed.\n"},{"id":"347793","messageId":"20180516154935.GA9712@chatter","threadId":"48504","inReplyTo":"87k1s3bomt.fsf@evledraar.gmail.com","subject":"Re: worktrees vs. alternates","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2018-05-16T15:49:35Z","receivedAt":"2018-05-16T15:49:42Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On Wed, May 16, 2018 at 05:34:34PM +0200, Ævar Arnfjörð Bjarmason wrote:\n>I may have missed some edge case, but I believe this entire workaround\n>isn't needed if you guarantee that the parent repo doesn't contain any\n>objects that will get un-referenced.\n\nYou can't guarantee that, because the parent repo can have its history\nrewritten either via a forced push, or via a rebase. Obviously, this\nwon't happen in something like torvalds/linux.git, which is why it's\npretty safe to alternate off of that repo for us, but codeaurora.org\nrepos aren't always strictly-ff (e.g. because they may rebase themselves\nbased on what is in upstream AOSP repos) -- so objects in them may\nbecome unreferenced and pruned away, corrupting any repos using them for\nalternates.\n\n>> I'm very interested in GVFS, because it would certainly make my life\n>> easier maintaining source.codeaurora.org, which is many thousands of\n>> repos that are mostly forks of the same stuff. However, GVFS appears to\n>> only exist for Windows (hint-hint, nudge-nudge). :)\n>\n>This should make you happy:\n>\n>https://arstechnica.com/gadgets/2017/11/microsoft-and-github-team-up-to-take-git-virtual-file-system-to-macos-linux/\n>\n>But I don't know what the current status is or where it can be followed.\n\nVery good to know, thanks!\n\n-K\n"},{"id":"347800","messageId":"5972145.OdP4kjFpBj@mfick-lnx","threadId":"48504","inReplyTo":"0f19f9f8-d215-622e-5090-1341c013babc@linuxfoundation.org","subject":"Re: worktrees vs. alternates","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2018-05-16T17:14:44Z","receivedAt":"2018-05-16T17:14:51Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Wednesday, May 16, 2018 10:58:19 AM Konstantin Ryabitsev \nwrote:\n> \n> 1. Find every repo mentioning the parent repository in\n> their alternates 2. Repack them without the -l switch\n> (which copies all the borrowed objects into those repos)\n> 3. Once all child repos have been repacked this way, prune\n> the parent repo (it's safe now)\n\nThis is probably only true if the repos are in read-only \nmode?  I suspect this is still racy on a busy server with no \ndowntime.\n\n> 4. Repack child repos again, this time with the -l flag,\n> to get your savings back.\n \n> I would heartily love a way to teach git-repack to\n> recognize when an object it's borrowing from the parent\n> repo is in danger of being pruned. The cheapest way of\n> doing this would probably be to hardlink loose objects\n> into its own objects directory and only consider \"safe\"\n> objects those that are part of the parent repository's\n> pack. This should make alternates a lot safer, just in\n> case git-prune happens to run by accident.\n\nI think that hard linking is generally a good approach to \nsolving many of the \"pruning\" races left in git.\n\nI have uploaded a \"hard linking\" proposal to jgit that could \npotentially solve a similar situation that is not alternate \nspecific, and only for packfiles, with the intent of \neventually also doing something similar for loose \nobjects.  You can see this here: \n\nhttps://git.eclipse.org/r/c/122288/2\n\nI think it would be good to fill in more of these pruning \ngaps!\n\n-Martin\n\n-- \nThe Qualcomm Innovation Center, Inc. is a member of Code \nAurora Forum, hosted by The Linux Foundation\n"},{"id":"347802","messageId":"099ff2bf-c0f8-60fc-7833-9b129dd4dffe@linuxfoundation.org","threadId":"48504","inReplyTo":"5972145.OdP4kjFpBj@mfick-lnx","subject":"Re: worktrees vs. alternates","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2018-05-16T17:41:55Z","receivedAt":"2018-05-16T17:42:03Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On 05/16/18 13:14, Martin Fick wrote:\n> On Wednesday, May 16, 2018 10:58:19 AM Konstantin Ryabitsev \n> wrote:\n>>\n>> 1. Find every repo mentioning the parent repository in\n>> their alternates 2. Repack them without the -l switch\n>> (which copies all the borrowed objects into those repos)\n>> 3. Once all child repos have been repacked this way, prune\n>> the parent repo (it's safe now)\n> \n> This is probably only true if the repos are in read-only \n> mode?  I suspect this is still racy on a busy server with no \n> downtime.\n\nWe don't actually do this anywhere. :) It's a feature I keep hoping to\nadd one day to grokmirror, but keep putting off because of various\nconsiderations. As you can imagine, if we have 300 forks of linux.git\nall using torvalds/linux.git as their alternates, then repacking them\nall without -l would balloon our disk usage 300-fold. At this time it's\njust cheaper to keep a bunch of loose objects around forever at the cost\nof decreased performance.\n\nMaybe git-repack can be told to only borrow parent objects if they are\nin packs. Anything not in packs should be hardlinked into the child\nrepo. That's my wishful think for the day. :)\n\nBest,\n-- \nKonstantin Ryabitsev\nDirector, IT Infrastructure Security\nThe Linux Foundation\n\n"},{"id":"347803","messageId":"87in7nbi5b.fsf@evledraar.gmail.com","threadId":"48504","inReplyTo":"20180516154935.GA9712@chatter","subject":"Re: worktrees vs. alternates","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-05-16T17:54:40Z","receivedAt":"2018-05-16T17:54:46Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, May 16 2018, Konstantin Ryabitsev wrote:\n\n> On Wed, May 16, 2018 at 05:34:34PM +0200, Ævar Arnfjörð Bjarmason wrote:\n>>I may have missed some edge case, but I believe this entire workaround\n>>isn't needed if you guarantee that the parent repo doesn't contain any\n>>objects that will get un-referenced.\n>\n> You can't guarantee that, because the parent repo can have its history\n> rewritten either via a forced push, or via a rebase. Obviously, this\n> won't happen in something like torvalds/linux.git, which is why it's\n> pretty safe to alternate off of that repo for us, but codeaurora.org\n> repos aren't always strictly-ff (e.g. because they may rebase themselves\n> based on what is in upstream AOSP repos) -- so objects in them may\n> become unreferenced and pruned away, corrupting any repos using them for\n> alternates.\n\nRight, it wouldn't work in the general case. I was thinking of the\nuse-case for doing this (say with known big monorepos) where you know a\ngiven branch won't be unwound.\n\nStill, there's a tiny variation on this that should work with arbitrary\nrepos whose master may be rewound, you just setup a refspec to fetch\ntheir upstream HEAD into master-1 without having \"+\" in the\nrefspec. Then if they never rewind you keep fetching to master-1\nforever.\n\nIf they do rewind you fetch that to master-2 and so forth, so you can\nfollow an upstream rewinding branch while still guaranteeing that no\nobjects ever disappear from your parent repo. This is still a lot\nsimpler than the juggling approach you noted, since it's just a tiny\nshellscript around the \"fetch\".\n\nThis assumes that:\n\n  1. Whenever this happens the history is still similar enough that the\n     parent won't balloon in size like this, or at least it won't be\n     worse than not using alternates at all.\n\n 2. You're getting most of the gains of the object sharing by just\n    grabbing the upstream HEAD branch, i.e. you don't have some repo\n    with huge and N unrelated histories.\n"},{"id":"347804","messageId":"87h8n7bhro.fsf@evledraar.gmail.com","threadId":"48504","inReplyTo":"099ff2bf-c0f8-60fc-7833-9b129dd4dffe@linuxfoundation.org","subject":"Re: worktrees vs. alternates","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-05-16T18:02:51Z","receivedAt":"2018-05-16T18:02:58Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, May 16 2018, Konstantin Ryabitsev wrote:\n\n> Maybe git-repack can be told to only borrow parent objects if they are\n> in packs. Anything not in packs should be hardlinked into the child\n> repo. That's my wishful think for the day. :)\n\nCan you elaborate on how this would help?\n\nWe're just going to create loose objects on interactive \"git commit\",\npresumably you're not adding someone's working copy as the alternate.\n\nOtherwise if it's just being pushed to all those pushes are going to be\nin packs, and the packs may contain e.g. pushes for the \"pu\" branch or\nwhatever, which are objects that'll go away.\n"},{"id":"347805","messageId":"a933cb3a-6c04-d963-aeda-b5850ca8994c@linuxfoundation.org","threadId":"48504","inReplyTo":"87h8n7bhro.fsf@evledraar.gmail.com","subject":"Re: worktrees vs. alternates","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2018-05-16T18:12:24Z","receivedAt":"2018-05-16T18:12:32Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On 05/16/18 14:02, Ævar Arnfjörð Bjarmason wrote:\n> \n> On Wed, May 16 2018, Konstantin Ryabitsev wrote:\n> \n>> Maybe git-repack can be told to only borrow parent objects if they are\n>> in packs. Anything not in packs should be hardlinked into the child\n>> repo. That's my wishful think for the day. :)\n> \n> Can you elaborate on how this would help?\n> \n> We're just going to create loose objects on interactive \"git commit\",\n> presumably you're not adding someone's working copy as the alternate.\n\nThe loose objects I'm thinking of are those that are generated when we\ndo \"git repack -Ad\" -- this takes all unreachable objects and loosens\nthem (see man git-repack for more info). Normally, these would be pruned\nafter a certain period, but we're deliberately keeping them around\nforever just in case another repo relies on them via alternates. I want\nthose repos to \"claim\" these loose objects via hardlinks, such that we\ncan run git-prune on the mother repo instead of dragging all the\nunreachable objects on forever just in case.\n\n> Otherwise if it's just being pushed to all those pushes are going to be\n> in packs, and the packs may contain e.g. pushes for the \"pu\" branch or\n> whatever, which are objects that'll go away.\n\nThere are lots of cases where unreachable objects in one repo would\nnever become unreachable in another -- for example, if the author had\nstopped updating it.\n\nHope this helps.\n\nBest,\n-- \nKonstantin Ryabitsev\nDirector, IT Infrastructure Security\nThe Linux Foundation\n\n"},{"id":"347808","messageId":"1950199.Z2x8tXoTfI@mfick-lnx","threadId":"48504","inReplyTo":"a933cb3a-6c04-d963-aeda-b5850ca8994c@linuxfoundation.org","subject":"Re: worktrees vs. alternates","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2018-05-16T18:26:25Z","receivedAt":"2018-05-16T18:26:30Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Wednesday, May 16, 2018 02:12:24 PM Konstantin Ryabitsev \nwrote:\n> The loose objects I'm thinking of are those that are\n> generated when we do \"git repack -Ad\" -- this takes all\n> unreachable objects and loosens them (see man git-repack\n> for more info). Normally, these would be pruned after a\n> certain period, but we're deliberately keeping them\n> around forever just in case another repo relies on them\n> via alternates. I want those repos to \"claim\" these loose\n> objects via hardlinks, such that we can run git-prune on\n> the mother repo instead of dragging all the unreachable\n> objects on forever just in case.\n\nIf you are going to keep the unreferenced objects around \nforever, it might be better to keep them around in packed \nform?  We currently do that because we don't think there is \na safe way to prune objects yet on a running server (which \nis why I am teaching jgit to be able to recover from a racy \npruning error),\n\n-Martin\n\n-- \nThe Qualcomm Innovation Center, Inc. is a member of Code \nAurora Forum, hosted by The Linux Foundation\n\n"},{"id":"347809","messageId":"e8776c83-ea57-456d-5bc8-ca2fc990bed0@linuxfoundation.org","threadId":"48504","inReplyTo":"1950199.Z2x8tXoTfI@mfick-lnx","subject":"Re: worktrees vs. alternates","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2018-05-16T19:01:13Z","receivedAt":"2018-05-16T19:01:20Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On 05/16/18 14:26, Martin Fick wrote:\n> If you are going to keep the unreferenced objects around \n> forever, it might be better to keep them around in packed \n> form?\n\nI'm undecided about that. On the one hand this does create lots of small\nfiles and inevitably causes (some) performance degradation. On the other\nhand, I don't want to keep useless objects in the pack, because that\nwould also cause performance degradation for people cloning the \"mother\nrepo.\" If my assumptions on any of that are incorrect, I'm happy to\nlearn more.\n\nBest,\n-- \nKonstantin Ryabitsev\nDirector, IT Infrastructure Security\nThe Linux Foundation\n\n"},{"id":"347810","messageId":"2828274.Q9q2dc6g5t@mfick-lnx","threadId":"48504","inReplyTo":"e8776c83-ea57-456d-5bc8-ca2fc990bed0@linuxfoundation.org","subject":"Re: worktrees vs. alternates","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2018-05-16T19:03:42Z","receivedAt":"2018-05-16T19:03:47Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Wednesday, May 16, 2018 03:01:13 PM Konstantin Ryabitsev \nwrote:\n> On 05/16/18 14:26, Martin Fick wrote:\n> > If you are going to keep the unreferenced objects around\n> > forever, it might be better to keep them around in\n> > packed\n> > form?\n> \n> I'm undecided about that. On the one hand this does create\n> lots of small files and inevitably causes (some)\n> performance degradation. On the other hand, I don't want\n> to keep useless objects in the pack, because that would\n> also cause performance degradation for people cloning the\n> \"mother repo.\" If my assumptions on any of that are\n> incorrect, I'm happy to learn more.\n\nMy suggestion is to use science, not logic or hearsay. :) \ni.e. test it!\n\n-Martin\n\n-- \nThe Qualcomm Innovation Center, Inc. is a member of Code \nAurora Forum, hosted by The Linux Foundation\n\n"},{"id":"347811","messageId":"4a18e167-8cec-0141-fe2c-4e0a35f16daf@linuxfoundation.org","threadId":"48504","inReplyTo":"2828274.Q9q2dc6g5t@mfick-lnx","subject":"Re: worktrees vs. alternates","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2018-05-16T19:11:47Z","receivedAt":"2018-05-16T19:11:54Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On 05/16/18 15:03, Martin Fick wrote:\n>> I'm undecided about that. On the one hand this does create\n>> lots of small files and inevitably causes (some)\n>> performance degradation. On the other hand, I don't want\n>> to keep useless objects in the pack, because that would\n>> also cause performance degradation for people cloning the\n>> \"mother repo.\" If my assumptions on any of that are\n>> incorrect, I'm happy to learn more.\n> My suggestion is to use science, not logic or hearsay. :) \n> i.e. test it!\n\nI think the answer will be \"it depends.\" In many of our cases the repos\nthat need those loose objects are rarely accessed -- usually because\nthey are forks with older data (hence why they need objects that are no\nlonger used by the mother repo). Therefore, performance impacts of\noccasionally touching a handful of loose objects will be fairly\nnegligible. This is especially true on non-spinning media where seek\ntimes are low anyway. Having slimmer packs for the mother repo would be\nmore beneficial in this case.\n\nOn the other hand, if the \"child repo\" is frequently used, then the\nimpact of needing a bunch of loose objects would be greater. For the\nsake of simplicity, I think I'll leave things as they are -- it's\ncheaper to fix this via reducing seek times than by applying complicated\nlogic trying to optimize on a per-repo basis.\n\nBest,\n-- \nKonstantin Ryabitsev\nDirector, IT Infrastructure Security\nThe Linux Foundation\n\n"},{"id":"347812","messageId":"20180516191410.GA3417@sigill.intra.peff.net","threadId":"48504","inReplyTo":"0f19f9f8-d215-622e-5090-1341c013babc@linuxfoundation.org","subject":"Re: worktrees vs. alternates","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-05-16T19:14:11Z","receivedAt":"2018-05-16T19:14:17Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, May 16, 2018 at 10:58:19AM -0400, Konstantin Ryabitsev wrote:\n\n> The parent repo is not keeping track of any other repositories that may\n> be using it for alternates, which is why you basically:\n> \n> 1. never run auto-gc in the parent repo\n> 2. repack it manually using -Ad to keep loose objects that other repos\n> may be borrowing (but we don't know if they are)\n> 3. never prune the parent repo, because this may delete objects other\n> repos are borrowing\n> \n> Very infrequently you may consider this extra set of maintenance steps:\n> \n> 1. Find every repo mentioning the parent repository in their alternates\n> 2. Repack them without the -l switch (which copies all the borrowed\n> objects into those repos)\n> 3. Once all child repos have been repacked this way, prune the parent\n> repo (it's safe now)\n> 4. Repack child repos again, this time with the -l flag, to get your\n> savings back.\n\nYou can also do periodic maintenance like:\n\n  1. Copy each ref in the forked repositories into the parent repository\n     (e.g., giving each child that borrows from the parent its own\n     hierarchy in refs/remotes/<child>/*).\n\n  2. Repack the parent as normal. It will retain any objects referenced\n     by the children (because they are now referenced by it).\n\nBut note that:\n\n  1. It's not atomic with respect to updates in the child repos (but\n     then, neither is the single-repo case!).\n\n  2. It doesn't know about reflogs or the index in the child\n     repositories.\n\nThis is more or less how we use alternates at GitHub.\n\n> I would heartily love a way to teach git-repack to recognize when an\n> object it's borrowing from the parent repo is in danger of being pruned.\n> The cheapest way of doing this would probably be to hardlink loose\n> objects into its own objects directory and only consider \"safe\" objects\n> those that are part of the parent repository's pack. This should make\n> alternates a lot safer, just in case git-prune happens to run by accident.\n\nIf you set:\n\n  git config core.repositoryformatversion 1\n  git config extensions.preciousObjects true\n\nin the parent, git-prune (repack -d) will refuse to run. That doesn't\nsolve the problem of how to repack, but it can help prevent accidental\nmisuse.\n\n-Peff\n"},{"id":"347813","messageId":"5484271.13heo1O2yY@mfick-lnx","threadId":"48504","inReplyTo":"4a18e167-8cec-0141-fe2c-4e0a35f16daf@linuxfoundation.org","subject":"Re: worktrees vs. alternates","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2018-05-16T19:18:20Z","receivedAt":"2018-05-16T19:18:27Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Wednesday, May 16, 2018 03:11:47 PM Konstantin Ryabitsev \nwrote:\n> On 05/16/18 15:03, Martin Fick wrote:\n> >> I'm undecided about that. On the one hand this does\n> >> create lots of small files and inevitably causes\n> >> (some) performance degradation. On the other hand, I\n> >> don't want to keep useless objects in the pack,\n> >> because that would also cause performance degradation\n> >> for people cloning the \"mother repo.\" If my\n> >> assumptions on any of that are incorrect, I'm happy to\n> >> learn more.\n> > \n> > My suggestion is to use science, not logic or hearsay.\n> > :)\n> > i.e. test it!\n> \n> I think the answer will be \"it depends.\" In many of our\n> cases the repos that need those loose objects are rarely\n> accessed -- usually because they are forks with older\n> data (hence why they need objects that are no longer used\n> by the mother repo). Therefore, performance impacts of\n> occasionally touching a handful of loose objects will be\n> fairly negligible. This is especially true on\n> non-spinning media where seek times are low anyway.\n> Having slimmer packs for the mother repo would be more\n> beneficial in this case.\n> \n> On the other hand, if the \"child repo\" is frequently used,\n> then the impact of needing a bunch of loose objects would\n> be greater. For the sake of simplicity, I think I'll\n> leave things as they are -- it's cheaper to fix this via\n> reducing seek times than by applying complicated logic\n> trying to optimize on a per-repo basis.\n\nI think a major performance issue with loose objects is not \njust the seek time, but also the fact that they are not \ndelta compressed.  This means that sending them over the \nwire will likely have a significant cost before sending it. \nUnlike the seek time, this cost is not mitigated across \nconcurrent fetches by the FS (or jgit if you were to use it) \ncaching,\n\n-Martin\n\n-- \nThe Qualcomm Innovation Center, Inc. is a member of Code \nAurora Forum, hosted by The Linux Foundation\n\n"},{"id":"347814","messageId":"20180516192343.GB3417@sigill.intra.peff.net","threadId":"48504","inReplyTo":"e8776c83-ea57-456d-5bc8-ca2fc990bed0@linuxfoundation.org","subject":"Re: worktrees vs. alternates","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-05-16T19:23:43Z","receivedAt":"2018-05-16T19:23:50Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, May 16, 2018 at 03:01:13PM -0400, Konstantin Ryabitsev wrote:\n\n> On 05/16/18 14:26, Martin Fick wrote:\n> > If you are going to keep the unreferenced objects around \n> > forever, it might be better to keep them around in packed \n> > form?\n> \n> I'm undecided about that. On the one hand this does create lots of small\n> files and inevitably causes (some) performance degradation. On the other\n> hand, I don't want to keep useless objects in the pack, because that\n> would also cause performance degradation for people cloning the \"mother\n> repo.\" If my assumptions on any of that are incorrect, I'm happy to\n> learn more.\n\nI implemented \"repack -k\", which keeps all objects and just rolls them\ninto the new pack (along with any currently-loose unreachable objects).\nAside from corner cases (e.g., where somebody accidentally added a 20GB\nfile to an otherwise 100MB-repo and then rolled it back), it usually\ndoesn't significantly affect the repository size.\n\nAnd it generally should not cause performance problems for people\ncloning, since Git will create a custom pack for each client with only\nthe reachable objects.\n\nThere _is_ an interesting corner case where a reachable object might be\na delta against an unreachable one, which can cause a clone to have to\nbreak that relationship and find a new delta. At GitHub we have some\ncustom code that tries to avoid these kind of delta dependencies (not\njust to unreachable objects, but to other forks that share object\nstorage). You can see the patch at:\n\n  https://github.com/peff/git jk/delta-islands\n\n-Peff\n"},{"id":"347816","messageId":"3289a942-3f0d-ff63-7eab-95fe06c4c0f6@linuxfoundation.org","threadId":"48504","inReplyTo":"20180516192343.GB3417@sigill.intra.peff.net","subject":"Re: worktrees vs. alternates","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2018-05-16T19:29:42Z","receivedAt":"2018-05-16T19:29:49Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On 05/16/18 15:23, Jeff King wrote:\n> I implemented \"repack -k\", which keeps all objects and just rolls them\n> into the new pack (along with any currently-loose unreachable objects).\n> Aside from corner cases (e.g., where somebody accidentally added a 20GB\n> file to an otherwise 100MB-repo and then rolled it back), it usually\n> doesn't significantly affect the repository size.\n\nHmm... I should read manpages more often! :)\n\nSo, do you suggest that this is a better approach:\n\n- mother repos: \"git repack -adk\"\n- child repos: \"git repack -Adl\" (followed by prune)\n\nCurrently, we do \"-Adl\" regardless, but we already track whether a repo\nis being used for alternates anywhere (so we don't prune it) and can do\ndifferent flags if that improves performance.\n\nBest,\n-- \nKonstantin Ryabitsev\nDirector, IT Infrastructure Security\nThe Linux Foundation\n\n"},{"id":"347817","messageId":"20180516193744.GA4036@sigill.intra.peff.net","threadId":"48504","inReplyTo":"3289a942-3f0d-ff63-7eab-95fe06c4c0f6@linuxfoundation.org","subject":"Re: worktrees vs. alternates","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-05-16T19:37:45Z","receivedAt":"2018-05-16T19:37:51Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, May 16, 2018 at 03:29:42PM -0400, Konstantin Ryabitsev wrote:\n\n> On 05/16/18 15:23, Jeff King wrote:\n> > I implemented \"repack -k\", which keeps all objects and just rolls them\n> > into the new pack (along with any currently-loose unreachable objects).\n> > Aside from corner cases (e.g., where somebody accidentally added a 20GB\n> > file to an otherwise 100MB-repo and then rolled it back), it usually\n> > doesn't significantly affect the repository size.\n> \n> Hmm... I should read manpages more often! :)\n> \n> So, do you suggest that this is a better approach:\n> \n> - mother repos: \"git repack -adk\"\n> - child repos: \"git repack -Adl\" (followed by prune)\n\nYes, that's pretty close to what we do at GitHub. Before doing any\nrepacking in the mother repo, we actually do the equivalent of:\n\n  git fetch --prune ../$id.git +refs/*:refs/remotes/$id/*\n  git repack -Adl\n\nfrom each child to pick up any new objects to de-duplicate (our \"mother\"\nrepos are not real repos at all, but just big shared-object stores).\n\nI say \"equivalent\" because those commands can actually be a bit slow. So\nwe do some hacky tricks like directly moving objects in the filesystem.\n\nIn theory the fetch means that it's safe to actually prune in the mother\nrepo, but in practice there are still races. They don't come up often,\nbut if you have enough repositories, they do eventually. :)\n\n-Peff\n"},{"id":"347818","messageId":"42435260.5sd4EuToWN@mfick-lnx","threadId":"48504","inReplyTo":"20180516193744.GA4036@sigill.intra.peff.net","subject":"Re: worktrees vs. alternates","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2018-05-16T19:40:56Z","receivedAt":"2018-05-16T19:41:01Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Wednesday, May 16, 2018 12:37:45 PM Jeff King wrote:\n> On Wed, May 16, 2018 at 03:29:42PM -0400, Konstantin \nRyabitsev wrote:\n> Yes, that's pretty close to what we do at GitHub. Before\n> doing any repacking in the mother repo, we actually do\n> the equivalent of:\n> \n>   git fetch --prune ../$id.git +refs/*:refs/remotes/$id/*\n>   git repack -Adl\n> \n> from each child to pick up any new objects to de-duplicate\n> (our \"mother\" repos are not real repos at all, but just\n> big shared-object stores).\n... \n> In theory the fetch means that it's safe to actually prune\n> in the mother repo, but in practice there are still\n> races. They don't come up often, but if you have enough\n> repositories, they do eventually. :)\n\nPeff,\n\nI would be very curious to hear what you think of this \napproach to mitigating the effect of those races?\n\nhttps://git.eclipse.org/r/c/122288/2\n\n-Martin\n-- \nThe Qualcomm Innovation Center, Inc. is a member of Code \nAurora Forum, hosted by The Linux Foundation\n\n"},{"id":"347819","messageId":"5156717b-6fc9-b792-dfa4-1ba48ac50333@linuxfoundation.org","threadId":"48504","inReplyTo":"20180516193744.GA4036@sigill.intra.peff.net","subject":"Re: worktrees vs. alternates","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2018-05-16T20:02:53Z","receivedAt":"2018-05-16T20:03:01Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On 05/16/18 15:37, Jeff King wrote:\n> Yes, that's pretty close to what we do at GitHub. Before doing any\n> repacking in the mother repo, we actually do the equivalent of:\n> \n>   git fetch --prune ../$id.git +refs/*:refs/remotes/$id/*\n>   git repack -Adl\n> \n> from each child to pick up any new objects to de-duplicate (our \"mother\"\n> repos are not real repos at all, but just big shared-object stores).\n\nYes, I keep thinking of doing the same, too -- instead of using\ntorvalds/linux.git for alternates, have an internal repo where objects\nfrom all forks are stored. This conversation may finally give me the\nshove I've been needing to poke at this. :)\n\nIs your delta-islands patch heading into upstream, or is that something\nthat's going to remain external?\n\n> I say \"equivalent\" because those commands can actually be a bit slow. So\n> we do some hacky tricks like directly moving objects in the filesystem.\n> \n> In theory the fetch means that it's safe to actually prune in the mother\n> repo, but in practice there are still races. They don't come up often,\n> but if you have enough repositories, they do eventually. :)\n\nI feel like a whitepaper on \"how we deal with bajillions of forks at\nGitHub\" would be nice. :) I was previously told that it's unlikely such\npaper could be written due to so many custom-built things at GH, but I\nwould be very happy if that turned out not to be the case.\n\nBest,\n-- \nKonstantin Ryabitsev\nDirector, IT Infrastructure Security\nThe Linux Foundation\n\n"},{"id":"347820","messageId":"20180516200658.GC4036@sigill.intra.peff.net","threadId":"48504","inReplyTo":"42435260.5sd4EuToWN@mfick-lnx","subject":"Re: worktrees vs. alternates","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-05-16T20:06:59Z","receivedAt":"2018-05-16T20:07:07Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, May 16, 2018 at 01:40:56PM -0600, Martin Fick wrote:\n\n> > In theory the fetch means that it's safe to actually prune\n> > in the mother repo, but in practice there are still\n> > races. They don't come up often, but if you have enough\n> > repositories, they do eventually. :)\n> \n> Peff,\n> \n> I would be very curious to hear what you think of this \n> approach to mitigating the effect of those races?\n> \n> https://git.eclipse.org/r/c/122288/2\n\nThe crux of the problem is that we have no way to atomically mark an\nobject as \"I am using this -- do not delete\" with respect to the actual\ndeletion. \n\nSo if I'm reading your approach correctly, you put objects into a\npurgatory rather than delete them, and let some operations rescue them\nfrom purgatory if we had a race.  That's certainly a direction we've\nconsidered, but I think there are some open questions, like:\n\n  1. When do you rescue from purgatory? Any time the object is\n     referenced? Do you then pull in all of its reachable objects too?\n\n  2. How do you decide when to drop an object from purgatory? And\n     specifically, how do you avoid racing with somebody using the\n     object as you're pruning purgatory?\n\n  3. How do you know that an operation has been run that will actually\n     rescue the object, as opposed to silently having a corrupted state\n     on disk?\n\n     E.g., imagine this sequence:\n\n       a. git-prune computes reachability and finds that commit X is\n          ready to be pruned\n\n       b. another process sees that commit X exists and builds a commit\n          that references it as a parent\n\n       c. git-prune drops the object into purgatory\n\n     Now we have a corrupt state created by the process in (b), since we\n     have a reachable object in purgatory. But what if nobody goes back\n     and tries to read those commits in the meantime?\n\nI think this might be solvable by using the purgatory as a kind of\n\"lock\", where prune does something like:\n\n  1. compute reachability\n\n  2. move candidate objects into purgatory; nobody can look into\n     purgatory except us\n\n  3. compute reachability _again_, making sure that no purgatory objects\n     are used (if so, rollback the deletion and try again)\n\nBut even that's not quite there, because you need to have some\nconsistent atomic view of what's \"used\". Just checking refs isn't\nenough, because some other process may be planning to reference a\npurgatory object but not yet have updated the ref. So you need some\natomic way of saying \"I am interested in using this object\".\n\n-Peff\n"},{"id":"347821","messageId":"20180516201727.GD4036@sigill.intra.peff.net","threadId":"48504","inReplyTo":"5156717b-6fc9-b792-dfa4-1ba48ac50333@linuxfoundation.org","subject":"Re: worktrees vs. alternates","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-05-16T20:17:28Z","receivedAt":"2018-05-16T20:17:40Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, May 16, 2018 at 04:02:53PM -0400, Konstantin Ryabitsev wrote:\n\n> On 05/16/18 15:37, Jeff King wrote:\n> > Yes, that's pretty close to what we do at GitHub. Before doing any\n> > repacking in the mother repo, we actually do the equivalent of:\n> > \n> >   git fetch --prune ../$id.git +refs/*:refs/remotes/$id/*\n> >   git repack -Adl\n> > \n> > from each child to pick up any new objects to de-duplicate (our \"mother\"\n> > repos are not real repos at all, but just big shared-object stores).\n> \n> Yes, I keep thinking of doing the same, too -- instead of using\n> torvalds/linux.git for alternates, have an internal repo where objects\n> from all forks are stored. This conversation may finally give me the\n> shove I've been needing to poke at this. :)\n> \n> Is your delta-islands patch heading into upstream, or is that something\n> that's going to remain external?\n\nI have vague plans to submit it upstream, but I'm still not convinced\nit's quite optimal. The resulting packs tend to be a fair bit larger\nthan they could be when packed by themselves, because we miss many delta\nopportunities (and it's important to \"repack -f --window=250\" once in a\nwhile, since we're throwing away so many delta candidates).\n\nThere's an alternative way of doing it, too, which I think git.or.cz\nuses: it \"layers\" forks in a hierarchy. So if I fork torvalds/linux.git,\nthen I get my own repo that uses torvalds/linux as an alternate. And if\nsomebody forks my repo, then I'm their alternate, and they recursively\ndepend on torvalds/linux. So each fork basically layers a slice of its\nown pack on top of the parent.\n\nThis is all from recollections of past discussions (which were sadly not\non the list -- I don't know if they've written up their scheme anywhere\npublic), so I may have some details wrong. But I think that their\nrepacking is done hierarchically, too: any objects which the root fork\nmight drop get migrated up to the children instead, and so forth, until\nthe leaf nodes can actually throw away objects.\n\nThe big problem with this is that Git tends to behave better when\nobjects are in the same pack:\n\n  1. We don't bother looking for new deltas within the same pack,\n     whereas a clone of a fork may actually try to find new deltas\n     between the layers.\n\n  2. Reachability bitmaps can't cross pack boundaries (due to the way\n     they're implemented, but also the current on-disk format). So you\n     can only bitmap the root repo, not any of the other layers.\n\n> I feel like a whitepaper on \"how we deal with bajillions of forks at\n> GitHub\" would be nice. :) I was previously told that it's unlikely such\n> paper could be written due to so many custom-built things at GH, but I\n> would be very happy if that turned out not to be the case.\n\nWe have a few engineering blog posts on the subject, like:\n\n  https://githubengineering.com/counting-objects/\n  https://githubengineering.com/introducing-dgit/\n  https://githubengineering.com/building-resilience-in-spokes/\n\nbut we haven't done a very good job of keeping that up. I think a\nsummary whitepaper would interesting. Maybe one day...:)\n\n-Peff\n"},{"id":"347822","messageId":"2848918.IL6YjaUS5T@mfick-lnx","threadId":"48504","inReplyTo":"20180516200658.GC4036@sigill.intra.peff.net","subject":"Re: worktrees vs. alternates","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2018-05-16T20:43:24Z","receivedAt":"2018-05-16T20:43:29Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Wednesday, May 16, 2018 01:06:59 PM Jeff King wrote:\n> On Wed, May 16, 2018 at 01:40:56PM -0600, Martin Fick \nwrote:\n> > > In theory the fetch means that it's safe to actually\n> > > prune in the mother repo, but in practice there are\n> > > still races. They don't come up often, but if you\n> > > have enough repositories, they do eventually. :)\n> > \n> > Peff,\n> > \n> > I would be very curious to hear what you think of this\n> > approach to mitigating the effect of those races?\n> > \n> > https://git.eclipse.org/r/c/122288/2\n> \n> The crux of the problem is that we have no way to\n> atomically mark an object as \"I am using this -- do not\n> delete\" with respect to the actual deletion.\n> \n> So if I'm reading your approach correctly, you put objects\n> into a purgatory rather than delete them, and let some\n> operations rescue them from purgatory if we had a race. \n\nYes.  This has the cost of extra disk space for a while, but \nonce I realized that we are incurring that cost already \nbecause for our repos, we already put things into purgatory \nto avoid getting stale NFS File handle errors during \nunrecoverable paths (while streaming an object).  So \neffectively this has no extra space cost then what is needed \nto run safely on NFS.\n\n>   1. When do you rescue from purgatory? Any time the\n> object is referenced? Do you then pull in all of its\n> reachable objects too?\n\nFor my approach, I decided a) Yes b) No\n\nBecause:\n\na) Rescue on reference is cheap and allows any other policy \nto be built upon it, just ensure that policy references it \nat some point before it is prune from the purgatory.\n\nb)  The other referenced objects will likely get pulled in \non reference anyway or by virtue of being in the same pack.\n\n>   2. How do you decide when to drop an object from\n> purgatory? And specifically, how do you avoid racing with\n> somebody using the object as you're pruning purgatory?\n\nIf you clean the purgatory during repacking after creating \nall the new packs and before deleting the old ones, you will \nhave a significant grace window to handle most longer running \noperations.  In this way, repacking will have re-referenced \nany missing objects from the purgatory before it gets pruned \ncausing them to be recovered if necessary.  Those missing \nobjects, believed to be in the exact packs in the purgatory \nat that time, should only ever have been referenced by write \noperations that started before those packs were moved to the \npurgatory, which was before the previous repacking round \nended.  This leaves write operations a full repacking cycle \nto complete in to avoid loosing objects.\n\n>   3. How do you know that an operation has been run that\n> will actually rescue the object, as opposed to silently\n> having a corrupted state on disk?\n> \n>      E.g., imagine this sequence:\n> \n>        a. git-prune computes reachability and finds that\n> commit X is ready to be pruned\n> \n>        b. another process sees that commit X exists and\n> builds a commit that references it as a parent\n> \n>        c. git-prune drops the object into purgatory\n> \n>      Now we have a corrupt state created by the process in\n> (b), since we have a reachable object in purgatory. But\n> what if nobody goes back and tries to read those commits\n> in the meantime?\n\nSee answer to #2, repacking itself should rescue any objects \nthat need to be rescued before pruning the purgatory.\n\n> I think this might be solvable by using the purgatory as a\n> kind of \"lock\", where prune does something like:\n> \n>   1. compute reachability\n> \n>   2. move candidate objects into purgatory; nobody can\n> look into purgatory except us\n\nI don't think this is needed.\n\nIt should be OK to let others see the objects in the \npurgatory after 1 and before 3 as long as \"seeing\" them, \ncauses them to be recovered!\n\n>   3. compute reachability _again_, making sure that no\n> purgatory objects are used (if so, rollback the deletion\n> and try again)\n\nYes, you laid out the formula, but nothing says this \nrecompute can't wait until the next repack (again see my \nanswer to #2)!  i.e. there is no rush to cause a recovery as \nlong as it gets recovered before it gets pruned from the \npurgatory.\n\n\n> But even that's not quite there, because you need to have\n> some consistent atomic view of what's \"used\". Just\n> checking refs isn't enough, because some other process\n> may be planning to reference a purgatory object but not\n> yet have updated the ref. So you need some atomic way of\n> saying \"I am interested in using this object\".\n\nAs long as all write paths also read the object first (I \nassume they do, or we would be in big trouble already), then \nthis should not be an issue.  The idea is to force all reads \n(and thus all writes also) to recover the object,\n\n-Martin\n\n-- \nThe Qualcomm Innovation Center, Inc. is a member of Code \nAurora Forum, hosted by The Linux Foundation\n\n"},{"id":"347825","messageId":"CAGZ79kaVroLhEu+QmTwLCpv1irst5zbnQBzg7xLkfFy8cC9owA@mail.gmail.com","threadId":"48504","inReplyTo":"20180516191410.GA3417@sigill.intra.peff.net","subject":"Re: worktrees vs. alternates","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2018-05-16T21:18:20Z","receivedAt":"2018-05-16T21:18:25Z","isPatch":false,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":">\n> You can also do periodic maintenance like:\n>\n>   1. Copy each ref in the forked repositories into the parent repository\n>      (e.g., giving each child that borrows from the parent its own\n>      hierarchy in refs/remotes/<child>/*).\n\nCan you just copy? I assume the mother repo doesn't know about\nall objects, hence by copying the ref, we have a \"spotty\" history.\n\nAnd to improve copying could permanent symlinking be used instead?\n"},{"id":"347886","messageId":"20180516234546.GA8521@sigill.intra.peff.net","threadId":"48504","inReplyTo":"CAGZ79kaVroLhEu+QmTwLCpv1irst5zbnQBzg7xLkfFy8cC9owA@mail.gmail.com","subject":"Re: worktrees vs. alternates","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-05-16T23:45:47Z","receivedAt":"2018-05-16T23:45:53Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, May 16, 2018 at 02:18:20PM -0700, Stefan Beller wrote:\n\n> >\n> > You can also do periodic maintenance like:\n> >\n> >   1. Copy each ref in the forked repositories into the parent repository\n> >      (e.g., giving each child that borrows from the parent its own\n> >      hierarchy in refs/remotes/<child>/*).\n> \n> Can you just copy? I assume the mother repo doesn't know about\n> all objects, hence by copying the ref, we have a \"spotty\" history.\n> \n> And to improve copying could permanent symlinking be used instead?\n\nSorry, by copying, I meant \"fetching\". I.e., migrating objects and refs.\n\n-Peff\n"},{"id":"347891","messageId":"20180517004355.GA9431@sita-lt.atc.tcs.com","threadId":"48504","inReplyTo":"5156717b-6fc9-b792-dfa4-1ba48ac50333@linuxfoundation.org","subject":"Re: worktrees vs. alternates","fromName":"Sitaram Chamarty","fromEmail":"sitaramc@gmail.com","sentAt":"2018-05-17T00:43:55Z","receivedAt":"2018-05-17T00:44:06Z","isPatch":false,"sender":{"key":"sitaramc@gmail.com","avatar":"https://avatars.githubusercontent.com/u/43316?v=4"},"body":"On Wed, May 16, 2018 at 04:02:53PM -0400, Konstantin Ryabitsev wrote:\n> On 05/16/18 15:37, Jeff King wrote:\n> > Yes, that's pretty close to what we do at GitHub. Before doing any\n> > repacking in the mother repo, we actually do the equivalent of:\n> > \n> >   git fetch --prune ../$id.git +refs/*:refs/remotes/$id/*\n> >   git repack -Adl\n> > \n> > from each child to pick up any new objects to de-duplicate (our \"mother\"\n> > repos are not real repos at all, but just big shared-object stores).\n> \n> Yes, I keep thinking of doing the same, too -- instead of using\n> torvalds/linux.git for alternates, have an internal repo where objects\n> from all forks are stored. This conversation may finally give me the\n> shove I've been needing to poke at this. :)\n\nI may have missed a few of the earlier messages, but in the last\n20 or so in this thread, I did not see namespaces mentioned by\nanyone. (I.e., apologies if it was addressed and discarded\nearlier!)\n\nI was under the impression that, as long as \"read\" access need\nnot be controlled (Konstantin's situation, at least, and maybe\nPeff's too, for public repos), namespaces are a good way to\ncreate and manage that \"mother repo\".\n\nIs that not true anymore?  Mind, I have not actually used them\nin anger anywhere, so I could be missing some really big point\nhere.\n\nsitaram\n"},{"id":"347896","messageId":"20180517033110.GA13235@sigill.intra.peff.net","threadId":"48504","inReplyTo":"20180517004355.GA9431@sita-lt.atc.tcs.com","subject":"Re: worktrees vs. alternates","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-05-17T03:31:11Z","receivedAt":"2018-05-17T03:31:17Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, May 17, 2018 at 06:13:55AM +0530, Sitaram Chamarty wrote:\n\n> I may have missed a few of the earlier messages, but in the last\n> 20 or so in this thread, I did not see namespaces mentioned by\n> anyone. (I.e., apologies if it was addressed and discarded\n> earlier!)\n> \n> I was under the impression that, as long as \"read\" access need\n> not be controlled (Konstantin's situation, at least, and maybe\n> Peff's too, for public repos), namespaces are a good way to\n> create and manage that \"mother repo\".\n> \n> Is that not true anymore?  Mind, I have not actually used them\n> in anger anywhere, so I could be missing some really big point\n> here.\n\nThe biggest problem with namespaces as they are currently implemented is\nthat they do not apply universally to all commands. If you only access\nthe repo via push/fetch, they may be fine. But as soon as you start\ndoing other operations (e.g., showing the history of a branch in a web\ninterface), you don't get to use the namespaced names anymore.\n\nI think a different implementation of namespaces could do this better.\nE.g., by controlling the view of the refs at the refs.c layer (or\nperhaps as a filtering backend).\n\n-Peff\n"},{"id":"348066","messageId":"CACsJy8AGugSYaPw9qxxXhGEz4RgawQ+moVtoGJcWhQ_=HRUOqA@mail.gmail.com","threadId":"48504","inReplyTo":"20180517033110.GA13235@sigill.intra.peff.net","subject":"Re: worktrees vs. alternates","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-05-19T05:45:56Z","receivedAt":"2018-05-19T05:46:30Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, May 17, 2018 at 5:31 AM, Jeff King <peff@peff.net> wrote:\n> On Thu, May 17, 2018 at 06:13:55AM +0530, Sitaram Chamarty wrote:\n>\n>> I may have missed a few of the earlier messages, but in the last\n>> 20 or so in this thread, I did not see namespaces mentioned by\n>> anyone. (I.e., apologies if it was addressed and discarded\n>> earlier!)\n>>\n>> I was under the impression that, as long as \"read\" access need\n>> not be controlled (Konstantin's situation, at least, and maybe\n>> Peff's too, for public repos), namespaces are a good way to\n>> create and manage that \"mother repo\".\n>>\n>> Is that not true anymore?  Mind, I have not actually used them\n>> in anger anywhere, so I could be missing some really big point\n>> here.\n>\n> The biggest problem with namespaces as they are currently implemented is\n> that they do not apply universally to all commands. If you only access\n> the repo via push/fetch, they may be fine. But as soon as you start\n> doing other operations (e.g., showing the history of a branch in a web\n> interface), you don't get to use the namespaced names anymore.\n>\n> I think a different implementation of namespaces could do this better.\n> E.g., by controlling the view of the refs at the refs.c layer (or\n> perhaps as a filtering backend).\n\nYeah. Namespaces (that work for all commands) + worktree was my plan\nfor centralizing repos (for one user). But I never got that far to\nlook into making ref namespaces work for everything.\n-- \nDuy\n"}]}