{"thread":{"id":"30758","subject":"Keeping unreachable objects in a separate pack instead of loose?","startedAt":"2012-06-10T12:31:51Z","lastAt":"2012-06-13T21:27:57Z","messageCount":48,"participants":["Theodore Ts'o","Hallvard B Furuseth","Thomas Rast","Ted Ts'o","Junio C Hamano","Jeff King","Nicolas Pitre","Hallvard Breien Furuseth","Shawn Pearce","Andreas Schwab","Martin Fick","Johan Herland"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"193252","messageId":"E1SdhJ9-0006B1-6p@tytso-glaptop.cam.corp.google.com","threadId":"30758","inReplyTo":null,"subject":"Keeping unreachable objects in a separate pack instead of loose?","fromName":"Theodore Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2012-06-10T12:31:51Z","receivedAt":"2012-06-10T12:31:51Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"I recently noticed that after a git gc, I had a huge number of loose\nobjects that were unreachable.  In fact, about 4.5 megabytes worth of\nobjects.\n\nWhen I packed them, via:\n\n   cd .git/objects ; find [0-9a-f][0-9a-f] -type f | git pack-objects pack\n\nthe resulting pack file was 244k.\n\nWhich got me thinking.... the whole point of leaving the objects loose\nis to make it easier to expire them, right?   But given how expensive it\nis to have loose objects lying around, why not:\n\na)  Have git-pack-objects have an option which writes the unreachable\n    objects into a separate pack file, instead of kicking them loose?\n\nb)  Have git-prune delete a pack only if *all* of the objects in the\n    pack meet the expiry deadline?\n\nWhat would be the downsides of pursueing such a strategy?  Is it worth\ntrying to implement as proof-of-concept?\n\n\t\t\t\t\t\t- Ted\n"},{"id":"193285","messageId":"bb7062f387c9348f702acb53803589f1.squirrel@webmail.uio.no","threadId":"30758","inReplyTo":"E1SdhJ9-0006B1-6p@tytso-glaptop.cam.corp.google.com","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Hallvard B Furuseth","fromEmail":"h.b.furuseth@usit.uio.no","sentAt":"2012-06-10T23:24:21Z","receivedAt":"2012-06-10T23:24:21Z","isPatch":false,"sender":{"key":"h.b.furuseth@usit.uio.no","avatar":null},"body":"On Sun, June 10, 2012 14:31, Theodore Ts'o wrote:\n> I recently noticed that after a git gc, I had a huge number of loose\n> objects that were unreachable.  In fact, about 4.5 megabytes worth of\n> objects.\n\nI got gigabytes once, and a full disk.  See thread\n\"git gc == git garbage-create from removed branch\", May 3 2012.\n\n> When I packed them, via:\n>\n>    cd .git/objects ; find [0-9a-f][0-9a-f] -type f | git pack-objects pack\n>\n> the resulting pack file was 244k.\n>\n> Which got me thinking.... the whole point of leaving the objects loose\n> is to make it easier to expire them, right?   But given how expensive it\n> is to have loose objects lying around, why not:\n>\n> a)  Have git-pack-objects have an option which writes the unreachable\n>     objects into a separate pack file, instead of kicking them loose?\n\nI think this should be the default.  It's very unintuitive that\ngc can eat up lots of disk space instead of saving space.\n\nUntil this is fixed, this behavior needs to be documented -\nalong with how to avoid it.\n\n> b)  Have git-prune delete a pack only if *all* of the objects in the\n>     pack meet the expiry deadline?\n>\n> What would be the downsides of pursueing such a strategy?  Is it worth\n> trying to implement as proof-of-concept?\n\nHallvard\n"},{"id":"193305","messageId":"87vcixaoxe.fsf@thomas.inf.ethz.ch","threadId":"30758","inReplyTo":"bb7062f387c9348f702acb53803589f1.squirrel@webmail.uio.no","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2012-06-11T14:44:29Z","receivedAt":"2012-06-11T14:44:29Z","isPatch":false,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"\"Hallvard B Furuseth\" <h.b.furuseth@usit.uio.no> writes:\n\n> On Sun, June 10, 2012 14:31, Theodore Ts'o wrote:\n>> I recently noticed that after a git gc, I had a huge number of loose\n>> objects that were unreachable.  In fact, about 4.5 megabytes worth of\n>> objects.\n>\n> I got gigabytes once, and a full disk.  See thread\n> \"git gc == git garbage-create from removed branch\", May 3 2012.\n>\n>> When I packed them, via:\n>>\n>>    cd .git/objects ; find [0-9a-f][0-9a-f] -type f | git pack-objects pack\n>>\n>> the resulting pack file was 244k.\n>>\n>> Which got me thinking.... the whole point of leaving the objects loose\n>> is to make it easier to expire them, right?   But given how expensive it\n>> is to have loose objects lying around, why not:\n>>\n>> a)  Have git-pack-objects have an option which writes the unreachable\n>>     objects into a separate pack file, instead of kicking them loose?\n>\n> I think this should be the default.  It's very unintuitive that\n> gc can eat up lots of disk space instead of saving space.\n>\n> Until this is fixed, this behavior needs to be documented -\n> along with how to avoid it.\n\nStarting with v1.7.10.2, and in the v1.7.11-rc versions, there's a\nchange by Peff: 7e52f56 (gc: do not explode objects which will be\nimmediately pruned, 2012-04-07).  Does it solve your problem?\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"193313","messageId":"20120611153103.GA16086@thunk.org","threadId":"30758","inReplyTo":"87vcixaoxe.fsf@thomas.inf.ethz.ch","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Ted Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2012-06-11T15:31:03Z","receivedAt":"2012-06-11T15:31:03Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Mon, Jun 11, 2012 at 04:44:29PM +0200, Thomas Rast wrote:\n> \n> Starting with v1.7.10.2, and in the v1.7.11-rc versions, there's a\n> change by Peff: 7e52f56 (gc: do not explode objects which will be\n> immediately pruned, 2012-04-07).  Does it solve your problem?\n\nI'm currently using 1.7.10.2.552.gaa3bb87, and a \"git gc\" still kicked\nloose a little over 4.5 megabytes of loose objects were not pruned via\n\"git prune\" (since they hadn't yet expired).  These loose objects\ncould be stored in a 244k pack file.\n\nSo while I'm sure that change has helped, if you happen to use a\nworkflow that uses git rebase and/or guilt and/or throwaway test\nintegration branches a lot, there will still be a large number of\nunexpired commits which still get kicked loose, and won't get pruned\nfor a week or two.\n\nWhat I think would make sense is for git pack-objects to have a new\noption which outputs a list of object id's which whould have been\nkicked out as loose objects if it had been given the (undocumented)\n--unpacked-unreachable option.  Then the git-repack shell script (if\ngiven the -A option) would use that new option instead of\n--unpacked-unreachable, and then using the list created by this new\noption, create another pack which contains all of these\nunreachable-but-not-yet-expired objects.\n\nRegards,\n\n\t\t\t\t\t\t- Ted\n"},{"id":"193316","messageId":"7vd35597qu.fsf@alter.siamese.dyndns.org","threadId":"30758","inReplyTo":"E1SdhJ9-0006B1-6p@tytso-glaptop.cam.corp.google.com","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-06-11T15:40:57Z","receivedAt":"2012-06-11T15:40:57Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Theodore Ts'o\" <tytso@mit.edu> writes:\n\n> Which got me thinking.... the whole point of leaving the objects loose\n> is to make it easier to expire them, right?   But given how expensive it\n> is to have loose objects lying around, why not:\n>\n> a)  Have git-pack-objects have an option which writes the unreachable\n>     objects into a separate pack file, instead of kicking them loose?\n>\n> b)  Have git-prune delete a pack only if *all* of the objects in the\n>     pack meet the expiry deadline?\n>\n> What would be the downsides of pursueing such a strategy?  Is it worth\n> trying to implement as proof-of-concept?\n\nI do not offhand see a downside; as a matter of fact, there has\nalready been the first-step change in the direction in v1.7.10.2\nand newer that avoids exploading the unreachable ones into loose\nobject if we know they are going to be immediately pruned.\n"},{"id":"193320","messageId":"20120611160824.GB12773@sigill.intra.peff.net","threadId":"30758","inReplyTo":"20120611153103.GA16086@thunk.org","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-06-11T16:08:24Z","receivedAt":"2012-06-11T16:08:24Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jun 11, 2012 at 11:31:03AM -0400, Ted Ts'o wrote:\n\n> I'm currently using 1.7.10.2.552.gaa3bb87, and a \"git gc\" still kicked\n> loose a little over 4.5 megabytes of loose objects were not pruned via\n> \"git prune\" (since they hadn't yet expired).  These loose objects\n> could be stored in a 244k pack file.\n\nOut of curiosity, what is the size of the whole repo? If it's a 500M\nkernel repository, then 4.5M is not all _that_ worrisome. Not that it\ncould not be better, or that it's not worth addressing (since there are\ncorner cases that behave way worse). But it gives a sense of the urgency\nof the problem, if that is the scope of the issue for average use.\n\n> What I think would make sense is for git pack-objects to have a new\n> option which outputs a list of object id's which whould have been\n> kicked out as loose objects if it had been given the (undocumented)\n> --unpacked-unreachable option.  Then the git-repack shell script (if\n> given the -A option) would use that new option instead of\n> --unpacked-unreachable, and then using the list created by this new\n> option, create another pack which contains all of these\n> unreachable-but-not-yet-expired objects.\n\nI don't think that will work, because we will keep repacking the\nunreachable bits into new packs. And the 2-week expiration is based on\nthe pack timestamp. So if your \"repack -Ad\" ends in two packs (the one\nyou actually want, and the pack of expired crap), then you would get\ninto this cycle:\n\n  1. You run \"git repack -Ad\". It makes A.pack, with stuff you want, and\n     B.pack, with unreachable junk. They both get a timestamp of \"now\".\n\n  2. A day passes. You run \"git repack -Ad\" again. It makes C.pack, the\n     new stuff you want, and repacks all of B.pack along with the\n     new expired cruft from A.pack, making D.pack. B.pack can go away.\n     D.pack gets a timestamp of \"now\".\n\nAnd so on, as long as you repack within the two week window, the objects\nfrom the cruft pack will never get ejected. So you might suggest that\nthe problem is that in step 2, we repack the items from B. But if you\ndon't, then you will accumulate a bunch of cruft packs (2 weeks worth),\nand those objects won't be delta'd against each other.  It's probably\nbetter than making them all loose, of course (you get chunks of delta'd\nobjects from each repack, instead of none at all), but it's far from a\nfull solution to the issue.\n\nI think solving it for good would involve a separate list of per-object\nexpiration dates. Obviously we get that easily with loose objects (since\nit is one object per file).\n\nAs a workaround, it might be worth lowering the default pruneExpire from\n2 weeks to 1 day or something. It is really about creating safety for\noperations in progress (e.g., you write the object, and then are _about_\nto add it to the index or update a ref when it gets pruned). I think the\n2 weeks number was pulled out of a hat as \"absurdly long for an\noperation to take\", and was never revisited because nobody cared or\ncomplained.\n\n-Peff\n"},{"id":"193330","messageId":"alpine.LFD.2.02.1206111249270.23555@xanadu.home","threadId":"30758","inReplyTo":"20120611160824.GB12773@sigill.intra.peff.net","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2012-06-11T17:04:07Z","receivedAt":"2012-06-11T17:04:07Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 11 Jun 2012, Jeff King wrote:\n\n> As a workaround, it might be worth lowering the default pruneExpire from\n> 2 weeks to 1 day or something. It is really about creating safety for\n> operations in progress (e.g., you write the object, and then are _about_\n> to add it to the index or update a ref when it gets pruned). I think the\n> 2 weeks number was pulled out of a hat as \"absurdly long for an\n> operation to take\", and was never revisited because nobody cared or\n> complained.\n\nAbsolutely.  I think this should even be considered a \"fix\" to lower \nthis value not a \"workaround\".\n\nIIRC, the 2 weeks number was instored when there wasn't any reflog on \nHEAD and the only way to salvage lost commits was to use 'git fsck \n--lost-found'.  These days, this is used only as a safety measure \nbecause there is always a tiny window during which objects are dangling \nbefore they're finally all referenced as you say.  But someone would \nhave to work hard to hit that race even if the delay was only 30 \nseconds.  So realistically this could even be set to 1 hour.\n\n\nNicolas\n"},{"id":"193333","messageId":"20120611172732.GB16086@thunk.org","threadId":"30758","inReplyTo":"20120611160824.GB12773@sigill.intra.peff.net","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Ted Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2012-06-11T17:27:32Z","receivedAt":"2012-06-11T17:27:32Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Mon, Jun 11, 2012 at 12:08:24PM -0400, Jeff King wrote:\n> On Mon, Jun 11, 2012 at 11:31:03AM -0400, Ted Ts'o wrote:\n> \n> > I'm currently using 1.7.10.2.552.gaa3bb87, and a \"git gc\" still kicked\n> > loose a little over 4.5 megabytes of loose objects were not pruned via\n> > \"git prune\" (since they hadn't yet expired).  These loose objects\n> > could be stored in a 244k pack file.\n> \n> Out of curiosity, what is the size of the whole repo? If it's a 500M\n> kernel repository, then 4.5M is not all _that_ worrisome. Not that it\n> could not be better, or that it's not worth addressing (since there are\n> corner cases that behave way worse). But it gives a sense of the urgency\n> of the problem, if that is the scope of the issue for average use.\n\nIt' my e2fsprogs development repo.  I have my \"base\" repo, which is\nwhat has been pushed out to the public (including a rewinding pu\nbranch).  The total size of that repo is a little over 15 megs:\n\n<tytso@tytso-glaptop.cam.corp.google.com> {/usr/projects/e2fsprogs/e2fsprogs}  [maint]\n899% ls ../base/objects/pack/\ntotal 16156\n  908 pack-6964a1516433f16e43dcdf4fcec1996052099f31.idx\n15248 pack-6964a1516433f16e43dcdf4fcec1996052099f31.pack\n\nI then have my development repo, which uses a\n.git/objects/info/alternates pointing at the bare \"base\" repo, so the\nonly thing in this repo are my private development branches, and other\nthings that haven't been pushed for public consumption.\n\n<tytso@tytso-glaptop.cam.corp.google.com> {/usr/projects/e2fsprogs/e2fsprogs}  [maint]\n900% ls .git/objects/pack/\ntotal 1048\n 28 5a486e6c2156109f7dfc725b36a201c10652803d.idx    28 pack-7b2a9cccab669338f61a681e34c39362976fb5de.idx\n224 5a486e6c2156109f7dfc725b36a201c10652803d.pack  768 pack-7b2a9cccab669338f61a681e34c39362976fb5de.pack\n\nThe 4.5 megabytes of loose objects packed down to a 224k \"cruft\" repo,\nand 768k worth of private development objects.\n\nSo depending on how you would want to do the comparison, probably the\nfairest thing to say is that I had a total \"good\" packs totally about\n16 megs, and the loose cruft objects was an additional 4.5 megabytes.\n\n> I don't think that will work, because we will keep repacking the\n> unreachable bits into new packs. And the 2-week expiration is based on\n> the pack timestamp. So if your \"repack -Ad\" ends in two packs (the one\n> you actually want, and the pack of expired crap), then you would get\n> into this cycle:\n> \n>   1. You run \"git repack -Ad\". It makes A.pack, with stuff you want, and\n>      B.pack, with unreachable junk. They both get a timestamp of \"now\".\n> \n>   2. A day passes. You run \"git repack -Ad\" again. It makes C.pack, the\n>      new stuff you want, and repacks all of B.pack along with the\n>      new expired cruft from A.pack, making D.pack. B.pack can go away.\n>      D.pack gets a timestamp of \"now\".\n\nHmm, yes.  What we'd really want to do is to make D.pack contain those\nitems that were are newly unreachable, not including the objects in\nB.pack, and keep B.pack around until the expiry window goes by.  But\nthat's a much more complicated thing, and the proof-of-concept\nalgorithm I had outlined wouldn't do that.\n\n> I think solving it for good would involve a separate list of per-object\n> expiration dates. Obviously we get that easily with loose objects (since\n> it is one object per file).\n\nWell, either that or we need to teach git-repack the difference\nbetween packs that are expected to contain good stuff, and packs that\ncontain cruft, and to not copy \"old cruft\" to new packs, so the old\npack can finally get nuked 2 weeks (or whatever the expire window\nmight happen to be) later.\n\n\t\t\t\t\t- Ted\n"},{"id":"193334","messageId":"20120611174507.GC16086@thunk.org","threadId":"30758","inReplyTo":"alpine.LFD.2.02.1206111249270.23555@xanadu.home","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Ted Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2012-06-11T17:45:07Z","receivedAt":"2012-06-11T17:45:07Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Mon, Jun 11, 2012 at 01:04:07PM -0400, Nicolas Pitre wrote:\n> \n> IIRC, the 2 weeks number was instored when there wasn't any reflog on \n> HEAD and the only way to salvage lost commits was to use 'git fsck \n> --lost-found'.  These days, this is used only as a safety measure \n> because there is always a tiny window during which objects are dangling \n> before they're finally all referenced as you say.  But someone would \n> have to work hard to hit that race even if the delay was only 30 \n> seconds.  So realistically this could even be set to 1 hour.\n\nThere's another useful part of the two week window, and it's as a\npartial workaround using .git/objects/info/alternates with one or more\nrewinding branches.\n\nMy /usr/projects/e2fsprogs/base repo is a bare repo that contains all\nof my public branches, including a rewinding \"pu\" branch.\n\nMy /usr/projects/e2fsprogs/e2fsprogs uses an alternates file to\nminimize disk usage, and it points at the base repo.\n\nThe problem comes when I need to gc the base repo, every 3 months or\nso.  When I do that, objects that belonged to older incarnations of\nthe rewinding pu branch disappear.  The two week window gives you a\npartial saving throw until development repo breaks due to objects that\nit depends upon disappearing.\n\nIt would be nice if a gc of the devel repo knew that some of the\nobjects it was depending on were \"expired cruft\", and copy them to the\nits local objects directory.  But of course things don't work that\nway.\n\nHere's what I do today (please don't barf; I know it's ugly):\n\n1)  cd to the base repository; run \"git repack -Adfl --window=300 --depth=100\"\n\n2)  cd to the objects directory, and create my base \"cruft\" pack:\n\tfind . --name [0-9a-f][0-9a-f] | tr -d / | git pack-objects pack-\n\n3)  hard link it into my devel repo's pack directory:\n\tln pack-* /usr/projects/e2fsprogs/e2fsprogs/.git/objects/pack\n\n4) to save space in my base repo, move it to the pack directory and\n   run git prune-packed:\n\tmv pack-* pack\n\tgit prune-packed\n\n4)  run \"git repack -Adfl --window=300 --depth=100\" in my devel repo\n\n5)  create a cruft pack in my devel repo (to save disk space):\n\tcd /usr/projects/e2fsprogs/e2fsprogs/.git/objects\n\tfind . --name [0-9a-f][0-9a-f] | tr -d / | git pack-objects pack-\n\tmv pack-* pack\n\tgit prune-packed\n\nThis probably falls in the \"don't use --shared unless you know what it\ndoes admonition in the git-clone man page.  :-)\n\nDon't worry, I don't recommend that anyone *else* do this.  But it\nworks for me (although it would be nice if I made this workflow be a\nbit more optimized; at the very least I should make a shell script\nthat does all this for me automatically, instead of typing all of the\nshell commands by hand.)\n\n   \t      \t     \t   - Ted\n"},{"id":"193335","messageId":"20120611174628.GA20134@sigill.intra.peff.net","threadId":"30758","inReplyTo":"alpine.LFD.2.02.1206111249270.23555@xanadu.home","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-06-11T17:46:28Z","receivedAt":"2012-06-11T17:46:28Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jun 11, 2012 at 01:04:07PM -0400, Nicolas Pitre wrote:\n\n> IIRC, the 2 weeks number was instored when there wasn't any reflog on \n> HEAD and the only way to salvage lost commits was to use 'git fsck \n> --lost-found'.  These days, this is used only as a safety measure \n> because there is always a tiny window during which objects are dangling \n> before they're finally all referenced as you say.  But someone would \n> have to work hard to hit that race even if the delay was only 30 \n> seconds.  So realistically this could even be set to 1 hour.\n\nNope, the HEAD reflog dates back to bd104db (enable separate reflog for\nHEAD, 2007-01-26), whereas the discussion leading to gc.pruneExpire was\nin 2008. See this sub-thread for the discussion on the time limit:\n\n  http://thread.gmane.org/gmane.comp.version-control.git/76899/focus=76943\n\nThe interesting points from that thread are:\n\n  1. You might care about object prune times if your session is\n     interrupted. So imagine you are working on something, do not\n     commit, go home for the weekend, and then come back.\n\n     I think we are OK with most in-progress operations, because any\n     blobs added to the index will of course be considered reachable and\n     not pruned. What would hurt most is that you could do:\n\n       $ hack hack hack\n       $ git add foo\n       $ hack hack hack\n       $ git add foo\n       $ hack hack hack\n       [oops! I realized I really wanted the initial version of foo!]\n       $ git fsck --lost-found\n\n     If your session is interrupted during the third \"hack hack hack\"\n     bit, a short prune might lose the initial version. Of course,\n     that's _if_ a gc is run in the middle. So I find it possible, but\n     somewhat unlikely.\n\n  2. If a branch is deleted, all of its reflogs go away immediately. A\n     short expiration time would mean that the objects will probably go\n     away at the next \"gc\".\n\n     In many cases, you will be saved by the HEAD reflog, which stays\n     around until the real expiration is reached. But not always (e.g.,\n     server-side reflogs).\n\n     I'd much rather address this in the general case by saving\n     deleted-branch reflogs, though I know you and I have disagreed on\n     that in the past.\n\nOne issue not brought up in that thread is that of bare repositories,\nwhich do not have reflogs enabled by default. In that case, the 2-week\nprune window is really doing something.\n\nI really wonder if there is a good reason not to turn on reflogs for all\nrepositories, bare included. We have them on for all of the server-side\nrepositories at github, and they have been invaluable for resolving\nmany support tickets from people who forced a push without realizing\nwhat they were doing.\n\n-Peff\n"},{"id":"193337","messageId":"20120611175419.GB20134@sigill.intra.peff.net","threadId":"30758","inReplyTo":"20120611174507.GC16086@thunk.org","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-06-11T17:54:19Z","receivedAt":"2012-06-11T17:54:19Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jun 11, 2012 at 01:45:07PM -0400, Ted Ts'o wrote:\n\n> There's another useful part of the two week window, and it's as a\n> partial workaround using .git/objects/info/alternates with one or more\n> rewinding branches.\n> \n> My /usr/projects/e2fsprogs/base repo is a bare repo that contains all\n> of my public branches, including a rewinding \"pu\" branch.\n> \n> My /usr/projects/e2fsprogs/e2fsprogs uses an alternates file to\n> minimize disk usage, and it points at the base repo.\n> \n> The problem comes when I need to gc the base repo, every 3 months or\n> so.  When I do that, objects that belonged to older incarnations of\n> the rewinding pu branch disappear.  The two week window gives you a\n> partial saving throw until development repo breaks due to objects that\n> it depends upon disappearing.\n\nYou're doing it wrong (but you can hardly be blamed, because there isn't\ngood tool support for doing it right). You should never prune or repack\nin the base repo without taking into account all of the refs of its\nchildren.\n\nWe have a similar setup at github (every time you \"fork\" a repo, it is\ncreating a new repo that links back to a project-wide \"network\" repo for\nits object store). We maintain a refs/remotes/XXX directory for each\nchild repo, which stores the complete refs/ hierarchy of that child.\n\nIt's all done by a tangled mass of shell scripts. I've considered trying\nto clean up and open source them, but I really doubt it would help\npeople generally. There's so much stuff specific to our setup.\n\n-Peff\n"},{"id":"193341","messageId":"20120611182012.GD16086@thunk.org","threadId":"30758","inReplyTo":"20120611175419.GB20134@sigill.intra.peff.net","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Ted Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2012-06-11T18:20:12Z","receivedAt":"2012-06-11T18:20:12Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Mon, Jun 11, 2012 at 01:54:19PM -0400, Jeff King wrote:\n> \n> You're doing it wrong (but you can hardly be blamed, because there isn't\n> good tool support for doing it right). You should never prune or repack\n> in the base repo without taking into account all of the refs of its\n> children.\n\nWell, I don't do a simple gc.  See the complicated set of steps I use\nto make sure I don't lose loose commits at the end of my last e-mail\nmessage on this thread.  It gets worse when I have multiple devel\nrepos, but I simplified things for the purposes of discussion.\n\n> We have a similar setup at github (every time you \"fork\" a repo, it is\n> creating a new repo that links back to a project-wide \"network\" repo for\n> its object store). We maintain a refs/remotes/XXX directory for each\n> child repo, which stores the complete refs/ hierarchy of that child.\n\nSo you basically are copying the refs around and making sure the\nparent repo has an uptodate pointer of all of the child repos, such\nthat when you do the repack, *all* of the commits end up in the parent\ncommit, correct?\n\nThe system that I'm using means that objects which are local to a\nchild repo stays in the child repo, and if an object is about to be\ndropped from the parent repo as a result of a gc, the child repo has\nan opportunity claim a copy of that object for itself in its object\ndatabase.\n\nYou can do things either way.  I like knowing that objects only used\nby a child repo are in the child repo's .git directory, but that's\narguably more of a question of taste than anything else.\n\n\t      \t   \t       \t     - Ted\n"},{"id":"193344","messageId":"20120611183414.GD20134@sigill.intra.peff.net","threadId":"30758","inReplyTo":"20120611172732.GB16086@thunk.org","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-06-11T18:34:14Z","receivedAt":"2012-06-11T18:34:14Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jun 11, 2012 at 01:27:32PM -0400, Ted Ts'o wrote:\n\n> The 4.5 megabytes of loose objects packed down to a 224k \"cruft\" repo,\n> and 768k worth of private development objects.\n> \n> So depending on how you would want to do the comparison, probably the\n> fairest thing to say is that I had a total \"good\" packs totally about\n> 16 megs, and the loose cruft objects was an additional 4.5 megabytes.\n\nOK, so that 4.5 is at least a respectable percentage of the total repo\nsize. I suspect it may be worse for small repos in that sense, because\nthe 4.5 megabytes is not \"how big is this repo\" but probably \"how much\nwork did you do in this repo in the last 2 weeks\". Which should be a\nconstant with respect to the total size of the repo.\n\nHowever, for a very busy repo (e.g., one used for automated integration\ntesting or similar), the \"how much work\" number could be quite high.\n\nWe ran into this at github due to our \"merge this\" button, which does\ntest-merges to see if a pull request can be merged cleanly. Whenever the\nupstream branch is updated, all of the outstanding pull requests get\nre-tested, and the old test-merge objects become unreferenced.  We ended\nup dropping pruneExpire to 1 day to keep the cruft to a minimum.\n\n> >   1. You run \"git repack -Ad\". It makes A.pack, with stuff you want,\n> >   and B.pack, with unreachable junk. They both get a timestamp of\n> >   \"now\".\n> > \n> >   2. A day passes. You run \"git repack -Ad\" again. It makes C.pack,\n> >   the new stuff you want, and repacks all of B.pack along with the\n> >   new expired cruft from A.pack, making D.pack. B.pack can go away.\n> >   D.pack gets a timestamp of \"now\".\n> \n> Hmm, yes.  What we'd really want to do is to make D.pack contain those\n> items that were are newly unreachable, not including the objects in\n> B.pack, and keep B.pack around until the expiry window goes by.  But\n> that's a much more complicated thing, and the proof-of-concept\n> algorithm I had outlined wouldn't do that.\n\nRight. When dumping the list of unreachable objects from pack-objects,\nyou'd want to tell it to ignore objects from any pack that contained\nonly unreachable objects in the first place.\n\nExcept that is perhaps not quite right, either. Because if a pack has\n100 objects, but you happen to re-reference 1 of them, you'd probably\nwant to leave it (even though that re-referenced one is now duplicated,\nthe savings aren't worth the trouble of repacking the cruft objects).\nOf course, what is the N that makes it worth the trouble? Now you're\ngetting into heuristics.\n\nYou _could_ make a separate cruft pack for each pack that you repack. So\nif I have A.pack and B.pack, I'd pack all of the reachable objects into\nC.pack, and then make D.pack containing the unreachable objects from\nA.pack, and E.pack with the unreachable objects from B.pack. And then\nset the mtime of the cruft packs to that of their parent packs.\n\nAnd then the next time you pack, repacking D and E would probably be a\nno-op that preserves mtime, but might create a new pack that ejects some\nnow-reachable object.\n\nTo implement that, I think your --list-unreachable would just have to\nprint a list of \"<pack-mtime> <sha1>\" pairs, and then you would pack\neach set with an identical mtime (or even a \"close enough\" mtime within\nsome slop).\n\nBut yet, this is all getting complicated. :)\n\n> > I think solving it for good would involve a separate list of\n> > per-object expiration dates. Obviously we get that easily with loose\n> > objects (since it is one object per file).\n> \n> Well, either that or we need to teach git-repack the difference\n> between packs that are expected to contain good stuff, and packs that\n> contain cruft, and to not copy \"old cruft\" to new packs, so the old\n> pack can finally get nuked 2 weeks (or whatever the expire window\n> might happen to be) later.\n\nThat is harder because those objects may become re-reachable during that\nwindow. So I think you don't want to deal with \"expected to contain...\"\nbut rather \"what does it contain now?\". The latter is easy to figure out\nby doing a reachability analysis (which we do as part of the repack\nanyway).\n\n-Peff\n"},{"id":"193345","messageId":"20120611184349.GE20134@sigill.intra.peff.net","threadId":"30758","inReplyTo":"20120611182012.GD16086@thunk.org","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-06-11T18:43:49Z","receivedAt":"2012-06-11T18:43:49Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jun 11, 2012 at 02:20:12PM -0400, Ted Ts'o wrote:\n\n> On Mon, Jun 11, 2012 at 01:54:19PM -0400, Jeff King wrote:\n> > \n> > You're doing it wrong (but you can hardly be blamed, because there isn't\n> > good tool support for doing it right). You should never prune or repack\n> > in the base repo without taking into account all of the refs of its\n> > children.\n> \n> Well, I don't do a simple gc.  See the complicated set of steps I use\n> to make sure I don't lose loose commits at the end of my last e-mail\n> message on this thread.  It gets worse when I have multiple devel\n> repos, but I simplified things for the purposes of discussion.\n\nAh, right. I was thinking that your first step, which is \"git repack\n-Adfl\", would throw out old objects rather than unpack them in recent\nversions of git.  But due to the way I implemented it (namely that you\nmust pass --unpack-unreachable yourself, so this feature only kicks in\nautomatically for \"git gc\"), that is not the case.\n\nI don't recall if that was an accident, or if I was very clever in\nmaintaining backwards compatibility for your case. Let's just assume the\nlatter. :)\n\n> > We have a similar setup at github (every time you \"fork\" a repo, it is\n> > creating a new repo that links back to a project-wide \"network\" repo for\n> > its object store). We maintain a refs/remotes/XXX directory for each\n> > child repo, which stores the complete refs/ hierarchy of that child.\n> \n> So you basically are copying the refs around and making sure the\n> parent repo has an uptodate pointer of all of the child repos, such\n> that when you do the repack, *all* of the commits end up in the parent\n> commit, correct?\n\nYes. The child repositories generally have no objects in them at all\n(they occasionally do for a period between runs of the migration\nscript).\n\n> The system that I'm using means that objects which are local to a\n> child repo stays in the child repo, and if an object is about to be\n> dropped from the parent repo as a result of a gc, the child repo has\n> an opportunity claim a copy of that object for itself in its object\n> database.\n\nThat implies the concept of \"local to a child repo\", which implies that\nyou have some set of \"common\" refs. I suspect in your case your base\nrepo represents the master branch, or something similar. We actually\ntreat our network repo as a pure parent; every repo, including the\noriginal one that everybody forks from, is a child. That makes it easier\nto treat the original repo as just another repo (e.g., the original\nowner is free to delete it, and the forks won't care).\n\n> You can do things either way.  I like knowing that objects only used\n> by a child repo are in the child repo's .git directory, but that's\n> arguably more of a question of taste than anything else.\n\nYeah, I don't think there is any real benefit to it.\n\n-Peff\n"},{"id":"193380","messageId":"0450a24b1f53420f36a3d864c50536cb@ulrik.uio.no","threadId":"30758","inReplyTo":"20120611183414.GD20134@sigill.intra.peff.net","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Hallvard Breien Furuseth","fromEmail":"h.b.furuseth@usit.uio.no","sentAt":"2012-06-11T20:44:39Z","receivedAt":"2012-06-11T20:44:39Z","isPatch":false,"sender":{"key":"h.b.furuseth@usit.uio.no","avatar":null},"body":" On Mon, 11 Jun 2012 14:34:14 -0400, Jeff King <peff@peff.net> wrote:\n> On Mon, Jun 11, 2012 at 01:27:32PM -0400, Ted Ts'o wrote:\n>> So depending on how you would want to do the comparison, probably \n>> the\n>> fairest thing to say is that I had a total \"good\" packs totally \n>> about\n>> 16 megs, and the loose cruft objects was an additional 4.5 \n>> megabytes.\n>\n> OK, so that 4.5 is at least a respectable percentage of the total \n> repo\n> size. I suspect it may be worse for small repos in that sense, (...)\n\n 'git gc' gave a 3100% increase with my example:\n\n     $ git clone --bare --branch linux-overhaul-20010122 \\\n         git://git.savannah.gnu.org/config.git\n     $ cd config.git/\n     $ git tag -d `git tag`; git branch -D master\n     $ du -s objects\n     624     objects\n     $ git gc\n     $ du -s objects\n     19840   objects\n\n Basically: Clone/fetch a repo, keep a small part of it, drop the\n rest, and gc.  Gc explodes all the objects you no longer want.\n\n This hits you hard if your small project tracks a big one, and\n later ceases doing so.  Maybe your project is to convert a small\n part of the remote into a library, or you're just tracking a few\n files from the remote - e.g. from the GNU Config repo.\n\n Not something one does every day, but it does happen.\n Tweaks to expiration time etc do not help here.\n\n Hallvard\n"},{"id":"193384","messageId":"20120611211401.GA21775@thunk.org","threadId":"30758","inReplyTo":"20120611183414.GD20134@sigill.intra.peff.net","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Ted Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2012-06-11T21:14:01Z","receivedAt":"2012-06-11T21:14:01Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Mon, Jun 11, 2012 at 02:34:14PM -0400, Jeff King wrote:\n> You _could_ make a separate cruft pack for each pack that you repack. So\n> if I have A.pack and B.pack, I'd pack all of the reachable objects into\n> C.pack, and then make D.pack containing the unreachable objects from\n> A.pack, and E.pack with the unreachable objects from B.pack. And then\n> set the mtime of the cruft packs to that of their parent packs.\n> \n> And then the next time you pack, repacking D and E would probably be a\n> no-op that preserves mtime, but might create a new pack that ejects some\n> now-reachable object.\n> \n> To implement that, I think your --list-unreachable would just have to\n> print a list of \"<pack-mtime> <sha1>\" pairs, and then you would pack\n> each set with an identical mtime (or even a \"close enough\" mtime within\n> some slop)....\n\nHow about this instead?  We distinguish between cruft packs and \"real\"\npacks by the filename.  So we have \"cruft-<SHA1>.{idx,pack}\" and\n\"pack-<SHA1>.{idx.pack}\".\n\nNormally, git will look at any pack in the pack directory that has an\n.idx and .pack extension, but during repack operation, it will by only\nlook in the pack-* packs first.  If it can't find an object there, it\nwill then fall back to trying to fetch the object from the cruft-*\npacks, and if it finds the object, it copies it into the new pack\nwhich is creating, thus \"rescueing\" an object which reappears during\nthe expiry window.  This should be a relatively rare event, and if it\nhappens, the object will be in two packs, a pack-* pack and a cruft-*\npack, but that's OK.\n\nSo since git pack-objects isn't even looking in the cruft-* packs\nexcept when it needs to rescue an object, the objects in the cruft-*\npacks won't get copied, and we won't need to have per-object mtimes.\nIt also means it will go faster since it's not copying the cruft-*\npacks at all, and possibly not even looking at them.\n\nNow all we need to do is delete any cruft-* packs which are older than\nthe expiry window.  We don't even need to look at their contents.\n\nIt does imply that we may accumulate a new cruft-<SHA1> pack each time\nwe run git gc, but users shouldn't be running git gc all that often\nanyway.  And even if they do run it all the time, it will still be\nmore efficient than keeping the unreachable objects as loose objects.\n\n     \t       \t    \t    \t\t    \t    - Ted\n"},{"id":"193385","messageId":"20120611211414.GA32061@sigill.intra.peff.net","threadId":"30758","inReplyTo":"0450a24b1f53420f36a3d864c50536cb@ulrik.uio.no","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-06-11T21:14:14Z","receivedAt":"2012-06-11T21:14:14Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jun 11, 2012 at 10:44:39PM +0200, Hallvard Breien Furuseth wrote:\n\n> >OK, so that 4.5 is at least a respectable percentage of the total\n> >repo\n> >size. I suspect it may be worse for small repos in that sense, (...)\n> \n> 'git gc' gave a 3100% increase with my example:\n> \n>     $ git clone --bare --branch linux-overhaul-20010122 \\\n>         git://git.savannah.gnu.org/config.git\n>     $ cd config.git/\n>     $ git tag -d `git tag`; git branch -D master\n>     $ du -s objects\n>     624     objects\n>     $ git gc\n>     $ du -s objects\n>     19840   objects\n> \n> Basically: Clone/fetch a repo, keep a small part of it, drop the\n> rest, and gc.  Gc explodes all the objects you no longer want.\n\nI would argue that this is not a very interesting case in the first\nplace, since the right thing to do is use \"clone --single-branch\"[1] to\nvoid transferring all of those objects in the first place.\n\nBut there are plenty of variant cases, where you are not just deleting\nall of the refs, but rather doing some manipulation of the branches,\ndiffing them to make sure it is safe to drop some bits, running\nfilter-branch, etc. And it would be nice to make those cases more\nefficient.\n\n-Peff\n\n[1] It looks like \"--single-branch\" does not actually work, and still\n    fetches master. I think this is a bug in the implementation of\n    single-branch (it looks like it fetches HEAD unconditionally).\n    +cc Duy.\n"},{"id":"193388","messageId":"20120611213948.GB32061@sigill.intra.peff.net","threadId":"30758","inReplyTo":"20120611211401.GA21775@thunk.org","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-06-11T21:39:48Z","receivedAt":"2012-06-11T21:39:48Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jun 11, 2012 at 05:14:01PM -0400, Ted Ts'o wrote:\n\n> On Mon, Jun 11, 2012 at 02:34:14PM -0400, Jeff King wrote:\n> > You _could_ make a separate cruft pack for each pack that you repack. So\n> > if I have A.pack and B.pack, I'd pack all of the reachable objects into\n> > C.pack, and then make D.pack containing the unreachable objects from\n> > A.pack, and E.pack with the unreachable objects from B.pack. And then\n> > set the mtime of the cruft packs to that of their parent packs.\n> > \n> > And then the next time you pack, repacking D and E would probably be a\n> > no-op that preserves mtime, but might create a new pack that ejects some\n> > now-reachable object.\n> > \n> > To implement that, I think your --list-unreachable would just have to\n> > print a list of \"<pack-mtime> <sha1>\" pairs, and then you would pack\n> > each set with an identical mtime (or even a \"close enough\" mtime within\n> > some slop)....\n> \n> How about this instead?  We distinguish between cruft packs and \"real\"\n> packs by the filename.  So we have \"cruft-<SHA1>.{idx,pack}\" and\n> \"pack-<SHA1>.{idx.pack}\".\n> \n> Normally, git will look at any pack in the pack directory that has an\n> .idx and .pack extension, but during repack operation, it will by only\n> look in the pack-* packs first.  If it can't find an object there, it\n> will then fall back to trying to fetch the object from the cruft-*\n> packs, and if it finds the object, it copies it into the new pack\n> which is creating, thus \"rescueing\" an object which reappears during\n> the expiry window.  This should be a relatively rare event, and if it\n> happens, the object will be in two packs, a pack-* pack and a cruft-*\n> pack, but that's OK.\n\nYou don't need to do anything magical for the object lookup process of\npack-objects. By definition, the unreachable objects will not be\nincluded in the new pack you are creating, because it is only packing\nreachable things. So the cruft pack does not have to be a fallback at\nall; it is a regular pack from the object lookup perspective.\n\nThe important differences from the current behavior would be:\n\n  1. In pack-objects, do not explode (or if just listing, do not list)\n     objects from cruft packs.\n\n  2. In repack, make a new pack from the list of unreachable objects.\n\n  3. During \"repack -Ad\", prune cruft packs only if they are older than\n     an expiration date.\n\nAnd then you'd end up with a new cruft pack each time. You could just\nmark it with a \".cruft\" file (similar to the existing \".keep\" files),\nand you don't have to worry about changing the pack-* name.\n\n> So since git pack-objects isn't even looking in the cruft-* packs\n> except when it needs to rescue an object, the objects in the cruft-*\n> packs won't get copied, and we won't need to have per-object mtimes.\n> It also means it will go faster since it's not copying the cruft-*\n> packs at all, and possibly not even looking at them.\n\nYeah. It doesn't eliminate duplicates, but that may not be worth caring\nabout. I find the \"cruft\" marking a little hacky, because it is only\n\"objects in here _may_ be cruft\", but as long as that is understood, it\nis OK (and it is understood in the sequence above; \"repack -Ad\" is safe\nbecause it knows that it would have repacked any non-cruft).\n\nYou would have to be careful with \"repack -d\" (without the \"-a\"). It\nwould not be necessarily be safe to remove cruft packs, because you\nmight not have rescued the objects. However, AFAICT \"repack -d\" does not\ncurrently delete packs at all, so this would be no different.\n\n> It does imply that we may accumulate a new cruft-<SHA1> pack each time\n> we run git gc, but users shouldn't be running git gc all that often\n> anyway.  And even if they do run it all the time, it will still be\n> more efficient than keeping the unreachable objects as loose objects.\n\nYeah, it would be nice to keep it all in a single pack, but that means\ndoing the I/O on rewriting the cruft packs each time. And figuring out\nsome way of handling the mtime in such a way that we don't keep\nrefreshing the age during each gc.\n\nSpeaking of which, what is the mtime of the newly created cruft pack? Is\nit the current mtime? Then those unreachable objects will stick for\nanother 2 weeks, instead of being back-dated to their pack's date. You\ncould back-date to the mtime of the most recent deleted pack, but that\nwould still prolong the life of objects from the older packs. It may be\nacceptable to just ignore the issue, though; they will expire\neventually.\n\n-Peff\n"},{"id":"193389","messageId":"97844573ed775a758d75eec5508b6c85@ulrik.uio.no","threadId":"30758","inReplyTo":"20120611211414.GA32061@sigill.intra.peff.net","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Hallvard Breien Furuseth","fromEmail":"h.b.furuseth@usit.uio.no","sentAt":"2012-06-11T21:41:56Z","receivedAt":"2012-06-11T21:41:56Z","isPatch":false,"sender":{"key":"h.b.furuseth@usit.uio.no","avatar":null},"body":" On Mon, 11 Jun 2012 17:14:14 -0400, Jeff King <peff@peff.net> wrote:\n> On Mon, Jun 11, 2012 at 10:44:39PM +0200, Hallvard Breien Furuseth \n> wrote:\n>\n>> >OK, so that 4.5 is at least a respectable percentage of the total\n>> >repo\n>> >size. I suspect it may be worse for small repos in that sense, \n>> (...)\n>>\n>> 'git gc' gave a 3100% increase with my example:\n>>\n>>     $ git clone --bare --branch linux-overhaul-20010122 \\\n>>         git://git.savannah.gnu.org/config.git\n>>     $ cd config.git/\n>>     $ git tag -d `git tag`; git branch -D master\n>>     $ du -s objects\n>>     624     objects\n>>     $ git gc\n>>     $ du -s objects\n>>     19840   objects\n>>\n>> Basically: Clone/fetch a repo, keep a small part of it, drop the\n>> rest, and gc.  Gc explodes all the objects you no longer want.\n>\n> I would argue that this is not a very interesting case in the first\n> place, since the right thing to do is use \"clone --single-branch\"[1] \n> to\n> void transferring all of those objects in the first place.\n\n Yeah, I just wanted a simple way to fetch a lot and drop most of\n it, since that's what triggers the problem.  A simple use case\n would be where you did a lot of work between cloing and pruning\n most refs.  I described some actual use cases below the command.\n\n> But there are plenty of variant cases, where you are not just \n> deleting\n> all of the refs, but rather doing some manipulation of the branches,\n> diffing them to make sure it is safe to drop some bits, running\n> filter-branch, etc. And it would be nice to make those cases more\n> efficient.\n\n-- \n Hallvard\n"},{"id":"193396","messageId":"20120611221439.GE21775@thunk.org","threadId":"30758","inReplyTo":"20120611213948.GB32061@sigill.intra.peff.net","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Ted Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2012-06-11T22:14:39Z","receivedAt":"2012-06-11T22:14:39Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Mon, Jun 11, 2012 at 05:39:48PM -0400, Jeff King wrote:\n> \n> Yeah. It doesn't eliminate duplicates, but that may not be worth caring\n> about. I find the \"cruft\" marking a little hacky, because it is only\n> \"objects in here _may_ be cruft\", but as long as that is understood, it\n> is OK (and it is understood in the sequence above; \"repack -Ad\" is safe\n> because it knows that it would have repacked any non-cruft).\n\nWell, all the objects in the file *were* cruft at the time that it was\ncreated.  And the reason why we are keeping them around is in case we\nwere wrong about their being cruft, so I guess I don't have that much\ntrouble with the name.  Something like \"KillShelter\" (as in the\nopposite of No-Kill Animal Shelters) would be more discriptive, but I\nthink it's a bit lacking in taste....\n\n> > It does imply that we may accumulate a new cruft-<SHA1> pack each time\n> > we run git gc, but users shouldn't be running git gc all that often\n> > anyway.  And even if they do run it all the time, it will still be\n> > more efficient than keeping the unreachable objects as loose objects.\n> \n> Yeah, it would be nice to keep it all in a single pack, but that means\n> doing the I/O on rewriting the cruft packs each time. And figuring out\n> some way of handling the mtime in such a way that we don't keep\n> refreshing the age during each gc.\n\nWell, I'd like to avoid doing the I/O because I want to minimize wear\non SSD drives; and given that it's unlikely that the cruft packs will\nbe referenced, the fact that we have a bunch of cruft packs shouldn't\nbe a big deal, especially if we teach git to search the cruft packs\nlast.\n\n> Speaking of which, what is the mtime of the newly created cruft pack? Is\n> it the current mtime? Then those unreachable objects will stick for\n> another 2 weeks, instead of being back-dated to their pack's date. You\n> could back-date to the mtime of the most recent deleted pack, but that\n> would still prolong the life of objects from the older packs. It may be\n> acceptable to just ignore the issue, though; they will expire\n> eventually.\n\nWell, we have that problem today when \"git pack-objects\n--unpack-unreachable\" explodes unreferenced objects --- they are\nwritten with the current mtime.  I assume you're worried about\npre-existing loose objects that get collected up into a new cruft\npack, since they would get the extra two weeks of life.  Given how\nmuch more efficient storing the cruft objects in a pack, I think\nignoring the issue is what makes the most amount of sense, since it's\na one-time extension, and the extra objects really won't do any harm.\n\nOne last thought: if a sysadmin is really hard up for space, (and if\nthe cruft objects include some really big sound or video files) one\nadvantage of labelling the cruft packs explicitly is that someone who\nreally needs the space could potentially find the oldest cruft files\nand delete them, since they would be tagged for easy findability.\n\n    \t   \t       \t    \t     \t    - Ted\n"},{"id":"193397","messageId":"20120611222308.GA10476@sigill.intra.peff.net","threadId":"30758","inReplyTo":"20120611221439.GE21775@thunk.org","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-06-11T22:23:08Z","receivedAt":"2012-06-11T22:23:08Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jun 11, 2012 at 06:14:39PM -0400, Ted Ts'o wrote:\n\n> > Speaking of which, what is the mtime of the newly created cruft pack? Is\n> > it the current mtime? Then those unreachable objects will stick for\n> > another 2 weeks, instead of being back-dated to their pack's date. You\n> > could back-date to the mtime of the most recent deleted pack, but that\n> > would still prolong the life of objects from the older packs. It may be\n> > acceptable to just ignore the issue, though; they will expire\n> > eventually.\n> \n> Well, we have that problem today when \"git pack-objects\n> --unpack-unreachable\" explodes unreferenced objects --- they are\n> written with the current mtime.\n\nNo, we don't; they get the mtime of the pack they are coming from (and\nif the pack is older than pruneExpire, they are not exploded at all,\nsince they would just be pruned immediately anyway).\n\nSo an exploded object might have only a day or an hour to live after the\nexplosion, but with your strategy they always get two weeks.\n\n> I assume you're worried about pre-existing loose objects that get\n> collected up into a new cruft pack, since they would get the extra two\n> weeks of life.  Given how much more efficient storing the cruft\n> objects in a pack, I think ignoring the issue is what makes the most\n> amount of sense, since it's a one-time extension, and the extra\n> objects really won't do any harm.\n\nI'm more specifically worried about large objects which are no better in\npacks than they are in loose form (e.g., video files). This strategy is\na regression, since we are not saving space by putting them in a pack,\nbut we are keeping them around much longer. It also makes it harder to\njust run \"git prune\" to get rid of large objects (since prune will never\nkill off a pack), or to manually delete files from the object database.\nYou have to run \"git gc --prune=now\" instead, so it can make a new pack\nand throw away the old bits (or run \"git repack -ad\").\n\n> One last thought: if a sysadmin is really hard up for space, (and if\n> the cruft objects include some really big sound or video files) one\n> advantage of labelling the cruft packs explicitly is that someone who\n> really needs the space could potentially find the oldest cruft files\n> and delete them, since they would be tagged for easy findability.\n\nNo! That's exactly what I was worried about with the name. It is _not_\nsafe to do so. It's only safe after you have done a full repack to\nrescue any non-cruft objects.\n\n-Peff\n"},{"id":"193398","messageId":"20120611222843.GF21775@thunk.org","threadId":"30758","inReplyTo":"20120611222308.GA10476@sigill.intra.peff.net","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Ted Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2012-06-11T22:28:43Z","receivedAt":"2012-06-11T22:28:43Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Mon, Jun 11, 2012 at 06:23:08PM -0400, Jeff King wrote:\n> \n> I'm more specifically worried about large objects which are no better in\n> packs than they are in loose form (e.g., video files). This strategy is\n> a regression, since we are not saving space by putting them in a pack,\n> but we are keeping them around much longer. It also makes it harder to\n> just run \"git prune\" to get rid of large objects (since prune will never\n> kill off a pack), or to manually delete files from the object database.\n> You have to run \"git gc --prune=now\" instead, so it can make a new pack\n> and throw away the old bits (or run \"git repack -ad\").\n\nIf we're really worried about this, we could set a threshold and only\npack small objects in the cruft packs.\n\n> > One last thought: if a sysadmin is really hard up for space, (and if\n> > the cruft objects include some really big sound or video files) one\n> > advantage of labelling the cruft packs explicitly is that someone who\n> > really needs the space could potentially find the oldest cruft files\n> > and delete them, since they would be tagged for easy findability.\n> \n> No! That's exactly what I was worried about with the name. It is _not_\n> safe to do so. It's only safe after you have done a full repack to\n> rescue any non-cruft objects.\n\nWell, yes.  I was thinking it would be safe thing to do after a \"git\ngc\" didn't result in enough space savings.  This would require that a\ngit repack always rescue objects from cruft packs even if the -a/-A\noptions are not specified, but since we're doing a full reachability\nscan, that should slow down git gc much, right?\n\n    \t\t\t\t\t- Ted\n"},{"id":"193399","messageId":"20120611223546.GA10619@sigill.intra.peff.net","threadId":"30758","inReplyTo":"20120611222843.GF21775@thunk.org","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-06-11T22:35:46Z","receivedAt":"2012-06-11T22:35:46Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jun 11, 2012 at 06:28:43PM -0400, Ted Ts'o wrote:\n\n> On Mon, Jun 11, 2012 at 06:23:08PM -0400, Jeff King wrote:\n> > \n> > I'm more specifically worried about large objects which are no better in\n> > packs than they are in loose form (e.g., video files). This strategy is\n> > a regression, since we are not saving space by putting them in a pack,\n> > but we are keeping them around much longer. It also makes it harder to\n> > just run \"git prune\" to get rid of large objects (since prune will never\n> > kill off a pack), or to manually delete files from the object database.\n> > You have to run \"git gc --prune=now\" instead, so it can make a new pack\n> > and throw away the old bits (or run \"git repack -ad\").\n> \n> If we're really worried about this, we could set a threshold and only\n> pack small objects in the cruft packs.\n\nI think I'd be more inclined to just ignore it. It is only prolonging\nthe lifetime of the files by a finite amount (and we are discussing\ndropping that finite amount anyway). And as a bonus, this strategy could\npotentially allow an optimization that would make large files better in\nthis case: if we notice that a pack has _only_ unreachable objects, we\ncan simply mark it as \".cruft\" without actually repacking it. Coupled\nwith the recent-ish code to stream large blobs directly to packs, that\nmeans a large blob which becomes unreachable would not ever be\nrewritten.\n\n> > No! That's exactly what I was worried about with the name. It is _not_\n> > safe to do so. It's only safe after you have done a full repack to\n> > rescue any non-cruft objects.\n> \n> Well, yes.  I was thinking it would be safe thing to do after a \"git\n> gc\" didn't result in enough space savings.  This would require that a\n> git repack always rescue objects from cruft packs even if the -a/-A\n> options are not specified, but since we're doing a full reachability\n> scan, that should slow down git gc much, right?\n\nDoing \"git gc\" will always repack everything, IIRC. It is \"git gc\n--auto\" which will make small incremental packs. I think we do a full\nreachability analysis so we can prune there, but that is something I\nthink we should stop doing. It is typically orders of magnitude slower\nthan the incremental repack.\n\n-Peff\n"},{"id":"193406","messageId":"alpine.LFD.2.02.1206112024110.23555@xanadu.home","threadId":"30758","inReplyTo":"20120611222308.GA10476@sigill.intra.peff.net","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2012-06-12T00:41:03Z","receivedAt":"2012-06-12T00:41:03Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 11 Jun 2012, Jeff King wrote:\n\n> On Mon, Jun 11, 2012 at 06:14:39PM -0400, Ted Ts'o wrote:\n> \n> > One last thought: if a sysadmin is really hard up for space, (and if\n> > the cruft objects include some really big sound or video files) one\n> > advantage of labelling the cruft packs explicitly is that someone who\n> > really needs the space could potentially find the oldest cruft files\n> > and delete them, since they would be tagged for easy findability.\n> \n> No! That's exactly what I was worried about with the name. It is _not_\n> safe to do so. It's only safe after you have done a full repack to\n> rescue any non-cruft objects.\n\nTo make it \"safe\", the cruft packs would have to be searchable for \nobject retrieval, but not during object creation.  That nuance would \naffect the core code in subtle ways and I'm not sure if that would be \nworth it ... just for the safe handling of cruft.\n\n\nNicolas\n"},{"id":"193464","messageId":"20120612171048.GB12706@sigill.intra.peff.net","threadId":"30758","inReplyTo":"alpine.LFD.2.02.1206112024110.23555@xanadu.home","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-06-12T17:10:48Z","receivedAt":"2012-06-12T17:10:48Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jun 11, 2012 at 08:41:03PM -0400, Nicolas Pitre wrote:\n\n> > On Mon, Jun 11, 2012 at 06:14:39PM -0400, Ted Ts'o wrote:\n> > \n> > > One last thought: if a sysadmin is really hard up for space, (and if\n> > > the cruft objects include some really big sound or video files) one\n> > > advantage of labelling the cruft packs explicitly is that someone who\n> > > really needs the space could potentially find the oldest cruft files\n> > > and delete them, since they would be tagged for easy findability.\n> > \n> > No! That's exactly what I was worried about with the name. It is _not_\n> > safe to do so. It's only safe after you have done a full repack to\n> > rescue any non-cruft objects.\n> \n> To make it \"safe\", the cruft packs would have to be searchable for \n> object retrieval, but not during object creation.  That nuance would \n> affect the core code in subtle ways and I'm not sure if that would be \n> worth it ... just for the safe handling of cruft.\n\nWhy is that? If you do a \"repack -Ad\", then any referenced objects will\nhave been retrieved and put into the new all-in-one pack. At that point,\nby deleting the cruft pack, you are guaranteed to be deleting only\nobjects that are either unreferenced, or are duplicated in another pack.\n\n-Peff\n"},{"id":"193465","messageId":"alpine.LFD.2.02.1206121326490.23555@xanadu.home","threadId":"30758","inReplyTo":"20120612171048.GB12706@sigill.intra.peff.net","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2012-06-12T17:30:07Z","receivedAt":"2012-06-12T17:30:07Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 12 Jun 2012, Jeff King wrote:\n\n> On Mon, Jun 11, 2012 at 08:41:03PM -0400, Nicolas Pitre wrote:\n> \n> > > On Mon, Jun 11, 2012 at 06:14:39PM -0400, Ted Ts'o wrote:\n> > > \n> > > > One last thought: if a sysadmin is really hard up for space, (and if\n> > > > the cruft objects include some really big sound or video files) one\n> > > > advantage of labelling the cruft packs explicitly is that someone who\n> > > > really needs the space could potentially find the oldest cruft files\n> > > > and delete them, since they would be tagged for easy findability.\n> > > \n> > > No! That's exactly what I was worried about with the name. It is _not_\n> > > safe to do so. It's only safe after you have done a full repack to\n> > > rescue any non-cruft objects.\n> > \n> > To make it \"safe\", the cruft packs would have to be searchable for \n> > object retrieval, but not during object creation.  That nuance would \n> > affect the core code in subtle ways and I'm not sure if that would be \n> > worth it ... just for the safe handling of cruft.\n> \n> Why is that? If you do a \"repack -Ad\", then any referenced objects will\n> have been retrieved and put into the new all-in-one pack. At that point,\n> by deleting the cruft pack, you are guaranteed to be deleting only\n> objects that are either unreferenced, or are duplicated in another pack.\n\nNow what if you fetch and a bunch of objects are already found in your \ncruft pack?  Right now, we search for the existence of any object before \ncreating them, and if the cruft packs are searchable then such objects \nwon't get uncruftified.\n\n\nNicolas\n"},{"id":"193466","messageId":"20120612173214.GA16014@sigill.intra.peff.net","threadId":"30758","inReplyTo":"alpine.LFD.2.02.1206121326490.23555@xanadu.home","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-06-12T17:32:14Z","receivedAt":"2012-06-12T17:32:14Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jun 12, 2012 at 01:30:07PM -0400, Nicolas Pitre wrote:\n\n> > > To make it \"safe\", the cruft packs would have to be searchable for \n> > > object retrieval, but not during object creation.  That nuance would \n> > > affect the core code in subtle ways and I'm not sure if that would be \n> > > worth it ... just for the safe handling of cruft.\n> > \n> > Why is that? If you do a \"repack -Ad\", then any referenced objects will\n> > have been retrieved and put into the new all-in-one pack. At that point,\n> > by deleting the cruft pack, you are guaranteed to be deleting only\n> > objects that are either unreferenced, or are duplicated in another pack.\n> \n> Now what if you fetch and a bunch of objects are already found in your \n> cruft pack?  Right now, we search for the existence of any object before \n> creating them, and if the cruft packs are searchable then such objects \n> won't get uncruftified.\n\nThen those objects will remain in the cruft pack. Which is why, as I\nsaid, it is not generally safe to just delete a cruft pack. However,\nwhen you do a full repack, those objects will be copied into the new\npack (because they are referenced). Which is why I am claiming that it\nis safe to remove cruft packs at that point.\n\n-Peff\n"},{"id":"193468","messageId":"CAJo=hJvMtfVhadYowvVE0zUhDpbViXqGsvkmHpJpuynySLwb3A@mail.gmail.com","threadId":"30758","inReplyTo":"20120612173214.GA16014@sigill.intra.peff.net","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2012-06-12T17:45:22Z","receivedAt":"2012-06-12T17:45:22Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Tue, Jun 12, 2012 at 10:32 AM, Jeff King <peff@peff.net> wrote:\n> On Tue, Jun 12, 2012 at 01:30:07PM -0400, Nicolas Pitre wrote:\n>\n>> > > To make it \"safe\", the cruft packs would have to be searchable for\n>> > > object retrieval, but not during object creation.  That nuance would\n>> > > affect the core code in subtle ways and I'm not sure if that would be\n>> > > worth it ... just for the safe handling of cruft.\n>> >\n>> > Why is that? If you do a \"repack -Ad\", then any referenced objects will\n>> > have been retrieved and put into the new all-in-one pack. At that point,\n>> > by deleting the cruft pack, you are guaranteed to be deleting only\n>> > objects that are either unreferenced, or are duplicated in another pack.\n>>\n>> Now what if you fetch and a bunch of objects are already found in your\n>> cruft pack?  Right now, we search for the existence of any object before\n>> creating them, and if the cruft packs are searchable then such objects\n>> won't get uncruftified.\n>\n> Then those objects will remain in the cruft pack. Which is why, as I\n> said, it is not generally safe to just delete a cruft pack. However,\n> when you do a full repack, those objects will be copied into the new\n> pack (because they are referenced). Which is why I am claiming that it\n> is safe to remove cruft packs at that point.\n\nBut there is a race condition with a concurrent fetch and a concurrent\nrepack. If that fetch needs those cruft objects, and sees them in the\ncruft pack, and the repack sees the references before the fetch, the\nrepacker might delete things the fetch is about to reference and that\nwill leave you with a corrupt repository.\n\nI think we already have this race condition with loose unreachable\nobjects whose mtimes are older than 2 weeks; they are removed by prune\nbut may have just become reachable by a concurrent fetch that doesn't\noverwrite them because they already exist, and doesn't update the\nmtime because they aren't writable.\n"},{"id":"193469","messageId":"alpine.LFD.2.02.1206121345500.23555@xanadu.home","threadId":"30758","inReplyTo":"20120612173214.GA16014@sigill.intra.peff.net","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2012-06-12T17:49:26Z","receivedAt":"2012-06-12T17:49:26Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 12 Jun 2012, Jeff King wrote:\n\n> On Tue, Jun 12, 2012 at 01:30:07PM -0400, Nicolas Pitre wrote:\n> \n> > > > To make it \"safe\", the cruft packs would have to be searchable for \n> > > > object retrieval, but not during object creation.  That nuance would \n> > > > affect the core code in subtle ways and I'm not sure if that would be \n> > > > worth it ... just for the safe handling of cruft.\n> > > \n> > > Why is that? If you do a \"repack -Ad\", then any referenced objects will\n> > > have been retrieved and put into the new all-in-one pack. At that point,\n> > > by deleting the cruft pack, you are guaranteed to be deleting only\n> > > objects that are either unreferenced, or are duplicated in another pack.\n> > \n> > Now what if you fetch and a bunch of objects are already found in your \n> > cruft pack?  Right now, we search for the existence of any object before \n> > creating them, and if the cruft packs are searchable then such objects \n> > won't get uncruftified.\n> \n> Then those objects will remain in the cruft pack. Which is why, as I\n> said, it is not generally safe to just delete a cruft pack.\n\n... and my reply was about the needed changes to still make cruft packs \nalways crufty even if some of its content suddenly becomes useful again.\n\n> However, when you do a full repack, those objects will be copied into \n> the new pack (because they are referenced). Which is why I am claiming \n> that it is safe to remove cruft packs at that point.\n\nYes, but then there is no point marking such packs as cruft if at any \nmoment they can become useful again.\n\n\nNicolas\n"},{"id":"193470","messageId":"20120612175046.GA16522@sigill.intra.peff.net","threadId":"30758","inReplyTo":"CAJo=hJvMtfVhadYowvVE0zUhDpbViXqGsvkmHpJpuynySLwb3A@mail.gmail.com","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-06-12T17:50:46Z","receivedAt":"2012-06-12T17:50:46Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jun 12, 2012 at 10:45:22AM -0700, Shawn O. Pearce wrote:\n\n> > Then those objects will remain in the cruft pack. Which is why, as I\n> > said, it is not generally safe to just delete a cruft pack. However,\n> > when you do a full repack, those objects will be copied into the new\n> > pack (because they are referenced). Which is why I am claiming that it\n> > is safe to remove cruft packs at that point.\n> \n> But there is a race condition with a concurrent fetch and a concurrent\n> repack. If that fetch needs those cruft objects, and sees them in the\n> cruft pack, and the repack sees the references before the fetch, the\n> repacker might delete things the fetch is about to reference and that\n> will leave you with a corrupt repository.\n> \n> I think we already have this race condition with loose unreachable\n> objects whose mtimes are older than 2 weeks; they are removed by prune\n> but may have just become reachable by a concurrent fetch that doesn't\n> overwrite them because they already exist, and doesn't update the\n> mtime because they aren't writable.\n\nCorrect. There is a race condition, but it is there already. I have\ndiscussed this with other GitHub folks, because we prune fairly\naggressively (in our case it would be a push, not a fetch, of course).\nSo far we have not had any record of it actually happening in practice.\n\nWe could close it in both cases by tweaking the mtime of the file\ncontaining the object when we decide not to write because the object\nalready exists.\n\n-Peff\n"},{"id":"193471","messageId":"20120612175438.GB16522@sigill.intra.peff.net","threadId":"30758","inReplyTo":"alpine.LFD.2.02.1206121345500.23555@xanadu.home","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-06-12T17:54:38Z","receivedAt":"2012-06-12T17:54:38Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jun 12, 2012 at 01:49:26PM -0400, Nicolas Pitre wrote:\n\n> > Then those objects will remain in the cruft pack. Which is why, as I\n> > said, it is not generally safe to just delete a cruft pack.\n> \n> ... and my reply was about the needed changes to still make cruft packs \n> always crufty even if some of its content suddenly becomes useful again.\n\nI think we are somehow missing each other's point, then. My point is\nthat you do not _need_ to make the cruft packs 100% cruft. You can\ntolerate the duplicated objects until they are pruned.\n\nEarlier in the thread, I outlined another scheme by which you could\nrepack and avoid the duplicates. It does not require changes to git's\nobject lookup process, because it would involve manually feeding the\nlist of cruft objects to pack-objects (which will pack what you ask it,\nregardless of whether the objects are in other packs).\n\n> > However, when you do a full repack, those objects will be copied into \n> > the new pack (because they are referenced). Which is why I am claiming \n> > that it is safe to remove cruft packs at that point.\n> \n> Yes, but then there is no point marking such packs as cruft if at any \n> moment they can become useful again.\n\nHow do you know to keep the packs around and expire them after 2 weeks\nif they are not marked in some way? Otherwise you would delete them as\npart of a \"git gc\", pushing the reachable objects into the new pack and\nthe unreachable objects into a new cruft pack. IOW, you need some way of\nkeeping the expiration date on the unreachable objects, or they will\nkeep getting \"refreshed\" by each gc.\n\n-Peff\n"},{"id":"193472","messageId":"alpine.LFD.2.02.1206121351470.23555@xanadu.home","threadId":"30758","inReplyTo":"CAJo=hJvMtfVhadYowvVE0zUhDpbViXqGsvkmHpJpuynySLwb3A@mail.gmail.com","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2012-06-12T17:55:22Z","receivedAt":"2012-06-12T17:55:22Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 12 Jun 2012, Shawn Pearce wrote:\n\n> On Tue, Jun 12, 2012 at 10:32 AM, Jeff King <peff@peff.net> wrote:\n> > On Tue, Jun 12, 2012 at 01:30:07PM -0400, Nicolas Pitre wrote:\n> >\n> >> > > To make it \"safe\", the cruft packs would have to be searchable for\n> >> > > object retrieval, but not during object creation.  That nuance would\n> >> > > affect the core code in subtle ways and I'm not sure if that would be\n> >> > > worth it ... just for the safe handling of cruft.\n> >> >\n> >> > Why is that? If you do a \"repack -Ad\", then any referenced objects will\n> >> > have been retrieved and put into the new all-in-one pack. At that point,\n> >> > by deleting the cruft pack, you are guaranteed to be deleting only\n> >> > objects that are either unreferenced, or are duplicated in another pack.\n> >>\n> >> Now what if you fetch and a bunch of objects are already found in your\n> >> cruft pack?  Right now, we search for the existence of any object before\n> >> creating them, and if the cruft packs are searchable then such objects\n> >> won't get uncruftified.\n> >\n> > Then those objects will remain in the cruft pack. Which is why, as I\n> > said, it is not generally safe to just delete a cruft pack. However,\n> > when you do a full repack, those objects will be copied into the new\n> > pack (because they are referenced). Which is why I am claiming that it\n> > is safe to remove cruft packs at that point.\n> \n> But there is a race condition with a concurrent fetch and a concurrent\n> repack. If that fetch needs those cruft objects, and sees them in the\n> cruft pack, and the repack sees the references before the fetch, the\n> repacker might delete things the fetch is about to reference and that\n> will leave you with a corrupt repository.\n> \n> I think we already have this race condition with loose unreachable\n> objects whose mtimes are older than 2 weeks; they are removed by prune\n> but may have just become reachable by a concurrent fetch that doesn't\n> overwrite them because they already exist, and doesn't update the\n> mtime because they aren't writable.\n\nSplat!\n\n\nNicolas\n"},{"id":"193473","messageId":"alpine.LFD.2.02.1206121356480.23555@xanadu.home","threadId":"30758","inReplyTo":"20120612175046.GA16522@sigill.intra.peff.net","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2012-06-12T17:57:53Z","receivedAt":"2012-06-12T17:57:53Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 12 Jun 2012, Jeff King wrote:\n\n> On Tue, Jun 12, 2012 at 10:45:22AM -0700, Shawn O. Pearce wrote:\n> \n> > > Then those objects will remain in the cruft pack. Which is why, as I\n> > > said, it is not generally safe to just delete a cruft pack. However,\n> > > when you do a full repack, those objects will be copied into the new\n> > > pack (because they are referenced). Which is why I am claiming that it\n> > > is safe to remove cruft packs at that point.\n> > \n> > But there is a race condition with a concurrent fetch and a concurrent\n> > repack. If that fetch needs those cruft objects, and sees them in the\n> > cruft pack, and the repack sees the references before the fetch, the\n> > repacker might delete things the fetch is about to reference and that\n> > will leave you with a corrupt repository.\n> > \n> > I think we already have this race condition with loose unreachable\n> > objects whose mtimes are older than 2 weeks; they are removed by prune\n> > but may have just become reachable by a concurrent fetch that doesn't\n> > overwrite them because they already exist, and doesn't update the\n> > mtime because they aren't writable.\n> \n> Correct. There is a race condition, but it is there already. I have\n> discussed this with other GitHub folks, because we prune fairly\n> aggressively (in our case it would be a push, not a fetch, of course).\n> So far we have not had any record of it actually happening in practice.\n> \n> We could close it in both cases by tweaking the mtime of the file\n> containing the object when we decide not to write because the object\n> already exists.\n\nYes, that is a worthwhile thing to do.\n\n\nNicolas\n"},{"id":"193478","messageId":"alpine.LFD.2.02.1206121359260.23555@xanadu.home","threadId":"30758","inReplyTo":"20120612175438.GB16522@sigill.intra.peff.net","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2012-06-12T18:25:47Z","receivedAt":"2012-06-12T18:25:47Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 12 Jun 2012, Jeff King wrote:\n\n> On Tue, Jun 12, 2012 at 01:49:26PM -0400, Nicolas Pitre wrote:\n> \n> > > Then those objects will remain in the cruft pack. Which is why, as I\n> > > said, it is not generally safe to just delete a cruft pack.\n> > \n> > ... and my reply was about the needed changes to still make cruft packs \n> > always crufty even if some of its content suddenly becomes useful again.\n> \n> I think we are somehow missing each other's point, then. My point is\n> that you do not _need_ to make the cruft packs 100% cruft. You can\n> tolerate the duplicated objects until they are pruned.\n\nAbsolutely.  Duplicated objectes are fine, and I was in fact suggesting \nto actively duplicate any needed object when it is to be found in a \ncruft pack only.\n\n> Earlier in the thread, I outlined another scheme by which you could\n> repack and avoid the duplicates. It does not require changes to git's\n> object lookup process, because it would involve manually feeding the\n> list of cruft objects to pack-objects (which will pack what you ask it,\n> regardless of whether the objects are in other packs).\n\nThat might be hard to achieve good delta compression though, as the main \nkey to sort those objects is their path name, and with unreferenced \nobjects you might not necessarily have that information.  The ability to \nreuse pack data might mitigate this though.\n\n> > > However, when you do a full repack, those objects will be copied into \n> > > the new pack (because they are referenced). Which is why I am claiming \n> > > that it is safe to remove cruft packs at that point.\n> > \n> > Yes, but then there is no point marking such packs as cruft if at any \n> > moment they can become useful again.\n> \n> How do you know to keep the packs around and expire them after 2 weeks\n> if they are not marked in some way? Otherwise you would delete them as\n> part of a \"git gc\", pushing the reachable objects into the new pack and\n> the unreachable objects into a new cruft pack. IOW, you need some way of\n> keeping the expiration date on the unreachable objects, or they will\n> keep getting \"refreshed\" by each gc.\n\nMy feeling is that we should make a step backward and consider if this \nis actually the right problem to solve.  I don't remember why I might \nhave been opposed to a reflog for deleted branches as you say I did, but \nthat is certainly a feature that could prove to be useful.\n\nThen having a repository that can be used as an alternate for other \nrepositories without knowing about it is also a problem that needs \nfixing and not only because of this object expiry issue.  This is not \neasy to fix though.\n\nThen, the creation of unreferenced objects from successive 'git add' \nshouldn't create that many objects in the first place.  They currently \nnever get the chance to be packed to start with.\n\nSo the problem is really about 'git gc' creating more data on disk which \nis counter productive for a garbage collecting task.  Maybe the trick is \nsimply not to delete any of the old pack which content was repacked into \na single new pack and let them age before deleting them, rather than \nexploding a bunch of loose objects.  But then we're back to the same \nissue I wanted to get away from i.e. identifying real cruft packs and \nmaking them safely deletable.\n\nOh well...\n\n\nNicolas\n"},{"id":"193479","messageId":"20120612183702.GD1803@thunk.org","threadId":"30758","inReplyTo":"alpine.LFD.2.02.1206121359260.23555@xanadu.home","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Ted Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2012-06-12T18:37:02Z","receivedAt":"2012-06-12T18:37:02Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Tue, Jun 12, 2012 at 02:25:47PM -0400, Nicolas Pitre wrote:\n> > Earlier in the thread, I outlined another scheme by which you could\n> > repack and avoid the duplicates. It does not require changes to git's\n> > object lookup process, because it would involve manually feeding the\n> > list of cruft objects to pack-objects (which will pack what you ask it,\n> > regardless of whether the objects are in other packs).\n> \n> That might be hard to achieve good delta compression though, as the main \n> key to sort those objects is their path name, and with unreferenced \n> objects you might not necessarily have that information.  The ability to \n> reuse pack data might mitigate this though.\n\nCompared to loose objects, even not-so-great delta compression is\nmanna from heaven.  Remember what originally got me to start this\nflag.  There was 4.5 megabytes worth of loose objects, that when I\ncreated the object id list and fed the result to git pack-object, the\nresulting pack was 244k.\n\nOK, maybe the delta compression wasn't optimal.  Compared to the 4.5\nmegabytes of loose objects --- I'll happily settle for that!  :-)\n\n> So the problem is really about 'git gc' creating more data on disk which \n> is counter productive for a garbage collecting task.  Maybe the trick is \n> simply not to delete any of the old pack which content was repacked into \n> a single new pack and let them age before deleting them, rather than \n> exploding a bunch of loose objects.  But then we're back to the same \n> issue I wanted to get away from i.e. identifying real cruft packs and \n> making them safely deletable.\n\nBut the old packs are huge; in my case, a full set of packs was around\n16 megabytes.  Right now, git gc *increased* my disk usage by 4.5\nmegabytes.  If we don't delete the old backs, then git gc would\nincrease disk usage by 16 megabytes --- which is far, far worse.\n\nWriting a 244k cruft pack is a soooooo much preferable.\n\n\t      \t       \t  \t    - Ted\n"},{"id":"193480","messageId":"m2fwa0fk0y.fsf@igel.home","threadId":"30758","inReplyTo":"20120612175046.GA16522@sigill.intra.peff.net","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Andreas Schwab","fromEmail":"schwab@linux-m68k.org","sentAt":"2012-06-12T18:43:41Z","receivedAt":"2012-06-12T18:43:41Z","isPatch":false,"sender":{"key":"schwab@linux-m68k.org","avatar":"https://avatars.githubusercontent.com/u/2175493?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> We could close it in both cases by tweaking the mtime of the file\n> containing the object when we decide not to write because the object\n> already exists.\n\nThough there is always the window between the existence check and the\nmtime update where pruning can hit you.\n\nAndreas.\n\n-- \nAndreas Schwab, schwab@linux-m68k.org\nGPG Key fingerprint = 58CA 54C7 6D53 942B 1756  01D3 44D5 214B 8276 4ED5\n\"And now for something completely different.\"\n"},{"id":"193483","messageId":"20120612190717.GA16911@sigill.intra.peff.net","threadId":"30758","inReplyTo":"m2fwa0fk0y.fsf@igel.home","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-06-12T19:07:17Z","receivedAt":"2012-06-12T19:07:17Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jun 12, 2012 at 08:43:41PM +0200, Andreas Schwab wrote:\n\n> Jeff King <peff@peff.net> writes:\n> \n> > We could close it in both cases by tweaking the mtime of the file\n> > containing the object when we decide not to write because the object\n> > already exists.\n> \n> Though there is always the window between the existence check and the\n> mtime update where pruning can hit you.\n\nFor the loose object case, you could do them both atomically by calling\nutime() on the object, and considering the object to exist only if it\nsucceeds.\n\nDoing it safely for packs would be harder, though; I think you'd\nhave to bump the mtime forward, do the search, and then bump it back.\nYou might err by causing a pack not to be pruned, but that is better\nthan the opposite.\n\nUnfortunately it gets trickier with network transfers. If somebody is\npushing to your repository, you might tell them you have some set of\nobjects, then they prepare a pack based on that assumption (which might\ntake minutes or hours to transfer), and then finally at the end you find\nthat you actually need the objects in question. Of course, that race is\neven harder to trigger, because we do not advertise unreachable objects.\nSo you would have to have a sequence where the objects are reachable,\nthe client connects and receives your ref advertisement, then the\nobjects become unreachable (e.g., due to a simultaneous non-ff push or\ndeletion), and you do a prune in that interval which removes the\nobjects.  Unlikely, but still possible.\n\n-Peff\n"},{"id":"193484","messageId":"alpine.LFD.2.02.1206121507120.23555@xanadu.home","threadId":"30758","inReplyTo":"m2fwa0fk0y.fsf@igel.home","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2012-06-12T19:09:25Z","receivedAt":"2012-06-12T19:09:25Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 12 Jun 2012, Andreas Schwab wrote:\n\n> Jeff King <peff@peff.net> writes:\n> \n> > We could close it in both cases by tweaking the mtime of the file\n> > containing the object when we decide not to write because the object\n> > already exists.\n> \n> Though there is always the window between the existence check and the\n> mtime update where pruning can hit you.\n\nThis is a tiny window compared to 2 weeks.\n\nThis could be avoided entirely with:\n\n1. check presence of object X\n\n2. update its mtime\n\n3. check presence of object X again\n\n4. create if doesn't exist\n\n\nNicolas\n"},{"id":"193485","messageId":"20120612191528.GB16911@sigill.intra.peff.net","threadId":"30758","inReplyTo":"alpine.LFD.2.02.1206121359260.23555@xanadu.home","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-06-12T19:15:28Z","receivedAt":"2012-06-12T19:15:28Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jun 12, 2012 at 02:25:47PM -0400, Nicolas Pitre wrote:\n\n> My feeling is that we should make a step backward and consider if this \n> is actually the right problem to solve.  I don't remember why I might \n> have been opposed to a reflog for deleted branches as you say I did, but \n> that is certainly a feature that could prove to be useful.\n\nI think your argument was along the lines of \"this information can be\nreconstructed from the HEAD reflog, anyway, so it is not worth the\neffort\". My counter to that is that the HEAD reflog is useless on bare\nrepositories (I have considered adding each pushed ref to a HEAD-like\nreflog with everything in it, but doing it without lock contention\nbetween pushes to different refs is tricky).\n\nBut keep in mind that a deletion reflog does not make this problem go\naway. It might make it less likely, but there are still cases where the\ngc can create a much larger object db.\n\n> Then having a repository that can be used as an alternate for other \n> repositories without knowing about it is also a problem that needs \n> fixing and not only because of this object expiry issue.  This is not \n> easy to fix though.\n\nYeah, I think that is an open problem, because you do not necessarily\nhave any write access at all to the alternates repository (however, that\ndoes not need to stop us from making it safer in the case that you _do_\nhave write access to the alternates repository).\n\n> Then, the creation of unreferenced objects from successive 'git add' \n> shouldn't create that many objects in the first place.  They currently \n> never get the chance to be packed to start with.\n\nI don't think these objects are necessarily from successive \"git add\"s.\nThat is one source, but they may also come from reflogs expiring. I\nguess in that case that they would typically be in an older pack,\nthough.\n\n> So the problem is really about 'git gc' creating more data on disk which \n> is counter productive for a garbage collecting task.  Maybe the trick is \n> simply not to delete any of the old pack which content was repacked into \n> a single new pack and let them age before deleting them, rather than \n> exploding a bunch of loose objects.  But then we're back to the same \n> issue I wanted to get away from i.e. identifying real cruft packs and \n> making them safely deletable.\n\nThat is satisfyingly simple, but the storage requirement is quite bad.\nThe unreachable objects are very much in the minority, and an occasional\nduplication there is not a big deal; duplicating all of the reachable\nobjects would double the object directory's size.\n\n-Peff\n"},{"id":"193486","messageId":"alpine.LFD.2.02.1206121509340.23555@xanadu.home","threadId":"30758","inReplyTo":"20120612183702.GD1803@thunk.org","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2012-06-12T19:15:46Z","receivedAt":"2012-06-12T19:15:46Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 12 Jun 2012, Ted Ts'o wrote:\n\n> On Tue, Jun 12, 2012 at 02:25:47PM -0400, Nicolas Pitre wrote:\n> > > Earlier in the thread, I outlined another scheme by which you could\n> > > repack and avoid the duplicates. It does not require changes to git's\n> > > object lookup process, because it would involve manually feeding the\n> > > list of cruft objects to pack-objects (which will pack what you ask it,\n> > > regardless of whether the objects are in other packs).\n> > \n> > That might be hard to achieve good delta compression though, as the main \n> > key to sort those objects is their path name, and with unreferenced \n> > objects you might not necessarily have that information.  The ability to \n> > reuse pack data might mitigate this though.\n> \n> Compared to loose objects, even not-so-great delta compression is\n> manna from heaven.  Remember what originally got me to start this\n> flag.  There was 4.5 megabytes worth of loose objects, that when I\n> created the object id list and fed the result to git pack-object, the\n> resulting pack was 244k.\n> \n> OK, maybe the delta compression wasn't optimal.  Compared to the 4.5\n> megabytes of loose objects --- I'll happily settle for that!  :-)\n\nSure.  However I would be even happier if we could delete those unneeded \nobjects outright.  The official reason why they're there for two weeks \nshould be to avoid some race conditions, and in this case two weeks is \nway over the top as in \"normal\" conditions the actual window for a race \nis in the order of a few seconds..  Any other use case should be \nconsidered abusive.\n\n> > So the problem is really about 'git gc' creating more data on disk which \n> > is counter productive for a garbage collecting task.  Maybe the trick is \n> > simply not to delete any of the old pack which content was repacked into \n> > a single new pack and let them age before deleting them, rather than \n> > exploding a bunch of loose objects.  But then we're back to the same \n> > issue I wanted to get away from i.e. identifying real cruft packs and \n> > making them safely deletable.\n> \n> But the old packs are huge; in my case, a full set of packs was around\n> 16 megabytes.  Right now, git gc *increased* my disk usage by 4.5\n> megabytes.  If we don't delete the old backs, then git gc would\n> increase disk usage by 16 megabytes --- which is far, far worse.\n> \n> Writing a 244k cruft pack is a soooooo much preferable.\n\nBut as you might have noticed, there are a bunch of semantic problems \nwith that as well.\n\n\nNicolas\n"},{"id":"193487","messageId":"20120612191929.GA12161@thunk.org","threadId":"30758","inReplyTo":"alpine.LFD.2.02.1206121509340.23555@xanadu.home","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Ted Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2012-06-12T19:19:29Z","receivedAt":"2012-06-12T19:19:29Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Tue, Jun 12, 2012 at 03:15:46PM -0400, Nicolas Pitre wrote:\n> > But the old packs are huge; in my case, a full set of packs was around\n> > 16 megabytes.  Right now, git gc *increased* my disk usage by 4.5\n> > megabytes.  If we don't delete the old backs, then git gc would\n> > increase disk usage by 16 megabytes --- which is far, far worse.\n> > \n> > Writing a 244k cruft pack is a soooooo much preferable.\n> \n> But as you might have noticed, there are a bunch of semantic problems \n> with that as well.\n\nI've proposed something (explicitly labelled cruft packs) which is no\nworse than before.  The one potential problem is that objects in the\ncruft pack might have their lifespan extended by two weeks (or\nwhatever the expire timeout might be), but Peff has agreed that it's\nsimple enough to ignore that, since the benefits far outweigh the\npotential that some objects in cruft packs will get to live a bit\nlonger.\n\nThe race condition you've pointed out exists today, with the git prune\nracing against the git fetch.\n\n\t\t\t\t\t- Ted\n"},{"id":"193488","messageId":"20120612192318.GC16911@sigill.intra.peff.net","threadId":"30758","inReplyTo":"alpine.LFD.2.02.1206121507120.23555@xanadu.home","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-06-12T19:23:18Z","receivedAt":"2012-06-12T19:23:18Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jun 12, 2012 at 03:09:25PM -0400, Nicolas Pitre wrote:\n\n> > Jeff King <peff@peff.net> writes:\n> > \n> > > We could close it in both cases by tweaking the mtime of the file\n> > > containing the object when we decide not to write because the object\n> > > already exists.\n> > \n> > Though there is always the window between the existence check and the\n> > mtime update where pruning can hit you.\n> \n> This is a tiny window compared to 2 weeks.\n\nI don't think the race window is actually 2 weeks long. If you have this\nsequence:\n\n  1. object X becomes unreferenced\n\n  2. 1 week later, you create a new ref that mentions X\n\n  3. >2 weeks later, you run \"git prune --expire=2.weeks.ago\"\n\nwe will not consider the object for pruning in step 3, because it is\nreachable. The race is more like:\n\n  1. object X becomes unreferenced\n\n  2. >2 weeks later, you run \"git prune --expire=2.weeks.ago\"\n\n  3. git-prune reads the list of refs\n\n  4. simultaneous to the git-prune, you reference X\n\n  5. git-prune removes X\n\n  6. your reference is now broken\n\nSo the race window depends on the time it takes \"git prune\" to run.\n\nI wonder if git-prune could do a double-check of the refs. Something\nlike:\n\n  1. calculate reachability on all refs\n\n  2. read list of objects to prune, and make a list of unreachable ones\n\n  3. calculate reachability again (which should be very cheap, because\n     you can stop when you get to an object you have already seen)\n\n  4. Drop any objects found in (3) from the list in (2), and delete\n     items from your list\n\nBut I think that still has a race where objects are created before\nstep 2, but are not actually referenced until after step 3. I think\ndoing it safely may actually require a repo-wide prune lock.\n\n-Peff\n"},{"id":"193490","messageId":"alpine.LFD.2.02.1206121533370.23555@xanadu.home","threadId":"30758","inReplyTo":"20120612191929.GA12161@thunk.org","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2012-06-12T19:35:17Z","receivedAt":"2012-06-12T19:35:17Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 12 Jun 2012, Ted Ts'o wrote:\n\n> On Tue, Jun 12, 2012 at 03:15:46PM -0400, Nicolas Pitre wrote:\n> > > But the old packs are huge; in my case, a full set of packs was around\n> > > 16 megabytes.  Right now, git gc *increased* my disk usage by 4.5\n> > > megabytes.  If we don't delete the old backs, then git gc would\n> > > increase disk usage by 16 megabytes --- which is far, far worse.\n> > > \n> > > Writing a 244k cruft pack is a soooooo much preferable.\n> > \n> > But as you might have noticed, there are a bunch of semantic problems \n> > with that as well.\n> \n> I've proposed something (explicitly labelled cruft packs) which is no\n> worse than before.  The one potential problem is that objects in the\n> cruft pack might have their lifespan extended by two weeks (or\n> whatever the expire timeout might be), but Peff has agreed that it's\n> simple enough to ignore that, since the benefits far outweigh the\n> potential that some objects in cruft packs will get to live a bit\n> longer.\n> \n> The race condition you've pointed out exists today, with the git prune\n> racing against the git fetch.\n\nYes, however that race is trivial to fix when loose objects are used.\n\n\nNicolas\n"},{"id":"193491","messageId":"alpine.LFD.2.02.1206121536570.23555@xanadu.home","threadId":"30758","inReplyTo":"20120612192318.GC16911@sigill.intra.peff.net","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2012-06-12T19:39:05Z","receivedAt":"2012-06-12T19:39:05Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 12 Jun 2012, Jeff King wrote:\n\n> So the race window depends on the time it takes \"git prune\" to run.\n> \n> I wonder if git-prune could do a double-check of the refs. Something\n> like:\n> \n>   1. calculate reachability on all refs\n> \n>   2. read list of objects to prune, and make a list of unreachable ones\n> \n>   3. calculate reachability again (which should be very cheap, because\n>      you can stop when you get to an object you have already seen)\n> \n>   4. Drop any objects found in (3) from the list in (2), and delete\n>      items from your list\n> \n> But I think that still has a race where objects are created before\n> step 2, but are not actually referenced until after step 3. I think\n> doing it safely may actually require a repo-wide prune lock.\n\nYeah... that's what I was thinking too.  Maybe we're making our life \noverly miserable by trying to avoid any locking here.\n\n\nNicolas\n"},{"id":"193492","messageId":"20120612194126.GA17519@sigill.intra.peff.net","threadId":"30758","inReplyTo":"alpine.LFD.2.02.1206121536570.23555@xanadu.home","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-06-12T19:41:26Z","receivedAt":"2012-06-12T19:41:26Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jun 12, 2012 at 03:39:05PM -0400, Nicolas Pitre wrote:\n\n> > So the race window depends on the time it takes \"git prune\" to run.\n> > \n> > I wonder if git-prune could do a double-check of the refs. Something\n> > like:\n> > \n> >   1. calculate reachability on all refs\n> > \n> >   2. read list of objects to prune, and make a list of unreachable ones\n> > \n> >   3. calculate reachability again (which should be very cheap, because\n> >      you can stop when you get to an object you have already seen)\n> > \n> >   4. Drop any objects found in (3) from the list in (2), and delete\n> >      items from your list\n> > \n> > But I think that still has a race where objects are created before\n> > step 2, but are not actually referenced until after step 3. I think\n> > doing it safely may actually require a repo-wide prune lock.\n> \n> Yeah... that's what I was thinking too.  Maybe we're making our life \n> overly miserable by trying to avoid any locking here.\n\nI think I would be OK with \"prune\" locking, as long as everything else\nwas able to happen simultaneously. Especially if we can keep prune's\nlock as short as possible through double-reads or similar tricks (like\nwe do for ref updates).\n\n-Peff\n"},{"id":"193494","messageId":"20120612194356.GD12161@thunk.org","threadId":"30758","inReplyTo":"alpine.LFD.2.02.1206121533370.23555@xanadu.home","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Ted Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2012-06-12T19:43:56Z","receivedAt":"2012-06-12T19:43:56Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Tue, Jun 12, 2012 at 03:35:17PM -0400, Nicolas Pitre wrote:\n> > \n> > The race condition you've pointed out exists today, with the git prune\n> > racing against the git fetch.\n> \n> Yes, however that race is trivial to fix when loose objects are used.\n\nWe can use the same trivial fix with the cruft pack, by touching the\nmtime of the pack.  Yes, it will extend the lifetime of the cruft pack\nby another 2 weeks but again, let me remind you: 244k versus 4.5\nmegabytes.  I can live with an extra 244k hanging around a wee bit\nlonger.  :-)\n\nThere are other fixes we could do involving flock() and removing the\ncruft label once we add a reference to a cruft pack, that I don't\nthink would be that complicated.\n\n\t\t\t\t\t\t- Ted\n"},{"id":"193564","messageId":"loom.20120613T185623-81@post.gmane.org","threadId":"30758","inReplyTo":"20120612191528.GB16911@sigill.intra.peff.net","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2012-06-13T18:17:46Z","receivedAt":"2012-06-13T18:17:46Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"Jeff King <peff <at> peff.net> writes:\n> > Then, the creation of unreferenced objects from successive 'git add' \n> > shouldn't create that many objects in the first place.  They currently \n> > never get the chance to be packed to start with.\n> \n> I don't think these objects are necessarily from successive \"git add\"s.\n> That is one source, but they may also come from reflogs expiring. I\n> guess in that case that they would typically be in an older pack,\n> though.\n...\n> That is satisfyingly simple, but the storage requirement is quite bad.\n> The unreachable objects are very much in the minority, and an \n> occasional duplication there is not a big deal; duplicating all of the \n> reachable objects would double the object directory's size.\n...\n(I don't think this is a valid generalization for servers)\n\nI am sorry to be coming a bit late into this discussion, but I think there\n is an even worse use case which can cause much worse loose object \nexplosions which does not seem to have been mentioned yet:   \"the \nserver upload rejected case\".  For example, think of a client pushing a \nchange from the wrong repository to a server.  Since there will be no \nhistory in common, the client will push the entire repository and if for\n some reason this gets rejected by the server (perhaps a pre-receive \nhook, or a gerrit server which says:  \"way too many new changes...\"), \nthen the pack file may stay abandonned on the server.  When gc runs: \nboom the entire history of that other project will explode but not get\n pruned since the pack file may be fairly new!\n\nI believe that this has happened to us several times fairly recently.  We\n have a tiny project which some people keep confusing for the kernel\nand they push a change destined for the kernel to it.  Gerrit rejects it and\ntheir massive packfile (larger than the entire project) stays around.  If gc \nruns, it almost becomes a DOS for us, the sheer number of loose object\nfiles makes the system crawl when accessing that repo, even on an SSD.\n We have been talking about moving to NFS soon (with packfiles git \nshould still perform fairly well on NFS), but this explosion really scares \nme.\n\nIt seems like the current design is a DOS just waiting to happen for\nservers.  While I would love to eliminate the races discussed in this\nthread, I think I agree with Ted in that the first fix should just focus on\nnever expanding loose objects for pruning (if certain objects simply don't \ndo well in pack files and the local gc policy says they should be loose, \ngo ahead: expand them, but that should be unrelated to pruning).  People\ncan DOS a server with unused packfiles too, but that rarely will have the\nsame impact that loose objects would have,\n\n-Martin\n\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a member \nof Code Aurora Forum\n"},{"id":"193594","messageId":"CALKQrgedkyA71Kr935Hdo6KGAg4c-1NbY4RB432zmuvSVU6mqw@mail.gmail.com","threadId":"30758","inReplyTo":"loom.20120613T185623-81@post.gmane.org","subject":"Re: Keeping unreachable objects in a separate pack instead of loose?","fromName":"Johan Herland","fromEmail":"johan@herland.net","sentAt":"2012-06-13T21:27:57Z","receivedAt":"2012-06-13T21:27:57Z","isPatch":false,"sender":{"key":"johan@herland.net","avatar":"https://avatars.githubusercontent.com/u/547031?v=4"},"body":"On Wed, Jun 13, 2012 at 8:17 PM, Martin Fick <mfick@codeaurora.org> wrote:\n> Jeff King <peff <at> peff.net> writes:\n>> > Then, the creation of unreferenced objects from successive 'git add'\n>> > shouldn't create that many objects in the first place.  They currently\n>> > never get the chance to be packed to start with.\n>>\n>> I don't think these objects are necessarily from successive \"git add\"s.\n>> That is one source, but they may also come from reflogs expiring. I\n>> guess in that case that they would typically be in an older pack,\n>> though.\n> ...\n>> That is satisfyingly simple, but the storage requirement is quite bad.\n>> The unreachable objects are very much in the minority, and an\n>> occasional duplication there is not a big deal; duplicating all of the\n>> reachable objects would double the object directory's size.\n> ...\n> (I don't think this is a valid generalization for servers)\n>\n> I am sorry to be coming a bit late into this discussion, but I think there\n>  is an even worse use case which can cause much worse loose object\n> explosions which does not seem to have been mentioned yet:   \"the\n> server upload rejected case\".  For example, think of a client pushing a\n> change from the wrong repository to a server.  Since there will be no\n> history in common, the client will push the entire repository and if for\n>  some reason this gets rejected by the server (perhaps a pre-receive\n> hook, or a gerrit server which says:  \"way too many new changes...\"),\n> then the pack file may stay abandonned on the server.  When gc runs:\n> boom the entire history of that other project will explode but not get\n>  pruned since the pack file may be fairly new!\n\n[...]\n\nJust a +1 from me. We had the same problem at my former $dayjob, and\nworked around it by running a \"git gc --prune=now\" in the server repo\n(which is a command you'd rather not want to run in a server repo) to\nremove exploded loose objects.\n\n\n...Johan\n\n-- \nJohan Herland, <johan@herland.net>\nwww.herland.net\n"}]}