{"thread":{"id":"18992","subject":"dangling commits and blobs: is this normal?","startedAt":"2009-04-21T21:46:16Z","lastAt":"2009-04-23T18:51:08Z","messageCount":21,"participants":["John Dlugosz","Jeff King","Brandon Casey","Nicolas Pitre","Matthieu Moy","Geert Bosch","Shawn O. Pearce","Matthias Andree"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"111908","messageId":"450196A1AAAE4B42A00A8B27A59278E70ACE0502@EXCHANGE.trad.tradestation.com","threadId":"18992","inReplyTo":null,"subject":"dangling commits and blobs: is this normal?","fromName":"John Dlugosz","fromEmail":"jdlugosz@tradestation.com","sentAt":"2009-04-21T21:46:16Z","receivedAt":"2009-04-21T21:46:16Z","isPatch":false,"sender":{"key":"jdlugosz@tradestation.com","avatar":null},"body":"Immediately after doing a git gc, a git fsck --full reports dangling\nobjects.  Is this normal?  What does dangling mean, if not those things\nthat gc finds?\n\n--John\n\nTradeStation Group, Inc. is a publicly-traded holding company (NASDAQ GS: TRAD) of three operating subsidiaries, TradeStation Securities, Inc. (Member NYSE, FINRA, SIPC and NFA), TradeStation Technologies, Inc., a trading software and subscription company, and TradeStation Europe Limited, a United Kingdom, FSA-authorized introducing brokerage firm. None of these companies provides trading or investment advice, recommendations or endorsements of any kind. The information transmitted is intended only for the person or entity to which it is addressed and may contain confidential and/or privileged material. Any review, retransmission, dissemination or other use of, or taking of any action in reliance upon, this information by persons or entities other than the intended recipient is prohibited.\n  If you received this in error, please contact the sender and delete the material from any computer.\n"},{"id":"111954","messageId":"20090422152719.GA12881@coredump.intra.peff.net","threadId":"18992","inReplyTo":"450196A1AAAE4B42A00A8B27A59278E70ACE0502@EXCHANGE.trad.tradestation.com","subject":"Re: dangling commits and blobs: is this normal?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-04-22T15:27:20Z","receivedAt":"2009-04-22T15:27:20Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Apr 21, 2009 at 05:46:16PM -0400, John Dlugosz wrote:\n\n> Immediately after doing a git gc, a git fsck --full reports dangling\n> objects.  Is this normal?  What does dangling mean, if not those things\n> that gc finds?\n\ngc will leave dangling loose objects for a set expiration time\n(defaulting to two weeks). This makes it safe to run even if there are\noperations in progress that want those dangling objects, but haven't yet\nadded a reference to them (as long as said operation takes less than two\nweeks).\n\nYou can also end up with dangling objects in packs. When that pack is\nrepacked, those objects will be loosened, and then eventually expired\nunder the rule mentioned above. However, I believe gc will not always\nrepack old packs; it will make new packs until you have a lot of packs,\nand then combine them all (at least that is what \"gc --auto\" will do; I\ndon't recall whether just \"git gc\" follows the same rule).\n\n-Peff\n"},{"id":"111965","messageId":"W0cjdA0pSHr_AbT2c-k5hDf7LyNvwkc38qIIhTtJJRwFnGBxaBsEiw@cipher.nrlssc.navy.mil","threadId":"18992","inReplyTo":"20090422152719.GA12881@coredump.intra.peff.net","subject":"Re: dangling commits and blobs: is this normal?","fromName":"Brandon Casey","fromEmail":"casey@nrlssc.navy.mil","sentAt":"2009-04-22T16:53:45Z","receivedAt":"2009-04-22T16:53:45Z","isPatch":false,"sender":{"key":"drafnel@gmail.com","avatar":"https://avatars.githubusercontent.com/u/921167?v=4"},"body":"Jeff King wrote:\n> On Tue, Apr 21, 2009 at 05:46:16PM -0400, John Dlugosz wrote:\n> \n>> Immediately after doing a git gc, a git fsck --full reports dangling\n>> objects.  Is this normal?  What does dangling mean, if not those things\n>> that gc finds?\n> \n> gc will leave dangling loose objects for a set expiration time\n> (defaulting to two weeks). This makes it safe to run even if there are\n> operations in progress that want those dangling objects, but haven't yet\n> added a reference to them (as long as said operation takes less than two\n> weeks).\n> \n> You can also end up with dangling objects in packs. When that pack is\n> repacked, those objects will be loosened, and then eventually expired\n> under the rule mentioned above. However, I believe gc will not always\n> repack old packs; it will make new packs until you have a lot of packs,\n> and then combine them all (at least that is what \"gc --auto\" will do; I\n> don't recall whether just \"git gc\" follows the same rule).\n\n'git gc' (without --auto) always creates one new pack.\n\nI've often wondered whether a plain 'git gc' should adopt the behavior\nof --auto with respect to the number of packs.  If there were few packs,\nthen 'git gc' would do an incremental repack, rather than a 'repack -A -d -l'.\n\nI'm still on the fence about it.  I think 'git gc' is supposed to be a\ndo-the-right-thing command, so in that sense I think it would be good\nbehavior and it would probably be what most less experienced users want.\nBut, 'git gc' is also used by experienced users who may expect the historical\nbehavior and may _want_ the \"pack into one pack\" behavior.  It could also be\nthat the more experienced users who want the \"pack into one pack\" behavior\nare actually the only users of 'git gc', and others just rely on the automatic\n'git gc --auto' calling.\n\nNot sure.\n\n-brandon\n"},{"id":"111969","messageId":"alpine.LFD.2.00.0904221331450.6741@xanadu.home","threadId":"18992","inReplyTo":"W0cjdA0pSHr_AbT2c-k5hDf7LyNvwkc38qIIhTtJJRwFnGBxaBsEiw@cipher.nrlssc.navy.mil","subject":"Re: dangling commits and blobs: is this normal?","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-22T17:39:21Z","receivedAt":"2009-04-22T17:39:21Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 22 Apr 2009, Brandon Casey wrote:\n\n> Jeff King wrote:\n> > On Tue, Apr 21, 2009 at 05:46:16PM -0400, John Dlugosz wrote:\n> > \n> >> Immediately after doing a git gc, a git fsck --full reports dangling\n> >> objects.  Is this normal?  What does dangling mean, if not those things\n> >> that gc finds?\n> > \n> > gc will leave dangling loose objects for a set expiration time\n> > (defaulting to two weeks). This makes it safe to run even if there are\n> > operations in progress that want those dangling objects, but haven't yet\n> > added a reference to them (as long as said operation takes less than two\n> > weeks).\n> > \n> > You can also end up with dangling objects in packs. When that pack is\n> > repacked, those objects will be loosened, and then eventually expired\n> > under the rule mentioned above. However, I believe gc will not always\n> > repack old packs; it will make new packs until you have a lot of packs,\n> > and then combine them all (at least that is what \"gc --auto\" will do; I\n> > don't recall whether just \"git gc\" follows the same rule).\n> \n> 'git gc' (without --auto) always creates one new pack.\n> \n> I've often wondered whether a plain 'git gc' should adopt the behavior\n> of --auto with respect to the number of packs.  If there were few packs,\n> then 'git gc' would do an incremental repack, rather than a 'repack -A -d -l'.\n\nWhy so?  Having fewer packs is always a good thing.  Having only one \npack is of course the optimal situation.  The --auto version doesn't do \nit in the hope of being lightter and less noticeable by the user.  \nHowever the user manually invoking gc should be expecting some work is \nactually happening.  If you don't want the whole repo read from one pack \njust to be written in another pack (say the repo is huge and waiting \nafter the IO is not worth it) then just mark such a pack with a .keep \nfile.\n\n\nNicolas\n"},{"id":"111975","messageId":"vpqws9cd06b.fsf@bauges.imag.fr","threadId":"18992","inReplyTo":"alpine.LFD.2.00.0904221331450.6741@xanadu.home","subject":"Re: dangling commits and blobs: is this normal?","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2009-04-22T18:15:56Z","receivedAt":"2009-04-22T18:15:56Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"Nicolas Pitre <nico@cam.org> writes:\n\n> Why so?  Having fewer packs is always a good thing.  Having only one \n> pack is of course the optimal situation. \n\nGood and optimal wrt Git, but not wrt an incremental backup system for\nexample. I have a \"git gc\" running daily in a cron job in each of my\nrepositories, but to be nice with my sysadmin, I don't want to rewrite\ntens of megabytes of data each night just because I commited a 2 lines\npatch somewhere.\n\n-- \nMatthieu\n"},{"id":"111979","messageId":"20090422190856.GB13424@coredump.intra.peff.net","threadId":"18992","inReplyTo":"vpqws9cd06b.fsf@bauges.imag.fr","subject":"Re: dangling commits and blobs: is this normal?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-04-22T19:08:56Z","receivedAt":"2009-04-22T19:08:56Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Apr 22, 2009 at 08:15:56PM +0200, Matthieu Moy wrote:\n\n> Nicolas Pitre <nico@cam.org> writes:\n> \n> > Why so?  Having fewer packs is always a good thing.  Having only one \n> > pack is of course the optimal situation. \n> \n> Good and optimal wrt Git, but not wrt an incremental backup system for\n> example. I have a \"git gc\" running daily in a cron job in each of my\n> repositories, but to be nice with my sysadmin, I don't want to rewrite\n> tens of megabytes of data each night just because I commited a 2 lines\n> patch somewhere.\n\nYou can mark your \"big\" pack with a .keep, then do your nightly gc as\nusual. You'll have a smaller pack being rewritten each night. When it\ngets big enough, drop the .keep, gc, and then .keep the new pack.\n\nYes, it's a bit more work for you, but having \"git gc\" optimize by\ndefault for git's performance seems to be the only sensible course.\nYour idea of what is \"big enough\" above is somewhat outside the realm of\ngit, so you have to pay the price to specify it by tweaking the\nkeep-files.\n\n-Peff\n"},{"id":"111981","messageId":"alpine.LFD.2.00.0904221509550.6741@xanadu.home","threadId":"18992","inReplyTo":"vpqws9cd06b.fsf@bauges.imag.fr","subject":"Re: dangling commits and blobs: is this normal?","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-22T19:14:18Z","receivedAt":"2009-04-22T19:14:18Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 22 Apr 2009, Matthieu Moy wrote:\n\n> Nicolas Pitre <nico@cam.org> writes:\n> \n> > Why so?  Having fewer packs is always a good thing.  Having only one \n> > pack is of course the optimal situation. \n> \n> Good and optimal wrt Git, but not wrt an incremental backup system for\n> example.\n\nThis goes without saying that git should optimize for its own usage by \ndefault, and not for a particular backup system.\n\n> I have a \"git gc\" running daily in a cron job in each of my\n> repositories, but to be nice with my sysadmin, I don't want to rewrite\n> tens of megabytes of data each night just because I commited a 2 lines\n> patch somewhere.\n\nJust add a .keep file along side your .pack file after repacking.\n\n\nNicolas\n"},{"id":"111986","messageId":"FcecxnoVg4H8G3MKjZgl2T6zCGDer4yYyScIgaweFTNgDCKG65Xiig@cipher.nrlssc.navy.mil","threadId":"18992","inReplyTo":"alpine.LFD.2.00.0904221331450.6741@xanadu.home","subject":"Re: dangling commits and blobs: is this normal?","fromName":"Brandon Casey","fromEmail":"casey@nrlssc.navy.mil","sentAt":"2009-04-22T19:26:26Z","receivedAt":"2009-04-22T19:26:26Z","isPatch":false,"sender":{"key":"drafnel@gmail.com","avatar":"https://avatars.githubusercontent.com/u/921167?v=4"},"body":"Nicolas Pitre wrote:\n> On Wed, 22 Apr 2009, Brandon Casey wrote:\n\n>> I've often wondered whether a plain 'git gc' should adopt the behavior\n>> of --auto with respect to the number of packs.  If there were few packs,\n>> then 'git gc' would do an incremental repack, rather than a 'repack -A -d -l'.\n> \n> Why so?  Having fewer packs is always a good thing.  Having only one \n> pack is of course the optimal situation.  The --auto version doesn't do \n> it in the hope of being lightter and less noticeable by the user.\n\nThe only reason for avoiding packing all packs into one would be speed in\nthis case also.  I recall reading complaints or surprise about gc\nrepacking all packs into one, so I'm only trying to think about how to\nmatch program behavior with user expectations.  gc does a lot already,\nand even Jeff wasn't sure what to expect from 'git gc' with respect to\npacks.  Possibly an acceptable trade off between speed and optimal packing\nwould be to adopt the --auto behavior for deciding when to use '-A' with\nrepack.\n\n> However the user manually invoking gc should be expecting some work is \n> actually happening.  If you don't want the whole repo read from one pack \n> just to be written in another pack (say the repo is huge and waiting \n> after the IO is not worth it) then just mark such a pack with a .keep \n> file.\n\nThat's true, but a user who knows about the .keep mechanism would also\nnot be afraid to run 'repack -d -l' (I'm ignoring the other operations\nof gc).\n\n-brandon\n"},{"id":"111990","messageId":"I5p8gPPuE_qW2RDhwiqxCWDuMtnuvvgtSkeTkxby6rlj_FKtpERaBA@cipher.nrlssc.navy.mil","threadId":"18992","inReplyTo":"20090422190856.GB13424@coredump.intra.peff.net","subject":"Re: dangling commits and blobs: is this normal?","fromName":"Brandon Casey","fromEmail":"casey@nrlssc.navy.mil","sentAt":"2009-04-22T19:45:29Z","receivedAt":"2009-04-22T19:45:29Z","isPatch":false,"sender":{"key":"drafnel@gmail.com","avatar":"https://avatars.githubusercontent.com/u/921167?v=4"},"body":"Jeff King wrote:\n> On Wed, Apr 22, 2009 at 08:15:56PM +0200, Matthieu Moy wrote:\n> \n>> Nicolas Pitre <nico@cam.org> writes:\n>>\n>>> Why so?  Having fewer packs is always a good thing.  Having only one \n>>> pack is of course the optimal situation. \n>> Good and optimal wrt Git, but not wrt an incremental backup system for\n>> example. I have a \"git gc\" running daily in a cron job in each of my\n>> repositories, but to be nice with my sysadmin, I don't want to rewrite\n>> tens of megabytes of data each night just because I commited a 2 lines\n>> patch somewhere.\n> \n> You can mark your \"big\" pack with a .keep, then do your nightly gc as\n> usual. You'll have a smaller pack being rewritten each night. When it\n> gets big enough, drop the .keep, gc, and then .keep the new pack.\n> \n> Yes, it's a bit more work for you, but having \"git gc\" optimize by\n> default for git's performance seems to be the only sensible course.\n> Your idea of what is \"big enough\" above is somewhat outside the realm of\n> git, so you have to pay the price to specify it by tweaking the\n> keep-files.\n\nBut isn't git-gc supposed to be the \"high-level\" command that just does\nthe right thing?  It doesn't seem to me to be outside the scope of this\ncommand to make a decision about trading off speed/io for optimal repo\nlayout.  In fact, it does do this already.  The default window, depth and\ncompression settings are chosen to be \"good enough\", not to produce the\nabsolute optimum repo.\n\nI'm just pointing out that everything is a trade off.  So I think saying\nsomething like \"gc must optimize for git's performance\" is not entirely\naccurate.  We make tradeoffs now.  Other tradeoffs may be helpful.\n\nAlso, don't interpret my comments as me being convinced that a change to\ngc should be made.  It's a trivial patch, but I'm not yet certain one\nway or the other.\n\n-brandon\n"},{"id":"111994","messageId":"20090422195854.GA14146@coredump.intra.peff.net","threadId":"18992","inReplyTo":"I5p8gPPuE_qW2RDhwiqxCWDuMtnuvvgtSkeTkxby6rlj_FKtpERaBA@cipher.nrlssc.navy.mil","subject":"Re: dangling commits and blobs: is this normal?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-04-22T19:58:54Z","receivedAt":"2009-04-22T19:58:54Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Apr 22, 2009 at 02:45:29PM -0500, Brandon Casey wrote:\n\n> > Yes, it's a bit more work for you, but having \"git gc\" optimize by\n> > default for git's performance seems to be the only sensible course.\n> > Your idea of what is \"big enough\" above is somewhat outside the realm of\n> > git, so you have to pay the price to specify it by tweaking the\n> > keep-files.\n> \n> But isn't git-gc supposed to be the \"high-level\" command that just does\n> the right thing?  It doesn't seem to me to be outside the scope of this\n> command to make a decision about trading off speed/io for optimal repo\n> layout.  In fact, it does do this already.  The default window, depth and\n> compression settings are chosen to be \"good enough\", not to produce the\n> absolute optimum repo.\n> \n> I'm just pointing out that everything is a trade off.  So I think saying\n> something like \"gc must optimize for git's performance\" is not entirely\n> accurate.  We make tradeoffs now.  Other tradeoffs may be helpful.\n\nSure, but my point was that git doesn't even know _how_ to make that\ntradeoff. It doesn't know what you consider a reasonable size of backup\nfor your incremental backups, how often you might want to rollover your\nkeep files, how often you expect to commit and how big the commits will\nbe, etc.\n\nSo it does the most reasonable thing, which is to optimize for git\nitself based on what it does know. If there is any improvement to be\nmade, it is probably to make a simpler way for the user to specify that\nexternal knowledge to git (because tweaking .keep files really is\nunnecessarily complex for Matthieu's scenario). And maybe that is just\nadding a config variable analagous to \"gc.autopacklimit\" to be used for\nregular gc, but that would default to 0 (i.e., default to the current\nbehavior of always repacking).\n\nBut I don't think it makes sense to change the default.\n\n-Peff\n"},{"id":"111995","messageId":"alpine.LFD.2.00.0904221548310.6741@xanadu.home","threadId":"18992","inReplyTo":"FcecxnoVg4H8G3MKjZgl2T6zCGDer4yYyScIgaweFTNgDCKG65Xiig@cipher.nrlssc.navy.mil","subject":"Re: dangling commits and blobs: is this normal?","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-22T20:00:06Z","receivedAt":"2009-04-22T20:00:06Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 22 Apr 2009, Brandon Casey wrote:\n\n> Nicolas Pitre wrote:\n> > On Wed, 22 Apr 2009, Brandon Casey wrote:\n> \n> >> I've often wondered whether a plain 'git gc' should adopt the behavior\n> >> of --auto with respect to the number of packs.  If there were few packs,\n> >> then 'git gc' would do an incremental repack, rather than a 'repack -A -d -l'.\n> > \n> > Why so?  Having fewer packs is always a good thing.  Having only one \n> > pack is of course the optimal situation.  The --auto version doesn't do \n> > it in the hope of being lightter and less noticeable by the user.\n> \n> The only reason for avoiding packing all packs into one would be speed in\n> this case also.  I recall reading complaints or surprise about gc\n> repacking all packs into one, so I'm only trying to think about how to\n> match program behavior with user expectations.\n\nIt's user's expectations that need adjusting then.  Making a single pack \nis indeed the job of an explicit gc invocation.\n\n> gc does a lot already, and even Jeff wasn't sure what to expect from \n> 'git gc' with respect to packs.  Possibly an acceptable trade off \n> between speed and optimal packing would be to adopt the --auto \n> behavior for deciding when to use '-A' with repack.\n\nAnd what would be the point of manually running 'git gc' then, given \nthat 'git gc --auto' is already invoked automatically after most commit \ncreating commands?\n\nI mean, if you consider explicit 'git gc' too long, then simply wait \nuntil you can spare the time, if at all.  This is not like a non gc'd \nrepository suddently becomes non functional.\n\nWRT trade offs, the current behavior is already a pretty good compromize \nbetween speed and optimal packing, the later implying -f to 'git \nrepack' which is far far slower.\n\n\nNicolas\n"},{"id":"111996","messageId":"20090422200502.GA14304@coredump.intra.peff.net","threadId":"18992","inReplyTo":"alpine.LFD.2.00.0904221548310.6741@xanadu.home","subject":"Re: dangling commits and blobs: is this normal?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-04-22T20:05:02Z","receivedAt":"2009-04-22T20:05:02Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Apr 22, 2009 at 04:00:06PM -0400, Nicolas Pitre wrote:\n\n> And what would be the point of manually running 'git gc' then, given \n> that 'git gc --auto' is already invoked automatically after most commit \n> creating commands?\n> \n> I mean, if you consider explicit 'git gc' too long, then simply wait \n> until you can spare the time, if at all.  This is not like a non gc'd \n> repository suddently becomes non functional.\n\nThe other tradeoff, mentioned by Matthieu, is not about speed, but about\nrollover of files on disk. I think he would be in favor of a less\noptimal pack setup if it meant rewriting the largest packfile less\nfrequently.\n\nHowever, it may be reasonable to suggest that he just not manually \"gc\"\nthen. If he is not generating enough commits to warrant an auto-gc, then\nhe is probably not losing much by having loose objects. And if he is,\nthen auto-gc is already taking care of it.\n\n-Peff\n"},{"id":"111997","messageId":"alpine.LFD.2.00.0904221601200.6741@xanadu.home","threadId":"18992","inReplyTo":"I5p8gPPuE_qW2RDhwiqxCWDuMtnuvvgtSkeTkxby6rlj_FKtpERaBA@cipher.nrlssc.navy.mil","subject":"Re: dangling commits and blobs: is this normal?","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-22T20:07:25Z","receivedAt":"2009-04-22T20:07:25Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 22 Apr 2009, Brandon Casey wrote:\n\n> But isn't git-gc supposed to be the \"high-level\" command that just does\n> the right thing?  It doesn't seem to me to be outside the scope of this\n> command to make a decision about trading off speed/io for optimal repo\n> layout.  In fact, it does do this already.  The default window, depth and\n> compression settings are chosen to be \"good enough\", not to produce the\n> absolute optimum repo.\n\nExact.\n\n> I'm just pointing out that everything is a trade off.  So I think saying\n> something like \"gc must optimize for git's performance\" is not entirely\n> accurate.  We make tradeoffs now.  Other tradeoffs may be helpful.\n\nGit makes tradeoffs for itself.  Trying to optimize by _default_ for \nsome random backup system, or any other environmental component not \ninvolved in git usage, is completely silly.\n\n> Also, don't interpret my comments as me being convinced that a change to\n> gc should be made.  It's a trivial patch, but I'm not yet certain one\n> way or the other.\n\nBe free to interpret my replies as me being certain of not doing such a \nchange.\n\n\nNicolas\n"},{"id":"111998","messageId":"alpine.LFD.2.00.0904221609250.6741@xanadu.home","threadId":"18992","inReplyTo":"20090422200502.GA14304@coredump.intra.peff.net","subject":"Re: dangling commits and blobs: is this normal?","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-22T20:11:44Z","receivedAt":"2009-04-22T20:11:44Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 22 Apr 2009, Jeff King wrote:\n\n> On Wed, Apr 22, 2009 at 04:00:06PM -0400, Nicolas Pitre wrote:\n> \n> > And what would be the point of manually running 'git gc' then, given \n> > that 'git gc --auto' is already invoked automatically after most commit \n> > creating commands?\n> > \n> > I mean, if you consider explicit 'git gc' too long, then simply wait \n> > until you can spare the time, if at all.  This is not like a non gc'd \n> > repository suddently becomes non functional.\n> \n> The other tradeoff, mentioned by Matthieu, is not about speed, but about\n> rollover of files on disk. I think he would be in favor of a less\n> optimal pack setup if it meant rewriting the largest packfile less\n> frequently.\n> \n> However, it may be reasonable to suggest that he just not manually \"gc\"\n> then. If he is not generating enough commits to warrant an auto-gc, then\n> he is probably not losing much by having loose objects. And if he is,\n> then auto-gc is already taking care of it.\n\nMy point exactly.\n\nAnd those people savvy enough to automate 'git gc' nightly should be \nable to cope with .keep files as well.\n\n\nNicolas\n"},{"id":"112000","messageId":"450196A1AAAE4B42A00A8B27A59278E70ACE07F3@EXCHANGE.trad.tradestation.com","threadId":"18992","inReplyTo":"20090422152719.GA12881@coredump.intra.peff.net","subject":"RE: dangling commits and blobs: is this normal?","fromName":"John Dlugosz","fromEmail":"jdlugosz@tradestation.com","sentAt":"2009-04-22T20:15:54Z","receivedAt":"2009-04-22T20:15:54Z","isPatch":false,"sender":{"key":"jdlugosz@tradestation.com","avatar":null},"body":"\n\n> -----Original Message-----\n> From: Jeff King [mailto:peff@peff.net]\n> Sent: Wednesday, April 22, 2009 10:27 AM\n> To: John Dlugosz\n> Cc: git@vger.kernel.org\n> Subject: Re: dangling commits and blobs: is this normal?\n> \n> \n> gc will leave dangling loose objects for a set expiration time\n> (defaulting to two weeks). This makes it safe to run even if there are\n> operations in progress that want those dangling objects, but haven't\n> yet\n> added a reference to them (as long as said operation takes less than\n> two\n> weeks).\n\n\nAh, very enlightening.  I see: it's not just reflog stuff (which gc should know are root entry points and not complain about), it really does leave uncollected garbage on purpose, in case something is in progress.\n\n--John\n\nTradeStation Group, Inc. is a publicly-traded holding company (NASDAQ GS: TRAD) of three operating subsidiaries, TradeStation Securities, Inc. (Member NYSE, FINRA, SIPC and NFA), TradeStation Technologies, Inc., a trading software and subscription company, and TradeStation Europe Limited, a United Kingdom, FSA-authorized introducing brokerage firm. None of these companies provides trading or investment advice, recommendations or endorsements of any kind. The information transmitted is intended only for the person or entity to which it is addressed and may contain confidential and/or privileged material. Any review, retransmission, dissemination or other use of, or taking of any action in reliance upon, this information by persons or entities other than the intended recipient is prohibited. If you received this in error, please contact the sender and delete the material from any computer.\n"},{"id":"112080","messageId":"vpq8wlr4mh0.fsf@bauges.imag.fr","threadId":"18992","inReplyTo":"20090422190856.GB13424@coredump.intra.peff.net","subject":"Re: dangling commits and blobs: is this normal?","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2009-04-23T11:51:23Z","receivedAt":"2009-04-23T11:51:23Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> On Wed, Apr 22, 2009 at 08:15:56PM +0200, Matthieu Moy wrote:\n>\n>> Nicolas Pitre <nico@cam.org> writes:\n>> \n>> > Why so?  Having fewer packs is always a good thing.  Having only one \n>> > pack is of course the optimal situation. \n>> \n>> Good and optimal wrt Git, but not wrt an incremental backup system for\n>> example. I have a \"git gc\" running daily in a cron job in each of my\n>> repositories, but to be nice with my sysadmin, I don't want to rewrite\n>> tens of megabytes of data each night just because I commited a 2 lines\n>> patch somewhere.\n>\n> You can mark your \"big\" pack with a .keep, then do your nightly gc as\n> usual. You'll have a smaller pack being rewritten each night. When it\n> gets big enough, drop the .keep, gc, and then .keep the new pack.\n\n(thanks, I wasn't aware of this .keep thing before reading this\nthread)\n\n> Yes, it's a bit more work for you, but having \"git gc\" optimize by\n> default for git's performance seems to be the only sensible course.\n\nSure. Sorry if my message read as \"git gc does the wrong thing\", I was\njust mentionning that it's not optimal with respect to everything.\n\n-- \nMatthieu\n"},{"id":"112098","messageId":"064C9132-2E72-4665-A44D-A2F4194DAC2B@adacore.com","threadId":"18992","inReplyTo":"20090422200502.GA14304@coredump.intra.peff.net","subject":"Re: dangling commits and blobs: is this normal?","fromName":"Geert Bosch","fromEmail":"bosch@adacore.com","sentAt":"2009-04-23T17:43:18Z","receivedAt":"2009-04-23T17:43:18Z","isPatch":false,"sender":{"key":"bosch@adacore.com","avatar":null},"body":"\nOn Apr 22, 2009, at 16:05, Jeff King wrote:\n> The other tradeoff, mentioned by Matthieu, is not about speed, but  \n> about\n> rollover of files on disk. I think he would be in favor of a less\n> optimal pack setup if it meant rewriting the largest packfile less\n> frequently.\n>\n> However, it may be reasonable to suggest that he just not manually  \n> \"gc\"\n> then. If he is not generating enough commits to warrant an auto-gc,  \n> then\n> he is probably not losing much by having loose objects. And if he is,\n> then auto-gc is already taking care of it.\n\nFor large repositories with lots of large files, git spends too much\ntime copying large packs for relatively little gain. This is obvious  \nwhen\nyou include a few dozen large objects in any repository.\nCurrently, there is no limit to the number of times this data may\nbe copied. In particular, the average amount of I/O needed for\nchanges of size X depends linearly on the size of the total repository.\nSo, the mere presence of a couple of large objects has an large  \ndistributed overhead.\n\nWouldn't it be better to have a maximum of N packs, named\npack_0 .. pack_(N - 1),  in the repository with each pack_i being\nbetween 2^i and 2^(i+1)-1 bytes large? We could even dispense\ncompletely with loose objects and instead have each git operation\ncreate a single new pack.\n\nThen the repacking rule simply becomes: if a new pack_i would\noverwrite one of the same name, both packs are merged into a new  \npack_(i+1).\n\nTo analyze performance, let's assume the worst case, where the\nsize of a pack is equal to the expanded size of all objects contained  \nin it\nand new packs only have unique objects. With these assumptions, an  \nobject\nresiding in pack_i can only be merged into a pack_j with j > i.\nSo, if any repository of size n has k objects, the maximum total I/O  \nrequired\nto create the repository (counting all operations in its history) is  \nO(n log k).\n\nThe current situation, the number of repacks required is linear in the  \nnumber of\nobjects, so the total work required is more like O(n k).\n\nWhile I understand that the above is a gross simplification, and actual\nperformance is dictated by packing efficiency and constant factors  \nrather\nthan asymptotic performance, I think the general idea of limiting the\nnumber of packs in the way described is useful and will lead to  \nsignificant\nspeedups, especially during large imports that currently require  \nfrequent\nrepacking of the entire repository.\n"},{"id":"112099","messageId":"20090423175612.GV23604@spearce.org","threadId":"18992","inReplyTo":"064C9132-2E72-4665-A44D-A2F4194DAC2B@adacore.com","subject":"Re: dangling commits and blobs: is this normal?","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2009-04-23T17:56:13Z","receivedAt":"2009-04-23T17:56:13Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Geert Bosch <bosch@adacore.com> wrote:\n> significant\n> speedups, especially during large imports that currently require  \n> frequent\n> repacking of the entire repository.\n\nLarge imports should be using fast-import, and then issue a single\nmassive `git repack -f --window=250 --depth=50` or some such repack\ncommand after the entire import is complete.\n\nIf your favorite import tool (*cough* git-svn *cough*) can't use\nfast-import, and you are importing a large enough repository that\nthis matters to you, use another importer that can use fast-import.\n\n-- \nShawn.\n"},{"id":"112102","messageId":"0F4B0DFB-A0A3-45FA-9BB7-801FFD2E8862@adacore.com","threadId":"18992","inReplyTo":"20090423175612.GV23604@spearce.org","subject":"Re: dangling commits and blobs: is this normal?","fromName":"Geert Bosch","fromEmail":"bosch@adacore.com","sentAt":"2009-04-23T18:10:56Z","receivedAt":"2009-04-23T18:10:56Z","isPatch":false,"sender":{"key":"bosch@adacore.com","avatar":null},"body":"\nOn Apr 23, 2009, at 13:56, Shawn O. Pearce wrote:\n\n> If your favorite import tool (*cough* git-svn *cough*) can't use\n> fast-import, and you are importing a large enough repository that\n> this matters to you, use another importer that can use fast-import.\n\nHow did you guess? :) You're right of course, except that I can't\nuse fast-import AFAIK. The issue is also more general, as the\nsame scenario of adding new objects and repacking occurs\noutside the context of git-svn.\n\n   -Geert\n"},{"id":"112104","messageId":"op.usuqfufi1e62zd@balu.cs.uni-paderborn.de","threadId":"18992","inReplyTo":"0F4B0DFB-A0A3-45FA-9BB7-801FFD2E8862@adacore.com","subject":"Re: dangling commits and blobs: is this normal?","fromName":"Matthias Andree","fromEmail":"matthias.andree@gmx.de","sentAt":"2009-04-23T18:17:44Z","receivedAt":"2009-04-23T18:17:44Z","isPatch":false,"sender":{"key":"matthias.andree@gmx.de","avatar":null},"body":"Am 23.04.2009, 20:10 Uhr, schrieb Geert Bosch <bosch@adacore.com>:\n\n>\n> On Apr 23, 2009, at 13:56, Shawn O. Pearce wrote:\n>\n>> If your favorite import tool (*cough* git-svn *cough*) can't use\n>> fast-import, and you are importing a large enough repository that\n>> this matters to you, use another importer that can use fast-import.\n>\n> How did you guess? :) You're right of course, except that I can't\n> use fast-import AFAIK. The issue is also more general, as the\n> same scenario of adding new objects and repacking occurs\n> outside the context of git-svn.\n\nIf you can't use fast-import for lack of access to the SVN repo, svnsync  \nmay help with that part. Or easier: ask the admin to upload a dump or  \nprovide one for download... :-)\n\n-- \nMatthias Andree\n"},{"id":"112110","messageId":"alpine.LFD.2.00.0904231438280.6741@xanadu.home","threadId":"18992","inReplyTo":"064C9132-2E72-4665-A44D-A2F4194DAC2B@adacore.com","subject":"Re: dangling commits and blobs: is this normal?","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-23T18:51:08Z","receivedAt":"2009-04-23T18:51:08Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 23 Apr 2009, Geert Bosch wrote:\n\n> \n> On Apr 22, 2009, at 16:05, Jeff King wrote:\n> > The other tradeoff, mentioned by Matthieu, is not about speed, but about\n> > rollover of files on disk. I think he would be in favor of a less\n> > optimal pack setup if it meant rewriting the largest packfile less\n> > frequently.\n> > \n> > However, it may be reasonable to suggest that he just not manually \"gc\"\n> > then. If he is not generating enough commits to warrant an auto-gc, then\n> > he is probably not losing much by having loose objects. And if he is,\n> > then auto-gc is already taking care of it.\n> \n> For large repositories with lots of large files, git spends too much\n> time copying large packs for relatively little gain. This is obvious when\n> you include a few dozen large objects in any repository.\n> Currently, there is no limit to the number of times this data may\n> be copied. In particular, the average amount of I/O needed for\n> changes of size X depends linearly on the size of the total repository.\n> So, the mere presence of a couple of large objects has an large distributed\n> overhead.\n\nYou can put a limit on the number of times this data is copied, and even \nset the limit to zero.  Just add a .keep file to your .pack file and \nthat data will remain in stone.  Any further repack will consider only \nthose newly added objects you may have.\n\n> Wouldn't it be better to have a maximum of N packs, named\n> pack_0 .. pack_(N - 1),  in the repository with each pack_i being\n> between 2^i and 2^(i+1)-1 bytes large? We could even dispense\n> completely with loose objects and instead have each git operation\n> create a single new pack.\n\nI suggested that already for large enough objects.  For small objects \nthis makes no sense as you may accumulate too many of them and each one \nwould need to be opened in order to find if it contains the desired \nobject whereas currently you need a simple directory lookup.\n\n> number of packs in the way described is useful and will lead to significant\n> speedups, especially during large imports that currently require frequent\n> repacking of the entire repository.\n\nOthers commented on that issue already.\n\n\nNicolas\n"}]}