{"thread":{"id":"9781","subject":"People unaware of the importance of \"git gc\"?","startedAt":"2007-09-05T07:09:27Z","lastAt":"2018-10-10T19:08:42Z","messageCount":97,"participants":["Linus Torvalds","Martin Langhoff","Junio C Hamano","Karl Hasselström","Pierre Habouzit","Tomash Brechko","David Kastrup","Johan Herland","Matthieu Moy","Steven Grimm","Wincent Colaiuta","Johan De Messemaeker","Govind Salinas","Carl Worth","J. Bruce Fields","Nix","Jing Xue","Brandon Casey","Nicolas Pitre","Mike Hommey","Alex Riesen","Jeff King","Carlos Rica","Shawn O. Pearce","Russ Dill","Andreas Ericsson","Johannes Schindelin","Johannes Sixt","Andy Parkins","Ævar Arnfjörð Bjarmason","Stefan Beller"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"52543","messageId":"alpine.LFD.0.999.0709042355030.19879@evo.linux-foundation.org","threadId":"9781","inReplyTo":null,"subject":"People unaware of the importance of \"git gc\"?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-09-05T07:09:27Z","receivedAt":"2007-09-05T07:09:27Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\nSo we had a git bof at linux.conf.eu yesterday, and I leart something \nnew: even people who have been using git for a long time apparently don't \nnecessarily realize the importance of repacking.\n\nJames Bottomley (the Linux SCSI maintainer) is an old-time BK user, and \nvery comfy using git. But when he was demonstrating things on his poor old \nlaptop, simple things like \"git branch\" literally took a long time, and \nJames didn't seem to realize that the fact that he had apparently never \never repacked his repository was a big deal.\n\nThe kernel archive is a 190MB pack for me fully repacked (I just checked - \nI had actually thought that it was somewhat larger than that), but because \nJames hadn't repacked, his .git directory was over a gigabyte in size, and \nhis laptop wasn't able to cache anything at all effectively as a result.\n\nRepacking it took over an hour, simply because everything was *so* \nunpacked, and James' kernel repository had something like 92 thousand \nloose objects, and several hundred packfiles. Simple operations that \nreally take much less than a second for me (\"git branch\" takes 0.022s on \nmy laptop, which has the same 512M that James had on his) took many many \nseconds as a result, and James seemed to think that this was all normal.\n\nAnd James didn't even want to repack, because it was so expensive (which \nhe knew - he claims to have never ever repacked at all, but maybe he had \nstarted it and just control-C'd it when it was really slow at some point).\n\nNow, it may be that James didn't realize how important the occasional \ngarbage collect is exactly *because* he is an old-timer and used BK long \nbefore he used git, and just continued using git simply as a BK \nreplacement, but it did make me wonder whether maybe this lack of \nrepacking awareness is fairly common. \n\nI've been against automatic repacking, but that was really based on what \nappears to be potentially a very wrong assumption, namely that people \nwould do the manual repack on their own. If it turns out that people don't \ndo it, maybe the right thing for git to do really is to at least notify \npeople when they have way too many pack-files and/or loose objects.\n\nI personally repack everything way more often than is necessary, and I had \nkind of assumed that people did it that way, but I was apparently wrong. \nComments?\n\n\t\tLinus\n"},{"id":"52544","messageId":"46a038f90709050021o5cdcc976xd099242aeb70643d@mail.gmail.com","threadId":"9781","inReplyTo":"alpine.LFD.0.999.0709042355030.19879@evo.linux-foundation.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2007-09-05T07:21:29Z","receivedAt":"2007-09-05T07:21:29Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 9/5/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> I personally repack everything way more often than is necessary, and I had\n> kind of assumed that people did it that way, but I was apparently wrong.\n> Comments?\n\n(resent with CC to git@)\n\nI never followed up on one of your suggestions back in the day -- that\nwe printed an informational msg along the lines of \"you have X loose\nobjects, it's about time to repack\" after some operations (fetch,\nmerge, commit). These days it's all C, so I'll pass the buck to people\nthat actually know how to do printf() ;-)\n\nAlso -- early users got everything exploded during clone, James is\nprobable one of them. It is the worst case scenario, really. Users of\na modern git will start off with a large packs, and accumulate little\npacks from pulls, so it's not as bad.\n\nIn fact, in James' case, it would have been way way way faster to\n\"steal\" the packs from git.kernel.org via http (or your laptop) and\n_then_ repack. He'd been sorted in a minute.\n\ncheers,\n\n\nmartin\n"},{"id":"52549","messageId":"20070905072628.GB4911@moonlight.home","threadId":"9781","inReplyTo":"7vsl5tk1r8.fsf@gitster.siamese.dyndns.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Tomash Brechko","fromEmail":"tomash.brechko@gmail.com","sentAt":"2007-09-05T07:26:28Z","receivedAt":"2007-09-05T07:26:28Z","isPatch":false,"sender":{"key":"tomash.brechko@gmail.com","avatar":null},"body":"Hi!\n\nOn Wed, Sep 05, 2007 at 00:30:35 -0700, Junio C Hamano wrote:\n> Perhaps _exiting_ \"git-commit\" and \"git-fetch\" before doing\n> anything, when the repository has more than 5000 loose objects\n> with a LOUD bang that instructs an immediate repack would be\n> good?\n\nThis may break automation.  I run git-gc monthly via cron, but that\ndoesn't guarantee I won't get 5000 loose objects before that.  And I\nagree that automatic run is annoying.  Perhaps simple BIG FAT WARNING\nis the best after all.\n\n\n-- \n   Tomash Brechko\n"},{"id":"52546","messageId":"7vsl5tk1r8.fsf@gitster.siamese.dyndns.org","threadId":"9781","inReplyTo":"alpine.LFD.0.999.0709042355030.19879@evo.linux-foundation.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-09-05T07:30:35Z","receivedAt":"2007-09-05T07:30:35Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> I personally repack everything way more often than is necessary, and I had \n> kind of assumed that people did it that way, but I was apparently wrong. \n> Comments?\n\nI am as old timer as you are so I am not qualified to add much\nvariety to the discussion, but I agree that excessive cruft is\nsomething we should warn the user about.\n\nI personally was _extremely_ annoyed by git-cvsimport\noccassionary deciding to repack whenever it finds more than\ncertain number of loose objects, not because it is a big import,\nbut because I happened to start the command to start a very\nsmall import after doing my own development for a while to\naccumulate loose objects, and I really hate automatic repacking\nfor any operation (or tool that thinks it knows better than I do\nin general).\n\nPerhaps _exiting_ \"git-commit\" and \"git-fetch\" before doing\nanything, when the repository has more than 5000 loose objects\nwith a LOUD bang that instructs an immediate repack would be\ngood?\n\nI really do not like the idea of automatically running a repack\nafter first interrupting the original command and then resuming.\nFor one thing it would make a horribly difficult situation to\ndebug if anything goes wrong.  You cannot reproduce such a\nsituation easily.\n"},{"id":"52547","messageId":"20070905073738.GA10570@diana.vm.bytemark.co.uk","threadId":"9781","inReplyTo":"46a038f90709050021o5cdcc976xd099242aeb70643d@mail.gmail.com","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Karl Hasselström","fromEmail":"kha@treskal.com","sentAt":"2007-09-05T07:37:38Z","receivedAt":"2007-09-05T07:37:38Z","isPatch":false,"sender":{"key":"kha@treskal.com","avatar":"https://gravatar.com/avatar/f0120c734b5279b345075a28521e1ac66acb20c9913ffe9bf6ae97e53f7f3f13?d=mp&s=160"},"body":"On 2007-09-05 19:21:29 +1200, Martin Langhoff wrote:\n\n> I never followed up on one of your suggestions back in the day --\n> that we printed an informational msg along the lines of \"you have X\n> loose objects, it's about time to repack\" after some operations\n> (fetch, merge, commit).\n\ngit-gui pops up a dialog that says precisely that, and gives you the\nchoice of repacking right then and there, or skip it.\n\nAs for truly automatic repacking after commands such as fetch, it\ncould probably be a config option (defaulting to \"on\"). It'd be\nimportant to have \"press any key to abort repacking (with no ill\neffects)\" type funtctionality, though.\n\n-- \nKarl Hasselström, kha@treskal.com\n      www.treskal.com/kalle\n"},{"id":"52548","messageId":"20070905074206.GA31750@artemis.corp","threadId":"9781","inReplyTo":"alpine.LFD.0.999.0709042355030.19879@evo.linux-foundation.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Pierre Habouzit","fromEmail":"madcoder@debian.org","sentAt":"2007-09-05T07:42:06Z","receivedAt":"2007-09-05T07:42:06Z","isPatch":false,"sender":{"key":"madcoder@debian.org","avatar":"https://avatars.githubusercontent.com/u/44708?v=4"},"body":"On Wed, Sep 05, 2007 at 07:09:27AM +0000, Linus Torvalds wrote:\n> I've been against automatic repacking, but that was really based on what \n> appears to be potentially a very wrong assumption, namely that people \n> would do the manual repack on their own. If it turns out that people don't \n> do it, maybe the right thing for git to do really is to at least notify \n> people when they have way too many pack-files and/or loose objects.\n\n  Well independently from the fact that one could suppose that users\nshould use gc on their own, the big nasty problem with repacking is that\nit's really slow. And I just can't imagine git that I use to commit\nblazingly fast, will then be unavailable for a very long time (repacks\non my projects -- that are not as big as the kernel but still -- usually\ntake more than 10 to 20 seconds each).\n\n> I personally repack everything way more often than is necessary, and I had \n> kind of assumed that people did it that way, but I was apparently wrong. \n> Comments?\n\n  I do, when I'm bored and that I can't get things done. you know, it\nhas become one of my many twitches when I have an empty tty in front of\nme and that I'm doing nothing useful. Though, when I'm in a hack-attack,\nwell I don't necessarily remember to repack. I'm in one of the (not so\nmany ?) very lucky companies (yay start-ups) where I could show that git\nwas very superior, and we now use it as our sole SCM. So when I'm in a\nhack attack, it's usually that it's a busy week, and that new patches,\ntrees, objects (and sometimes with large binary things in it) flows like\nhell. And the repository grows larger and larger. Well, the way we chose\nto avoid the \"I'm coding don't bother me with administrivia\"-attitude is\nthat our users use a small cron that basically runs git gc each day, and\nan aggressive repack (with a window of 50 or 100 I don't remember) each\nWeek-end in a cron. Because the best criterion to repack a repository\nis: when there is no-one on the computer.\n\n  It has proven quite good, as we have never seen a repository explode\nin a day, even after some funny mistakes where people rebase some big\nparts of the tree many times, generating very large number of loose\nobjets.\n\n\n  I know I don't really answer the question, but the point I try to make\nis that yeah, some kind of automated way to run the gc is great, but I'm\nnot sure that _git_ is the tool to automate that, because when *I* use\ngit, I expect it to be just plain fast, and I don't want it to\noccasionally hang.\n\n-- \n·O·  Pierre Habouzit\n··O                                                madcoder@debian.org\nOOO                                                http://www.madism.org\n"},{"id":"52556","messageId":"200709051013.39910.johan@herland.net","threadId":"9781","inReplyTo":"7vsl5tk1r8.fsf@gitster.siamese.dyndns.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Johan Herland","fromEmail":"johan@herland.net","sentAt":"2007-09-05T08:13:31Z","receivedAt":"2007-09-05T08:13:31Z","isPatch":false,"sender":{"key":"johan@herland.net","avatar":"https://avatars.githubusercontent.com/u/547031?v=4"},"body":"On Wednesday 05 September 2007, Junio C Hamano wrote:\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n> \n> > I personally repack everything way more often than is necessary, and I had \n> > kind of assumed that people did it that way, but I was apparently wrong. \n> > Comments?\n> \n> I am as old timer as you are so I am not qualified to add much\n> variety to the discussion, but I agree that excessive cruft is\n> something we should warn the user about.\n> \n> I personally was _extremely_ annoyed by git-cvsimport\n> occassionary deciding to repack whenever it finds more than\n> certain number of loose objects, not because it is a big import,\n> but because I happened to start the command to start a very\n> small import after doing my own development for a while to\n> accumulate loose objects, and I really hate automatic repacking\n> for any operation (or tool that thinks it knows better than I do\n> in general).\n> \n> Perhaps _exiting_ \"git-commit\" and \"git-fetch\" before doing\n> anything, when the repository has more than 5000 loose objects\n> with a LOUD bang that instructs an immediate repack would be\n> good?\n> \n> I really do not like the idea of automatically running a repack\n> after first interrupting the original command and then resuming.\n> For one thing it would make a horribly difficult situation to\n> debug if anything goes wrong.  You cannot reproduce such a\n> situation easily.\n\nWhat about some sort of middle ground:\n\nWhen git-fetch and git-commit has done its job and is about to exit, it checks \nthe number of loose object, and if too high tells the user something \nlike \"There are too many loose objects in the repo, do you want me to repack? \n(y/N)\". If the user answers \"n\" or simply <Enter>, it exits immediately \nwithout doing anything, but if the user answers \"y\", or if there is no \nresponse, say, within a minute (i.e. the user went to lunch), the repack is \ninitiated. (Of course, the user should be told that a Ctrl-C will abort the \nrepack and not be harmful in any way.)\n\nIf the user answers \"n\" (or aborts the repack), the question will keep popping \nup on the next git-{commit,fetch} to remind/annoy the user until a repack is \ndone.\n\n\n...Johan\n\n-- \nJohan Herland, <johan@herland.net>\nwww.herland.net\n"},{"id":"52554","messageId":"7vfy1tjzmu.fsf@gitster.siamese.dyndns.org","threadId":"9781","inReplyTo":"20070905074206.GA31750@artemis.corp","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-09-05T08:16:25Z","receivedAt":"2007-09-05T08:16:25Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Pierre Habouzit <madcoder@debian.org> writes:\n\n>   I do, when I'm bored and that I can't get things done. you know, it\n> has become one of my many twitches when I have an empty tty in front of\n> me and that I'm doing nothing useful.\n\nVery well said ;-)\n"},{"id":"52555","messageId":"861wddedce.fsf@lola.quinscape.zz","threadId":"9781","inReplyTo":"alpine.LFD.0.999.0709042355030.19879@evo.linux-foundation.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-05T08:16:49Z","receivedAt":"2007-09-05T08:16:49Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> Now, it may be that James didn't realize how important the\n> occasional garbage collect is exactly *because* he is an old-timer\n> and used BK long before he used git, and just continued using git\n> simply as a BK replacement, but it did make me wonder whether maybe\n> this lack of repacking awareness is fairly common.\n>\n> I've been against automatic repacking, but that was really based on\n> what appears to be potentially a very wrong assumption, namely that\n> people would do the manual repack on their own. If it turns out that\n> people don't do it, maybe the right thing for git to do really is to\n> at least notify people when they have way too many pack-files and/or\n> loose objects.\n>\n> I personally repack everything way more often than is necessary, and\n> I had kind of assumed that people did it that way, but I was\n> apparently wrong.  Comments?\n\nCan it be that getting rid of unused objects is harder once they are\npacked?  If that is the case, an automatic pack while mucking about\nwith temporary branches and/or confidential files would be quite a\nnuisance.\n\nAutomatic packing maybe would be acceptable if packing was really\ntransparent to what you do with your repo (including janitoring work).\nAnd it would be nice if automatic packing could be done in an\nincremental manner, not bogging down normal work.\n\n-- \nDavid Kastrup\n"},{"id":"52558","messageId":"vpqtzq91p5z.fsf@bauges.imag.fr","threadId":"9781","inReplyTo":"200709051013.39910.johan@herland.net","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2007-09-05T08:39:52Z","receivedAt":"2007-09-05T08:39:52Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"Johan Herland <johan@herland.net> writes:\n\n> When git-fetch and git-commit has done its job and is about to exit, it checks \n> the number of loose object, and if too high tells the user something \n> like \"There are too many loose objects in the repo, do you want me to repack? \n> (y/N)\". If the user answers \"n\" or simply <Enter>,\n\nI don't like commands to be interactive if they don't _need_ to be so.\nIt kills scripting, it makes it hard for a front-end (git gui or so)\nto use the command, ...\n\n-- \nMatthieu\n"},{"id":"52559","messageId":"200709051042.04345.johan@herland.net","threadId":"9781","inReplyTo":"vpqtzq91p5z.fsf@bauges.imag.fr","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Johan Herland","fromEmail":"johan@herland.net","sentAt":"2007-09-05T08:41:57Z","receivedAt":"2007-09-05T08:41:57Z","isPatch":false,"sender":{"key":"johan@herland.net","avatar":"https://avatars.githubusercontent.com/u/547031?v=4"},"body":"On Wednesday 05 September 2007, Matthieu Moy wrote:\n> Johan Herland <johan@herland.net> writes:\n> \n> > When git-fetch and git-commit has done its job and is about to exit, it checks \n> > the number of loose object, and if too high tells the user something \n> > like \"There are too many loose objects in the repo, do you want me to repack? \n> > (y/N)\". If the user answers \"n\" or simply <Enter>,\n> \n> I don't like commands to be interactive if they don't _need_ to be so.\n> It kills scripting, it makes it hard for a front-end (git gui or so)\n> to use the command, ...\n\nOk, so add an option or config variable to turn on/off this behaviour.\n\n...Johan\n\n-- \nJohan Herland, <johan@herland.net>\nwww.herland.net\n"},{"id":"52560","messageId":"86tzq9cxcb.fsf@lola.quinscape.zz","threadId":"9781","inReplyTo":"200709051042.04345.johan@herland.net","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-05T08:47:48Z","receivedAt":"2007-09-05T08:47:48Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Johan Herland <johan@herland.net> writes:\n\n> On Wednesday 05 September 2007, Matthieu Moy wrote:\n>> Johan Herland <johan@herland.net> writes:\n>> \n>> > When git-fetch and git-commit has done its job and is about to exit, it checks \n>> > the number of loose object, and if too high tells the user something \n>> > like \"There are too many loose objects in the repo, do you want me to repack? \n>> > (y/N)\". If the user answers \"n\" or simply <Enter>,\n>> \n>> I don't like commands to be interactive if they don't _need_ to be so.\n>> It kills scripting, it makes it hard for a front-end (git gui or so)\n>> to use the command, ...\n>\n> Ok, so add an option or config variable to turn on/off this behaviour.\n\nA bad idea which one can turn optionally off remains a bad idea for\neveryone that has not been bitten enough by it already to actually\nlook up the problem and remedy.\n\nMake this a warning.\n\n-- \nDavid Kastrup\n"},{"id":"52561","messageId":"46DE6DBC.30704@midwinter.com","threadId":"9781","inReplyTo":"20070905074206.GA31750@artemis.corp","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Steven Grimm","fromEmail":"koreth@midwinter.com","sentAt":"2007-09-05T08:50:04Z","receivedAt":"2007-09-05T08:50:04Z","isPatch":false,"sender":{"key":"koreth@midwinter.com","avatar":"https://gravatar.com/avatar/71b4d2e8b62f168bdc9e9205341159e3567003b4f9e2127c617c5fa0a1f5bad2?d=mp&s=160"},"body":"Pierre Habouzit wrote:\n>   Well independently from the fact that one could suppose that users\n> should use gc on their own, the big nasty problem with repacking is that\n> it's really slow. And I just can't imagine git that I use to commit\n> blazingly fast, will then be unavailable for a very long time (repacks\n> on my projects -- that are not as big as the kernel but still -- usually\n> take more than 10 to 20 seconds each).\n>   \n\nWhat about kicking off a repack in the background at the ends of certain \ncommands? With an option to disable, of course. It could run at a low \npriority and could even sleep a lot to avoid saturating the system's \ndisks -- since it'd be running asynchronously there should be no problem \nif it takes longer to run.\n\nAlternately, if it's possible to break the repack work up into chunks \nthat can be executed a bit at a time, you could do a small amount of \nrepacking very frequently (possibly still in the background) rather than \nthe whole thing at once. I suspect the nature of a repack, where you \npresumably want everything loaded at once, would make that a challenge, \nbut it might not be impossible.\n\nOn the more general question...\n\nIMO expecting end users to regularly perform what are essentially \ndatabase administration tasks (running git-gc is akin to rebuilding \nindexes or packing tables on a DBMS) is naive. Heck, even database \nadministrators don't like to run database administration commands; \nPostgreSQL added the \"autovacuum\" feature precisely because manual \nperiodic repacking (and the associated monitoring to figure out when to \ndo it) was too annoying for developers and DBAs. But you don't have to \nlook that far; anyone who has worked in IT can tell you horror stories \nof users, including developers, whose computers have slowed to a crawl \nbecause the users never bothered to defrag their hard disks. And that \naffects *everything* the users do, not just version control operations!\n\nIt'll get worse as better UIs and tool integration become available and \ngit gains large numbers of users who are neither software developers nor \nsystem administrators, and wouldn't know a packfile from a hole in the \nground. I'm talking web designers, graphic artists, mechanical \nengineers, even managers and secretaries -- all of those people are in \ngit's ultimate target audience, even if it's not ready for them today. \nNone of them is going to be interested in doing random housekeeping \noperations by hand, but they'll all appreciate a fast environment.\n\nThe fact that git sometimes stores your files individually in the .git \ndirectory and sometimes bundles them together into big archives should \nbe an implementation detail that end-users don't have to worry about day \nto day; git should do the right thing to remain fast under typical usage \nscenarios, while leaving the plumbing exposed so people with atypical \nusage can get their stuff done too.\n\n-Steve\n"},{"id":"52562","messageId":"65C61F04-EF05-4CB4-A4E2-BFF6601F46B9@wincent.com","threadId":"9781","inReplyTo":"7vsl5tk1r8.fsf@gitster.siamese.dyndns.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Wincent Colaiuta","fromEmail":"win@wincent.com","sentAt":"2007-09-05T08:51:09Z","receivedAt":"2007-09-05T08:51:09Z","isPatch":false,"sender":{"key":"greg@hurrell.net","avatar":"https://avatars.githubusercontent.com/u/7074?v=4"},"body":"El 5/9/2007, a las 9:30, Junio C Hamano escribió:\n\n> Perhaps _exiting_ \"git-commit\" and \"git-fetch\" before doing\n> anything, when the repository has more than 5000 loose objects\n> with a LOUD bang that instructs an immediate repack would be\n> good?\n>\n> I really do not like the idea of automatically running a repack\n> after first interrupting the original command and then resuming.\n> For one thing it would make a horribly difficult situation to\n> debug if anything goes wrong.  You cannot reproduce such a\n> situation easily.\n\nI would strongly oppose any *automatic* repacking and strongly  \nsupport any *advisory* recommandation to repack when the loose object  \ncount exceeds a certain threshold. I don't think *exiting* a command  \nin such cases is a good idea; worse than automatic repacking this  \nwould be *forced* manual repacking, which isn't very user-friendly.\n\nCheers,\nWincent\n"},{"id":"52563","messageId":"20070905085158.GC31750@artemis.corp","threadId":"9781","inReplyTo":"vpqtzq91p5z.fsf@bauges.imag.fr","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Pierre Habouzit","fromEmail":"madcoder@debian.org","sentAt":"2007-09-05T08:51:58Z","receivedAt":"2007-09-05T08:51:58Z","isPatch":false,"sender":{"key":"madcoder@debian.org","avatar":"https://avatars.githubusercontent.com/u/44708?v=4"},"body":"On Wed, Sep 05, 2007 at 08:39:52AM +0000, Matthieu Moy wrote:\n> Johan Herland <johan@herland.net> writes:\n> \n> > When git-fetch and git-commit has done its job and is about to exit, it checks \n> > the number of loose object, and if too high tells the user something \n> > like \"There are too many loose objects in the repo, do you want me to repack? \n> > (y/N)\". If the user answers \"n\" or simply <Enter>,\n> \n> I don't like commands to be interactive if they don't _need_ to be so.\n> It kills scripting, it makes it hard for a front-end (git gui or so)\n> to use the command, ...\n\n  There is absolutely no problem here, as it can be avoided if the\noutput is not a tty. It's not _that_ hard to guess if you're currently\nrunning in a script or in an interactive shell after all.\n\n  Really, git commit/fetch/... whatever suggesting to repack/gc when it\nbelieves it begins to be critical to performance is not a bad idea.\nThough the risk is that the warning could be printed very often, but\nthat can be avoided trivially by just writing to a state file in the\n.git directory that the warning was printed not so long time ago, and\nthat git should STFU for some more commits/time.\n\n-- \n·O·  Pierre Habouzit\n··O                                                madcoder@debian.org\nOOO                                                http://www.madism.org\n"},{"id":"52565","messageId":"86k5r5cwol.fsf@lola.quinscape.zz","threadId":"9781","inReplyTo":"20070905085158.GC31750@artemis.corp","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-05T09:02:02Z","receivedAt":"2007-09-05T09:02:02Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Pierre Habouzit <madcoder@debian.org> writes:\n\n> On Wed, Sep 05, 2007 at 08:39:52AM +0000, Matthieu Moy wrote:\n>> Johan Herland <johan@herland.net> writes:\n>> \n>> > When git-fetch and git-commit has done its job and is about to exit, it checks \n>> > the number of loose object, and if too high tells the user something \n>> > like \"There are too many loose objects in the repo, do you want me to repack? \n>> > (y/N)\". If the user answers \"n\" or simply <Enter>,\n>> \n>> I don't like commands to be interactive if they don't _need_ to be so.\n>> It kills scripting, it makes it hard for a front-end (git gui or so)\n>> to use the command, ...\n>\n>   There is absolutely no problem here, as it can be avoided if the\n> output is not a tty.\n\nWhich output?  stdout?  stderr?  Where is the question appearing?\nWhat if the command has been started in the background?  What if stdin\n(not stdout) is from a pipe, maybe for taking a commit message?  What\nif stdin is from a pseudo-tty because the commit has been started with\nan internal shell command inside of Emacs, and the command/message\nwill only get echoed once git-commit completes?\n\n> It's not _that_ hard to guess if you're currently running in a\n> script or in an interactive shell after all.\n\nOh, it is not hard to _guess_.  Just throw a die.  What is hard is to\n_know_ 100% sure that one is doing the right thing and not breaking\nany legitimate use.\n\n-- \nDavid Kastrup\n"},{"id":"52566","messageId":"vpqbqchxz3j.fsf@bauges.imag.fr","threadId":"9781","inReplyTo":"20070905085158.GC31750@artemis.corp","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2007-09-05T09:04:16Z","receivedAt":"2007-09-05T09:04:16Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"Pierre Habouzit <madcoder@debian.org> writes:\n\n> On Wed, Sep 05, 2007 at 08:39:52AM +0000, Matthieu Moy wrote:\n>> Johan Herland <johan@herland.net> writes:\n>> \n>> > When git-fetch and git-commit has done its job and is about to exit, it checks \n>> > the number of loose object, and if too high tells the user something \n>> > like \"There are too many loose objects in the repo, do you want me to repack? \n>> > (y/N)\". If the user answers \"n\" or simply <Enter>,\n>> \n>> I don't like commands to be interactive if they don't _need_ to be so.\n>> It kills scripting, it makes it hard for a front-end (git gui or so)\n>> to use the command, ...\n>\n>   There is absolutely no problem here, as it can be avoided if the\n> output is not a tty. It's not _that_ hard to guess if you're currently\n> running in a script or in an interactive shell after all.\n\nI do find it hard to guess _reliably_ if you're running interactively\nor not. For example, I've been bitten recently by \"git log\" running\ninside a pager while I was launching it non-interactively inside Emacs\n(as part of DVC). I don't know whether this was git's or Emacs's\nfault, and the fix was not too hard (GIT_PAGER=cat), but it took some\nof my time to get it working.\n\nAdding more interactive stuff means adding more opportunities for this\nkind of problems. None will be a huge problem, but each problem will\ntake some time to be fixed (I'm pretty sure adding an interactive\nprompt in git-commit will break DVC's commit functionality, and we'll\nhave to fix it).\n\n>   Really, git commit/fetch/... whatever suggesting to repack/gc when it\n> believes it begins to be critical to performance is not a bad idea.\n\n_Suggesting_ is a good idea, definitely. Something like\n\nif (number_of_unpacked > 1000 && number_of_unpacked < 10000) {\n\tprintf (\"more than 1000 unpacked objects. Think of running git-gc\\n\");\n} else if (number_of_unpacked >= 10000) {\n\tprintf (\"HEY, WHAT THE HELL ARE YOU DOING WITH >10000 UNPACKED OBJECTS???\\n\"\n                \"I TOLD YOU TO REPACK\\n\");\n}\n\nwould be fine with me. The proposal to run git-gc in the background,\nwith low priority seems to be a good idea too.\n\nBut please, don't put an interactive prompt where it's not needed.\n\n-- \nMatthieu\n"},{"id":"52567","messageId":"46DE71BC.5040008@midwinter.com","threadId":"9781","inReplyTo":"86ps0xcwxo.fsf@lola.quinscape.zz","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Steven Grimm","fromEmail":"koreth@midwinter.com","sentAt":"2007-09-05T09:07:08Z","receivedAt":"2007-09-05T09:07:08Z","isPatch":false,"sender":{"key":"koreth@midwinter.com","avatar":"https://gravatar.com/avatar/71b4d2e8b62f168bdc9e9205341159e3567003b4f9e2127c617c5fa0a1f5bad2?d=mp&s=160"},"body":"David Kastrup wrote:\n> You'll potentially get accumulating unfinished files from\n> aborted/killed repack processes.\n\nWhich can get cleaned up when the next repack starts. This is no \ndifferent from unfinished files accumulating from aborted/killed manual \nrepacks.\n\n> If communication fails, you'll get a\n> new repack session for every command you start.\n\nGit handles this already:\n\n$ git-gc\nfatal: unable to create '.git/packed-refs.lock': File exists\nerror: failed to run pack-refs\n\nPresumably in that case you would simply not fire up a new repack.\n\n>   If a repository is used by multiple people...\n>   \n\nThen the first one will kick off the repack, and subsequent ones won't.\n\n> And so on.  The multiuser aspect makes it a bad idea to do any\n> janitorial tasks automatically.  You don't really want every user to\n> start a repack at the same time.\n>   \n\nQuite true, but that's already impossible, so not a problem.\n\nOne other thing: The heuristics for this can be such that users who are \nalready regularly running git-gc by hand will see no change in behavior. \nTheir repos will never get to a bad enough state that the automatic \ngit-gc is invoked. Old-timers who run git-gc might, in theory, never \neven notice a change like this.\n\n-Steve\n"},{"id":"52568","messageId":"7vbqchjx9f.fsf@gitster.siamese.dyndns.org","threadId":"9781","inReplyTo":"46DE6DBC.30704@midwinter.com","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-09-05T09:07:40Z","receivedAt":"2007-09-05T09:07:40Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Steven Grimm <koreth@midwinter.com> writes:\n\n> The fact that git sometimes stores your files individually in the .git\n> directory and sometimes bundles them together into big archives should\n> be an implementation detail that end-users don't have to worry about\n> day to day...\n\n[alias]\n        begin = gc\n\tleave = gc\n\nThat is, the user's manual says 'at the beginning of the day,\nrun \"git begin\" to start the day, and at the end of day, run\n\"git leave\" to conclude your day', without saying why ;-)\n"},{"id":"52569","messageId":"86abs1cw5g.fsf@lola.quinscape.zz","threadId":"9781","inReplyTo":"46DE71BC.5040008@midwinter.com","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-05T09:13:31Z","receivedAt":"2007-09-05T09:13:31Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Steven Grimm <koreth@midwinter.com> writes:\n\n> David Kastrup wrote:\n>> You'll potentially get accumulating unfinished files from\n>> aborted/killed repack processes.\n>\n> Which can get cleaned up when the next repack starts. This is no\n> different from unfinished files accumulating from aborted/killed\n> manual repacks.\n>\n>> If communication fails, you'll get a\n>> new repack session for every command you start.\n>\n> Git handles this already:\n>\n> $ git-gc\n> fatal: unable to create '.git/packed-refs.lock': File exists\n> error: failed to run pack-refs\n>\n> Presumably in that case you would simply not fire up a new repack.\n>\n>>   If a repository is used by multiple people...\n>>   \n>\n> Then the first one will kick off the repack, and subsequent ones won't.\n\nAnd the first one might get habitually killed by the user unwittingly\nhaving started it (because he really only logs in for shorter amounts\nof times than needed for git-gc to finish), wasting disk space and\ntime all the while.\n\n-- \nDavid Kastrup\n"},{"id":"52571","messageId":"864pi9cw4w.fsf@lola.quinscape.zz","threadId":"9781","inReplyTo":"46DE6DBC.30704@midwinter.com","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-05T09:13:51Z","receivedAt":"2007-09-05T09:13:51Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Steven Grimm <koreth@midwinter.com> writes:\n\n> Pierre Habouzit wrote:\n>>   Well independently from the fact that one could suppose that users\n>> should use gc on their own, the big nasty problem with repacking is that\n>> it's really slow. And I just can't imagine git that I use to commit\n>> blazingly fast, will then be unavailable for a very long time (repacks\n>> on my projects -- that are not as big as the kernel but still -- usually\n>> take more than 10 to 20 seconds each).\n>>   \n>\n> What about kicking off a repack in the background at the ends of\n> certain commands? With an option to disable, of course. It could run\n> at a low priority and could even sleep a lot to avoid saturating the\n> system's disks -- since it'd be running asynchronously there should\n> be no problem if it takes longer to run.\n\nYou'll potentially get accumulating unfinished files from\naborted/killed repack processes.  If communication fails, you'll get a\nnew repack session for every command you start.  If a repository is\nused by multiple people...\n\nAnd so on.  The multiuser aspect makes it a bad idea to do any\njanitorial tasks automatically.  You don't really want every user to\nstart a repack at the same time.\n\n-- \nDavid Kastrup\n"},{"id":"52570","messageId":"20070905091400.GE31750@artemis.corp","threadId":"9781","inReplyTo":"46DE6DBC.30704@midwinter.com","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Pierre Habouzit","fromEmail":"madcoder@debian.org","sentAt":"2007-09-05T09:14:00Z","receivedAt":"2007-09-05T09:14:00Z","isPatch":false,"sender":{"key":"madcoder@debian.org","avatar":"https://avatars.githubusercontent.com/u/44708?v=4"},"body":"On Wed, Sep 05, 2007 at 08:50:04AM +0000, Steven Grimm wrote:\n> Pierre Habouzit wrote:\n> >  Well independently from the fact that one could suppose that users\n> >should use gc on their own, the big nasty problem with repacking is that\n> >it's really slow. And I just can't imagine git that I use to commit\n> >blazingly fast, will then be unavailable for a very long time (repacks\n> >on my projects -- that are not as big as the kernel but still -- usually\n> >take more than 10 to 20 seconds each).\n> >  \n> \n> What about kicking off a repack in the background at the ends of certain \n> commands? With an option to disable, of course. It could run at a low \n> priority and could even sleep a lot to avoid saturating the system's \n> disks -- since it'd be running asynchronously there should be no problem \n> if it takes longer to run.\n\n  there is an issue with that: repack is memory and CPU intensive. Of\ncourse renicing the process deals with the CPU issue, but not with the\nmemory one. I've often seen repacks eat more than 300 to 400Mo of memory\non not so big repositories: it seems (and experience tells me that, not\nlooking at the code) that if you have some big binary blobs (we have\n.swf's and .fla's in our repository) it can consume quite a lot of RAM\nto (presumably) compute efficient deltas.\n\n  Sadly there is no way to \"renice\" the ram usage of a process. Once a\nrepack is launched, it will make your system swap, and put the whole\ncomputer on its knees.\n\n> IMO expecting end users to regularly perform what are essentially \n> database administration tasks (running git-gc is akin to rebuilding \n> indexes or packing tables on a DBMS) is naive. Heck, even database \n> administrators don't like to run database administration commands; \n\n  Well that's what crons are for. When you install a SGBD in a\nreasonable enough distro, it comes with the optimizing scripts in crons,\nlaunched at a reasonable period of the day (localtime). So the\ncomparison doesn't hold. And that's exactly the problem: it's quite hard\nto ship git with an optimizing cron task, because we can't know where\nthe user will keep his repositories, and when he works, so you have\nsomehow to do it yourself.\n\n  Or you can deal with that with a \"rule\". At work, we have our devel\ntrees under $HOME/dev/, so the cron we use is just a (roughly):\n\n    find $HOME/dev/ -name .git -type d -maxdepth 4 | while read repo\n    do\n        GIT_DIR=\"$repo\" git gc\n    done\n\n  As we work on NFS, with a new developper, we can just setup the cron\nfor him at a date where he's not supposed to be at work, and that's it.\nI'm not sure there is a good solution at all.\n\n  Or we could also provide a: git-coffee-break command that would tell\ngit: do whatever you want with this computer in the next 10 minutes,\nthere won't be anyone watching, but I assume tea-lovers will feel\nexcluded.\n\n-- \n·O·  Pierre Habouzit\n··O                                                madcoder@debian.org\nOOO                                                http://www.madism.org\n"},{"id":"52572","messageId":"46a038f90709050227u777ed7b9w23dc3bab13c7b09b@mail.gmail.com","threadId":"9781","inReplyTo":"7vbqchjx9f.fsf@gitster.siamese.dyndns.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2007-09-05T09:27:23Z","receivedAt":"2007-09-05T09:27:23Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 9/5/07, Junio C Hamano <gitster@pobox.com> wrote:\n> [alias]\n>         begin = gc\n>         leave = gc\n>\n> That is, the user's manual says 'at the beginning of the day,\n> run \"git begin\" to start the day, and at the end of day, run\n> \"git leave\" to conclude your day', without saying why ;-)\n\nI actually like that one ;-)\n\nAnyway - this is turning out to be a bit of a bikeshed-painting event.\nYou guys should google earlier discussions on this very same subject.\nThey have always ended in \"automatic=bad\", \"warning=good\", and\n\"careful or you might be called an idiot\" before ;-)\n\ncheers,\n\n\nmartin\n"},{"id":"52573","messageId":"vpqzm01v4li.fsf@bauges.imag.fr","threadId":"9781","inReplyTo":"46a038f90709050227u777ed7b9w23dc3bab13c7b09b@mail.gmail.com","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2007-09-05T09:33:45Z","receivedAt":"2007-09-05T09:33:45Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"\"Martin Langhoff\" <martin.langhoff@gmail.com> writes:\n\n> On 9/5/07, Junio C Hamano <gitster@pobox.com> wrote:\n>> [alias]\n>>         begin = gc\n>>         leave = gc\n>>\n>> That is, the user's manual says 'at the beginning of the day,\n>> run \"git begin\" to start the day, and at the end of day, run\n>> \"git leave\" to conclude your day', without saying why ;-)\n>\n> I actually like that one ;-)\n\nThere's indeed a real idea behind that. The issue is that the alias\nshouldn't be just \"gc\", but \"find-all-repositories-and-do-gc-there\".\n\nCurrently, AFAIK, that can only be done with a (trivial) script\nexternal to git. I suppose this can easily be added to the core git\nporcelain. Perhaps a \"git gc --recursive\" would do.\n\nIt doesn't solve the problem, but makes it easier to solve it (git gc\n--recursive in cron for example).\n\n-- \nMatthieu\n"},{"id":"52583","messageId":"D32A7C27-EAF5-4156-BE0E-99FE3D948AE8@wgaf.org","threadId":"9781","inReplyTo":"vpqzm01v4li.fsf@bauges.imag.fr","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Johan De Messemaeker","fromEmail":"johan.demessemaeker@wgaf.org","sentAt":"2007-09-05T14:17:01Z","receivedAt":"2007-09-05T14:17:01Z","isPatch":false,"sender":{"key":"johan.demessemaeker@wgaf.org","avatar":null},"body":"\nOn 05 Sep 2007, at 11:33, Matthieu Moy wrote:\n>\n> There's indeed a real idea behind that. The issue is that the alias\n> shouldn't be just \"gc\", but \"find-all-repositories-and-do-gc-there\".\n>\n> Currently, AFAIK, that can only be done with a (trivial) script\n> external to git. I suppose this can easily be added to the core git\n> porcelain. Perhaps a \"git gc --recursive\" would do.\n>\n> It doesn't solve the problem, but makes it easier to solve it (git gc\n> --recursive in cron for example).\n\nI'm a git newb so I can be wrong here but ...\n\nWhy --recursive? Why not use the submodule-information ?\n\nJohan\n"},{"id":"52591","messageId":"69b0c0350709050947k5e32ba7fj38924a0968569d9a@mail.gmail.com","threadId":"9781","inReplyTo":"alpine.LFD.0.999.0709042355030.19879@evo.linux-foundation.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Govind Salinas","fromEmail":"govindsalinas@gmail.com","sentAt":"2007-09-05T16:47:45Z","receivedAt":"2007-09-05T16:47:45Z","isPatch":false,"sender":{"key":"govindsalinas@gmail.com","avatar":null},"body":"Hi,\n\nI am very new to git but I have thought about this a bit from a user's\nperspective.  I have several thoughts on the matter.\n\nFirst, I would like to point out that the hg folks like to compare\nthemselves to git a lot and they list the need for manual gc as a\nreason to choose hg over git.  This may not be something that the git\ncommunity cares about but I thought I would point it out.\n\nSecond, it *is* a hassle.  When trying to figure out what I could\nconvince my co-workers to use, having to gc was something that I did\nnot think they would be conscious of or care enough about to do.  It\nmakes git more of a PITA than it could be.  Similarly, I have no idea\nwhen it is a good time to do a gc.  After every commit?  Before push?\nWhat if I never push a repo?  What if it is a remote repo only used to\nsync up with my co-workers, do I have to go there and periodically gc?\n This is one reason why I really think that gc should be *plumbing*\nand *not* porcelain.\n\nThe user should never have to trigger a gc, they should even be\ndiscouraged from doing so.  That is how other gc systems are.  Can you\nimagine if you had a Java app that had a button on it to do a gc?\nWhen should I push it?  Should I wait till the system is getting slow\nor just start spamming the button whenever I'm bored?  I know that\nJava/c#/py GC are different than git gc, but they fulfill the same\nbasic purpose as git gc.  IE to clean up unused items and free up\nresources.  Git additionally may do some re-optimization, but that is\nnot relevant to a user.\n\nI know this goes against the general mood here (which seems to be\nagainst auto-gc) but I thought I would give my $.02 as a user of git.\n\nThanks,\nGovind.\n\nOn 9/5/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>\n> So we had a git bof at linux.conf.eu yesterday, and I leart something\n> new: even people who have been using git for a long time apparently don't\n> necessarily realize the importance of repacking.\n>\n> James Bottomley (the Linux SCSI maintainer) is an old-time BK user, and\n> very comfy using git. But when he was demonstrating things on his poor old\n> laptop, simple things like \"git branch\" literally took a long time, and\n> James didn't seem to realize that the fact that he had apparently never\n> ever repacked his repository was a big deal.\n>\n> The kernel archive is a 190MB pack for me fully repacked (I just checked -\n> I had actually thought that it was somewhat larger than that), but because\n> James hadn't repacked, his .git directory was over a gigabyte in size, and\n> his laptop wasn't able to cache anything at all effectively as a result.\n>\n> Repacking it took over an hour, simply because everything was *so*\n> unpacked, and James' kernel repository had something like 92 thousand\n> loose objects, and several hundred packfiles. Simple operations that\n> really take much less than a second for me (\"git branch\" takes 0.022s on\n> my laptop, which has the same 512M that James had on his) took many many\n> seconds as a result, and James seemed to think that this was all normal.\n>\n> And James didn't even want to repack, because it was so expensive (which\n> he knew - he claims to have never ever repacked at all, but maybe he had\n> started it and just control-C'd it when it was really slow at some point).\n>\n> Now, it may be that James didn't realize how important the occasional\n> garbage collect is exactly *because* he is an old-timer and used BK long\n> before he used git, and just continued using git simply as a BK\n> replacement, but it did make me wonder whether maybe this lack of\n> repacking awareness is fairly common.\n>\n> I've been against automatic repacking, but that was really based on what\n> appears to be potentially a very wrong assumption, namely that people\n> would do the manual repack on their own. If it turns out that people don't\n> do it, maybe the right thing for git to do really is to at least notify\n> people when they have way too many pack-files and/or loose objects.\n>\n> I personally repack everything way more often than is necessary, and I had\n> kind of assumed that people did it that way, but I was apparently wrong.\n> Comments?\n>\n>                 Linus\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n"},{"id":"52593","messageId":"87ir6pc9n1.wl%cworth@cworth.org","threadId":"9781","inReplyTo":"69b0c0350709050947k5e32ba7fj38924a0968569d9a@mail.gmail.com","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Carl Worth","fromEmail":"cworth@cworth.org","sentAt":"2007-09-05T17:19:46Z","receivedAt":"2007-09-05T17:19:46Z","isPatch":false,"sender":{"key":"cworth@cworth.org","avatar":"https://gravatar.com/avatar/3746dc28cde609bdbd7f939058356e7e2bbd16d21e32274df0725eb3d998bc5b?d=mp&s=160"},"body":"On Wed, 5 Sep 2007 11:47:45 -0500, \"Govind Salinas\" wrote:\n> I know this goes against the general mood here (which seems to be\n> against auto-gc) but I thought I would give my $.02 as a user of git.\n\nI'll throw my opinion in here as well. I think git should\nautomatically do repacking by default, (once loose objects exceed some\nthreshold). There have several posts in this thread from people who\ndon't want auto-gc, but these same people should be able to avoid it,\nand likely without changing habits. That's because:\n\n  * They're already in the habit of manually repacking every once in a\n    while, (or like, Linus, much more often than strictly necessary).\n\n  * They've already got cron jobs setup to do the repacking.\n\nAnd one could augment this with an option to disable the repacking of\ncourse.\n\nAnd if you're really concerned about people that don't want this\ngetting it anyway, just determine some useful threshold and then\ndouble it or so before it triggers automatic repacking, (so the\nautomatic repacking hits only us idiots that completely neglect it).\n\n[Pardon me for continuing to quote in the original top-posted order,\nbut I like the flow here.]\n\nOn 9/5/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> > James didn't seem to realize that the fact that he had apparently never\n> > ever repacked his repository was a big deal.\n\nI know it was surprising to you, Linus, but I'm glad you noticed\nit. I've seen the same thing from many users. And git actually\ndiscourages users from learning about repacking. If the user starts\nwith a small (or new) project, then everything performs well, and\nthere's no performance problem whatsoever.\n\nSo then the problems creep up gradually, and the user has no idea that\nhe should be doing anything different than he's always done. Instead\nthe user is left to just conclude that git's performance isn't scaling\nwell as the project grows. That's a bad conclusion of course, and it's\nbad that git sets things up so the user reaches that conclusion.\nInstead, git should just fix things up itself in this case.\n\n> > I've been against automatic repacking, but that was really based on what\n> > appears to be potentially a very wrong assumption, namely that people\n> > would do the manual repack on their own. If it turns out that people don't\n> > do it, maybe the right thing for git to do really is to at least notify\n> > people when they have way too many pack-files and/or loose\n> > objects.\n\nI don't think the warning message alone is a good fix. I think the\npeople who would understand the warning and appreciate that they could\nthen take care of repacking as convenient are the same people that\nalready understand the repacking concept, and are likely already\nrepacking occasionally, (so would likely never see the warning).\n\nBut the problematic case is the user who knows nothing of the\nissue. And in that case, giving this warning isn't useful education,\nit's just forcing the user to learn more and do more work. \"If git\nnotices it has too many 'loose object' and 'git gc' would fix the\nproblem, then why didn't it do that itself? And what the heck is a\n'loose object' anyway?\"\n\nIn general, git has always printed too many obscure messages that\ndon't actually help a new user get his work done, (and the work is\n_not_ to learn more about git internals). From 1.4 to 1.5 much of that\nwas improved. But please let's not go backwards by adding more of\nthese.\n\nSo one vote from me for auto repacking, (but feel free to make the\nthreshold so high that anyone that actually _cares_ about loose\nobjects and repacking will never get the auto repack).\n\n-Carl\n"},{"id":"52596","messageId":"vpq1wddkohr.fsf@bauges.imag.fr","threadId":"9781","inReplyTo":"D32A7C27-EAF5-4156-BE0E-99FE3D948AE8@wgaf.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2007-09-05T17:31:44Z","receivedAt":"2007-09-05T17:31:44Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"Johan De Messemaeker <johan.demessemaeker@wgaf.org> writes:\n\n>> Currently, AFAIK, that can only be done with a (trivial) script\n>> external to git. I suppose this can easily be added to the core git\n>> porcelain. Perhaps a \"git gc --recursive\" would do.\n>>\n>> It doesn't solve the problem, but makes it easier to solve it (git gc\n>> --recursive in cron for example).\n>\n> I'm a git newb so I can be wrong here but ...\n>\n> Why --recursive? Why not use the submodule-information ?\n\nall projects are not necessarily subprojects of each others.\n\nI have ~/teaching/some-course/.git (well, almost) and ~/etc/.git which\nare two unrelated projects, and to \"git gc\" both of them, I need\neither a script, or two manual invocations.\n\n(yes, I'm really talking about something trivial)\n\n-- \nMatthieu\n"},{"id":"52597","messageId":"46DEE8E8.2000801@midwinter.com","threadId":"9781","inReplyTo":"69b0c0350709050947k5e32ba7fj38924a0968569d9a@mail.gmail.com","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Steven Grimm","fromEmail":"koreth@midwinter.com","sentAt":"2007-09-05T17:35:36Z","receivedAt":"2007-09-05T17:35:36Z","isPatch":false,"sender":{"key":"koreth@midwinter.com","avatar":"https://gravatar.com/avatar/71b4d2e8b62f168bdc9e9205341159e3567003b4f9e2127c617c5fa0a1f5bad2?d=mp&s=160"},"body":"Govind Salinas wrote:\n> This is one reason why I really think that gc should be *plumbing*\n> and *not* porcelain.\n>   \n\nThat's a good way to think of it IMO. It's a low-level operation (albeit \none that encapsulates other, lower-level ones) that tells git to \nrearrange its internal data structures. It is not something that has any \nuser-visible effect. Every other porcelain-level git command *does \nsomething* from the user's point of view. Running git-gc is basically a \nno-op, which from the user's point of view makes it a waste of \nkeystrokes and an annoying distraction from focusing on the stuff \nthey're using git to help them build.\n\n> The user should never have to trigger a gc, they should even be\n> discouraged from doing so.  That is how other gc systems are.  Can you\n> imagine if you had a Java app that had a button on it to do a gc?\n> When should I push it?  Should I wait till the system is getting slow\n> or just start spamming the button whenever I'm bored?  I know that\n> Java/c#/py GC are different than git gc, but they fulfill the same\n> basic purpose as git gc.  IE to clean up unused items and free up\n> resources.  Git additionally may do some re-optimization, but that is\n> not relevant to a user.\n>   \n\nI'll play devil's advocate for a moment here, though, and say that, as \nothers have suggested in this thread, git could be made to tell you when \nit's appropriate to run gc. So the \"I don't know when to run it\" \nargument isn't a hard one to address.\n\nWith that in mind, here's what the message should look like IMO:\n\n---\nYour repository can be optimized for better performance and lower disk \nusage.\nPlease run \"git gc\" to optimize it now, or run \"git config gc.auto true\" \nto tell\ngit to automatically optimize it in the future (this will launch \nprocesses in the\nbackground.) For more information, \"man git-gc\".\n---\n\nAnd that \"gc.auto\" config option (just an arbitrary name, call it \nsomething else if that's no good) actually has four settings:\n\nwarn (the default) - prints the warning message, at most once every N \nminutes (we can determine a good value for N)\ntrue - launches git-gc in the background as needed\nfalse - suppresses the warning and the check that triggers the warning\nforeground - launches git-gc in the foreground as needed (to make it \neasier to abort)\n\n\nI don't buy the \"git gc takes too much memory to run in the background\" \nargument as a reason automatic git-gc is a bad idea. Many of us (me \nincluded) work on machines with plenty of memory to launch a background \ngit-gc without hampering our development work, and/or on repositories \nsmall enough that it doesn't eat that much memory in the first place. \nAnd if you make it an option that the user has to enable, people on \nlow-memory machines can simply not enable it, end of problem.\n\nOne big problem with git-gc now is that it's not discoverable. Or \nrather, the need for it isn't discoverable. So at the very least we \nshould print the warning, IMO -- and if we're already going to all the \ntrouble to determine whether or not git-gc needs to be run, it will \nreduce the \"why are you telling me to run something when you could just \ndo it for me, you stupid machine?\" factor if there's an easily \ndiscoverable way to just do it as needed.\n\n-Steve\n"},{"id":"52599","messageId":"20070905174427.GC13314@fieldses.org","threadId":"9781","inReplyTo":"alpine.LFD.0.999.0709042355030.19879@evo.linux-foundation.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"J. Bruce Fields","fromEmail":"bfields@fieldses.org","sentAt":"2007-09-05T17:44:27Z","receivedAt":"2007-09-05T17:44:27Z","isPatch":false,"sender":{"key":"bfields@citi.umich.edu","avatar":null},"body":"On Wed, Sep 05, 2007 at 12:09:27AM -0700, Linus Torvalds wrote:\n> I personally repack everything way more often than is necessary, and I had \n> kind of assumed that people did it that way, but I was apparently wrong. \n> Comments?\n\nWell, this may just prove I'm an idiot, but one of the reasons I rarely\nrun it is that I have trouble remembering exactly what it does; in\nparticular,\n\n\t- does it prune anything that might be needed by a repo I\n\t  cloned with -s?\n\t- is there anything that's unsafe to do while the git-gc is\n\t  running?\n\t- what are the implications for http users if this is a public\n\t  repo?\n\t- is git-gc enough on its own or should I be running something\n\t  more agressive ocassionally too?\n\nNo doubt they all have simple answers, which probably amount to \"just\ndon't worry about it\", and which I could have found in less time than\nit'd take to write this email.  But when I've got other work to do,\nreading \"man git-gc\" is just enough effort for me to postpone the whole\nthing to another day.\n\nSo, anyway, your message reminded me to run git-gc on my main working\nrepo.  At which point one of my personal scripts immediately started\nfailing--it was assuming it could find any ref under .git/refs/, and I\nhadn't realized (or maybe I had once, and I'd forgotten) that git-gc\npacks refs by default now.\n\nBah.  I don't know what the moral of that story is.\n\n--b.\n"},{"id":"52600","messageId":"87odgh0zn6.fsf@hades.wkstn.nix","threadId":"9781","inReplyTo":"20070905074206.GA31750@artemis.corp","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Nix","fromEmail":"nix@esperi.org.uk","sentAt":"2007-09-05T17:51:09Z","receivedAt":"2007-09-05T17:51:09Z","isPatch":false,"sender":{"key":"nix@esperi.org.uk","avatar":"https://avatars.githubusercontent.com/u/6503005?v=4"},"body":"On 5 Sep 2007, Pierre Habouzit said:\n>   I know I don't really answer the question, but the point I try to make\n> is that yeah, some kind of automated way to run the gc is great, but I'm\n> not sure that _git_ is the tool to automate that, because when *I* use\n> git, I expect it to be just plain fast, and I don't want it to\n> occasionally hang.\n\nIndeed. I repack all our git trees in the middle of the night, and our\nincremental backup script drops .keep files corresponding to every\nexisting pack before running the backup.\n\nThis is probably a good job for cron :)\n"},{"id":"52602","messageId":"20070905135549.b5k3etn94wos4g4o@intranet.digizenstudio.com","threadId":"9781","inReplyTo":"87ir6pc9n1.wl%cworth@cworth.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Jing Xue","fromEmail":"jingxue@digizenstudio.com","sentAt":"2007-09-05T17:55:49Z","receivedAt":"2007-09-05T17:55:49Z","isPatch":false,"sender":{"key":"jingxue@digizenstudio.com","avatar":null},"body":"\nQuoting Carl Worth <cworth@cworth.org>:\n\n> I don't think the warning message alone is a good fix. I think the\n> people who would understand the warning and appreciate that they could\n> then take care of repacking as convenient are the same people that\n> already understand the repacking concept, and are likely already\n> repacking occasionally, (so would likely never see the warning).\n>\n> But the problematic case is the user who knows nothing of the\n> issue. And in that case, giving this warning isn't useful education,\n> it's just forcing the user to learn more and do more work. \"If git\n> notices it has too many 'loose object' and 'git gc' would fix the\n> problem, then why didn't it do that itself? And what the heck is a\n> 'loose object' anyway?\"\n\n(my 2 cents as another ordinary new git user)\nHmm, not necessarily. That a system knows what the best action is  \ndoesn't meant that _right now_ is the best time to take that action.   \nOne subtle difference I think between git's gc and Java/python/etc.'s  \ngc is that in the latter case it is, at least metaphorically, a life  \nand death situation - if gc isn't run, the application will run out of  \nmemory, where as in git, it's more of a performance degradation issue,  \nwhich, sort of, can wait.\n\nOn the issue of implementation awareness, a warning message saying  \nsomething along the lines of \"your repository is getting slower. You  \nmight want to consider running 'git gc', and remember to do that from  \ntime to time.\" is not much different from \"your file system is getting  \nslower. You might want to consider running <whatever-defrag-tool>, and  \nremember to do that from time to time.\"\n\nNeither these messages nor the actions they propose _require_ users to  \nlearn what \"repacking\", \"loose object\", or \"file fragments\" are about  \nbefore they can proceed.\n\nCheers.\n-- \nJing Xue\n"},{"id":"52604","messageId":"46DEF1FA.4050500@midwinter.com","threadId":"9781","inReplyTo":"87odgh0zn6.fsf@hades.wkstn.nix","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Steven Grimm","fromEmail":"koreth@midwinter.com","sentAt":"2007-09-05T18:14:18Z","receivedAt":"2007-09-05T18:14:18Z","isPatch":false,"sender":{"key":"koreth@midwinter.com","avatar":"https://gravatar.com/avatar/71b4d2e8b62f168bdc9e9205341159e3567003b4f9e2127c617c5fa0a1f5bad2?d=mp&s=160"},"body":"Nix wrote:\n> Indeed. I repack all our git trees in the middle of the night, and our\n> incremental backup script drops .keep files corresponding to every\n> existing pack before running the backup.\n>\n> This is probably a good job for cron :)\n>   \n\nIf you are setting up cron jobs to repack multiple git trees, you are \nnot the kind of novice or casual git user who this proposal would \nprimarily be aimed at.\n\nBut in any event, since you are doing that, your repos will never \naccumulate a high enough percentage of loose objects (whatever the \nthreshold is) to trigger the warning and/or automatic launch. So you can \ncontinue to operate as before, no difference in behavior, while people \nwho don't know how / want to set up cron jobs will have their \nrepositories cleaned too.\n\ngit-gc can leave behind a \"last completed\" timestamp and we can suppress \nthe check for excess loose objects until some minimum amount of time has \npassed since last git-gc. If that amount is greater than the interval \nbetween your cron jobs, you won't even get any (measurable) overhead \nfrom the detection to see if the warning is needed.\n\n-Steve\n"},{"id":"52605","messageId":"877in50y7p.fsf@hades.wkstn.nix","threadId":"9781","inReplyTo":"46DEF1FA.4050500@midwinter.com","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Nix","fromEmail":"nix@esperi.org.uk","sentAt":"2007-09-05T18:22:02Z","receivedAt":"2007-09-05T18:22:02Z","isPatch":false,"sender":{"key":"nix@esperi.org.uk","avatar":"https://avatars.githubusercontent.com/u/6503005?v=4"},"body":"On 5 Sep 2007, Steven Grimm stated:\n\n> Nix wrote:\n>> Indeed. I repack all our git trees in the middle of the night, and our\n>> incremental backup script drops .keep files corresponding to every\n>> existing pack before running the backup.\n>>\n>> This is probably a good job for cron :)\n>\n> If you are setting up cron jobs to repack multiple git trees, you are\n> not the kind of novice or casual git user who this proposal would\n> primarily be aimed at.\n\nTrue enough: but the point is that it was only about three lines of code\n(a locate and git-gc pipeline). We could just put that in the\ndocumentation...\n\n... which people then won't read. Oh well. Sorry for the mindless\noptimism.\n\n> git-gc can leave behind a \"last completed\" timestamp and we can\n> suppress the check for excess loose objects until some minimum amount\n> of time has passed since last git-gc. If that amount is greater than\n> the interval between your cron jobs, you won't even get any\n> (measurable) overhead from the detection to see if the warning is\n> needed.\n\nI personally wonder if git-gc shouldn't use a proportional scheme, so\nthat only some packs get repacked, maybe the smallest ones (and when\nthey grow to the same size as the next largest one, the two get repacked\ninto one). This has the singular advantage that you won't have to\ncarefully drop .keep files everywhere or have to worry about your git-gc\nof 50K of loose objects suddenly deciding to repack 100Mb of packfiles\nand taking ages.\n\nIt's probably not hard to implement, but I don't need it because I keep\neverything packed anyway...\n"},{"id":"52606","messageId":"873axt0xxe.fsf@hades.wkstn.nix","threadId":"9781","inReplyTo":"46DEE8E8.2000801@midwinter.com","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Nix","fromEmail":"nix@esperi.org.uk","sentAt":"2007-09-05T18:28:13Z","receivedAt":"2007-09-05T18:28:13Z","isPatch":false,"sender":{"key":"nix@esperi.org.uk","avatar":"https://avatars.githubusercontent.com/u/6503005?v=4"},"body":"On 5 Sep 2007, Steven Grimm stated:\n\n> Govind Salinas wrote:\n>> This is one reason why I really think that gc should be *plumbing*\n>> and *not* porcelain.\n>\n> That's a good way to think of it IMO. It's a low-level operation\n> (albeit one that encapsulates other, lower-level ones) that tells git\n> to rearrange its internal data structures. It is not something that\n> has any user-visible effect.\n\nIt certainly has a sysadmin-visible effect. Repack a couple of big git\nrepositories and that's a backup tape gone if you do incremental\nbackups: and you can't *not* back up the pack files, even though a lot\nof the state in them is recoverable from elsewhere on the net: the stuff\nwhich is not recoverable is tangled up with the stuff which is.\n\n(of course the solution here was .keep files. I cheered when they were\nintroduced and started rolling git out everywhere I could. There's just\none last vast repository maintained by a horrible shell script layered\natop SCCS which I have to find some way to convert...)\n"},{"id":"52608","messageId":"Pine.LNX.4.64.0709051339420.30020@torch.nrlssc.navy.mil","threadId":"9781","inReplyTo":"20070905174427.GC13314@fieldses.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Brandon Casey","fromEmail":"casey@nrlssc.navy.mil","sentAt":"2007-09-05T18:46:42Z","receivedAt":"2007-09-05T18:46:42Z","isPatch":false,"sender":{"key":"drafnel@gmail.com","avatar":"https://avatars.githubusercontent.com/u/921167?v=4"},"body":"On Wed, 5 Sep 2007, J. Bruce Fields wrote:\n\n> Well, this may just prove I'm an idiot, but one of the reasons I rarely\n> run it is that I have trouble remembering exactly what it does; in\n> particular,\n>\n> \t- does it prune anything that might be needed by a repo I\n> \t  cloned with -s?\n\n     YES! yikes.\n\nThis is about the best argument put forth so far for not automatically\nrunning git-gc. Personally, I think git-gc should not remove unreferenced\nobjects without --prune (but I haven't done anything about it). But even\nif git-gc was modified in this way, an occasional git-gc --prune would\nstill be necessary to remove all of the unreferenced and dangling objects\nsafely with a human thinking about the shared repo implications (unless\nshared repo handling is modified).\n\n-brandon\n"},{"id":"52610","messageId":"alpine.LFD.0.9999.0709051438460.21186@xanadu.home","threadId":"9781","inReplyTo":"877in50y7p.fsf@hades.wkstn.nix","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-09-05T18:54:40Z","receivedAt":"2007-09-05T18:54:40Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 5 Sep 2007, Nix wrote:\n\n> I personally wonder if git-gc shouldn't use a proportional scheme, so\n> that only some packs get repacked, maybe the smallest ones (and when\n> they grow to the same size as the next largest one, the two get repacked\n> into one). This has the singular advantage that you won't have to\n> carefully drop .keep files everywhere or have to worry about your git-gc\n> of 50K of loose objects suddenly deciding to repack 100Mb of packfiles\n> and taking ages.\n\nNot only that.  Currently the \"Counting objects\" phase when running \ngit-gc on the Linux repo takes a significant amount of time, even if \nthere is little to repack.\n\nIf any kind of automatic repack is implemented, it should be an \nincremental repacking only, not the full thing, i.e. git-repack without \n-a, or git-pack-objects with --unpacked.  The idea is to be the least \nintrusive as possible.  Also, object walking should be limited to \nobjects linked to a commit object which is itself unpacked in order to \ncut on the time required to fully enumerate all objects.\n\nThis way a semi-packed state will always be preserved and should be good \nenough.  The full repacking should probably be left to manual execution \nof git-gc.\n\n\nNicolas\n"},{"id":"52612","messageId":"85642phqtn.fsf@lola.goethe.zz","threadId":"9781","inReplyTo":"Pine.LNX.4.64.0709051339420.30020@torch.nrlssc.navy.mil","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-05T19:09:40Z","receivedAt":"2007-09-05T19:09:40Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Brandon Casey <casey@nrlssc.navy.mil> writes:\n\n> On Wed, 5 Sep 2007, J. Bruce Fields wrote:\n>\n>> Well, this may just prove I'm an idiot, but one of the reasons I rarely\n>> run it is that I have trouble remembering exactly what it does; in\n>> particular,\n>>\n>> \t- does it prune anything that might be needed by a repo I\n>> \t  cloned with -s?\n>\n>     YES! yikes.\n>\n> This is about the best argument put forth so far for not\n> automatically running git-gc.\n\nWell, it could also mean that if git finds a dead symbolic link when\nlooking up an object, it should check the corresponding link target\ndirectory for a pack file with the respective object...  and if it\nfinds such a pack file, create a link to it and use it.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"52613","messageId":"20070905191310.GE13314@fieldses.org","threadId":"9781","inReplyTo":"85642phqtn.fsf@lola.goethe.zz","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"J. Bruce Fields","fromEmail":"bfields@fieldses.org","sentAt":"2007-09-05T19:13:10Z","receivedAt":"2007-09-05T19:13:10Z","isPatch":false,"sender":{"key":"bfields@citi.umich.edu","avatar":null},"body":"On Wed, Sep 05, 2007 at 09:09:40PM +0200, David Kastrup wrote:\n> Brandon Casey <casey@nrlssc.navy.mil> writes:\n> \n> > On Wed, 5 Sep 2007, J. Bruce Fields wrote:\n> >\n> >> Well, this may just prove I'm an idiot, but one of the reasons I rarely\n> >> run it is that I have trouble remembering exactly what it does; in\n> >> particular,\n> >>\n> >> \t- does it prune anything that might be needed by a repo I\n> >> \t  cloned with -s?\n> >\n> >     YES! yikes.\n> >\n> > This is about the best argument put forth so far for not\n> > automatically running git-gc.\n> \n> Well, it could also mean that if git finds a dead symbolic link when\n> looking up an object, it should check the corresponding link target\n> directory for a pack file with the respective object...  and if it\n> finds such a pack file, create a link to it and use it.\n\nOne of the two of us is very confused about what \"git-clone -s\" does.\nSee the git-clone man page.  I don't think symlinks are involved.\n\n--b.\n"},{"id":"52623","messageId":"20070905192033.GA4681@glandium.org","threadId":"9781","inReplyTo":"85642phqtn.fsf@lola.goethe.zz","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2007-09-05T19:20:33Z","receivedAt":"2007-09-05T19:20:33Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Wed, Sep 05, 2007 at 09:09:40PM +0200, David Kastrup <dak@gnu.org> wrote:\n> Brandon Casey <casey@nrlssc.navy.mil> writes:\n> \n> > On Wed, 5 Sep 2007, J. Bruce Fields wrote:\n> >\n> >> Well, this may just prove I'm an idiot, but one of the reasons I rarely\n> >> run it is that I have trouble remembering exactly what it does; in\n> >> particular,\n> >>\n> >> \t- does it prune anything that might be needed by a repo I\n> >> \t  cloned with -s?\n> >\n> >     YES! yikes.\n> >\n> > This is about the best argument put forth so far for not\n> > automatically running git-gc.\n> \n> Well, it could also mean that if git finds a dead symbolic link when\n> looking up an object, it should check the corresponding link target\n> directory for a pack file with the respective object...  and if it\n> finds such a pack file, create a link to it and use it.\n\nThe problem here is that the clone could be having refs on objects from\nthe origin that don't have refs left there. git-gc might, at some point,\nprune these refs, and the clone would have dangling refs. That could\neasily happen, for example, if you rebase a branch in the origin, but\nstill have a clone with the original branch.\n\nMike\n"},{"id":"52624","messageId":"85sl5shp9j.fsf@lola.goethe.zz","threadId":"9781","inReplyTo":"20070905191310.GE13314@fieldses.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-05T19:43:20Z","receivedAt":"2007-09-05T19:43:20Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"\"J. Bruce Fields\" <bfields@fieldses.org> writes:\n\n> On Wed, Sep 05, 2007 at 09:09:40PM +0200, David Kastrup wrote:\n>> Brandon Casey <casey@nrlssc.navy.mil> writes:\n>> \n>> > On Wed, 5 Sep 2007, J. Bruce Fields wrote:\n>> >\n>> >> Well, this may just prove I'm an idiot, but one of the reasons I rarely\n>> >> run it is that I have trouble remembering exactly what it does; in\n>> >> particular,\n>> >>\n>> >> \t- does it prune anything that might be needed by a repo I\n>> >> \t  cloned with -s?\n>> >\n>> >     YES! yikes.\n>> >\n>> > This is about the best argument put forth so far for not\n>> > automatically running git-gc.\n>> \n>> Well, it could also mean that if git finds a dead symbolic link when\n>> looking up an object, it should check the corresponding link target\n>> directory for a pack file with the respective object...  and if it\n>> finds such a pack file, create a link to it and use it.\n>\n> One of the two of us is very confused about what \"git-clone -s\" does.\n> See the git-clone man page.  I don't think symlinks are involved.\n\nGuilty as charged.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"52628","messageId":"7vr6lcj2zi.fsf@gitster.siamese.dyndns.org","threadId":"9781","inReplyTo":"alpine.LFD.0.9999.0709051438460.21186@xanadu.home","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-09-05T20:01:37Z","receivedAt":"2007-09-05T20:01:37Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nicolas Pitre <nico@cam.org> writes:\n\n> Not only that.  Currently the \"Counting objects\" phase when running \n> git-gc on the Linux repo takes a significant amount of time, even if \n> there is little to repack.\n>\n> If any kind of automatic repack is implemented, it should be an \n> incremental repacking only, not the full thing, i.e. git-repack without \n> -a, or git-pack-objects with --unpacked.  The idea is to be the least \n> intrusive as possible.  Also, object walking should be limited to \n> objects linked to a commit object which is itself unpacked in order to \n> cut on the time required to fully enumerate all objects.\n>\n> This way a semi-packed state will always be preserved and should be good \n> enough.  The full repacking should probably be left to manual execution \n> of git-gc.\n\nOk, how about doing something like this?\n\n-- >8 -- snipsnap -- >8 -- clipcrap -- >8 --\nImplement git gc --auto\n\nThis implements a new option \"git gc --auto\".  When gc.auto is\nset to a positive value, and the object database has accumulated\nroughly that many number of loose objects, this runs a\nlightweight version of \"git gc\".  The primary difference from\nthe full \"git gc\" is that it does not pass \"-a\" option to \"git\nrepack\", which means we do not try to repack _everything_, but\nonly repack incrementally.  We still do \"git prune-packed\".  The\ndefault threshold is arbitrarily set by yours truly to:\n\n - not trigger it for fully unpacked git v0.99 history;\n\n - do trigger it for fully unpacked git v1.0.0 history;\n\n - not trigger it for incremental update to git v1.0.0 starting\n   from fully packed git v0.99 history.\n\nThis patch does not add invocation of the \"auto repacking\".  It\nis left to key Porcelain commands that could produce tons of\nloose objects to add a call to \"git gc --auto\" after they are\ndone their work.  Obvious candidates are:\n\n\tgit add\n\tgit fetch\n        git merge\n        git rebase        \n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n builtin-gc.c |   64 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++-\n 1 files changed, 63 insertions(+), 1 deletions(-)\n\ndiff --git a/builtin-gc.c b/builtin-gc.c\nindex 9397482..093b3dd 100644\n--- a/builtin-gc.c\n+++ b/builtin-gc.c\n@@ -20,6 +20,7 @@ static const char builtin_gc_usage[] = \"git-gc [--prune] [--aggressive]\";\n \n static int pack_refs = 1;\n static int aggressive_window = -1;\n+static int gc_auto_threshold = 6700;\n \n #define MAX_ADD 10\n static const char *argv_pack_refs[] = {\"pack-refs\", \"--all\", \"--prune\", NULL};\n@@ -28,6 +29,8 @@ static const char *argv_repack[MAX_ADD] = {\"repack\", \"-a\", \"-d\", \"-l\", NULL};\n static const char *argv_prune[] = {\"prune\", NULL};\n static const char *argv_rerere[] = {\"rerere\", \"gc\", NULL};\n \n+static const char *argv_repack_auto[] = {\"repack\", \"-d\", \"-l\", NULL};\n+\n static int gc_config(const char *var, const char *value)\n {\n \tif (!strcmp(var, \"gc.packrefs\")) {\n@@ -41,6 +44,10 @@ static int gc_config(const char *var, const char *value)\n \t\taggressive_window = git_config_int(var, value);\n \t\treturn 0;\n \t}\n+\tif (!strcmp(var, \"gc.auto\")) {\n+\t\tgc_auto_threshold = git_config_int(var, value);\n+\t\treturn 0;\n+\t}\n \treturn git_default_config(var, value);\n }\n \n@@ -57,10 +64,49 @@ static void append_option(const char **cmd, const char *opt, int max_length)\n \tcmd[i] = NULL;\n }\n \n+static int need_to_gc(void)\n+{\n+\t/*\n+\t * Quickly check if a \"gc\" is needed, by estimating how\n+\t * many loose objects there are.  Because SHA-1 is evenly\n+\t * distributed, we can check only one and get a reasonable\n+\t * estimate.\n+\t */\n+\tchar path[PATH_MAX];\n+\tconst char *objdir = get_object_directory();\n+\tDIR *dir;\n+\tstruct dirent *ent;\n+\tint auto_threshold;\n+\tint num_loose = 0;\n+\tint needed = 0;\n+\n+\tif (sizeof(path) <= snprintf(path, sizeof(path), \"%s/17\", objdir)) {\n+\t\twarning(\"insanely long object directory %.*s\", 50, objdir);\n+\t\treturn 0;\n+\t}\n+\tdir = opendir(path);\n+\tif (!dir)\n+\t\treturn 0;\n+\n+\tauto_threshold = (gc_auto_threshold + 255) / 256;\n+\twhile ((ent = readdir(dir)) != NULL) {\n+\t\tif (strspn(ent->d_name, \"0123456789abcdef\") != 38 ||\n+\t\t    ent->d_name[38] != '\\0')\n+\t\t\tcontinue;\n+\t\tif (++num_loose > auto_threshold) {\n+\t\t\tneeded = 1;\n+\t\t\tbreak;\n+\t\t}\n+\t}\n+\tclosedir(dir);\n+\treturn needed;\n+}\n+\n int cmd_gc(int argc, const char **argv, const char *prefix)\n {\n \tint i;\n \tint prune = 0;\n+\tint auto_gc = 0;\n \tchar buf[80];\n \n \tgit_config(gc_config);\n@@ -82,12 +128,28 @@ int cmd_gc(int argc, const char **argv, const char *prefix)\n \t\t\t}\n \t\t\tcontinue;\n \t\t}\n-\t\t/* perhaps other parameters later... */\n+\t\tif (!strcmp(arg, \"--auto\")) {\n+\t\t\tif (gc_auto_threshold <= 0)\n+\t\t\t\treturn 0;\n+\t\t\tauto_gc = 1;\n+\t\t\tcontinue;\n+\t\t}\n \t\tbreak;\n \t}\n \tif (i != argc)\n \t\tusage(builtin_gc_usage);\n \n+\tif (auto_gc) {\n+\t\t/*\n+\t\t * Auto-gc should be least intrusive as possible.\n+\t\t */\n+\t\tprune = 0;\n+\t\tfor (i = 0; i < ARRAY_SIZE(argv_repack_auto); i++)\n+\t\t\targv_repack[i] = argv_repack_auto[i];\n+\t\tif (!need_to_gc())\n+\t\t\treturn 0;\n+\t}\n+\n \tif (pack_refs && run_command_v_opt(argv_pack_refs, RUN_GIT_CMD))\n \t\treturn error(FAILED_RUN, argv_pack_refs[0]);\n \n"},{"id":"52630","messageId":"alpine.LFD.0.9999.0709051634190.21186@xanadu.home","threadId":"9781","inReplyTo":"7vr6lcj2zi.fsf@gitster.siamese.dyndns.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-09-05T20:35:19Z","receivedAt":"2007-09-05T20:35:19Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 5 Sep 2007, Junio C Hamano wrote:\n\n> Implement git gc --auto\n> \n> This implements a new option \"git gc --auto\".  When gc.auto is\n> set to a positive value, and the object database has accumulated\n> roughly that many number of loose objects, this runs a\n> lightweight version of \"git gc\".  The primary difference from\n> the full \"git gc\" is that it does not pass \"-a\" option to \"git\n> repack\", which means we do not try to repack _everything_, but\n> only repack incrementally.  We still do \"git prune-packed\".  \n\nA big part of the repack cost is the counting of objects. I don't know \nif --unpacked to git-pack-objects skips walking trees of a packed commit \nobject.  If no then it probably should to gain a significant speed up, \nor maybe a separate option should be created to actually imply this \nloosened semantic.\n\n> This patch does not add invocation of the \"auto repacking\".  It\n> is left to key Porcelain commands that could produce tons of\n> loose objects to add a call to \"git gc --auto\" after they are\n> done their work.  Obvious candidates are:\n> \n> \tgit add\n\nNope!  'git add' creates loose objects which are not yet reachable from \nanywhere.  They won't get repacked until a commit is made.\n\n> \tgit fetch\n\nI think that would be a much better idea to simply decrease the \nfetch.unpackLimit default value.\n\n>         git merge\n>         git rebase        \n\nand git commit.  Which resumes it to commit creating operation.\n\n\nNicolas\n"},{"id":"52631","messageId":"7vhcm8j1bp.fsf_-_@gitster.siamese.dyndns.org","threadId":"9781","inReplyTo":"7vr6lcj2zi.fsf@gitster.siamese.dyndns.org","subject":"[PATCH] Invoke \"git gc --auto\" from \"git add\" and \"git fetch\"","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-09-05T20:37:30Z","receivedAt":"2007-09-05T20:37:30Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"This makes the two commands to call \"git gc --auto\" when they\nare done.\n\nI earlier said that obvious candidates also include merge and\nrebase, but these are lot less frequent operations compared to\nadd, and more importantly, in a normal workflow they would\nalmost always happen after \"git fetch\" is done.\n\nIn other words, if you are downstream developer, the automatic\ninvocation in \"git fetch\" will take care of things for you, and\notherwise if you do not have an upstream, you would be doing\nyour own development, so \"git add\" to add your changes will take\ncare of the auto invocation for you.\n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n * This is obviously a follow-up to the previous one that allows\n   you to say \"git gc --auto\".  I somewhat feel dirty about\n   calling cmd_gc() bypassing fork & exec from \"git add\",\n   though...\n\n builtin-add.c |    2 ++\n git-fetch.sh  |    1 +\n 2 files changed, 3 insertions(+), 0 deletions(-)\n\ndiff --git a/builtin-add.c b/builtin-add.c\nindex 105a9f0..8431c16 100644\n--- a/builtin-add.c\n+++ b/builtin-add.c\n@@ -263,9 +263,11 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n \n  finish:\n \tif (active_cache_changed) {\n+\t\tconst char *args[] = { \"gc\", \"--auto\", NULL };\n \t\tif (write_cache(newfd, active_cache, active_nr) ||\n \t\t    close(newfd) || commit_locked_index(&lock_file))\n \t\t\tdie(\"Unable to write new index file\");\n+\t\tcmd_gc(2, args, NULL);\n \t}\n \n \treturn 0;\ndiff --git a/git-fetch.sh b/git-fetch.sh\nindex c3a2001..86050eb 100755\n--- a/git-fetch.sh\n+++ b/git-fetch.sh\n@@ -375,3 +375,4 @@ case \"$orig_head\" in\n \tfi\n \t;;\n esac\n+git gc --auto\n-- \n1.5.3.1.840.g0fedbc\n"},{"id":"52633","messageId":"69b0c0350709051359o343ce517md19bda824d84852b@mail.gmail.com","threadId":"9781","inReplyTo":"69b0c0350709051357ifa547aarfe3e0b36cf9be98f@mail.gmail.com","subject":"Fwd: [PATCH] Invoke \"git gc --auto\" from \"git add\" and \"git fetch\"","fromName":"Govind Salinas","fromEmail":"govindsalinas@gmail.com","sentAt":"2007-09-05T20:59:17Z","receivedAt":"2007-09-05T20:59:17Z","isPatch":true,"sender":{"key":"govindsalinas@gmail.com","avatar":null},"body":"Forgot to cc the list.\n\n---------- Forwarded message ----------\nFrom: Govind Salinas <govindsalinas@gmail.com>\nDate: Sep 5, 2007 3:57 PM\nSubject: Re: [PATCH] Invoke \"git gc --auto\" from \"git add\" and \"git fetch\"\nTo: Junio C Hamano <gitster@pobox.com>\n\n\nI have a completely uninformed question...\n\nCan git-add/rm/etc create dangling object or objects that would\nbe cleaned up by git-gc --auto?  I would think (and I could be\ncompletely off base here) that you would only want to call gc\nafter an operation that could create stuff that needs to be gc'ed,\nsince only then could the threshold be reached.\n\nAnyways, just curious.  One day I should actually go in and read\nsome git code.\n\n-Govind\n\nOn 9/5/07, Junio C Hamano <gitster@pobox.com> wrote:\n> This makes the two commands to call \"git gc --auto\" when they\n> are done.\n>\n> I earlier said that obvious candidates also include merge and\n> rebase, but these are lot less frequent operations compared to\n> add, and more importantly, in a normal workflow they would\n> almost always happen after \"git fetch\" is done.\n>\n> In other words, if you are downstream developer, the automatic\n> invocation in \"git fetch\" will take care of things for you, and\n> otherwise if you do not have an upstream, you would be doing\n> your own development, so \"git add\" to add your changes will take\n> care of the auto invocation for you.\n>\n> Signed-off-by: Junio C Hamano <gitster@pobox.com>\n> ---\n>  * This is obviously a follow-up to the previous one that allows\n>    you to say \"git gc --auto\".  I somewhat feel dirty about\n>    calling cmd_gc() bypassing fork & exec from \"git add\",\n>    though...\n>\n>  builtin-add.c |    2 ++\n>  git-fetch.sh  |    1 +\n>  2 files changed, 3 insertions(+), 0 deletions(-)\n>\n> diff --git a/builtin-add.c b/builtin-add.c\n> index 105a9f0..8431c16 100644\n> --- a/builtin-add.c\n> +++ b/builtin-add.c\n> @@ -263,9 +263,11 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n>\n>   finish:\n>         if (active_cache_changed) {\n> +               const char *args[] = { \"gc\", \"--auto\", NULL };\n>                 if (write_cache(newfd, active_cache, active_nr) ||\n>                     close(newfd) || commit_locked_index(&lock_file))\n>                         die(\"Unable to write new index file\");\n> +               cmd_gc(2, args, NULL);\n>         }\n>\n>         return 0;\n> diff --git a/git-fetch.sh b/git-fetch.sh\n> index c3a2001..86050eb 100755\n> --- a/git-fetch.sh\n> +++ b/git-fetch.sh\n> @@ -375,3 +375,4 @@ case \"$orig_head\" in\n>         fi\n>         ;;\n>  esac\n> +git gc --auto\n> --\n> 1.5.3.1.840.g0fedbc\n>\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n"},{"id":"52634","messageId":"20070905210741.GA3770@steel.home","threadId":"9781","inReplyTo":"alpine.LFD.0.999.0709042355030.19879@evo.linux-foundation.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2007-09-05T21:07:41Z","receivedAt":"2007-09-05T21:07:41Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"Linus Torvalds, Wed, Sep 05, 2007 09:09:27 +0200:\n> I personally repack everything way more often than is necessary, and I had \n> kind of assumed that people did it that way, but I was apparently wrong. \n> Comments?\n\nI do it from time to time. Seldom in working repositories, because\nthey usually come and go before they have a chance to accumulate\nenough of loose objects. I do a partial repack (git repack -d) after\nevery import from p4 repo, because every snapshot of it is an ugly\nmess changing files all over the tree. Sometimes, after I merged a big\nchunk with the p4 repo and sent it over (the process involves rebase).\n\nIt is usually concious decision when to do a repack or gc. The repack\ntime is seldom a problem: it is fast enough even on windows (and I do\nhave big repos and binary objects). The gc causes my machines to swap,\nthough. Some of them heavily, so there my repos stay longer partially\npacked. I do use .keep packs for this reason (and because windows or\ncygwin or both have more problems with big files the they have with\nsmall).\n\nI used to clone repos with \"-s\", but quickly stopped after a few\nbroken histories.  This also tought me to think before running\n\"git gc\" or \"git repack -a -d\".\n\nOn a rare occurance I even use \"git repack -a -d -l\" and \"git\npack-refs\" separately.\n\nThis was all specific to my day-job. At home, on linux systems I just\nrun git-gc whenever I please, without even thinking why. It finishes\nmostly in less than a minute (the kernel: ~40-50 sec on my P4 2.6GHz, 1Gb).\n"},{"id":"52635","messageId":"87ps0wzufo.fsf@hades.wkstn.nix","threadId":"9781","inReplyTo":"alpine.LFD.0.9999.0709051634190.21186@xanadu.home","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Nix","fromEmail":"nix@esperi.org.uk","sentAt":"2007-09-05T21:14:19Z","receivedAt":"2007-09-05T21:14:19Z","isPatch":false,"sender":{"key":"nix@esperi.org.uk","avatar":"https://avatars.githubusercontent.com/u/6503005?v=4"},"body":"On 5 Sep 2007, Nicolas Pitre stated:\n\n> On Wed, 5 Sep 2007, Junio C Hamano wrote:\n>> \tgit fetch\n>\n> I think that would be a much better idea to simply decrease the \n> fetch.unpackLimit default value.\n\nI think `git fetch' works reasonably well as is: unless you're fetching\nevery five minutes you often find you get packs anyway. There's no point\npacking incrementally *too* often, or you replace a lots-of-objects\nproblem with a lots-of-packs problem, after which you're worse off than\nwhen you started.\n"},{"id":"52636","messageId":"20070905211838.GB3770@steel.home","threadId":"9781","inReplyTo":"7vr6lcj2zi.fsf@gitster.siamese.dyndns.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2007-09-05T21:18:38Z","receivedAt":"2007-09-05T21:18:38Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"Junio C Hamano, Wed, Sep 05, 2007 22:01:37 +0200:\n> +\t/*\n> +\t * Quickly check if a \"gc\" is needed, by estimating how\n> +\t * many loose objects there are.  Because SHA-1 is evenly\n> +\t * distributed, we can check only one and get a reasonable\n> +\t * estimate.\n> +\t */\n\n:))\n\n> +\tif (sizeof(path) <= snprintf(path, sizeof(path), \"%s/17\", objdir)) {\n> +\t\twarning(\"insanely long object directory %.*s\", 50, objdir);\n\nor a non-POSIX snprintf returning \"negative value\" (Microsoft)\n"},{"id":"52640","messageId":"7v1wdciy3w.fsf@gitster.siamese.dyndns.org","threadId":"9781","inReplyTo":"alpine.LFD.0.9999.0709051634190.21186@xanadu.home","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-09-05T21:46:59Z","receivedAt":"2007-09-05T21:46:59Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nicolas Pitre <nico@cam.org> writes:\n\n>> This patch does not add invocation of the \"auto repacking\".  It\n>> is left to key Porcelain commands that could produce tons of\n>> loose objects to add a call to \"git gc --auto\" after they are\n>> done their work.  Obvious candidates are:\n>> \n>> \tgit add\n>\n> Nope!  'git add' creates loose objects which are not yet reachable from \n> anywhere.  They won't get repacked until a commit is made.\n\nBzzt, I am releaved to see you are sometimes wrong ;-)\n\nThey are reachable from the index and are not subject to\npruning.\n\n>> \tgit fetch\n>\n> I think that would be a much better idea to simply decrease the \n> fetch.unpackLimit default value.\n\nOne thing that I find lacking in that auto patch is actually\nthat we should sometimes consolidate multiple small packs into a\nsingle larger one.  Any behaviour change to encourage creation\nof many tiny packs should be avoided until it materializes.\n\nProbably we should introduce a built-in minimum value for a\npositive gc.auto, somewhere around 1000 or so, for this reason.\n"},{"id":"52641","messageId":"7vwsv4hjfi.fsf@gitster.siamese.dyndns.org","threadId":"9781","inReplyTo":"alpine.LFD.0.9999.0709051634190.21186@xanadu.home","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-09-05T21:49:21Z","receivedAt":"2007-09-05T21:49:21Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nicolas Pitre <nico@cam.org> writes:\n\n> and git commit.  Which resumes it to commit creating operation.\n\nGood point.  I think that makes sense.\n"},{"id":"52642","messageId":"7vmyw0hixs.fsf_-_@gitster.siamese.dyndns.org","threadId":"9781","inReplyTo":"7vwsv4hjfi.fsf@gitster.siamese.dyndns.org","subject":"Invoke \"git gc --auto\" from commit, merge, am and rebase.","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-09-05T21:59:59Z","receivedAt":"2007-09-05T21:59:59Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"The point of auto gc is to pack new objects created in loose\nformat, so a good rule of thumb is where we do update-ref after\ncreating a new commit.\n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n  Let's chuck the previous \"git add/git fetch\" one, and replace it\n  with this.\n\n  Also I realize I misread your earlier comment about \"git add\".\n  You are still among the only few people on the list that I\n  consider are always more right than I am ;-).\n\n git-am.sh                  |    2 ++\n git-commit.sh              |    1 +\n git-merge.sh               |    1 +\n git-rebase--interactive.sh |    2 ++\n 4 files changed, 6 insertions(+), 0 deletions(-)\n\ndiff --git a/git-am.sh b/git-am.sh\nindex 6809aa0..4db4701 100755\n--- a/git-am.sh\n+++ b/git-am.sh\n@@ -466,6 +466,8 @@ do\n \t\t\"$GIT_DIR\"/hooks/post-applypatch\n \tfi\n \n+\tgit gc --auto\n+\n \tgo_next\n done\n \ndiff --git a/git-commit.sh b/git-commit.sh\nindex 1d04f1f..d22d35e 100755\n--- a/git-commit.sh\n+++ b/git-commit.sh\n@@ -652,6 +652,7 @@ git rerere\n \n if test \"$ret\" = 0\n then\n+\tgit gc --auto\n \tif test -x \"$GIT_DIR\"/hooks/post-commit\n \tthen\n \t\t\"$GIT_DIR\"/hooks/post-commit\ndiff --git a/git-merge.sh b/git-merge.sh\nindex 3a01db0..697bec2 100755\n--- a/git-merge.sh\n+++ b/git-merge.sh\n@@ -82,6 +82,7 @@ finish () {\n \t\t\t;;\n \t\t*)\n \t\t\tgit update-ref -m \"$rlogm\" HEAD \"$1\" \"$head\" || exit 1\n+\t\t\tgit gc --auto\n \t\t\t;;\n \t\tesac\n \t\t;;\ndiff --git a/git-rebase--interactive.sh b/git-rebase--interactive.sh\nindex abc2b1c..8258b7a 100755\n--- a/git-rebase--interactive.sh\n+++ b/git-rebase--interactive.sh\n@@ -307,6 +307,8 @@ do_next () {\n \trm -rf \"$DOTEST\" &&\n \twarn \"Successfully rebased and updated $HEADNAME.\"\n \n+\tgit gc --auto\n+\n \texit\n }\n \n"},{"id":"52643","messageId":"alpine.LFD.0.9999.0709051858060.21186@xanadu.home","threadId":"9781","inReplyTo":"7v1wdciy3w.fsf@gitster.siamese.dyndns.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-09-05T23:04:27Z","receivedAt":"2007-09-05T23:04:27Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 5 Sep 2007, Junio C Hamano wrote:\n\n> Nicolas Pitre <nico@cam.org> writes:\n> \n> >> This patch does not add invocation of the \"auto repacking\".  It\n> >> is left to key Porcelain commands that could produce tons of\n> >> loose objects to add a call to \"git gc --auto\" after they are\n> >> done their work.  Obvious candidates are:\n> >> \n> >> \tgit add\n> >\n> > Nope!  'git add' creates loose objects which are not yet reachable from \n> > anywhere.  They won't get repacked until a commit is made.\n> \n> Bzzt, I am releaved to see you are sometimes wrong ;-)\n> \n> They are reachable from the index and are not subject to\n> pruning.\n\nThe index?  What's that?  ;-)\n\n> >> \tgit fetch\n> >\n> > I think that would be a much better idea to simply decrease the \n> > fetch.unpackLimit default value.\n> \n> One thing that I find lacking in that auto patch is actually\n> that we should sometimes consolidate multiple small packs into a\n> single larger one.  Any behaviour change to encourage creation\n> of many tiny packs should be avoided until it materializes.\n> \n> Probably we should introduce a built-in minimum value for a\n> positive gc.auto, somewhere around 1000 or so, for this reason.\n\nWhy not just let the default value take care of it?  If someone really \nwants to set gc.auto to 50, why prevent it?\n\nThe more I think of it, the less I like automatic repack.  There is \nalways a bad case for it somewhere.\n\n\nNicolas\n"},{"id":"52644","messageId":"7v3axshe6q.fsf@gitster.siamese.dyndns.org","threadId":"9781","inReplyTo":"alpine.LFD.0.9999.0709051858060.21186@xanadu.home","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-09-05T23:42:37Z","receivedAt":"2007-09-05T23:42:37Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nicolas Pitre <nico@cam.org> writes:\n\n> On Wed, 5 Sep 2007, Junio C Hamano wrote:\n>\n>> Nicolas Pitre <nico@cam.org> writes:\n>> \n> The index?  What's that?  ;-)\n\nSorry, my mistake.  You are always more right than I am [tm] ;-)\n\n> The more I think of it, the less I like automatic repack.  There is \n> always a bad case for it somewhere.\n\nI tend to agree, but at the same time, I think the long term\ngoal should be not to have bad cases.\n\nOld timers like ourselves learned to run \"repack -a -d\" when not\ndoing real work (i.e. beginning of the day while fetching\ncoffee, before leaving to lunch break, end of the day before\nleaving) and we have been _trained_ not to feel that a choir,\nbut I think that is wrong.  \"Sync freezes I/O for and causes my\nreal-time databasy job undue latency --- I would want to disable\nswapper/bdflush/whatever machine-wide and prefer typing 'sync'\nfrom the command line when it is convenient for me\" is fine for\nan experienced user working on a single user machine, but it\nstill feels wrong (we do not have \"multi-user\" issues in git\nrepository, so this analogy is not quite right, though).\n"},{"id":"52650","messageId":"20070905235653.GB25001@coredump.intra.peff.net","threadId":"9781","inReplyTo":"vpq1wddkohr.fsf@bauges.imag.fr","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-09-05T23:56:53Z","receivedAt":"2007-09-05T23:56:53Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Sep 05, 2007 at 07:31:44PM +0200, Matthieu Moy wrote:\n\n> I have ~/teaching/some-course/.git (well, almost) and ~/etc/.git which\n> are two unrelated projects, and to \"git gc\" both of them, I need\n> either a script, or two manual invocations.\n>\n> (yes, I'm really talking about something trivial)\n\nI tend to have a lot of small projects, so I have on the order of 80 git\nrepositories on each machine I use, most of which have a 'mothership'\norigin on a central, backed-up machine.\n\nWhen I sit down to work, I want to see which repositories\nhave changes that need to be pulled. And when I get up to leave, I want\nto see which repositories have changes that need to be pushed. Not to\nmention files that need committed, loose objects that need packed, etc.\n\nSo I wrote the 'git-stale' script, included below. It's not especially\nuser-friendly, but you might find it useful, as it solves the exact\nproblem you are talking about (and much more).\n\nIt reads 'repository specifications' from ~/.gitstale, one per line,\nwhich are either of the form:\n\n  /path/to/repo\n\nwhich specifies a repo to check, or:\n\n  r:/path/to/many/repos\n\nwhich specifies a hierarchy in which to recursively find repos.\n\nMy .gitstale looks something like this:\n\n  /home/peff/compile/git\n  /home/peff/compile/tig\n  r:/home/peff/work\n\nand I get output something like this (edited for brevity):\n\nChecking (1/77) /home/peff/compile/git...\nChecking (2/77) /home/peff/compile/tig...\n[...]\nChecking (77/77) /home/peff/work/foo...\nMERGE:next /home/peff/compile/git\nCOMMIT: /home/peff/work/foo\nPACK: /home/peff/work/foo\nPUSH:master /home/peff/work/bar\n\nwhich translates to:\n  - the git repo has commits in 'origin/next' that are not in 'next'\n    (and you might want to merge them in)\n  - there are uncommitted files in 'foo'\n  - 'foo' needs packing\n  - in the 'bar' repo there are commits in master that are not in origin\n    (and you might want to push)\n\nHopefully it will be useful to you, though I think it is probably too\nspecific to my workflow to be part of git.\n\n-Peff\n\n-- >8 --\n#!/usr/bin/perl\n\nuse strict;\nuse Getopt::Long;\n\nmy $CONFIG_FILE = \"$ENV{HOME}/.gitstale\";\n\nmy $nofetch = $ENV{GITSTALE_NOFETCH};\nGetopt::Long::Configure(qw(bundling));\nGetOptions('nofetch|n!' => \\$nofetch) or exit 100;\n\nmy @projects = process_spec(@ARGV ? @ARGV : cat($CONFIG_FILE));\n\nmy $n = 1;\nmy $total = @projects;\nmy %errors;\nforeach my $p (@projects) {\n  print \"Checking ($n/$total) $p...\\n\";\n  $errors{$p} = [check_git($p)];\n  $n++;\n}\n\nmy $errcount;\nforeach my $p (@projects) {\n  foreach my $e (@{$errors{$p}}) {\n    print \"$e: $p\\n\";\n  }\n}\n\nexit $errcount ? 1 : 0;\n\nsub cat {\n  my $fn = shift;\n  open(my $fh, '<', $fn)\n    or die \"unable to open $fn: $!\\n\";\n  return map { chomp; length($_) ? $_ : () } <$fh>;\n}\n\nsub process_spec {\n  my @dirs;\n  my @roots;\n  my @exclude;\n\n  foreach (@_) {\n    if(/^r:(.*)/) { push @roots, $1 }\n    elsif(/^d:(.*)/) { push @dirs, $1 }\n    elsif(/^-(.*)/) { push @exclude, qr#(^|/)$1($|/)# }\n    else { push @dirs, $_ }\n  }\n\n  use File::Find;\n  find({\n      no_chdir => 1,\n      preprocess => sub { sort @_ },\n      wanted => sub {\n        return unless -d $_ && $_ =~ m#/.git$#;\n        foreach my $e (@exclude) { return if $_ =~ $e }\n        my $d = $_;\n        $d =~ s#/\\.git$##;\n        push @dirs, $d;\n      }\n    }, @roots) if @roots;\n  return @dirs;\n}\n\nsub count_zero {\n  open(my $fh, '-|', @_) or die \"unable to fork: $!\\n\";\n  my $line = <$fh>;\n  return length($line) == 0;\n}\n\nsub check_git {\n  my $d = shift;\n\n  chdir($d) or return 'CHDIR';\n\n  my @r;\n  count_zero(qw(\n        git-ls-files -m -o -d --exclude-per-directory=.gitignore\n        --directory --no-empty-directory\n  )) or push @r, 'COMMIT';\n\n  if(has_origin()) {\n    push @r, 'FETCH' if !$nofetch && system('git-fetch');\n\n    foreach my $p (branch_pairs()) {\n      count_zero('git-rev-list', \"$p->[0]..$p->[1]\")\n        or push @r, \"MERGE:$p->[0]\";\n      count_zero('git-rev-list', \"$p->[1]..$p->[0]\")\n        or push @r, \"PUSH:$p->[0]\";\n    }\n  }\n  else {\n    push @r, 'ORIGIN';\n  }\n\n  push @r, 'PACK' if unpacked_objects() > 1000;\n\n  return @r;\n}\n\nsub unpacked_objects {\n  my $objects = `git-count-objects`;\n  $objects =~ /^(\\d+)/;\n  return $1;\n}\n\nsub branch_pairs {\n  my %config;\n  foreach my $line (`git-repo-config --get-regexp 'branch..*..*'`) {\n    $line =~ m#^branch\\.([^.]+)\\.([^ ]+) (?:refs/heads/)?(.*)#\n      or die \"confusing git-repo-config output: $line\\n\";\n    $config{$1}{$2} = $3;\n  }\n\n  return [qw(master origin)] if -e '.git/refs/heads/origin';\n\n  return\n    (-e '.git/refs/heads/origin' ? [qw(master origin)] : ()),\n    map {\n      $config{$_}{remote} && $config{$_}{merge} ?\n        [$_, $config{$_}{remote} . '/' . $config{$_}{merge}] :\n        ()\n    } sort keys(%config);\n}\n\nsub has_origin {\n  return\n    -e '.git/branches/origin' ||\n    -e '.git/remotes/origin' ||\n    !count_zero(qw(git-repo-config --get remote.origin.url));\n}\n__END__\n"},{"id":"52652","messageId":"1b46aba20709051727od0644d7t16eaa348af86952a@mail.gmail.com","threadId":"9781","inReplyTo":"7v3axshe6q.fsf@gitster.siamese.dyndns.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Carlos Rica","fromEmail":"jasampler@gmail.com","sentAt":"2007-09-06T00:27:05Z","receivedAt":"2007-09-06T00:27:05Z","isPatch":false,"sender":{"key":"jasampler@gmail.com","avatar":null},"body":"2007/9/6, Junio C Hamano <gitster@pobox.com>:\n> Nicolas Pitre <nico@cam.org> writes:\n> > The more I think of it, the less I like automatic repack.  There is\n> > always a bad case for it somewhere.\n>\n> I tend to agree, but at the same time, I think the long term\n> goal should be not to have bad cases.\n\nThe best solution is make \"git gc\" unnecessary.\nAt the long term, and without loss of efficiency.\n"},{"id":"52675","messageId":"20070906023934.GI18160@spearce.org","threadId":"9781","inReplyTo":"7vmyw0hixs.fsf_-_@gitster.siamese.dyndns.org","subject":"Re: Invoke \"git gc --auto\" from commit, merge, am and rebase.","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-09-06T02:39:34Z","receivedAt":"2007-09-06T02:39:34Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Junio C Hamano <gitster@pobox.com> wrote:\n> The point of auto gc is to pack new objects created in loose\n> format, so a good rule of thumb is where we do update-ref after\n> creating a new commit.\n...\n>  git-am.sh                  |    2 ++\n>  git-commit.sh              |    1 +\n>  git-merge.sh               |    1 +\n>  git-rebase--interactive.sh |    2 ++\n>  4 files changed, 6 insertions(+), 0 deletions(-)\n...\n> diff --git a/git-rebase--interactive.sh b/git-rebase--interactive.sh\n> index abc2b1c..8258b7a 100755\n> --- a/git-rebase--interactive.sh\n> +++ b/git-rebase--interactive.sh\n> @@ -307,6 +307,8 @@ do_next () {\n>  \trm -rf \"$DOTEST\" &&\n>  \twarn \"Successfully rebased and updated $HEADNAME.\"\n>  \n> +\tgit gc --auto\n> +\n>  \texit\n>  }\n\nWhy bother with git-rebase--interactive.sh?  It calls two tools,\ngit-cherry-pick (which calls git-commit) and git-commit to do its\nper-commit dirty work.  So on every step of `git rebase -i` we are\nnow running `git gc --auto`.  No need to also run it at the end.\n\nNote this is also true of `git rebase -m` as that uses the wonderful\nfeature of `git commit -C $oldid` per commit to make the new commit.\n \n-- \nShawn.\n"},{"id":"52676","messageId":"loom.20070906T044017-727@post.gmane.org","threadId":"9781","inReplyTo":"7vr6lcj2zi.fsf@gitster.siamese.dyndns.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Russ Dill","fromEmail":"russ.dill@gmail.com","sentAt":"2007-09-06T02:44:00Z","receivedAt":"2007-09-06T02:44:00Z","isPatch":false,"sender":{"key":"russ.dill@gmail.com","avatar":"https://gravatar.com/avatar/989b24f3fa63126a35d7c74069e2626e715a3b1a86616be9962f7ceae04ff9c5?d=mp&s=160"},"body":"\n> Ok, how about doing something like this?\n> \n\ngit add? merge? rebase? No, I have a sneakier place to invoke gc.\n\nWhenever $EDITOR gets invoked. Heck, whenever git is waiting for any user input,\ndo some gc in the background, it'd just have to be incremental so that we could\npick up where we left off.\n\nSimilarly, you could mix it in with git pull/push so that while we are waiting\non the network, we can do some packing.\n\nCourse, this wouldn't work for all repositories.\n"},{"id":"52677","messageId":"20070906024555.GJ18160@spearce.org","threadId":"9781","inReplyTo":"7vr6lcj2zi.fsf@gitster.siamese.dyndns.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-09-06T02:45:55Z","receivedAt":"2007-09-06T02:45:55Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Junio C Hamano <gitster@pobox.com> wrote:\n> Implement git gc --auto\n... \n\nDanger...  If the user sets `gc.auto` to a low enough value and\nthey are also unlucky enough to have a few truely unreachable (thus\npruneable) objects in .git/objects/17/ then this is going to run\na bunch of gc work on every commit they make.\n\nI'm actually running into this problem in git-gui.  On Windows\nit suggests a repack if there is one object in .git/objects/42/.\nSome users have been unlucky enough to stage a file, have it\nhash into that directory, then restage a different version of it.\nThe prior one is never considered reachable (it was never committed),\nbut will now *always* cause git-gui to suggest a repack on every\nstartup.  For all time.\n\nYea, I need to fix that.\n\nBut this suffers from the same fate if the user sets gc.auto too\nsmall and doesn't realize that the reason Git is always repacking\nis because over the last 6 months they have been unlucky enough to\nstage the magic number of unreachable blobs into the 17 directory\nand they have *never* run `git gc --prune` because the auto thing\nis working just fine for them and they don't realize they need to\nprune every once in a blue moon.\n\n-- \nShawn.\n"},{"id":"52678","messageId":"46DF6AA8.60804@midwinter.com","threadId":"9781","inReplyTo":"20070906024555.GJ18160@spearce.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Steven Grimm","fromEmail":"koreth@midwinter.com","sentAt":"2007-09-06T02:49:12Z","receivedAt":"2007-09-06T02:49:12Z","isPatch":false,"sender":{"key":"koreth@midwinter.com","avatar":"https://gravatar.com/avatar/71b4d2e8b62f168bdc9e9205341159e3567003b4f9e2127c617c5fa0a1f5bad2?d=mp&s=160"},"body":"Shawn O. Pearce wrote:\n> But this suffers from the same fate if the user sets gc.auto too\n> small and doesn't realize that the reason Git is always repacking\n> is because over the last 6 months they have been unlucky enough to\n> stage the magic number of unreachable blobs into the 17 directory\n> and they have *never* run `git gc --prune` because the auto thing\n> is working just fine for them and they don't realize they need to\n> prune every once in a blue moon.\n>   \n\nCheck the modification times on those files and don't count ones that \nare older than the last git-gc run, maybe? That'd take care of the problem.\n\n-Steve\n"},{"id":"52679","messageId":"20070906025224.GK18160@spearce.org","threadId":"9781","inReplyTo":"loom.20070906T044017-727@post.gmane.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-09-06T02:52:24Z","receivedAt":"2007-09-06T02:52:24Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Russ Dill <Russ.Dill@gmail.com> wrote:\n> > Ok, how about doing something like this?\n> \n> git add? merge? rebase? No, I have a sneakier place to invoke gc.\n> \n> Whenever $EDITOR gets invoked. Heck, whenever git is waiting for any user input,\n> do some gc in the background, it'd just have to be incremental so that we could\n> pick up where we left off.\n\nHeh.  That is a really good idea.  I've been thinking about doing\nsome automatic generational style GC type repacking controls in\ngit-gui, and doing them when git-gui is sitting idle and has not\nbeen used in the past couple of minutes.\n\nThis is along the same vein of thought.  I like it.  Often it\ntakes me a while to come up with a good commit message even if\nI am using command line commit.\n\nBut git-rebase/git-am can cause a huge number of objects to be\ncreated, especially if you are pushing a large stack of patches\naround.  So it may still be a good idea to trigger `gc --auto`\nat the end of those operations.\n \n> Similarly, you could mix it in with git pull/push so that while we are waiting\n> on the network, we can do some packing.\n\nHere's a better thought:\n\nIf we are pushing somewhere, and the push size is \"large-ish\" and\nwe aren't pushing a thin pack (its currently considered not nice\nto the remote end so it doesn't happen by default) and the objects\nwe are packing are mostly all loose maybe we should also save a\ncopy of that packfile locally, then prune *only* those loose objects\nback.\n\nNot every git user pushes their work.  But many do.  And those\nthat push usually will do so in bursts, are already expecting to\nwait for the network latency, and usually are pushing the majority\nof the things that are loose.  Such users will probably never see\nthe `gc --auto` trip in places like commit/am/merge as they would\nalready be clearing their ODB with the push.\n\n-- \nShawn.\n"},{"id":"52680","messageId":"20070906025640.GL18160@spearce.org","threadId":"9781","inReplyTo":"46DF6AA8.60804@midwinter.com","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-09-06T02:56:40Z","receivedAt":"2007-09-06T02:56:40Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Steven Grimm <koreth@midwinter.com> wrote:\n> Shawn O. Pearce wrote:\n> >But this suffers from the same fate if the user sets gc.auto too\n> >small and doesn't realize that the reason Git is always repacking\n> >is because over the last 6 months they have been unlucky enough to\n> >stage the magic number of unreachable blobs into the 17 directory\n> >and they have *never* run `git gc --prune` because the auto thing\n> >is working just fine for them and they don't realize they need to\n> >prune every once in a blue moon.\n> \n> Check the modification times on those files and don't count ones that \n> are older than the last git-gc run, maybe? That'd take care of the problem.\n\nEh, that could mean a bunch of stat calls that it would be nice\nto avoid.  The counter Junio (and git-gui) implements just does\na readdir().  Reasonably cheap.\n\nMaybe just save a \".git/gc_last_auto\" with the last object count\nof .git/objects/17, after repacking.  If the count is over the\ngc.auto limit *and* is still over the limit after subtracting the\n\".git/gc_last_auto\" value then consider that auto is required.\n\nThis way the file is only consulted if we are really thinking\nabout running a repack, and its only written to if we actually do\nthe repack.  So we only take the extra penalty if we are going to\nbe taking a *really* big extra penalty by repacking.\n\n-- \nShawn.\n"},{"id":"52694","messageId":"85veaofid9.fsf@lola.goethe.zz","threadId":"9781","inReplyTo":"7v1wdciy3w.fsf@gitster.siamese.dyndns.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-06T05:55:14Z","receivedAt":"2007-09-06T05:55:14Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Nicolas Pitre <nico@cam.org> writes:\n>\n>>> This patch does not add invocation of the \"auto repacking\".  It\n>>> is left to key Porcelain commands that could produce tons of\n>>> loose objects to add a call to \"git gc --auto\" after they are\n>>> done their work.  Obvious candidates are:\n>>> \n>>> \tgit add\n>>\n>> Nope!  'git add' creates loose objects which are not yet reachable from \n>> anywhere.  They won't get repacked until a commit is made.\n>\n> Bzzt, I am releaved to see you are sometimes wrong ;-)\n>\n> They are reachable from the index and are not subject to\n> pruning.\n\nHm.  Isn't it possible to work with several index files at once?  I\nseem to remember that even git-add does this itself.  So what is it\nthat protects objects in such a temporary index from being garbage\ncollected by a different git process running on the same repository?\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"52718","messageId":"46DFC848.60201@op5.se","threadId":"9781","inReplyTo":"loom.20070906T044017-727@post.gmane.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2007-09-06T09:28:40Z","receivedAt":"2007-09-06T09:28:40Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Russ Dill wrote:\n>> Ok, how about doing something like this?\n>>\n> \n> git add? merge? rebase? No, I have a sneakier place to invoke gc.\n> \n> Whenever $EDITOR gets invoked. Heck, whenever git is waiting for any user input,\n> do some gc in the background, it'd just have to be incremental so that we could\n> pick up where we left off.\n> \n\nI like it. Writing a commit-message takes anywhere from 30 seconds to 5 minutes\nfor me (sometimes having to check up bug id's, or verifying details in the code).\nSneaking in a repack here would be absolutely stellar :)\n\nIt's also nice in that it won't affect people who just follow a project's tip to\nget the bleeding edge. For them it shouldn't matter much that they have multiple\nsmall packs obtained while fetching, or if it's all bungled together in a big one.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"52742","messageId":"Pine.LNX.4.64.0709061301250.28586@racer.site","threadId":"9781","inReplyTo":"7vhcm8j1bp.fsf_-_@gitster.siamese.dyndns.org","subject":"Re: [PATCH] Invoke \"git gc --auto\" from \"git add\" and \"git fetch\"","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-09-06T12:02:09Z","receivedAt":"2007-09-06T12:02:09Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 5 Sep 2007, Junio C Hamano wrote:\n\n>  * This is obviously a follow-up to the previous one that allows\n>    you to say \"git gc --auto\".  I somewhat feel dirty about\n>    calling cmd_gc() bypassing fork & exec from \"git add\",\n>    though...\n\nSince all git-gc seems to do is to fork() and exec() other git programs, \nthis should be fine (have not looked at cmd_gc() in a while, though).\n\nCiao,\nDscho\n"},{"id":"52765","messageId":"Pine.LNX.4.64.0709061651550.28586@racer.site","threadId":"9781","inReplyTo":"7vr6lcj2zi.fsf@gitster.siamese.dyndns.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-09-06T15:54:56Z","receivedAt":"2007-09-06T15:54:56Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 5 Sep 2007, Junio C Hamano wrote:\n\n> @@ -20,6 +20,7 @@ static const char builtin_gc_usage[] = \"git-gc [--prune] [--aggressive]\";\n>  \n>  static int pack_refs = 1;\n>  static int aggressive_window = -1;\n> +static int gc_auto_threshold = 6700;\n\nPlease don't do that.\n\nWhen you share objects with another git directory, git-gc --auto can get \nrid of the objects when some objects go away in the referenced repository.  \n\nSo we need _at least_ check gc.auto not being set in the repo when \"git \nclone --share\"ing it (and fail otherwise).\n\nMy preferred way would be to set it in \"git init\" so that existing setups \nare not affected, and put some big red message on top of the next release \nnotes that people might want to set gc.auto in their existing setups.\n\nCiao,\nDscho\n"},{"id":"52786","messageId":"7vk5r3adlx.fsf@gitster.siamese.dyndns.org","threadId":"9781","inReplyTo":"Pine.LNX.4.64.0709061651550.28586@racer.site","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-09-06T17:49:14Z","receivedAt":"2007-09-06T17:49:14Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> On Wed, 5 Sep 2007, Junio C Hamano wrote:\n>\n>> @@ -20,6 +20,7 @@ static const char builtin_gc_usage[] = \"git-gc [--prune] [--aggressive]\";\n>>  \n>>  static int pack_refs = 1;\n>>  static int aggressive_window = -1;\n>> +static int gc_auto_threshold = 6700;\n>\n> Please don't do that.\n>\n> When you share objects with another git directory, git-gc --auto can get \n> rid of the objects when some objects go away in the referenced repository.  \n\nI thought the whole point of \"gc --auto\" was to have something\nthat does not lose/prune any objects, even the ones that do not\nseem to be referenced from anywhere.  That is why invocations of\n\"git gc --auto\" do not say --prune as you saw the second patch,\nand the repack command \"gc --auto\" runs is \"repack -d -l\"\ninstead of \"repack -a -d -l\", which means that it does run\ngit-prune-packed after repacking but not git-prune.\n\nMaybe I am missing something...\n"},{"id":"52793","messageId":"alpine.LFD.0.999.0709061906010.5626@evo.linux-foundation.org","threadId":"9781","inReplyTo":"7vk5r3adlx.fsf@gitster.siamese.dyndns.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-09-06T18:15:58Z","receivedAt":"2007-09-06T18:15:58Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 6 Sep 2007, Junio C Hamano wrote:\n> \n> I thought the whole point of \"gc --auto\" was to have something\n> that does not lose/prune any objects, even the ones that do not\n> seem to be referenced from anywhere.  That is why invocations of\n> \"git gc --auto\" do not say --prune as you saw the second patch,\n> and the repack command \"gc --auto\" runs is \"repack -d -l\"\n> instead of \"repack -a -d -l\", which means that it does run\n> git-prune-packed after repacking but not git-prune.\n\nI think \"repack -d -l\" should be ok from a safety perspective, but I'd \nalso like to say that always running it incrementally is going to largely \nsuck after a time.\n\nIOW, if you get lots of small incrmental packs, after a while you really \n*do* need to do \"git gc\" to get the real pack generated.\n\nIn the case I saw, James really had hundreds of pack-files. That makes all \nour object lookups suck. Yes, not having loose objects at all is a big \ndeal too, and yes, we try to start from the last pack-file we found (for \nthe locality that we hope is there), but it's still pretty bad from a \ncache usage standpoint, and when we create a new object, we'll first \nsearch (in vain) in all the hundreds of pack-files.\n\nSo would \"git gc --auto\" have helped James? I'm sure it would have. But he \nalready had lots of pack-files from doing \"git fetch/pull\", and while \ndoing the \"git gc --auto\" will likely *delay* the point where you need to \ndo a full repack, it doesn't make it go away.\n\nWe still need to tell people to do a full git gc at some point, or do it \nfor them. And the longer you delay doing it, the more expensive it's going \nto get to do and/or the worse the final packing is going to be (especially \nif it ends up reusing non-optimal packing decisions from the smaller \npacks).\n\nSo I think the --auto stuff is still worth it, but it's really just \npushing the pain somewhat further out.\n\n(In the kernel community, if you fetch my tree daily, you really *are* \ngoing to have hundreds and hundreds of packfiles just from doing that).\n\nSo I'd really like us to also remind people to do a *real* and full \"git \ngc\", not just the incremental ones.\n\n\t\tLinus\n"},{"id":"52796","messageId":"46E046F3.9000703@midwinter.com","threadId":"9781","inReplyTo":"alpine.LFD.0.999.0709061906010.5626@evo.linux-foundation.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Steven Grimm","fromEmail":"koreth@midwinter.com","sentAt":"2007-09-06T18:29:07Z","receivedAt":"2007-09-06T18:29:07Z","isPatch":false,"sender":{"key":"koreth@midwinter.com","avatar":"https://gravatar.com/avatar/71b4d2e8b62f168bdc9e9205341159e3567003b4f9e2127c617c5fa0a1f5bad2?d=mp&s=160"},"body":"Linus Torvalds wrote:\n> IOW, if you get lots of small incrmental packs, after a while you really \n> *do* need to do \"git gc\" to get the real pack generated.\n>   \n\nI wonder if it makes sense to repack just the small incremental packs \ninto a large (but still incremental) pack, rather than repacking the \nentire repository. Presumably that would be a lot faster than a full \n\"git gc\", while still giving you reasonably good packing (at least, if \nthe threshold is set to a hugh enough number of small packs) and keeping \nthings fast. That could run as a second phase of \"git gc --auto\" -- it \nshould be quick enough to not be too terribly annoying since we're not \nrunning it in the background.\n\nYeah, if you use the same repo for a long time, you'll accumulate a ton \nof medium-sized packs this way, but (a) that's much better than the \nsituation we have today, and (b) it puts off the performance degradation \nfor long enough that it becomes more reasonable to expect people to find \nout about running the full \"git gc\" in the meantime, or for git to \nfurther evolve to not need it.\n\n-Steve\n"},{"id":"52814","messageId":"7v1wdb9ymf.fsf_-_@gitster.siamese.dyndns.org","threadId":"9781","inReplyTo":"alpine.LFD.0.999.0709061906010.5626@evo.linux-foundation.org","subject":"Subject: [PATCH] git-merge-pack","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-09-06T23:12:56Z","receivedAt":"2007-09-06T23:12:56Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"This is a beginning of \"git-merge-pack\" that combines smaller\npacks into one.  Currently it does not actually create a new\npack, but pretends that it is a (dumb) \"git-rev-list --objects\"\nthat lists the objects in the affected packs.  You have to pipe\nits output to \"git-pack-objects\".\n\nThe command reads names of pack-*.pack files from the standard\ninput, outputs the objects' names in the order they are stored\nin the original packs (i.e. the offset order).  This sorting is\ndone in order to emulate the traversal order the original\n\"git-rev-list --objects\" that was used to create the existing\npack listed the objects.\n\nWhile this approach would give the resulting packfile very\nsimilar locality of access as the original, it does not give the\n\"name\" component you would see in \"git-rev-list --objects\"\noutput.  This information is used as the clustering cue while\ncomputing delta, and the lack of it means you can get horrible\ndelta selection.  You do _not_ want to run the downstream\n\"git-pack-objects\" without the optimization/heuristics to reuse\ndelta.  IOW, do not run it with --no-reuse-delta.\n\nTo consolidate all packs that are smaller than a megabytes into\none, you would use it in its current form like this:\n\n    $ old=$(find .git/objects/pack -type f -name '*.pack' -size 1M)\n    $ new=$(echo \"$old\" | git merge-pack | git pack-objects pack)\n    $ for p in $old; do rm -f $p ${p%.pack}.idx; done\n    $ for s in pack idx; do mv pack-$new.$s .git/objects/pack/; done\n\nAn obvious next steps that can be done in parallel by interested\nparties would be:\n\n (1) come up with a way to give \"name\" aka \"clustering cue\" (I\n     think this is very hard);\n\n (2) run the above four command sequence internally without\n     having to resort to shell wrapper (easy).\n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n  Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n  > IOW, if you get lots of small incrmental packs, after a while you really \n  > *do* need to do \"git gc\" to get the real pack generated.\n\n  'auto' should do a lessor impact repack than the usual one.\n  Especially we do not want to lose objects that do not look like\n  they are reachable from this reopsitory, to help people with\n  alternate object stores, aka \"repo.or.cz style _forked_\n  repositories\".  However, a full repack with \"-a -d\" discards\n  unreferenced objects that are only in packs.\n\n  We need a middle ground between \"pack and prune-pack only loose\n  ones\" and \"full repack.\n\n  Here is one.\n\n Makefile             |    1 +\n builtin-merge-pack.c |   87 ++++++++++++++++++++++++++++++++++++++++++++++++++\n builtin.h            |    1 +\n git.c                |    1 +\n 4 files changed, 90 insertions(+), 0 deletions(-)\n create mode 100644 builtin-merge-pack.c\n\ndiff --git a/Makefile b/Makefile\nindex dace211..cdff756 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -343,6 +343,7 @@ BUILTIN_OBJS = \\\n \tbuiltin-mailsplit.o \\\n \tbuiltin-merge-base.o \\\n \tbuiltin-merge-file.o \\\n+\tbuiltin-merge-pack.o \\\n \tbuiltin-mv.o \\\n \tbuiltin-name-rev.o \\\n \tbuiltin-pack-objects.o \\\ndiff --git a/builtin-merge-pack.c b/builtin-merge-pack.c\nnew file mode 100644\nindex 0000000..c98da80\n--- /dev/null\n+++ b/builtin-merge-pack.c\n@@ -0,0 +1,87 @@\n+#include \"builtin.h\"\n+#include \"cache.h\"\n+#include \"pack.h\"\n+\n+struct in_pack_object {\n+\toff_t offset;\n+\tconst unsigned char *sha1;\n+};\n+\n+static uint32_t get_packed_object_list(struct packed_git *p, struct in_pack_object *list, uint32_t loc)\n+{\n+\tuint32_t n;\n+\n+\tfor (n = 0; n < p->num_objects; n++) {\n+\t\tlist[loc].sha1 = nth_packed_object_sha1(p, n);\n+\t\tlist[loc].offset = find_pack_entry_one(list[loc].sha1, p);\n+\t\tloc++;\n+\t}\n+\treturn loc;\n+}\n+\n+static int ofscmp(const void *a_, const void *b_)\n+{\n+\tstruct in_pack_object *a = (struct in_pack_object *)a_;\n+\tstruct in_pack_object *b = (struct in_pack_object *)b_;\n+\tif (a->offset < b->offset)\n+\t\treturn -1;\n+\telse if (a->offset > b->offset)\n+\t\treturn 1;\n+\telse\n+\t\treturn hashcmp(a->sha1, b->sha1);\n+}\n+\n+int cmd_merge_pack(int ac, const char **av, const char *prefix)\n+{\n+\tchar filename[PATH_MAX];\n+\tstruct packed_git **pack = NULL;\n+\tint pack_nr = 0;\n+\tint pack_alloc = 0;\n+\tuint32_t max_objs, cnt;\n+\tstruct in_pack_object *objs;\n+\tint i;\n+\n+\twhile (fgets(filename, sizeof(filename), stdin) != NULL) {\n+\t\tint len = strlen(filename);\n+\t\tstruct packed_git *p;\n+\n+\t\twhile (0 < len) {\n+\t\t\tif (filename[len-1] != '\\n' &&\n+\t\t\t    filename[len-1] != '\\r')\n+\t\t\t\tbreak;\n+\t\t\tfilename[--len] = '\\0';\n+\t\t}\n+\t\tif (strcmp(filename + len - 5, \".pack\"))\n+\t\t\tgoto error;\n+\n+\t\t/* add-packed-git wants the name of .idx file */\n+\t\tstrcpy(filename + len - 5, \".idx\");\n+\t\tlen--;\n+\t\tp = add_packed_git(filename, len, 1);\n+\t\tif (!p)\n+\t\t\tgoto error;\n+\t\tif (open_pack_index(p))\n+\t\t\tgoto error;\n+\n+\t\tif (pack_alloc <= pack_nr) {\n+\t\t\tpack_alloc = alloc_nr(pack_nr);\n+\t\t\tpack = xrealloc(pack, pack_alloc * sizeof(*pack));\n+\t\t}\n+\t\tpack[pack_nr++] = p;\n+\t\tcontinue;\n+\terror:\n+\t\tdie(\"Cannot add a pack .idx file: %s\", filename);\n+\t}\n+\n+\tmax_objs = 0;\n+\tfor (i = 0; i < pack_nr; i++)\n+\t\tmax_objs += pack[i]->num_objects;\n+\tobjs = xmalloc(sizeof(*objs) * max_objs);\n+\tcnt = 0;\n+\tfor (i = 0; i < pack_nr; i++)\n+\t\tcnt = get_packed_object_list(pack[i], objs, cnt);\n+\tqsort(objs, cnt, sizeof(*objs), ofscmp);\n+\tfor (cnt = 0; cnt < max_objs; cnt++)\n+\t\tprintf(\"%s\\n\", sha1_to_hex(objs[cnt].sha1));\n+\treturn 0;\n+}\ndiff --git a/builtin.h b/builtin.h\nindex bb72000..aff28ca 100644\n--- a/builtin.h\n+++ b/builtin.h\n@@ -49,6 +49,7 @@ extern int cmd_mailinfo(int argc, const char **argv, const char *prefix);\n extern int cmd_mailsplit(int argc, const char **argv, const char *prefix);\n extern int cmd_merge_base(int argc, const char **argv, const char *prefix);\n extern int cmd_merge_file(int argc, const char **argv, const char *prefix);\n+extern int cmd_merge_pack(int argc, const char **argv, const char *prefix);\n extern int cmd_mv(int argc, const char **argv, const char *prefix);\n extern int cmd_name_rev(int argc, const char **argv, const char *prefix);\n extern int cmd_pack_objects(int argc, const char **argv, const char *prefix);\ndiff --git a/git.c b/git.c\nindex fd3d83c..69e86bc 100644\n--- a/git.c\n+++ b/git.c\n@@ -353,6 +353,7 @@ static void handle_internal_command(int argc, const char **argv)\n \t\t{ \"mailsplit\", cmd_mailsplit },\n \t\t{ \"merge-base\", cmd_merge_base, RUN_SETUP },\n \t\t{ \"merge-file\", cmd_merge_file },\n+\t\t{ \"merge-pack\", cmd_merge_pack },\n \t\t{ \"mv\", cmd_mv, RUN_SETUP | NEED_WORK_TREE },\n \t\t{ \"name-rev\", cmd_name_rev, RUN_SETUP },\n \t\t{ \"pack-objects\", cmd_pack_objects, RUN_SETUP },\n-- \n1.5.3.1.860.g2cce2\n"},{"id":"52819","messageId":"alpine.LFD.0.999.0709070027520.5626@evo.linux-foundation.org","threadId":"9781","inReplyTo":"7v1wdb9ymf.fsf_-_@gitster.siamese.dyndns.org","subject":"Re: Subject: [PATCH] git-merge-pack","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-09-06T23:35:11Z","receivedAt":"2007-09-06T23:35:11Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 6 Sep 2007, Junio C Hamano wrote:\n>\n> This is a beginning of \"git-merge-pack\" that combines smaller\n> packs into one.  Currently it does not actually create a new\n> pack, but pretends that it is a (dumb) \"git-rev-list --objects\"\n> that lists the objects in the affected packs.  You have to pipe\n> its output to \"git-pack-objects\".\n\nOk, so I had to double-check that builtin-pack-objects then deals properly \nwith duplicate object names (which it does seem to do), so maybe it's \nworth adding a comment to that effect.\n\nBut ACK, this seems to be the right thing to do to generate a single \nbigger pack from many smaller ones.\n\n\t\tLinus\n"},{"id":"52829","messageId":"alpine.LFD.0.9999.0709061942320.21186@xanadu.home","threadId":"9781","inReplyTo":"7v1wdb9ymf.fsf_-_@gitster.siamese.dyndns.org","subject":"Re: Subject: [PATCH] git-merge-pack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-09-07T00:51:58Z","receivedAt":"2007-09-07T00:51:58Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 6 Sep 2007, Junio C Hamano wrote:\n\n> This is a beginning of \"git-merge-pack\" that combines smaller\n> packs into one.  Currently it does not actually create a new\n> pack, but pretends that it is a (dumb) \"git-rev-list --objects\"\n> that lists the objects in the affected packs.  You have to pipe\n> its output to \"git-pack-objects\".\n> \n> The command reads names of pack-*.pack files from the standard\n> input, outputs the objects' names in the order they are stored\n> in the original packs (i.e. the offset order).  This sorting is\n> done in order to emulate the traversal order the original\n> \"git-rev-list --objects\" that was used to create the existing\n> pack listed the objects.\n> \n> While this approach would give the resulting packfile very\n> similar locality of access as the original, it does not give the\n> \"name\" component you would see in \"git-rev-list --objects\"\n> output.  This information is used as the clustering cue while\n> computing delta, and the lack of it means you can get horrible\n> delta selection.  You do _not_ want to run the downstream\n> \"git-pack-objects\" without the optimization/heuristics to reuse\n> delta.  IOW, do not run it with --no-reuse-delta.\n\nI wonder if this is the best way to go.  In the context of a really fast \nrepack happening automatically after (or during) user interactive \noperations, the above seems a bit heavyweight and slow to me.\n\nI would have concatenated all packs provided on the command line into a \nsingle one, simply by reading data from existing packs and writing it \nback without any processing at all.  The offset for OBJ_OFS_DELTA is \nrelative so a simple concatenation will just work.\n\nThen the index for that pack can be created just as easily by reading \nexisting pack index files and storing the data into an array of struct \npack_idx_entry, adding the appropriate offset to object offsets, then \ncall write_idx_file().\n\nAll data is read once and written once making it no more costly than a \nsimple file copy.  On the flip side it wouldn't get rid of duplicated \nobjects (I don't know if that matters i.e. if something might break with \nthe same object twice in a pack).\n\n> To consolidate all packs that are smaller than a megabytes into\n> one, you would use it in its current form like this:\n> \n>     $ old=$(find .git/objects/pack -type f -name '*.pack' -size 1M)\n>     $ new=$(echo \"$old\" | git merge-pack | git pack-objects pack)\n>     $ for p in $old; do rm -f $p ${p%.pack}.idx; done\n>     $ for s in pack idx; do mv pack-$new.$s .git/objects/pack/; done\n\nYou might want to move the new pack before removing the old ones though.\n\n> An obvious next steps that can be done in parallel by interested\n> parties would be:\n> \n>  (1) come up with a way to give \"name\" aka \"clustering cue\" (I\n>      think this is very hard);\n\nIt is, and IMHO not worth it.  If you do it separately from the usual \npack-objects process you'll perform extra IO and decompression when \nwalking tree objects just to reconstruct those paths, becoming really \nslow by the context definition I provided above.\n\nIf you really want to do it then the best way might simply to reverse \nyour find result above, in order to use pack-objects as if the larger \npacks, i.e. the ones that you don't want to merge, simply had an \nassociated .keep file.\n\nIn fact, since we want to _also_ perform a repack of loose objects in \nthe context of automatic repacking, I wonder why we wouldn't use that \n--unpacked= argument to also repack smallish packs at the same time in \nonly one pack-objects pass.  Or maybe I'm missing something?\n\n\nNicolas\n"},{"id":"52834","messageId":"7v7in38ce6.fsf@gitster.siamese.dyndns.org","threadId":"9781","inReplyTo":"alpine.LFD.0.9999.0709061942320.21186@xanadu.home","subject":"Re: Subject: [PATCH] git-merge-pack","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-09-07T01:58:25Z","receivedAt":"2007-09-07T01:58:25Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nicolas Pitre <nico@cam.org> writes:\n\n> I wonder if this is the best way to go.  In the context of a really fast \n> repack happening automatically after (or during) user interactive \n> operations, the above seems a bit heavyweight and slow to me.\n\nHonestly, I do not believe in that mode of operation that much.\n\n\"While the user is waiting for the EDITOR\"?\n\nBecause you do not know how much time you will be given before\nyou start, unless\n\n (1) your process can be snapshotted and you can restart at the\n     next chance; or\n\n (2) it is so cheap and you can afford to abort and start over\n     from scratch at the next chance; or\n\n (3) it is so quick that you can simply have the user wait until\n     you are done without adding too much latency to be annoying,\n     when you cannnot finish before the EDITOR come back;\n\nI think that is a false sense of \"ok, we will be able to do\nsomething else in the background meantime\", which is not so\nuseful in practice.\n\n>> An obvious next steps that can be done in parallel by interested\n>> parties would be:\n>> \n>>  (1) come up with a way to give \"name\" aka \"clustering cue\" (I\n>>      think this is very hard);\n>\n> It is, and IMHO not worth it.  If you do it separately from the usual \n> pack-objects process you'll perform extra IO and decompression when \n> walking tree objects just to reconstruct those paths, becoming really \n> slow by the context definition I provided above.\n\nWell, I said \"name\" in quotes because you do _NOT_ have to give\nthe real name.  I was not thinking about doing the actual tree\ntraversal at all.  What you need to do is to come up with a\ntoken that is the same for the objects in the same deltification\nchain so that they cluster together, and that should be doable\nby looking at the delta chain patterns inside a packfile.\n"},{"id":"52835","messageId":"alpine.LFD.0.9999.0709062220450.21186@xanadu.home","threadId":"9781","inReplyTo":"7v7in38ce6.fsf@gitster.siamese.dyndns.org","subject":"Re: Subject: [PATCH] git-merge-pack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-09-07T02:32:01Z","receivedAt":"2007-09-07T02:32:01Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 6 Sep 2007, Junio C Hamano wrote:\n\n> Nicolas Pitre <nico@cam.org> writes:\n> \n> > I wonder if this is the best way to go.  In the context of a really fast \n> > repack happening automatically after (or during) user interactive \n> > operations, the above seems a bit heavyweight and slow to me.\n> \n> Honestly, I do not believe in that mode of operation that much.\n> \n> \"While the user is waiting for the EDITOR\"?\n> \n> Because you do not know how much time you will be given before\n> you start, unless\n> \n>  (1) your process can be snapshotted and you can restart at the\n>      next chance; or\n> \n>  (2) it is so cheap and you can afford to abort and start over\n>      from scratch at the next chance; or\n> \n>  (3) it is so quick that you can simply have the user wait until\n>      you are done without adding too much latency to be annoying,\n>      when you cannnot finish before the EDITOR come back;\n\nI think we have to aim for #3.  \"Automatic\" certainly doesn't imply \"can \nbe slow\".  It should be reasonably instantaneous, otherwise it'll become \nannoying quickly enough.  If it can't be (almost) instantaneous in 99% \nof normal cases, then I think it simply should be remain asynchronously \nthrougha manual invokation of 'git gc' and we only need to teach/remind \npeople about it more strongly.\n\n> >> An obvious next steps that can be done in parallel by interested\n> >> parties would be:\n> >> \n> >>  (1) come up with a way to give \"name\" aka \"clustering cue\" (I\n> >>      think this is very hard);\n> >\n> > It is, and IMHO not worth it.  If you do it separately from the usual \n> > pack-objects process you'll perform extra IO and decompression when \n> > walking tree objects just to reconstruct those paths, becoming really \n> > slow by the context definition I provided above.\n> \n> Well, I said \"name\" in quotes because you do _NOT_ have to give\n> the real name.  I was not thinking about doing the actual tree\n> traversal at all.  What you need to do is to come up with a\n> token that is the same for the objects in the same deltification\n> chain so that they cluster together, and that should be doable\n> by looking at the delta chain patterns inside a packfile.\n\nObviously!  Sorry for being slow.\n\nBut I still think that a single repack pass should already be able to \npick loose objects and selected (small) packs, and produce a pack with \nthem all.  No need for a separate merge-pack I'd say.\n\n\nNicolas\n"},{"id":"52841","messageId":"20070907040731.GT18160@spearce.org","threadId":"9781","inReplyTo":"alpine.LFD.0.9999.0709061942320.21186@xanadu.home","subject":"Re: Subject: [PATCH] git-merge-pack","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-09-07T04:07:31Z","receivedAt":"2007-09-07T04:07:31Z","isPatch":true,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Nicolas Pitre <nico@cam.org> wrote:\n> I would have concatenated all packs provided on the command line into a \n> single one, simply by reading data from existing packs and writing it \n> back without any processing at all.  The offset for OBJ_OFS_DELTA is \n> relative so a simple concatenation will just work.\n> \n> Then the index for that pack can be created just as easily by reading \n> existing pack index files and storing the data into an array of struct \n> pack_idx_entry, adding the appropriate offset to object offsets, then \n> call write_idx_file().\n> \n> All data is read once and written once making it no more costly than a \n> simple file copy.  On the flip side it wouldn't get rid of duplicated \n> objects (I don't know if that matters i.e. if something might break with \n> the same object twice in a pack).\n\nYea, that's a really quick repack.  :-)  Plus its actually something\nthat can be easily halted in the middle and resumed later.  Just need\nto save the list of packfiles you are concatenating so you can pick\nup later when you get more time.\n\nThere shouldn't be a problem with having duplicates in the packfile.\nYou can do one of two things:\n\n  a) Omit the duplicates from the .idx when you merge the .idx tables\n     together to produce the new one.  Just take the object with the\n\t earliest offset.\n\n  b) Leave the duplicates in the final .idx.  In this case the\n     binary search may pick any of them, but it wouldn't matter\n     which it finds.\n\nAbout the only process that might care about duplicates would be\nindex-pack.  I don't think it makes sense to run index-pack on a\npackfile you already have a .idx for.  I don't think it would have\na problem with the duplicate SHA-1s either, but it wouldn't be hard\nto make it do something reasonable when it finds them.\n \n> > To consolidate all packs that are smaller than a megabytes into\n> > one, you would use it in its current form like this:\n> > \n> >     $ old=$(find .git/objects/pack -type f -name '*.pack' -size 1M)\n> >     $ new=$(echo \"$old\" | git merge-pack | git pack-objects pack)\n> >     $ for p in $old; do rm -f $p ${p%.pack}.idx; done\n> >     $ for s in pack idx; do mv pack-$new.$s .git/objects/pack/; done\n> \n> You might want to move the new pack before removing the old ones though.\n\nNot might, *must*.  If you delete the old ones before the new\nones are ready then readers can run into problems trying to access\nthe objects.  We've spent some effort trying to make these sorts\nof operations safe.  No sense in destroying that by getting the\norder wrong here.  :)\n\n-- \nShawn.\n"},{"id":"52845","messageId":"7vwsv36q6p.fsf@gitster.siamese.dyndns.org","threadId":"9781","inReplyTo":"alpine.LFD.0.9999.0709061942320.21186@xanadu.home","subject":"Re: Subject: [PATCH] git-merge-pack","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-09-07T04:43:26Z","receivedAt":"2007-09-07T04:43:26Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nicolas Pitre <nico@cam.org> writes:\n\n> I would have concatenated all packs provided on the command line into a \n> single one, simply by reading data from existing packs and writing it \n> back without any processing at all.  The offset for OBJ_OFS_DELTA is \n> relative so a simple concatenation will just work.\n\nAs I was planning to do this outside of pack-objects, I did not\nwant to write something that intimately knows the details of\npackfile format, but see below.\n\n> All data is read once and written once making it no more costly than a \n> simple file copy.  On the flip side it wouldn't get rid of duplicated \n> objects (I don't know if that matters i.e. if something might break with \n> the same object twice in a pack).\n\nI do not think duplicates create problems, as long as the pack\nidx remains sane.  But a bigger issue is for people who fetch\nover dumb protocols, from a repository that repacks with \"-a -d\"\nevery once in a while.  There, many duplicates are norm.\n\n> In fact, since we want to _also_ perform a repack of loose objects in \n> the context of automatic repacking, I wonder why we wouldn't use that \n> --unpacked= argument to also repack smallish packs at the same time in \n> only one pack-objects pass.  Or maybe I'm missing something?\n\nI think this is a much better idea.  You obviously need some\ntwist to the pack-objects, and being lazy that was the reason I\ndid not want to do this that way.\n\nWhen a new parameter, perhaps --lossless, is given, together\nwith the --unpacked= parameters, we can change pack-objects to\niterate over all objects in the --unpacked= packs, and add the\nones that are not marked for inclusion to the set of objects to\nbe packed, after doing the usual \"objects to be packed\"\ndiscovery.\n\nI am not sure --lossless is a good option name from marketing\npoint of view, though.\n"},{"id":"52846","messageId":"20070907044841.GX18160@spearce.org","threadId":"9781","inReplyTo":"7vk5r3adlx.fsf@gitster.siamese.dyndns.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-09-07T04:48:41Z","receivedAt":"2007-09-07T04:48:41Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Junio C Hamano <gitster@pobox.com> wrote:\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> > On Wed, 5 Sep 2007, Junio C Hamano wrote:\n> >>  static int aggressive_window = -1;\n> >> +static int gc_auto_threshold = 6700;\n> >\n> > Please don't do that.\n> >\n> > When you share objects with another git directory, git-gc --auto can get \n> > rid of the objects when some objects go away in the referenced repository.  \n> \n> I thought the whole point of \"gc --auto\" was to have something\n> that does not lose/prune any objects, even the ones that do not\n> seem to be referenced from anywhere.  That is why invocations of\n> \"git gc --auto\" do not say --prune as you saw the second patch,\n> and the repack command \"gc --auto\" runs is \"repack -d -l\"\n> instead of \"repack -a -d -l\", which means that it does run\n> git-prune-packed after repacking but not git-prune.\n> \n> Maybe I am missing something...\n\nNo, you aren't Junio.  `gc --auto` as you defined it is safe.\nIt won't delete objects from the database.  So it won't impact shared\nrepositories, or readers that are actively running in parallel with\nthe gc.  Both of which are important.\n\n-- \nShawn.\n"},{"id":"52864","messageId":"46E0F998.1080202@eudaptics.com","threadId":"9781","inReplyTo":"7v1wdb9ymf.fsf_-_@gitster.siamese.dyndns.org","subject":"Re: Subject: [PATCH] git-merge-pack","fromName":"Johannes Sixt","fromEmail":"j.sixt@eudaptics.com","sentAt":"2007-09-07T07:11:20Z","receivedAt":"2007-09-07T07:11:20Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Junio C Hamano schrieb:\n> This is a beginning of \"git-merge-pack\" that combines smaller\n> packs into one.\n\nThis gives a new meaning to the term \"merge\". IMHO, \"git-combine-pack\" would \nbe a better name.\n\n-- Hannes\n"},{"id":"52866","messageId":"200709070824.23541.andyparkins@gmail.com","threadId":"9781","inReplyTo":"7v1wdb9ymf.fsf_-_@gitster.siamese.dyndns.org","subject":"Re: Subject: [PATCH] git-merge-pack","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-09-07T07:24:20Z","receivedAt":"2007-09-07T07:24:20Z","isPatch":true,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Friday 2007 September 07, Junio C Hamano wrote:\n\n>  builtin-merge-pack.c |   87\n\nCan I suggest not calling it git-merge-pack?  It makes it look like it's a new \nmerge strategy called \"pack\"...\n\ngit-merge-base\ngit-merge-file\ngit-merge-index\ngit-merge-octopus\ngit-merge-one-file\ngit-merge-ours\ngit-merge-recur\ngit-merge-recursive\ngit-merge-resolve\ngit-merge-stupid\ngit-merge-subtree\ngit-merge-tree\n\n\n  \nAndy\n\n-- \nDr Andy Parkins, M Eng (hons), MIET\nandyparkins@gmail.com\n"},{"id":"52869","messageId":"7vd4wv6i8w.fsf@gitster.siamese.dyndns.org","threadId":"9781","inReplyTo":"46E0F998.1080202@eudaptics.com","subject":"Re: Subject: [PATCH] git-merge-pack","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-09-07T07:34:55Z","receivedAt":"2007-09-07T07:34:55Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Sixt <j.sixt@eudaptics.com> writes:\n\n> Junio C Hamano schrieb:\n>> This is a beginning of \"git-merge-pack\" that combines smaller\n>> packs into one.\n>\n> This gives a new meaning to the term \"merge\". IMHO, \"git-combine-pack\"\n> would be a better name.\n\nYeah, that makes sense, but I think this can and should be done\nas part of pack-objects itself as Nico suggested.\n\nSo consider that patch scrapped for now.\n"},{"id":"52885","messageId":"Pine.LNX.4.64.0709071111400.28586@racer.site","threadId":"9781","inReplyTo":"7vk5r3adlx.fsf@gitster.siamese.dyndns.org","subject":"Re: People unaware of the importance of \"git gc\"?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-09-07T10:12:30Z","receivedAt":"2007-09-07T10:12:30Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 6 Sep 2007, Junio C Hamano wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> > On Wed, 5 Sep 2007, Junio C Hamano wrote:\n> >\n> >> @@ -20,6 +20,7 @@ static const char builtin_gc_usage[] = \"git-gc [--prune] [--aggressive]\";\n> >>  \n> >>  static int pack_refs = 1;\n> >>  static int aggressive_window = -1;\n> >> +static int gc_auto_threshold = 6700;\n> >\n> > Please don't do that.\n> >\n> > When you share objects with another git directory, git-gc --auto can \n> > get rid of the objects when some objects go away in the referenced \n> > repository.\n> \n> I thought the whole point of \"gc --auto\" was to have something\n> that does not lose/prune any objects, even the ones that do not\n> seem to be referenced from anywhere.  That is why invocations of\n> \"git gc --auto\" do not say --prune as you saw the second patch,\n> and the repack command \"gc --auto\" runs is \"repack -d -l\"\n> instead of \"repack -a -d -l\", which means that it does run\n> git-prune-packed after repacking but not git-prune.\n> \n> Maybe I am missing something...\n\nNo, _I_ missed the fact that no pack is rewritten...\n\nSorry for the line noise,\nDscho\n"},{"id":"52985","messageId":"7v4pi51o6p.fsf_-_@gitster.siamese.dyndns.org","threadId":"9781","inReplyTo":"7vwsv36q6p.fsf@gitster.siamese.dyndns.org","subject":"[PATCH] make sha1_file.c::matches_pack_name() available to others","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-09-08T09:50:06Z","receivedAt":"2007-09-08T09:50:06Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Even though our convention is \"zero return means good\", it goes a\nbit too far for matches_pack_name() to return 0 when it found\nthe pack is what the name refers to.  This fixes that silly and\nobvious interface bug.\n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n Junio C Hamano <gitster@pobox.com> writes:\n\n > Nicolas Pitre <nico@cam.org> writes:\n > ...\n >> In fact, since we want to _also_ perform a repack of loose objects in \n >> the context of automatic repacking, I wonder why we wouldn't use that \n >> --unpacked= argument to also repack smallish packs at the same time in \n >> only one pack-objects pass.  Or maybe I'm missing something?\n >\n > I think this is a much better idea.  You obviously need some\n > twist to the pack-objects, and being lazy that was the reason I\n > did not want to do this that way.\n\n So what follows is two-patch series, which still is a rough\n sketch, as I am feeling a bit too tired to do tests and\n documentation (help is always welcomed, hint hint).\n\n This message contains the first one, which is more or less\n independent, that exposes matches_pack_name() function from\n sha1_file.c, while fixing a silly and obvious interface bug.\n\n cache.h     |    1 +\n sha1_file.c |   14 +++++++-------\n 2 files changed, 8 insertions(+), 7 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex 70abbd5..3fa5b8e 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -529,6 +529,7 @@ extern void *unpack_entry(struct packed_git *, off_t, enum object_type *, unsign\n extern unsigned long unpack_object_header_gently(const unsigned char *buf, unsigned long len, enum object_type *type, unsigned long *sizep);\n extern unsigned long get_size_from_delta(struct packed_git *, struct pack_window **, off_t);\n extern const char *packed_object_info_detail(struct packed_git *, off_t, unsigned long *, unsigned long *, unsigned int *, unsigned char *);\n+extern int matches_pack_name(struct packed_git *p, const char *name);\n \n /* Dumb servers support */\n extern int update_server_info(int);\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 9978a58..5801c3e 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -1684,22 +1684,22 @@ off_t find_pack_entry_one(const unsigned char *sha1,\n \treturn 0;\n }\n \n-static int matches_pack_name(struct packed_git *p, const char *ig)\n+int matches_pack_name(struct packed_git *p, const char *name)\n {\n \tconst char *last_c, *c;\n \n-\tif (!strcmp(p->pack_name, ig))\n-\t\treturn 0;\n+\tif (!strcmp(p->pack_name, name))\n+\t\treturn 1;\n \n \tfor (c = p->pack_name, last_c = c; *c;)\n \t\tif (*c == '/')\n \t\t\tlast_c = ++c;\n \t\telse\n \t\t\t++c;\n-\tif (!strcmp(last_c, ig))\n-\t\treturn 0;\n+\tif (!strcmp(last_c, name))\n+\t\treturn 1;\n \n-\treturn 1;\n+\treturn 0;\n }\n \n static int find_pack_entry(const unsigned char *sha1, struct pack_entry *e, const char **ignore_packed)\n@@ -1717,7 +1717,7 @@ static int find_pack_entry(const unsigned char *sha1, struct pack_entry *e, cons\n \t\tif (ignore_packed) {\n \t\t\tconst char **ig;\n \t\t\tfor (ig = ignore_packed; *ig; ig++)\n-\t\t\t\tif (!matches_pack_name(p, *ig))\n+\t\t\t\tif (matches_pack_name(p, *ig))\n \t\t\t\t\tbreak;\n \t\t\tif (*ig)\n \t\t\t\tgoto next;\n"},{"id":"52986","messageId":"7vodgdzdaj.fsf_-_@gitster.siamese.dyndns.org","threadId":"9781","inReplyTo":"7vwsv36q6p.fsf@gitster.siamese.dyndns.org","subject":"[PATCH] pack-objects --repack-unpacked","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-09-08T10:01:24Z","receivedAt":"2007-09-08T10:01:24Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"The usual command line that uses \"--unpacked=<existing>\" option\nlooks like this:\n\n\tgit pack-objects --non-empty --all --reflog \\\n        \t--unpacked --unpacked=<existing> \\\n                packname-prefix\n\nThis packs loose objects and objects in the named existing\npacks that are reachable from any and all refs and reflog\nentries.  It is typically used by \"git repack -a -d\", which\nthen removes the named existing packs from the repository, and\nhas an effect of getting rid of unreachable objects these packs\nhold.\n\nThis adds \"--repack-unpacked\" option to pack-objects to help\ncombining small packs into one, without losing unreferenced\nobjects that are in the packs.  When this option is given in\naddition to the above command line, we also make sure all the\nobjects in the named existing packs are included in the result.\n\nThis allows us to safely remove the packs that were named on the\ncommand line after installing the resulting pack in the\nrepository.\n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n I am too tired to keep staring at this code now.  Fixes,\n improvements, replacements and enhancements, in the code,\n documentation and tests, are very much welcomed.\n\n builtin-pack-objects.c |   95 +++++++++++++++++++++++++++++++++++++++++++++++-\n 1 files changed, 93 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin-pack-objects.c b/builtin-pack-objects.c\nindex 12509fa..9bc2faa 100644\n--- a/builtin-pack-objects.c\n+++ b/builtin-pack-objects.c\n@@ -21,7 +21,7 @@ git-pack-objects [{ -q | --progress | --all-progress }] \\n\\\n \t[--window=N] [--window-memory=N] [--depth=N] \\n\\\n \t[--no-reuse-delta] [--no-reuse-object] [--delta-base-offset] \\n\\\n \t[--non-empty] [--revs [--unpacked | --all]*] [--reflog] \\n\\\n-\t[--stdout | base-name] [<ref-list | <object-list]\";\n+\t[--stdout | base-name] [--repack-unpacked] [<ref-list | <object-list]\";\n \n struct object_entry {\n \tstruct pack_idx_entry idx;\n@@ -57,7 +57,7 @@ static struct object_entry **written_list;\n static uint32_t nr_objects, nr_alloc, nr_result, nr_written;\n \n static int non_empty;\n-static int no_reuse_delta, no_reuse_object;\n+static int no_reuse_delta, no_reuse_object, repack_unpacked;\n static int local;\n static int incremental;\n static int allow_ofs_delta;\n@@ -1625,15 +1625,21 @@ static void read_object_list_from_stdin(void)\n \t}\n }\n \n+#define OBJECT_ADDED (1u<<20)\n+\n static void show_commit(struct commit *commit)\n {\n \tadd_object_entry(commit->object.sha1, OBJ_COMMIT, NULL, 0);\n+\tcommit->object.flags |= OBJECT_ADDED;\n }\n \n static void show_object(struct object_array_entry *p)\n {\n+\tstruct object *o = lookup_unknown_object(p->item->sha1);\n+\n \tadd_preferred_base_object(p->name);\n \tadd_object_entry(p->item->sha1, p->item->type, p->name, 0);\n+\to->flags |= OBJECT_ADDED;\n }\n \n static void show_edge(struct commit *commit)\n@@ -1641,6 +1647,84 @@ static void show_edge(struct commit *commit)\n \tadd_preferred_base(commit->object.sha1);\n }\n \n+struct in_pack_object {\n+\toff_t offset;\n+\tconst unsigned char *sha1;\n+};\n+\n+struct in_pack {\n+\tint alloc;\n+\tint nr;\n+\tstruct in_pack_object *array;\n+};\n+\n+static void mark_in_pack_object(const unsigned char *sha1, struct packed_git *p, struct in_pack *in_pack)\n+{\n+\tin_pack->array[in_pack->nr].offset = find_pack_entry_one(sha1, p);\n+\tin_pack->array[in_pack->nr].sha1 = sha1;\n+\tin_pack->nr++;\n+}\n+\n+/*\n+ * Compare the objects in the offset order, in order to emulate the\n+ * \"git-rev-list --objects\" output that produced the pack originally.\n+ */\n+static int ofscmp(const void *a_, const void *b_)\n+{\n+\tstruct in_pack_object *a = (struct in_pack_object *)a_;\n+\tstruct in_pack_object *b = (struct in_pack_object *)b_;\n+\n+\tif (a->offset < b->offset)\n+\t\treturn -1;\n+\telse if (a->offset > b->offset)\n+\t\treturn 1;\n+\telse\n+\t\treturn hashcmp(a->sha1, b->sha1);\n+}\n+\n+static void add_objects_in_unpacked_packs(struct rev_info *revs)\n+{\n+\tstruct packed_git *p;\n+\n+\tfor (p = packed_git; p; p = p->next) {\n+\t\tstruct in_pack in_pack;\n+\t\tconst unsigned char *sha1;\n+\t\tstruct object *o;\n+\t\tuint32_t i;\n+\n+\t\tfor (i = 0; i < revs->num_ignore_packed; i++) {\n+\t\t\tif (matches_pack_name(p, revs->ignore_packed[i]))\n+\t\t\t\tbreak;\n+\t\t}\n+\t\tif (revs->num_ignore_packed <= i)\n+\t\t\tcontinue;\n+\t\tif (open_pack_index(p))\n+\t\t\tdie(\"cannot open pack index\");\n+\n+\t\tin_pack.alloc = p->num_objects;\n+\t\tin_pack.nr = 0;\n+\t\tin_pack.array = xmalloc(sizeof(in_pack.array[0]) *\n+\t\t\t\t\tp->num_objects);\n+\t\tfor (i = 0; i < p->num_objects; i++) {\n+\t\t\tsha1 = nth_packed_object_sha1(p, i);\n+\t\t\to = lookup_unknown_object(sha1);\n+\t\t\tif (!(o->flags & OBJECT_ADDED))\n+\t\t\t\tmark_in_pack_object(sha1, p, &in_pack);\n+\t\t\to->flags |= OBJECT_ADDED;\n+\t\t}\n+\t\tif (!in_pack.nr)\n+\t\t\tcontinue;\n+\t\tqsort(in_pack.array, in_pack.nr, sizeof(in_pack.array[0]),\n+\t\t      ofscmp);\n+\t\tfor (i = 0; i < in_pack.nr; i++) {\n+\t\t\tsha1 = in_pack.array[i].sha1;\n+\t\t\to = lookup_unknown_object(sha1);\n+\t\t\tadd_object_entry(sha1, o->type, \"\", 0);\n+\t\t}\n+\t\tfree(in_pack.array);\n+\t}\n+}\n+\n static void get_object_list(int ac, const char **av)\n {\n \tstruct rev_info revs;\n@@ -1672,6 +1756,9 @@ static void get_object_list(int ac, const char **av)\n \tprepare_revision_walk(&revs);\n \tmark_edges_uninteresting(revs.commits, &revs, show_edge);\n \ttraverse_commit_list(&revs, show_commit, show_object);\n+\n+\tif (repack_unpacked)\n+\t\tadd_objects_in_unpacked_packs(&revs);\n }\n \n static int adjust_perm(const char *path, mode_t mode)\n@@ -1789,6 +1876,10 @@ int cmd_pack_objects(int argc, const char **argv, const char *prefix)\n \t\t\tuse_internal_rev_list = 1;\n \t\t\tcontinue;\n \t\t}\n+\t\tif (!strcmp(\"--repack-unpacked\", arg)) {\n+\t\t\trepack_unpacked = 1;\n+\t\t\tcontinue;\n+\t\t}\n \t\tif (!strcmp(\"--unpacked\", arg) ||\n \t\t    !prefixcmp(arg, \"--unpacked=\") ||\n \t\t    !strcmp(\"--reflog\", arg) ||\n"},{"id":"359792","messageId":"87k1mta9x5.fsf@evledraar.gmail.com","threadId":"9781","inReplyTo":"7vr6lcj2zi.fsf@gitster.siamese.dyndns.org","subject":"What's so special about objects/17/ ?","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-10-07T18:28:22Z","receivedAt":"2018-10-07T18:28:27Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"In 2007 Junio wrote\n(https://public-inbox.org/git/7vr6lcj2zi.fsf@gitster.siamese.dyndns.org/):\n\n    +static int need_to_gc(void)\n    +{\n    +\t/*\n    +\t * Quickly check if a \"gc\" is needed, by estimating how\n    +\t * many loose objects there are.  Because SHA-1 is evenly\n    +\t * distributed, we can check only one and get a reasonable\n    +\t * estimate.\n    +\t */\n    +\tchar path[PATH_MAX];\n    +\tconst char *objdir = get_object_directory();\n    +\tDIR *dir;\n    +\tstruct dirent *ent;\n    +\tint auto_threshold;\n    +\tint num_loose = 0;\n    +\tint needed = 0;\n    +\n    +\tif (sizeof(path) <= snprintf(path, sizeof(path), \"%s/17\", objdir)) {\n    +\t\twarning(\"insanely long object directory %.*s\", 50, objdir);\n    +\t\treturn 0;\n    +\t}\n    +\tdir = opendir(path);\n    +\tif (!dir)\n    +\t\treturn 0;\n    +\n    +\tauto_threshold = (gc_auto_threshold + 255) / 256;\n    +\twhile ((ent = readdir(dir)) != NULL) {\n    +\t\tif (strspn(ent->d_name, \"0123456789abcdef\") != 38 ||\n    +\t\t    ent->d_name[38] != '\\0')\n    +\t\t\tcontinue;\n    +\t\tif (++num_loose > auto_threshold) {\n    +\t\t\tneeded = 1;\n    +\t\t\tbreak;\n    +\t\t}\n    +\t}\n\nA couple of questions about this patch, which is in git.git as\n2c3c439947 (\"Implement git gc --auto\", 2007-09-05)\n\n1. We still have this check of objects/17/ in builtin/gc.c today. Why\n   objects/17/ and not e.g. objects/00/ to go with other 000* magic such\n   as the 0000000000000000000000000000000000000000 SHA-1?  Statistically\n   it doesn't matter, but 17 seems like an odd thing to pick at random\n   out of 00..ff, does it have any significance?\n\n2. It seems overly paranoid to be checking that the files in\n  .git/objects/17/ look like a SHA-1. If we have stuff not generated by\n  git in .git/objects/??/ we probably have bigger problems than\n  prematurely triggering auto gc, can this just be removed as\n  redundant. Was this some check e.g. expecting that this would need to\n  deal with tempfiles in these directories that we created at the time\n  (but no longer do?)?\n"},{"id":"359793","messageId":"f64b5c5d-ef72-a347-bd0f-7b1669a8c10d@kdbg.org","threadId":"9781","inReplyTo":"87k1mta9x5.fsf@evledraar.gmail.com","subject":"Re: What's so special about objects/17/ ?","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2018-10-07T18:35:57Z","receivedAt":"2018-10-07T18:36:05Z","isPatch":false,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 07.10.18 um 20:28 schrieb Ævar Arnfjörð Bjarmason:\n> In 2007 Junio wrote\n> (https://public-inbox.org/git/7vr6lcj2zi.fsf@gitster.siamese.dyndns.org/):\n> \n>      +static int need_to_gc(void)\n>      +{\n>      +\t/*\n>      +\t * Quickly check if a \"gc\" is needed, by estimating how\n>      +\t * many loose objects there are.  Because SHA-1 is evenly\n>      +\t * distributed, we can check only one and get a reasonable\n>      +\t * estimate.\n>      +\t */\n\n> 1. We still have this check of objects/17/ in builtin/gc.c today. Why\n>     objects/17/ and not e.g. objects/00/ to go with other 000* magic such\n>     as the 0000000000000000000000000000000000000000 SHA-1?  Statistically\n>     it doesn't matter, but 17 seems like an odd thing to pick at random\n>     out of 00..ff, does it have any significance?\n\nThe reason is explained in the comment. And, BTW, you do know about this \none: https://xkcd.com/221/ don't you? (TLDR: the title is \"Random Number\")\n\n> 2. It seems overly paranoid to be checking that the files in\n>    .git/objects/17/ look like a SHA-1. If we have stuff not generated by\n>    git in .git/objects/??/ we probably have bigger problems than\n>    prematurely triggering auto gc, can this just be removed as\n>    redundant. Was this some check e.g. expecting that this would need to\n>    deal with tempfiles in these directories that we created at the time\n>    (but no longer do?)?\n\nIt's not about that there are SHA-1s in there, it's about how many there \nare.\n\n-- Hannes\n"},{"id":"359794","messageId":"87in2da862.fsf@evledraar.gmail.com","threadId":"9781","inReplyTo":"f64b5c5d-ef72-a347-bd0f-7b1669a8c10d@kdbg.org","subject":"Re: What's so special about objects/17/ ?","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-10-07T19:06:13Z","receivedAt":"2018-10-07T19:06:23Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Sun, Oct 07 2018, Johannes Sixt wrote:\n\n> Am 07.10.18 um 20:28 schrieb Ævar Arnfjörð Bjarmason:\n>> In 2007 Junio wrote\n>> (https://public-inbox.org/git/7vr6lcj2zi.fsf@gitster.siamese.dyndns.org/):\n>>\n>>      +static int need_to_gc(void)\n>>      +{\n>>      +\t/*\n>>      +\t * Quickly check if a \"gc\" is needed, by estimating how\n>>      +\t * many loose objects there are.  Because SHA-1 is evenly\n>>      +\t * distributed, we can check only one and get a reasonable\n>>      +\t * estimate.\n>>      +\t */\n>\n>> 1. We still have this check of objects/17/ in builtin/gc.c today. Why\n>>     objects/17/ and not e.g. objects/00/ to go with other 000* magic such\n>>     as the 0000000000000000000000000000000000000000 SHA-1?  Statistically\n>>     it doesn't matter, but 17 seems like an odd thing to pick at random\n>>     out of 00..ff, does it have any significance?\n>\n> The reason is explained in the comment. And, BTW, you do know about\n> this one: https://xkcd.com/221/ don't you? (TLDR: the title is \"Random\n> Number\")\n\nPicking any one number is explained in the comment. I'm asking why 17 in\nparticular not for correctness reasons but as a bit of historical lore,\nand because my ulterior is to improve the GC docs.\n\nThe number in that comic is 4 (and no datestamp on when it was\npublished). Are you saying Junio's patch is somehow a reference to that\nxkcd in particular, or that it's just a funny reference in this context?\n\n>> 2. It seems overly paranoid to be checking that the files in\n>>    .git/objects/17/ look like a SHA-1. If we have stuff not generated by\n>>    git in .git/objects/??/ we probably have bigger problems than\n>>    prematurely triggering auto gc, can this just be removed as\n>>    redundant. Was this some check e.g. expecting that this would need to\n>>    deal with tempfiles in these directories that we created at the time\n>>    (but no longer do?)?\n>\n> It's not about that there are SHA-1s in there, it's about how many\n> there are.\n\nRight, I'm wondering if it couldn't be replaced by some general path.c\n\"number_of_files_in_dir\" helper. I.e. why this code is being paranoid\nabout ignoring the likes of\n.git/objects/17/{foo,bar,some-other-garbage}. A number_of_files_in_dir()\nwould obviously need to ignore \".\" and \"..\".\n"},{"id":"359795","messageId":"xmqqpnwltu8s.fsf@gitster-ct.c.googlers.com","threadId":"9781","inReplyTo":"87k1mta9x5.fsf@evledraar.gmail.com","subject":"Re: What's so special about objects/17/ ?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-10-07T19:46:43Z","receivedAt":"2018-10-07T19:46:51Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n\n> 1. We still have this check of objects/17/ in builtin/gc.c today. Why\n>    objects/17/ and not e.g. objects/00/ to go with other 000* magic such\n>    as the 0000000000000000000000000000000000000000 SHA-1?d  Statistically\n>    it doesn't matter, but 17 seems like an odd thing to pick at random\n>    out of 00..ff, does it have any significance?\n\nThere is no \"other 000* magic such as ...\". There is only one 0{40}\nmagic and that one must be memorable and explainable.\n\nThe 1/256 sample can be any one among 256.  Just like the date\nstring on the first line of the output to be used as the /etc/magic\nsignature by format-patch, it was an arbitrary choice, rather than a\nrandom choice, and unlike 0{40} this does not have to be memorable\nby general public and I do not have to explain the choice to the\ngeneral public ;-)\n\n> 2. It seems overly paranoid to be checking that the files in\n>   .git/objects/17/ look like a SHA-1.\n\nThere is no other reason than futureproofing.  We were paying cost\nto open and scan the directory anyway, and checking that we only\ncount the loose object files was (and still is) a sensible thing to\ndo to allow us not even worry about the other kind of things we\nmight end up creating there.\n"},{"id":"359812","messageId":"xmqqlg79tta8.fsf@gitster-ct.c.googlers.com","threadId":"9781","inReplyTo":"xmqqpnwltu8s.fsf@gitster-ct.c.googlers.com","subject":"Re: What's so special about objects/17/ ?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-10-07T20:07:27Z","receivedAt":"2018-10-07T20:07:32Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n>\n>> 1. We still have this check of objects/17/ in builtin/gc.c today. Why\n>>    objects/17/ and not e.g. objects/00/ to go with other 000* magic such\n>>    as the 0000000000000000000000000000000000000000 SHA-1?d  Statistically\n>>    it doesn't matter, but 17 seems like an odd thing to pick at random\n>>    out of 00..ff, does it have any significance?\n>\n> ...\n> by general public and I do not have to explain the choice to the\n> general public ;-)\n\nOne thing that is more important than \"why not 00 but 17?\" to answer\nis why a hardcoded number rather than a runtime random.  It is for\nrepeatability.\n\n"},{"id":"359814","messageId":"c904d10e-d6a1-b1d4-73eb-fb93f5caecdb@kdbg.org","threadId":"9781","inReplyTo":"87in2da862.fsf@evledraar.gmail.com","subject":"Re: What's so special about objects/17/ ?","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2018-10-07T22:39:47Z","receivedAt":"2018-10-07T22:39:52Z","isPatch":false,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 07.10.18 um 21:06 schrieb Ævar Arnfjörð Bjarmason:\n> Picking any one number is explained in the comment. I'm asking why 17 in\n> particular not for correctness reasons but as a bit of historical lore,\n> and because my ulterior is to improve the GC docs.\n> \n> The number in that comic is 4 (and no datestamp on when it was\n> published). Are you saying Junio's patch is somehow a reference to that\n> xkcd in particular, or that it's just a funny reference in this context?\n\nNo lore, AFAIR. It's just a random number, determined by a fair dice \nroll or something ;)\n\n-- Hannes\n"},{"id":"359815","messageId":"xmqqh8hxtfzn.fsf@gitster-ct.c.googlers.com","threadId":"9781","inReplyTo":"c904d10e-d6a1-b1d4-73eb-fb93f5caecdb@kdbg.org","subject":"Re: What's so special about objects/17/ ?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-10-08T00:54:36Z","receivedAt":"2018-10-08T00:54:44Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Sixt <j6t@kdbg.org> writes:\n\n> Am 07.10.18 um 21:06 schrieb Ævar Arnfjörð Bjarmason:\n>> Picking any one number is explained in the comment. I'm asking why 17 in\n>> particular not for correctness reasons but as a bit of historical lore,\n>> and because my ulterior is to improve the GC docs.\n>>\n>> The number in that comic is 4 (and no datestamp on when it was\n>> published). Are you saying Junio's patch is somehow a reference to that\n>> xkcd in particular, or that it's just a funny reference in this context?\n>\n> No lore, AFAIR. It's just a random number, determined by a fair dice\n> roll or something ;)\n\nAs I already said, I did not pick the number randomly, but rather\narbitrarily, and it is not 00 because the chosen number (unlike the\n0{40} magic we use elsewhere) does not have to be memorable, and the\nchoice does not have to be explainable.\n\nSo people will not get any further explanation as to the reason\nbehind that arbitrary choice, but it was not random.\n"},{"id":"359817","messageId":"87h8hwafof.fsf@evledraar.gmail.com","threadId":"9781","inReplyTo":"xmqqpnwltu8s.fsf@gitster-ct.c.googlers.com","subject":"Re: What's so special about objects/17/ ?","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-10-08T10:36:16Z","receivedAt":"2018-10-08T10:36:21Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Sun, Oct 07 2018, Junio C Hamano wrote:\n\n> Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n>\n>> 1. We still have this check of objects/17/ in builtin/gc.c today. Why\n>>    objects/17/ and not e.g. objects/00/ to go with other 000* magic such\n>>    as the 0000000000000000000000000000000000000000 SHA-1?d  Statistically\n>>    it doesn't matter, but 17 seems like an odd thing to pick at random\n>>    out of 00..ff, does it have any significance?\n>\n> There is no \"other 000* magic such as ...\". There is only one 0{40}\n> magic and that one must be memorable and explainable.\n\nDepending on how we're counting there's at least two. We also use\n0000000000000000000000000000000000000000 as a placeholder for \"couldn't\nread a ref\" in addition or \"this is a placeholder for an invalid ref\" in\naddition to how it's used to signify creation/deletion to the in the\nlikes of the pre-receive hook:\n\n    $ echo hello > .git/refs/something\n    $ git fsck\n    [...]\n    error: refs/something: invalid sha1 pointer 0000000000000000000000000000000000000000\n    $ > .git/refs/something\n    $ git fsck\n    [...]\n    error: refs/something: invalid sha1 pointer 0000000000000000000000000000000000000000\n\nThis is because the refs backend will memzero the oid struct, and if we\nfail to read things it'll still be zero'd out.\n\nThis manifests e.g. in this confusing fsck output, due to a bug where\nGitLab will write empty refs/keep-around/* refs sometimes:\nhttps://gitlab.com/gitlab-org/gitlab-ce/issues/44431\n\n> The 1/256 sample can be any one among 256.  Just like the date\n> string on the first line of the output to be used as the /etc/magic\n> signature by format-patch, it was an arbitrary choice, rather than a\n> random choice, and unlike 0{40} this does not have to be memorable\n> by general public and I do not have to explain the choice to the\n> general public ;-)\n\nI wanted to elaborate on the explanation for \"gc.auto\" in\ngit-config. Now we just say \"approximately 6700\". Since this behavior\nhas been really stable for a long time we could say we sample 1/256 of\nthe .git/objects/?? dirs, and this explains any perceived discrepancies\nbetween the 6700 number and $(find .git/objects/?? -type f | wc -l).\n\n>> 2. It seems overly paranoid to be checking that the files in\n>>   .git/objects/17/ look like a SHA-1.\n>\n> There is no other reason than futureproofing.  We were paying cost\n> to open and scan the directory anyway, and checking that we only\n> count the loose object files was (and still is) a sensible thing to\n> do to allow us not even worry about the other kind of things we\n> might end up creating there.\n\nMakes sense. Just wanted to ask if it was that or some workaround for\nhistorical files being there.\n"},{"id":"359849","messageId":"CAGZ79kZq3xtsbscrRFD8CSn++yrvdM6Ux+nkQ3AamgabXtPL+w@mail.gmail.com","threadId":"9781","inReplyTo":"xmqqlg79tta8.fsf@gitster-ct.c.googlers.com","subject":"Re: What's so special about objects/17/ ?","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2018-10-08T19:17:40Z","receivedAt":"2018-10-08T19:17:55Z","isPatch":false,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Sun, Oct 7, 2018 at 1:07 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Junio C Hamano <gitster@pobox.com> writes:\n\n> > ...\n> > by general public and I do not have to explain the choice to the\n> > general public ;-)\n>\n> One thing that is more important than \"why not 00 but 17?\" to answer\n> is why a hardcoded number rather than a runtime random.  It is for\n> repeatability.\n\nLet's talk about repeatability vs statistics for a second. ;-)\n\nIf I am a user and I were really into optimizing my loose object count\nfor some reason, so I would want to choose a low number of\ngc.auto. Let's say I go with 128.\n\nAt the low end of loose objects the approximation is yielding\nsome high relative errors. This is because of the granularity, i.e.\ngc would implicitly estimate the loose objects to be 0 or 256 or 512, (or more)\nif there is 0, 1, 2 (or more) loose objects in the objects/17.\n\nAs each object can be viewed following an unfair coin flip\n(With a chance of 1/256 it is in objects/17), the distribution in\nobjects/17 (and hence any other objects/XX bin) follows the\nBernoulli distribution.\n\nIf I do have say about 157 loose objects (and having auto.gc\nconfigured anywhere in 1..255), then the probability to not\ngc is 54% (as that is the probability to have 0 objects in /17,\nfollowing probability mass function of the Bernoulli distribution,\n(i.e. Pr(0 objects) = (157 over 0) x (1/256)^0 x (255/256)^157))\n\nAs it is repeatable (by picking the same /17 every time), I can run\n\"gc --auto\" multiple times and still have 157 loose objects, despite\nwanting to have only 128 loose objects at a 54% chance.\n\nIf we'd roll the 256 dice every time to pick a different bin,\nthen we might hit another bin and gc in the second or third\ngc, which would be more precise on average.\n\nBy having repeatability we allow for these numbers to be far off\nmore often when configuring small numbers.\n\nI think that is the right choice, as we probably do not care about the\nexactness of auto-gc for small numbers, as it is a performance\nthing anyway. Although documenting it properly might be a challenge.\n\nThe current wording of auto.gc seems to suggest that we are right\nfor the number as we compute it via the implying the expected value,\n(i.e. we pick a bin and multiply the fullness of the bin by the number\nof bins to estimate the whole fullness, see the mean=n p on [1])\nI think a user would be far more interested in giving an upper bound,\ni.e. expressing something like \"I will have at most $auto.gc objects\nbefore gc kicks in\" or \"The likelihood to exceed the $auto.gc number\nof loose objects by $this much is less than 5%\", for which the math\nwould be more complicated, but easier to document with the words of\nstatistics.\n\n[1]  https://en.wikipedia.org/wiki/Binomial_distribution\n\nStefan\n"},{"id":"359886","messageId":"xmqq4ldwszh8.fsf@gitster-ct.c.googlers.com","threadId":"9781","inReplyTo":"CAGZ79kZq3xtsbscrRFD8CSn++yrvdM6Ux+nkQ3AamgabXtPL+w@mail.gmail.com","subject":"Re: What's so special about objects/17/ ?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-10-09T01:03:31Z","receivedAt":"2018-10-09T01:03:37Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Stefan Beller <sbeller@google.com> writes:\n\n> On Sun, Oct 7, 2018 at 1:07 PM Junio C Hamano <gitster@pobox.com> wrote:\n>>\n>> Junio C Hamano <gitster@pobox.com> writes:\n>\n>> > ...\n>> > by general public and I do not have to explain the choice to the\n>> > general public ;-)\n>>\n>> One thing that is more important than \"why not 00 but 17?\" to answer\n>> is why a hardcoded number rather than a runtime random.  It is for\n>> repeatability.\n>\n> Let's talk about repeatability vs statistics for a second. ;-)\n\nOh, I think I misled you by saying \"more important\".  \n\nI didn't mean that it is more important to stick to the \"use\nhardcoded value\" design decision than sticking to \"use 17\".  I've\nmade sure that everybody would understnd choosing any arbitrary byte\nvalue other than \"17\" does not make the resulting Git any better nor\nworse.  But discussing the design decision to use hardcoded value is\n\"more important\", as that affects the balance between the end-user\nexperience and debuggability, and I tried to help those who do not\nknow the history by giving the fact that choice was made for the\nlatter and not for other hidden reasons, that those who would\npropose to change the system may have to keep in mind.\n\nSorry if you mistook it as if I were saying that it is important to\nkeep the design to use a hardcoded byte value.  That wasn't what the\nmessage was about.\n"},{"id":"359887","messageId":"xmqqzhvorkq7.fsf@gitster-ct.c.googlers.com","threadId":"9781","inReplyTo":"87h8hwafof.fsf@evledraar.gmail.com","subject":"Re: What's so special about objects/17/ ?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-10-09T01:07:28Z","receivedAt":"2018-10-09T01:07:34Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n\n> Depending on how we're counting there's at least two.\n\nI thought you were asking \"why the special sentinel is not 0{40}?\"\nYou counted the number of reasons why 0{40} is used to stand in for\na real value, but that was the number I didn't find interesting in\nthe scope of this discussion, i.e. \"why the special sample is 17?\"\n\nI vaguely recall we also used 0{39}1 for something else long time\nago; I offhand do not recall if we still do, or we got rid of it.\n"},{"id":"359928","messageId":"CAGZ79ka5kKrSqPCWFMDetRLYxDqcguJUzJXDex9q-VMwT-ABAw@mail.gmail.com","threadId":"9781","inReplyTo":"xmqq4ldwszh8.fsf@gitster-ct.c.googlers.com","subject":"Re: What's so special about objects/17/ ?","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2018-10-09T17:37:59Z","receivedAt":"2018-10-09T17:38:14Z","isPatch":false,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Mon, Oct 8, 2018 at 6:03 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Stefan Beller <sbeller@google.com> writes:\n>\n> > On Sun, Oct 7, 2018 at 1:07 PM Junio C Hamano <gitster@pobox.com> wrote:\n> >>\n> >> Junio C Hamano <gitster@pobox.com> writes:\n> >\n> >> > ...\n> >> > by general public and I do not have to explain the choice to the\n> >> > general public ;-)\n> >>\n> >> One thing that is more important than \"why not 00 but 17?\" to answer\n> >> is why a hardcoded number rather than a runtime random.  It is for\n> >> repeatability.\n> >\n> > Let's talk about repeatability vs statistics for a second. ;-)\n>\n> Oh, I think I misled you by saying \"more important\".\n>\n> I didn't mean that it is more important to stick to the \"use\n> hardcoded value\" design decision than sticking to \"use 17\".  I've\n> made sure that everybody would understnd choosing any arbitrary byte\n> value other than \"17\" does not make the resulting Git any better nor\n> worse.\n\nYes, I totally get that. We could have chosen 42 just because.\n\n\n>  But discussing the design decision to use hardcoded value is\n> \"more important\", as that affects the balance between the end-user\n> experience and debuggability, and I tried to help those who do not\n> know the history by giving the fact that choice was made for the\n> latter and not for other hidden reasons, that those who would\n> propose to change the system may have to keep in mind.\n\nFrom an end users point of view, the auto gc kicks in at random.\n(Maybe it's just me, but I don't keep track of the loose object count ;-)\n\nFor debuggability, we could design a system that allows for debugging,\ne.g. \"When GIT_AUTO_GC_BIN is set, use the number as set, otherwise\ntake a random slot\".\n\n> Sorry if you mistook it as if I were saying that it is important to\n> keep the design to use a hardcoded byte value.  That wasn't what the\n> message was about.\n\nI understood very well that the choice of value was arbitrary and you\ndo not have a convincing story as to why 17 (and not say 23, but such\na story is not required, as all slots are equal from a design perspective).\n\nI do challenge the decision to take a hardcoded value, though, as it\nyields better properties for the end users IMHO, whereas debugging\nthis specific case does not seem to be important to me.\n"},{"id":"359929","messageId":"CAGZ79kbASAXMfKR95CNiJpRFTTq6DKho2v=UfpLG=_jnVCTTUA@mail.gmail.com","threadId":"9781","inReplyTo":"xmqqzhvorkq7.fsf@gitster-ct.c.googlers.com","subject":"Re: What's so special about objects/17/ ?","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2018-10-09T17:40:50Z","receivedAt":"2018-10-09T17:41:05Z","isPatch":false,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Mon, Oct 8, 2018 at 6:07 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n>\n> > Depending on how we're counting there's at least two.\n>\n> I thought you were asking \"why the special sentinel is not 0{40}?\"\n> You counted the number of reasons why 0{40} is used to stand in for\n> a real value, but that was the number I didn't find interesting in\n> the scope of this discussion, i.e. \"why the special sample is 17?\"\n>\n> I vaguely recall we also used 0{39}1 for something else long time\n> ago; I offhand do not recall if we still do, or we got rid of it.\n\ngitk still shows changes added to the index as 0{39}1, whereas\nchanges not added yet are marked as 0{40}.\n"},{"id":"359971","messageId":"xmqqmurmobd3.fsf@gitster-ct.c.googlers.com","threadId":"9781","inReplyTo":"CAGZ79ka5kKrSqPCWFMDetRLYxDqcguJUzJXDex9q-VMwT-ABAw@mail.gmail.com","subject":"Re: What's so special about objects/17/ ?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-10-10T01:10:16Z","receivedAt":"2018-10-10T01:10:22Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Stefan Beller <sbeller@google.com> writes:\n>> Oh, I think I misled you by saying \"more important\".\n>> ...\n> I do challenge the decision to take a hardcoded value, though, ...\n\nI do not find any reason why you need to say \"though\" here.  If you\nunderstood the message you are responding to that use of hardcoded\nvalue was chosen not to help the end-user experience, it should have\nbeen clear that we are in agreement.\n\nI also sometimes find certain people here are unnecessarily\ncombative in their discussion.  It this just some language issue?\n\n\n"},{"id":"360058","messageId":"CAGZ79kbz6MmmCYGbNaK7hr9U6Kaq9ANiw69M2_WPdMvX_w3o6w@mail.gmail.com","threadId":"9781","inReplyTo":"xmqqmurmobd3.fsf@gitster-ct.c.googlers.com","subject":"Re: What's so special about objects/17/ ?","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2018-10-10T19:08:26Z","receivedAt":"2018-10-10T19:08:42Z","isPatch":false,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Tue, Oct 9, 2018 at 6:10 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Stefan Beller <sbeller@google.com> writes:\n> >> Oh, I think I misled you by saying \"more important\".\n> >> ...\n> > I do challenge the decision to take a hardcoded value, though, ...\n>\n> I do not find any reason why you need to say \"though\" here.\n\nI caught myself using lots of filler-words lately.\n\n  Though, however, I think, I would guess, IMHO....\n  fills a lot of space without saying much.\n\nI'll reduce that.\n\n>  If you\n> understood the message you are responding to that use of hardcoded\n> value was chosen not to help the end-user experience, it should have\n> been clear that we are in agreement.\n\nWe are, but for different reasons.\n\n> I also sometimes find certain people here are unnecessarily\n> combative in their discussion.  It this just some language issue?\n\ncertain people? ;-)\nI have issues with ambiguity in communication directed towards me,\nwhich is why I sometimes try to be very direct and blunt.\nOther times I strive on ambiguity as well (mostly in my reviews).\n"}]}