{"thread":{"id":"11241","subject":"git annotate runs out of memory","startedAt":"2007-12-11T17:33:56Z","lastAt":"2007-12-18T00:05:49Z","messageCount":51,"participants":["Daniel Berlin","Nicolas Pitre","Marco Costalba","Linus Torvalds","Matthieu Moy","Daniel Barkalow","Jason Sewall","Steven Grimm","Pierre Habouzit","Junio C Hamano","Jakub Narebski","Jon Smirl","Davide Libenzi","Shawn O. Pearce","Jeff King","Florian Weimer","Jan Hudec"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"62751","messageId":"4aca3dc20712110933i636342fbifb15171d3e3cafb3@mail.gmail.com","threadId":"11241","inReplyTo":null,"subject":"git annotate runs out of memory","fromName":"Daniel Berlin","fromEmail":"dberlin@dberlin.org","sentAt":"2007-12-11T17:33:56Z","receivedAt":"2007-12-11T17:33:56Z","isPatch":false,"sender":{"key":"dberlin@dberlin.org","avatar":null},"body":"On the gcc repository (which is now a 234 meg pack for me), git\nannotate ChangeLog takes > 800 meg of memory (I stopped it at about\n1.6 gig, since it started swapping my machine).\nI assume it will run out of memory.  I stopped it after 2 minutes.\n\nMercurial, on the same file, takes 50 meg and 30 seconds.\n\n\ngit annotate fold-const.c takes 300 meg of memory and takes > 30 seconds.\nMercurial, on the same file takes 50 meg of memory and 10 seconds.\nsvn takes 15 seconds and 20 meg of memory.\n\nI have excluded the mmap memory from mmap'ing the pack/file (in\ngit/mercurial respectively).\n\nAnnotate is treasured by gcc developers (this was a key sticking point\nin svn conversion).\nHaving an annotate that is 2x slower and takes 15x memory would not\nfly (regardless of how good the results are).\n\nThis seems to be a common problem with git. It seems to use a lot of\nmemory to perform common operations on the gcc repository (even though\nit is faster in some cases than hg).\n\n--Dan\n"},{"id":"62753","messageId":"alpine.LFD.0.99999.0712111245260.555@xanadu.home","threadId":"11241","inReplyTo":"4aca3dc20712110933i636342fbifb15171d3e3cafb3@mail.gmail.com","subject":"Re: git annotate runs out of memory","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-11T17:47:41Z","receivedAt":"2007-12-11T17:47:41Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 11 Dec 2007, Daniel Berlin wrote:\n\n> On the gcc repository (which is now a 234 meg pack for me), git\n> annotate ChangeLog takes > 800 meg of memory (I stopped it at about\n> 1.6 gig, since it started swapping my machine).\n> I assume it will run out of memory.  I stopped it after 2 minutes.\n\nAnd I bet this is the exact same issue as the repack one.\n\nDo you still have the 2.1GB pack around?  I bet annotate would eat much \nless memory in that case.\n\n\nNicolas\n"},{"id":"62754","messageId":"4aca3dc20712110953h13e3c33ftb310609bbac6a0a8@mail.gmail.com","threadId":"11241","inReplyTo":"alpine.LFD.0.99999.0712111245260.555@xanadu.home","subject":"Re: git annotate runs out of memory","fromName":"Daniel Berlin","fromEmail":"dberlin@dberlin.org","sentAt":"2007-12-11T17:53:30Z","receivedAt":"2007-12-11T17:53:30Z","isPatch":false,"sender":{"key":"dberlin@dberlin.org","avatar":null},"body":"On 12/11/07, Nicolas Pitre <nico@cam.org> wrote:\n> On Tue, 11 Dec 2007, Daniel Berlin wrote:\n>\n> > On the gcc repository (which is now a 234 meg pack for me), git\n> > annotate ChangeLog takes > 800 meg of memory (I stopped it at about\n> > 1.6 gig, since it started swapping my machine).\n> > I assume it will run out of memory.  I stopped it after 2 minutes.\n>\n> And I bet this is the exact same issue as the repack one.\n>\n> Do you still have the 2.1GB pack around?  I bet annotate would eat much\n> less memory in that case.\n\nI do not, but i could remake it in a few days if it would help\n"},{"id":"62756","messageId":"alpine.LFD.0.99999.0712111259120.555@xanadu.home","threadId":"11241","inReplyTo":"4aca3dc20712110953h13e3c33ftb310609bbac6a0a8@mail.gmail.com","subject":"Re: git annotate runs out of memory","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-11T18:01:41Z","receivedAt":"2007-12-11T18:01:41Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 11 Dec 2007, Daniel Berlin wrote:\n\n> On 12/11/07, Nicolas Pitre <nico@cam.org> wrote:\n> > On Tue, 11 Dec 2007, Daniel Berlin wrote:\n> >\n> > > On the gcc repository (which is now a 234 meg pack for me), git\n> > > annotate ChangeLog takes > 800 meg of memory (I stopped it at about\n> > > 1.6 gig, since it started swapping my machine).\n> > > I assume it will run out of memory.  I stopped it after 2 minutes.\n> >\n> > And I bet this is the exact same issue as the repack one.\n> >\n> > Do you still have the 2.1GB pack around?  I bet annotate would eat much\n> > less memory in that case.\n> \n> I do not, but i could remake it in a few days if it would help\n\nWell, depending on the amount of RAM in your machine, you might even not \nbe able to remake it at the moment.  I currently can't reproduce it \nmyself due to the same out-of-memory issue.\n\n\nNicolas\n"},{"id":"62761","messageId":"e5bfff550712111032p60fedbedu304cab834ce86eb9@mail.gmail.com","threadId":"11241","inReplyTo":"4aca3dc20712110933i636342fbifb15171d3e3cafb3@mail.gmail.com","subject":"Re: git annotate runs out of memory","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2007-12-11T18:32:12Z","receivedAt":"2007-12-11T18:32:12Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"On Dec 11, 2007 6:33 PM, Daniel Berlin <dberlin@dberlin.org> wrote:\n>\n> Annotate is treasured by gcc developers (this was a key sticking point\n> in svn conversion).\n> Having an annotate that is 2x slower and takes 15x memory would not\n> fly (regardless of how good the results are).\n>\n\nSpeed of annotation is mainly due to getting the file history more\nthen calculating the actual annotation.\n\nI don't know *how* file history is stored in the others scm, perhaps\nis easier to retrieve, i.e. without a full walk across the\nrevisions...\n\nIn case you have qgit (especially the 2.0 version that is much faster\nin this feature) I would be very interested to have annotation times\non this file. Indeed annotation times are shown splitted between file\nhistory retrieval, based on something along the lines of \"git log -p\n-- <path>\", and actual annotation calculation (fully internal at\nqgit).\n\nI would be interested in cold start and warm cache start (close the\nannotation tab and start annotation again).\n\n\nThanks (a lot)\nMarco\n"},{"id":"62762","messageId":"alpine.LFD.0.9999.0712111018540.25032@woody.linux-foundation.org","threadId":"11241","inReplyTo":"4aca3dc20712110933i636342fbifb15171d3e3cafb3@mail.gmail.com","subject":"Re: git annotate runs out of memory","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-11T18:40:36Z","receivedAt":"2007-12-11T18:40:36Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 11 Dec 2007, Daniel Berlin wrote:\n>\n> This seems to be a common problem with git. It seems to use a lot of\n> memory to perform common operations on the gcc repository (even though\n> it is faster in some cases than hg).\n\nThe thing is, git has a very different notion of \"common operations\" than \nyou do.\n\nTo git, \"git annotate\" is just about the *last* thing you ever want to do. \nIt's not a common operation, it's a \"last resort\" operation. In git, the \nwhole workflow is designed for \"git log -p <pathnamepattern>\" rather than \nannotate/blame.\n\nIn fact, we didn't support annotate at all for the first year or so of \ngit.\n\nThe reason for git being relatively slow is exactly that git doesn't have \n\"file history\" at all, and only tracks full snapshots. So \"git blame\" is \nreally a very complex operation that basically looks at the global history \n(because nothing else exists) and will basically generate a totally \ndifferent \"view\" of local history from that one.\n\nThe disadvantage is that it's much slower and much more costly than just \nhaving a local history view to begin with.\n\nHowever, the absolutely *huge* advantage is that it isn't then limited to \nlocal history.\n\nSo where git shines is when you actually use the global history, and do \nmerges or when you track more than one file (which others find hard, but \ngit finds much more natural).\n\nAn examples of this is content that actually comes from multiple files. \nFile-based systems simply cannot do this at all. They aren't just slower, \nthey are totally unable to do it sanely. For git, it's all the same: it \nnever really cares about file boundaries in the first place.\n\nThe other example is doing things like \"git log -p drivers/char\", where \nyou don't ask for the log of a single file, but a general file pattern, \nand get (still atomic!) commits as the result.\n\nAnd perhaps the best example is just tracking code when you have two files \nthat merge into one (possibly because the \"same\" file was created \nindependently in two different branches). git gets things like that right \nwithout even thinking about it. Others tend to just flounder about and \ncan't do anything at all about it.\n\nThat said, I'll see if I can speed up \"git blame\" on the gcc repository. \nIt _is_ a fundamentally much more expensive operation than it is for \nsystems that do single-file things.\n\n\t\t\tLinus\n"},{"id":"62765","messageId":"vpq4pepcaz5.fsf@bauges.imag.fr","threadId":"11241","inReplyTo":"alpine.LFD.0.9999.0712111018540.25032@woody.linux-foundation.org","subject":"Re: git annotate runs out of memory","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2007-12-11T19:01:02Z","receivedAt":"2007-12-11T19:01:02Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> The other example is doing things like \"git log -p drivers/char\", where \n> you don't ask for the log of a single file, but a general file pattern, \n> and get (still atomic!) commits as the result.\n\nI've seen you pointing this kind of examples many times, but is that\nreally different from what even SVN does? \"svn log drivers/char\" will\nalso list atomic commits, and give me a filtered view of the global\nlog.\n\nSo, yes, that's cool, but I don't see a real difference between git\nand almost anything else (except CVS which really got this wrong, no\nbig surprise).\n\n-- \nMatthieu\n"},{"id":"62766","messageId":"4aca3dc20712111103s1af3b045h484ea749378c6282@mail.gmail.com","threadId":"11241","inReplyTo":"e5bfff550712111032p60fedbedu304cab834ce86eb9@mail.gmail.com","subject":"Re: git annotate runs out of memory","fromName":"Daniel Berlin","fromEmail":"dberlin@dberlin.org","sentAt":"2007-12-11T19:03:49Z","receivedAt":"2007-12-11T19:03:49Z","isPatch":false,"sender":{"key":"dberlin@dberlin.org","avatar":null},"body":"On 12/11/07, Marco Costalba <mcostalba@gmail.com> wrote:\n> On Dec 11, 2007 6:33 PM, Daniel Berlin <dberlin@dberlin.org> wrote:\n> >\n> > Annotate is treasured by gcc developers (this was a key sticking point\n> > in svn conversion).\n> > Having an annotate that is 2x slower and takes 15x memory would not\n> > fly (regardless of how good the results are).\n> >\n>\n> Speed of annotation is mainly due to getting the file history more\n> then calculating the actual annotation.\n>\n\nYes, i figured as much.\n\n> I don't know *how* file history is stored in the others scm, perhaps\n> is easier to retrieve, i.e. without a full walk across the\n> revisions...\n\nIt is stored in an easier format. However, can you not simply provide\nside-indexes to do the annotation?\n\nI guess that own't work in git because you can change history (in\nother scm's, history is readonly so you could know the results for\ncommitted revisions will never change).\n\n> I would be interested in cold start and warm cache start (close the\n> annotation tab and start annotation again).\n\nI will try to do this.\n"},{"id":"62767","messageId":"alpine.LFD.0.99999.0712111403080.555@xanadu.home","threadId":"11241","inReplyTo":"alpine.LFD.0.9999.0712111018540.25032@woody.linux-foundation.org","subject":"Re: git annotate runs out of memory","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-11T19:06:15Z","receivedAt":"2007-12-11T19:06:15Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 11 Dec 2007, Linus Torvalds wrote:\n\n> That said, I'll see if I can speed up \"git blame\" on the gcc repository. \n> It _is_ a fundamentally much more expensive operation than it is for \n> systems that do single-file things.\n\nIt has no excuse for eating up to 1.6GB or RAM though.  That's plainly \nwrong.\n\n\nNicolas\n"},{"id":"62768","messageId":"4aca3dc20712111109y5d74a292rf29be6308932393c@mail.gmail.com","threadId":"11241","inReplyTo":"alpine.LFD.0.9999.0712111018540.25032@woody.linux-foundation.org","subject":"Re: git annotate runs out of memory","fromName":"Daniel Berlin","fromEmail":"dberlin@dberlin.org","sentAt":"2007-12-11T19:09:03Z","receivedAt":"2007-12-11T19:09:03Z","isPatch":false,"sender":{"key":"dberlin@dberlin.org","avatar":null},"body":"On 12/11/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>\n>\n> On Tue, 11 Dec 2007, Daniel Berlin wrote:\n> >\n> > This seems to be a common problem with git. It seems to use a lot of\n> > memory to perform common operations on the gcc repository (even though\n> > it is faster in some cases than hg).\n>\n> The thing is, git has a very different notion of \"common operations\" than\n> you do.\n>\n> To git, \"git annotate\" is just about the *last* thing you ever want to do.\n> It's not a common operation, it's a \"last resort\" operation. In git, the\n> whole workflow is designed for \"git log -p <pathnamepattern>\" rather than\n> annotate/blame.\n>\nI understand this, and completely agree with you.\nHowever, I cannot force GCC people to adopt completely new workflow in\nthis regard.\nThe changelog's are not useful enough (and we've had huge fights over\nthis) to do git log -p and figure out the info we want.\nLooking through thousands of diffs to find the one that happened to\nyour line is also pretty annoying.\nAnnotate is a major use for gcc developers as a result\nI wish I could fix this silliness, but i can't :)\n\n> That said, I'll see if I can speed up \"git blame\" on the gcc repository.\n> It _is_ a fundamentally much more expensive operation than it is for\n> systems that do single-file things.\n\nSVN had the same problem (the file retrieval was the most expensive op\non FSFS). One of the things i did to speed it up tremendously was to\ndo the annotate from newest to oldest (IE in reverse), and stop\nannotating when we had come up with annotate info for all the lines.\nIf you can't speed up file retrieval itself, you can make it need less\nfiles :)\nIn GCC history, it is likely you will be able to cut off at least 30%\nof the time if you do this, because files often have changed entirely\nmultiple times.\n"},{"id":"62770","messageId":"e5bfff550712111114s4e9c31cxb7aed4da70d23382@mail.gmail.com","threadId":"11241","inReplyTo":"4aca3dc20712111103s1af3b045h484ea749378c6282@mail.gmail.com","subject":"Re: git annotate runs out of memory","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2007-12-11T19:14:41Z","receivedAt":"2007-12-11T19:14:41Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"On Dec 11, 2007 8:03 PM, Daniel Berlin <dberlin@dberlin.org> wrote:\n>\n> > I don't know *how* file history is stored in the others scm, perhaps\n> > is easier to retrieve, i.e. without a full walk across the\n> > revisions...\n>\n> It is stored in an easier format. However, can you not simply provide\n> side-indexes to do the annotation?\n>\n> I guess that own't work in git because you can change history (in\n> other scm's, history is readonly so you could know the results for\n> committed revisions will never change).\n>\n\nAs Linus pointed out annotation in git is \"much slower and much more\ncostly than just\nhaving a local history view to begin with\".\n\nIndeed to annotate say kernel/sched.c\n\nthe time is spent by git while executing\n\ngit log -p -- kernel/sched.c\n\ncould be also 10X higher the the following annotation processing time\nstarting from the git log output.\n\nUnfortunately my knowledge of git internals falls far far shorter then\nguessing what could be done to increase the *one file* history case\nthat _seems_ to be the common one.\n\n\n> > I would be interested in cold start and warm cache start (close the\n> > annotation tab and start annotation again).\n>\n> I will try to do this.\n>\n\nThanks. Very appreciated.\n"},{"id":"62773","messageId":"alpine.LFD.0.9999.0712111119310.25032@woody.linux-foundation.org","threadId":"11241","inReplyTo":"vpq4pepcaz5.fsf@bauges.imag.fr","subject":"Re: git annotate runs out of memory","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-11T19:22:20Z","receivedAt":"2007-12-11T19:22:20Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 11 Dec 2007, Matthieu Moy wrote:\n> \n> I've seen you pointing this kind of examples many times, but is that\n> really different from what even SVN does? \"svn log drivers/char\" will\n> also list atomic commits, and give me a filtered view of the global\n> log.\n\nOk, BK and CVS both got this horribly wrong, which is why I care. Maybe \nthis is one of the things SVN gets right.\n\nI seriously doubt it, though. Do you get *history* right, or do you just \nget a random list of commits?\n\nOf course, to see the difference, you need to do \"gitk drivers/char\" or \nuse another of the log viewers that actually show you history too. A plain \n\"git log\" won't make it obvious (unless you actually ask for parent \ninformation and then just track the history in your head, in which case \nyou don't really need an SCM in the first place ;)\n\n\t\t\tLinus\n"},{"id":"62774","messageId":"4aca3dc20712111124y1d9171eem4d2c4f0872703786@mail.gmail.com","threadId":"11241","inReplyTo":"alpine.LFD.0.9999.0712111119310.25032@woody.linux-foundation.org","subject":"Re: git annotate runs out of memory","fromName":"Daniel Berlin","fromEmail":"dberlin@dberlin.org","sentAt":"2007-12-11T19:24:54Z","receivedAt":"2007-12-11T19:24:54Z","isPatch":false,"sender":{"key":"dberlin@dberlin.org","avatar":null},"body":"On 12/11/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>\n>\n> On Tue, 11 Dec 2007, Matthieu Moy wrote:\n> >\n> > I've seen you pointing this kind of examples many times, but is that\n> > really different from what even SVN does? \"svn log drivers/char\" will\n> > also list atomic commits, and give me a filtered view of the global\n> > log.\n>\n> Ok, BK and CVS both got this horribly wrong, which is why I care. Maybe\n> this is one of the things SVN gets right.\n>\n> I seriously doubt it, though. Do you get *history* right, or do you just\n> get a random list of commits?\n\nNo, it will get actual history (IE not just things that happen to have\nthat path in the repository)\n"},{"id":"62775","messageId":"Pine.LNX.4.64.0712111410370.5349@iabervon.org","threadId":"11241","inReplyTo":"4aca3dc20712111109y5d74a292rf29be6308932393c@mail.gmail.com","subject":"Re: git annotate runs out of memory","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2007-12-11T19:26:46Z","receivedAt":"2007-12-11T19:26:46Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Tue, 11 Dec 2007, Daniel Berlin wrote:\n\n> On 12/11/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> >\n> >\n> > On Tue, 11 Dec 2007, Daniel Berlin wrote:\n> > >\n> > > This seems to be a common problem with git. It seems to use a lot of\n> > > memory to perform common operations on the gcc repository (even though\n> > > it is faster in some cases than hg).\n> >\n> > The thing is, git has a very different notion of \"common operations\" than\n> > you do.\n> >\n> > To git, \"git annotate\" is just about the *last* thing you ever want to do.\n> > It's not a common operation, it's a \"last resort\" operation. In git, the\n> > whole workflow is designed for \"git log -p <pathnamepattern>\" rather than\n> > annotate/blame.\n> >\n> I understand this, and completely agree with you.\n> However, I cannot force GCC people to adopt completely new workflow in\n> this regard.\n> The changelog's are not useful enough (and we've had huge fights over\n> this) to do git log -p and figure out the info we want.\n> Looking through thousands of diffs to find the one that happened to\n> your line is also pretty annoying.\n> Annotate is a major use for gcc developers as a result\n> I wish I could fix this silliness, but i can't :)\n> \n> > That said, I'll see if I can speed up \"git blame\" on the gcc repository.\n> > It _is_ a fundamentally much more expensive operation than it is for\n> > systems that do single-file things.\n> \n> SVN had the same problem (the file retrieval was the most expensive op\n> on FSFS). One of the things i did to speed it up tremendously was to\n> do the annotate from newest to oldest (IE in reverse), and stop\n> annotating when we had come up with annotate info for all the lines.\n> If you can't speed up file retrieval itself, you can make it need less\n> files :)\n> In GCC history, it is likely you will be able to cut off at least 30%\n> of the time if you do this, because files often have changed entirely\n> multiple times.\n\nUnfortunately, we're doing that already. One improvement that is already \navailable is that we can do progressive annotate: we can output lines we \nfind in the order we find them, such that lines that changed recently \n(which are usually the more interesting ones) get annotated quicker. \nObviously, you need a GUI-ish thing to do this, because pagers don't like \nhaving stuff written out of order, but there's a good chance that a user \nannotating fold-const.c will have the info for the interesting lines in a \nfew seconds, and go on while git is still trying to find where the boring \nold lines came from.\n\nThere's also the possibility of generating caches of commit:file pairs \nyou've annotated, which would make generating the annotation for something \nyou'd annotated for a recent commit blindingly fast.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"62777","messageId":"31e9dd080712111127h72ca18a4o574f3e65ff1acb16@mail.gmail.com","threadId":"11241","inReplyTo":"4aca3dc20712111103s1af3b045h484ea749378c6282@mail.gmail.com","subject":"Re: git annotate runs out of memory","fromName":"Jason Sewall","fromEmail":"jasonsewall@gmail.com","sentAt":"2007-12-11T19:27:37Z","receivedAt":"2007-12-11T19:27:37Z","isPatch":false,"sender":{"key":"jasonsewall@gmail.com","avatar":null},"body":"On Dec 11, 2007 2:03 PM, Daniel Berlin <dberlin@dberlin.org> wrote:\n> It is stored in an easier format. However, can you not simply provide\n> side-indexes to do the annotation?\n>\n> I guess that own't work in git because you can change history (in\n> other scm's, history is readonly so you could know the results for\n> committed revisions will never change).\n>\n\nI don't know how other scms work, but history is definitely readonly\nin git - whatever sha1 you have that describes a commit was calculated\nbased on its ancestor commits.\n\nIf you have a commit's id, it will *always* refer to the same thing -\na tree state and its complete ancestry.\n"},{"id":"62778","messageId":"74ED838F-4966-42A9-BC8A-906FD0B4B46F@midwinter.com","threadId":"11241","inReplyTo":"alpine.LFD.0.9999.0712111018540.25032@woody.linux-foundation.org","subject":"Re: git annotate runs out of memory","fromName":"Steven Grimm","fromEmail":"koreth@midwinter.com","sentAt":"2007-12-11T19:29:59Z","receivedAt":"2007-12-11T19:29:59Z","isPatch":false,"sender":{"key":"koreth@midwinter.com","avatar":"https://gravatar.com/avatar/71b4d2e8b62f168bdc9e9205341159e3567003b4f9e2127c617c5fa0a1f5bad2?d=mp&s=160"},"body":"On Dec 11, 2007, at 10:40 AM, Linus Torvalds wrote:\n> To git, \"git annotate\" is just about the *last* thing you ever want  \n> to do.\n> It's not a common operation, it's a \"last resort\" operation. In git,  \n> the\n> whole workflow is designed for \"git log -p <pathnamepattern>\" rather  \n> than\n> annotate/blame.\n\nMy use of \"git blame\" is perhaps not typical, but I use it fairly  \noften when I'm looking at a part of my company's code base that I'm  \nnot terribly familiar with. I've found it's the fastest way to figure  \nout who to go ask about a particular block of code that I think is  \nresponsible for a bug, or more commonly, who to ask to review a change  \nI'm making.\n\n\"git log\" is too coarse-grained to be useful for that purpose; it  \nusually doesn't tell me which of the 500 revisions to the file I'm  \nlooking at introduced the actual line of code I want to change.\n\nTo me that really has nothing whatsoever to do with git workflow or  \nsvn workflow; it happens well before I'm ready to do any kind of  \nintegration or commit or even, sometimes, before I've made any changes  \nto any code at all.\n\nGiven infinite spare time, one of the things I'd be strongly tempted  \nto try to build would be some kind of blame cache. You could  \ntheoretically make blame pretty much instantaneous by doing something  \nas simple as caching the per-line revision ID for each file in each  \nrevision in a shadow repository (or a shadow branch in the main repo)  \nand keeping a map between shadow-repo revisions and real-repo ones. If  \nthe cache was of the form \"one SHA1 hash per line in the original  \nfile\" it would delta-compress pretty well. It'd be easy to update  \nincrementally since you only need to walk back in history until you  \nget to the most recently cached revision for each file, at which point  \nyou use the cached value for all the lines that haven't changed.\n\nYeah, I know, code talks louder than words...\n\n-Steve\n"},{"id":"62780","messageId":"20071211193407.GC20644@artemis.madism.org","threadId":"11241","inReplyTo":"4aca3dc20712111109y5d74a292rf29be6308932393c@mail.gmail.com","subject":"Re: git annotate runs out of memory","fromName":"Pierre Habouzit","fromEmail":"madcoder@debian.org","sentAt":"2007-12-11T19:34:07Z","receivedAt":"2007-12-11T19:34:07Z","isPatch":false,"sender":{"key":"madcoder@debian.org","avatar":"https://avatars.githubusercontent.com/u/44708?v=4"},"body":"On Tue, Dec 11, 2007 at 07:09:03PM +0000, Daniel Berlin wrote:\n> On 12/11/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> >\n> >\n> > On Tue, 11 Dec 2007, Daniel Berlin wrote:\n> > >\n> > > This seems to be a common problem with git. It seems to use a lot of\n> > > memory to perform common operations on the gcc repository (even though\n> > > it is faster in some cases than hg).\n> >\n> > The thing is, git has a very different notion of \"common operations\" than\n> > you do.\n> >\n> > To git, \"git annotate\" is just about the *last* thing you ever want to do.\n> > It's not a common operation, it's a \"last resort\" operation. In git, the\n> > whole workflow is designed for \"git log -p <pathnamepattern>\" rather than\n> > annotate/blame.\n> >\n> I understand this, and completely agree with you.\n> However, I cannot force GCC people to adopt completely new workflow in\n> this regard.\n> The changelog's are not useful enough (and we've had huge fights over\n> this) to do git log -p and figure out the info we want.\n\n> Looking through thousands of diffs to find the one that happened to\n> your line is also pretty annoying.\n\n  If the question you want to answer is \"what happened to that line\"\nthen using git annotate is using a big hammer for no good reason.\n\ngit log -S'<put the content of the line here>' -- path/to/file.c\n\nwill give you the very same answer, pointing you to the changes that\nadded or removed that line directly. It's not a fast command either, but\nit should be less resource hungry than annotate that has to do roughly\nthe same for all lines whereas you're interested in one only.\n\nThe direct plus here, is that git log output is incremental, so you have\nanswers about the first diffs quite quick, which let you examine the\nfirst answers while the rest is still being computed.\n\nUnlike git annotate, this also allow you to restrict the revisions\nwhere it searches to a range where you know this happened, which makes\nit almost instantaneous in most cases.\n\nOf course, if the line is '    free(p);\\n' then you will probably have\nquite a few false positives, but with the path restriction, I assume\nthis will still be quite accurate.\n\nWhat is important here is to know what is the real question the GCC\nprogrammers want to answer to. It seems to me that `blame` is an\noverkill for the underlying issue.\n\n\nNote that it does not justifies the current memory consumption that just\nlooks bad and wrong to me, but this aims at finding a way to answer your\nquestion doing just what you need to answer it and not gazillions of\nother things :)\n-- \n·O·  Pierre Habouzit\n··O                                                madcoder@debian.org\nOOO                                                http://www.madism.org\n"},{"id":"62782","messageId":"alpine.LFD.0.9999.0712111122400.25032@woody.linux-foundation.org","threadId":"11241","inReplyTo":"4aca3dc20712111109y5d74a292rf29be6308932393c@mail.gmail.com","subject":"Re: git annotate runs out of memory","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-11T19:42:01Z","receivedAt":"2007-12-11T19:42:01Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 11 Dec 2007, Daniel Berlin wrote:\n>\n> I understand this, and completely agree with you.\n> However, I cannot force GCC people to adopt completely new workflow in\n> this regard.\n\nOh, I agree. It's why we do have \"git blame\" these days, and it's why I've \ntried to make people use the nicer incremental mode, which is not at all \nfaster, but it's a hell of a lot more pleasant to use because you get some \noutput immediately.\n\nIn other words,\n\n\tgit blame gcc/ChangeLog\n\nis virtually useless because it's too expensive, but try doing\n\n\tgit gui blame gcc ChangeLog\n\ninstead, and doesn't that just seem nicer? (*)\n\nThe difference is that the GUI one does it incrementally, and doesn't have \nto get _all_ the results before it can start reporting blame.\n\nNot that I claim that the gui blame is perfect either (I dunno why it \ndelays the nice coloring so long, for example), but it was something I \npushed - and others made the gui for - exactly to help people with the \nfact that git interally really does it that incremental way.\n\n> SVN had the same problem (the file retrieval was the most expensive op\n> on FSFS). One of the things i did to speed it up tremendously was to\n> do the annotate from newest to oldest (IE in reverse), and stop\n> annotating when we had come up with annotate info for all the lines.\n\nWe do that. The expense for git is that we don't do the revisions as a \nsingle file at all. We'll look through each commit, check whether the \n\"gcc\" directory changed, if it did, we'll go into it, and check whether \nthe \"ChangeLog\" file changed - and if it did, we'll actually diff it \nagainst the previous version.\n\n> In GCC history, it is likely you will be able to cut off at least 30%\n> of the time if you do this, because files often have changed entirely\n> multiple times.\n\nNot gcc/ChangeLog, though (apart from the renames that happen \noccasionally).\n\nBtw, an example of something git *should* do right, but is just too damn \nexpensive, is doing\n\n\tgit gui blame gcc/ChangeLog-2000\n\nand have it actually be able to track the original source of each of those \nannotations across that \"ChangeLog split from hell\". \n\nI bet it would eventually get it right, but that's a large file, way back \nin history, and it will try to do a non-whitespace blame with copy \ndetection.\n\nThat's *expensive*, although it is an amusing thing to try to do ;)\n\n\t\t\tLinus\n\nPS. I also do agree that we seem to use an excessive amount of memory \nthere. As to whether it's the same issue or not, I'd not go as far as Nico \nand say \"yes\" yet. But it's interesting.\n\nIt's not entirely surprising that we see multiple issues with the gcc \nrepo, simply because it's not the kind of repo that people have ever \nreally worked on. So I don't think it's necessarily related at all, except \nin the sense of it being a different load and showing issues.\n"},{"id":"62783","messageId":"20071211194238.GD20644@artemis.madism.org","threadId":"11241","inReplyTo":"4aca3dc20712111124y1d9171eem4d2c4f0872703786@mail.gmail.com","subject":"Re: git annotate runs out of memory","fromName":"Pierre Habouzit","fromEmail":"madcoder@debian.org","sentAt":"2007-12-11T19:42:38Z","receivedAt":"2007-12-11T19:42:38Z","isPatch":false,"sender":{"key":"madcoder@debian.org","avatar":"https://avatars.githubusercontent.com/u/44708?v=4"},"body":"On Tue, Dec 11, 2007 at 07:24:54PM +0000, Daniel Berlin wrote:\n> On 12/11/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> >\n> >\n> > On Tue, 11 Dec 2007, Matthieu Moy wrote:\n> > >\n> > > I've seen you pointing this kind of examples many times, but is that\n> > > really different from what even SVN does? \"svn log drivers/char\" will\n> > > also list atomic commits, and give me a filtered view of the global\n> > > log.\n> >\n> > Ok, BK and CVS both got this horribly wrong, which is why I care. Maybe\n> > this is one of the things SVN gets right.\n> >\n> > I seriously doubt it, though. Do you get *history* right, or do you just\n> > get a random list of commits?\n> \n> No, it will get actual history (IE not just things that happen to have\n> that path in the repository)\n\nOTOH svn has the result right, but the way it does that is horrible.\nWhen you svn log some/path, I think it just (basically) ask svn log for\neach file in that directory, and merge the logs together. This is \"easy\"\nfor svn since it remembers \"where this specific file\" came from.\n\nSo for svn it's just a matter of merging the individual files histories\ntogether. It may have a more clever implementation, but basically I\nbelieve it would be similar to that in the end.\n\nOf course, if you do something as stupid as:\n  svn cp Makefile some/path/foo.c\n  # completely rewrite foo.c\n  svn commit\nthen you'll have the history of `Makefile` melded into the\nsome/path/foo.c svn log, which is completely horribly wrong.\n\nor if you do (which unlike the previous example isn't silly for so\nmany good reasons):\n  cp bar.c foo.c\n  svn add foo.c\n  svn commit\nthen foo.c won't have bar.c history in its svn log.\n\n-- \n·O·  Pierre Habouzit\n··O                                                madcoder@debian.org\nOOO                                                http://www.madism.org\n"},{"id":"62784","messageId":"Pine.LNX.4.64.0712111428040.5349@iabervon.org","threadId":"11241","inReplyTo":"4aca3dc20712111103s1af3b045h484ea749378c6282@mail.gmail.com","subject":"Re: git annotate runs out of memory","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2007-12-11T19:46:09Z","receivedAt":"2007-12-11T19:46:09Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Tue, 11 Dec 2007, Daniel Berlin wrote:\n\n> It is stored in an easier format. However, can you not simply provide\n> side-indexes to do the annotation?\n> \n> I guess that own't work in git because you can change history (in\n> other scm's, history is readonly so you could know the results for\n> committed revisions will never change).\n\nHistory in git is read-only. It's just that git lets you fork and move \nforward with something different. Each commit can never change (and, in \nfact, you'd have to badly break SHA1 to change it), but which commits are \nrelevant to the history can change.\n\nKeeping extra information is fine; at worst, it'll go irrelevant.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"62786","messageId":"alpine.LFD.0.9999.0712111146200.25032@woody.linux-foundation.org","threadId":"11241","inReplyTo":"alpine.LFD.0.9999.0712111122400.25032@woody.linux-foundation.org","subject":"Re: git annotate runs out of memory","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-11T19:50:08Z","receivedAt":"2007-12-11T19:50:08Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 11 Dec 2007, Linus Torvalds wrote:\n> \n> We do that. The expense for git is that we don't do the revisions as a \n> single file at all. We'll look through each commit, check whether the \n> \"gcc\" directory changed, if it did, we'll go into it, and check whether \n> the \"ChangeLog\" file changed - and if it did, we'll actually diff it \n> against the previous version.\n\nAnd, btw: the diff is totally different from the xdelta we have, so even \nif we have an already prepared nice xdelta between the two versions, we'll \nend up re-generating the files in full, and then do a diff on the end \nresult.\n\nOf course, part of that is that git logically *never* works with deltas, \nexcept in the actual code-paths that generate objects (or generate packs, \nof course). So even if we had used a delta algorithm that would be \namenable to be turned into a diff directly, it would have been a layering \nviolation to actually do that.\n\nOther systems can sometimes just re-use their deltas to generate the \ndiffs and/or blame information. I dunno whether SVN does that. CVS does, \nafaik.\n\n\t\t\tLinus\n"},{"id":"62787","messageId":"7vve75c89q.fsf@gitster.siamese.dyndns.org","threadId":"11241","inReplyTo":"20071211193407.GC20644@artemis.madism.org","subject":"Re: git annotate runs out of memory","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-12-11T19:59:29Z","receivedAt":"2007-12-11T19:59:29Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Pierre Habouzit <madcoder@debian.org> writes:\n\n>> Looking through thousands of diffs to find the one that happened to\n>> your line is also pretty annoying.\n>\n>   If the question you want to answer is \"what happened to that line\"\n> then using git annotate is using a big hammer for no good reason.\n>\n> git log -S'<put the content of the line here>' -- path/to/file.c\n>\n> will give you the very same answer, pointing you to the changes that\n> added or removed that line directly. It's not a fast command either, but\n> it should be less resource hungry than annotate that has to do roughly\n> the same for all lines whereas you're interested in one only.\n>\n> The direct plus here, is that git log output is incremental, so you have\n> answers about the first diffs quite quick, which let you examine the\n> first answers while the rest is still being computed.\n\nYes.\n\n> Unlike git annotate, this also allow you to restrict the revisions\n> where it searches to a range where you know this happened, which makes\n> it almost instantaneous in most cases.\n\nYes, but blame also takes revision bottoms (obviously you have to start\ndigging from a single revision so \"blame master..next pu\" would not\nwork, but \"blame ^foo ^bar baz\" would).\n\n> Of course, if the line is '    free(p);\\n' then you will probably have\n> quite a few false positives,...\n\nYou can feed more than a line from -S, and the assumed and recommended\ntypical use case is to do so.\n\n> Note that it does not justifies the current memory consumption that just\n> looks bad and wrong to me,...\n\nRight.\n"},{"id":"62790","messageId":"e5bfff550712111214p189945e5g38e85e11bf7af68@mail.gmail.com","threadId":"11241","inReplyTo":"Pine.LNX.4.64.0712111428040.5349@iabervon.org","subject":"Re: git annotate runs out of memory","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2007-12-11T20:14:02Z","receivedAt":"2007-12-11T20:14:02Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"On Dec 11, 2007 8:46 PM, Daniel Barkalow <barkalow@iabervon.org> wrote:\n> On Tue, 11 Dec 2007, Daniel Berlin wrote:\n>\n> > It is stored in an easier format. However, can you not simply provide\n> > side-indexes to do the annotation?\n> >\n> > I guess that own't work in git because you can change history (in\n> > other scm's, history is readonly so you could know the results for\n> > committed revisions will never change).\n>\n> History in git is read-only. It's just that git lets you fork and move\n> forward with something different. Each commit can never change (and, in\n> fact, you'd have to badly break SHA1 to change it), but which commits are\n> relevant to the history can change.\n>\n\nWell, revisions never change, but history intended as revision's\nparent information could and do changes when you use a path delimiter.\nSo does the graph that is a direct visualization of parent\ninformation.\n\nFor a single revision (that modifies say 3 files) you can have at leat\n3 different histories and acutally more if you want to visualize also\nthe history of the directories trees that owns the modified files.\n\nYou end up with a quite big number of different histories all showing\nyour revisions in different ways, according to the path delimiter you\nuse.\n\nPerhaps the intended meaning of \"changing histories\" is this, and in\nany case is this the reason you cannot (or has no sense to do) \"save\"\na single file history in git.\n\nMarco\n"},{"id":"62791","messageId":"m3bq8xrntc.fsf@roke.D-201","threadId":"11241","inReplyTo":"74ED838F-4966-42A9-BC8A-906FD0B4B46F@midwinter.com","subject":"Re: git annotate runs out of memory","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-12-11T20:14:56Z","receivedAt":"2007-12-11T20:14:56Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Steven Grimm <koreth@midwinter.com> writes:\n> On Dec 11, 2007, at 10:40 AM, Linus Torvalds wrote:\n\n> > To git, \"git annotate\" is just about the *last* thing you ever want\n> > to do.\n> > It's not a common operation, it's a \"last resort\" operation. In git,\n> > the\n> > whole workflow is designed for \"git log -p <pathnamepattern>\" rather\n> > than\n> > annotate/blame.\n> \n> My use of \"git blame\" is perhaps not typical, but I use it fairly\n> often when I'm looking at a part of my company's code base that I'm\n> not terribly familiar with. I've found it's the fastest way to figure\n> out who to go ask about a particular block of code that I think is\n> responsible for a bug, or more commonly, who to ask to review a change\n> I'm making.\n> \n> \"git log\" is too coarse-grained to be useful for that purpose; it\n> usually doesn't tell me which of the 500 revisions to the file I'm\n> looking at introduced the actual line of code I want to change.\n\nThere is always \"pickaxe\" search, i.e. \n  $ git log -p -S'<string>' -- <file or pathspec>\nwhich can be used instead of blame (perhaps with --follow).\n\nAnd you can limit blame to the interesting region of file, and to\ninteresting (important) range of revisions.\n\n\n[about blame cache]\n\n\"git gui blame\" uses incremental blame; if only it accepted range\n(file fragment) limiting, and if \"reblame\" (blame --reference=<rev>,\nblaming incrementally only lines which changed wrt. given revision)\nwas implemented.\n\nBTW. qgit actually does blame using it's own \"multiple files bottom-up\nblame\" code (it would be nice to have it in core-git if possible,\nhint, hint), and does some caching, although I'm not sure if blame\ninfo also. You should try it, I think.\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"62794","messageId":"e5bfff550712111229i227361e9s1a6dcbed9a13019d@mail.gmail.com","threadId":"11241","inReplyTo":"4aca3dc20712111109y5d74a292rf29be6308932393c@mail.gmail.com","subject":"Re: git annotate runs out of memory","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2007-12-11T20:29:59Z","receivedAt":"2007-12-11T20:29:59Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"On Dec 11, 2007 8:09 PM, Daniel Berlin <dberlin@dberlin.org> wrote:\n>\n> In GCC history, it is likely you will be able to cut off at least 30%\n> of the time if you do this, because files often have changed entirely\n> multiple times.\n>\n\nThis could be useful for a command line tool but for a GUI the top\ndown approach is a myth IMHO.\n\nIn the GUI case what you actually end up doing (because a GUI allows\nit) is to start from the latest file version, check the code region\nyou are interested then when you find the changed lines you _may_ want\nto double click and go to see how it was the file before that change\nand then perhaps start a new digging.\n\nI found this is my typical workflow with annotation info because I'm\nmore interested not in what lines have changed but _why_ have changed\nand to do this you naturally end up digging in the past (and checking\nalso the corresponding revisions patch as example in another tab)\n\nIn this case the advantage of oldest to newest annotation algorithm is\nthat you have _already_ annotated all the history so you can walk and\ndig back and forth among the different file versions without *any*\nadditional delay.\n\nMarco\n"},{"id":"62795","messageId":"9e4733910712111231x1bbe181ew4f90fc5bb0e87039@mail.gmail.com","threadId":"11241","inReplyTo":"alpine.LFD.0.99999.0712111403080.555@xanadu.home","subject":"Re: git annotate runs out of memory","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-11T20:31:23Z","receivedAt":"2007-12-11T20:31:23Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/11/07, Nicolas Pitre <nico@cam.org> wrote:\n> On Tue, 11 Dec 2007, Linus Torvalds wrote:\n>\n> > That said, I'll see if I can speed up \"git blame\" on the gcc repository.\n> > It _is_ a fundamentally much more expensive operation than it is for\n> > systems that do single-file things.\n>\n> It has no excuse for eating up to 1.6GB or RAM though.  That's plainly\n> wrong.\n\n git blame gcc/ChangeLog\nIt needs 2.25GB of RAM to run without swapping\n\nThat is pretty close to the same number the repack needs.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62801","messageId":"4aca3dc20712111309yf43179dh6c43ac84dcaf38e8@mail.gmail.com","threadId":"11241","inReplyTo":"20071211194238.GD20644@artemis.madism.org","subject":"Re: git annotate runs out of memory","fromName":"Daniel Berlin","fromEmail":"dberlin@dberlin.org","sentAt":"2007-12-11T21:09:08Z","receivedAt":"2007-12-11T21:09:08Z","isPatch":false,"sender":{"key":"dberlin@dberlin.org","avatar":null},"body":"On 12/11/07, Pierre Habouzit <madcoder@debian.org> wrote:\n> On Tue, Dec 11, 2007 at 07:24:54PM +0000, Daniel Berlin wrote:\n> > On 12/11/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> > >\n> > >\n> > > On Tue, 11 Dec 2007, Matthieu Moy wrote:\n> > > >\n> > > > I've seen you pointing this kind of examples many times, but is that\n> > > > really different from what even SVN does? \"svn log drivers/char\" will\n> > > > also list atomic commits, and give me a filtered view of the global\n> > > > log.\n> > >\n> > > Ok, BK and CVS both got this horribly wrong, which is why I care. Maybe\n> > > this is one of the things SVN gets right.\n> > >\n> > > I seriously doubt it, though. Do you get *history* right, or do you just\n> > > get a random list of commits?\n> >\n> > No, it will get actual history (IE not just things that happen to have\n> > that path in the repository)\n>\n> OTOH svn has the result right, but the way it does that is horrible.\n> When you svn log some/path, I think it just (basically) ask svn log for\n> each file in that directory, and merge the logs together. This is \"easy\"\n> for svn since it remembers \"where this specific file\" came from.\n\nWhat?\nWe version directories too.\nWe don't do svn log for each file in the directory when you request a path.\nWe look at the history of the path, follow renames, etc.\n\nWhen you change foo/bar/fred.c, we consider it a change to foo/bar and\nfoo/, and thus, they have new versions.\n\nI'm not sure where you get this crazy notion that we do anything with\nfiles when you ask about directories.\n"},{"id":"62803","messageId":"alpine.LFD.0.9999.0712111300440.25032@woody.linux-foundation.org","threadId":"11241","inReplyTo":"alpine.LFD.0.9999.0712111122400.25032@woody.linux-foundation.org","subject":"Re: git annotate runs out of memory","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-11T21:14:18Z","receivedAt":"2007-12-11T21:14:18Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 11 Dec 2007, Linus Torvalds wrote:\n> \n> PS. I also do agree that we seem to use an excessive amount of memory \n> there. As to whether it's the same issue or not, I'd not go as far as Nico \n> and say \"yes\" yet. But it's interesting.\n\nI think the answer here is that git-annotate is a totally different issue.\n\nThe blame machinery keeps around all the blobs it has ever needed to do a \ndiff, which explains why something like gcc/ChangeLog blows up badly.\n\nTry this trivial patch.\n\nIt will cause us to potentially re-generate some blobs much more, but \nthat's a reasonably cheap operation, and our delta base cache will get the \nexpensive cases.\n\nIt's still not a free operation, but I get\n\n\t[torvalds@woody gcc]$ /usr/bin/time ~/git/git-blame gcc/ChangeLog > /dev/null\n\t20.68user 1.25system 0:21.94elapsed 100%CPU (0avgtext+0avgdata 0maxresident)k\n\t0inputs+0outputs (0major+599833minor)pagefaults 0swaps\n\nso it took 22s and I never saw it grow very large either (it grew to 72M \nresident, but I don't know how much of that was the mmap of the \npack-file, so that number is pretty meaningless). Valgrind reports that \nit used a maximum heap of about 24M, and almost all of that seems to have \nbeen in the delta cache (which is all good).\n\n\t\tLinus\n\n----\n builtin-blame.c |   10 ++++++++++\n 1 files changed, 10 insertions(+), 0 deletions(-)\n\ndiff --git a/builtin-blame.c b/builtin-blame.c\nindex c158d31..18f9924 100644\n--- a/builtin-blame.c\n+++ b/builtin-blame.c\n@@ -87,6 +87,14 @@ struct origin {\n \tchar path[FLEX_ARRAY];\n };\n \n+static void drop_origin_blob(struct origin *o)\n+{\n+\tif (o->file.ptr) {\n+\t\tfree(o->file.ptr);\n+\t\to->file.ptr = NULL;\n+\t}\n+}\n+\n /*\n  * Given an origin, prepare mmfile_t structure to be used by the\n  * diff machinery\n@@ -558,6 +566,8 @@ static struct patch *get_patch(struct origin *parent, struct origin *origin)\n \tif (!file_p.ptr || !file_o.ptr)\n \t\treturn NULL;\n \tpatch = compare_buffer(&file_p, &file_o, 0);\n+\tdrop_origin_blob(parent);\n+\tdrop_origin_blob(origin);\n \tnum_get_patch++;\n \treturn patch;\n }\n"},{"id":"62804","messageId":"4aca3dc20712111314wf4525l790120dce29a9bc5@mail.gmail.com","threadId":"11241","inReplyTo":"alpine.LFD.0.9999.0712111146200.25032@woody.linux-foundation.org","subject":"Re: git annotate runs out of memory","fromName":"Daniel Berlin","fromEmail":"dberlin@dberlin.org","sentAt":"2007-12-11T21:14:59Z","receivedAt":"2007-12-11T21:14:59Z","isPatch":false,"sender":{"key":"dberlin@dberlin.org","avatar":null},"body":"On 12/11/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>\n>\n> On Tue, 11 Dec 2007, Linus Torvalds wrote:\n> >\n> > We do that. The expense for git is that we don't do the revisions as a\n> > single file at all. We'll look through each commit, check whether the\n> > \"gcc\" directory changed, if it did, we'll go into it, and check whether\n> > the \"ChangeLog\" file changed - and if it did, we'll actually diff it\n> > against the previous version.\n>\n> And, btw: the diff is totally different from the xdelta we have, so even\n> if we have an already prepared nice xdelta between the two versions, we'll\n> end up re-generating the files in full, and then do a diff on the end\n> result.\n\nThis is what SVN does as well.\n\n>\n> Of course, part of that is that git logically *never* works with deltas,\n> except in the actual code-paths that generate objects (or generate packs,\n> of course). So even if we had used a delta algorithm that would be\n> amenable to be turned into a diff directly, it would have been a layering\n> violation to actually do that.\n\nRight. SVN has the same problem.\n\n>\n> Other systems can sometimes just re-use their deltas to generate the\n> diffs and/or blame information. I dunno whether SVN does that. CVS does,\n> afaik.\n\nCVS does because it's delta is line based, so it's easy.\n\nYou theroetically can generate blame info from SVN/GIT's block deltas,\nbut you of course, have the problem GIT does, which is that the delta\nis not meant to represent the actual changes that occurred, but\ninstead, the smallest way to reconstruct data x from data y.\nThis only sometimes has any relation to how the file actually changed\n"},{"id":"62809","messageId":"4aca3dc20712111324n51411550r6a35fbf45b4a3b49@mail.gmail.com","threadId":"11241","inReplyTo":"alpine.LFD.0.9999.0712111122400.25032@woody.linux-foundation.org","subject":"Re: git annotate runs out of memory","fromName":"Daniel Berlin","fromEmail":"dberlin@dberlin.org","sentAt":"2007-12-11T21:24:17Z","receivedAt":"2007-12-11T21:24:17Z","isPatch":false,"sender":{"key":"dberlin@dberlin.org","avatar":null},"body":"On 12/11/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>\n>\n> It's not entirely surprising that we see multiple issues with the gcc\n> repo, simply because it's not the kind of repo that people have ever\n> really worked on. So I don't think it's necessarily related at all, except\n> in the sense of it being a different load and showing issues.\n>\n\nI'm not surprised at all.\nWe had a number of issues with SVN that needed to be resolved.\nI'm basically trying to get issues worked (both on git and mercurial)\nout to the point where it is fair for our users to try their branch\nand trunk workflows with git and mercurial,  and see which they like\nmore.\n:)\n"},{"id":"62812","messageId":"alpine.LFD.0.9999.0712111323270.25032@woody.linux-foundation.org","threadId":"11241","inReplyTo":"4aca3dc20712111314wf4525l790120dce29a9bc5@mail.gmail.com","subject":"Re: git annotate runs out of memory","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-11T21:34:10Z","receivedAt":"2007-12-11T21:34:10Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 11 Dec 2007, Daniel Berlin wrote:\n> \n> You theroetically can generate blame info from SVN/GIT's block deltas,\n> but you of course, have the problem GIT does, which is that the delta\n> is not meant to represent the actual changes that occurred, but\n> instead, the smallest way to reconstruct data x from data y.\n> This only sometimes has any relation to how the file actually changed\n\nExactly. Git objects in themselves have no history or relationships, and \nbeing a delta against another object means nothing at all except for the \nfact that the data seems to resemble that other object (which has a \n_correlation_ with being related, but nothign more).\n\nAnyway, I think the git annotate memory usage was simpyl just a real bug \nthat nobody had noticed before because the memory leak wasn't all that \nnoticeable with smaller files and/or less deep histories. Can'you verify \nthat it works for you with the patch I sent out?\n\nWith that fix, I could even run \n\n\tgit blame -C gcc/ChangeLog-2000\n\nto see the blame machinery work past the strange \"combine many different \nchangelogs into year-based ones\" commit. Now, I cannot honestly claim that \nit was really *usable* (it did take three minutes to run!), but sometimes \nthose three minutes of CPU time may be worth it, if it shows the real \nhistorical context it came from. \n\nIn the case of the ChangeLog-2000 file, all the original lines obviously \ncame from older versions of a file called \"gcc/ChangeLog\", so the end \nresult doesn't really show what an involved situation it was to track the \nsources back through not just renames, but actually file splits and \nmerges. Sad, but once you know what it did it's still a bit cool to see \nthat it worked ;)\n\n\t\t\tLinus\n"},{"id":"62814","messageId":"7vprxcdhis.fsf@gitster.siamese.dyndns.org","threadId":"11241","inReplyTo":"alpine.LFD.0.9999.0712111300440.25032@woody.linux-foundation.org","subject":"Re: git annotate runs out of memory","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-12-11T21:54:19Z","receivedAt":"2007-12-11T21:54:19Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n>  builtin-blame.c |   10 ++++++++++\n>  1 files changed, 10 insertions(+), 0 deletions(-)\n>\n> diff --git a/builtin-blame.c b/builtin-blame.c\n> index c158d31..18f9924 100644\n> --- a/builtin-blame.c\n> +++ b/builtin-blame.c\n> @@ -87,6 +87,14 @@ struct origin {\n>  \tchar path[FLEX_ARRAY];\n>  };\n>  \n>  /*\n>   * Given an origin, prepare mmfile_t structure to be used by the\n>   * diff machinery\n> @@ -558,6 +566,8 @@ static struct patch *get_patch(struct origin *parent, struct origin *origin)\n>  \tif (!file_p.ptr || !file_o.ptr)\n>  \t\treturn NULL;\n>  \tpatch = compare_buffer(&file_p, &file_o, 0);\n> +\tdrop_origin_blob(parent);\n> +\tdrop_origin_blob(origin);\n>  \tnum_get_patch++;\n>  \treturn patch;\n>  }\n\nWhile this should be safe (because the user of blob lazily re-fetches),\nit feels a bit too aggressive, especially when -C or other \"retry and\ntry harder to assign blame elsewhere\" option is used.\n\nInstead, how about discarding after we are done with each origin, like\nthis?\n\n---\n builtin-blame.c |   17 +++++++++++++++--\n 1 files changed, 15 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin-blame.c b/builtin-blame.c\nindex c158d31..eda79d0 100644\n--- a/builtin-blame.c\n+++ b/builtin-blame.c\n@@ -130,6 +130,14 @@ static void origin_decref(struct origin *o)\n \t}\n }\n \n+static void drop_origin_blob(struct origin *o)\n+{\n+\tif (o->file.ptr) {\n+\t\tfree(o->file.ptr);\n+\t\to->file.ptr = NULL;\n+\t}\n+}\n+\n /*\n  * Each group of lines is described by a blame_entry; it can be split\n  * as we pass blame to the parents.  They form a linked list in the\n@@ -1274,8 +1282,13 @@ static void pass_blame(struct scoreboard *sb, struct origin *origin, int opt)\n \t\t}\n \n  finish:\n-\tfor (i = 0; i < MAXPARENT; i++)\n-\t\torigin_decref(parent_origin[i]);\n+\tfor (i = 0; i < MAXPARENT; i++) {\n+\t\tif (parent_origin[i]) {\n+\t\t\tdrop_origin_blob(parent_origin[i]);\n+\t\t\torigin_decref(parent_origin[i]);\n+\t\t}\n+\t}\n+\tdrop_origin_blob(origin);\n }\n \n /*\n"},{"id":"62829","messageId":"alpine.LFD.0.9999.0712111523210.25032@woody.linux-foundation.org","threadId":"11241","inReplyTo":"7vprxcdhis.fsf@gitster.siamese.dyndns.org","subject":"Re: git annotate runs out of memory","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-11T23:36:19Z","receivedAt":"2007-12-11T23:36:19Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 11 Dec 2007, Junio C Hamano wrote:\n> \n> Instead, how about discarding after we are done with each origin, like\n> this?\n\nSure, looks fine to me. With either of these patches, all of the cost is \nin the diffing routines:\n\n\tsamples  %        image name               app name                 symbol name\n\t191317   31.4074  git                      git                      xdl_hash_record\n\t120060   19.7096  git                      git                      xdl_recmatch\n\t99286    16.2992  git                      git                      xdl_prepare_ctx\n\t56370     9.2539  libc-2.7.so              libc-2.7.so              memcpy\n\t23315     3.8275  git                      git                      xdl_prepare_env\n\t..\n\nand while I suspect xdiff could be optimized a bit more for the cases \nwhere we have no changes at the end, that's beyond my skills.\n\n\t\tLinus\n"},{"id":"62830","messageId":"vpq63z49511.fsf@bauges.imag.fr","threadId":"11241","inReplyTo":"alpine.LFD.0.9999.0712111119310.25032@woody.linux-foundation.org","subject":"Re: git annotate runs out of memory","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2007-12-11T23:37:46Z","receivedAt":"2007-12-11T23:37:46Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Tue, 11 Dec 2007, Matthieu Moy wrote:\n>> \n>> I've seen you pointing this kind of examples many times, but is that\n>> really different from what even SVN does? \"svn log drivers/char\" will\n>> also list atomic commits, and give me a filtered view of the global\n>> log.\n>\n> Ok, BK and CVS both got this horribly wrong, which is why I care. Maybe \n> this is one of the things SVN gets right.\n>\n> I seriously doubt it, though. Do you get *history* right, or do you just \n> get a random list of commits?\n\nWell, you don't get merge commit right with SVN, but that's a\ndifferent issue (svn 1.5 is supposed to have something about merge\nhistory, I don't know how it's done ...). So, if by \"history\", you\nmean how branches interferred together, obviously, SVN is bad at this.\nBut it's equally bad at \"svn log dir/\" and plain \"svn log\".\n\nBut to simplify, if you take a linear history (no merge commits),\n\"svn log dir/\" give you the list of commits which changed something\ninside \"dir/\". As pointed out in other messages, the way it's done is\nreally different from what git does. SVN does know a lot about\ndirectories, and records a lot about them at commit time, while git\njust considers them as file containers.\n\nYear, CVS got this terribly wrong. IIRC, it just took the log for\nindividual messages, and mix them together, so a commit touching\nmultiple files would appear several times.\n\nI've taken SVN as an extreme example, but at least bzr and mercurial\nhave an approach very similar to git.\n\nSo, to me, this particular point is something git obviously got right,\nbut not a point where git is so different from the others.\n\n-- \nMatthieu\n"},{"id":"62832","messageId":"alpine.LFD.0.9999.0712111544390.25032@woody.linux-foundation.org","threadId":"11241","inReplyTo":"vpq63z49511.fsf@bauges.imag.fr","subject":"Re: git annotate runs out of memory","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-11T23:48:14Z","receivedAt":"2007-12-11T23:48:14Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 12 Dec 2007, Matthieu Moy wrote:\n>\n> > I seriously doubt it, though. Do you get *history* right, or do you just \n> > get a random list of commits?\n> \n> Well, you don't get merge commit right with SVN, but that's a\n> different issue (svn 1.5 is supposed to have something about merge\n> history, I don't know how it's done ...). So, if by \"history\", you\n> mean how branches interferred together, obviously, SVN is bad at this.\n> But it's equally bad at \"svn log dir/\" and plain \"svn log\".\n\nYeah, git just has higher goals.\n\nThe time history really matters (or rather, what I call the \"shape\" of \nhistory) is when you are trying to merge, and you get a merge conflict. \nThat's when you want to do\n\n\tgitk master merge ^merge-base -- files-that-are-unmerged\n\nand in fact this is such an important thing for me that there is a \nshorthand argument to do exactly that, ie:\n\n\tgitk --merge\n\nwhich shows the commits that touched the unmerged files graphically *with* \nthe history being correct (ie you don't just get a random log of \"these \nchanges happened\", you get the real history of the two branches as it \npertains to the files you care about!)\n\n> But to simplify, if you take a linear history (no merge commits),\n> \"svn log dir/\" give you the list of commits which changed something\n> inside \"dir/\"\n\nSure, linear history is trivial. But it's also almost totally \nuninteresting.\n\n\t\t\tLinus\n"},{"id":"62833","messageId":"alpine.LFD.0.9999.0712111548200.25032@woody.linux-foundation.org","threadId":"11241","inReplyTo":"alpine.LFD.0.9999.0712111523210.25032@woody.linux-foundation.org","subject":"Re: git annotate runs out of memory","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-12T00:02:45Z","receivedAt":"2007-12-12T00:02:45Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 11 Dec 2007, Linus Torvalds wrote:\n> \n> and while I suspect xdiff could be optimized a bit more for the cases \n> where we have no changes at the end, that's beyond my skills.\n\nOk, I lied.\n\nNothing is beyond my skills. My mad k0der skillz are unbeatable.\n\nThis speeds up git-blame on ChangeLog-style files by a big amount, by just \nignoring the common end that we don't care about, since we don't want any \ncontext anyway at that point. So I now get:\n\n\t[torvalds@woody gcc]$ time git blame gcc/ChangeLog > /dev/null\n\n\treal    0m7.031s\n\tuser    0m6.852s\n\tsys     0m0.180s\n\nwhich seems quite reasonable, and is about three times faster than trying \nto diff those big files.\n\nDavide: this really _does_ make a huge difference. Maybe xdiff itself \nshould do this optimization on its own, rather than have the caller hack \naround the fact that xdiff doesn't handle this common case all that well?\n\nThe same thing obviously works for the beginning-of-file too, but then you \nhave to play games with line numbers being affected etc, so the end is the \nrather much easier case and is the case that a ChangeLog-style file cares \nabout.\n\nDaniel, this is obviously on top of the patches that fix the memory leak.\n\n\t\t\tLinus\n\n---\ndiff --git a/builtin-blame.c b/builtin-blame.c\nindex c158d31..677188c 100644\n--- a/builtin-blame.c\n+++ b/builtin-blame.c\n@@ -543,6 +551,20 @@ static struct patch *compare_buffer(mmfile_t *file_p, mmfile_t *file_o,\n \treturn state.ret;\n }\n \n+#define BLOCK 1024\n+\n+static void truncate_common_data(mmfile_t *a, mmfile_t *b)\n+{\n+\tlong l1 = a->size, l2 = b->size;\n+\n+\twhile ((l1 -= BLOCK) > 0 && (l2 -= BLOCK) > 0) {\n+\t\tif (memcmp(a->ptr + l1, b->ptr + l2, BLOCK))\n+\t\t\tbreak;\n+\t\ta->size = l1;\n+\t\tb->size = l2;\n+\t}\n+}\n+\n /*\n  * Run diff between two origins and grab the patch output, so that\n  * we can pass blame for lines origin is currently suspected for\n@@ -557,6 +579,7 @@ static struct patch *get_patch(struct origin *parent, struct origin *origin)\n \tfill_origin_blob(origin, &file_o);\n \tif (!file_p.ptr || !file_o.ptr)\n \t\treturn NULL;\n+\ttruncate_common_data(&file_p, &file_o);\n \tpatch = compare_buffer(&file_p, &file_o, 0);\n \tnum_get_patch++;\n \treturn patch;\n"},{"id":"62835","messageId":"Pine.LNX.4.64.0712111611570.1671@alien.or.mcafeemobile.com","threadId":"11241","inReplyTo":"alpine.LFD.0.9999.0712111548200.25032@woody.linux-foundation.org","subject":"Re: git annotate runs out of memory","fromName":"Davide Libenzi","fromEmail":"davidel@xmailserver.org","sentAt":"2007-12-12T00:22:40Z","receivedAt":"2007-12-12T00:22:40Z","isPatch":false,"sender":{"key":"davidel@xmailserver.org","avatar":null},"body":"On Tue, 11 Dec 2007, Linus Torvalds wrote:\n\n> On Tue, 11 Dec 2007, Linus Torvalds wrote:\n> > \n> > and while I suspect xdiff could be optimized a bit more for the cases \n> > where we have no changes at the end, that's beyond my skills.\n> \n> Ok, I lied.\n> \n> Nothing is beyond my skills. My mad k0der skillz are unbeatable.\n> \n> This speeds up git-blame on ChangeLog-style files by a big amount, by just \n> ignoring the common end that we don't care about, since we don't want any \n> context anyway at that point. So I now get:\n> \n> \t[torvalds@woody gcc]$ time git blame gcc/ChangeLog > /dev/null\n> \n> \treal    0m7.031s\n> \tuser    0m6.852s\n> \tsys     0m0.180s\n> \n> which seems quite reasonable, and is about three times faster than trying \n> to diff those big files.\n> \n> Davide: this really _does_ make a huge difference. Maybe xdiff itself \n> should do this optimization on its own, rather than have the caller hack \n> around the fact that xdiff doesn't handle this common case all that well?\n\nI didn't follow the thread, but I can guess from the subject that this is \nabout memory, isn't it?\nLibxdiff already has a xdl_trim_ends() that strips all the common \nbeginning and ending records, but at that point files are already loaded.\nSince libxdiff works with memory files in order to keep any sort of \nsystem dependency out of the window, so the optimization would be \nuseless on libxdiff side. This because the user would have to have \nalready the file loaded in memory, to pass it to libxdiff.\nIf this is really about memory, this better be kept on the libxdiff caller \nside, so that it can avoid loading the terminal file sections altogether.\nAbout your code, you may want to have an extend-till-next-eol code after \nthe trimming part, since the last line may be used for context in the \ndiffs.\n\n\n\n\n- Davide\n"},{"id":"62839","messageId":"alpine.LFD.0.9999.0712111648180.25032@woody.linux-foundation.org","threadId":"11241","inReplyTo":"Pine.LNX.4.64.0712111611570.1671@alien.or.mcafeemobile.com","subject":"Re: git annotate runs out of memory","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-12T00:50:34Z","receivedAt":"2007-12-12T00:50:34Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 11 Dec 2007, Davide Libenzi wrote:\n> \n> I didn't follow the thread, but I can guess from the subject that this is \n> about memory, isn't it?\n\nNo, it started out that way, but now it's about performance.\n\n> Libxdiff already has a xdl_trim_ends() that strips all the common \n> beginning and ending records, but at that point files are already loaded.\n\nThat's not the problem. The problem with xdl_trim_ends() is that it \nhappens *after* you have done all the hashing, so as an optimization it's \nfairly useless, because it still leaves the real cost (the per-line \nhashing) on the table.\n\nSo doing the trimming of the ends before you do even that, allows you to \njust do the trivial \"let's see if the ends are identical\" with a plain \nmemcmp, which is much faster.\n\n\t\t\tLinus\n"},{"id":"62840","messageId":"7vwsrkbuje.fsf@gitster.siamese.dyndns.org","threadId":"11241","inReplyTo":"alpine.LFD.0.9999.0712111548200.25032@woody.linux-foundation.org","subject":"Re: git annotate runs out of memory","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-12-12T00:56:05Z","receivedAt":"2007-12-12T00:56:05Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Tue, 11 Dec 2007, Linus Torvalds wrote:\n>> \n>> and while I suspect xdiff could be optimized a bit more for the cases \n>> where we have no changes at the end, that's beyond my skills.\n>\n> Ok, I lied.\n>\n> Nothing is beyond my skills. My mad k0der skillz are unbeatable.\n>\n> This speeds up git-blame on ChangeLog-style files by a big amount, by just \n> ignoring the common end that we don't care about, since we don't want any \n> context anyway at that point. So I now get:\n>\n> \t[torvalds@woody gcc]$ time git blame gcc/ChangeLog > /dev/null\n>\n> \treal    0m7.031s\n> \tuser    0m6.852s\n> \tsys     0m0.180s\n>\n> which seems quite reasonable, and is about three times faster than trying \n> to diff those big files.\n\nFunny.  I did not understand what you were talking about \"no changes at\nthe end\" when I read it ('cause I am at work and do not have the data\nyou are looking at handy), but now I see what you meant.  It is a cute\nhack that optimizes for a very special case of \"prepend only\" files (aka\n\"ChangeLog\").\n\nI suspect that this optimization has an interesting corner case, though.\nWhat happens if you chomp at the middle of the last line that is\ndifferent between the two files?  xdiff will report the line number but\nwouldn't its (now artificial) \"No newline at the end of the file\" affect\nthe blame logic?\n\nBesides, \"prepend only\" (or \"append only\") files would be good\ncandidates for the original -S\"pickaxe\" search, I would imagine, and\nunless you are looking at that ChangeLog-2000 consolidated log, isn't\nblame way overkill?\n"},{"id":"62841","messageId":"Pine.LNX.4.64.0712111653520.1671@alien.or.mcafeemobile.com","threadId":"11241","inReplyTo":"alpine.LFD.0.9999.0712111648180.25032@woody.linux-foundation.org","subject":"Re: git annotate runs out of memory","fromName":"Davide Libenzi","fromEmail":"davidel@xmailserver.org","sentAt":"2007-12-12T01:12:21Z","receivedAt":"2007-12-12T01:12:21Z","isPatch":false,"sender":{"key":"davidel@xmailserver.org","avatar":null},"body":"On Tue, 11 Dec 2007, Linus Torvalds wrote:\n\n> > Libxdiff already has a xdl_trim_ends() that strips all the common \n> > beginning and ending records, but at that point files are already loaded.\n> \n> That's not the problem. The problem with xdl_trim_ends() is that it \n> happens *after* you have done all the hashing, so as an optimization it's \n> fairly useless, because it still leaves the real cost (the per-line \n> hashing) on the table.\n\nCareful. The real cost of diffing, is not the O(1) pass of the prepare \nphase. It's the potentially O(N*M) worst case of the cross-record compare. \nSo that optimization is far from useless. That optimization is indeed \nmainly targeted to avoid such worst case.\n\n\n\n> So doing the trimming of the ends before you do even that, allows you to \n> just do the trivial \"let's see if the ends are identical\" with a plain \n> memcmp, which is much faster.\n\nYes, tail trimming done on a block-basis is faster and does not consume \nmemory. The code for libxdiff would have to be a bit more complex though, \nsince memory files can be composed by many sections, of different sizes \n(so you cannot just assume it's a single block you're trimming the end). \nAlso, you'd need some code at the end that hands you back at least the N \nlines you want for context.\n\n\n\n- Davide\n"},{"id":"62844","messageId":"alpine.LFD.0.9999.0712111806320.25032@woody.linux-foundation.org","threadId":"11241","inReplyTo":"Pine.LNX.4.64.0712111653520.1671@alien.or.mcafeemobile.com","subject":"Re: git annotate runs out of memory","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-12T02:10:19Z","receivedAt":"2007-12-12T02:10:19Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 11 Dec 2007, Davide Libenzi wrote:\n>\n> > That's not the problem. The problem with xdl_trim_ends() is that it \n> > happens *after* you have done all the hashing, so as an optimization it's \n> > fairly useless, because it still leaves the real cost (the per-line \n> > hashing) on the table.\n> \n> Careful. The real cost of diffing, is not the O(1) pass of the prepare \n> phase. It's the potentially O(N*M) worst case of the cross-record compare. \n> So that optimization is far from useless. That optimization is indeed \n> mainly targeted to avoid such worst case.\n\nI'm not saying it's useless. I'm saying it's ineffective.\n\nMy simple patch that you saw, speeded up a real-life case by A FACTOR OF \nTHREE. We're not talking small potatoes here.\n\n> Also, you'd need some code at the end that hands you back at least the N \n> lines you want for context.\n\nSure. The special case I added it to specifically wanted a context of zero \nin the caller, so I could just ignore that.\n\nBut doing this in general and handing back the context is a simple matter \nof\n\n\twhile (size < orig && context_lines) {\n\t\tif (src->buffer[size++] == '\\n')\n\t\t\tcontext_lines--;\n\t}\n\nwhich will usually hit in a really short time (ie three lines by default, \njust a few tens of bytes).\n\t\t\t\n\t\tLinus\n"},{"id":"62845","messageId":"alpine.LFD.0.9999.0712111816220.25032@woody.linux-foundation.org","threadId":"11241","inReplyTo":"7vwsrkbuje.fsf@gitster.siamese.dyndns.org","subject":"Re: git annotate runs out of memory","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-12T02:20:52Z","receivedAt":"2007-12-12T02:20:52Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 11 Dec 2007, Junio C Hamano wrote:\n> \n> I suspect that this optimization has an interesting corner case, though.\n> What happens if you chomp at the middle of the last line that is\n> different between the two files?  xdiff will report the line number but\n> wouldn't its (now artificial) \"No newline at the end of the file\" affect\n> the blame logic?\n\nIt shouldn't. I thought about it, but there doesn't seem to be any reason \nwhy blame could possibly care - the message can come at the end of a \n_real_ file, of course, so if the extra message confuses the blame logic, \nthere's already a bug there. \n\nBut no, I didn't create a test-case.\n\n> Besides, \"prepend only\" (or \"append only\") files would be good\n> candidates for the original -S\"pickaxe\" search, I would imagine, and\n> unless you are looking at that ChangeLog-2000 consolidated log, isn't\n> blame way overkill?\n\nActually, I suspect that this makes a difference for totally normal files \ntoo. I bet it cuts the size of the files to be tested for the common case \n(ie just a few small changes) down by 30-50% even on average. The fact \nthat it cuts it down by 99.9% on ChangeLog files is just an added bonus.\n\nAs Davide mentioned, xdiff actually does something like that hack for the \nbeginning and end of files internally _anyway_, the problem with that is \nthat it does it so late that it's already done a fairly expensive hash for \nthe file (and allocated space for it based on guesses that are in turn \nbased on the original size) that it doesn't actually get the full effect \nof the optimization.\n\n\t\t\tLinus\n"},{"id":"62846","messageId":"alpine.LFD.0.9999.0712111834510.25032@woody.linux-foundation.org","threadId":"11241","inReplyTo":"alpine.LFD.0.9999.0712111816220.25032@woody.linux-foundation.org","subject":"Re: git annotate runs out of memory","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-12T02:39:38Z","receivedAt":"2007-12-12T02:39:38Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 11 Dec 2007, Linus Torvalds wrote:\n> \n> But no, I didn't create a test-case.\n\nThis *should* trigger the special case:\n\n\tmkdir test-dir\n\tcd test-dir\n\tgit init\n\t(echo -n a ; yes '' | dd count=2) > file\n\tgit add file\n\tgit commit -m \"'a' + 1k newlines\"\n\t(echo -n b ; yes '' | dd count=2) > file\n\tgit add file\n\tgit commit -m \"'b' + 1k newlines\"\n\nand it all seems to work fine.\n\nBut I didn't actually check that it really triggered, this is just \ncreating a 1025-byte file that has a single character and then 1024 \nnewlines. So when the logic removes the shared tail (all the newlines), it \nleaves a single-character newlineless buffer for diff, and no, git-blame \ndidn't care, and got the right answer.\n\n\t\t\tLinus\n"},{"id":"62849","messageId":"alpine.LFD.0.9999.0712111933500.25032@woody.linux-foundation.org","threadId":"11241","inReplyTo":"alpine.LFD.0.9999.0712111806320.25032@woody.linux-foundation.org","subject":"Re: git annotate runs out of memory","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-12T03:35:42Z","receivedAt":"2007-12-12T03:35:42Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 11 Dec 2007, Linus Torvalds wrote:\n> \n> I'm not saying it's useless. I'm saying it's ineffective.\n\nSorry, I _did_ call it \"fairly useless\". \n\nThe rest of the comment stands. I'm sure the trimming that xdiff does is \ngood at avoiding some common O(n*m) cases, it's just not as good as it \ncould be, and leaves a big constant factor of the O(n) case on the table.\n\n\t\t\tLinus\n"},{"id":"62850","messageId":"20071212035737.GL14735@spearce.org","threadId":"11241","inReplyTo":"alpine.LFD.0.9999.0712111122400.25032@woody.linux-foundation.org","subject":"Re: git annotate runs out of memory","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-12-12T03:57:37Z","receivedAt":"2007-12-12T03:57:37Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> wrote:\n...\n> is virtually useless because it's too expensive, but try doing\n> \n> \tgit gui blame gcc ChangeLog\n> \n> instead, and doesn't that just seem nicer? (*)\n> \n> The difference is that the GUI one does it incrementally, and doesn't have \n> to get _all_ the results before it can start reporting blame.\n> \n> Not that I claim that the gui blame is perfect either (I dunno why it \n> delays the nice coloring so long ...\n\ngit-gui waits to color until after it gets the move/copy annotations\nback from the -C -C -w second pass it does.  This way the coloring\nis based on the original source location, not on the move/copy that\ncaused it to be placed where it is now.\n\nI played around with this for a while and finally made it work the\nway it does as I assumed most users would want to see where something\noriginally came from more than how it got moved to where it is now.\n\nIOW the (very expensive) -C -C -w pass is usually much more\ninteresting than the default (fast) pass, so that is the line\nannotation data we color with.  But it takes longer to get and\nis run second, so yea, coloring takes a while.\n\n-- \nShawn.\n"},{"id":"62852","messageId":"7v63z4bjrq.fsf@gitster.siamese.dyndns.org","threadId":"11241","inReplyTo":"7vprxcdhis.fsf@gitster.siamese.dyndns.org","subject":"Re: git annotate runs out of memory","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-12-12T04:48:41Z","receivedAt":"2007-12-12T04:48:41Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> While this should be safe (because the user of blob lazily re-fetches),\n> it feels a bit too aggressive, especially when -C or other \"retry and\n> try harder to assign blame elsewhere\" option is used.\n>\n> Instead, how about discarding after we are done with each origin, like\n> this?\n\nIt's been a while for me to look at the blame engine, and it hit me that\nit would be interesting to run assign_blame() loop on multi-core machine\nin parallel threads.\n"},{"id":"62866","messageId":"20071212075725.GA7676@coredump.intra.peff.net","threadId":"11241","inReplyTo":"alpine.LFD.0.9999.0712111146200.25032@woody.linux-foundation.org","subject":"Re: git annotate runs out of memory","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-12-12T07:57:25Z","receivedAt":"2007-12-12T07:57:25Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Dec 11, 2007 at 11:50:08AM -0800, Linus Torvalds wrote:\n\n> And, btw: the diff is totally different from the xdelta we have, so even \n> if we have an already prepared nice xdelta between the two versions, we'll \n> end up re-generating the files in full, and then do a diff on the end \n> result.\n> \n> Of course, part of that is that git logically *never* works with deltas, \n> except in the actual code-paths that generate objects (or generate packs, \n> of course). So even if we had used a delta algorithm that would be \n> amenable to be turned into a diff directly, it would have been a layering \n> violation to actually do that.\n\nThat doesn't mean we can't opportunistically jump layers when available,\nand fall back on the regular behavior otherwise. The nice thing about\nclean and simple layers is that you can always add optimizations later\nby poking sane holes.\n\nLet's assume for the sake of argument that we can convert an xdelta into\na diff fairly cheaply.  Using the patch below, we can count the places\nwhere we are diffing two blobs, and one blob is a delta base of the\nother (assuming our magical conversion function can also reverse diffs.\n;) ).\n\nFor a \"git log -p\" on git.git, I get:\n\n   9951 diffs could be optimized\n  10958 diffs could not be optimized\n\nor about 48%. It would be nice if we could drop the cost by almost 50%\n(if our magical function is free to call, too!).\n\nOf course, I haven't even looked at whether converting xdeltas to\nunified diffs is possible. I suspect in some cases it is (e.g., pure\naddition of text) and in some cases it isn't (I assume xdelta doesn't\nhave any context lines, which might hurt). And it's possible that a\nspecialized diff user like git-blame can just learn to use the xdeltas\nby itself (I didn't get a \"could optimize\" count for git-blame since\nit seems to follow a different codepath for its diffs).\n\n---\ndiff --git a/cache.h b/cache.h\nindex 27d90fe..0d672be 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -569,6 +569,7 @@ extern void *unpack_entry(struct packed_git *, off_t, enum object_type *, unsign\n extern unsigned long unpack_object_header_gently(const unsigned char *buf, unsigned long len, enum object_type *type, unsigned long *sizep);\n extern unsigned long get_size_from_delta(struct packed_git *, struct pack_window **, off_t);\n extern const char *packed_object_info_detail(struct packed_git *, off_t, unsigned long *, unsigned long *, unsigned int *, unsigned char *);\n+extern int have_xdelta(unsigned char from[20], unsigned char to[20]);\n extern int matches_pack_name(struct packed_git *p, const char *name);\n \n /* Dumb servers support */\ndiff --git a/diff.c b/diff.c\nindex f780e3e..5402900 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -1299,6 +1299,10 @@ static void builtin_diff(const char *name_a,\n \t\t}\n \t}\n \n+\tfprintf(stderr, \"could optimize: %s\\n\",\n+\t\t\t(have_xdelta(one->sha1, two->sha1) ||\n+\t\t\thave_xdelta(two->sha1, one->sha1)) ? \"yes\" : \"no\");\n+\n \tif (fill_mmfile(&mf1, one) < 0 || fill_mmfile(&mf2, two) < 0)\n \t\tdie(\"unable to read files to diff\");\n \ndiff --git a/sha1_file.c b/sha1_file.c\nindex b0c2435..f811ddc 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -2422,3 +2422,20 @@ int read_pack_header(int fd, struct pack_header *header)\n \t\treturn PH_ERROR_PROTOCOL;\n \treturn 0;\n }\n+\n+int have_xdelta(unsigned char from[20], unsigned char to[20])\n+{\n+\tstruct pack_entry e;\n+\tunsigned char base_sha1[20];\n+\tconst char *type;\n+\tunsigned long size;\n+\tunsigned long store_size;\n+\tunsigned int delta_chain_length;\n+\n+\tif (!find_pack_entry(to, &e, NULL))\n+\t\treturn 0;\n+\n+\ttype = packed_object_info_detail(e.p, e.offset, &size, &store_size,\n+\t\t\t\t\t &delta_chain_length, base_sha1);\n+\treturn !hashcmp(base_sha1, from);\n+}\n"},{"id":"62895","messageId":"82bq8wi4j2.fsf@mid.bfk.de","threadId":"11241","inReplyTo":"4aca3dc20712110933i636342fbifb15171d3e3cafb3@mail.gmail.com","subject":"Re: git annotate runs out of memory","fromName":"Florian Weimer","fromEmail":"fweimer@bfk.de","sentAt":"2007-12-12T10:36:01Z","receivedAt":"2007-12-12T10:36:01Z","isPatch":false,"sender":{"key":"fweimer@bfk.de","avatar":null},"body":"* Daniel Berlin:\n\n> On the gcc repository (which is now a 234 meg pack for me), git\n> annotate ChangeLog takes > 800 meg of memory (I stopped it at about\n> 1.6 gig, since it started swapping my machine).\n> I assume it will run out of memory.  I stopped it after 2 minutes.\n\nA less unwieldy repository that shows the same problem is:\n\n  svn://svn.debian.org/secure-testing/\n\nIt's annotating the data/CVE/list file that uses tons of memory.  I\nguess you don't need to clone the full history to exhibit the problem.\n\n-- \nFlorian Weimer                <fweimer@bfk.de>\nBFK edv-consulting GmbH       http://www.bfk.de/\nKriegsstraße 100              tel: +49-721-96201-1\nD-76133 Karlsruhe             fax: +49-721-96201-99\n"},{"id":"62955","messageId":"4aca3dc20712121143q4d7ccd9n5408ad7199981164@mail.gmail.com","threadId":"11241","inReplyTo":"alpine.LFD.0.9999.0712111548200.25032@woody.linux-foundation.org","subject":"Re: git annotate runs out of memory","fromName":"Daniel Berlin","fromEmail":"dberlin@dberlin.org","sentAt":"2007-12-12T19:43:20Z","receivedAt":"2007-12-12T19:43:20Z","isPatch":false,"sender":{"key":"dberlin@dberlin.org","avatar":null},"body":"On 12/11/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>\n>\n> Daniel, this is obviously on top of the patches that fix the memory leak.\n\nThanks, these patches work *great*.\n\nI'm starting to have a few users who have no experience with git or hg\ntry their daily workflow with it, to see what UI issues they come up\nwith :)\n"},{"id":"63520","messageId":"20071217232450.GA13012@efreet.light.src","threadId":"11241","inReplyTo":"20071212075725.GA7676@coredump.intra.peff.net","subject":"Re: git annotate runs out of memory","fromName":"Jan Hudec","fromEmail":"bulb@ucw.cz","sentAt":"2007-12-17T23:24:50Z","receivedAt":"2007-12-17T23:24:50Z","isPatch":false,"sender":{"key":"bulb@ucw.cz","avatar":null},"body":"On Wed, Dec 12, 2007 at 02:57:25 -0500, Jeff King wrote:\n> On Tue, Dec 11, 2007 at 11:50:08AM -0800, Linus Torvalds wrote:\n> > And, btw: the diff is totally different from the xdelta we have, so even \n> > if we have an already prepared nice xdelta between the two versions, we'll \n> > end up re-generating the files in full, and then do a diff on the end \n> > result.\n\nThe problem is whether git does not end-up re-generating the same file\nmultiple times. When it needs to construct the diff between two versions of\na file and one is delta-base (even indirect) of the other, does it know to\ncreate the first, remember it, continue to the other and calculate the diff?\n\n> > Of course, part of that is that git logically *never* works with deltas, \n> > except in the actual code-paths that generate objects (or generate packs, \n> > of course). So even if we had used a delta algorithm that would be \n> > amenable to be turned into a diff directly, it would have been a layering \n> > violation to actually do that.\n> \n> That doesn't mean we can't opportunistically jump layers when available,\n> and fall back on the regular behavior otherwise. The nice thing about\n> clean and simple layers is that you can always add optimizations later\n> by poking sane holes.\n> \n> Let's assume for the sake of argument that we can convert an xdelta into\n> a diff fairly cheaply.  Using the patch below, we can count the places\n> where we are diffing two blobs, and one blob is a delta base of the\n> other (assuming our magical conversion function can also reverse diffs.\n> ;) ).\n> \n> For a \"git log -p\" on git.git, I get:\n> \n>    9951 diffs could be optimized\n>   10958 diffs could not be optimized\n> \n> or about 48%. It would be nice if we could drop the cost by almost 50%\n> (if our magical function is free to call, too!).\n\nThis is actually a gross underestimation. The idea would be to know all the\ndiffs we need to calculate and than remember all useful results. Ie. if we\nknow we'll want objects A and C, A's delta base is B and B's delta base is C,\nstart calculating A and when it turns out to need C at some point, just\nremember it for purpose of doing the final diff. On the other hand B can be\nthrown away early (because we don't need it) to save memory.\n\nNow git can know the list of deltas it will need in advance. First generate\nthe list of revisions -- nothing helps there, but their delta bases are\nlikely to be randomish anyway -- and than with the knowledge of full list of\ntrees, start doing the diffs to see which touched the subtree in question.\nRepeat for each level.\n\nSince the list of deltas that will be needed is known, the objects from\nwhich all deltas were already generated can be expired from cache (but not\nthrown away immediately, as they may help building other objects).\n\n> Of course, I haven't even looked at whether converting xdeltas to\n> unified diffs is possible. I suspect in some cases it is (e.g., pure\n> addition of text) and in some cases it isn't (I assume xdelta doesn't\n> have any context lines, which might hurt). And it's possible that a\n> specialized diff user like git-blame can just learn to use the xdeltas\n> by itself (I didn't get a \"could optimize\" count for git-blame since\n> it seems to follow a different codepath for its diffs).\n\nWell, it's about as hard as applying them, because you can remember the\nnecessary stuff when applying. The imporant bit would be to avoid applying\nthe same delta more than once during the whole annotate.\n\n-- \n\t\t\t\t\t\t Jan 'Bulb' Hudec <bulb@ucw.cz>\n"},{"id":"63527","messageId":"alpine.LFD.0.9999.0712171558350.21557@woody.linux-foundation.org","threadId":"11241","inReplyTo":"20071217232450.GA13012@efreet.light.src","subject":"Re: git annotate runs out of memory","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-18T00:05:49Z","receivedAt":"2007-12-18T00:05:49Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 18 Dec 2007, Jan Hudec wrote:\n> On Tue, Dec 11, 2007 at 11:50:08AM -0800, Linus Torvalds wrote:\n> > And, btw: the diff is totally different from the xdelta we have, so even \n> > if we have an already prepared nice xdelta between the two versions, we'll \n> > end up re-generating the files in full, and then do a diff on the end \n> > result.\n> \n> The problem is whether git does not end-up re-generating the same file\n> multiple times. When it needs to construct the diff between two versions of\n> a file and one is delta-base (even indirect) of the other, does it know to\n> create the first, remember it, continue to the other and calculate the diff?\n\nYes.\n\nActually, it doesn't \"know\" anything at all - what happens is that git \ninternally has a simple \"delta-cache\", which just caches the latest \nobjects we've generated from deltas, and which automatically handles this \ncommon case (and others).\n\nSo when we tend to work with multiple versions of the same file (which is \nobviously very common with diff, and even more so with something like \n\"annotate\"), those multiple versions will obviously also tend to be deltas \nagainst each other and/or against some shared base object, and when we see \na delta, we'll look the base object up in the delta cache, and if it has \nbeen generated earlier we'll be able to short-circuit the whole delta \nchain and just use the whole object we already cached.\n\nSo if you compare two objects that each have a very deep delta chain, you \nwill obviously have to walk the whole delta chain _once_ (to generate \nwhichever version of the file you happen to look up first), but you won't \nneed to do it twice, because the second time you'll end up hitting in the \ndelta cache.\n\n\t\t\tLinus\n"}]}