git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: git annotate runs out of memory

From
Pierre Habouzit <madcoder@debian.org>
Date
Dec 11, 2007, 19:34 UTC
Message-ID
<20071211193407.GC20644@artemis.madism.org>
In-Reply-To
<4aca3dc20712111109y5d74a292rf29be6308932393c@mail.gmail.com>
On Tue, Dec 11, 2007 at 07:09:03PM +0000, Daniel Berlin wrote:
Show 22 quoted lines
> On 12/11/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:
> >
> >
> > On Tue, 11 Dec 2007, Daniel Berlin wrote:
> > >
> > > This seems to be a common problem with git. It seems to use a lot of
> > > memory to perform common operations on the gcc repository (even though
> > > it is faster in some cases than hg).
> >
> > The thing is, git has a very different notion of "common operations" than
> > you do.
> >
> > To git, "git annotate" is just about the *last* thing you ever want to do.
> > It's not a common operation, it's a "last resort" operation. In git, the
> > whole workflow is designed for "git log -p <pathnamepattern>" rather than
> > annotate/blame.
> >
> I understand this, and completely agree with you.
> However, I cannot force GCC people to adopt completely new workflow in
> this regard.
> The changelog's are not useful enough (and we've had huge fights over
> this) to do git log -p and figure out the info we want.
> Looking through thousands of diffs to find the one that happened to
> your line is also pretty annoying.
  If the question you want to answer is "what happened to that line"
then using git annotate is using a big hammer for no good reason.
git log -S'<put the content of the line here>' -- path/to/file.c

will give you the very same answer, pointing you to the changes that added or removed that line directly. It's not a fast command either, but it should be less resource hungry than annotate that has to do roughly the same for all lines whereas you're interested in one only.

The direct plus here, is that git log output is incremental, so you have answers about the first diffs quite quick, which let you examine the first answers while the rest is still being computed.

Unlike git annotate, this also allow you to restrict the revisions where it searches to a range where you know this happened, which makes it almost instantaneous in most cases.

Of course, if the line is ' free(p);\n' then you will probably have quite a few false positives, but with the path restriction, I assume this will still be quite accurate.

What is important here is to know what is the real question the GCC programmers want to answer to. It seems to me that `blame` is an overkill for the underlying issue.

Note that it does not justifies the current memory consumption that just looks bad and wrong to me, but this aims at finding a way to answer your question doing just what you need to answer it and not gazillions of other things :)

-- 
·O·  Pierre Habouzit
··O                                                madcoder@debian.org
OOO                                                http://www.madism.org
Previous: Daniel BarkalowNext: Junio C Hamano
Message 23 of 51 in “git annotate runs out of memory”
  1. Daniel BerlinDec 11, 2007
  2. Nicolas PitreDec 11, 2007
  3. Daniel BerlinDec 11, 2007
  4. Nicolas PitreDec 11, 2007
  5. Marco CostalbaDec 11, 2007
  6. Daniel BerlinDec 11, 2007
  7. Marco CostalbaDec 11, 2007
  8. Jason SewallDec 11, 2007
  9. Daniel BarkalowDec 11, 2007
  10. Marco CostalbaDec 11, 2007
  11. Linus TorvaldsDec 11, 2007
  12. Matthieu MoyDec 11, 2007
  13. Linus TorvaldsDec 11, 2007
  14. Daniel BerlinDec 11, 2007
  15. Pierre HabouzitDec 11, 2007
  16. Daniel BerlinDec 11, 2007
  17. Matthieu MoyDec 11, 2007
  18. Linus TorvaldsDec 11, 2007
  19. Nicolas PitreDec 11, 2007
  20. Jon SmirlDec 11, 2007
  21. Daniel BerlinDec 11, 2007
  22. Daniel BarkalowDec 11, 2007
  23. Pierre HabouzitDec 11, 2007
  24. Junio C HamanoDec 11, 2007
  25. Linus TorvaldsDec 11, 2007
  26. Linus TorvaldsDec 11, 2007
  27. Daniel BerlinDec 11, 2007
  28. Linus TorvaldsDec 11, 2007
  29. Jeff KingDec 12, 2007
  30. Jan HudecDec 17, 2007
  31. Linus TorvaldsDec 18, 2007
  32. Linus TorvaldsDec 11, 2007
  33. Junio C HamanoDec 11, 2007
  34. Linus TorvaldsDec 11, 2007
  35. Linus TorvaldsDec 12, 2007
  36. Davide LibenziDec 12, 2007
  37. Linus TorvaldsDec 12, 2007
  38. Davide LibenziDec 12, 2007
  39. Linus TorvaldsDec 12, 2007
  40. Linus TorvaldsDec 12, 2007
  41. Junio C HamanoDec 12, 2007
  42. Linus TorvaldsDec 12, 2007
  43. Linus TorvaldsDec 12, 2007
  44. Daniel BerlinDec 12, 2007
  45. Junio C HamanoDec 12, 2007
  46. Daniel BerlinDec 11, 2007
  47. Shawn O. PearceDec 12, 2007
  48. Marco CostalbaDec 11, 2007
  49. Steven GrimmDec 11, 2007
  50. Jakub NarebskiDec 11, 2007
  51. Florian WeimerDec 12, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.