git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: git annotate runs out of memory

From
Linus Torvalds <torvalds@linux-foundation.org>
Date
Dec 11, 2007, 19:42 UTC
Message-ID
<alpine.LFD.0.9999.0712111122400.25032@woody.linux-foundation.org>
In-Reply-To
<4aca3dc20712111109y5d74a292rf29be6308932393c@mail.gmail.com>
On Tue, 11 Dec 2007, Daniel Berlin wrote:
>
> I understand this, and completely agree with you.
> However, I cannot force GCC people to adopt completely new workflow in
> this regard.

Oh, I agree. It's why we do have "git blame" these days, and it's why I've tried to make people use the nicer incremental mode, which is not at all faster, but it's a hell of a lot more pleasant to use because you get some output immediately.

In other words,
	git blame gcc/ChangeLog
is virtually useless because it's too expensive, but try doing
	git gui blame gcc ChangeLog
instead, and doesn't that just seem nicer? (*)

The difference is that the GUI one does it incrementally, and doesn't have to get _all_ the results before it can start reporting blame.

Not that I claim that the gui blame is perfect either (I dunno why it delays the nice coloring so long, for example), but it was something I pushed - and others made the gui for - exactly to help people with the fact that git interally really does it that incremental way.

> SVN had the same problem (the file retrieval was the most expensive op
> on FSFS). One of the things i did to speed it up tremendously was to
> do the annotate from newest to oldest (IE in reverse), and stop
> annotating when we had come up with annotate info for all the lines.

We do that. The expense for git is that we don't do the revisions as a single file at all. We'll look through each commit, check whether the "gcc" directory changed, if it did, we'll go into it, and check whether the "ChangeLog" file changed - and if it did, we'll actually diff it against the previous version.

> In GCC history, it is likely you will be able to cut off at least 30%
> of the time if you do this, because files often have changed entirely
> multiple times.

Not gcc/ChangeLog, though (apart from the renames that happen occasionally).

Btw, an example of something git *should* do right, but is just too damn expensive, is doing

	git gui blame gcc/ChangeLog-2000

and have it actually be able to track the original source of each of those annotations across that "ChangeLog split from hell".

I bet it would eventually get it right, but that's a large file, way back in history, and it will try to do a non-whitespace blame with copy detection.

That's *expensive*, although it is an amusing thing to try to do ;)
			Linus

PS. I also do agree that we seem to use an excessive amount of memory there. As to whether it's the same issue or not, I'd not go as far as Nico and say "yes" yet. But it's interesting.

It's not entirely surprising that we see multiple issues with the gcc repo, simply because it's not the kind of repo that people have ever really worked on. So I don't think it's necessarily related at all, except in the sense of it being a different load and showing issues.

Previous: Junio C HamanoNext: Linus Torvalds
Message 25 of 51 in “git annotate runs out of memory”
  1. Daniel BerlinDec 11, 2007
  2. Nicolas PitreDec 11, 2007
  3. Daniel BerlinDec 11, 2007
  4. Nicolas PitreDec 11, 2007
  5. Marco CostalbaDec 11, 2007
  6. Daniel BerlinDec 11, 2007
  7. Marco CostalbaDec 11, 2007
  8. Jason SewallDec 11, 2007
  9. Daniel BarkalowDec 11, 2007
  10. Marco CostalbaDec 11, 2007
  11. Linus TorvaldsDec 11, 2007
  12. Matthieu MoyDec 11, 2007
  13. Linus TorvaldsDec 11, 2007
  14. Daniel BerlinDec 11, 2007
  15. Pierre HabouzitDec 11, 2007
  16. Daniel BerlinDec 11, 2007
  17. Matthieu MoyDec 11, 2007
  18. Linus TorvaldsDec 11, 2007
  19. Nicolas PitreDec 11, 2007
  20. Jon SmirlDec 11, 2007
  21. Daniel BerlinDec 11, 2007
  22. Daniel BarkalowDec 11, 2007
  23. Pierre HabouzitDec 11, 2007
  24. Junio C HamanoDec 11, 2007
  25. Linus TorvaldsDec 11, 2007
  26. Linus TorvaldsDec 11, 2007
  27. Daniel BerlinDec 11, 2007
  28. Linus TorvaldsDec 11, 2007
  29. Jeff KingDec 12, 2007
  30. Jan HudecDec 17, 2007
  31. Linus TorvaldsDec 18, 2007
  32. Linus TorvaldsDec 11, 2007
  33. Junio C HamanoDec 11, 2007
  34. Linus TorvaldsDec 11, 2007
  35. Linus TorvaldsDec 12, 2007
  36. Davide LibenziDec 12, 2007
  37. Linus TorvaldsDec 12, 2007
  38. Davide LibenziDec 12, 2007
  39. Linus TorvaldsDec 12, 2007
  40. Linus TorvaldsDec 12, 2007
  41. Junio C HamanoDec 12, 2007
  42. Linus TorvaldsDec 12, 2007
  43. Linus TorvaldsDec 12, 2007
  44. Daniel BerlinDec 12, 2007
  45. Junio C HamanoDec 12, 2007
  46. Daniel BerlinDec 11, 2007
  47. Shawn O. PearceDec 12, 2007
  48. Marco CostalbaDec 11, 2007
  49. Steven GrimmDec 11, 2007
  50. Jakub NarebskiDec 11, 2007
  51. Florian WeimerDec 12, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.