git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: git annotate runs out of memory

From
Steven Grimm <koreth@midwinter.com>
Date
Dec 11, 2007, 19:29 UTC
Message-ID
<74ED838F-4966-42A9-BC8A-906FD0B4B46F@midwinter.com>
In-Reply-To
<alpine.LFD.0.9999.0712111018540.25032@woody.linux-foundation.org>
On Dec 11, 2007, at 10:40 AM, Linus Torvalds wrote:
Show 7 quoted lines
> To git, "git annotate" is just about the *last* thing you ever want  
> to do.
> It's not a common operation, it's a "last resort" operation. In git,  
> the
> whole workflow is designed for "git log -p <pathnamepattern>" rather  
> than
> annotate/blame.

My use of "git blame" is perhaps not typical, but I use it fairly often when I'm looking at a part of my company's code base that I'm not terribly familiar with. I've found it's the fastest way to figure out who to go ask about a particular block of code that I think is responsible for a bug, or more commonly, who to ask to review a change I'm making.

"git log" is too coarse-grained to be useful for that purpose; it usually doesn't tell me which of the 500 revisions to the file I'm looking at introduced the actual line of code I want to change.

To me that really has nothing whatsoever to do with git workflow or svn workflow; it happens well before I'm ready to do any kind of integration or commit or even, sometimes, before I've made any changes to any code at all.

Given infinite spare time, one of the things I'd be strongly tempted to try to build would be some kind of blame cache. You could theoretically make blame pretty much instantaneous by doing something as simple as caching the per-line revision ID for each file in each revision in a shadow repository (or a shadow branch in the main repo) and keeping a map between shadow-repo revisions and real-repo ones. If the cache was of the form "one SHA1 hash per line in the original file" it would delta-compress pretty well. It'd be easy to update incrementally since you only need to walk back in history until you get to the most recently cached revision for each file, at which point you use the cached value for all the lines that haven't changed.

Yeah, I know, code talks louder than words...
-Steve
Previous: Marco CostalbaNext: Jakub Narebski
Message 49 of 51 in “git annotate runs out of memory”
  1. Daniel BerlinDec 11, 2007
  2. Nicolas PitreDec 11, 2007
  3. Daniel BerlinDec 11, 2007
  4. Nicolas PitreDec 11, 2007
  5. Marco CostalbaDec 11, 2007
  6. Daniel BerlinDec 11, 2007
  7. Marco CostalbaDec 11, 2007
  8. Jason SewallDec 11, 2007
  9. Daniel BarkalowDec 11, 2007
  10. Marco CostalbaDec 11, 2007
  11. Linus TorvaldsDec 11, 2007
  12. Matthieu MoyDec 11, 2007
  13. Linus TorvaldsDec 11, 2007
  14. Daniel BerlinDec 11, 2007
  15. Pierre HabouzitDec 11, 2007
  16. Daniel BerlinDec 11, 2007
  17. Matthieu MoyDec 11, 2007
  18. Linus TorvaldsDec 11, 2007
  19. Nicolas PitreDec 11, 2007
  20. Jon SmirlDec 11, 2007
  21. Daniel BerlinDec 11, 2007
  22. Daniel BarkalowDec 11, 2007
  23. Pierre HabouzitDec 11, 2007
  24. Junio C HamanoDec 11, 2007
  25. Linus TorvaldsDec 11, 2007
  26. Linus TorvaldsDec 11, 2007
  27. Daniel BerlinDec 11, 2007
  28. Linus TorvaldsDec 11, 2007
  29. Jeff KingDec 12, 2007
  30. Jan HudecDec 17, 2007
  31. Linus TorvaldsDec 18, 2007
  32. Linus TorvaldsDec 11, 2007
  33. Junio C HamanoDec 11, 2007
  34. Linus TorvaldsDec 11, 2007
  35. Linus TorvaldsDec 12, 2007
  36. Davide LibenziDec 12, 2007
  37. Linus TorvaldsDec 12, 2007
  38. Davide LibenziDec 12, 2007
  39. Linus TorvaldsDec 12, 2007
  40. Linus TorvaldsDec 12, 2007
  41. Junio C HamanoDec 12, 2007
  42. Linus TorvaldsDec 12, 2007
  43. Linus TorvaldsDec 12, 2007
  44. Daniel BerlinDec 12, 2007
  45. Junio C HamanoDec 12, 2007
  46. Daniel BerlinDec 11, 2007
  47. Shawn O. PearceDec 12, 2007
  48. Marco CostalbaDec 11, 2007
  49. Steven GrimmDec 11, 2007
  50. Jakub NarebskiDec 11, 2007
  51. Florian WeimerDec 12, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.