git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: git annotate runs out of memory

From
DLDavide Libenzi <davidel@xmailserver.org>
Date
Dec 12, 2007, 01:12 UTC
Message-ID
<Pine.LNX.4.64.0712111653520.1671@alien.or.mcafeemobile.com>
In-Reply-To
<alpine.LFD.0.9999.0712111648180.25032@woody.linux-foundation.org>
On Tue, 11 Dec 2007, Linus Torvalds wrote:
Show 7 quoted lines
> > Libxdiff already has a xdl_trim_ends() that strips all the common 
> > beginning and ending records, but at that point files are already loaded.
> 
> That's not the problem. The problem with xdl_trim_ends() is that it 
> happens *after* you have done all the hashing, so as an optimization it's 
> fairly useless, because it still leaves the real cost (the per-line 
> hashing) on the table.

Careful. The real cost of diffing, is not the O(1) pass of the prepare phase. It's the potentially O(N*M) worst case of the cross-record compare. So that optimization is far from useless. That optimization is indeed mainly targeted to avoid such worst case.

> So doing the trimming of the ends before you do even that, allows you to 
> just do the trivial "let's see if the ends are identical" with a plain 
> memcmp, which is much faster.

Yes, tail trimming done on a block-basis is faster and does not consume memory. The code for libxdiff would have to be a bit more complex though, since memory files can be composed by many sections, of different sizes (so you cannot just assume it's a single block you're trimming the end). Also, you'd need some code at the end that hands you back at least the N lines you want for context.

- Davide
Previous: Linus TorvaldsNext: Linus Torvalds
Message 38 of 51 in “git annotate runs out of memory”
  1. Daniel BerlinDec 11, 2007
  2. Nicolas PitreDec 11, 2007
  3. Daniel BerlinDec 11, 2007
  4. Nicolas PitreDec 11, 2007
  5. Marco CostalbaDec 11, 2007
  6. Daniel BerlinDec 11, 2007
  7. Marco CostalbaDec 11, 2007
  8. Jason SewallDec 11, 2007
  9. Daniel BarkalowDec 11, 2007
  10. Marco CostalbaDec 11, 2007
  11. Linus TorvaldsDec 11, 2007
  12. Matthieu MoyDec 11, 2007
  13. Linus TorvaldsDec 11, 2007
  14. Daniel BerlinDec 11, 2007
  15. Pierre HabouzitDec 11, 2007
  16. Daniel BerlinDec 11, 2007
  17. Matthieu MoyDec 11, 2007
  18. Linus TorvaldsDec 11, 2007
  19. Nicolas PitreDec 11, 2007
  20. Jon SmirlDec 11, 2007
  21. Daniel BerlinDec 11, 2007
  22. Daniel BarkalowDec 11, 2007
  23. Pierre HabouzitDec 11, 2007
  24. Junio C HamanoDec 11, 2007
  25. Linus TorvaldsDec 11, 2007
  26. Linus TorvaldsDec 11, 2007
  27. Daniel BerlinDec 11, 2007
  28. Linus TorvaldsDec 11, 2007
  29. Jeff KingDec 12, 2007
  30. Jan HudecDec 17, 2007
  31. Linus TorvaldsDec 18, 2007
  32. Linus TorvaldsDec 11, 2007
  33. Junio C HamanoDec 11, 2007
  34. Linus TorvaldsDec 11, 2007
  35. Linus TorvaldsDec 12, 2007
  36. Davide LibenziDec 12, 2007
  37. Linus TorvaldsDec 12, 2007
  38. Davide LibenziDec 12, 2007
  39. Linus TorvaldsDec 12, 2007
  40. Linus TorvaldsDec 12, 2007
  41. Junio C HamanoDec 12, 2007
  42. Linus TorvaldsDec 12, 2007
  43. Linus TorvaldsDec 12, 2007
  44. Daniel BerlinDec 12, 2007
  45. Junio C HamanoDec 12, 2007
  46. Daniel BerlinDec 11, 2007
  47. Shawn O. PearceDec 12, 2007
  48. Marco CostalbaDec 11, 2007
  49. Steven GrimmDec 11, 2007
  50. Jakub NarebskiDec 11, 2007
  51. Florian WeimerDec 12, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.