git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH 2/2] Implement a simple delta_base cache

From
Nicolas Pitre <nico@cam.org>
Date
Mar 18, 2007, 03:00 UTC
Message-ID
<alpine.LFD.0.83.0703172228220.18328@xanadu.home>
In-Reply-To
<Pine.LNX.4.64.0703171911120.6730@woody.linux-foundation.org>
On Sat, 17 Mar 2007, Linus Torvalds wrote:
Show 9 quoted lines
> 
> 
> On Sat, 17 Mar 2007, Nicolas Pitre wrote:
> > 
> > Well... in my opinion it is the _current_ tree walker that is quite ugly 
> > and complex.  It is always messier to parse strings than fixed width 
> > binary fields.
> 
> Sure. On the other hand, text is what made things easy to do initially,
Oh indeed.  No argument there.
Show 5 quoted lines
> and you're missing one *BIG* clue: you cannot remote the support without 
> losing compatibility with all traditional object formats.
> 
> So you have no choice. You need to support the text representation. As a 
> result, *your* code will now be way more ugly and messy.

Depends. We currently have separate parsers for trees, commits, tags, etc. That should be easy enough to add another (separate) parser for new tree objects while still having a common higher level accessor interface like tree_entry().

But right now we only regenerate the text representation whenever the binary representation is encountered just to make things easy to do, and yet we still have a performance gain already in _addition_ to a net saving in disk footprint.

> The thing is, parsing some little text may sound expensive, but if the 
> expense is in finding the end of the string, we're doing really well.

Of course the current tree parser will remain, probably forever. And it is always a good thing to optimize it further when ever possible.

> But what you're ignoring here is that "16%" may sound like a huge deal, 
> but it's 16% of somethng that takes 1 second, and that other SCM's cannot 
> do AT ALL.

Sure. But at this point the reference to compare GIT performance against might be GIT itself. And while 1 second is really nice in this case, there are some repos where it could be (and has already been reported to be) much more.

I still have a feeling that we can do even better than we do now. Much much better than 16% actually. But that require a new data format that is designed for speed.

We'll see.
Nicolas
Previous: Linus TorvaldsNext: Linus Torvalds
Message 37 of 79 in “cleaner/better zlib sources?”
  1. Linus TorvaldsMar 16, 2007
  2. Shawn O. PearceMar 16, 2007
  3. Jeff GarzikMar 16, 2007
  4. Matt MackallMar 16, 2007
  5. Linus TorvaldsMar 16, 2007
  6. Linus TorvaldsMar 16, 2007
  7. Davide LibenziMar 16, 2007
  8. Linus TorvaldsMar 16, 2007
  9. Davide LibenziMar 16, 2007
  10. Linus TorvaldsMar 16, 2007
  11. Davide LibenziMar 16, 2007
  12. Linus TorvaldsMar 16, 2007
  13. Davide LibenziMar 16, 2007
  14. Linus TorvaldsMar 17, 2007
  15. Linus TorvaldsMar 17, 2007
  16. Nicolas PitreMar 17, 2007
  17. Shawn O. PearceMar 17, 2007
  18. Linus TorvaldsMar 17, 2007
  19. Linus TorvaldsMar 17, 2007
  20. 1/2 Make trivial wrapper functions around delta base generation and freeingLinus Torvalds, Mar 17, 2007
  21. 2/2 Implement a simple delta_base cacheLinus Torvalds, Mar 17, 2007
  22. Linus TorvaldsMar 17, 2007
  23. Junio C HamanoMar 17, 2007
  24. Linus TorvaldsMar 17, 2007
  25. Linus TorvaldsMar 17, 2007
  26. Nicolas PitreMar 18, 2007
  27. Junio C HamanoMar 18, 2007
  28. Junio C HamanoMar 17, 2007
  29. Linus TorvaldsMar 17, 2007
  30. Jon SmirlMar 17, 2007
  31. Morten WelinderMar 18, 2007
  32. Linus TorvaldsMar 18, 2007
  33. Nicolas PitreMar 18, 2007
  34. Linus TorvaldsMar 18, 2007
  35. Nicolas PitreMar 18, 2007
  36. Linus TorvaldsMar 18, 2007
  37. Nicolas PitreMar 18, 2007
  38. Linus TorvaldsMar 18, 2007
  39. Julian PhillipsMar 18, 2007
  40. Linus TorvaldsMar 18, 2007
  41. Robin RosenbergMar 18, 2007
  42. Linus TorvaldsMar 18, 2007
  43. Robin RosenbergMar 18, 2007
  44. Shawn O. PearceMar 18, 2007
  45. David BrodskyMar 19, 2007
  46. Robin RosenbergMar 20, 2007
  47. David BrodskyMar 20, 2007
  48. Linus TorvaldsMar 21, 2007
  49. Nicolas PitreMar 21, 2007
  50. 3/2 Avoid unnecessary strlen() callsLinus Torvalds, Mar 18, 2007
  51. Junio C HamanoMar 18, 2007
  52. Linus TorvaldsMar 18, 2007
  53. Linus TorvaldsMar 18, 2007
  54. Shawn O. PearceMar 18, 2007
  55. Linus TorvaldsMar 18, 2007
  56. Johannes SchindelinMar 20, 2007
  57. Shawn O. PearceMar 20, 2007
  58. Shawn O. PearceMar 20, 2007
  59. Linus TorvaldsMar 20, 2007
  60. Shawn O. PearceMar 20, 2007
  61. Linus TorvaldsMar 20, 2007
  62. Junio C HamanoMar 20, 2007
  63. Junio C HamanoMar 20, 2007
  64. Linus TorvaldsMar 20, 2007
  65. Shawn O. PearceMar 20, 2007
  66. Linus TorvaldsMar 20, 2007
  67. Linus TorvaldsMar 18, 2007
  68. Avi KivityMar 18, 2007
  69. Linus TorvaldsMar 17, 2007
  70. Jeff GarzikMar 16, 2007
  71. Matt MackallMar 16, 2007
  72. Linus TorvaldsMar 16, 2007
  73. Nicolas PitreMar 16, 2007
  74. Shawn O. PearceMar 16, 2007
  75. Nicolas PitreMar 16, 2007
  76. Linus TorvaldsMar 16, 2007
  77. Nicolas PitreMar 16, 2007
  78. Davide LibenziMar 16, 2007
  79. Davide LibenziMar 16, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.