git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: libgit2 - a true git library

From
Andreas Ericsson <ae@op5.se>
Date
Nov 3, 2008, 10:17 UTC
Message-ID
<490ECFC2.1070901@op5.se>
In-Reply-To
<20081102015041.GG15463@spearce.org>
Shawn O. Pearce wrote:
Show 25 quoted lines
> Andreas Ericsson <ae@op5.se> wrote:
>> Shawn O. Pearce wrote:
>>> Eh, I disagree here.  In git.git today "struct commit" exposes its
>>> buffer with the canonical commit encoding.  Having that visible
>>> wrecks what Nico and I were thinking about doing with pack v4 and
>>> encoding commits in a non-canonical format when stored in packs.
>>> Ditto with trees.
>> Err... isn't that backwards?
> 
> No.
> 
>> Surely you want to store stuff in the
>> canonical format so you're forced to do as few translations as
>> possible?
> 
> No.  We suspect that canonical format is harder to decompress and
> parse during revision traversal.  Other encodings in the pack file
> may produce much faster runtime performance, and reduce page faults
> (due to smaller pack sizes).
> 
> We hardly ever use the canonical format for actual output; most
> output rips the canonical format apart and then formats the data
> the way it was requested.  If we have the data *already* parsed in
> the pack its much faster to output.
> 

I'll have to look into the pack v4 stuff, as I can't get what you're saying to make sense to me. Not canonicalizing the data when storing it means you'll have to have conversion routines from all the various encodings, unless you first canonicalize it and then encode it in the way you need it in the pack way. Never using canonical format means it has no potential for combinatorial explosion, and every converter needs to know about every other format.

Show 19 quoted lines
>> Or are you trying to speed up packing by skipping the
>> canonicalization part?
> 
> Wrong; we're trying to speed up reading.  Packing may go slower,
> especially during the first conversion of v2->v4 for any given
> repository, but packing is infrequent so the minor (if any) drop
> in performance here is probably worth the reading performance gains.
> 
>> Well, if macro usage is adhered to one wouldn't have to worry,
>> since the macro can just be rewritten with a function later (if,
>> for example, translation or some such happens to be required).
>> Older code linking to a newer library would work (assuming the
>> size of the commit object doesn't change anyway),
> 
> You are assuming too much magic.  If the older ABI used a macro
> and the newer one (which supports pack v4) organized struct commit
> differently and the user upgrades libgit2.so the older applications
> just broke, horribly.
> 

Naturally. Re-sizing objects that weren't previously protected by accessor functions will always break the ABI unless the layout of the previously existing items in the object doesn't change, which is exactly what I said.

Show 5 quoted lines
> We know we want to do pack v4 in the near future.  Or at least
> experiment with it and see if it works.  If it does, we don't
> want to have to cause a major ABI breakage across all those newly
> installed libgit2s... yikes.
> 

That's not necessary, although it requires a bit more thought when changing the objects.

Show 7 quoted lines
> I'm really in favor of accessor functions for the first version of
> the library.  They can always be converted to macros once someone
> shows that their git visualizer program saves 10 ms on a 8,000 ms
> render operation by avoiding accessor functions.  I'd rather spend
> our brain cycles optimizing the runtime and the in-core data so
> we spend less time in our tight revision traversal loops.
> 

I agree that for early versions they should definitely be functions, but since we've been talking about "final design" earlier in the thread, that's what I was referring to here to.

Show 11 quoted lines
> Seriously.  We make at least 10 or 11 function calls *per commit*
> that comes out of get_revision().  If the formatting application is
> really suffering from its 4 or 5 accessor function calls in order
> to get that returned data, we probably should also be looking at
> how we can avoid function cals in the library.
> 
> Oh, and even with 4 or 5 accessor functions per commit in the
> application that is *still* better than the 10 or so calls the
> application probably makes today scraping "git log --format=raw"
> off a pipe and segment it into the different fields it needs.
> 
Right.
Show 6 quoted lines
> Unless pipes in Linux somehow allow negative time warping with
> CPU counters.  Though on dual-core systems they might, since the
> two processes can run on different cores.  But oh, you didn't want
> to worry about threading support too much in libgit2, so I guess
> you also don't want to use multi-core systems.
> 

All comps where I use git today are multi-core systems, so I'm very much in favour of parallell processing. Otoh, I also don't think it makes sense to jump through hoops to make sure one can, fe, read and parse all tags in multiple threads simultaneously. In other words, I don't think we need to bother adding thread-safety for stuff that really only should happen if the application is poorly designed (unless it's straightforward to do so, ofcourse).

-- 
Andreas Ericsson                   andreas.ericsson@op5.se
OP5 AB                             www.op5.se
Tel: +46 8-230225                  Fax: +46 8-230231
Previous: Shawn O. PearceNext: Shawn O. Pearce
Message 17 of 83 in “libgit2 - a true git library”
  1. Shawn O. PearceOct 31, 2008
  2. Pieter de BieOct 31, 2008
  3. Pieter de BieOct 31, 2008
  4. Pierre HabouzitOct 31, 2008
  5. Shawn O. PearceOct 31, 2008
  6. Pierre HabouzitOct 31, 2008
  7. Shawn O. PearceOct 31, 2008
  8. Pierre HabouzitOct 31, 2008
  9. Junio C HamanoOct 31, 2008
  10. Shawn O. PearceOct 31, 2008
  11. Pierre HabouzitNov 1, 2008
  12. Andreas EricssonNov 1, 2008
  13. Pierre HabouzitNov 1, 2008
  14. Shawn O. PearceNov 1, 2008
  15. Andreas EricssonNov 1, 2008
  16. Shawn O. PearceNov 2, 2008
  17. Andreas EricssonNov 3, 2008
  18. Shawn O. PearceNov 2, 2008
  19. Pierre HabouzitNov 2, 2008
  20. Nicolas PitreOct 31, 2008
  21. david@lang.hmOct 31, 2008
  22. Nicolas PitreOct 31, 2008
  23. Shawn O. PearceOct 31, 2008
  24. Shawn O. PearceOct 31, 2008
  25. Pierre HabouzitOct 31, 2008
  26. Pierre HabouzitOct 31, 2008
  27. Nicolas PitreOct 31, 2008
  28. Andreas EricssonNov 1, 2008
  29. Pieter de BieOct 31, 2008
  30. Shawn O. PearceOct 31, 2008
  31. Junio C HamanoOct 31, 2008
  32. Pierre HabouzitNov 1, 2008
  33. Shawn O. PearceNov 1, 2008
  34. Pierre HabouzitNov 1, 2008
  35. Shawn O. PearceNov 1, 2008
  36. Nicolas PitreNov 1, 2008
  37. Shawn O. PearceNov 1, 2008
  38. Nicolas PitreNov 1, 2008
  39. Shawn O. PearceNov 1, 2008
  40. Johannes SchindelinNov 1, 2008
  41. Pierre HabouzitNov 1, 2008
  42. Nicolas PitreNov 1, 2008
  43. Pierre HabouzitNov 1, 2008
  44. Johannes SchindelinNov 1, 2008
  45. Junio C HamanoOct 31, 2008
  46. Pierre HabouzitOct 31, 2008
  47. Shawn O. PearceOct 31, 2008
  48. Jakub NarebskiOct 31, 2008
  49. david@lang.hmNov 1, 2008
  50. Shawn O. PearceNov 1, 2008
  51. david@lang.hmNov 1, 2008
  52. Pierre HabouzitNov 1, 2008
  53. Nicolas PitreNov 1, 2008
  54. Pierre HabouzitNov 1, 2008
  55. Nicolas PitreNov 1, 2008
  56. Shawn O. PearceNov 1, 2008
  57. Nicolas PitreNov 1, 2008
  58. Shawn O. PearceNov 1, 2008
  59. Scott ChaconNov 2, 2008
  60. Scott ChaconNov 2, 2008
  61. Shawn O. PearceNov 2, 2008
  62. David BrownNov 2, 2008
  63. Shawn O. PearceNov 3, 2008
  64. Pierre HabouzitNov 1, 2008
  65. david@lang.hmNov 1, 2008
  66. Brian GernhardtOct 31, 2008
  67. Andreas EricssonOct 31, 2008
  68. Shawn O. PearceOct 31, 2008
  69. Junio C HamanoOct 31, 2008
  70. Andreas EricssonNov 1, 2008
  71. Johannes SchindelinOct 31, 2008
  72. Bruno SantosOct 31, 2008
  73. Shawn O. PearceOct 31, 2008
  74. Andreas EricssonNov 1, 2008
  75. Shawn O. PearceNov 1, 2008
  76. Johannes SchindelinNov 2, 2008
  77. Pierre HabouzitNov 2, 2008
  78. Andreas EricssonNov 3, 2008
  79. Steve FrécinauxNov 8, 2008
  80. Andreas EricssonNov 8, 2008
  81. Pierre HabouzitNov 8, 2008
  82. Andreas EricssonNov 9, 2008
  83. Shawn O. PearceNov 9, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.