git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Figured out how to get Mozilla into git

From
RDRogan Dawes <lists@dawes.za.net>
Date
Jun 10, 2006, 18:36 UTC
Message-ID
<448B1130.8020005@dawes.za.net>
In-Reply-To
<Pine.LNX.4.64.0606101041490.5498@g5.osdl.org>
Linus Torvalds wrote:
Show 23 quoted lines
> 
> On Sat, 10 Jun 2006, Rogan Dawes wrote:
>> Here's an idea. How about separating trees and commits from the actual blobs
>> (e.g. in separate packs)? My reasoning is that the commits and trees should
>> only be a small portion of the overall repository size, and should not be that
>> expensive to transfer. (Of course, this is only a guess, and needs some
>> numbers to back it up.)
> 
> The trees in particular are actually a pretty big part of the history. 
> 
> More importantly, the blobs compress horribly badly in the absense of 
> history - a _lot_ of the compression in git packing comes very much from 
> the fact that we do a good job at delta-compression.
> 
> So if you get all of the commit/tree history, but none of the blob 
> history, you're actually not going to win that much space. As already 
> discussed, the _whole_ history packed with git is usually not insanely 
> bigger than just the whole unpacked tree (with no history at all).
> 
> So you'd think that getting just the top version of the tree would be a 
> much bigger space-saving that it actually is. If you _also_ get all the 
> tree and commit objects, the space saving is even less.
> 

One possibility, given that the full commit and tree history is so large, is simply to get the HEAD commit and the trees that the commit depends directly on, rather than fetching them all up front.

Show 14 quoted lines
> I actually suspect that the most realistic way to handle this is to use 
> the "fetch.c" logic (ie the incremental fetcher used by http), and add 
> some mode to the git daemon where you fetch literally one object at a time 
> (ie this would be totally _separate_ from the pack-file thing: you'd not 
> ask for "git-upload-pack", you'd ask for something like 
> "git-serve-objects" instead).
> 
> The fetch.c logic really does allow for on-demand object fetching, and is 
> thus much more suitable for incomplete repositories.
> 
> HOWEVER. The fetch.c logic - by necessity - works on a object-by-object 
> level. That means that you'd get no delta compression AT ALL, and I 
> suspect that the downside of that would be a factor of ten expansion or 
> more, which means that it would really not work that well in practice.

Would it be possible to add a mode where fetch.c is given a list of desired objects, and returns a list of pointers to those objects? Then callers that already have such a list could be modified to pass the whole list at once, allowing at least SOME compression, and optimisation of round trips, etc? There would be a tradeoff in memory use, though, I guess.

Rogan
Previous: Jon SmirlNext: Linus Torvalds
Message 47 of 67 in “Figured out how to get Mozilla into git”
  1. Jon SmirlJun 9, 2006
  2. Nicolas PitreJun 9, 2006
  3. Martin LanghoffJun 9, 2006
  4. Jon SmirlJun 9, 2006
  5. Jakub NarebskiJun 9, 2006
  6. Linus TorvaldsJun 9, 2006
  7. Nicolas PitreJun 9, 2006
  8. Linus TorvaldsJun 9, 2006
  9. Nicolas PitreJun 9, 2006
  10. Linus TorvaldsJun 9, 2006
  11. Jakub NarebskiJun 9, 2006
  12. Jon SmirlJun 9, 2006
  13. Linus TorvaldsJun 9, 2006
  14. Jon SmirlJun 9, 2006
  15. Linus TorvaldsJun 9, 2006
  16. Jon SmirlJun 9, 2006
  17. Linus TorvaldsJun 9, 2006
  18. Linus TorvaldsJun 9, 2006
  19. Greg KHJun 9, 2006
  20. Martin LanghoffJun 9, 2006
  21. Linus TorvaldsJun 9, 2006
  22. Jon SmirlJun 10, 2006
  23. Linus TorvaldsJun 10, 2006
  24. Jon SmirlJun 10, 2006
  25. Jon SmirlJun 10, 2006
  26. Jakub NarebskiJun 9, 2006
  27. Nicolas PitreJun 9, 2006
  28. Jon SmirlJun 9, 2006
  29. Martin LanghoffJun 10, 2006
  30. Martin LanghoffJun 10, 2006
  31. Linus TorvaldsJun 10, 2006
  32. Linus TorvaldsJun 10, 2006
  33. Jon SmirlJun 10, 2006
  34. Linus TorvaldsJun 10, 2006
  35. Jon SmirlJun 10, 2006
  36. Carl WorthJun 10, 2006
  37. Linus TorvaldsJun 10, 2006
  38. Jakub NarebskiJun 10, 2006
  39. Junio C HamanoJun 10, 2006
  40. Rogan DawesJun 10, 2006
  41. Junio C HamanoJun 10, 2006
  42. Rogan DawesJun 10, 2006
  43. Jakub NarebskiJun 10, 2006
  44. Nicolas PitreJun 10, 2006
  45. Linus TorvaldsJun 10, 2006
  46. Jon SmirlJun 10, 2006
  47. Rogan DawesJun 10, 2006
  48. Linus TorvaldsJun 10, 2006
  49. Jon SmirlJun 10, 2006
  50. Martin LanghoffJun 10, 2006
  51. Junio C HamanoJun 10, 2006
  52. Linus TorvaldsJun 10, 2006
  53. Linus TorvaldsJun 10, 2006
  54. Jon SmirlJun 10, 2006
  55. Junio C HamanoJun 10, 2006
  56. Jon SmirlJun 10, 2006
  57. Timo HirvonenJun 10, 2006
  58. Petr BaudisJun 10, 2006
  59. Lars JohannsenJun 10, 2006
  60. Nicolas PitreJun 11, 2006
  61. Linus TorvaldsJun 18, 2006
  62. Martin LanghoffJun 18, 2006
  63. Linus TorvaldsJun 18, 2006
  64. Broken PPC sha1.. (Re: Figured out how to get Mozilla into git)Linus Torvalds, Jun 18, 2006
  65. Fix PPC SHA1 routine for large input buffersPaul Mackerras, Jun 18, 2006
  66. Linus TorvaldsJun 19, 2006
  67. Pavel RoskinJun 9, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.