git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Unresolved issues #2 (shallow clone again)

From
Linus Torvalds <torvalds@osdl.org>
Date
May 7, 2006, 15:27 UTC
Message-ID
<Pine.LNX.4.64.0605070802590.16343@g5.osdl.org>
In-Reply-To
<20060507075631.GA24423@coredump.intra.peff.net>
On Sun, 7 May 2006, Jeff King wrote:
Show 7 quoted lines
>
>   - Total savings by going shallow: 10.7%
> 
> So basically, trees and commits DON'T compress as well as historical
> blobs (potentially because git-pack-objects isn't currently optimized
> for this -- I haven't checked). As a result, we're saving only 10% by
> going shallow instead of a potential 50%.

The biggest size savers from packing is (in rough order of relevance, if I recall the rough statistics I did):

 - avoiding block boundaries. 
 - delta packing of blobs
 - delta packing of trees
 - regular compression

The block boundaries are huge, we have tons of small objects, and that was one of the primary reasons for packing. I'd suspect that this is a 3:1 factor for a lot of things for many "common" filesystem setups. You probably didn't even account for the size of inodes in your "du" setup.

And blobs with history generally delta very well (_much_ better than regular compression).

Trees should _delta_ very well, but they basically don't compress, especially after deltaing. The SHA1's are totally incompressible (in a tree they aren't even ASCII), and as a deta, the names won't compress much either because they are short.

Commits are fairly small, shouldn't delta all that much, and they don't even compress _that_ well either (they're normal text and often have some redundancy with the committer and author being the same, but they are short and have some fairly incompressible elements, so..)

The thing with trees in particular is that they are very common for the kernel (and probably not so much for many other projects). A single commit ends up quite commonly being just one commit object, one blob (that deltas really well), and three or four trees. Merges often have no new blobs at all, just several new trees and the commit object.

So a huge amount of the wins from packing come from the file _history_, the part that a shallow clone (on purpose) leaves behind.

The regular compression will pick up a fair amount of slack with the blobs, but it's a much smaller factor than the delta compression for something that has a long history.

It's somewhat interesting to note that over the year that we've used git, the kernel pack-size hasn't even increased all that much. I forget exactly what it was when we started packing, but it was on the order of ~75M. It is now 115M for me. And the old linux-history thing (full BK history over three years) is 177M - not much more than twice the size of just a few kernel versions - with some higher packing ratios..

Exactly because blobs delta so incredibly well.
		Linus
Previous: Jeff KingNext: Theodore Tso
Message 44 of 81 in “Recent unresolved issues”
  1. Junio C HamanoApr 14, 2006
  2. Petr BaudisApr 14, 2006
  3. seanApr 14, 2006
  4. Petr BaudisApr 14, 2006
  5. Carl WorthApr 14, 2006
  6. Johannes SchindelinApr 15, 2006
  7. Junio C HamanoApr 15, 2006
  8. Junio C HamanoApr 15, 2006
  9. Linus TorvaldsApr 14, 2006
  10. Linus TorvaldsApr 15, 2006
  11. Linus TorvaldsApr 15, 2006
  12. Junio C HamanoApr 15, 2006
  13. Linus TorvaldsApr 15, 2006
  14. Linus TorvaldsApr 15, 2006
  15. Linus TorvaldsApr 15, 2006
  16. Junio C HamanoApr 15, 2006
  17. Junio C HamanoApr 15, 2006
  18. Junio C HamanoApr 15, 2006
  19. Johannes SchindelinApr 15, 2006
  20. Linus TorvaldsApr 15, 2006
  21. Linus TorvaldsApr 15, 2006
  22. Junio C HamanoApr 16, 2006
  23. Junio C HamanoApr 15, 2006
  24. Linus TorvaldsApr 15, 2006
  25. Junio C HamanoApr 15, 2006
  26. Unresolved issues #2Junio C Hamano, May 4, 2006
  27. Jakub NarebskiMay 4, 2006
  28. Junio C HamanoMay 4, 2006
  29. Jakub NarebskiMay 4, 2006
  30. Petr BaudisMay 4, 2006
  31. Pavel RoskinMay 4, 2006
  32. Carl WorthMay 4, 2006
  33. Junio C HamanoMay 5, 2006
  34. Martin LanghoffMay 5, 2006
  35. Carl WorthMay 5, 2006
  36. Jakub NarebskiMay 5, 2006
  37. Linus TorvaldsMay 5, 2006
  38. Jakub NarebskiMay 5, 2006
  39. Linus TorvaldsMay 5, 2006
  40. Martin LanghoffMay 6, 2006
  41. Junio C HamanoMay 6, 2006
  42. Martin LanghoffMay 7, 2006
  43. Jeff KingMay 7, 2006
  44. Linus TorvaldsMay 7, 2006
  45. Theodore TsoMay 8, 2006
  46. Linus TorvaldsMay 8, 2006
  47. Theodore TsoMay 8, 2006
  48. Linus TorvaldsMay 8, 2006
  49. Theodore TsoMay 8, 2006
  50. Linus TorvaldsMay 8, 2006
  51. Jeff KingMay 8, 2006
  52. Linus TorvaldsMay 8, 2006
  53. Sergey VlasovMay 7, 2006
  54. Martin LanghoffMay 7, 2006
  55. Junio C HamanoMay 7, 2006
  56. Martin LanghoffMay 7, 2006
  57. Carl WorthMay 5, 2006
  58. Jakub NarebskiMay 7, 2006
  59. Junio C HamanoMay 8, 2006
  60. Jakub NarebskiMay 8, 2006
  61. Jakub NarebskiMay 8, 2006
  62. Daniel BarkalowMay 4, 2006
  63. Linus TorvaldsMay 4, 2006
  64. Junio C HamanoMay 6, 2006
  65. Linus TorvaldsMay 6, 2006
  66. seanMay 6, 2006
  67. Linus TorvaldsMay 6, 2006
  68. seanMay 6, 2006
  69. Linus TorvaldsMay 6, 2006
  70. Junio C HamanoMay 6, 2006
  71. Johannes SchindelinMay 6, 2006
  72. Linus TorvaldsMay 6, 2006
  73. Junio C HamanoMay 7, 2006
  74. Junio C HamanoMay 7, 2006
  75. Johannes SchindelinMay 7, 2006
  76. Jakub NarebskiMay 7, 2006
  77. Junio C HamanoMay 8, 2006
  78. Jakub NarebskiMay 7, 2006
  79. David WoodhouseMay 9, 2006
  80. Bertrand JacquinMay 9, 2006
  81. Nicolas PitreMay 9, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.