git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Unresolved issues #2 (shallow clone again)

From
Linus Torvalds <torvalds@osdl.org>
Date
May 8, 2006, 15:32 UTC
Message-ID
<Pine.LNX.4.64.0605080813590.3718@g5.osdl.org>
In-Reply-To
<20060508042429.GA20249@coredump.intra.peff.net>
On Mon, 8 May 2006, Jeff King wrote:
Show 12 quoted lines
>
> On Sun, May 07, 2006 at 08:27:02AM -0700, Linus Torvalds wrote:
> 
> > factor for a lot of things for many "common" filesystem setups. You 
> > probably didn't even account for the size of inodes in your "du" setup.
> 
> My numbers came from git-count-objects, which uses the st_blocks sum for
> all objects. The actual du numbers showing space wasted by block
> boundaries are:
>   du -c ??: 1429216
>   du -c --apparent-size ??: 792277
> So it's about 45% wasted space.

And that's actually ignoring inode sizes and directory sizes (well, it doesn't "ignore" directory sizes - it counts them - but if you compare it to a straight packed format, it's still overhead).

Anyway, looks like it's about 2:1, not 3:1 like I claimed, but the point being that blocking factors tend to be at least on the same order of magnitude as just plain compression (which also tends to be in the 2:1 area for normal, fairly easily compressible, stuff).

The delta-packing obviously is much bigger for any project with real history. In traditional setups (where you always delta-pack within one thing, ie at the level of individual SCCS/RCS files), the delta packing obviously _also_ avoids blocking issues, since it means that a thousand revisions of the same file will all share the same inode.

So because git uses a whole-file model, it obviously makes the blocking issues with its unpacked format _much_ higher than for any traditional medium - no conglomeration of different versions of the file in the same filesystem object. On the other hand, the packed format also tends to be even _more_ efficient than a traditional one, so the end result of it all is apparently a pretty big net win even in space consumption).

Side note: I realize that some people think the packs are ugly and strange. They aren't linear versions of a file, and instead appear as a fairly random "jumble". And they can't be incrementally re-packed, and you have to generate a whole new pack-file (which can be incremental in _content_, of course). So people think they are ugly.

I'd argue that they are beautiful. They are beautiful because they _don't_ contain history in themselves (the objects they contain encode the history of course, but the pack-file itself does not).

And they are beautiful because we can use the exact same format for streaming data over the network as for the database itself (that, of course, was just about _the_ design consideration). Show me another system that has exactly the same (not "similar", not "same concepts": _same_) network protocol as it internal database.

And they are beautiful exactly because their lack of any internal structure allows you to pack things by criteria _you_ care about, ie the whole "sort things by recency" thing, so that commonly accessed data can be packed at the head of the pack-file - exactly because the pack-file doesn't have any internal structure of its own that you need to worry about and that constrains your sorting.

The same thing is what allows you to delta any blob against any other blob - without worrying about history or other random pack-file rules. You can do packign purely by how well you want to pack, not by any secondary constraints.

And the "no incremental updates" may sound like a huge downside, but it's all the same basic git logic: objects and filesystem contents are immutable, and that allows us to avoid a lot of locking overhead. Locking is _hard_. Locking is _inefficient_. And locking really really screws you when you miss it.

So I'll happily say that pack-files are strange, and that you have to get a bit used to the notion that they should be repacked "asynchronously". But it's really a matter of "getting used to it", because once you do, you'll see that it's actually an absolutely huge deal, and you'll learn to love the bomb^H^H^H^Hpack-file.

			Linus "pack-files rule" Torvalds
Previous: Jeff KingNext: Sergey Vlasov
Message 52 of 81 in “Recent unresolved issues”
  1. Junio C HamanoApr 14, 2006
  2. Petr BaudisApr 14, 2006
  3. seanApr 14, 2006
  4. Petr BaudisApr 14, 2006
  5. Carl WorthApr 14, 2006
  6. Johannes SchindelinApr 15, 2006
  7. Junio C HamanoApr 15, 2006
  8. Junio C HamanoApr 15, 2006
  9. Linus TorvaldsApr 14, 2006
  10. Linus TorvaldsApr 15, 2006
  11. Linus TorvaldsApr 15, 2006
  12. Junio C HamanoApr 15, 2006
  13. Linus TorvaldsApr 15, 2006
  14. Linus TorvaldsApr 15, 2006
  15. Linus TorvaldsApr 15, 2006
  16. Junio C HamanoApr 15, 2006
  17. Junio C HamanoApr 15, 2006
  18. Junio C HamanoApr 15, 2006
  19. Johannes SchindelinApr 15, 2006
  20. Linus TorvaldsApr 15, 2006
  21. Linus TorvaldsApr 15, 2006
  22. Junio C HamanoApr 16, 2006
  23. Junio C HamanoApr 15, 2006
  24. Linus TorvaldsApr 15, 2006
  25. Junio C HamanoApr 15, 2006
  26. Unresolved issues #2Junio C Hamano, May 4, 2006
  27. Jakub NarebskiMay 4, 2006
  28. Junio C HamanoMay 4, 2006
  29. Jakub NarebskiMay 4, 2006
  30. Petr BaudisMay 4, 2006
  31. Pavel RoskinMay 4, 2006
  32. Carl WorthMay 4, 2006
  33. Junio C HamanoMay 5, 2006
  34. Martin LanghoffMay 5, 2006
  35. Carl WorthMay 5, 2006
  36. Jakub NarebskiMay 5, 2006
  37. Linus TorvaldsMay 5, 2006
  38. Jakub NarebskiMay 5, 2006
  39. Linus TorvaldsMay 5, 2006
  40. Martin LanghoffMay 6, 2006
  41. Junio C HamanoMay 6, 2006
  42. Martin LanghoffMay 7, 2006
  43. Jeff KingMay 7, 2006
  44. Linus TorvaldsMay 7, 2006
  45. Theodore TsoMay 8, 2006
  46. Linus TorvaldsMay 8, 2006
  47. Theodore TsoMay 8, 2006
  48. Linus TorvaldsMay 8, 2006
  49. Theodore TsoMay 8, 2006
  50. Linus TorvaldsMay 8, 2006
  51. Jeff KingMay 8, 2006
  52. Linus TorvaldsMay 8, 2006
  53. Sergey VlasovMay 7, 2006
  54. Martin LanghoffMay 7, 2006
  55. Junio C HamanoMay 7, 2006
  56. Martin LanghoffMay 7, 2006
  57. Carl WorthMay 5, 2006
  58. Jakub NarebskiMay 7, 2006
  59. Junio C HamanoMay 8, 2006
  60. Jakub NarebskiMay 8, 2006
  61. Jakub NarebskiMay 8, 2006
  62. Daniel BarkalowMay 4, 2006
  63. Linus TorvaldsMay 4, 2006
  64. Junio C HamanoMay 6, 2006
  65. Linus TorvaldsMay 6, 2006
  66. seanMay 6, 2006
  67. Linus TorvaldsMay 6, 2006
  68. seanMay 6, 2006
  69. Linus TorvaldsMay 6, 2006
  70. Junio C HamanoMay 6, 2006
  71. Johannes SchindelinMay 6, 2006
  72. Linus TorvaldsMay 6, 2006
  73. Junio C HamanoMay 7, 2006
  74. Junio C HamanoMay 7, 2006
  75. Johannes SchindelinMay 7, 2006
  76. Jakub NarebskiMay 7, 2006
  77. Junio C HamanoMay 8, 2006
  78. Jakub NarebskiMay 7, 2006
  79. David WoodhouseMay 9, 2006
  80. Bertrand JacquinMay 9, 2006
  81. Nicolas PitreMay 9, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.