git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [RFC PATCH] Re: Empty directories...

From
Linus Torvalds <torvalds@linux-foundation.org>
Date
Jul 23, 2007, 23:57 UTC
Message-ID
<alpine.LFD.0.999.0707231638020.3607@woody.linux-foundation.org>
In-Reply-To
<873azen1c7.fsf@hades.wkstn.nix>
On Tue, 24 Jul 2007, Nix wrote:
Show 7 quoted lines
>
> On 23 Jul 2007, Linus Torvalds spake thusly:
> > So practically speaking, you want to track the *minimal* possible state, 
> > not the maximal one. 
> 
> I think it depends on your use case. For source code and indeed anything
> with heavy merges, this is true

Yes, very obviously. Git is targeted towards source code and working in a distributed manner across a very wide variety of users and setups, while something that would be more targeted towards a special scenario and much stricter usage would find that the "minimum" set is much bigger, and might well include ACL's and usr information.

Show 6 quoted lines
> but I'm increasingly using git as a sort of `merged historical tar' to 
> store images of entire random filesystem trees across time, and gaining 
> the benefit of the packer's lovely space-efficiency as well (doing this 
> with svn would be a lost cause, twice the space usage before you even 
> think about the repository). And in that case, preserving everything you 
> can makes sense.

On the other hand, almost all the space-efficiency comes from things that delta well, and change quickly. That includes the file data itself (and very much the tree contents), but it doesn't necessarily include things like permissions and user information - mainly because that doesn't actually delta at all (not because it can't, but because it hardly ever changes, and when it does change, it often changes all over the map).

To make an example of your "tar" situation: if you want to be space- efficient in a tar-like setting, you should *not* make user information be something that is per-file at all! Why? Because in 99% of all tar-files, there is a single user name.

So even your usage *may* actually be much better off using git as a "data backend", and using something totally different for "user/group" information. Yes, you'd have to make a "shim layer" on top of git to hide the fact that the user information is handled separately, but that shouldn't be that hard per se.

Show 5 quoted lines
> (Perhaps what I should be doing is tarring the directory tree up and
> storing the *tarball* in git. I'll try that and see what it does to pack
> sizes. These are version-controlled backups of my mother's magnum opus
> in progress so you can understand that I don't want to destroy them
> accidentally: I'd never hear the end of it! ;) )
You don't want to do this. 
There's a few reasons, but the two big ones are:
 - the git delta logic is strictly a "single delta base" thing.
   Yes, git would be able to find the delta's between two tar-files (as 
   long as you don't compress them), and express one tar-file in terms of 
   the other, and it would probably save a fair amount of disk.
   But it would not be able to do _nearly_ as well as it can if you store 
   individual files, and let git just find the best delta per-file (and 
   not just "one delta base for the whole tar-ball")
 - git is very much optimized for "many small files". Yes, you can check 
   in large files, and it works fine, but quite frankly, all the design 
   and heavy optimizations have been about having trees with tens of 
   thousands of files, but the files individually reasonably small.
   A lot of the speed advantages of git come from efficiently pruning away 
   whole sub-directory structures, for example, and not even touching the 
   data at all!
   So if you track just one file that changes in every version, all the 
   things that make git fly are basically disabled, and you won't take 
   full advantage of what git does.
Show 6 quoted lines
> Yes indeed: that's why I proposed doing this using a couple of new hooks
> driving entirely optional permissions-preservation stuff. Most use cases
> really won't want to track this, so this sort of stuff shouldn't impose
> upon the git core or upon anyone who doesn't want it. (However, the
> ability to have alternative file merging strategies *may* be useful
> elsewhere, perhaps.)

The ".gitattributes" file really could be used for some of that. Using it to track ownership and full permissions would not be impossible, and it could have interesting semantics (especially as .gitattibutes is path pattern based - so you could literally do a "user" attribute, and say that everything in a particular subdirectory is owned by a particular user).

That wouldn't be UNIX-like semantics, of course, but it can be very useful for certain things.

Taking an example of something totally independent of git, look at how "udev" handles permissions, for example. In situations like that, static user information is useless, and it actually ends up setting up modes and ownership based on name-based patterns rather than having each file have a permission/user (because individual files appear and disappear, the name-based patterns are the things that matter).

So if you *just* want to track a regular filesystem layout, that's not the right thing, but "udev" does show an example of a totally different way of describing ownership and permissions, and one which wouldn't actually be at all foreign to git.

		Linus
Previous: NixNext: David Kastrup
Message 37 of 137 in “Empty directories...”
  1. David KastrupJul 18, 2007
  2. Johannes SchindelinJul 18, 2007
  3. David KastrupJul 18, 2007
  4. Johannes SchindelinJul 18, 2007
  5. Linus TorvaldsJul 18, 2007
  6. Linus TorvaldsJul 18, 2007
  7. David KastrupJul 18, 2007
  8. Linus TorvaldsJul 18, 2007
  9. Matthieu MoyJul 18, 2007
  10. Linus TorvaldsJul 18, 2007
  11. David KastrupJul 18, 2007
  12. Linus TorvaldsJul 18, 2007
  13. David KastrupJul 18, 2007
  14. Re: Empty directories...Linus Torvalds, Jul 18, 2007
  15. Linus TorvaldsJul 18, 2007
  16. David KastrupJul 18, 2007
  17. Linus TorvaldsJul 19, 2007
  18. Junio C HamanoJul 19, 2007
  19. Shawn O. PearceJul 19, 2007
  20. David KastrupJul 19, 2007
  21. Geoff RussellJul 19, 2007
  22. Shawn O. PearceJul 19, 2007
  23. Matthieu MoyJul 19, 2007
  24. Tomash BrechkoJul 19, 2007
  25. David KastrupJul 19, 2007
  26. Tomash BrechkoJul 19, 2007
  27. David KastrupJul 19, 2007
  28. NixJul 23, 2007
  29. David KastrupJul 23, 2007
  30. NixJul 23, 2007
  31. NixJul 23, 2007
  32. Jakub NarebskiJul 23, 2007
  33. NixJul 25, 2007
  34. David KastrupJul 23, 2007
  35. Linus TorvaldsJul 23, 2007
  36. NixJul 23, 2007
  37. Linus TorvaldsJul 23, 2007
  38. David KastrupJul 19, 2007
  39. David KastrupJul 19, 2007
  40. Johannes SchindelinJul 19, 2007
  41. David KastrupJul 19, 2007
  42. Brian GernhardtJul 19, 2007
  43. Johannes SchindelinJul 19, 2007
  44. Brian GernhardtJul 19, 2007
  45. Johannes SchindelinJul 19, 2007
  46. David KastrupJul 19, 2007
  47. Brian GernhardtJul 19, 2007
  48. Johannes SchindelinJul 19, 2007
  49. David KastrupJul 19, 2007
  50. Matthieu MoyJul 19, 2007
  51. David KastrupJul 19, 2007
  52. David KastrupJul 19, 2007
  53. David KastrupJul 19, 2007
  54. David KastrupJul 21, 2007
  55. Linus TorvaldsJul 21, 2007
  56. Linus TorvaldsJul 21, 2007
  57. David KastrupJul 21, 2007
  58. Linus TorvaldsJul 21, 2007
  59. David KastrupJul 21, 2007
  60. Simon 'corecode' SchubertJul 21, 2007
  61. David KastrupJul 21, 2007
  62. Linus TorvaldsJul 21, 2007
  63. David KastrupJul 22, 2007
  64. Linus TorvaldsJul 22, 2007
  65. David KastrupJul 22, 2007
  66. Linus TorvaldsJul 22, 2007
  67. David KastrupJul 22, 2007
  68. Linus TorvaldsJul 22, 2007
  69. David KastrupJul 22, 2007
  70. david@lang.hmJul 22, 2007
  71. David KastrupJul 22, 2007
  72. Linus TorvaldsJul 22, 2007
  73. David KastrupJul 22, 2007
  74. Linus TorvaldsJul 22, 2007
  75. Linus TorvaldsJul 22, 2007
  76. David KastrupJul 22, 2007
  77. Jakub NarebskiJul 22, 2007
  78. David KastrupJul 22, 2007
  79. Jakub NarebskiJul 22, 2007
  80. David KastrupJul 22, 2007
  81. Jakub NarebskiJul 22, 2007
  82. David KastrupJul 22, 2007
  83. David KastrupJul 23, 2007
  84. David KastrupJul 23, 2007
  85. David KastrupJul 22, 2007
  86. Brian GernhardtJul 22, 2007
  87. David KastrupJul 28, 2007
  88. David KastrupJul 18, 2007
  89. Matthieu MoyJul 18, 2007
  90. David KastrupJul 18, 2007
  91. Shawn O. PearceJul 18, 2007
  92. Junio C HamanoJul 18, 2007
  93. David KastrupJul 18, 2007
  94. Wincent ColaiutaJul 18, 2007
  95. Junio C HamanoJul 18, 2007
  96. Johan HerlandJul 20, 2007
  97. David KastrupJul 20, 2007
  98. Johan HerlandJul 20, 2007
  99. David KastrupJul 20, 2007
  100. Johan HerlandJul 20, 2007
  101. David KastrupJul 22, 2007
  102. Robin RosenbergJul 26, 2007
  103. David KastrupJul 27, 2007
  104. Johannes SchindelinJul 18, 2007
  105. Matthieu MoyJul 18, 2007
  106. David KastrupJul 18, 2007
  107. Junio C HamanoJul 18, 2007
  108. Brian GernhardtJul 19, 2007
  109. David KastrupJul 19, 2007
  110. Brian GernhardtJul 19, 2007
  111. Junio C HamanoJul 20, 2007
  112. Linus TorvaldsJul 20, 2007
  113. Linus TorvaldsJul 20, 2007
  114. Junio C HamanoJul 20, 2007
  115. Linus TorvaldsJul 20, 2007
  116. David KastrupJul 20, 2007
  117. David KastrupJul 20, 2007
  118. Linus TorvaldsJul 20, 2007
  119. David KastrupJul 20, 2007
  120. Simon 'corecode' SchubertJul 20, 2007
  121. David KastrupJul 20, 2007
  122. Junio C HamanoJul 20, 2007
  123. David KastrupJul 20, 2007
  124. Linus TorvaldsJul 20, 2007
  125. Johan HerlandJul 20, 2007
  126. Linus TorvaldsJul 20, 2007
  127. Julian PhillipsJul 20, 2007
  128. Linus TorvaldsJul 21, 2007
  129. David KastrupJul 21, 2007
  130. David KastrupJul 21, 2007
  131. David KastrupJul 20, 2007
  132. Olivier GalibertJul 20, 2007
  133. Johan HerlandJul 20, 2007
  134. David KastrupJul 20, 2007
  135. David KastrupJul 21, 2007
  136. David KastrupJul 22, 2007
  137. NixJul 24, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.