git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [RFC PATCH] Re: Empty directories...

From
David Kastrup <dak@gnu.org>
Date
Jul 22, 2007, 13:53 UTC
Message-ID
<857iosmto0.fsf@lola.goethe.zz>
In-Reply-To
<200707221406.25541.jnareb@gmail.com>

Jakub, this mail is too long already, and it does not make sense to tack a changed proposal to its end since then the readers will be exhausted at the time they come there. So I'll instead tack a followup to the "big picture" mail instead where I outline a modified approach which is presumably easier to understand and completely backwards-compatible, incorporating your feedback.

There is probably little sense in wasting your time on a detailed response: feel free to point out where you don't see myself making sense. I have no problem with people coming to different conclusions that I do, but I would prefer it if it is not because they consider myself a raving lunatic, but because they have different opinions regarding the details.

"I can follow you, but I disagree with your conclusion" is perfectly fine for now since I am going to propose something else, anyway.

Thanks for the feedback.  It gave me some good ideas.
Jakub Narebski <jnareb@gmail.com> writes:
Show 8 quoted lines
> On Sun, 22 July 2007, David Kastrup wrote:
>> Jakub Narebski <jnareb@gmail.com> writes:
>>> David Kastrup wrote:
>>>
>>>> I must be really bad at explaining things, or I am losing a fight
>>>> against preconceptions fixed beyond my imagination.
>
> Or you are wrong...

Well, there is little reason for you to take my word on it, but I happen to have a history of designing and implementing systems where I have been responsible for every single byte, bootloader, firmware, applications, target compiler, assembler, whatever. I have been exposed to Unix and working with it several years before Linux even existed. I also have a track record of being not exactly stupid.

So I pretty much can rule out that I am wrong on the factual side.

But where I may be wrong is in estimating the how obvious the design can appear to others, and how useful and maintainable for others it may be in the long run. Linus says "code talks", but that's actually not half the story. If my code says that it works and the evidence is there, but nobody is able to understand _why_ it works, it has no place in a project where I am not permanently around.

If smart people don't get what I am talking about, it does not matter that the patch is surprisingly well-contained: it will be a maintenance nightmare because people will never figure out why something stopped working after some particular change.

Show 7 quoted lines
>> I disagree here.  The object database _can_ represent an _empty_
>> directory that has been added explicitly, because up to now no
>> operations existed that actually left an empty tree.  But it can't
>> distinguish a _non_-empty directory that has been added explicitly
>> from non-empty directory that has not been added explicitly.
>
> True. I forgot about that.

Thanks. It is almost a revelation that anybody can agree on any point with me at the moment.

Show 5 quoted lines
> IMHO it would be best to first provide plumbing infrastructure (as
> e.g.  it was the case of submodule support), then add option to
> git-update-index to change the "stickiness"/"autoremoval" status of
> a directory (of a tree), and _last_ think about how to change the
> porcelain (git-add and git-rm).

Sure. It does no harm to think about reducing the amount of breaking porcelain, though.

Show 12 quoted lines
> [...]
>
>> And a perfectly consistent way is to make those trees with an
>> explicitly added directory _non-empty_, by virtue of putting a file
>> "." in them.  This file, of course, exists in every physical
>> directory, but we may or may not decide to let it be tracked by
>> git, using the gitignore mechanism on the pattern ".".  Perfectly
>> expedient.
>
> Here we disagree. I think putting "." in a tree as marker of having
> it not be automatically deleted when empty, as opposed to marking
> tree using filemode in the parent, is not a good idea.

Well, "not a good idea" is a far step forward from "stupid idiot babbling nonsense", so we may make progress towards actually being able to _weigh_ different options. I can actually associate with "not a good idea", not least because nobody else seems to get the idea, and that makes it infeasible for maintenance.

So I'll address some points and then propose a different way of implementing what will in the end amount to rather similar semantics, but with a different view of looking at those semantics, one that corresponds well with the implementation.

Show 8 quoted lines
> The only advantage to the "." idea is that it can use gitignore
> mechanism (both in-tree .gitignore, tracked or not, and info/exclude
> file). But I also think that the fact that gitignore mechanism is
> recursive is more of disadvantage than advantage.
>
> First, it is _not_ consistent. Working directory trees _always_ have
> '.'  in them, while trees would have or would have not it, depending
> if they would be "sticky" or "autoremoved".

Let me point out again that this inconsistency is already present in the difference of tracked and untracked _files_: they are always in the working directory, while trees have or not have them, depending on whether they are "registered" or "not".

There is no inconsistency involved here, but it seems to make people _very_ uncomfortable to factor out the "stays around even if empty" functionality and call it "dir/." from the "can hold content" functionality which is in effect called "dir/", and basically associate tracked physical existence just with the former.

The recursiveness of the gitignore mechanism has the advantage that when maintaining a large repository with actual or logical subprojects, one does not need to pick a single policy for all subprojects. I think that is quite important. It could possibly be achieved with some other method of having per-subproject configuration, but I see little wrong in using what is there and documented already.

Show 7 quoted lines
> Second, the "easy implementation" is anything but easy. "git add ."
> as a way to mark directory as "sticky" is not backward compatibile:
> currently it mean to add _all contents_ of current directory.
> Implementation is tricky: as we have seen trying to unlink '.' or
> create '.' can unfortunately succeed on [some Sun OS, and UFS
> filesystem] (which follows POSIX stupidly to the letter) f**king up
> the filesystem.

I was not suggesting actually leaving any such calls in place: after all, they would presumably lead to error messages. But I agree that this could lead to nasty surprises when somebody with a legacy version of git worked with a repository containing "." as explicit entries of some file type.

> The alternative proposal of adding "magic mode" to mark directory as
> "not remove when empty" is largely tested; it is very similar to the
> subproject support.
Good.  Because it is what I converged to last night.
Show 8 quoted lines
> Third, is contrary to the git philosophy of tracking contents.
> "Stickiness" is an attribute; the fact that directory is explicitely
> tracked or not does not change contents of a directory. Compare to
> 'blob' which contains only contents of a file: not a filename, not a
> pathname, not [subset of] filemode.
>
> Fourth, is very artificial. What would you put for filemode for '.'?
> 040000 (i.e. directory)?

Taken already. By something very artificial, namely a tree... Yes, this was a wart in my proposal.

> What would you put for sha1?  Sha1 of an empty directory?
Some fixed value.  Everywhere the same.  Not really relevant.
Show 9 quoted lines
>> That basically implies that no information about directories could
>> be tracked in the repository.  And yes, we need appropriate
>> information in the index.  Again, the information whether a
>> directory was added explicitly.
>
> Whether directory is automatically managed by git (automatically
> removed or untracked). But we need directory entry in index for
> git-diff, for example to recognize if there is or there is not empty
> directory, or if a directory is automanaged or not.

One conclusion that I have come to (and I think I am in agreement with Linus here) is that the information "empty or not" is actually useless separately: when I add files below a directory to the repository, the directory _can't_ be empty. And git has no way of knowing whether it is non-empty because I wanted the directory to be there, or whether it is non-empty because I could not have checked in the files into the tree below it otherwise.

Show 9 quoted lines
>> And the repository is a versioned and hierarchically hashed version
>> of the index, but its trees contain _no_ information that is not
>> already inherently represented by the files alone. [...]
>
> The above sentence is nonsensical. Index is helper for repository,
> and can be derived from repository. Not vice versa.
>
> Trees do contain information which is not inherently present by the 
> blobs.

Could you give examples for such information? As long as we are not talking about _history_, I am at a loss at what else you mean. File names and permissions?

-- 
David Kastrup, Kriemhildstr. 15, 44793 Bochum
Previous: Jakub NarebskiNext: Jakub Narebski
Message 80 of 137 in “Empty directories...”
  1. David KastrupJul 18, 2007
  2. Johannes SchindelinJul 18, 2007
  3. David KastrupJul 18, 2007
  4. Johannes SchindelinJul 18, 2007
  5. Linus TorvaldsJul 18, 2007
  6. Linus TorvaldsJul 18, 2007
  7. David KastrupJul 18, 2007
  8. Linus TorvaldsJul 18, 2007
  9. Matthieu MoyJul 18, 2007
  10. Linus TorvaldsJul 18, 2007
  11. David KastrupJul 18, 2007
  12. Linus TorvaldsJul 18, 2007
  13. David KastrupJul 18, 2007
  14. Re: Empty directories...Linus Torvalds, Jul 18, 2007
  15. Linus TorvaldsJul 18, 2007
  16. David KastrupJul 18, 2007
  17. Linus TorvaldsJul 19, 2007
  18. Junio C HamanoJul 19, 2007
  19. Shawn O. PearceJul 19, 2007
  20. David KastrupJul 19, 2007
  21. Geoff RussellJul 19, 2007
  22. Shawn O. PearceJul 19, 2007
  23. Matthieu MoyJul 19, 2007
  24. Tomash BrechkoJul 19, 2007
  25. David KastrupJul 19, 2007
  26. Tomash BrechkoJul 19, 2007
  27. David KastrupJul 19, 2007
  28. NixJul 23, 2007
  29. David KastrupJul 23, 2007
  30. NixJul 23, 2007
  31. NixJul 23, 2007
  32. Jakub NarebskiJul 23, 2007
  33. NixJul 25, 2007
  34. David KastrupJul 23, 2007
  35. Linus TorvaldsJul 23, 2007
  36. NixJul 23, 2007
  37. Linus TorvaldsJul 23, 2007
  38. David KastrupJul 19, 2007
  39. David KastrupJul 19, 2007
  40. Johannes SchindelinJul 19, 2007
  41. David KastrupJul 19, 2007
  42. Brian GernhardtJul 19, 2007
  43. Johannes SchindelinJul 19, 2007
  44. Brian GernhardtJul 19, 2007
  45. Johannes SchindelinJul 19, 2007
  46. David KastrupJul 19, 2007
  47. Brian GernhardtJul 19, 2007
  48. Johannes SchindelinJul 19, 2007
  49. David KastrupJul 19, 2007
  50. Matthieu MoyJul 19, 2007
  51. David KastrupJul 19, 2007
  52. David KastrupJul 19, 2007
  53. David KastrupJul 19, 2007
  54. David KastrupJul 21, 2007
  55. Linus TorvaldsJul 21, 2007
  56. Linus TorvaldsJul 21, 2007
  57. David KastrupJul 21, 2007
  58. Linus TorvaldsJul 21, 2007
  59. David KastrupJul 21, 2007
  60. Simon 'corecode' SchubertJul 21, 2007
  61. David KastrupJul 21, 2007
  62. Linus TorvaldsJul 21, 2007
  63. David KastrupJul 22, 2007
  64. Linus TorvaldsJul 22, 2007
  65. David KastrupJul 22, 2007
  66. Linus TorvaldsJul 22, 2007
  67. David KastrupJul 22, 2007
  68. Linus TorvaldsJul 22, 2007
  69. David KastrupJul 22, 2007
  70. david@lang.hmJul 22, 2007
  71. David KastrupJul 22, 2007
  72. Linus TorvaldsJul 22, 2007
  73. David KastrupJul 22, 2007
  74. Linus TorvaldsJul 22, 2007
  75. Linus TorvaldsJul 22, 2007
  76. David KastrupJul 22, 2007
  77. Jakub NarebskiJul 22, 2007
  78. David KastrupJul 22, 2007
  79. Jakub NarebskiJul 22, 2007
  80. David KastrupJul 22, 2007
  81. Jakub NarebskiJul 22, 2007
  82. David KastrupJul 22, 2007
  83. David KastrupJul 23, 2007
  84. David KastrupJul 23, 2007
  85. David KastrupJul 22, 2007
  86. Brian GernhardtJul 22, 2007
  87. David KastrupJul 28, 2007
  88. David KastrupJul 18, 2007
  89. Matthieu MoyJul 18, 2007
  90. David KastrupJul 18, 2007
  91. Shawn O. PearceJul 18, 2007
  92. Junio C HamanoJul 18, 2007
  93. David KastrupJul 18, 2007
  94. Wincent ColaiutaJul 18, 2007
  95. Junio C HamanoJul 18, 2007
  96. Johan HerlandJul 20, 2007
  97. David KastrupJul 20, 2007
  98. Johan HerlandJul 20, 2007
  99. David KastrupJul 20, 2007
  100. Johan HerlandJul 20, 2007
  101. David KastrupJul 22, 2007
  102. Robin RosenbergJul 26, 2007
  103. David KastrupJul 27, 2007
  104. Johannes SchindelinJul 18, 2007
  105. Matthieu MoyJul 18, 2007
  106. David KastrupJul 18, 2007
  107. Junio C HamanoJul 18, 2007
  108. Brian GernhardtJul 19, 2007
  109. David KastrupJul 19, 2007
  110. Brian GernhardtJul 19, 2007
  111. Junio C HamanoJul 20, 2007
  112. Linus TorvaldsJul 20, 2007
  113. Linus TorvaldsJul 20, 2007
  114. Junio C HamanoJul 20, 2007
  115. Linus TorvaldsJul 20, 2007
  116. David KastrupJul 20, 2007
  117. David KastrupJul 20, 2007
  118. Linus TorvaldsJul 20, 2007
  119. David KastrupJul 20, 2007
  120. Simon 'corecode' SchubertJul 20, 2007
  121. David KastrupJul 20, 2007
  122. Junio C HamanoJul 20, 2007
  123. David KastrupJul 20, 2007
  124. Linus TorvaldsJul 20, 2007
  125. Johan HerlandJul 20, 2007
  126. Linus TorvaldsJul 20, 2007
  127. Julian PhillipsJul 20, 2007
  128. Linus TorvaldsJul 21, 2007
  129. David KastrupJul 21, 2007
  130. David KastrupJul 21, 2007
  131. David KastrupJul 20, 2007
  132. Olivier GalibertJul 20, 2007
  133. Johan HerlandJul 20, 2007
  134. David KastrupJul 20, 2007
  135. David KastrupJul 21, 2007
  136. David KastrupJul 22, 2007
  137. NixJul 24, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.