git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Empty directories...

From
David Kastrup <dak@gnu.org>
Date
Jul 22, 2007, 21:35 UTC
Message-ID
<85sl7gf7g1.fsf@lola.goethe.zz>
In-Reply-To
<7vhco28aoq.fsf@assigned-by-dhcp.cox.net>
Coming full circle...
Junio C Hamano <gitster@pobox.com> writes:
Show 14 quoted lines
> The right approach to take probably would be to allow entries of
> mode 040000 in the index.  Traditionally, we allowed only 100644
> (blobs as regular files) and 120000 (blobs as symlinks).  We
> recently added 160000 (commit from outer space, aka subproject).
>
> And we do that for all directories, not just empty ones.  So if
> you have fileA, empty/, sub/fileB tracked, your index would
> probably have these four entries, immediately after read-tree
> of an existing tree object:
>
> 	100644 15db6f1f27ef7a... 0	fileA
> 	040000 4b825dc642cb6e... 0	empty
> 	040000 e125e11d3b63e3... 0	sub
> 	100644 52054201c2a872... 0	sub/fileB
This would be very much what I am proposing now, except that instead
of 040000 we would have 040755 usually, so that when the index makes
it into the repository where 040000 already has a meaning (a
disappear-when-empty tree) we get the right information.  Also note
that the above comes about when doing
git-add *
but not when doing
git-add fileA empty sub/fileB (in the latter case, the entry for sub
                               would be missing)
Show 13 quoted lines
> If you add sub/fileC, with "update-index" (and "add"), you
> invalidate the SHA-1 object name you stored for "sub" (because
> there is no point recomputing the tree object until you know you
> need a subtree for "sub" part, which does not happen until the
> next "write-tree"), and end up with something like:
>
> 	100644 15db6f1f27ef7a... 0	fileA
> 	040000 4b825dc642cb6e... 0	empty
> 	040000 00000000000000... 0	sub
> 	100644 52054201c2a872... 0	sub/fileB
> 	100644 705bf16c546f32... 0	sub/fileC
>
> These "missing" SHA-1 would need to be recomputed on-demand.

Ah, ok. Does it even make sense to compute the SHA-1 values in the index in advance? What would they be useful for?

Show 5 quoted lines
> We have had necessary infrastructure to do this "keeping
> untouched tree object names in the index" for quite some time,
> but it is not a part of the index proper (it is stored in an
> extension section in the index file, to keep the index
> compatible with older versions of git).
What is the application for which this is being used?
Show 9 quoted lines
> Having made it sound so easy, here are the issues I would expect
> to be nontrivial (but probably not rocket surgery either).
>
>  * unpack-trees, which is the workhorse for twoway merge (aka
>    "switching branches") and threeway merge, has a convoluted
>    logic to avoid D/F conflicts; it can probably be cleaned up
>    once we do the above conversion so that the index starts
>    saying "Hey, I have a directory here" more explicitly.  The
>    end result would probably be a code easier to follow.

I am afraid that this is unlikely to happen, and that is because directory tracking remains optional at a fundamental level as long as we want to support the current behavior as an option. However, one could conceivably add 040000 entries (rather than 040755) for directories that have not been passed into tracking but are required by git, if this simplifies matters. But it sounds like something that might complicate working with several different git versions on the same index.

Show 6 quoted lines
>  * status, update-index --refresh, and diff-files cares about
>    the information cached in the index from the last time
>    lstat(2) is run on each entry.  What we should store there
>    for "tree" entries is very unclear to me, but probably we
>    should teach them to ignore the stat-matching logic for
>    these entries.

At the current point of time, git tracks just the u+x bit for normal files, and for directories, there is really nothing worth tracking as long as no attempt of restoring more mode bits is done. Modification times are probably a bit too risky to pay attention to.

>  * diff-index walks the index and a tree in parallel but does
>    not currently expect to see a tree object in the index.  It
>    needs to be taught to ignore these "tree" entries.
Or do something sensible when comparing.  Understood.
>  * merge-recursive and merge-index walk the index, coming up
>    with the merge results one path at a time.  They also need to
>    be taught to ignore these "tree" entries.
Same here.
Show 18 quoted lines
>  * diff-index and "read-tree -m" should be taught to take
>    advantage of the "tree" entries in the index.  For example,
>    if diff-index finds the "tree" entry in the index and the
>    subtree found from the tree object exactly match, it does not
>    even have to descend into the tree, which would be a huge
>    performance win (because you do not have to open the subtree
>    and its subtrees from the tree side; you already have read
>    everything on the index side, and still have to skip the
>    entries in the directory).  "read-tree -m" also should be
>    able to optimize two identical subtrees in the 2 or 3 trees
>    involved.
>
>    Even if we follow the "lazy invalidate" strategy to maintain
>    the "tree" entries in the normal codepath, we could have a
>    special operation that says "now update all the tree entries
>    by recomputing the tree object names as needed".  Perhaps we
>    might want to initiate such an operation before "read-tree
>    -m" automatically.
Over my head, but it would appear that it can safely left for later.
-- 
David Kastrup, Kriemhildstr. 15, 44793 Bochum
Previous: Johan HerlandNext: Robin Rosenberg
Message 101 of 137 in “Empty directories...”
  1. David KastrupJul 18, 2007
  2. Johannes SchindelinJul 18, 2007
  3. David KastrupJul 18, 2007
  4. Johannes SchindelinJul 18, 2007
  5. Linus TorvaldsJul 18, 2007
  6. Linus TorvaldsJul 18, 2007
  7. David KastrupJul 18, 2007
  8. Linus TorvaldsJul 18, 2007
  9. Matthieu MoyJul 18, 2007
  10. Linus TorvaldsJul 18, 2007
  11. David KastrupJul 18, 2007
  12. Linus TorvaldsJul 18, 2007
  13. David KastrupJul 18, 2007
  14. Re: Empty directories...Linus Torvalds, Jul 18, 2007
  15. Linus TorvaldsJul 18, 2007
  16. David KastrupJul 18, 2007
  17. Linus TorvaldsJul 19, 2007
  18. Junio C HamanoJul 19, 2007
  19. Shawn O. PearceJul 19, 2007
  20. David KastrupJul 19, 2007
  21. Geoff RussellJul 19, 2007
  22. Shawn O. PearceJul 19, 2007
  23. Matthieu MoyJul 19, 2007
  24. Tomash BrechkoJul 19, 2007
  25. David KastrupJul 19, 2007
  26. Tomash BrechkoJul 19, 2007
  27. David KastrupJul 19, 2007
  28. NixJul 23, 2007
  29. David KastrupJul 23, 2007
  30. NixJul 23, 2007
  31. NixJul 23, 2007
  32. Jakub NarebskiJul 23, 2007
  33. NixJul 25, 2007
  34. David KastrupJul 23, 2007
  35. Linus TorvaldsJul 23, 2007
  36. NixJul 23, 2007
  37. Linus TorvaldsJul 23, 2007
  38. David KastrupJul 19, 2007
  39. David KastrupJul 19, 2007
  40. Johannes SchindelinJul 19, 2007
  41. David KastrupJul 19, 2007
  42. Brian GernhardtJul 19, 2007
  43. Johannes SchindelinJul 19, 2007
  44. Brian GernhardtJul 19, 2007
  45. Johannes SchindelinJul 19, 2007
  46. David KastrupJul 19, 2007
  47. Brian GernhardtJul 19, 2007
  48. Johannes SchindelinJul 19, 2007
  49. David KastrupJul 19, 2007
  50. Matthieu MoyJul 19, 2007
  51. David KastrupJul 19, 2007
  52. David KastrupJul 19, 2007
  53. David KastrupJul 19, 2007
  54. David KastrupJul 21, 2007
  55. Linus TorvaldsJul 21, 2007
  56. Linus TorvaldsJul 21, 2007
  57. David KastrupJul 21, 2007
  58. Linus TorvaldsJul 21, 2007
  59. David KastrupJul 21, 2007
  60. Simon 'corecode' SchubertJul 21, 2007
  61. David KastrupJul 21, 2007
  62. Linus TorvaldsJul 21, 2007
  63. David KastrupJul 22, 2007
  64. Linus TorvaldsJul 22, 2007
  65. David KastrupJul 22, 2007
  66. Linus TorvaldsJul 22, 2007
  67. David KastrupJul 22, 2007
  68. Linus TorvaldsJul 22, 2007
  69. David KastrupJul 22, 2007
  70. david@lang.hmJul 22, 2007
  71. David KastrupJul 22, 2007
  72. Linus TorvaldsJul 22, 2007
  73. David KastrupJul 22, 2007
  74. Linus TorvaldsJul 22, 2007
  75. Linus TorvaldsJul 22, 2007
  76. David KastrupJul 22, 2007
  77. Jakub NarebskiJul 22, 2007
  78. David KastrupJul 22, 2007
  79. Jakub NarebskiJul 22, 2007
  80. David KastrupJul 22, 2007
  81. Jakub NarebskiJul 22, 2007
  82. David KastrupJul 22, 2007
  83. David KastrupJul 23, 2007
  84. David KastrupJul 23, 2007
  85. David KastrupJul 22, 2007
  86. Brian GernhardtJul 22, 2007
  87. David KastrupJul 28, 2007
  88. David KastrupJul 18, 2007
  89. Matthieu MoyJul 18, 2007
  90. David KastrupJul 18, 2007
  91. Shawn O. PearceJul 18, 2007
  92. Junio C HamanoJul 18, 2007
  93. David KastrupJul 18, 2007
  94. Wincent ColaiutaJul 18, 2007
  95. Junio C HamanoJul 18, 2007
  96. Johan HerlandJul 20, 2007
  97. David KastrupJul 20, 2007
  98. Johan HerlandJul 20, 2007
  99. David KastrupJul 20, 2007
  100. Johan HerlandJul 20, 2007
  101. David KastrupJul 22, 2007
  102. Robin RosenbergJul 26, 2007
  103. David KastrupJul 27, 2007
  104. Johannes SchindelinJul 18, 2007
  105. Matthieu MoyJul 18, 2007
  106. David KastrupJul 18, 2007
  107. Junio C HamanoJul 18, 2007
  108. Brian GernhardtJul 19, 2007
  109. David KastrupJul 19, 2007
  110. Brian GernhardtJul 19, 2007
  111. Junio C HamanoJul 20, 2007
  112. Linus TorvaldsJul 20, 2007
  113. Linus TorvaldsJul 20, 2007
  114. Junio C HamanoJul 20, 2007
  115. Linus TorvaldsJul 20, 2007
  116. David KastrupJul 20, 2007
  117. David KastrupJul 20, 2007
  118. Linus TorvaldsJul 20, 2007
  119. David KastrupJul 20, 2007
  120. Simon 'corecode' SchubertJul 20, 2007
  121. David KastrupJul 20, 2007
  122. Junio C HamanoJul 20, 2007
  123. David KastrupJul 20, 2007
  124. Linus TorvaldsJul 20, 2007
  125. Johan HerlandJul 20, 2007
  126. Linus TorvaldsJul 20, 2007
  127. Julian PhillipsJul 20, 2007
  128. Linus TorvaldsJul 21, 2007
  129. David KastrupJul 21, 2007
  130. David KastrupJul 21, 2007
  131. David KastrupJul 20, 2007
  132. Olivier GalibertJul 20, 2007
  133. Johan HerlandJul 20, 2007
  134. David KastrupJul 20, 2007
  135. David KastrupJul 21, 2007
  136. David KastrupJul 22, 2007
  137. NixJul 24, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.