{"thread":{"id":"9086","subject":"Empty directories...","startedAt":"2007-07-18T00:13:11Z","lastAt":"2007-07-28T08:44:56Z","messageCount":137,"participants":["David Kastrup","Johannes Schindelin","Matthieu Moy","Junio C Hamano","Shawn O. Pearce","Wincent Colaiuta","Linus Torvalds","Geoff Russell","Tomash Brechko","Brian Gernhardt","Johan Herland","Simon 'corecode' Schubert","Olivier Galibert","Julian Phillips","Jakub Narebski","david@lang.hm","Nix","Robin Rosenberg"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"47685","messageId":"85lkdezi08.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":null,"subject":"Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-18T00:13:11Z","receivedAt":"2007-07-18T00:13:11Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"\nGIT(7) -- 03/05/2007\n\nNAME\n\tgit - the stupid content tracker\n\n\nWell, I use git for tracking contents.  That means, for example,\ninstallation trees for some application.  Let's take a typical TeXlive\ntree as an example.  Those trees contain, among other things,\ndirectories where new fonts/formats/whatever get placed as things run.\nQuite a few of them start out empty, but their permissions have to\ncorrespond to their purpose (for example, some are world-writable).\n\nI see little chance to get this achieved without doing something like\n\nfind -type d -empty -execdir touch {}/.git-this-is-empty +\n\nbefore every checkin and\n\nfind -name .git-this-is-empty -exec rm -- {} +\n\nafter every checkout.  Which is pretty stupid.\n\nAs some anecdotal stuff, I did something like\n\nmkdir test\ncd test\ngit-init\ntouch README\ngit-add README # another peeve: why is no empty reference point possible?\ngit-commit -a -m \"Initial branch\"\ngit checkout -b newbranch master\nunzip ../somearchive -d subdir\ngit add subdir\ngit commit -a -m \"Add subdir\"\ngit checkout -b newbranch2 master\n\nand expect to have a clean slate.  No such luck: without warning, all\nempty directories in the zip file are still remaining within subdir,\nwhich as a consequence has not been cleaned up.\n\nSo even if one is of the opinion that empty directories are not worth\nputting into the repository: if I check in an entire subdirectory\nhierarchy and then switch to a branch where this subdirectory is not\nexistent, I expect the subdirectory to be _gone_, and not have some\nlittering of empty directories lying around.\n\nAnd that git-diff can see nothing wrong with that does not really\nimprove things.\n\nSo if git is supposed to be a content tracker, I can't see a way\naround it actually being able to track content, and empty directories\n_are_ content.  It can't let them flying around with arbitrary\npermissions on them when I switch branches or tags.  And the\nworkaround using \"touch\" mentioned above is really awful to do\nmanually all the time.\n\nCould git technically track a file with a zero-length filename in\nempty directories if one tells it explicitly to include it, like with\ngit-add \\! -x \"\" subdir\nor has somebody a better idea or interface or rationale?  I understand\nthat there are use cases where one does not bother about empty\ndirectories, but for a _content_ tracker, not tracking directories\nbecause they are empty seems quite serious.\n\nOk, kill me.  This must likely be the most common FAQ/rant/whatever\nconcerning git.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"47686","messageId":"Pine.LNX.4.64.0707180135200.14781@racer.site","threadId":"9086","inReplyTo":"85lkdezi08.fsf@lola.goethe.zz","subject":"Re: Empty directories...","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-18T00:35:35Z","receivedAt":"2007-07-18T00:35:35Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 18 Jul 2007, David Kastrup wrote:\n\n> This must likely be the most common FAQ/rant/whatever concerning git.\n\nIf you had the idea already, I wonder why you did not find it.  It's not \nreally anything like hard to find:\n\nhttp://git.or.cz/gitwiki/GitFaq#head-1fbd4a018d45259c197b169e87dafce2a3c6b5f9\n\nCiao,\nDscho\n"},{"id":"47687","messageId":"vpqfy3m7dex.fsf@bauges.imag.fr","threadId":"9086","inReplyTo":"85lkdezi08.fsf@lola.goethe.zz","subject":"Re: Empty directories...","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2007-07-18T00:39:50Z","receivedAt":"2007-07-18T00:39:50Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"David Kastrup <dak@gnu.org> writes:\n\n> or has somebody a better idea or interface or rationale?  I understand\n> that there are use cases where one does not bother about empty\n> directories, but for a _content_ tracker, not tracking directories\n> because they are empty seems quite serious.\n\n,----[ http://www.spinics.net/lists/git/msg30730.html ]\n| From: Linus Torvalds <torvalds@xxxxxxxxxxxxxxxxxxxx>\n| \n| I wouldn't personally mind if somebody taught git to just track empty\n| directories too.\n| \n| There is no fundamental git database reason not to allow them: it's in\n| fact quite easy to create an empty tree object. The problems with\n| empty directories are in the *index*, and they shouldn't be\n| insurmountable.\n| \n| [...]\n`----\n\n-- \nMatthieu\n"},{"id":"47692","messageId":"7v8x9ea1rg.fsf@assigned-by-dhcp.cox.net","threadId":"9086","inReplyTo":"85lkdezi08.fsf@lola.goethe.zz","subject":"Re: Empty directories...","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-07-18T02:23:15Z","receivedAt":"2007-07-18T02:23:15Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"David Kastrup <dak@gnu.org> writes:\n\n> or has somebody a better idea or interface or rationale?  I understand\n> that there are use cases where one does not bother about empty\n> directories, but for a _content_ tracker, not tracking directories\n> because they are empty seems quite serious.\n\nNo objections as long as a patch is cleanly made without\nregression.  It's just nobody agreed that it is \"quite serious\"\nyet so far, and no fundamental reason against it.\n"},{"id":"47700","messageId":"85d4yqz24s.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"7v8x9ea1rg.fsf@assigned-by-dhcp.cox.net","subject":"Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-18T05:56:03Z","receivedAt":"2007-07-18T05:56:03Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> David Kastrup <dak@gnu.org> writes:\n>\n>> or has somebody a better idea or interface or rationale?  I understand\n>> that there are use cases where one does not bother about empty\n>> directories, but for a _content_ tracker, not tracking directories\n>> because they are empty seems quite serious.\n>\n> No objections as long as a patch is cleanly made without\n> regression.  It's just nobody agreed that it is \"quite serious\"\n> yet so far, and no fundamental reason against it.\n\nThanks.  It certainly is not serious for the Linux kernel source, but\nseems awkward for quite a few situations.  Anyway, what is your take\non the situation I described?\n\nThat creating some directory hierarchy (happening to contain empty\ndirectories) with some external program, adding and committing it,\nthen switching to a different branch (or maybe doing a git-reset\n--hard) leaves a skeleton of empty directories around?\n\nI find this almost worse than not being able to put them into the\nrepository: you can't get rid of them anymore either!\n\nI'd be tempted to propose that git should remove empty subdirectories\nwhen cleaning up a removed tree in the working directory, even though\nthat violates the principle to not delete anything it isn't tracking.\nBut since you can't get it to track the stuff in the first place...\n\nBut the real fix would be to track them.\n\nDoes some trick work possibly at checkin time, like putting an empty\nfile into every empty directory, adding to the index, then removing\nall empty files explicitly from the index and then checking in, or is\nthis hopeless to work around with from the user side without affecting\nthe repository itself?\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"47701","messageId":"858x9ez1li.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"Pine.LNX.4.64.0707180135200.14781@racer.site","subject":"Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-18T06:07:37Z","receivedAt":"2007-07-18T06:07:37Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> Hi,\n>\n> On Wed, 18 Jul 2007, David Kastrup wrote:\n>\n>> This must likely be the most common FAQ/rant/whatever concerning git.\n>\n> If you had the idea already, I wonder why you did not find it.  It's not \n> really anything like hard to find:\n>\n> http://git.or.cz/gitwiki/GitFaq#head-1fbd4a018d45259c197b169e87dafce2a3c6b5f9\n\nThe FAQ answer is weazeling on several accounts:\n\na) No, git only cares about files, or rather git tracks content and\n   empty directories have no content.\n\nIn the same manner as empty regular files have no contents, and git\ntracks those.  Existence and permissions are important.\n\nb) The problem is not just that empty directories don't get added into\nthe repository.  They also don't get removed again when switching to a\ndifferent checkout.  When git-diff returns zero, I expect a subsequent\ncheckout to not leave complete empty hierarchies around because git\ncan't delete any empty leaves which it chose not to track.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"47703","messageId":"85zm1uxmmw.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"vpqfy3m7dex.fsf@bauges.imag.fr","subject":"Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-18T06:16:07Z","receivedAt":"2007-07-18T06:16:07Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Matthieu Moy <Matthieu.Moy@imag.fr> writes:\n\n> David Kastrup <dak@gnu.org> writes:\n>\n>> or has somebody a better idea or interface or rationale?  I understand\n>> that there are use cases where one does not bother about empty\n>> directories, but for a _content_ tracker, not tracking directories\n>> because they are empty seems quite serious.\n>\n> ,----[ http://www.spinics.net/lists/git/msg30730.html ]\n> | From: Linus Torvalds <torvalds@xxxxxxxxxxxxxxxxxxxx>\n> | \n> | I wouldn't personally mind if somebody taught git to just track empty\n> | directories too.\n> | \n> | There is no fundamental git database reason not to allow them:\n> | it's in fact quite easy to create an empty tree object.\n> | The problems with empty directories are in the *index*, and they\n> | shouldn't be insurmountable.\n\nStop right here: does that mean that I can script some \"put empty\ndirectories into the last commit manually\" procedure bypassing the\nindex?\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"47704","messageId":"20070718063047.GA32566@spearce.org","threadId":"9086","inReplyTo":"85zm1uxmmw.fsf@lola.goethe.zz","subject":"Re: Empty directories...","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-07-18T06:30:47Z","receivedAt":"2007-07-18T06:30:47Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"David Kastrup <dak@gnu.org> wrote:\n> > ,----[ http://www.spinics.net/lists/git/msg30730.html ]\n> > | From: Linus Torvalds <torvalds@xxxxxxxxxxxxxxxxxxxx>\n> > | \n> > | I wouldn't personally mind if somebody taught git to just track empty\n> > | directories too.\n> > | \n> > | There is no fundamental git database reason not to allow them:\n> > | it's in fact quite easy to create an empty tree object.\n> > | The problems with empty directories are in the *index*, and they\n> > | shouldn't be insurmountable.\n> \n> Stop right here: does that mean that I can script some \"put empty\n> directories into the last commit manually\" procedure bypassing the\n> index?\n\nYes.  But when you read that tree into the index later (by say\nchecking out a branch that points to it) the empty directories\nwill not be created, as they have no files to cause their creation.\nCommitting changes on that branch will remove the empty directories.\n;-)\n\nOh, and the above question from you sounds like you think you can\nmodify the last commit to include new directories that weren't\nthere before.  You cannot do that without changing the tree SHA-1,\nwhich will cause the commit SHA-1 to change.  That in turns means you\nare not actually adding to the last commit but instead are creating\nan entirely different commit.  History in Git is always immutable.\n\n-- \nShawn.\n"},{"id":"47705","messageId":"BFBE8924-5F60-4D1F-9260-29545BEA0790@wincent.com","threadId":"9086","inReplyTo":"85d4yqz24s.fsf@lola.goethe.zz","subject":"Re: Empty directories...","fromName":"Wincent Colaiuta","fromEmail":"win@wincent.com","sentAt":"2007-07-18T06:34:27Z","receivedAt":"2007-07-18T06:34:27Z","isPatch":false,"sender":{"key":"greg@hurrell.net","avatar":"https://avatars.githubusercontent.com/u/7074?v=4"},"body":"El 18/7/2007, a las 7:56, David Kastrup escribió:\n\n> That creating some directory hierarchy (happening to contain empty\n> directories) with some external program, adding and committing it,\n> then switching to a different branch (or maybe doing a git-reset\n> --hard) leaves a skeleton of empty directories around?\n>\n> I find this almost worse than not being able to put them into the\n> repository: you can't get rid of them anymore either!\n>\n> I'd be tempted to propose that git should remove empty subdirectories\n> when cleaning up a removed tree in the working directory, even though\n> that violates the principle to not delete anything it isn't tracking.\n> But since you can't get it to track the stuff in the first place...\n>\n> But the real fix would be to track them.\n\nAlthough I haven't yet been \"bitten\" by this issue I understand where  \nyou're coming from. This could confuse users and appear inconsistent  \nto them (seeing as empty *files* can be tracked). I think it's  \nprobably worth tackling for that reason alone, but it will have the  \nadditional benefit of enabling other workflows like the one you  \ndescribe (\"installation trees for some application\").\n\n> Does some trick work possibly at checkin time, like putting an empty\n> file into every empty directory, adding to the index, then removing\n> all empty files explicitly from the index and then checking in, or is\n> this hopeless to work around with from the user side without affecting\n> the repository itself?\n\nI wouldn't recommend any \"tricks\" here. I think the real solution is  \nto allow the tracking of empty trees; everything else seems like a  \nkludge. And then, as you've noted already that will allow Git to  \nhandle the \"skeleton of empty directories\" left behind problem that  \nyou describe.\n\nCheers,\nWincent\n"},{"id":"47706","messageId":"7vhco28aoq.fsf@assigned-by-dhcp.cox.net","threadId":"9086","inReplyTo":"85d4yqz24s.fsf@lola.goethe.zz","subject":"Re: Empty directories...","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-07-18T06:53:25Z","receivedAt":"2007-07-18T06:53:25Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"David Kastrup <dak@gnu.org> writes:\n\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n>> No objections as long as a patch is cleanly made without\n>> regression.  It's just nobody agreed that it is \"quite serious\"\n>> yet so far, and no fundamental reason against it.\n>\n> Thanks.  It certainly is not serious for the Linux kernel source, but\n> seems awkward for quite a few situations.  Anyway, what is your take\n> on the situation I described?\n\nDidn't I say I do not have an objection for somebody who wants\nto track empty directories, already?  I probably would not do\nthat myself but I do not see a reason to forbid it, either.\n\nThe right approach to take probably would be to allow entries of\nmode 040000 in the index.  Traditionally, we allowed only 100644\n(blobs as regular files) and 120000 (blobs as symlinks).  We\nrecently added 160000 (commit from outer space, aka subproject).\n\nAnd we do that for all directories, not just empty ones.  So if\nyou have fileA, empty/, sub/fileB tracked, your index would\nprobably have these four entries, immediately after read-tree\nof an existing tree object:\n\n\t100644 15db6f1f27ef7a... 0\tfileA\n\t040000 4b825dc642cb6e... 0\tempty\n\t040000 e125e11d3b63e3... 0\tsub\n\t100644 52054201c2a872... 0\tsub/fileB\n\nMaking sure that empty/ directory exists in the working tree is\nprobably done in entry.c; we have been touching that area in an\nunrelated thread in the past few days.\n\nIf you add sub/fileC, with \"update-index\" (and \"add\"), you\ninvalidate the SHA-1 object name you stored for \"sub\" (because\nthere is no point recomputing the tree object until you know you\nneed a subtree for \"sub\" part, which does not happen until the\nnext \"write-tree\"), and end up with something like:\n\n\t100644 15db6f1f27ef7a... 0\tfileA\n\t040000 4b825dc642cb6e... 0\tempty\n\t040000 00000000000000... 0\tsub\n\t100644 52054201c2a872... 0\tsub/fileB\n\t100644 705bf16c546f32... 0\tsub/fileC\n\nThese \"missing\" SHA-1 would need to be recomputed on-demand.\n\nWe have had necessary infrastructure to do this \"keeping\nuntouched tree object names in the index\" for quite some time,\nbut it is not a part of the index proper (it is stored in an\nextension section in the index file, to keep the index\ncompatible with older versions of git).\n\nHaving made it sound so easy, here are the issues I would expect\nto be nontrivial (but probably not rocket surgery either).\n\n * unpack-trees, which is the workhorse for twoway merge (aka\n   \"switching branches\") and threeway merge, has a convoluted\n   logic to avoid D/F conflicts; it can probably be cleaned up\n   once we do the above conversion so that the index starts\n   saying \"Hey, I have a directory here\" more explicitly.  The\n   end result would probably be a code easier to follow.\n\n * status, update-index --refresh, and diff-files cares about\n   the information cached in the index from the last time\n   lstat(2) is run on each entry.  What we should store there\n   for \"tree\" entries is very unclear to me, but probably we\n   should teach them to ignore the stat-matching logic for\n   these entries.\n\n * diff-index walks the index and a tree in parallel but does\n   not currently expect to see a tree object in the index.  It\n   needs to be taught to ignore these \"tree\" entries.\n\n * merge-recursive and merge-index walk the index, coming up\n   with the merge results one path at a time.  They also need to\n   be taught to ignore these \"tree\" entries.\n\n * diff-index and \"read-tree -m\" should be taught to take\n   advantage of the \"tree\" entries in the index.  For example,\n   if diff-index finds the \"tree\" entry in the index and the\n   subtree found from the tree object exactly match, it does not\n   even have to descend into the tree, which would be a huge\n   performance win (because you do not have to open the subtree\n   and its subtrees from the tree side; you already have read\n   everything on the index side, and still have to skip the\n   entries in the directory).  \"read-tree -m\" also should be\n   able to optimize two identical subtrees in the 2 or 3 trees\n   involved.\n\n   Even if we follow the \"lazy invalidate\" strategy to maintain\n   the \"tree\" entries in the normal codepath, we could have a\n   special operation that says \"now update all the tree entries\n   by recomputing the tree object names as needed\".  Perhaps we\n   might want to initiate such an operation before \"read-tree\n   -m\" automatically.\n"},{"id":"47720","messageId":"Pine.LNX.4.64.0707181121520.14781@racer.site","threadId":"9086","inReplyTo":"858x9ez1li.fsf@lola.goethe.zz","subject":"Re: Empty directories...","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-18T10:26:15Z","receivedAt":"2007-07-18T10:26:15Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 18 Jul 2007, David Kastrup wrote:\n\n> The FAQ answer is weazeling on several accounts:\n> \n> a) No, git only cares about files, or rather git tracks content and\n>    empty directories have no content.\n> \n> In the same manner as empty regular files have no contents, and git\n> tracks those.  Existence and permissions are important.\n\nWe do not track permissions of directories at all.  This is because Git is \nprimarily meant to track source code, and most \"permissions\" (i.e. \nrestrictions) do not make any sense there.\n\n> b) The problem is not just that empty directories don't get added into\n> the repository.  They also don't get removed again when switching to a\n> different checkout.  When git-diff returns zero, I expect a subsequent\n> checkout to not leave complete empty hierarchies around because git\n> can't delete any empty leaves which it chose not to track.\n\nI _like_ the behaviour that Git does not remove a directory it added, when \nI put some untracked file into it.  And switching back to that branch, Git \nhas no problems, because it sees that the directory is already there.  In \ncase of a file, it would complain, and rightfully so.\n\nSee the fundamental difference between a file and a directory now?  I \nthink it boils down to \"an empty directory has _no_ contents, but an empty \nfile has an _empty_ content\".\n\nCiao,\nDscho\n"},{"id":"47725","messageId":"Pine.LNX.4.64.0707181218090.14781@racer.site","threadId":"9086","inReplyTo":"86tzs2m1h7.fsf@lola.quinscape.zz","subject":"Re: Empty directories...","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-18T11:24:46Z","receivedAt":"2007-07-18T11:24:46Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 18 Jul 2007, David Kastrup wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> > On Wed, 18 Jul 2007, David Kastrup wrote:\n> >\n> >> The FAQ answer is weazeling on several accounts:\n> >> \n> >> a) No, git only cares about files, or rather git tracks content and\n> >>    empty directories have no content.\n> >> \n> >> In the same manner as empty regular files have no contents, and git\n> >> tracks those.  Existence and permissions are important.\n> >\n> > We do not track permissions of directories at all.\n> \n> Ok, this seems like something that should be done as well, even if we\n> can stipulate at first that a directory should have rwx for the user\n> in question if you hope to track it.\n\nNo, no, no.  It should not be tracked.  It is the responsibility of the \n_user_ to set it to something sane, be that by a umask or by sticky \ngroups, or by setting the permissions of the parent directory.\n\nIt is _nothing_ we want to put into the repository.  That is the _wrong_ \nplace to put it.\n\n> > This is because Git is primarily meant to track source code,\n> \n> Tell that to the man page.  It declares git to be \"a content tracker\" \n> right at the front.\n\nWhy don't you?  I have no problems with the title.\n\n> > and most \"permissions\" (i.e.  restrictions) do not make any sense\n> > there.\n> \n> So why are permissions for files being tracked, then?\n\nThis question is invalid.  Git only tracks the _executable_ bit.  And \nagain, it is the users' responsibility, by setting the umask, to have the \nappropriate bits set for group and others.\n\n> >> b) The problem is not just that empty directories don't get added \n> >> into the repository.  They also don't get removed again when \n> >> switching to a different checkout.  When git-diff returns zero, I \n> >> expect a subsequent checkout to not leave complete empty hierarchies \n> >> around because git can't delete any empty leaves which it chose not \n> >> to track.\n> >\n> > I _like_ the behaviour that Git does not remove a directory it\n> > added, when I put some untracked file into it.\n> \n> But it does not remove a directory it _refused_ to add when there were\n> no files at all in it ever.  You probably have not read the problem\n> description carefully.\n\nI have.  But that does not apply here, because I used the term \"to add a \ndirectory\" in the sense of \"mkdir\".\n\n> > And switching back to that branch, Git has no problems, because it \n> > sees that the directory is already there.  In case of a file, it would \n> > complain, and rightfully so.\n> \n> And if you switch to a branch where the directory it did not remove now \n> is a file?\n\nGit already throws an error, and rightfully so.  I am pleased by the \ncurrent behaviour.\n\n> > See the fundamental difference between a file and a directory now?\n> \n> Condescension is not really solving a problem.\n\nHey, I only tried to help clarify things.\n\nBut since I seem to be unable to, I'll end my efforts with this \nsuggestion:\n\nIf you want to track empty directories, the best thing would be to\n\n- teach git-add to automatically create an empty .gitignore (and error out \n  if that already exists), and\n\n- teach git-archive to not put .gitignore files into the output by default \n  (but the directories).  This might be a sensible change regardless if \n  you want to add empty directories to the repository or not.\n\nCiao,\nDscho\n"},{"id":"47726","messageId":"vpqejj6c52u.fsf@bauges.imag.fr","threadId":"9086","inReplyTo":"Pine.LNX.4.64.0707181218090.14781@racer.site","subject":"Re: Empty directories...","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2007-07-18T11:40:57Z","receivedAt":"2007-07-18T11:40:57Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n>> > We do not track permissions of directories at all.\n>> \n>> Ok, this seems like something that should be done as well, even if we\n>> can stipulate at first that a directory should have rwx for the user\n>> in question if you hope to track it.\n>\n> No, no, no.  It should not be tracked.  It is the responsibility of the \n> _user_ to set it to something sane, be that by a umask or by sticky \n> groups, or by setting the permissions of the parent directory.\n>\n> It is _nothing_ we want to put into the repository.  That is the _wrong_ \n> place to put it.\n\nI'm not sure it's wrong to be able to track permissions, but it's\ndefinitely wrong to track them by default.\n\nGNU Arch had some permission tracking, and I got hit by it several\ntimes. You have several things you might have wanted to track:\n\n* read/write for the user. But I can't imagine a case where you\n  wouldn't want to be able to read and write your own files.\n\n* permissions for group. But that doesn't make any sense when several\n  persons work on the same project, and don't share the same\n  /etc/group.\n\n* permissions for others. But that, again, doesn't make sense when\n  several persons work on the same project with different setups. I\n  sometimes work at home, where I'm basically the only user, I don't\n  care at all about permissions for others. At work, it's totally\n  different, since it's a big NFS shared by all the lab. And I might\n  very well disclose my work to the rest of the lab, and work with\n  someone who do not want to do so.\n\n* Execute bit. This one is relevant. Indeed, it's more a kind of\n  metadata than really a permission (you can still execute the file\n  with /lib/ld-linux.so.2 /path/to/file or such kind of things).\n\nUsing GNU Arch, I got the cases in real life of a project in which\nsome files had group read permission, some other not, because they\nwere created by developers having different umask. Worse than this, I\ngot some group-writable files in my $HOME without noticing it, which\nis basically a security hole.\n\n-- \nMatthieu\n"},{"id":"47728","messageId":"861wf5nc6e.fsf@lola.quinscape.zz","threadId":"9086","inReplyTo":"vpqejj6c52u.fsf@bauges.imag.fr","subject":"Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-18T12:12:09Z","receivedAt":"2007-07-18T12:12:09Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Matthieu Moy <Matthieu.Moy@imag.fr> writes:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n>\n>>> > We do not track permissions of directories at all.\n>>> \n>>> Ok, this seems like something that should be done as well, even if we\n>>> can stipulate at first that a directory should have rwx for the user\n>>> in question if you hope to track it.\n>>\n>> No, no, no.  It should not be tracked.  It is the responsibility of the \n>> _user_ to set it to something sane, be that by a umask or by sticky \n>> groups, or by setting the permissions of the parent directory.\n>>\n>> It is _nothing_ we want to put into the repository.  That is the _wrong_ \n>> place to put it.\n>\n> I'm not sure it's wrong to be able to track permissions, but it's\n> definitely wrong to track them by default.\n\nI am not sure about \"definitely\", but there certainly are applications\nwhere it is appropriate.\n\n> * Execute bit. This one is relevant. Indeed, it's more a kind of\n>   metadata than really a permission (you can still execute the file\n>   with /lib/ld-linux.so.2 /path/to/file or such kind of things).\n\nPlease spare us the sophistry.  Probably the most flexible approach\nwould be to be able to specify a checkout umask, defaulting to 700\n(the other bits are then filled in from the normal user umask).  For\narchival purposes, one would then set it to 777 instead.\n\nThere is the question how to deal with checkins.  While there is no\nharm in checking in the full permissions in case one would need them,\nit would likely be a nuisance to track the individual contributor's\nsettings.\n\n-- \nDavid Kastrup\n"},{"id":"47747","messageId":"alpine.LFD.0.999.0707180912430.27353@woody.linux-foundation.org","threadId":"9086","inReplyTo":"858x9ez1li.fsf@lola.goethe.zz","subject":"Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-18T16:23:46Z","receivedAt":"2007-07-18T16:23:46Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 18 Jul 2007, David Kastrup wrote:\n> \n> In the same manner as empty regular files have no contents, and git\n> tracks those.  Existence and permissions are important.\n\nYes, but directories really are different.\n\nFirst off, git wouldn't track the permissions anyway (git tracks execute \nbits, but for directories that _has_ to be set or git couldn't use them \nitself, so that's not going to happen).\n\nSecond, and much more important, the directories will exist or not \n*regardless* of what git does.\n\n> b) The problem is not just that empty directories don't get added into\n> the repository.  They also don't get removed again when switching to a\n> different checkout.\n\nBzzt. Wrong.\n\nWe *do* remove directories when all files under them go away.\n\nHOWEVER (and this is where one of the reasons for not tracking them comes \nin):\n\n   ** YOU CANNOT REMOVE A DIRECTORY IF IT HAS SOME UNTRACKED CONTENTS **\n\nThink about that for five seconds, then think about it some more. Ponder \nit.\n\nSo the fact is, git *already* does ass good of a job as it could possibly \ndo wrt directories that go away: it tries to remove them if all the files \nthat are tracked in it have gone away.\n\nBut that leaves a very common case, namely switching to another branch \nwithout those files, and the directory still having stale object files etc \nbuild crud in it.\n\nA SCM *must*not* just remove that directory. It would be horrible. The \nfact that it has untracked files in it does not make those untracked files \n\"unimportant\". Maybe you feel that way about object files, but what about \ntracking some important parts of your home directory - does the fact that \nyou don't necessarily track *all* of it mean that the rest is totally \nunimportant adn that git should just remove it? HELL NO!\n\nSo directories really _are_ problematic. You cannot (and should not) track \nthem the same way as you track a file.\n\nAnd the difference is very fundamental indeed: when you track a regular \nfile, you track *all* of its content. But when you track a directory, \nyou don't track it's content *at*all*.\n\nThink about that, and then think about the fact that git is defined as a \n\"content tracker\", and it's not \"weasely\" at all to say that you don't \ntrack directories.\n\nSo your argument is totally bogus. When you track an empty file, you very \nmuch track the *content* of that file, and \"empty\" just happens to be a \nvery valid content.\n\nBut when you track a \"directory\", you don't actually track its content at \nall, you track it's *existence*, which is a very very very different \nthing. I hope you understand from the above what is so different.\n\n(A true \"directory content\" tracker by definition would have to track \nevery single file under that directory. You can claim that for the case of \nan empty directory the \"existence tracking\" is 100% equivalent with \n\"content tracking\", but that's simply not true. It becomes non-true the \nmoment there are any files at all inside that directory, and be honest \nnow: the only _point_ of an empty directory is that you expect it to \npotentially get files under it).\n\nSo \"existence\" != \"content\". Git very much does not track \"existence\" of \nfiles, it tracks the total content of them too.\n\n\t\t\tLinus\n"},{"id":"47750","messageId":"alpine.LFD.0.999.0707180925030.27353@woody.linux-foundation.org","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707180912430.27353@woody.linux-foundation.org","subject":"Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-18T16:33:43Z","receivedAt":"2007-07-18T16:33:43Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 18 Jul 2007, Linus Torvalds wrote:\n> \n> So \"existence\" != \"content\". Git very much does not track \"existence\" of \n> files, it tracks the total content of them too.\n\nBtw, don't get me wrong: I think that in order to be better at tracking \nother SCM's idiotic choices, we could (and I foresee that we eventually \nhave to) try to track empty directories as a special case too.\n\nSo I'm not _against_ the notion of tracking empty directories, and I would \nwelcome patches that do so. As I mentioned in some earlier thread when \nthis came up a few weeks ago, I actually suspect that the \"subproject\" \nsupport probably ended up making it easier, because in many ways an \"empty \ndirectory\" is very close to a \"anonymous subproject\" from a low-level \nplumbing standpoint (even if it is *not* so from a high-level standpoint).\n\nSo I suspect that adding support for empty directories ends up being about \njust slightly extending the places that now have subproject support to \nknow about a new situation.\n\nBut I do want to point out that \"tracking a directory\" is not at all the \nsame thing as \"tracking a file\", no matter how much you try to argue \notherwise. The semantics are totally different, and it all boils down to \nthe fact that when you track a file, you are always talking about the \n*full* content of the file, while tracking a directory is always about \ntracking just a *subset* of the contents of the directory.\n\nOf course, with directories, there's the trivial case where the subset \nhappens to be everything, but that is neither the common nor the \ninteresting case. All the interesting and complex cases happen exactly \nwhen the directory has untracked files in it, and at that point \n\n - you really aren't tracking \"contents\" any more\n - you can no longer recreate the directory from the data you have (so you \n   cannot remove it on branch switches etc)\n - ergo: you're not a content tracker any more, you're a \"container\" \n   tracker.\n\nAnd really, the \"nontracked files in a directory\" is the *default* thing, \nnot some really unusual thing that we could disallow.\n\nBut I'm not against adding support for \"container tracking\". I just want \npeople to understand that it's something totally different from what we do \nnow. It's much more like subproject support than tracking files.\n\n\t\tLinus\n"},{"id":"47752","messageId":"vpq4pk1vf7q.fsf@bauges.imag.fr","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707180912430.27353@woody.linux-foundation.org","subject":"Re: Empty directories...","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2007-07-18T16:39:21Z","receivedAt":"2007-07-18T16:39:21Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n>> b) The problem is not just that empty directories don't get added into\n>> the repository.  They also don't get removed again when switching to a\n>> different checkout.\n>\n> Bzzt. Wrong.\n>\n> We *do* remove directories when all files under them go away.\n>\n> HOWEVER (and this is where one of the reasons for not tracking them comes \n> in):\n>\n>    ** YOU CANNOT REMOVE A DIRECTORY IF IT HAS SOME UNTRACKED CONTENTS **\n\nI believe David's point was different.\n\nIf you checkout a branch, create an empty directory in this branch\n(probably a placeholder, either for future versionned files, or for\ngenerated files), you cannot tell git \"this empty directory is in this\nbranch, but not in other ones\" without adding a file in it.\n\nSo, doing \"git-checkout anotherbranch\", this empty directory doesn't\ngo away. It's just unversionned in both branches, git won't touch it.\n\n-- \nMatthieu\n"},{"id":"47754","messageId":"alpine.LFD.0.999.0707181004330.27353@woody.linux-foundation.org","threadId":"9086","inReplyTo":"vpq4pk1vf7q.fsf@bauges.imag.fr","subject":"Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-18T17:06:13Z","receivedAt":"2007-07-18T17:06:13Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 18 Jul 2007, Matthieu Moy wrote:\n> \n> If you checkout a branch, create an empty directory in this branch\n> (probably a placeholder, either for future versionned files, or for\n> generated files), you cannot tell git \"this empty directory is in this\n> branch, but not in other ones\" without adding a file in it.\n\nRight. Which is the suggested setup: add an empty \".gitignore\" file to the \ndirectory, and you're done. It now acts \"as if\" git tracked the directory \n(git will remove the directory when switching branches), but without the \nlie that we really track any directory contents.\n\n\t\t\tLinus\n"},{"id":"47755","messageId":"85vechy5su.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707180912430.27353@woody.linux-foundation.org","subject":"Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-18T17:34:25Z","receivedAt":"2007-07-18T17:34:25Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Wed, 18 Jul 2007, David Kastrup wrote:\n>\n>> b) The problem is not just that empty directories don't get added\n>> into the repository.  They also don't get removed again when\n>> switching to a different checkout.\n>\n> Bzzt. Wrong.\n>\n> We *do* remove directories when all files under them go away.\n\nBut empty directories which were empty to start with don't go away\nsince they are not tracked.  And that means that their parents don't\ngo away.\n\nGit will remove directories which _had_ git-tracked content prior to\nthe checkout.  But it will not register empty directories created\noutside of git, and consequently will not remove them.\n\n> HOWEVER (and this is where one of the reasons for not tracking them\n> comes in):\n>\n>    ** YOU CANNOT REMOVE A DIRECTORY IF IT HAS SOME UNTRACKED CONTENTS **\n>\n> Think about that for five seconds, then think about it some\n> more. Ponder it.\n\nLinus, condescension is all very nice, but I already told you: I had a\ndirectory hierarchy created outside of git's control (every file comes\ninto being first outside of git).  This hierarchy contained empty\ndirectories.  The while hierarchy was committed into git.  git\nsilently skipped registering empty directories.  Then a different\nversion got checked out which did not contain the directory hierarchy\nin question.  And git left the (unregistered) empty directories in, as\nwell as all their parent directories.\n\nAnd that is just plain wrong.\n\n> So the fact is, git *already* does ass good of a job as it could\n> possibly do wrt directories that go away: it tries to remove them if\n> all the files that are tracked in it have gone away.\n\nBut I told git to track the whole directory tree recursively.  There\nwere no uncommitted files it complained about.  It is not reasonable\nthat it is afterwards unable to remove this when I checkout some other\ntag.\n\n> A SCM *must*not* just remove that directory. It would be\n> horrible. The fact that it has untracked files in it does not make\n> those untracked files \"unimportant\".\n\nSure.  But that it refuses to track the files makes the total behavior\nan annoyance.  I don't complain _how_ git handles not being able to\ntrack empty directories.  I complain about it not being able to track\nthem in the first place.  The consequences are hideous.\n\n> Maybe you feel that way about object files, but what about tracking\n> some important parts of your home directory - does the fact that you\n> don't necessarily track *all* of it mean that the rest is totally\n> unimportant adn that git should just remove it? HELL NO!\n\nWhen I tell it to track it, it should not refuse.  Even if it is\nempty.  Because if it _stayed_ empty, git can then remove it (and\npossibly the parents) when I checkout something else.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"47756","messageId":"85r6n5y5m3.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707180925030.27353@woody.linux-foundation.org","subject":"Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-18T17:38:28Z","receivedAt":"2007-07-18T17:38:28Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> But I do want to point out that \"tracking a directory\" is not at all\n> the same thing as \"tracking a file\", no matter how much you try to\n> argue otherwise.\n\nSince I did not try to argue this, could you beat another strawman?\nI have seen this prepackaged rant already, but it does not really\naddress the problem I have been experiencing.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"47758","messageId":"alpine.LFD.0.999.0707181103580.27353@woody.linux-foundation.org","threadId":"9086","inReplyTo":"85r6n5y5m3.fsf@lola.goethe.zz","subject":"Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-18T18:05:55Z","receivedAt":"2007-07-18T18:05:55Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 18 Jul 2007, David Kastrup wrote:\n> \n> Since I did not try to argue this, could you beat another strawman?\n\nHow about a bit of honesty?\n\nHere's the quote:\n\n \"The FAQ answer is weazeling on several accounts:\n\n  a) No, git only cares about files, or rather git tracks content and\n     empty directories have no content.\n\n  In the same manner as empty regular files have no contents, and git\n  tracks those.  Existence and permissions are important.\"\n\nYou called it \"weaselly\" to say that git tracks only content, and then \nvery much tried to equate \"existence and permissions\" with content.\n\nThat's the part I answered.\n\nSo it wasn't a strawman, it was a direct answer to your assertion. Now go \naway and either come back with the patch to implement it (that I have \nencouraged you to do), or add a \".gitignore\" file to the directory (that \nothers have told you will solve your problems).\n\nDon't bother talking crap.\n\n\t\t\tLinus\n"},{"id":"47773","messageId":"85644hxujp.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707181004330.27353@woody.linux-foundation.org","subject":"Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-18T21:37:30Z","receivedAt":"2007-07-18T21:37:30Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Wed, 18 Jul 2007, Matthieu Moy wrote:\n>> \n>> If you checkout a branch, create an empty directory in this branch\n>> (probably a placeholder, either for future versionned files, or for\n>> generated files), you cannot tell git \"this empty directory is in this\n>> branch, but not in other ones\" without adding a file in it.\n>\n> Right. Which is the suggested setup: add an empty \".gitignore\" file\n> to the directory, and you're done.\n\nThat implies that every directory in a versioned tree will exclusively\nbe created under manual and conscious control.  Not by running some\ninstaller or script, unpacking some archive and so on.  But if every\ncontent on a disk was created and put there under manual control of\nthe disk owner, we could still get along with floppy disks quite fine.\nIn practice, much more content gets sent around and juggled than what\nis under immediate supervision of the user.\n\nThis is getting silly: you don't need to pull out rabbits out of your\nhead.  You said that you are not inclined to do any work in that area\nsince it does not touch _your_ use cases (well, at least not to a\ndegree that you consider worth bothering about) but that is no reason\nto get into ridiculous arguments about other usage.  No code will come\nof that.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"47775","messageId":"alpine.LFD.0.999.0707181444070.27353@woody.linux-foundation.org","threadId":"9086","inReplyTo":"85644hxujp.fsf@lola.goethe.zz","subject":"Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-18T21:45:06Z","receivedAt":"2007-07-18T21:45:06Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 18 Jul 2007, David Kastrup wrote:\n>\n> You said that you are not inclined to do any work in that area\n> since it does not touch _your_ use cases (well, at least not to a\n> degree that you consider worth bothering about) but that is no reason\n> to get into ridiculous arguments about other usage.\n\nHow hard is it for you to admit that I also said \"please send in a patch\".\n\nI don't need it. You do. You do the work. I'm just explaining why the work \nhasn't been done.\n\n\t\tLinus\n"},{"id":"47778","messageId":"85k5sxwbii.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707181444070.27353@woody.linux-foundation.org","subject":"Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-18T23:13:57Z","receivedAt":"2007-07-18T23:13:57Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Wed, 18 Jul 2007, David Kastrup wrote:\n>>\n>> You said that you are not inclined to do any work in that area\n>> since it does not touch _your_ use cases (well, at least not to a\n>> degree that you consider worth bothering about) but that is no reason\n>> to get into ridiculous arguments about other usage.\n>\n> How hard is it for you to admit that I also said \"please send in a\n> patch\".\n\nYup, that was one sentence in about 5 pages of bile.  In contrast,\nJunio gave a good overview of the technical areas involved here, and\nestimates about what to do there best.\n\nThat's a constructive way to encite somebody to delve into the task\nand try to see whether he can come up with something.\n\nBut 5 pages of what amounts to \"you are an idiot, come up with a\npatch\" is not leading anywhere.\n\n> I don't need it. You do. You do the work. I'm just explaining why\n> the work hasn't been done.\n\nNo, you are _defending_ why the work has not been done.  This\nrationalizing around the bush is a waste of time.  You probably have\nspent quite more time with your venting than Junio did with his\ntechnical analysis, and the latter has been much more helpful.\n\nSo why waste all that time and adrenaline on something where you have\nalready said all you consider relevant?  The arguments don't get any\nstronger by shouting, and it is not like you are inconvenienced in any\nmanner if somebody takes a look at the matter.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"47780","messageId":"alpine.LFD.0.999.0707181557270.27353@woody.linux-foundation.org","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707181444070.27353@woody.linux-foundation.org","subject":"[RFC PATCH] Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-18T23:16:26Z","receivedAt":"2007-07-18T23:16:26Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\nGaah.\n\nI'm a damn softie (and soft in the head too, for writing the code).\n\nOk, here's a trivial patch to start the ball rolling. I'm really not \ninterested in taking this patch any further personally, but I'm hoping \nthat maybe it can make somebody else who is actually _interested_ in \ntrackign empty directories (hint hint) decide that it's a good enough \nstart that they can fill in the details.\n\nThis really updates three different areas, which are nicely separated into \nthree different files, so while it's one single patch, you can actually \nfollow along the changes by just looking at the differences in each file, \nwhich directly translate to separate conceptual changes:\n\n - builtin-update-index.c\n\n   This simply contains the changes to update the index file. As usual, \n   there are multiple different cases, and they boil down to:\n\n\t(a) No index entry existed at all previously. If so, a directory \n\t    will first go through the \"index_path()\" logic, which tries to \n\t    create a GITLINK entry for it, if the subdirectory is a git \n\t    directory. However, the new thing is that if that fails, it \n\t    will instead just create a fake empty tree entry for it, and \n\t    set the index mode to S_IFDIR.\n\n\t(b) It was a gitlink entry before. It stays as a gitlink entry, \n\t    even if it cannot be indexed, and a file/symlink entry in \n\t    the working tree is a conflict error.\n\n\t(c) It was a empty directory entry before. A directory stays as an \n\t    empty directory entry, and a file/symlink entry in the working \n\t    tree is a conflict error.\n\n   Somebody should check that we properly delete the directory entry if we \n   add a file under it, I honestly didn't bother to go through all the \n   logic. I *think* we do it correctly just thanks to all the previous \n   code for gitlinks. Whatever.\n\n   What I'm trying to say is that the changes are fairly straightforward, \n   but if somebody decides to push this, they need to think about it a lot \n   more than I'm ready to right now.\n\n - read-cache.c: match the new index type with the filesystem.\n\n   This is pretty damn obvious. A S_ISDIR() always matches, and nothing \n   else matches at all. \n\n - unpack-trees.c: unpack empty directories not by unpacking them \n   recursively into the index, but by adding them directly to the index as \n   a S_IFDIR entry instead.\n\n   This one almost certainly needs more work, in particular when merging \n   trees where one has an empty directory, and the other has files _in_ \n   that directory! But the trivial approach makes a simple \"git read-tree\"\n   with an empty directory unpack it into the index as a S_IFDIR entry, so \n   now doing git-write-tree + git-read-tree should result in the original \n   index contents.\n\nI think the patch itself is pretty simple, but the subtle interactions \nthat flow out of this all are anything but. It may \"just work\" almost \nas-is, but quite frankly, I think people need to think about all the \nissues that can happen a lot!\n\nSo see this as a basis for further work. The \"further work\" may be pretty \nsimple, or it may not be. I'm personally not that interested, but like my \noriginal \"subprojects\" series, hopefully somebody else ends up running \nwith this (or alternatively just proving that trying to track empty \ndirectories is a total nightmare).\n\n\t\t\tLinus\n\n---\n builtin-update-index.c |   33 +++++++++++++++++++++++----------\n read-cache.c           |    4 ++++\n unpack-trees.c         |   12 +++++++++---\n 3 files changed, 36 insertions(+), 13 deletions(-)\n\ndiff --git a/builtin-update-index.c b/builtin-update-index.c\nindex 509369e..2eb2a46 100644\n--- a/builtin-update-index.c\n+++ b/builtin-update-index.c\n@@ -94,8 +94,16 @@ static int add_one_path(struct cache_entry *old, const char *path, int len, stru\n \tfill_stat_cache_info(ce, st);\n \tce->ce_mode = ce_mode_from_stat(old, st->st_mode);\n \n-\tif (index_path(ce->sha1, path, st, !info_only))\n-\t\treturn -1;\n+\tif (index_path(ce->sha1, path, st, !info_only)) {\n+\t\t/*\n+\t\t * If we weren't able to index the directory as a GITLINK,\n+\t\t * see if we can just add it as a plain directory instead.\n+\t\t */\n+\t\tif (!S_ISDIR(st->st_mode))\n+\t\t\treturn -1;\n+\t\tce->ce_mode = htonl(S_IFDIR);\n+\t\tpretend_sha1_file(NULL, 0, OBJ_TREE, ce->sha1);\n+\t}\n \toption = allow_add ? ADD_CACHE_OK_TO_ADD : 0;\n \toption |= allow_replace ? ADD_CACHE_OK_TO_REPLACE : 0;\n \tif (add_cache_entry(ce, option))\n@@ -134,6 +142,11 @@ static int process_directory(const char *path, int len, struct stat *st)\n \t/* Exact match: file or existing gitlink */\n \tif (pos >= 0) {\n \t\tstruct cache_entry *ce = active_cache[pos];\n+\n+\t\t/* Was it a directory before? */\n+\t\tif (S_ISDIR(ntohl(ce->ce_mode)))\n+\t\t\treturn 0;\n+\n \t\tif (S_ISGITLINK(ntohl(ce->ce_mode))) {\n \n \t\t\t/* Do nothing to the index if there is no HEAD! */\n@@ -162,12 +175,8 @@ static int process_directory(const char *path, int len, struct stat *st)\n \t\treturn error(\"%s: is a directory - add individual files instead\", path);\n \t}\n \n-\t/* No match - should we add it as a gitlink? */\n-\tif (!resolve_gitlink_ref(path, \"HEAD\", sha1))\n-\t\treturn add_one_path(NULL, path, len, st);\n-\n-\t/* Error out. */\n-\treturn error(\"%s: is a directory - add files inside instead\", path);\n+\t/* No match - try to just add it as-is */\n+\treturn add_one_path(NULL, path, len, st);\n }\n \n /*\n@@ -178,8 +187,12 @@ static int process_file(const char *path, int len, struct stat *st)\n \tint pos = cache_name_pos(path, len);\n \tstruct cache_entry *ce = pos < 0 ? NULL : active_cache[pos];\n \n-\tif (ce && S_ISGITLINK(ntohl(ce->ce_mode)))\n-\t\treturn error(\"%s is already a gitlink, not replacing\", path);\n+\tif (ce) {\n+\t\tif (S_ISGITLINK(ntohl(ce->ce_mode)))\n+\t\t\treturn error(\"%s is already a gitlink, not replacing\", path);\n+\t\tif (S_ISDIR(ntohl(ce->ce_mode)))\n+\t\t\treturn error(\"%s is already a directory entry, not replacing\", path);\n+\t}\n \n \treturn add_one_path(ce, path, len, st);\n }\ndiff --git a/read-cache.c b/read-cache.c\nindex a363f31..d3d2cc0 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -142,6 +142,10 @@ static int ce_match_stat_basic(struct cache_entry *ce, struct stat *st)\n \t\t    (has_symlinks || !S_ISREG(st->st_mode)))\n \t\t\tchanged |= TYPE_CHANGED;\n \t\tbreak;\n+\tcase S_IFDIR:\n+\t\tif (!S_ISDIR(st->st_mode))\n+\t\t\tchanged |= TYPE_CHANGED;\n+\t\treturn changed;\n \tcase S_IFGITLINK:\n \t\tif (!S_ISDIR(st->st_mode))\n \t\t\tchanged |= TYPE_CHANGED;\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 89dd279..22e452b 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -181,9 +181,13 @@ static int unpack_trees_rec(struct tree_entry_list **posns, int len,\n \t\t\t\tany_dirs = 1;\n \t\t\t\tparse_tree(tree);\n \t\t\t\tsubposns[i] = create_tree_entry_list(tree);\n-\t\t\t\tposns[i] = posns[i]->next;\n-\t\t\t\tsrc[i + o->merge] = o->df_conflict_entry;\n-\t\t\t\tcontinue;\n+\n+\t\t\t\t/* If it wasn't empty, recurse into it */\n+\t\t\t\tif (subposns[i]) {\n+\t\t\t\t\tposns[i] = posns[i]->next;\n+\t\t\t\t\tsrc[i + o->merge] = o->df_conflict_entry;\n+\t\t\t\t\tcontinue;\n+\t\t\t\t}\n \t\t\t}\n \n \t\t\tif (!o->merge)\n@@ -197,6 +201,8 @@ static int unpack_trees_rec(struct tree_entry_list **posns, int len,\n \n \t\t\tce = xcalloc(1, ce_size);\n \t\t\tce->ce_mode = create_ce_mode(posns[i]->mode);\n+\t\t\tif (posns[i]->directory)\n+\t\t\t\tce->ce_mode = htonl(S_IFDIR);\n \t\t\tce->ce_flags = create_ce_flags(baselen + pathlen,\n \t\t\t\t\t\t       ce_stage);\n \t\t\tmemcpy(ce->name, base, baselen);\n"},{"id":"47783","messageId":"7vy7hd5lri.fsf@assigned-by-dhcp.cox.net","threadId":"9086","inReplyTo":"867ioyqhgc.fsf@lola.quinscape.zz","subject":"Re: Empty directories...","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-07-18T23:34:41Z","receivedAt":"2007-07-18T23:34:41Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"David Kastrup <dak@gnu.org> writes:\n\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n>> Having made it sound so easy, here are the issues I would expect\n>> to be nontrivial (but probably not rocket surgery either).\n>> ...\n> This would seem to imply that the index does not need to be\n> upwards-compatible: simplifying the code means that old indexes won't\n> be treated all too well.\n\nI did not imply any such thing, by the way.  These are off the\ntop of my head technical issues and there probably are more, but\nI limited the list to technical side of the things.\n\nYou of course have social side to take care of.  If you are\nbreaking everybody else's index, you would need to tell\neverybody: \"I am sorry but if you upgrade your git to this\nversion that does what I want, you have to nuke your index and\nstart over, so commit all changes first, and then update the\ngit.  Sorry for causing you a minor inconvenience\".  Everybody\nat this point involves (obviously) the kernel folks, wine,\nx.org, among many others.\n\nI suspect your saying that to them is probably not good enough\nfor them to forgive the minor inconveniences, which means you\nneed to convince _me_ to join you in defending, in the release\nnotes, that this is a feature worth having even though there is\na minor inconvenience to redo everybody's index files.  Which I\nsuspect is quite unlikely to happen at this moment, though...\n\nA much less troublesome approach might be to do things\ndifferently from what I outlined, to keep the index compatible\nas long as it does not contain an empty directory, which is what\nwe did for subprojects support.\n"},{"id":"47785","messageId":"alpine.LFD.0.999.0707181628270.27353@woody.linux-foundation.org","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707181557270.27353@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-18T23:40:32Z","receivedAt":"2007-07-18T23:40:32Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 18 Jul 2007, Linus Torvalds wrote:\n>\n> +\t\tif (!S_ISDIR(st->st_mode))\n> +\t\t\treturn -1;\n> +\t\tce->ce_mode = htonl(S_IFDIR);\n> +\t\tpretend_sha1_file(NULL, 0, OBJ_TREE, ce->sha1);\n\nOh, one word of warning: that whole \"pretend_sha1_file()\" thing won't \ncreate the object itself, and when I did the limited testing that I did, I \nactually made sure had a magic zero-sized tree object in my object \ndirectory.\n\nIf you don't, some things will complain, because they end up getting a \nSHA1 that they cannot look up, becasue *they* didn't create that pretend \nentry.\n\nI didn't know which way I wanted to go with that thing. I was kind of \nthinking that maybe we would just have the zero-sized OBJ_BLOB and \nOBJ_TREE objects as special magical things, and have all git programs just \ndo that \"pretend\" at the beginning.\n\nBut that kind of thing is probably just a totally unnecessary special \ncase, and instead, that \"pretend_sha1_file()\" should have just been a\n\n\twrite_sha1_file(NULL, 0, \"tree\", ce->sha1);\n\ninstead.\n\nAnyway, if there are issues with not finding an object called \n4b825dc642cb6eb9a060e54bf8d69288fbee4904, then that's the empty tree \nobject, and that pretend thing was the cause.\n\n(The git repo itself has the empty tree as an object in it, because one of \nthe commits has that - probably as a result of a bug, but there you have \nit)\n\n\t\tLinus\n"},{"id":"47784","messageId":"85abttwa7m.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707181557270.27353@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-18T23:42:05Z","receivedAt":"2007-07-18T23:42:05Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> Gaah.\n>\n> I'm a damn softie (and soft in the head too, for writing the code).\n>\n> Ok, here's a trivial patch to start the ball rolling. I'm really not \n> interested in taking this patch any further personally, but I'm hoping \n> that maybe it can make somebody else who is actually _interested_ in \n> trackign empty directories (hint hint) decide that it's a good enough \n> start that they can fill in the details.\n\nWell, kudos.  Together with the analysis from Junio, this seems like a\ngood start.  Would you have any recommendations about what stuff one\nshould really read in order to get up to scratch about git internals?\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"47789","messageId":"alpine.LFD.0.999.0707181710271.27353@woody.linux-foundation.org","threadId":"9086","inReplyTo":"85abttwa7m.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-19T00:22:57Z","receivedAt":"2007-07-19T00:22:57Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 19 Jul 2007, David Kastrup wrote:\n> \n> Well, kudos.  Together with the analysis from Junio, this seems like a\n> good start.  Would you have any recommendations about what stuff one\n> should really read in order to get up to scratch about git internals?\n\nWell, you do need to understand the index. That's where all the new \nsubtlety happens.\n\nThe data structures themselves are trivial, and we've supported empty \ntrees (at the top level) from the beginning, so that part is not anything \nnew.\n\nHowever, now having a new entry type in the index (S_IFDIR) means that \nanything that interacts with the index needs to think twice. But a lot of \nthat is just testing what happens, and so the first thing to do is to have \na test-suite.\n\nThere's also the question about how to show an empty tree in a diff. We've \nnever had that: the only time we had empty trees was when we compared a \ntotally empty \"root\" tree against another tree, and then it was obvious. \nBut what if the empty tree is a subdirectory of another tree - how do you \nexpress that in a diff? Do you care? Right now, since we always recurse \ninto the tree (and then not find anything), empty trees will simply not \nshow up _at_all_ in any diffs.\n\nAnd what about usability issues elsewhere? With my patch, doing something \nlike a\n\n\tgit add directory/\n\nstill won't do anything, because the behaviour of \"git add\" has always \nbeen to recurse into directories. So to add a new empty directory, you'd \nhave to do\n\n\tgit update-index --add directory\n\nand that's not exactly user-friendly.\n\nSo do you add a \"-n\" flag to \"git add\" to tell it to not recurse? Or do \nyou always recurse, but then if you notice that the end result is empty, \nyou add it as a directory?\n\nAll the logic for that whole directory lookup is in git/dir.c, and that \ncode takes various flags because different programs want different things \n(show \"ignored\" files, or ignore them? Show empty directories or ignore \nthem? etc).\n\nSo primarily, I think the job is:\n\n - thinking about the index, and the interactions when adding a directory \n   or adding files under a directory that already exists.\n\n   I *think* we get all the corner cases right, because they should be \n   exactly the same as with subprojects, but hey, maybe there's some piece \n   that tests S_ISGITLINK() and now needs a S_ISDIR() test too..\n\n - adding test cases\n\n - thinking about the user interfaces for this, and adding code to handle \n   directories where needed (eg the above \"git add\" issue).\n\n - thinking about merges (which is largely about the index too, but is a \n   whole 'nother set of issues, with multiple stages in the same index at \n   the same time)\n\nIt might all be trivial. The directory traversal already knows that empty \ndirectories are special, so getting the right behaviour to \"git add\" may \nbe really really easy. Or maybe it's not. I think a lot of it is just \nfinding what needs to be done, seeign if we already do it, and if not, \nseeign how to do it. Boring test-cases, in other words.\n\n\t\tLinus\n"},{"id":"47807","messageId":"7vbqe93qtv.fsf@assigned-by-dhcp.cox.net","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707181710271.27353@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-07-19T05:28:12Z","receivedAt":"2007-07-19T05:28:12Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Thu, 19 Jul 2007, David Kastrup wrote:\n>> \n>> Well, kudos.  Together with the analysis from Junio, this seems like a\n>> good start.  Would you have any recommendations about what stuff one\n>> should really read in order to get up to scratch about git internals?\n>\n> Well, you do need to understand the index. That's where all the new \n> subtlety happens.\n>\n> The data structures themselves are trivial, and we've supported empty \n> trees (at the top level) from the beginning, so that part is not anything \n> new.\n>\n> However, now having a new entry type in the index (S_IFDIR) means that \n> anything that interacts with the index needs to think twice. But a lot of \n> that is just testing what happens, and so the first thing to do is to have \n> a test-suite.\n>\n> There's also the question about how to show an empty tree in a diff. We've \n> never had that: the only time we had empty trees was when we compared a \n> totally empty \"root\" tree against another tree, and then it was obvious. \n> But what if the empty tree is a subdirectory of another tree - how do you \n> express that in a diff? Do you care? Right now, since we always recurse \n> into the tree (and then not find anything), empty trees will simply not \n> show up _at_all_ in any diffs.\n>\n> And what about usability issues elsewhere? With my patch, doing something \n> like a\n>\n> \tgit add directory/\n>\n> still won't do anything, because the behaviour of \"git add\" has always \n> been to recurse into directories. So to add a new empty directory, you'd \n> have to do\n>\n> \tgit update-index --add directory\n>\n> and that's not exactly user-friendly.\n>\n> So do you add a \"-n\" flag to \"git add\" to tell it to not recurse? Or do \n> you always recurse, but then if you notice that the end result is empty, \n> you add it as a directory?\n\nAnother issue I thought about was what you would do in the step\n3 in the following:\n\n 1. David says \"mkdir D; git add D\"; you add S_IFDIR entry in\n    the index at D;\n\n 2. David says \"date >D/F; git add D/F\"; presumably you drop D\n    from the index (to keep the index more backward compatible)\n    and add S_IFREG entry at D/F.\n\n 3. David says \"git rm D/F\".\n\nHave we stopped keeping track of the \"empty directory\" at this\npoint?\n"},{"id":"47808","messageId":"20070719053858.GE32566@spearce.org","threadId":"9086","inReplyTo":"7vbqe93qtv.fsf@assigned-by-dhcp.cox.net","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-07-19T05:38:58Z","receivedAt":"2007-07-19T05:38:58Z","isPatch":true,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Junio C Hamano <gitster@pobox.com> wrote:\n> Another issue I thought about was what you would do in the step\n> 3 in the following:\n> \n>  1. David says \"mkdir D; git add D\"; you add S_IFDIR entry in\n>     the index at D;\n> \n>  2. David says \"date >D/F; git add D/F\"; presumably you drop D\n>     from the index (to keep the index more backward compatible)\n>     and add S_IFREG entry at D/F.\n> \n>  3. David says \"git rm D/F\".\n> \n> Have we stopped keeping track of the \"empty directory\" at this\n> point?\n\nSadly yes.  But I don't think that's what the folks who want to\ntrack empty directories want to have happen here.\n\nWhich is why I'm thinking we just need to track the directory, as a\nnode in the index, even if there are files in it, and even if we got\nthat directory and its contained files there by just unpacking trees.\n\n-- \nShawn.\n"},{"id":"47810","messageId":"85zm1tue6a.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"7vbqe93qtv.fsf@assigned-by-dhcp.cox.net","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-19T05:59:25Z","receivedAt":"2007-07-19T05:59:25Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Another issue I thought about was what you would do in the step\n> 3 in the following:\n>\n>  1. David says \"mkdir D; git add D\"; you add S_IFDIR entry in\n>     the index at D;\n>\n>  2. David says \"date >D/F; git add D/F\"; presumably you drop D\n>     from the index (to keep the index more backward compatible)\n>     and add S_IFREG entry at D/F.\n\nI don't think that one should drop D here.  Operation 1 _is_ not\nbackward compatible, so if you want to revert it, you should\nexplicitly remove D.  And we can't \"keep\" the index backward\ncompatible if it isn't so after step 1.\n\n>  3. David says \"git rm D/F\".\n>\n> Have we stopped keeping track of the \"empty directory\" at this\n> point?\n\nThe case I am worrying about is rather\n\nmkdir D\nmkdir D/E\ntouch D/E/file\ngit add D\n[*]\ngit rm D/E/file\n\n>From a user perspective, E should be registered still.  Compare this\nwith\n\nmkdir D\nmkdir D/E\ntouch D/E/file\ngit add D/E/file\n[*]\ngit rm D/E/file\n\nWhere likely both D and E should now be considered unregistered.  So\nthe situation is different between the first or the second [*], and\nthe difference might be impossible to express completely in the frame\nof a backwards-compatible index, even though we don't track an empty\ndirectory at the point [*] at all, and the only registered _file_ is\nD/E/file.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"47812","messageId":"85vechudrl.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"20070719053858.GE32566@spearce.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-19T06:08:14Z","receivedAt":"2007-07-19T06:08:14Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"\"Shawn O. Pearce\" <spearce@spearce.org> writes:\n\n> Sadly yes.  But I don't think that's what the folks who want to\n> track empty directories want to have happen here.\n>\n> Which is why I'm thinking we just need to track the directory, as a\n> node in the index, even if there are files in it, and even if we got\n> that directory and its contained files there by just unpacking\n> trees.\n\nI have come to about the same conclusion.  So if\nbackward-compatibility is any concern, one needs to work with some\nsort of extension records, and designing them in a way that\n\nnew-git add tree\nold-git rm tree\n\nwill not leave empty subdirectories in the index will be tricky, to\nsay the least.  One will likely have to add an extension record\n\"directory\" for each directory as well as \"my containing dir takes\ncare of itself\" to each file that has been added with new-git and has\nhad its parent directory entered by other means.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"47813","messageId":"20070719060922.GF32566@spearce.org","threadId":"9086","inReplyTo":"20070719053858.GE32566@spearce.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-07-19T06:09:22Z","receivedAt":"2007-07-19T06:09:22Z","isPatch":true,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"Shawn O. Pearce\" <spearce@spearce.org> wrote:\n> Junio C Hamano <gitster@pobox.com> wrote:\n> > Another issue I thought about was what you would do in the step\n> > 3 in the following:\n> > \n> >  1. David says \"mkdir D; git add D\"; you add S_IFDIR entry in\n> >     the index at D;\n> > \n> >  2. David says \"date >D/F; git add D/F\"; presumably you drop D\n> >     from the index (to keep the index more backward compatible)\n> >     and add S_IFREG entry at D/F.\n> > \n> >  3. David says \"git rm D/F\".\n> > \n> > Have we stopped keeping track of the \"empty directory\" at this\n> > point?\n> \n> Sadly yes.  But I don't think that's what the folks who want to\n> track empty directories want to have happen here.\n> \n> Which is why I'm thinking we just need to track the directory, as a\n> node in the index, even if there are files in it, and even if we got\n> that directory and its contained files there by just unpacking trees.\n\nI take this back.  I really don't want that behavior.\n\nIf I do:\n\n  mkdir -p foo/bar\n  echo hello >foo/bar/world\n  git add foo\n  git -f rm foo/bar/world\n\nI never asked for foo/bar or foo to stay.  In fact I want them\nto disappear from Git entirely, as foo/bar is now empty and has\nno content.\n\n\nBut we also cannot do a special --mkdir option for update-index\neither, because how do we know that the user designated subtree is\na directory we must always keep in the index?\n\nSo I think the only way this works is to have a new mode that we use\nin tree (04755 ?) that tells us not only is this thing a subtree,\nbut also that the user wants it to stay here, even if it is empty.\nThose trees are always in the index as a real tree entry, even if\nthere are files contained in it.\n\nAnd as far as getting that directory entry created/removed from\nthe index, well, I think a special flag to update-index would be\nin order, much like --chmod=[+-]x.\n\nJust my $0.0002 USD, which really ain't worth much at all.\n\n-- \nShawn.\n"},{"id":"47819","messageId":"93c3eada0707190010r10d048q72353304b9751770@mail.gmail.com","threadId":"9086","inReplyTo":"85vechudrl.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Geoff Russell","fromEmail":"geoffrey.russell@gmail.com","sentAt":"2007-07-19T07:10:33Z","receivedAt":"2007-07-19T07:10:33Z","isPatch":true,"sender":{"key":"geoffrey.russell@gmail.com","avatar":"https://gravatar.com/avatar/c30f497ccfa6bf06d86f30bd2ba092a2dd124c61c6bc902f7cb5c3f6486947de?d=mp&s=160"},"body":"Dear gits,\n\nWhen I first started using git, I naively did\n\n           $ mkdir NEWDIR && chmod BLAH NEWDIR\n           $ git add NEWDIR\n\nI just expected that this was content in the current directory that I\nwanted tracked\ntogether with the permissions.\n\nIt wasn't ... I spent a day or 2 thinking I was stupid, my version of git was\ncorrupt, my machine was busted, .... etc.  Eventually of course, I read the\ndocumentation (when all else fails) and realised that this perfectly obvious\nbehaviour was not supported.  The behaviour was obviously so obvious\nthat eventually\nan error message was added telling all the people who hadn't\nread the documentation that trying to add a directory was 'fatal'.\n\nI put up with and work around this behaviour because git is so bloody\nbrilliant at everything else.  But it would be nice if it worked.\n\nCheers,\nGeoff Russell\n"},{"id":"47823","messageId":"vpqvecgvmjh.fsf@bauges.imag.fr","threadId":"9086","inReplyTo":"20070719060922.GF32566@spearce.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2007-07-19T08:13:22Z","receivedAt":"2007-07-19T08:13:22Z","isPatch":true,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"\"Shawn O. Pearce\" <spearce@spearce.org> writes:\n\n> If I do:\n>\n>   mkdir -p foo/bar\n>   echo hello >foo/bar/world\n>   git add foo\n>   git -f rm foo/bar/world\n>\n> I never asked for foo/bar or foo to stay.\n\nWell, outside git, if you do\n\n$ mkdir -p foo/bar\n$ echo hello > foo/bar/world\n$ rm -f foo/bar/world\n\nYou didn't ask foo/bar to stay either, and still, it's quite natural\nto have it stay in your filesystem. So, the same way you'd have ran\n\"rm -r foo\", it seems reasonable to me to ask for \"git-rm -r foo\" if\nthe user wants to get rid of foo/ itself.\n\n-- \nMatthieu\n"},{"id":"47830","messageId":"86k5swen1w.fsf@lola.quinscape.zz","threadId":"9086","inReplyTo":"85zm1tue6a.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-19T09:54:19Z","receivedAt":"2007-07-19T09:54:19Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"David Kastrup <dak@gnu.org> writes:\n\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n>> Another issue I thought about was what you would do in the step\n>> 3 in the following:\n>>\n>>  1. David says \"mkdir D; git add D\"; you add S_IFDIR entry in\n>>     the index at D;\n>>\n>>  2. David says \"date >D/F; git add D/F\"; presumably you drop D\n>>     from the index (to keep the index more backward compatible)\n>>     and add S_IFREG entry at D/F.\n>\n> I don't think that one should drop D here.  Operation 1 _is_ not\n> backward compatible, so if you want to revert it, you should\n> explicitly remove D.  And we can't \"keep\" the index backward\n> compatible if it isn't so after step 1.\n>\n>>  3. David says \"git rm D/F\".\n>>\n>> Have we stopped keeping track of the \"empty directory\" at this\n>> point?\n>\n> The case I am worrying about is rather\n>\n> mkdir D\n> mkdir D/E\n> touch D/E/file\n> git add D\n> [*]\n> git rm D/E/file\n>\n> From a user perspective, E should be registered still.  Compare this\n> with\n>\n> mkdir D\n> mkdir D/E\n> touch D/E/file\n> git add D/E/file\n> [*]\n> git rm D/E/file\n\nLet's take this through the motions with my last proposal: at the\nfirst [*], the index now contains\n\nD/.        [dir]\nD/E/.      [dir]\nD/E/file   [file]\n\nAfter git rm D/E/file, it contains\n\nD/.        [dir]\nD/E/.      [dir]\n\nCompared with the second, where we just have in the index\n\nD/E/file   [file]\n\nand it is gone again after the remove.\n\nAfter commiting in the first case, we have in the repository\nD          [tree]\nD/.        [dir]\nD/E        [tree]\nD/E/.      [dir]\nD/E/file   [file]\n\nNow we do\ngit rm D/E, and the index contains\n\nD/E/.      [remove dir]\nD/E/file   [remove file]\n\nIf we commit now,\nD/E        [tree]\nbecomes empty and is removed.  All that stays is\n\nD          [tree]\nD/.        [dir]\n\nSo we still have [tree] items only in the repository, not in the\nindex, and there is no such thing as an empty tree.  But directories\nhave a presence in index and repository.  They are not containers of\nfiles, that role is retained by trees.  Rather they are siblings of\nthe files in their associated tree.\n\nAs a note aside: if one wanted to track directory permissions, one\nwould track them in the [dir] entries, not in the [tree] entries.\nTrees remain abstract structuring entities in the repository that\ndon't have an outside representation.  Directories will be\nauto-created and deleted as necessary in the work directory to\nfacilitate having a place for checking tree elements out and in.\n\nThis means that\ngit add D/E/file\nwould _not_ track permissions of D and E (nor their existence).\n\nHowever, Linus is right that permissions are something to be discussed\nseparately.  But separating [tree] and [dir] makes for a plausible and\nunderstandable way of treating them.\n\n-- \nDavid Kastrup\n"},{"id":"47837","messageId":"20070719105105.GA4929@moonlight.home","threadId":"9086","inReplyTo":"vpqvecgvmjh.fsf@bauges.imag.fr","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Tomash Brechko","fromEmail":"tomash.brechko@gmail.com","sentAt":"2007-07-19T10:51:05Z","receivedAt":"2007-07-19T10:51:05Z","isPatch":true,"sender":{"key":"tomash.brechko@gmail.com","avatar":null},"body":"Dear Git fellows,\n\nA year or so ago I too would strongly advocate the need of tracking\nempty directories, permissions et al., it seemed so \"natural\" and\n\"plain obvious\" to me back then.  But since that time I learned to\nappreciate the \"contents tracking\" approach, and now view directories\n(paths in general) only as the means for Git to know where to put the\ncontents on checkout.  This, BTW, is consistent with how Git figures\ncontainer copies/renames.\n\nNo doubt mighty Git developers can add support for empty directories,\nmanage to stay backward compatible, think out consistent user\ninterface etc.  But there's no end to how much information one may\nwant to store in Git to make it \"_file system_ contents tracking\nsoftware\".  Starting with empty directories, one may argue then that\ncertain installation trees also need particular file ownership, so\nlets store user/group names like tar does.  It was mentioned already\nin this thread that in addition to 'rwx' we also would have to store\nACLs (some OSes have only one of these concepts, some both), SELinux\nsecurity contexts, perhaps other arbitrary file attributes that may be\npart of file system state.\n\nWouldn't it be better to preserve Git as a contents tracking system,\nand add some tools on top of it that can translate file system state\ninto textual (or binary) form, so it can be stored in current Git?\nAnd then use this textual representation to restore actual file system\nattributes/layout on checkout?  And the only change in Git itself\nwould be some more hooks, for instance one hook before checking out\nover the old work tree, and one after the checkout.  Or one can simply\nwrap certain Git commands to implement such hooks.\n\nIn any case, no one is going to be against the new feature if it won't\nbreak anything for those of us who find the pure contents tracking the\nright thing.  And storing empty directories by default may not be\nnatural for everyone.  So before going into technical details of how\nthis can possibly be implemented, could someone answer the following:\n\n1 Is Git going to track directories _always_?  Looks like not, because\n  in this thread there seems to be a distinction between 'git add DIR'\n  and 'git add DIR/FILE', i.e. not everyone is sure if in the last\n  case Git should track DIR or not.\n\n2 If Git will track only explicitly mentioned directories, then what\n  about recursive operations?  Will it add only files by default, or\n  directories too?  Perhaps there will be some --add-dirs option to\n  'git add'.\n\n3 Since in certain recursive operations one will want to affect\n  directories too, how .gitignore will look?  Most files have a notion\n  of extension, so me may say '*.o', but with directories things a bit\n  more complicated.  One would want to say \"exclude DIR2 only if under\n  DIR1 at any hierarchy depth\", i.e. exclude paths matching\n  qr%DIR1/(.+/)?DIR1/%, and shell wildcards aren't that expressive,\n  '*' doesn't cross hierarchy.  Note that we live without this now,\n  but this will be the next \"natural\" demand once directories become\n  first class citizens.\n\n\nThis list is surely incomplete.  The point is that before we go into\ntechnical details, let's consider what exactly we are going to\nimplement, how this will affect current usage model, how (empty)\ndirectory handling will extend to future similar demands, etc.  My\nfear is that once some patch is around, it's very tempting to accept\nit.  And once it is in, it's almost impossible to remove the feature\nlater.\n\n\nRegards,\n\n-- \n   Tomash Brechko\n"},{"id":"47843","messageId":"86zm1sbpeh.fsf@lola.quinscape.zz","threadId":"9086","inReplyTo":"20070719105105.GA4929@moonlight.home","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-19T11:31:50Z","receivedAt":"2007-07-19T11:31:50Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Tomash Brechko <tomash.brechko@gmail.com> writes:\n\n> Dear Git fellows,\n>\n> A year or so ago I too would strongly advocate the need of tracking\n> empty directories, permissions et al., it seemed so \"natural\" and\n> \"plain obvious\" to me back then.  But since that time I learned to\n> appreciate the \"contents tracking\" approach, and now view\n> directories (paths in general) only as the means for Git to know\n> where to put the contents on checkout.  This, BTW, is consistent\n> with how Git figures container copies/renames.\n\nI'll answer to this based on my proposal of adding \"A/B/. [dir]\" as a\nseparate entity to index and repository, keeping \"[tree]\" out of\nindices, and don't allow an empty \"[tree]\" into repositories.\n\nThis is a very natural abstraction.\n\n> But there's no end to how much information one may want to store in\n> Git to make it \"_file system_ contents tracking software\".  Starting\n> with empty directories, one may argue then that certain installation\n> trees also need particular file ownership, so lets store user/group\n> names like tar does.  It was mentioned already in this thread that\n> in addition to 'rwx' we also would have to store ACLs (some OSes\n> have only one of these concepts, some both), SELinux security\n> contexts, perhaps other arbitrary file attributes that may be part\n> of file system state.\n\nA [dir] entry may be eventually be made to track any of this, like a\n[file] entry could.  If one wished to do this.\n\n> Wouldn't it be better to preserve Git as a contents tracking system,\n> and add some tools on top of it that can translate file system state\n> into textual (or binary) form, so it can be stored in current Git?\n> And then use this textual representation to restore actual file\n> system attributes/layout on checkout?  And the only change in Git\n> itself would be some more hooks, for instance one hook before\n> checking out over the old work tree, and one after the checkout.  Or\n> one can simply wrap certain Git commands to implement such hooks.\n\nThis is not good since \"tracking\" means \"tracking\".  With your model,\nthe metainformation would be dissociated from the information.\nRenames and moves would make ground beef of the metadata.\n\n> In any case, no one is going to be against the new feature if it\n> won't break anything for those of us who find the pure contents\n> tracking the right thing.\n\nMy proposal would allow setting an option to track or not track\ndirectories implicitly by default.\n\n> And storing empty directories by default may not be natural for\n> everyone.  So before going into technical details of how this can\n> possibly be implemented, could someone answer the following:\n\nI'll answer assuming the proposed model.\n\n> 1 Is Git going to track directories _always_?  Looks like not, because\n>   in this thread there seems to be a distinction between 'git add DIR'\n>   and 'git add DIR/FILE', i.e. not everyone is sure if in the last\n>   case Git should track DIR or not.\n\nLet's have a variable\n\ncore.adddirs\n\nIf you set core.adddirs to false, git will not enter directories into\nthe index for addition.  Consequently, they will not end up in the\nrepository.  If you git-rm a directory, the index will contain a\nnotice to delete the directory along with deletion notices for all\nregistered other elements of the directory.  Committing this means\nthat the directory will no longer be separately controlled by git,\neven if for some reason the repository has other files remaining in\nthe tree.\n\nSomething like the Linux kernel repository which may be accessed by\nancient git versions would naturally contain \"core.adddirs: false\" in\nits default configuration file, and this would be passed around when\ncloning.  So directory elements would stay out of it.\n\n> 2 If Git will track only explicitly mentioned directories, then what\n>   about recursive operations?  Will it add only files by default, or\n>   directories too?  Perhaps there will be some --add-dirs option to\n>   'git add'.\n\nThere could be a commandline override for \"core.adddirs\".\n\n> 3 Since in certain recursive operations one will want to affect\n>   directories too, how .gitignore will look?  Most files have a notion\n>   of extension, so me may say '*.o', but with directories things a bit\n>   more complicated.  One would want to say \"exclude DIR2 only if under\n>   DIR1 at any hierarchy depth\", i.e. exclude paths matching\n>   qr%DIR1/(.+/)?DIR1/%, and shell wildcards aren't that expressive,\n>   '*' doesn't cross hierarchy.  Note that we live without this now,\n>   but this will be the next \"natural\" demand once directories become\n>   first class citizens.\n\nHuh?  I don't get this.  It's like \"we can't allow people to buy\nchocolate, or they'll demand next to have nuclear weapons delivered at\ntheir house\".  Deal with the demands as they come up.  If a directory\nhas a tree-local name \".\", it can be dealt with in patterns if really\nneeded.  I don't see much of a necessity however.\n\nAlthough it would be natural to have\ncore.adddirs: false\nbe equivalent to\ncore.excludefile: .\n\nAnd so it might be possible to actually not need a separate\ncore.adddirs option at all, technically.\n\n-- \nDavid Kastrup\n"},{"id":"47850","messageId":"Pine.LNX.4.64.0707191310430.14781@racer.site","threadId":"9086","inReplyTo":"20070719105105.GA4929@moonlight.home","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-19T12:16:26Z","receivedAt":"2007-07-19T12:16:26Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 19 Jul 2007, Tomash Brechko wrote:\n\n> A year or so ago I too would strongly advocate the need of tracking \n> empty directories, permissions et al., it seemed so \"natural\" and \"plain \n> obvious\" to me back then.  But since that time I learned to appreciate \n> the \"contents tracking\" approach, and now view directories (paths in \n> general) only as the means for Git to know where to put the contents on \n> checkout.  This, BTW, is consistent with how Git figures container \n> copies/renames.\n\nThank you.  It is my impression, too, that after a while it becomes \nobvious what is good and what is not.\n\nFWIW I just whipped up a proof-of-concept patch (so at least _I_ cannot be \naccused of chickening out of writing code):\n\nThis adds the command line option \"--add-empty-dirs\" to \"git add\", which \ndoes the only sane thing: putting a placeholder into that directory, and \nadding that.  Since \".gitignore\" is already a reserved file name in git, \nit is used as the name of this place holder.\n\n---\n\n\tIt is probably not fool-proof yet, needs documentation and a test \n\tcase.  But I am really sick and tired of this discussion.\n\n builtin-add.c |   25 +++++++++++++++++++++----\n dir.c         |   16 +++++++++++++++-\n dir.h         |    3 ++-\n 3 files changed, 38 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin-add.c b/builtin-add.c\nindex 7345479..1294840 100644\n--- a/builtin-add.c\n+++ b/builtin-add.c\n@@ -47,7 +47,7 @@ static void prune_directory(struct dir_struct *dir, const char **pathspec, int p\n }\n \n static void fill_directory(struct dir_struct *dir, const char **pathspec,\n-\t\tint ignored_too)\n+\t\tint ignored_too, int substitute_empty_dirs)\n {\n \tconst char *path, *base;\n \tint baselen;\n@@ -63,6 +63,7 @@ static void fill_directory(struct dir_struct *dir, const char **pathspec,\n \t\tif (!access(excludes_file, R_OK))\n \t\t\tadd_excludes_from_file(dir, excludes_file);\n \t}\n+\tdir->substitute_empty_directories = substitute_empty_dirs;\n \n \t/*\n \t * Calculate common prefix for the pathspec, and\n@@ -143,7 +144,8 @@ static const char ignore_warning[] =\n int cmd_add(int argc, const char **argv, const char *prefix)\n {\n \tint i, newfd;\n-\tint verbose = 0, show_only = 0, ignored_too = 0;\n+\tint verbose = 0, show_only = 0, ignored_too = 0,\n+\t\tsubstitute_empty_dirs = 0;\n \tconst char **pathspec;\n \tstruct dir_struct dir;\n \tint add_interactive = 0;\n@@ -191,6 +193,10 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n \t\t\ttake_worktree_changes = 1;\n \t\t\tcontinue;\n \t\t}\n+\t\tif (!strcmp(arg, \"--add-empty-dirs\")) {\n+\t\t\tsubstitute_empty_dirs = 1;\n+\t\t\tcontinue;\n+\t\t}\n \t\tusage(builtin_add_usage);\n \t}\n \n@@ -206,7 +212,7 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n \t}\n \tpathspec = get_pathspec(prefix, argv + i);\n \n-\tfill_directory(&dir, pathspec, ignored_too);\n+\tfill_directory(&dir, pathspec, ignored_too, substitute_empty_dirs);\n \n \tif (show_only) {\n \t\tconst char *sep = \"\", *eof = \"\";\n@@ -231,8 +237,19 @@ int cmd_add(int argc, const char **argv, const char *prefix)\n \t\texit(1);\n \t}\n \n-\tfor (i = 0; i < dir.nr; i++)\n+\tfor (i = 0; i < dir.nr; i++) {\n+\t\tconst char *name = dir.entries[i]->name;\n+\t\tconst char *slash;\n+\t\tif (substitute_empty_dirs && (slash = strrchr(name, '/')) &&\n+\t\t\t\t!strcmp(slash, \"/.gitignore\") &&\n+\t\t\t\taccess(name, R_OK)) {\n+\t\t\tint fd = open(name, O_WRONLY | O_CREAT | O_EXCL, 0666);\n+\t\t\tif (fd < 0)\n+\t\t\t\treturn error(\"Could not create %s\", name);\n+\t\t\tclose(fd);\n+\t\t}\n \t\tadd_file_to_cache(dir.entries[i]->name, verbose);\n+\t}\n \n  finish:\n \tif (active_cache_changed) {\ndiff --git a/dir.c b/dir.c\nindex 8d8faf5..b0b4628 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -456,11 +456,11 @@ static int read_directory_recursive(struct dir_struct *dir, const char *path, co\n {\n \tDIR *fdir = opendir(path);\n \tint contents = 0;\n+\tchar fullname[PATH_MAX + 1];\n \n \tif (fdir) {\n \t\tint exclude_stk;\n \t\tstruct dirent *de;\n-\t\tchar fullname[PATH_MAX + 1];\n \t\tmemcpy(fullname, base, baselen);\n \n \t\texclude_stk = push_exclude_per_directory(dir, base, baselen);\n@@ -536,6 +536,20 @@ exit_early:\n \t\tpop_exclude_per_directory(dir, exclude_stk);\n \t}\n \n+\tif (!contents && dir->substitute_empty_directories) {\n+\t\tconst char *name = \".gitignore\";\n+\t\tint len = strlen(name);\n+\t\t/* Ignore overly long pathnames! */\n+\t\tif (len + baselen + 8 > sizeof(fullname))\n+\t\t\treturn 0;\n+\t\tmemcpy(fullname + baselen, name, len+1);\n+\t\tif (simplify_away(fullname, baselen + len, simplify)\n+\t\t\t\t|| excluded(dir, fullname))\n+\t\t\treturn 0;\n+\t\tdir_add_name(dir, fullname, baselen + len);\n+\t\treturn 1;\n+\t}\n+\n \treturn contents;\n }\n \ndiff --git a/dir.h b/dir.h\nindex ec0e8ab..0099718 100644\n--- a/dir.h\n+++ b/dir.h\n@@ -34,7 +34,8 @@ struct dir_struct {\n \t\t     show_other_directories:1,\n \t\t     hide_empty_directories:1,\n \t\t     no_gitlinks:1,\n-\t\t     collect_ignored:1;\n+\t\t     collect_ignored:1,\n+\t\t     substitute_empty_directories:1;\n \tstruct dir_entry **entries;\n \tstruct dir_entry **ignored;\n \n"},{"id":"47852","messageId":"86wswwa8ej.fsf@lola.quinscape.zz","threadId":"9086","inReplyTo":"Pine.LNX.4.64.0707191310430.14781@racer.site","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-19T12:24:20Z","receivedAt":"2007-07-19T12:24:20Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> Hi,\n>\n> On Thu, 19 Jul 2007, Tomash Brechko wrote:\n>\n>> A year or so ago I too would strongly advocate the need of tracking \n>> empty directories, permissions et al., it seemed so \"natural\" and \"plain \n>> obvious\" to me back then.  But since that time I learned to appreciate \n>> the \"contents tracking\" approach, and now view directories (paths in \n>> general) only as the means for Git to know where to put the contents on \n>> checkout.  This, BTW, is consistent with how Git figures container \n>> copies/renames.\n>\n> Thank you.  It is my impression, too, that after a while it becomes \n> obvious what is good and what is not.\n>\n> FWIW I just whipped up a proof-of-concept patch (so at least _I_ cannot be \n> accused of chickening out of writing code):\n>\n> This adds the command line option \"--add-empty-dirs\" to \"git add\", which \n> does the only sane thing: putting a placeholder into that directory, and \n> adding that.  Since \".gitignore\" is already a reserved file name in git, \n> it is used as the name of this place holder.\n\nBut that means that checkout will create a file .gitignore in\npreviously empty directories, doesn't it?\n\nI think that the placeholder name should rather be \".\".\n\n-- \nDavid Kastrup\n"},{"id":"47855","messageId":"20070719123214.GB4929@moonlight.home","threadId":"9086","inReplyTo":"86zm1sbpeh.fsf@lola.quinscape.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Tomash Brechko","fromEmail":"tomash.brechko@gmail.com","sentAt":"2007-07-19T12:32:14Z","receivedAt":"2007-07-19T12:32:14Z","isPatch":true,"sender":{"key":"tomash.brechko@gmail.com","avatar":null},"body":"Hi David,\n\nOn Thu, Jul 19, 2007 at 13:31:50 +0200, David Kastrup wrote:\n> core.excludefile: .\n\nReally nice idea to give directories 'DIR/.' name.  I'm sure there are\nseveral other ways to implement your proposal.  But why to put in in\nGit itself?  Decomposition and abstraction principle tells me that\nthis should go to some other place.\n\nPlease consider this: I myself use Git to track my own local projects,\nand for this usage you proposal have no value for me, i.e. as a\n_Source_ Code Management system Git is rather complete.  But I also\ntrack /etc and ~/ in Git, and for this I'd love to have directories,\npermissions, ownership, other attributes, to be tracked.  I have Perl\nscript wrapping Git that allows me to filter tracked paths by full\nregexps instead of Git's file globs, and also to filter out too big\nfiles assuming that they are binary anyway.  Most people solving the\nsame problem moved further and implemented tools to store part of file\nsystem state (permissions and ownership) in a textual representation,\nto track that in Git.  I'm sure you've seen such posts in the list.\nAnd my point is that rather than building the support for all of it\ninto core Git, and then implementing sophisticated configuration to\ndisable parts of it, wouldn't it be better to have a separate tools\northogonal to Git itself?\n\nAt the extreme case (probably not really seriously), consider the\nfollowing design: there are two layers, file system layer, and\ncontents layer.  On checkout file system layer creates (or examines\nexisting) directory tree along with all files and their file system\nstate (permissions, ownership, ACLs, attributes, ...), and then asks\ncontents layer to update the contents.  This way layers are\nindependent, and file system layer may be implemented on top of pure\ncontents tracking.  File system layer may be extended to be made\nparticular OS/FS dependent if some development team wishes so.  Even\nhard links may be supported: since file system layer may deside to\nremember that two paths really reference the same inode\n(i.e. contents), contents layer may be asked to update the data only\nonce with either file name/descriptor.\n\nThis, BTW, is why I think not tracking file attributes when\nversioning, say, /etc, is not a big loss.  When I will move to the new\nsystem, I will mostly be interested in contents diffs of the same\nconfiguration files in /etc.  I will trust their new attributes, and\nwill not want to restore them to what they were on the old system.\n\nSo the essence of my objection is that we should not pollute core Git\nwith file system state tracking more than it's required to know where\nto put the contents to.  Everything else should go elsewhere.\n\nAgain, I'd love to have your proposal be implemented, but only in a\nway that won't interfere with pure SCM's operations.\n\n\n-- \n   Tomash Brechko\n"},{"id":"47856","messageId":"86bqe8a7ql.fsf@lola.quinscape.zz","threadId":"9086","inReplyTo":"86zm1sbpeh.fsf@lola.quinscape.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-19T12:38:42Z","receivedAt":"2007-07-19T12:38:42Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"David Kastrup <dak@gnu.org> writes:\n\n> Although it would be natural to have\n> core.adddirs: false\n> be equivalent to\n> core.excludefile: .\n>\n> And so it might be possible to actually not need a separate\n> core.adddirs option at all, technically.\n\nTo followup on myself here:\n\nA project such as the linux kernel which presumably does not want to\nhave directories tracked will put the single pattern\n.\ninto its top-level .gitignore file.  That is all.  At least if it does\nnot confuse current versions of git to do ugly things.\n\nA separate option core.adddirs is still necessary because\nman gitignore\nstates:\n\n       When deciding whether to ignore a path, git normally  checks  gitignore\n       patterns from multiple sources, with the following order of precedence:\n\n       ·  Patterns read from the file specified by the configuration  variable\n          core.excludesfile.\n\n       ·  Patterns read from $GIT_DIR/info/exclude.\n\n       ·  Patterns  read  from  a .gitignore file in the same directory as the\n          path, or in any parent directory, ordered from the deepest such file\n          to  a file in the root of the repository. These patterns match rela‐\n          tive to the location of the  .gitignore  file.  A  project  normally\n          includes  such  .gitignore  files in its repository, containing pat‐\n          terns for files generated as part of the project build.\n\nThe priority for \"core.adddirs\", however, should be below that so that\npreferences set in the repository's .gitignore files take precedence.\nSo core.excludesfile seems to be the wrong place.\n\nA project with the policy of always tracking directories would place\n!.\ninto its top-level .gitignore file.\n\n-- \nDavid Kastrup\n"},{"id":"47861","messageId":"863azka7d4.fsf@lola.quinscape.zz","threadId":"9086","inReplyTo":"20070719123214.GB4929@moonlight.home","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-19T12:46:47Z","receivedAt":"2007-07-19T12:46:47Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Tomash Brechko <tomash.brechko@gmail.com> writes:\n\n> Hi David,\n>\n> On Thu, Jul 19, 2007 at 13:31:50 +0200, David Kastrup wrote:\n>> core.excludefile: .\n>\n> Really nice idea to give directories 'DIR/.' name.  I'm sure there are\n> several other ways to implement your proposal.  But why to put in in\n> Git itself?  Decomposition and abstraction principle tells me that\n> this should go to some other place.\n\nBecause of a fundamental law of computation: information maintained in\ntwo separate places will get out of synch eventually.\n\n> Please consider this: I myself use Git to track my own local\n> projects, and for this usage you proposal have no value for me,\n> i.e. as a _Source_ Code Management system Git is rather complete.\n> But I also track /etc and ~/ in Git, and for this I'd love to have\n> directories, permissions, ownership, other attributes, to be\n> tracked.  I have Perl script wrapping Git that allows me to filter\n> tracked paths by full regexps instead of Git's file globs, and also\n> to filter out too big files assuming that they are binary anyway.\n\nLook, git _tracks_ contents.  Your permissions managements needs to be\ntold explicitly when and how things change.  So you end up with git\n_tracking_ material and your permissions/directory management needing\nthe level of manual handholding Subversion demands.\n\n> And my point is that rather than building the support for all of it\n> into core Git, and then implementing sophisticated configuration to\n> disable parts of it, wouldn't it be better to have a separate tools\n> orthogonal to Git itself?\n\nAnd my personal answer to that is \"no\".  We don't want orthogonality\nfor intimately related things, because it forces us to work the\n\"orthogonal\" things in lockstep.  And if you force git to operate in\nlockstep with manual explicit tracking, then git becomes useless for\ntracking stuff automatically.\n\n> So the essence of my objection is that we should not pollute core\n> Git with file system state tracking more than it's required to know\n> where to put the contents to.  Everything else should go elsewhere.\n>\n> Again, I'd love to have your proposal be implemented, but only in a\n> way that won't interfere with pure SCM's operations.\n\nTell git to ignore \".\" and it won't \"interfere\".\n\n-- \nDavid Kastrup\n"},{"id":"47862","messageId":"86abts8r6z.fsf@lola.quinscape.zz","threadId":"9086","inReplyTo":"86bqe8a7ql.fsf@lola.quinscape.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-19T13:21:24Z","receivedAt":"2007-07-19T13:21:24Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"David Kastrup <dak@gnu.org> writes:\n\n> David Kastrup <dak@gnu.org> writes:\n>\n>> Although it would be natural to have\n>> core.adddirs: false\n>> be equivalent to\n>> core.excludefile: .\n>>\n>> And so it might be possible to actually not need a separate\n>> core.adddirs option at all, technically.\n>\n> To followup on myself here:\n>\n> A project such as the linux kernel which presumably does not want to\n> have directories tracked will put the single pattern\n> .\n> into its top-level .gitignore file.  That is all.  At least if it does\n> not confuse current versions of git to do ugly things.\n\nAnother followup: it doesn't.  I placed a single line\n.\ninto a .gitignore file.  This did not cause git to ignore the contents\nof ., and even\ngit-add .\nworked as previously, namely adding the contents of the current\ndirectory and subdirectories to the index.\n\nIn short: the gitignore idea for policing directory management is\nperfectly upwards-compatible with current versions of git.\n\n-- \nDavid Kastrup\n"},{"id":"47867","messageId":"7FE87F7A-53AD-4B92-8F33-ECDFAE6A7EFB@silverinsanity.com","threadId":"9086","inReplyTo":"86wswwa8ej.fsf@lola.quinscape.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Brian Gernhardt","fromEmail":"benji@silverinsanity.com","sentAt":"2007-07-19T14:44:12Z","receivedAt":"2007-07-19T14:44:12Z","isPatch":true,"sender":{"key":"benji@silverinsanity.com","avatar":"https://gravatar.com/avatar/e06c101dbc25c68114d859b4a9ec7cf8a2c52fd2b0270ef0eac0e2e63ff22311?d=mp&s=160"},"body":"\nOn Jul 19, 2007, at 8:24 AM, David Kastrup wrote:\n\n> I think that the placeholder name should rather be \".\".\n\nFor what it's worth, the more this gets discussed, the more I think  \nyour idea is a good one.\n\n~~ Brian\n"},{"id":"47869","messageId":"FA38709A-7C68-4D66-BA26-B5ED49DFA85A@silverinsanity.com","threadId":"9086","inReplyTo":"863azk78yp.fsf@lola.quinscape.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Brian Gernhardt","fromEmail":"benji@silverinsanity.com","sentAt":"2007-07-19T15:08:11Z","receivedAt":"2007-07-19T15:08:11Z","isPatch":true,"sender":{"key":"benji@silverinsanity.com","avatar":"https://gravatar.com/avatar/e06c101dbc25c68114d859b4a9ec7cf8a2c52fd2b0270ef0eac0e2e63ff22311?d=mp&s=160"},"body":"\nOn Jul 19, 2007, at 10:40 AM, David Kastrup wrote:\n\n> Have you synched with the current state of my proposals posted to the\n> mailing list before posting this note?  Perhaps your concerns have\n> already been addressed in them.\n\nMail.app split the thread into two or three pieces.  I wrote this  \nafter reading the first part, but had missed the rest.  I very much  \nlike the proposals of separating trees from directories and the \".\"  \nentries.\n\nMy apologies for the wasted bandwidth arguing for things that had  \nalready been decided.\n\n~~ Brian\n"},{"id":"47870","messageId":"86ejj45s8i.fsf@lola.quinscape.zz","threadId":"9086","inReplyTo":"FA38709A-7C68-4D66-BA26-B5ED49DFA85A@silverinsanity.com","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-19T15:27:09Z","receivedAt":"2007-07-19T15:27:09Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Brian Gernhardt <benji@silverinsanity.com> writes:\n\n> On Jul 19, 2007, at 10:40 AM, David Kastrup wrote:\n>\n>> Have you synched with the current state of my proposals posted to the\n>> mailing list before posting this note?  Perhaps your concerns have\n>> already been addressed in them.\n>\n> Mail.app split the thread into two or three pieces.  I wrote this\n> after reading the first part, but had missed the rest.  I very much\n> like the proposals of separating trees from directories and the \".\"\n> entries.\n>\n> My apologies for the wasted bandwidth arguing for things that had\n> already been decided.\n\n\"decided\"!  Now that's a strong word for my wild brainstorming if I\never heard one, in particular considering my well-near non-existent\nrecord of contributions and popularity here: most of the recent\n\"discussion\" has been me following up on myself.\n\nAnyway, thanks for the heads-up: very much appreciated.  I'll probably\nbadly need it when people in Pacific Standard Time get to work again\nand tear me to pieces.\n\n-- \nDavid Kastrup\n"},{"id":"47871","messageId":"Pine.LNX.4.64.0707191642270.14781@racer.site","threadId":"9086","inReplyTo":"7FE87F7A-53AD-4B92-8F33-ECDFAE6A7EFB@silverinsanity.com","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-19T15:43:30Z","receivedAt":"2007-07-19T15:43:30Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 19 Jul 2007, Brian Gernhardt wrote:\n\n> \n> On Jul 19, 2007, at 8:24 AM, David Kastrup wrote:\n> \n> > I think that the placeholder name should rather be \".\".\n> \n> For what it's worth, the more this gets discussed, the more I think your \n> idea is a good one.\n\nI do not like it at all. \".\" already has a very special meaning.  It is a \n_directory_, no place holder.\n\nMore and more I get the impression that this thread is just not worth it.  \nThe problem was solved long ago, and all that is talked about here is how \nto complicate things.\n\nUnhappy,\nDscho\n"},{"id":"47872","messageId":"9F03934D-5695-48C9-8F11-E2D94E0B6FC5@silverinsanity.com","threadId":"9086","inReplyTo":"86ejj45s8i.fsf@lola.quinscape.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Brian Gernhardt","fromEmail":"benji@silverinsanity.com","sentAt":"2007-07-19T15:50:41Z","receivedAt":"2007-07-19T15:50:41Z","isPatch":true,"sender":{"key":"benji@silverinsanity.com","avatar":"https://gravatar.com/avatar/e06c101dbc25c68114d859b4a9ec7cf8a2c52fd2b0270ef0eac0e2e63ff22311?d=mp&s=160"},"body":"\nOn Jul 19, 2007, at 11:27 AM, David Kastrup wrote:\n\n> Brian Gernhardt <benji@silverinsanity.com> writes:\n>\n>> My apologies for the wasted bandwidth arguing for things that had\n>> already been decided.\n>\n> \"decided\"!  Now that's a strong word for my wild brainstorming if I\n> ever heard one, in particular considering my well-near non-existent\n> record of contributions and popularity here: most of the recent\n> \"discussion\" has been me following up on myself.\n\nMeh.  I suppose I meant \"talked about\" or \"brought up\" here.  Trying  \nto be quick and terse, and ended up losing meaning like usual.\n\n~~ Brian\n"},{"id":"47873","messageId":"6C96EBA9-CDCE-40EA-B0EC-F9195DBE83DB@silverinsanity.com","threadId":"9086","inReplyTo":"Pine.LNX.4.64.0707191642270.14781@racer.site","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Brian Gernhardt","fromEmail":"benji@silverinsanity.com","sentAt":"2007-07-19T16:06:27Z","receivedAt":"2007-07-19T16:06:27Z","isPatch":true,"sender":{"key":"benji@silverinsanity.com","avatar":"https://gravatar.com/avatar/e06c101dbc25c68114d859b4a9ec7cf8a2c52fd2b0270ef0eac0e2e63ff22311?d=mp&s=160"},"body":"\nOn Jul 19, 2007, at 11:43 AM, Johannes Schindelin wrote:\n\n> I do not like it at all. \".\" already has a very special meaning.   \n> It is a\n> _directory_, no place holder.\n\nAnd we're talking about using it to describe the directory.\n\n> More and more I get the impression that this thread is just not  \n> worth it.\n> The problem was solved long ago, and all that is talked about here  \n> is how\n> to complicate things.\n\nBy solved, you mean ignored?  There is no reason for git not to track  \nempty directories other than \"we don't like it\".\n\nSome projects I work on require certain directories to exist in order  \nto run properly, but tend to occasionally do things like delete all  \nfiles in this required directory.  So far, it hasn't been an issue  \nbecause I'm working solo and using git just to bar against  \nstupidity.  Git's policy of \"don't touch things I don't know about\"  \nworks.  But if I ever had to have someone clone it, they'd need to re- \ncreate the directories.  In this case, empty directories are part of  \nthe content I care about.  Yes, I could have a script do it, but  \nthat's a work around, not a solution.\n\nIn another case, I'm using creating a git repository out of source  \nthat is distributed as occasional tarballs with patches in between.   \nGit's lack of ability to track the empty directories means that I can  \nNOT re-create appropriate tarballs for the states distributed only as  \npatches.  Yes, I could add placeholder files, but then the state is  \nnot identical.\n\nThere are use cases for tracking directories.  I'll agree that it  \nshouldn't be used for every source tree.  But there are cases where  \nit is useful and there's no reason to simply forbid it.\n\n~~ Brian\n"},{"id":"47874","messageId":"Pine.LNX.4.64.0707191715000.14781@racer.site","threadId":"9086","inReplyTo":"6C96EBA9-CDCE-40EA-B0EC-F9195DBE83DB@silverinsanity.com","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-19T16:17:52Z","receivedAt":"2007-07-19T16:17:52Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 19 Jul 2007, Brian Gernhardt wrote:\n\n> On Jul 19, 2007, at 11:43 AM, Johannes Schindelin wrote:\n> \n> > I do not like it at all. \".\" already has a very special meaning.  It \n> > is a _directory_, no place holder.\n> \n> And we're talking about using it to describe the directory.\n> \n> > More and more I get the impression that this thread is just not worth \n> > it. The problem was solved long ago, and all that is talked about here \n> > is how to complicate things.\n> \n> By solved, you mean ignored?  There is no reason for git not to track \n> empty directories other than \"we don't like it\".\n\nNo, no, no, no, no!\n\nYou are really trying to annoy me, right?\n\nHere a short description, which you should read until you understand it \nand then leave me alone:\n\nTo add a directory to the tracked content, you have to _mark_ it as \ntracked.  So that when you remove the _real_ content of the directory, Git \nwill not remove it.\n\nAlas, we already have such a marker.  It is called \".gitignore\", and has \nbeen ignored by _you_.  There is _nothing_ wrong, from a technical \nstandpoint, to call this marker \".gitignore\", and it is _also_ not wrong \nto put this marker into the file system _in addition_ to the index.\n\nSo go and add your directories via that marker, and _be done with it_.\n"},{"id":"47876","messageId":"vpqsl7kiczz.fsf@bauges.imag.fr","threadId":"9086","inReplyTo":"Pine.LNX.4.64.0707191642270.14781@racer.site","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2007-07-19T16:17:52Z","receivedAt":"2007-07-19T16:17:52Z","isPatch":true,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> More and more I get the impression that this thread is just not worth it.  \n> The problem was solved long ago, and all that is talked about here is how \n> to complicate things.\n\nThe problem was not _solved_, it was _worked around_.\n\nAdding a .gitignore or whatever other file to mean \"the directory\nexists\" is clearly a good workaround, but still, you have to use\n\"git-add $dir/.gitignore\" where you really _mean_ \"git-add $dir/\". I\ncan see no reason for the presence of this .gitignore file other than\n\"err, I've put it here because git doesn't manage empty directories\".\n\nThe fact that you need a FAQ entry for that actually shows there is a\nproblem. You don't have a FAQ for \"Q: How to I add a file? A: Use\ngit-add file\", you shouldn't need a FAQ for \"How do I add a\ndirectory\", it should just work as expected.\n\nYou claim it \"solves\" the problem, but have you ever used an importer\nlike git-svn on a project that uses empty directories as placeholders\n(I do have this problem in daily life because my colleagues still use\nSVN)? What is the meaning of this .gitignore file the day you export\nit to anything outside git?\n\nIf you ignore problems because they have a workaround, then even CVS\ncan be usable. People have been working around CVS's problems for\nyears, and many people are happy with CVS because they didn't realise\nthat solving problems is better than working around them (See the\nOpenCVS project ...). Fortunately, git doesn't have as many problems\nto work around as CVS ;-).\n\nI'm happy with the answer \"it should be done, but not by me, send a\npatch\", and I can't really complain myself since I did not send a\npatch, but here, you're complaining about someone who actually starts\nvolunteering to solve the problem, which I can't agree with.\n\n-- \nMatthieu\n"},{"id":"47875","messageId":"86sl7k4b57.fsf@lola.quinscape.zz","threadId":"9086","inReplyTo":"Pine.LNX.4.64.0707191642270.14781@racer.site","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-19T16:21:40Z","receivedAt":"2007-07-19T16:21:40Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> On Thu, 19 Jul 2007, Brian Gernhardt wrote:\n>\n>> \n>> On Jul 19, 2007, at 8:24 AM, David Kastrup wrote:\n>> \n>> > I think that the placeholder name should rather be \".\".\n>> \n>> For what it's worth, the more this gets discussed, the more I think your \n>> idea is a good one.\n>\n> I do not like it at all. \".\" already has a very special meaning.  It is a \n> _directory_, no place holder.\n\nAnd this is what it will be under my scheme: a directory.  It is just\nthat \"directory\" is differentiated from a \"tree\".  Both are tracked in\nthe repository (directory tracking is optional), and there is no such\nthing as an empty tree, a tree being defined by its contents and\nnothing else, as previously.  A \"directory\" has no contents, but only\nexistence in index and repository.  A \"tree\" only exists in the\nrepository, not in index or work directory.  It is mapped to physical\ndirectories in the work directory.  If no corresponding \"directory\"\nexists in index and/or repository, the work directories are created\nand deleted on the fly as before in order to represent the state of\nthe \"tree\" in the repository.  So here are the concepts:\n\nentity     working directory        index           repository\n--------------------------------------------------------------\nfile       mapped to files          file            [blob]\ndir        mapped to dir existence  dir             [dir]\ntree       mapped to dir tree       unrepresented   [tree] (non-empty container)\n\n> More and more I get the impression that this thread is just not\n> worth it.  The problem was solved long ago, and all that is talked\n> about here is how to complicate things.\n\nI disagree on both accounts: that the problem has been solved (the\nexistence of a workaround involving constant manual intervention is\nnot a solution for me), and that my proposal will constitute a\ncomplication to the user.\n\nFor projects setting a \".\" into the top level .gitignore, nothing at\nall will change, even when \"core.adddirs: true\" will become the\ndefault at some point of time.  Once this is the default, new users\nwith new projects will not notice anything surprising, at least until\nthe time that they pull from somebody with a repository with different\nnon-explicit conventions.\n\nThis is something which may still require thought in order to result\nin the least complicated handling of cooperation.  But with regard to\nthe internals itself, I don't see that there is too much non-obvious\ncomplexity involved here, and the framework appears very consistent,\nlogical, and compatible with git's ideas to me.\n\n-- \nDavid Kastrup\n"},{"id":"47877","messageId":"86ir8g4au7.fsf@lola.quinscape.zz","threadId":"9086","inReplyTo":"Pine.LNX.4.64.0707191715000.14781@racer.site","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-19T16:28:16Z","receivedAt":"2007-07-19T16:28:16Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> Here a short description, which you should read until you understand\n> it and then leave me alone:\n>\n> To add a directory to the tracked content, you have to _mark_ it as\n> tracked.  So that when you remove the _real_ content of the\n> directory, Git will not remove it.\n\nCorrect.  That is what my proposal is about.\n\n> Alas, we already have such a marker.  It is called \".gitignore\", and\n> has been ignored by _you_.  There is _nothing_ wrong, from a\n> technical standpoint, to call this marker \".gitignore\", and it is\n> _also_ not wrong to put this marker into the file system _in\n> addition_ to the index.\n\nUh, then the directories are no longer empty.\n\n> So go and add your directories via that marker, and _be done with\n> it_.\n\nBut one is not done before running\n\nfind -name .gitignore -delete\n\nand then the next recursive add will remove the .gitignore \"markers\".\nThe idea of \".\" is to have a marker that does _not_ appear in the work\ndirectory.\n\n-- \nDavid Kastrup\n"},{"id":"47878","messageId":"CC669745-4434-478E-9A24-E474071578C6@silverinsanity.com","threadId":"9086","inReplyTo":"Pine.LNX.4.64.0707191715000.14781@racer.site","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Brian Gernhardt","fromEmail":"benji@silverinsanity.com","sentAt":"2007-07-19T16:34:12Z","receivedAt":"2007-07-19T16:34:12Z","isPatch":true,"sender":{"key":"benji@silverinsanity.com","avatar":"https://gravatar.com/avatar/e06c101dbc25c68114d859b4a9ec7cf8a2c52fd2b0270ef0eac0e2e63ff22311?d=mp&s=160"},"body":"\nOn Jul 19, 2007, at 12:17 PM, Johannes Schindelin wrote:\n\n> Alas, we already have such a marker.  It is called \".gitignore\",  \n> and has\n> been ignored by _you_.  There is _nothing_ wrong, from a technical\n> standpoint, to call this marker \".gitignore\", and it is _also_ not  \n> wrong\n> to put this marker into the file system _in addition_ to the index.\n>\n> So go and add your directories via that marker, and _be done with it_.\n\nBut this alters the content of the directory away from what I want it  \nto be, namely empty.  You aren't addressing the concept of tracking  \nan empty directory, you're just saying you won't do it.\n\n~~ Brian\n"},{"id":"47880","messageId":"Pine.LNX.4.64.0707191829530.14781@racer.site","threadId":"9086","inReplyTo":"CC669745-4434-478E-9A24-E474071578C6@silverinsanity.com","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-19T17:30:31Z","receivedAt":"2007-07-19T17:30:31Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 19 Jul 2007, Brian Gernhardt wrote:\n\n> On Jul 19, 2007, at 12:17 PM, Johannes Schindelin wrote:\n> \n> > Alas, we already have such a marker.  It is called \".gitignore\", and has\n> > been ignored by _you_.  There is _nothing_ wrong, from a technical\n> > standpoint, to call this marker \".gitignore\", and it is _also_ not wrong\n> > to put this marker into the file system _in addition_ to the index.\n> > \n> > So go and add your directories via that marker, and _be done with it_.\n> \n> But this alters the content of the directory away from what I want it to be,\n> namely empty.  You aren't addressing the concept of tracking an empty\n> directory, you're just saying you won't do it.\n\nOMG last time I checked, my _empty_ directory contained \".\" and \"..\".  \nWhat do I do now?\n\nReally,\nDscho\n"},{"id":"47884","messageId":"867iow2smh.fsf@lola.quinscape.zz","threadId":"9086","inReplyTo":"CC669745-4434-478E-9A24-E474071578C6@silverinsanity.com","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-19T17:47:02Z","receivedAt":"2007-07-19T17:47:02Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> Hi,\n>\n> On Thu, 19 Jul 2007, Brian Gernhardt wrote:\n>\n>> On Jul 19, 2007, at 12:17 PM, Johannes Schindelin wrote:\n>> \n>> > Alas, we already have such a marker.  It is called \".gitignore\", and has\n>> > been ignored by _you_.  There is _nothing_ wrong, from a technical\n>> > standpoint, to call this marker \".gitignore\", and it is _also_ not wrong\n>> > to put this marker into the file system _in addition_ to the index.\n>> > \n>> > So go and add your directories via that marker, and _be done with it_.\n>> \n>> But this alters the content of the directory away from what I want it to be,\n>> namely empty.  You aren't addressing the concept of tracking an empty\n>> directory, you're just saying you won't do it.\n>\n> OMG last time I checked, my _empty_ directory contained \".\" and \"..\".  \n> What do I do now?\n\nIf you have a suitable Solaris system, you could try\n\nsudo unlink .\nsudo unlink ..\n\nand have a chance that this will work until the next file system\ncheck.\n\nI don't think that adding tracking of \"..\" would be easy to implement\nin git, but I seem to remember that somebody recently proposed a plan\nof at least tracking \".\" which would seem better than nothing and\npossibly more useful than \"sudo unlink .\".\n\nAll the best,\n\n-- \nDavid Kastrup\n"},{"id":"47909","messageId":"7vk5sw2ba7.fsf@assigned-by-dhcp.cox.net","threadId":"9086","inReplyTo":"FA38709A-7C68-4D66-BA26-B5ED49DFA85A@silverinsanity.com","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-07-20T00:01:36Z","receivedAt":"2007-07-20T00:01:36Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Brian Gernhardt <benji@silverinsanity.com> writes:\n\n> My apologies for the wasted bandwidth arguing for things that had\n> already been decided.\n\nSorry, who decided what?\n"},{"id":"47912","messageId":"alpine.LFD.0.999.0707191706120.27249@woody.linux-foundation.org","threadId":"9086","inReplyTo":"7vk5sw2ba7.fsf@assigned-by-dhcp.cox.net","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-20T00:15:28Z","receivedAt":"2007-07-20T00:15:28Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 19 Jul 2007, Junio C Hamano wrote:\n>\n> Brian Gernhardt <benji@silverinsanity.com> writes:\n> \n> > My apologies for the wasted bandwidth arguing for things that had\n> > already been decided.\n> \n> Sorry, who decided what?\n\nI think people who didn't know how the world works decided that \ndirectories that were added manually as directories would stay as \ndirectories even after the last file was removed.\n\nThat's physically impossible with the git data-structures (since there is \nno way of saving \"this directory was added empty\" in the tree structures, \nnor any point to it), so I think it's just insane rambling.\n\nI dunno. I think empty directories are worth supporting, mainly to be able \nto capture other SCM's notion of what _they_ track, but quite frankly, the \nlevel of discussion about them hasn't been exactly inspiring. It seems to \nbe more about \"this is what we'd like to see, without really having a \nreason for it, nor necessarily understanding what we're talking about\" \nthan \"this is realistic and useful and here are patches\".\n\nI *do* think that it's a very valid argument that if you import something \nfrom SVN that has an empty directory, the git import should show that.  \n\nThat's about the only valid argument I've ever seen for them, though, and \nI think that's totally irrelevant to such issues as to whether \"git rm \nfile/in/directory\" should remove the directory(*) from being tracked by \ngit when the file goes away or not.\n\n\t\t\tLinus\n \n(*) And, for anybody confused about the issue, the answer to the latter \nquestion is an emphatic: \"Yes it should, live with it, and if you want the \ndirectory back, you had better add it back as an empty directory\"\n"},{"id":"47916","messageId":"alpine.LFD.0.999.0707191726510.27249@woody.linux-foundation.org","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707191706120.27249@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-20T00:33:29Z","receivedAt":"2007-07-20T00:33:29Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 19 Jul 2007, Linus Torvalds wrote:\n> \n> That's physically impossible with the git data-structures (since there is \n> no way of saving \"this directory was added empty\" in the tree structures, \n> nor any point to it), so I think it's just insane rambling.\n\nOf course, it's physically *possible* to have a tree that contains two \nentries for the same name: first the \"empty tree\" and then the \"real \ntree\", and yeah, in theory you could track things that way.\n\nSo I guess the \"physically impossible\" was a bit strong. You'd have to \nhave a totally insane format, and you'd have to violate deeply seated \nrules about what trees look like (and the index too, for that matter: we'd \nhave to do the same for the index, and keep the S_IFDIR entry alive \ndespite having other entries that are children of it), but it's \n*possible*.\n\nIt's just a really bad idea.\n\nSo to be sane, when you add files, the empty directory entry has to go \naway. Otherwise you could have two very different trees that encode the \nsame *content* (just with different ways of getting there - depending on \nwhether you have a history with empty trees or not), and that's very much \nagainst the philosophy of git, and breaks some fundamental rules (like the \nfact that \"same content == same SHA1\").\n\nIn fact, that may be the best way to explain why it's *not* an option to \nhave \"empty trees remain empty trees if we remove the last file from \nthem\": git fundamnetally tracks \"content snapshots\", and anything that \nimplies the content containing any history is against the rules.\n\nSo the whole notion of \"remembering\" whether a directory was added \nexplicitly as an empty directory or not is just not a sensible concept in \ngit. \n\n\t\tLinus\n"},{"id":"47919","messageId":"7vir8f24o2.fsf@assigned-by-dhcp.cox.net","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707191726510.27249@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-07-20T02:24:29Z","receivedAt":"2007-07-20T02:24:29Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> So the whole notion of \"remembering\" whether a directory was added \n> explicitly as an empty directory or not is just not a sensible concept in \n> git. \n\nThat is true if it is implemented as David suggested, to have a\nphony \".\" entry in the tree object itself.  The object name of\nsuch a tree (when it contains blobs and trees underneath) will\nbe different from a tree that contains the same set of blobs and\ntrees.  It would destroy the fundamental concepts of git.\n\nBut you _could_ treat that \"should-be-kept-even-when-empty\"-ness\njust like we treat executable bit on blobs, I think.\n\nWhen blobs with the same contents but of different type (REG vs\nLNK) and regular file with or without executable bit are entered\nin git, they all get the same SHA-1 but we can still tell them\napart because the index and the tree entry have mode bits.  So\nhypothetically, you could introduce \"sticky\" directory in tree\nentries to mark \"this will not go away when emptied\".\n\nIn a 'tree' object, they might appear as:\n\n        40000 ordinary-directory '\\0' 20-byte SHA-1\n        41000 directory-dontremove-even-if-empty '\\0' 20-byte SHA-1\n\nIn 'index', as your \"I'm soft\" patch, we do not have to add\nnonsticky kind of tree nodes, but for \"empty\" ones, we can add:\n\n\t041000 directory-dontremove-even-if-empty '\\0' 20-byte SHA-1\n\nin the index and (unlike your patch) keep it there even after a\nblob or a tree is added underneath it.\n\nThe \"sticky\" bit on such a directory would have to obey the\nusual rule of 3-way merge, which would be a huge change to do\nso, but I do no see there is anything fundamental that prevents\nyou from doing this.  Other than the fact that probably no git\nlong timer is interested in spending time on such a feature,\nthat is.\n\nObviously, this \"sticky\" bit will cascade up and make your\notherwise equivalent parent tree's different, but I think that\nis just as a sane behaviour as two trees that contain the same\nblob with only executable-bit differences have different names.\n\nThis will involve a lot of changes, so I would not recommend\nanybody doing so, though.\n"},{"id":"47920","messageId":"alpine.LFD.0.999.0707191930030.27249@woody.linux-foundation.org","threadId":"9086","inReplyTo":"7vir8f24o2.fsf@assigned-by-dhcp.cox.net","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-20T02:31:44Z","receivedAt":"2007-07-20T02:31:44Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 19 Jul 2007, Junio C Hamano wrote:\n> \n> But you _could_ treat that \"should-be-kept-even-when-empty\"-ness\n> just like we treat executable bit on blobs, I think.\n\nTrue. Or you could make it a path attribute and/or a per-repository \ndecision, so that while the data wouldn't necessarily be in the database \nitself, the user could specify the behaviour he wanted.\n\n> This will involve a lot of changes, so I would not recommend\n> anybody doing so, though.\n\nAgreed. The upside just isn't there.\n\n\t\tLinus\n"},{"id":"47927","messageId":"85644fvdrn.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707191726510.27249@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-20T05:35:08Z","receivedAt":"2007-07-20T05:35:08Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Thu, 19 Jul 2007, Linus Torvalds wrote:\n>> \n>> That's physically impossible with the git data-structures (since\n>> there is no way of saving \"this directory was added empty\" in the\n>> tree structures, nor any point to it), so I think it's just insane\n>> rambling.\n>\n> Of course, it's physically *possible* to have a tree that contains\n> two entries for the same name: first the \"empty tree\" and then the\n> \"real tree\", and yeah, in theory you could track things that way.\n>\n> So I guess the \"physically impossible\" was a bit strong. You'd have\n> to have a totally insane format, and you'd have to violate deeply\n> seated rules about what trees look like (and the index too, for that\n> matter: we'd have to do the same for the index, and keep the S_IFDIR\n> entry alive despite having other entries that are children of it),\n> but it's *possible*.\n\nExcuse me?  You don't need a \"totally insane format\".  You need an\nentry \".\" of a new type \"directory\" that can be part of the current\nconcept of a \"tree\".  This new type does _not_ have children.  It is\nnot a container for files.  It would be the thing that would carry\npermissions or other properties if git were to store them.  It can be\nput into .gitignore files like other files.\n\nOne drawback is that adding and removing it alone is not supported\nwith the current git-add and git-remove commands: they would require\nan additional argument \"-d\" like \"ls\" does.\n\nAll of this is a straightforward extension fitting very well the\ncurrent paradigms and also existing file systems and their usage.\n\n> It's just a really bad idea.\n\n> So to be sane, when you add files, the empty directory entry has to\n> go away.\n\nYou really have not followed the discussion at all.  This is not\npossible since otherwise you could not distinguish the cases\n\nmkdir A\ntouch A/B\ngit-add A\ngit-rm A/B\n\nwhere A was added and not removed and should stay and\n\nmkdir A\ntouch A/B\ngit-add A/B\ngit-rm A/B\n\nwhere a single file was added and removed and nothing should stay.\n\n> Otherwise you could have two very different trees that encode the\n> same *content* (just with different ways of getting there -\n> depending on whether you have a history with empty trees or not),\n> and that's very much against the philosophy of git, and breaks some\n> fundamental rules (like the fact that \"same content == same SHA1\").\n\nNo, the content is _different_.  One tree contains a tracked\ndirectory, the other does not.  That means that the trees behave\n_differently_ when you manipulate them, and that means that they are\n_not_ the same tree.\n\n> In fact, that may be the best way to explain why it's *not* an\n> option to have \"empty trees remain empty trees if we remove the last\n> file from them\": git fundamnetally tracks \"content snapshots\", and\n> anything that implies the content containing any history is against\n> the rules.\n>\n> So the whole notion of \"remembering\" whether a directory was added\n> explicitly as an empty directory or not is just not a sensible\n> concept in git.\n\nCertainly.  That is why we instead remember whether or not a directory\nentry \".\" was added or not.  It will be added (unless the defaults and\ngitignore settings ask \".\" to be non-tracked) when git adds the\ncorresponding tree or subtree, and it will get removed when git\nremoves the corresponding tree or subtree.  Emptiness is not a special\ncase, and it can't be.  Currently, the main information associated\nwith \".\" is \"stay around even if tree becomes empty\".\n\nNow you can do\n\n    unlink .\n\nin Solaris and have the name \".\" vanish while the directory still\nworks as a container by other names.\n\nI don't propose that git be able to track this difference, though, and\nI doubt that most file archivers would.\n\nBut git can or cannot ignore files, and in a similar way it can or\ncannot ignore what a directory has more than being an abstract\ncontainer.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"47929","messageId":"851wf3vcxu.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707191726510.27249@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-20T05:53:01Z","receivedAt":"2007-07-20T05:53:01Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n>\n>> So the whole notion of \"remembering\" whether a directory was added \n>> explicitly as an empty directory or not is just not a sensible concept in \n>> git. \n>\n> That is true if it is implemented as David suggested, to have a\n> phony \".\" entry in the tree object itself.  The object name of such\n> a tree (when it contains blobs and trees underneath) will be\n> different from a tree that contains the same set of blobs and trees.\n> It would destroy the fundamental concepts of git.\n\nHow so?\n\n> But you _could_ treat that \"should-be-kept-even-when-empty\"-ness\n> just like we treat executable bit on blobs, I think.\n>\n> When blobs with the same contents but of different type (REG vs LNK)\n> and regular file with or without executable bit are entered in git,\n> they all get the same SHA-1 but we can still tell them apart because\n> the index and the tree entry have mode bits.  So hypothetically, you\n> could introduce \"sticky\" directory in tree entries to mark \"this\n> will not go away when emptied\".\n\nA tree containing files with and without executable bits will show\ndifferent SHA-1 sums.  There is no reason that this should be\ndifferent for a tree containing the conceptual \".\" or not.  I won't\nfight for a specific implementation but if I am going to implement\nthis (and the current lack of enthusiasm points to that) I will not go\nand duplicate the entire ignore/add/rm/index/repository machinery in\norder to have a bit rather than an actual \".\" directory entry.\n\nMost Unix file systems have an honest, physical, down-to-Earth\ndirectory entry \".\" even on disk because it _simplifies_ matters, even\nthough one could special-case \".\" all throughout and make do without a\nphysical entry in theory.\n\nAnd, as I explained, \".\" lends itself perfectly to the gitignore\nmachinery in order to policy projects to track or not track\ndirectories.\n\n> In a 'tree' object, they might appear as:\n>\n>         40000 ordinary-directory '\\0' 20-byte SHA-1\n>         41000 directory-dontremove-even-if-empty '\\0' 20-byte SHA-1\n>\n> In 'index', as your \"I'm soft\" patch, we do not have to add\n> nonsticky kind of tree nodes,\n\nIt does not work, since then you can't distinguish\n\nmkdir A\ntouch B\ngit-add A/B\n\nfrom\n\nmkdir A\ntouch B\ngit-add A\n\nIt is very clear that git-rm A/B _mustn't_ leave an empty directory in\nthe first case, and _must_ leave an empty directory in the second case\n_if_ and only if one tracks directories.\n\n> Obviously, this \"sticky\" bit will cascade up and make your otherwise\n> equivalent parent tree's different,\n\nNo, it must not \"cascade up\".  After\n\nmkdir -p A/B\ntouch A/B/C\ngit-add A/B\ngit-rm A/B\n\nthere must be nothing tracked by git.  The \"sticky\" bit does not\n\"cascade up\".  Its upward effect is only changing the SHA-1 of the\ntree, like any change below does.\n\n> This will involve a lot of changes, so I would not recommend anybody\n> doing so, though.\n\nNeither would I.  Why people want to complicate the code base\neverywhere by avoiding to treat \".\" like a legitimate entry (as Unix\nfile systems do for a _reason_) is simply a miracle to me.\n\nThe framework is pretty much _there_.  There is no point in not making\nuse of it and duplicating the whole machinery because we want a \"bit\nset\" implementation instead of a file name.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"47932","messageId":"85wswvty8u.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707191930030.27249@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-20T05:55:45Z","receivedAt":"2007-07-20T05:55:45Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Thu, 19 Jul 2007, Junio C Hamano wrote:\n>> \n>> But you _could_ treat that \"should-be-kept-even-when-empty\"-ness\n>> just like we treat executable bit on blobs, I think.\n>\n> True. Or you could make it a path attribute and/or a per-repository\n> decision, so that while the data wouldn't necessarily be in the\n> database itself, the user could specify the behaviour he wanted.\n\nNo, one can't.  Once can decide per repository whether one wants to\npermit this kind of information in.  But if one does, the information\nneeds to there for _every_ tree.  And a \".\" entry is a natural and\nintuitive way to do that.  \".\" has been used as a directory entry for\ndecades in Unix.\n\n>> This will involve a lot of changes, so I would not recommend\n>> anybody doing so, though.\n>\n> Agreed. The upside just isn't there.\n\nIt is a good thing that you did not design the Unix file systems.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"47931","messageId":"85sl7jty43.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"7vir8f24o2.fsf@assigned-by-dhcp.cox.net","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-20T05:58:36Z","receivedAt":"2007-07-20T05:58:36Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n>\n>> So the whole notion of \"remembering\" whether a directory was added \n>> explicitly as an empty directory or not is just not a sensible concept in \n>> git. \n>\n> That is true if it is implemented as David suggested, to have a\n> phony \".\" entry in the tree object itself.\n\nUnix file systems contain a phony \".\" entry in the directory itself,\nand have survived in spite of this.\n\n> The object name of such a tree (when it contains blobs and trees\n> underneath) will be different from a tree that contains the same set\n> of blobs and trees.  It would destroy the fundamental concepts of\n> git.\n\nLike \".\" destroyed the fundamental concepts of Unix filesystems.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"47941","messageId":"200707201029.10358.johan@herland.net","threadId":"9086","inReplyTo":"7vhco28aoq.fsf@assigned-by-dhcp.cox.net","subject":"Re: Empty directories...","fromName":"Johan Herland","fromEmail":"johan@herland.net","sentAt":"2007-07-20T08:29:10Z","receivedAt":"2007-07-20T08:29:10Z","isPatch":false,"sender":{"key":"johan@herland.net","avatar":"https://avatars.githubusercontent.com/u/547031?v=4"},"body":"On Wednesday 18 July 2007, Junio C Hamano wrote:\n> Didn't I say I do not have an objection for somebody who wants\n> to track empty directories, already?  I probably would not do\n> that myself but I do not see a reason to forbid it, either.\n> \n> The right approach to take probably would be to allow entries of\n> mode 040000 in the index.  Traditionally, we allowed only 100644\n> (blobs as regular files) and 120000 (blobs as symlinks).  We\n> recently added 160000 (commit from outer space, aka subproject).\n> \n> And we do that for all directories, not just empty ones.  So if\n> you have fileA, empty/, sub/fileB tracked, your index would\n> probably have these four entries, immediately after read-tree\n> of an existing tree object:\n\nSorry for jumping in late...\n\nWhy do you want to add _all_ directories, and not just the ones we want to \nexplicitly track (independent of whether they're empty or not).\n\nBasically, add a \"--dir\" flag to git-add, git-rm and friends, to tell them \nyou're acting on the directory itself (rather than its (recursive) \ncontents). \"git-add --dir foo\" will add the \"040000 123abc... 0 foo\" to the \nindex/tree whether or not foo is an empty directory. \"git-rm --dir foo\" will \nremove that entry (or fail if it doesn't exist), but _not_ the contents of \nfoo.\n\nSince we're making directory tracking _explicit_, this should all be trivially \nbackward-compatible.\n\n\n...Johan\n\n-- \nJohan Herland, <johan@herland.net>\nwww.herland.net\n"},{"id":"47943","messageId":"86hcnzxy9a.fsf@lola.quinscape.zz","threadId":"9086","inReplyTo":"200707201029.10358.johan@herland.net","subject":"Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-20T08:41:53Z","receivedAt":"2007-07-20T08:41:53Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Johan Herland <johan@herland.net> writes:\n\n> On Wednesday 18 July 2007, Junio C Hamano wrote:\n>> Didn't I say I do not have an objection for somebody who wants\n>> to track empty directories, already?  I probably would not do\n>> that myself but I do not see a reason to forbid it, either.\n>> \n>> The right approach to take probably would be to allow entries of\n>> mode 040000 in the index.  Traditionally, we allowed only 100644\n>> (blobs as regular files) and 120000 (blobs as symlinks).  We\n>> recently added 160000 (commit from outer space, aka subproject).\n>> \n>> And we do that for all directories, not just empty ones.  So if\n>> you have fileA, empty/, sub/fileB tracked, your index would\n>> probably have these four entries, immediately after read-tree\n>> of an existing tree object:\n>\n> Sorry for jumping in late...\n\nIt could have given you a chance to read up on what has already been\ndiscussed.\n\n> Why do you want to add _all_ directories, and not just the ones we\n> want to explicitly track (independent of whether they're empty or\n> not).\n\nBecause the problematic cases are more often than not the _implicit_\ncases.  Do you check a directory tree for empty directories before you\narchive it?  In order to archive every empty directory explicitly?\n\nIf you did that, you could equally maintain a script that manually\ndoes mkdir/rmdir.\n\n> Basically, add a \"--dir\" flag to git-add, git-rm and friends, to\n> tell them you're acting on the directory itself (rather than its\n> (recursive) contents). \"git-add --dir foo\" will add the \"040000\n> 123abc... 0 foo\" to the index/tree whether or not foo is an empty\n> directory. \"git-rm --dir foo\" will remove that entry (or fail if it\n> doesn't exist), but _not_ the contents of foo.\n\nThere is nothing wrong with implementing something like this in\n_addition_ to treating directory entries implicitly.  For example, ls\nhas an option -d which does just that, and even git-ls-files has an\noption --directory.  Heck, I even have\n\nrm --help\nUsage: rm [OPTION]... FILE...\nRemove (unlink) the FILE(s).\n\n  -d, --directory       unlink FILE, even if it is a non-empty directory\n                          (super-user only; this works only if your system\n                           supports `unlink' for nonempty directories)\n[...]\n\nwhich works on just the directory and not on the contents.\n\nSo a --directory option for appropriate commands would be natural for\n_explicit_ manipulation of such entries.\n\nBut the important, the _really_ important thing are the implicit\nbehaviors.  If I have to hassle with every directory myself, I don't\nneed a content tracking system.\n\nThe --directory stuff, in contrast, are things nice to have when the\nframework is in place (and may be even necessary for some direct\nmanual maintenance tasks), but they don't really concern the\nframework.\n\n-- \nDavid Kastrup\n"},{"id":"47947","messageId":"46A08006.4020500@fs.ei.tum.de","threadId":"9086","inReplyTo":"85644fvdrn.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Simon 'corecode' Schubert","fromEmail":"corecode@fs.ei.tum.de","sentAt":"2007-07-20T09:27:34Z","receivedAt":"2007-07-20T09:27:34Z","isPatch":true,"sender":{"key":"corecode@fs.ei.tum.de","avatar":"https://gravatar.com/avatar/eff9dbf0cdac0d1e6a6cd7ed0e50763edcb376b493b5253a35ff167918ad79e1?d=mp&s=160"},"body":"David Kastrup wrote:\n>> Otherwise you could have two very different trees that encode the\n>> same *content* (just with different ways of getting there -\n>> depending on whether you have a history with empty trees or not),\n>> and that's very much against the philosophy of git, and breaks some\n>> fundamental rules (like the fact that \"same content == same SHA1\").\n> \n> No, the content is _different_.  One tree contains a tracked\n> directory, the other does not.  That means that the trees behave\n> _differently_ when you manipulate them, and that means that they are\n> _not_ the same tree.\n\nYou are mistaking things.  Like the executable bit on a file is not content, the fact that a directory should be kept despite being empty is also an *attribute* of the directory.  This is meta-data, not actual data (content).  So no matter how elegant tracking the \".\" entry might be (and I think it is, because it covers a lot of corner cases already), it puts the information at the wrong place.\n\nThat's sad, because otherwise it would be really elegant.\n\ncheers\n  simon\n\n-- \nServe - BSD     +++  RENT this banner advert  +++    ASCII Ribbon   /\"\\\nWork - Mac      +++  space for low €€€ NOW!1  +++      Campaign     \\ /\nParty Enjoy Relax   |   http://dragonflybsd.org      Against  HTML   \\\nDude 2c 2 the max   !   http://golden-apple.biz       Mail + News   / \\\n"},{"id":"47948","messageId":"86k5svwfj4.fsf@lola.quinscape.zz","threadId":"9086","inReplyTo":"46A08006.4020500@fs.ei.tum.de","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-20T10:11:43Z","receivedAt":"2007-07-20T10:11:43Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Simon 'corecode' Schubert <corecode@fs.ei.tum.de> writes:\n\n> David Kastrup wrote:\n>>> Otherwise you could have two very different trees that encode the\n>>> same *content* (just with different ways of getting there -\n>>> depending on whether you have a history with empty trees or not),\n>>> and that's very much against the philosophy of git, and breaks some\n>>> fundamental rules (like the fact that \"same content == same SHA1\").\n>>\n>> No, the content is _different_.  One tree contains a tracked\n>> directory, the other does not.  That means that the trees behave\n>> _differently_ when you manipulate them, and that means that they are\n>> _not_ the same tree.\n>\n> You are mistaking things.\n\nNo, I am redefining them, or rather the view on them.  Subtle\ndifference.\n\n> Like the executable bit on a file is not content, the fact that a\n> directory should be kept despite being empty is also an *attribute*\n> of the directory.  This is meta-data, not actual data (content).\n\nWe need to track it, anyway.  So there is little point in not using\nthe existing infrastructure for handling named entities.\n\n> So no matter how elegant tracking the \".\" entry might be (and I\n> think it is, because it covers a lot of corner cases already), it\n> puts the information at the wrong place.\n\nI don't see that the place is wrong: after all, that is where Unix\nplaces \".\" too, and for good reason.  I was arguing for _separating_\nthe concept of \"directory\" and \"tree\" in the repository.  The tree is\na container entity defined exclusively by its contents (which\ndetermine its hash).  That is how git already does things.  There is\n_no_ connection with the physical existence of a directory: in the\nwork directory, git creates and deletes directories as a _side-effect_\nof storing and removing trees.  But git itself does not track\ndirectories as a physical entity at _all_.  If you had a flat\nfilesystem allowing slashes in filenames, git would get along better\nthan it does now, without ever creating or removing a directory.\nTrees are just a convenient selection and pattern matching mechanism\nfor files as far as git is concerned.  The correspondence to physical\ndirectories in the work directory is a nuisance rather than an asset\nas far as git is concerned.\n\nIn a recent thread here, tags with slashes were supported by\nessentially doing\n\n    mkdir -p \"`dirname $TAG`\"\n    touch $TAG\n\nwhere directory creation is just a side effect of supporting slashes.\nAnd that, if you look closely, is git's current relation with\ndirectories altogether.  The directories in the work file system are\ncreated by git just as a side effect for representing slashes, which\nin turn facilitate a certain manner of pattern matching.\n\nAnd \".\" seems perfectly well suited to bring across the point that\nthere actually is _physical_ existence associated with a directory,\nexistence that remains when the rest of the tree is gone and _makes_ a\ndifference to what the tree is, because it has a _different_\nrepresentation in the work file system.\n\nStoring it as an _attribute_ of the tree is a bad idea, since then the\nsimple rule \"a tree without contents is empty\" needs an exception.\nAnd a tree stops becoming just a container of its contents and all\nsort of new exceptions creep up.\n\nThere are some systems where the difference between directory as a\nfile and directory as a structuring method are more apparent than\nunder Unix (some utilities like rsync differentiate between A/B and\nA/B/ to bring across that difference).\n\nHere is an example for some Emacs function concerned with the concept:\n\n    directory-file-name is a built-in function in `C source code'.\n    (directory-file-name DIRECTORY)\n\n    Returns the file name of the directory named DIRECTORY.\n    This is the name of the file that holds the data for the directory DIRECTORY.\n    This operation exists because a directory is also a file, but its name as\n    a directory is different from its name as a file.\n    In Unix-syntax, this function just removes the final slash.\n    On VMS, given a VMS-syntax directory name such as \"[X.Y]\",\n    it returns a file name such as \"[X]Y.DIR.1\".\n\n    [back]\n\n> That's sad, because otherwise it would be really elegant.\n\nIf something is not elegant because of the angle of view, change the\nview.  And it is not like the different angle has no predecessors or\nno consistency.\n\n-- \nDavid Kastrup\n"},{"id":"47950","messageId":"20070720101942.GA77248@dspnet.fr.eu.org","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707191706120.27249@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Olivier Galibert","fromEmail":"galibert@pobox.com","sentAt":"2007-07-20T10:19:42Z","receivedAt":"2007-07-20T10:19:42Z","isPatch":true,"sender":{"key":"galibert@pobox.com","avatar":null},"body":"On Thu, Jul 19, 2007 at 05:15:28PM -0700, Linus Torvalds wrote:\n> (*) And, for anybody confused about the issue, the answer to the latter \n> question is an emphatic: \"Yes it should, live with it, and if you want the \n> directory back, you had better add it back as an empty directory\"\n\nWouldn't it be perfectly reasonable for git rm to re-add emptied\ndirectories as empty transparently if the appropriate\nflag/configuration is set?  rm is porcelain after all.\n\n  OG.\n"},{"id":"47951","messageId":"200707201220.15114.johan@herland.net","threadId":"9086","inReplyTo":"86hcnzxy9a.fsf@lola.quinscape.zz","subject":"Re: Empty directories...","fromName":"Johan Herland","fromEmail":"johan@herland.net","sentAt":"2007-07-20T10:20:15Z","receivedAt":"2007-07-20T10:20:15Z","isPatch":false,"sender":{"key":"johan@herland.net","avatar":"https://avatars.githubusercontent.com/u/547031?v=4"},"body":"On Friday 20 July 2007, David Kastrup wrote:\n> Johan Herland <johan@herland.net> writes:\n> > Sorry for jumping in late...\n> \n> It could have given you a chance to read up on what has already been\n> discussed.\n\nI have tried to keep on top of the discussion so far.\n\n> > Why do you want to add _all_ directories, and not just the ones we\n> > want to explicitly track (independent of whether they're empty or\n> > not).\n> \n> Because the problematic cases are more often than not the _implicit_\n> cases.  Do you check a directory tree for empty directories before you\n> archive it?  In order to archive every empty directory explicitly?\n\nNo, of course I don't. But then archiving (as in tar) is intended to recreate \nthe \"working copy\" exactly as it was. Git (and other SCMs), however, is only \ninterested in recreating the part of the working copy it explicitly tracks.\n\nGiven the following working copy:\n/\n/tracked/\n/tracked/file\n/tracked/dir/\n/untracked/\n/untracked/file\n/untracked/dir/\n\nand the following commands:\n$ git add tracked\n\n$ git clone\n\nThe cloned result could be any of the following:\n\n(1)\n/\n/tracked/\n/tracked/file\n\nThis is the current behaviour; directories are not tracked at all, but only \nadded as necessary to support files.\n\n(2)\n/\n/tracked/\n/tracked/file\n/tracked/dir/\n/untracked/\n/untracked/dir/\n\ni.e. implicitly tracking _all_ directories. This is what you literally ask \nfor, but I think most would find this unreasonable.\n\n(3)\n/\n/tracked/\n/tracked/file\n/tracked/dir/\n\ni.e. recursively tracking directories (and files). This seems useful, but \nthere is nothing _implicit_ about this.\n\n\nI have a feeling that you're actually arguing for doing (3) by default. What I \nam arguing is to do (1) by default, and (3) if given a suitable command-line \noption (i.e. \"git add --with-dirs tracked\").\n\nNote that this is really an interface question. How these entries are actually \nstored in the repo is a different discussion.\n\n\nFinally, let's look at the case of \"git add tracked/file\" followed by \"git rm \ntracked/file\". I'm arguing that \"tracked/\" should be automatically removed, \nsince I never asked for it to be tracked by git. On the other \nhand, \"git-add --non-recursive tracked\" followed by the above two commands, \nshould of course leave \"tracked/\" in place, since I now actually asked \nexplicitly for the directory to be tracked.\n\nMy point is fundamentally that selectively tracking directories is a more \npowerful concept than just tracking _all_ directories by default. Note that \nif we support selectively tracking directories, tracking _everything_ (like \nyou seem to want) is trivially implemented by _always_ supplying the \nappropriate option to git-add. If we track everything by design, we don't \nhave the option of selectively tracking some directories.\n\n\n> > Basically, add a \"--dir\" flag to git-add, git-rm and friends, to\n> > tell them you're acting on the directory itself (rather than its\n> > (recursive) contents). \"git-add --dir foo\" will add the \"040000\n> > 123abc... 0 foo\" to the index/tree whether or not foo is an empty\n> > directory. \"git-rm --dir foo\" will remove that entry (or fail if it\n> > doesn't exist), but _not_ the contents of foo.\n> \n> There is nothing wrong with implementing something like this in\n> _addition_ to treating directory entries implicitly.\n\nI don't agree. By _selectively_ tracking directories you can implement any \npolicy you want on top of it.\n\n> For example, ls \n> has an option -d which does just that, and even git-ls-files has an\n> option --directory.  Heck, I even have\n\nYes, having commandline options for explicitly specifying directories (and not \ntheir contents) is _exactly_ what I want.\n\n> But the important, the _really_ important thing are the implicit\n> behaviors.  If I have to hassle with every directory myself, I don't\n> need a content tracking system.\n\nI disagree. Just as you have to decide which files to track, you similarly \nshould have to decide which directories to track. Of course, the tools make \nthis easier for you by being able to recursively handle files. In the same \nway they should be able to do the same thing for directories.\n\n\nHave fun!\n\n...Johan\n\n-- \nJohan Herland, <johan@herland.net>\nwww.herland.net\n"},{"id":"47952","messageId":"7vk5svxt1f.fsf@assigned-by-dhcp.cox.net","threadId":"9086","inReplyTo":"46A08006.4020500@fs.ei.tum.de","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-07-20T10:34:36Z","receivedAt":"2007-07-20T10:34:36Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Simon 'corecode' Schubert <corecode@fs.ei.tum.de> writes:\n\n> You are mistaking things.  Like the executable bit on a file\n> is not content, the fact that a directory should be kept\n> despite being empty is also an *attribute* of the directory.\n> This is meta-data, not actual data (content).  So no matter\n> how elegant tracking the \".\" entry might be (and I think it\n> is, because it covers a lot of corner cases already), it puts\n> the information at the wrong place.\n\nActually, I do not think there is absolute right or wrong here.\nThe difference is not that the information is at the \"right\" or\n\"wrong\" place, but one approach places the information at more\nefficient-to-use place than the other.  In that sense, the\nattribute approach _is_ a more elegant solution between the two.\n\nMaking it an attribute has a huge practical advantage.\n\nBy treating executable bit as a piece metadata, we can compare\nthe \"contents\" quickly.  If you \"chmod +x\" a blob without\nchanging anything else, we can detect that fact, because blob\nobject names are equal.  At the philosophical level, you _could_\nargue that the executable-ness is one bit of content and include\nthat in the object name computation for the blob.  There is\nnothing fundamentally wrong about that approach, but that\ndestroys the nice \"cheap comparability\" between blobs that\ndiffer only by executable-ness.\n\nDavid's \".\" in tree is essentially the same argument as treating\nthe executable-ness as one extra bit of content.  The fact that\na particular tree wants to stay even after emptied can be\ntreated as part of contents (thereby reflected in its object\nname).  There is nothing fundamentally wrong there, either.  But\nthat means two trees that contain otherwise identical set of\nblobs and subtrees, but differ only in the behaviour of when\nthey are emptied, would get different object names, hence you\nneed to descend into them to see if they are different.\n\nUsing attribute that is detached from the content itself allows\nyou to hoist that one bit one level up.  By treating\nexecutable-ness not as part of content, we can compare two blobs\nwith different executable bits cheaply.  You can avoid\ndescending into such a tree when comparing it with another tree\nthat is different only by the \"will-stay-when-emptied\"-ness the\nsame way.\n"},{"id":"47954","messageId":"86tzrzuyyy.fsf@lola.quinscape.zz","threadId":"9086","inReplyTo":"200707201220.15114.johan@herland.net","subject":"Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-20T10:54:45Z","receivedAt":"2007-07-20T10:54:45Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Johan Herland <johan@herland.net> writes:\n\n> On Friday 20 July 2007, David Kastrup wrote:\n>> Johan Herland <johan@herland.net> writes:\n>> > Sorry for jumping in late...\n>> \n>> It could have given you a chance to read up on what has already been\n>> discussed.\n>\n> I have tried to keep on top of the discussion so far.\n>\n>> > Why do you want to add _all_ directories, and not just the ones we\n>> > want to explicitly track (independent of whether they're empty or\n>> > not).\n>> \n>> Because the problematic cases are more often than not the\n>> _implicit_ cases.  Do you check a directory tree for empty\n>> directories before you archive it?  In order to archive every empty\n>> directory explicitly?\n>\n> No, of course I don't. But then archiving (as in tar) is intended to\n> recreate the \"working copy\" exactly as it was. Git (and other SCMs),\n> however, is only interested in recreating the part of the working\n> copy it explicitly tracks.\n\nYes, and\ngit-add some-dir\ntells it to track _everything_ inside some-dir.  Which means that the\nincluded files are tracked _implicitly_.  The included directories\n(including some-dir itself) are not.\n\n> Given the following working copy:\n> /\n> /tracked/\n> /tracked/file\n> /tracked/dir/\n> /untracked/\n> /untracked/file\n> /untracked/dir/\n>\n> and the following commands:\n> $ git add tracked\n>\n> $ git clone\n>\n> The cloned result could be any of the following:\n>\n> (1)\n> /\n> /tracked/\n> /tracked/file\n>\n> This is the current behaviour; directories are not tracked at all, but only \n> added as necessary to support files.\n\nAnd so your case (1) actually rather is a single line:\n\n/tracked/file\n\nEverything else is just part of representing /tracked/file and\ndisappears as soon as /tracked/file disappears.\n\n> (2)\n> /\n> /tracked/\n> /tracked/file\n> /tracked/dir/\n> /untracked/\n> /untracked/dir/\n>\n> i.e. implicitly tracking _all_ directories. This is what you literally ask \n> for,\n\nI don't see how you can possibly conclude that from what I have been\nwriting.\n\n> but I think most would find this unreasonable.\n\nAnd it is.  So please _don't_ put words into my mouth.  In my\nproposal, the following (and nothing else) would get tracked:\n\n/tracked/.\n/tracked/file\n\nand that's it.  That is what was requested, and that is what is\ntracked.  There will be, incidentally, a tree \"/tracked/\" and a tree\n\"/\" in the _repository_, but those collapse as soon as they are empty.\nThey are just an _abstract_ data structuring tool in the repository\nthat is _mapped_ to directories on checkout.\n\n> /\n> /tracked/\n> /tracked/file\n> /tracked/dir/\n>\n> i.e. recursively tracking directories (and files). This seems useful, but \n> there is nothing _implicit_ about this.\n\nYou did not ask for \"/tracked/file\" and you did not ask for\n\"/tracked/dir/\" (whatever they may be).  That you wanted to track them\nwas _implied_ by your request of \"/tracked/\".\n\n> I have a feeling that you're actually arguing for doing (3) by\n> default.  What I am arguing is to do (1) by default, and (3) if\n> given a suitable command-line option (i.e. \"git add --with-dirs\n> tracked\").\n>\n> Note that this is really an interface question.\n\nNot at all.  It is a _conceptual_ question: in order for this to work\nat _all_ (instead of being an inconsistent heap of ugly surprises),\ndirectories need a representation in the repo.  This representation,\nas opposed to in the work file system, is _optional_: the repository\ngot perfectly well along without it up to now, and the fallback is\nalready implemented when there is a tree without corresponding\ndirectory.\n\n> How these entries are actually stored in the repo is a different\n> discussion.\n\nSure.  But anything that requires four dozens of special cases instead\nof four because one wanted to keep \"things that are under some\nspecialized view separate separate\" is not something I am going to\nimplement.  I am too old to juggle with complexity for the sake of\ncomplexity.  I can make much more use of the existing infrastructure\nby actually making file and directory entries quite similar.\n\nls -la\nalso has no special cases for \".\" and \"..\" because they are, at a very\nfundamental level, very special in achieving a special purpose\n_without_ being special-cased.\n\n> Finally, let's look at the case of \"git add tracked/file\" followed\n> by \"git rm tracked/file\". I'm arguing that \"tracked/\" should be\n> automatically removed, since I never asked for it to be tracked by\n> git.\n\nSure.  And nobody ever said otherwise.  In fact, I gave about a dozen\nexamples in that line and more special in the thread up to now.\n\n> On the other hand, \"git-add --non-recursive tracked\" followed by the\n> above two commands, should of course leave \"tracked/\" in place,\n> since I now actually asked explicitly for the directory to be\n> tracked.\n\nSure.  Use \"--directory\" instead of \"--non-recursive\" and you have a\nsomewhat more special option for that.\n\n> My point is fundamentally that selectively tracking directories is a\n> more powerful concept than just tracking _all_ directories by\n> default.\n\nPerhaps you might read up on some of the past discussion before\nbeating dead horses.  This has been covered already, and more than\nonce.  I never asked for \"all directories\" to be tracked.  I outlined\ncases where they are tracked and where not, and I tested that the\nmechanisms in \"man gitignore\" already work _perfectly_ with the\npattern \".\" for configuring the _implied_ tracking at directory,\nrepository, project, and user preference level.\n\n> Note that if we support selectively tracking directories, tracking\n> _everything_ (like you seem to want) is trivially implemented by\n> _always_ supplying the appropriate option to git-add. If we track\n> everything by design, we don't have the option of selectively\n> tracking some directories.\n\nBut that means manual intervention all of the time.  It is fine when a\ntool provides an option to shoot you in the arm instead of in the foot\nas usual, but that's not really a fix, but an acerbation of the\nproblem.\n\n>> > Basically, add a \"--dir\" flag to git-add, git-rm and friends, to\n>> > tell them you're acting on the directory itself (rather than its\n>> > (recursive) contents). \"git-add --dir foo\" will add the \"040000\n>> > 123abc... 0 foo\" to the index/tree whether or not foo is an empty\n>> > directory. \"git-rm --dir foo\" will remove that entry (or fail if\n>> > it doesn't exist), but _not_ the contents of foo.\n>> \n>> There is nothing wrong with implementing something like this in\n>> _addition_ to treating directory entries implicitly.\n>\n> I don't agree. By _selectively_ tracking directories you can\n> implement any policy you want on top of it.\n\nNo, you can't.  Because a \"policy\" means that things are _implied_.\nBeing able to do everything manually is not a policy.  It may be a\nlifesaver at times, but then you have little business drifting in the\nriver in the first place.\n\n>> But the important, the _really_ important thing are the implicit\n>> behaviors.  If I have to hassle with every directory myself, I\n>> don't need a content tracking system.\n>\n> I disagree. Just as you have to decide which files to track, you\n>similarly should have to decide which directories to track. Of\n>course, the tools make this easier for you by being able to\n>recursively handle files. In the same way they should be able to do\n>the same thing for directories.\n\n--directory _explicitly_ is not working recursively, so it does not\nsolve that problem.\n\n-- \nDavid Kastrup\n"},{"id":"47961","messageId":"200707201418.26534.johan@herland.net","threadId":"9086","inReplyTo":"86tzrzuyyy.fsf@lola.quinscape.zz","subject":"Re: Empty directories...","fromName":"Johan Herland","fromEmail":"johan@herland.net","sentAt":"2007-07-20T12:18:26Z","receivedAt":"2007-07-20T12:18:26Z","isPatch":false,"sender":{"key":"johan@herland.net","avatar":"https://avatars.githubusercontent.com/u/547031?v=4"},"body":"On Friday 20 July 2007, David Kastrup wrote:\n> Johan Herland <johan@herland.net> writes:\n> > My point is fundamentally that selectively tracking directories is a\n> > more powerful concept than just tracking _all_ directories by\n> > default.\n> \n> Perhaps you might read up on some of the past discussion before\n> beating dead horses.  This has been covered already, and more than\n> once.  I never asked for \"all directories\" to be tracked.  I outlined\n> cases where they are tracked and where not, and I tested that the\n> mechanisms in \"man gitignore\" already work _perfectly_ with the\n> pattern \".\" for configuring the _implied_ tracking at directory,\n> repository, project, and user preference level.\n\nIt seems our discussion is based on so many misunderstandings of each other \nthat it's not very useful to reply to specific parts of it.\n\nAFAICS, from a high-level POV, we're pretty much in agreement on the following \npoints:\n\n1. Git should be able to track directories.\n\n2. Tracked directories should be kept alive, even if empty.\n\n3. Git must not necessarily track _all_ directories.\n\n\nConversely, we seem to disagree on these points:\n\n4. Whether or not git should track directories by default. You say yes, I say \nno.\n\n5. How the tracking of directories should be implemented in git's object \ndatabase. I want to keep the index/tree as-is except for adding directory \nentries (w/mode 040000) for the tracked directories only. You seem to want to \nadd directory entries for _all_ directories and then additional \".\" entries \nfor directories you don't want deleted if/when empty.\n\n\nAm I making sense, or have I misunderstood our misunderstandings?\n\n\n...Johan\n\n\n-- \nJohan Herland, <johan@herland.net>\nwww.herland.net\n"},{"id":"47963","messageId":"200707201520.55911.johan@herland.net","threadId":"9086","inReplyTo":"86odi7utdj.fsf@lola.quinscape.zz","subject":"Re: Empty directories...","fromName":"Johan Herland","fromEmail":"johan@herland.net","sentAt":"2007-07-20T13:20:55Z","receivedAt":"2007-07-20T13:20:55Z","isPatch":false,"sender":{"key":"johan@herland.net","avatar":"https://avatars.githubusercontent.com/u/547031?v=4"},"body":"On Friday 20 July 2007, David Kastrup wrote:\n> Johan Herland <johan@herland.net> writes:\n> \n> > AFAICS, from a high-level POV, we're pretty much in agreement on the\n> > following points:\n> >\n> > 1. Git should be able to track directories.\n> >\n> > 2. Tracked directories should be kept alive, even if empty.\n> >\n> > 3. Git must not necessarily track _all_ directories.\n> >\n> >\n> > Conversely, we seem to disagree on these points:\n> >\n> > 4. Whether or not git should track directories by default. You say\n> > yes, I say no.\n> \n> Element of least surprise.  But since my proposal allows easy and\n> intuitive declaration of the preference at user, project, and\n> directory level without one choice messing with the choice of other\n> projects and contributors with mixed preferences, this is quite\n> unimportant.\n> \n> We are in agreement that adding or removing the tracking explicitly\n> for a single directory might be useful to have.  But it can't be the\n> only way.\n\nAs long as you can add/remove tracking recursively for a whole (sub)tree, I \ndon't see what's the problem. Of course, if you want to change the default \nbehaviour, you should be able either set a config variable somewhere, or - as \na last resort - alias git-add and git-rm to always supply the appropriate \ncommand-line option.\n\n> > 5. How the tracking of directories should be implemented in git's\n> > object database. I want to keep the index/tree as-is except for\n> > adding directory entries (w/mode 040000) for the tracked directories\n> > only. You seem to want to add directory entries for _all_\n> > directories and then additional \".\" entries for directories you\n> > don't want deleted if/when empty.\n> \n> No.  I don't want to change _anything_ for untracked directories.\n> They are, as previously, implied by the contents and have a \"tree\"\n> entry for efficiency reasons.  Nothing new here.\n> \n> The directory mode entries are named \".\" and are for tracked\n> directories only.\n\nOk. So our difference in opinion on implementation is even smaller than I \nimagined; basically only whether the directory is tracked by a mode \"040000\" \nentry, or by a \".\" entry.\n\n> > Am I making sense, or have I misunderstood our misunderstandings?\n> \n> The latter.  You are violently arguing for what I outlined.  Which\n> probably shows that I am not the best at explaining my ideas, and that\n> it reflects badly upon them.\n\nThat probably goes for both of us :)\n\n\nWell, as long as we have this clarified, I don't see much point in continuing \nthis part of the thread. I feel confident that the git community as a whole \nwill converge on the best technical solution, once it surfaces.\n\n\nHave fun!\n\n...Johan\n\n\n-- \nJohan Herland, <johan@herland.net>\nwww.herland.net\n"},{"id":"47964","messageId":"86k5svus2k.fsf@lola.quinscape.zz","threadId":"9086","inReplyTo":"7vk5svxt1f.fsf@assigned-by-dhcp.cox.net","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-20T13:23:47Z","receivedAt":"2007-07-20T13:23:47Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Actually, I do not think there is absolute right or wrong here.  The\n> difference is not that the information is at the \"right\" or \"wrong\"\n> place, but one approach places the information at more\n> efficient-to-use place than the other.\n\nAgreed.\n\n> In that sense, the attribute approach _is_ a more elegant solution\n> between the two.\n\nDisagreed.  See below.\n\n> Making it an attribute has a huge practical advantage.\n>\n> By treating executable bit as a piece metadata, we can compare the\n> \"contents\" quickly.  If you \"chmod +x\" a blob without changing\n> anything else, we can detect that fact, because blob object names\n> are equal.  At the philosophical level, you _could_ argue that the\n> executable-ness is one bit of content and include that in the object\n> name computation for the blob.  There is nothing fundamentally wrong\n> about that approach, but that destroys the nice \"cheap\n> comparability\" between blobs that differ only by executable-ness.\n>\n> David's \".\" in tree is essentially the same argument as treating the\n> executable-ness as one extra bit of content.  The fact that a\n> particular tree wants to stay even after emptied can be treated as\n> part of contents (thereby reflected in its object name).\n\nSmall nit here: the tree does not want to stay after emptied, since it\nis not empty as long as it contains \".\".\n\n> There is nothing fundamentally wrong there, either.  But that means\n> two trees that contain otherwise identical set of blobs and\n> subtrees, but differ only in the behaviour of when they are emptied,\n> would get different object names, hence you need to descend into\n> them to see if they are different.\n\nAnd here we disagree in our assessment, and where I find the example\nof the execute bit unfitting.  We are talking about _trees_ here, not\nfiles.  So this is only relevant if we have a _huge_, _flat_ tree with\n_lots_ of entries at _bottom_ level.\n\nHow often does it occur in practice that a _large_ tree has \".\"  added\nor removed and nothing else changes?  Never, because the normal use\ncase is that a directory is either tracked from the start, or not\ntracked at all.  And even if you change the tracking for a whole\nproject at once (which is a one-time job): the cost difference is\nlooking at all _tree_ leaf entries, not at all the involved files.\n\n> Using attribute that is detached from the content itself allows you\n> to hoist that one bit one level up.  By treating executable-ness not\n> as part of content, we can compare two blobs with different\n> executable bits cheaply.  You can avoid descending into such a tree\n> when comparing it with another tree that is different only by the\n> \"will-stay-when-emptied\"-ness the same way.\n\nBut changing the executable bit of a file will happen often during\ndevelopment.  Adding or removing \".\" will never usually be done _ever_\nexcept when the tree is first created or removed, and then the cost is\nnegligible.\n\nSo \"performance\" is not an issue for making this an attribute or a\nflat entry.  While the user level abstraction need not match the\nactual representation, I think that it will make for lot less special\ncases and problematic behavior to pull through with \".\" as a directory\nentry that mostly behaves like other files and, like other files,\nrequires git to create a directory to contain it.  All the logic for\ncreating and deleting directories and creating and adding and ignoring\nfiles can _perfectly_ stay the same.\n\nThere are just two differences:\n\na) git always sees \".\" as a file in every directory in the work tree\n   and considers it a file.\nb) when it comes to actually creating or modifying or reading the\n   actual file in the work directory, it silently skips the\n   operation.\n\nIt would not even be necessary to give the directory entry any special\nattributes or permissions to make this scheme work: declaring it a\nnormal file and just special-casing the name \".\" on those operations\nwould lead to consistent and working behavior, with no change of\nformat in index and repository at all.\n\nPossibly even a) alone would suffice, at the cost of letting git\ncomplain and continue at every operation (or making a _really_ royal\nmess for Solaris root users).\n\nI might be tempted to make a proof-of-concept patch for that.\n\nBut for backward-compatibility, it will be better to use an entry type\nwhich old versions of git will be able to ignore when checking out or\nin.  And for user-friendliness, one does not really want to list such\nentries as regular files.\n\n-- \nDavid Kastrup\n"},{"id":"47966","messageId":"86fy3jurlw.fsf@lola.quinscape.zz","threadId":"9086","inReplyTo":"200707201520.55911.johan@herland.net","subject":"Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-20T13:33:47Z","receivedAt":"2007-07-20T13:33:47Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Johan Herland <johan@herland.net> writes:\n\n> On Friday 20 July 2007, David Kastrup wrote:\n>> Johan Herland <johan@herland.net> writes:\n>\n>> > 4. Whether or not git should track directories by default. You\n>> > say yes, I say no.\n>> \n>> Element of least surprise.  But since my proposal allows easy and\n>> intuitive declaration of the preference at user, project, and\n>> directory level without one choice messing with the choice of other\n>> projects and contributors with mixed preferences, this is quite\n>> unimportant.\n>> \n>> We are in agreement that adding or removing the tracking explicitly\n>> for a single directory might be useful to have.  But it can't be\n>> the only way.\n>\n> As long as you can add/remove tracking recursively for a whole\n> (sub)tree, I don't see what's the problem.\n\nNeither do I.  But a --directory option never is recursive.  That is\nthe whole point.\n\nProbably we are in violent agreement again.\n\n> Of course, if you want to change the default behaviour, you should\n> be able either set a config variable somewhere, or - as a last\n> resort - alias git-add and git-rm to always supply the appropriate\n> command-line option.\n\nOr declare diverging behaviors using a !. or . entry in the gitignore\nmechanisms.  Which work everywhere where we need them.\n\n>> > 5. How the tracking of directories should be implemented in git's\n>> > object database. I want to keep the index/tree as-is except for\n>> > adding directory entries (w/mode 040000) for the tracked\n>> > directories only. You seem to want to add directory entries for\n>> > _all_ directories and then additional \".\" entries for directories\n>> > you don't want deleted if/when empty.\n>> \n>> No.  I don't want to change _anything_ for untracked directories.\n>> They are, as previously, implied by the contents and have a \"tree\"\n>> entry for efficiency reasons.  Nothing new here.\n>> \n>> The directory mode entries are named \".\" and are for tracked\n>> directories only.\n>\n> Ok. So our difference in opinion on implementation is even smaller\n> than I imagined; basically only whether the directory is tracked by\n> a mode \"040000\" entry, or by a \".\" entry.\n\nActually, even smaller: I'd track them by a \".\" entry with mode\n1777755755 or whatever is the natural expression for \"this is a\ndirectory\".  The mode would be different from the existing \"this is a\ntree\".\n\n_If_ one wants at one time track permissions of files apart from \"x\",\nthe \".\" entry would be natural for carrying directory permissions.\nWithout \".\", you basically tell git \"I don't care about the existence\nof this directory.  Just do what is necessary for checking out my\nfiles\".\n\n>> > Am I making sense, or have I misunderstood our misunderstandings?\n>> \n>> The latter.  You are violently arguing for what I outlined.  Which\n>> probably shows that I am not the best at explaining my ideas, and\n>> that it reflects badly upon them.\n>\n> That probably goes for both of us :)\n>\n> Well, as long as we have this clarified, I don't see much point in\n> continuing this part of the thread. I feel confident that the git\n> community as a whole will converge on the best technical solution,\n> once it surfaces.\n\nI'll probably crank out some insolently primitive proof of concept\neventually.\n\n-- \nDavid Kastrup\n"},{"id":"47970","messageId":"alpine.LFD.0.999.0707200827270.27249@woody.linux-foundation.org","threadId":"9086","inReplyTo":"85sl7jty43.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-20T15:31:18Z","receivedAt":"2007-07-20T15:31:18Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 20 Jul 2007, David Kastrup wrote:\n> \n> Like \".\" destroyed the fundamental concepts of Unix filesystems.\n\nDavid, I'd suggest you just be quiet and learn, instead of spouting \nidiotic nonsense.\n\nWhen Junio talks about fundamental concepts of git, you should sit back, \nrelax, and ponder. And maybe realize that the git filesystem isn't a \"unix \nfilesystem\". It's a content-addressable one, it's not POSIX, and yes, it \nreally does have totally different fundamental concepts.\n\nSo your arguments are just inane and stupid, and show that you aren't \nworth discussing with, because you don't even understand what you are \ntalking about.\n\nSo here's a suggestion: how about trying to *understand* git first. After \nthat, you can talk.\n\nIn fact, at this point, I have an even better suggestion: how about you \njust shut the hell up until you have a tested patch? Code talks, bullshit \nwalks. And right now you are nothing but bullshit.\n\n\t\tLinus\n"},{"id":"47990","messageId":"alpine.LFD.0.999.0707201210550.27249@woody.linux-foundation.org","threadId":"9086","inReplyTo":"7vk5svxt1f.fsf@assigned-by-dhcp.cox.net","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-20T19:24:02Z","receivedAt":"2007-07-20T19:24:02Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 20 Jul 2007, Junio C Hamano wrote:\n> \n> Using attribute that is detached from the content itself allows\n> you to hoist that one bit one level up.  By treating\n> executable-ness not as part of content, we can compare two blobs\n> with different executable bits cheaply.  You can avoid\n> descending into such a tree when comparing it with another tree\n> that is different only by the \"will-stay-when-emptied\"-ness the\n> same way.\n\nHaving thought about it a bit more, I would absolutely *detest* any kind \nof \"executable bit\" like behaviour.\n\nWhy? \n\nMerging. I think one of the fundamental issues in merging is that you do \nit \"in the working tree\". This is something that pretty much *everybody* \nelse gets wrong, and it's somethign where git absolutely shines.\n\nBut git shines here exactly because git never tracks \"history\" or the \nstate in the tree, and only ever tracks things that are indubitably real \ncontent. Which is why you never *ever* have to tell git about \"I moved \nfile X to file Y\" - because git only tracks things that it can see right \nin front of it, in the tree.\n\nThe \"sticky directory\" bit simply would not be something like that. It \nsimply isn't \"content\", and as such, it should not be tracked. It's as \neasy as that. We don't want a merge of two branches to have to specify any \nextra data \"outside\" the tree as to how it should be merged.\n\nSo the issue about whether a directory *exists* or not can be merged (just \nlook at the tree), but the issue about whether the directory is supposed \nto be sticky is something that you'd have to tell git about *outside* of \nthe tree, and that violates the whole point of working tree merges.\n\nI do realize that if you use inferior operating systems, we already have \nthese kinds of \"outside the tree\" data entries, thanks to issues like \nsymlinks and normal file executable bits that you would have to explicitly \ntell git about when you're working in a broken environment. So in that \nsense, it wouldn't be anything technically new for git. \n\nBut that doesn't change the fundamental issue: the limitation with \nexecutable bits and symlinks is a limitation of the broken environment, \nnot of git. But \"directories stay around after the last file is gone\" is \nnot that, it would simply be a design mistake in git itself.\n\nThere are other reasons to not do it. What about file renames? Maybe the \ndirectory got *renamed*. From a pure content angle, this is \"all the files \nin that directory went away\". If you have stupid rules like \"directories \nstay around even though all the files went away\", you would again have \nproblems with this common case.\n\nIn other words: I don't care one whit about the whiners. What's MUCH more \nimportant than some random whiny person saying \"Daddy, daddy, I want a \npony\" is whether you can afford to maintain that pony in the future. And \nthis pony is just stupid.\n\nSo here:\n\n\tNo, you cannot have a pony. NOT YOURS.\n\nbut I still think we should support the concept of importing things from \nother systems, and thus eventually support empty directories. Just not any \ncrazy semantics with sticky histories.\n\n\t\t\tLinus\n\nPS. As usual, per-user or per-repository *local* attributes are something \nelse. They aren't \"sticky history\", they are just purely behavioural \ndefaults. Those kinds of things may make sense. But that's not a \"tracking \ncontent\" issue.\n"},{"id":"47995","messageId":"200707202302.57788.johan@herland.net","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707201210550.27249@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Johan Herland","fromEmail":"johan@herland.net","sentAt":"2007-07-20T21:02:57Z","receivedAt":"2007-07-20T21:02:57Z","isPatch":true,"sender":{"key":"johan@herland.net","avatar":"https://avatars.githubusercontent.com/u/547031?v=4"},"body":"On Friday 20 July 2007, Linus Torvalds wrote:\n> [...]\n> \n> But that doesn't change the fundamental issue: the limitation with \n> executable bits and symlinks is a limitation of the broken environment, \n> not of git. But \"directories stay around after the last file is gone\" is \n> not that, it would simply be a design mistake in git itself.\n> \n> There are other reasons to not do it. What about file renames? Maybe the \n> directory got *renamed*. From a pure content angle, this is \"all the files \n> in that directory went away\". If you have stupid rules like \"directories \n> stay around even though all the files went away\", you would again have \n> problems with this common case.\n> \n> In other words: I don't care one whit about the whiners. What's MUCH more \n> important than some random whiny person saying \"Daddy, daddy, I want a \n> pony\" is whether you can afford to maintain that pony in the future. And \n> this pony is just stupid.\n> \n> So here:\n> \n> \tNo, you cannot have a pony. NOT YOURS.\n> \n> but I still think we should support the concept of importing things from \n> other systems, and thus eventually support empty directories. Just not any \n> crazy semantics with sticky histories.\n\nDoes this mean that you are firmly opposed to the concept of storing \ndirectories in the index/tree as such, or that you are only opposed to \n(some of) the implementation ideas that have been discussed so far?\n\nIf the former is the case, does this mean that there will be no support for \nempty directories in git, alternatively that such support is limited to \nincorporating e.g. Dscho's .gitignore workaround into porcelain commands \n(i.e. \"git add --directory some_dir\" will be mangled/transformed \ninto \"touch some_dir/.gitignore && git add some_dir/.gitignore\")?\n\n(Granted, Dscho's .gitignore workaround is fairly elegant as workarounds go, \nbut it still reeks of inheriting a CVS misfeature.)\n\n\nHave fun!\n\n...Johan\n\n-- \nJohan Herland, <johan@herland.net>\nwww.herland.net\n"},{"id":"47994","messageId":"alpine.LFD.0.999.0707201421110.27249@woody.linux-foundation.org","threadId":"9086","inReplyTo":"200707202302.57788.johan@herland.net","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-20T21:48:51Z","receivedAt":"2007-07-20T21:48:51Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 20 Jul 2007, Johan Herland wrote:\n> \n> Does this mean that you are firmly opposed to the concept of storing \n> directories in the index/tree as such, or that you are only opposed to \n> (some of) the implementation ideas that have been discussed so far?\n\nI've already sent out a *patch* to do so, for chissake. It handled all \nthese cases perfectly fine, as far as I know, but I didn't test it all \nthat deeply (and made it clear when I sent that patch out).\n\nIn fact, in this whole pointless discussion, I think I'm so far the only \none to have done anything constructive at all. Sad.\n\nSo here's my standpoint:\n\n - people who use git natively might as well use the \".gitignore\" trick. \n   It really *does* work, and there really aren't any downsides. Those \n   directories will stay around forever, until you decide that you don't \n   want them any more. Problem solved.\n\n   Sure, if you export the git archive into some other format, you might \n   well want to do something about the \".gitignore\" files (like just \n   delete them, since they won't be meaningful in an SVN environment, for \n   example, but you might also just convert them into SVN's \"attributes\" \n   or whatever it is that SVN uses to ignore files).\n\n - If you don't use git natively, but just to track another thing, you \n   could easily use the patches that I already sent out. Yes, they need \n   more testing. Yes, you'd also probably like some user interface updates \n   (notably \"git add/rm\" should be taught about directories).\n\n   And yes, I probably (almost certainly) didn't handle all cases, but the \n   patch I sent out was actually a working one. It really *did* pass my \n   trivial tests.\n\nBut once you start tracking empty directories *without* a .gitignore file, \nsome things fall out of that:\n\n - git really *really* is designed to track \"snapshots in time\". You \n   generate history from these snapshots. This is a very fundmanetal \n   issue, and a lot of people seem to have trouble understanding the \n   deeper implications.\n\n   For example, git and hg may look similar, but git tracks \"snapshots in \n   time\", and hg tracks \"file histories tied together in snapshots\". That \n   really is a fundamentally different thing. \n\n   And one of the fundamental results of git's approach is that content is \n   content. There is *never* any notion of \"history\".  A snapshot really \n   is just that: it's a standalone thing. It *has* no history. The history \n   comes entirely from outside.\n\n   This means that the whole notion of \"this directory will not go away \n   because I added it explicitly\" is a totally broken notion in git. It \n   has a notion of \"history\" - something that simply DOES NOT EXIST, \n   unless you seriously break the whole notion of \"snapshots in time\".\n\n   In other words, when I say that git is a \"content tracker\", I'm \n   serious. It tracks nothing *but* content. If some concept doesn't exist \n   in the working tree, git doesn't track it. If it cannot be seen in the \n   filesystem, it doesn't exist.\n\n - Contrast this with a lot of totally broken SCM's, that track \"history\" \n   of files. As a result, they have absolutely *horrid* merge problems, \n   because you can no longer just merge things in the working directory, \n   and \"the result\" is the result. No, if you track history, you now have \n   to tell the SCM about how the *history* moved, not just the content.\n\nSo this is why git MUST NOT make the difference between\n\n - a directory was was created explicitly and then had a few files added \n   to it, and then had those files deleted from it\n\nand\n\n - we added a few files, we removed them\n\nThe end result MUST BE the same, because the  state IN THE WORKING TREE is \nthe same!\n\nIf the contents are the same, the end result must be the same. It's that \nsimple. And it all comes down to: \"git tracks contents\".\n\nNow, having said that, it doesn't matter *what* the end result is, as long \nas it's the same for both cases. What we do now is that when the files go \naway, the directory is no longer tracked.\n\nBut we *could* say that when we remove files, we always add back the \ndirectory they were in if that directory still exists in the filesystem.\n\nSee? Both are consistent with the \"git tracks contents\" notion. The only \nthing that is *not* consistent with that notion is to have a flag that we \ncarry along that says \"keep this directory\". That's no longer content, and \nnow you'd be tracking some internal SCM history instead. And that is a \nmistake. It may sound like a small mistake (and it is), but down that path \nlies madness. It's much better to teach people _why_ git doesn't do it, \nthan to say \"ok, git tracks content, but we have this special case where \nwe also track something else, namely a git internal \"stickiness\" notion\".\n\nSCM is too important to play games with. Git gets things right, and I \ndoubt people really _realize_ that the \"tracks content\" is why git is so \nmuch better, and why git can do merges so much faster and more reliably \nthan anybody else.\n\nSo the rule really *must* be:\n\n - if two trees look the same in the filesystem, they *must* have the same \n   git SHA1, because by definition, they have the same content.\n\nAnything that breaks that very simple statement is fundamentally broken.\n\n\t\t\tLinus\n\nPS. I realize that nobody actually seems to be writing code, and that this \nis a \"paint the bike shed\" discussion for everybody else, but just in case \nthere are people who don't just masturbate about the color of the shed, \nI'd like to point out that we really *do* need to enhance the \"diff\" rules \ntoo, so that you can express the changes in a tree as a diff too. Because \nif we track empty directories, then we need to be able to also *show* the \ndifference between a tree that has an empty directory, and one that does \nnot.\n"},{"id":"47997","messageId":"Pine.LNX.4.64.0707202320300.16498@reaper.quantumfyre.co.uk","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707201421110.27249@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2007-07-20T22:36:25Z","receivedAt":"2007-07-20T22:36:25Z","isPatch":true,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Fri, 20 Jul 2007, Linus Torvalds wrote:\n\n>\n>\n> On Fri, 20 Jul 2007, Johan Herland wrote:\n>>\n>> Does this mean that you are firmly opposed to the concept of storing\n>> directories in the index/tree as such, or that you are only opposed to\n>> (some of) the implementation ideas that have been discussed so far?\n>\n> I've already sent out a *patch* to do so, for chissake. It handled all\n> these cases perfectly fine, as far as I know, but I didn't test it all\n> that deeply (and made it clear when I sent that patch out).\n>\n> In fact, in this whole pointless discussion, I think I'm so far the only\n> one to have done anything constructive at all. Sad.\n\nThere was Dscho's .gitignore based patch too ...\n\n>\n> So here's my standpoint:\n>\n> - people who use git natively might as well use the \".gitignore\" trick.\n>   It really *does* work, and there really aren't any downsides. Those\n>   directories will stay around forever, until you decide that you don't\n>   want them any more. Problem solved.\n>\n>   Sure, if you export the git archive into some other format, you might\n>   well want to do something about the \".gitignore\" files (like just\n>   delete them, since they won't be meaningful in an SVN environment, for\n>   example, but you might also just convert them into SVN's \"attributes\"\n>   or whatever it is that SVN uses to ignore files).\n\nPersonally I quite like this approach - I'm going to use it to keep all \nthe empty directories from Subversion in my importer.  It seems to address \neverthing quite neatly.\n\nI don't really understand the objections ... especially since I can't see \nwhy you want an empty directory if you're not going to put _something_ in \nit - in which case, presumably you want to ignore it (so maybe a \n.gitignore containing * would be better than an empty one)?  However, I'm \nsure that if people want it, they have a reason.\n\n> SCM is too important to play games with. Git gets things right, and I\n> doubt people really _realize_ that the \"tracks content\" is why git is so\n> much better, and why git can do merges so much faster and more reliably\n> than anybody else.\n\nThis is the thing that made me interested in git back in April '05.  I \ncouldn't see what we were going to end up with at that point - but I was \n_convinced_ that due to the underlying design it was worth watching. \nBeing a python type (sorry ... :$) hg looked interesting when it sprang up \n- but they threw away what I considered to be one of the most compelling \nfeatures of git (at the time there wasn't the wealth of really nice tools \nthat we now have).\n\nIn fact, I really should say \"Thank you Linus\", since I came that close to \nwriting an SCM from scratch myself - having been using Subversion with \nbranches for quite some time (and CVS before that - and yes I do mean \nbranches + CVS).  Now I no longer feel the need to write an SCM - just a \nlonging to use git.  git is probably better than anything I would have \ncome up with too. :D\n\n-- \nJulian\n\n  ---\nShe is descended from a long line that her mother listened to.\n \t\t-- Gypsy Rose Lee\n"},{"id":"47999","messageId":"alpine.LFD.0.999.0707201712150.27249@woody.linux-foundation.org","threadId":"9086","inReplyTo":"Pine.LNX.4.64.0707202320300.16498@reaper.quantumfyre.co.uk","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-21T00:18:42Z","receivedAt":"2007-07-21T00:18:42Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 20 Jul 2007, Julian Phillips wrote:\n\n> On Fri, 20 Jul 2007, Linus Torvalds wrote:\n> > \n> > So here's my standpoint:\n> > \n> > - people who use git natively might as well use the \".gitignore\" trick.\n> >   It really *does* work, and there really aren't any downsides. Those\n> >   directories will stay around forever, until you decide that you don't\n> >   want them any more. Problem solved.\n> \n> Personally I quite like this approach - I'm going to use it to keep all the\n> empty directories from Subversion in my importer.  It seems to address\n> everthing quite neatly.\n\nThe really sad part about this discussion is that the \".gitignore trick\" \nis really technically no different at all from the one that David Kastrup \nhas been advocating a few times, except he calls his \".gitignore\" just \n\".\", and seems to think that it's somehow different.\n\nIt is true that \".gitignore\" and \".\" _are_ different.\n\nBut they are actually different in the sense that the \".gitignore\" thing \nis something you can control, while the \".\" thing is something that is in \nall directories on UNIX, which is exactly why it _must_not_ be used by git \nto mark existence. Exactly because it has thus lost its ability to be \nsomething you can tune per-directory in the working tree!\n\nThat said, I actually like my patch, because the git tree structures \nactually lend themselves very naturally to the \"empty tree\", and I know \npeople have even built up those kinds of trees on purpose, even if the \nindex doesn't support that notion.\n\nSo in that sense, teaching the index about an empty tree is in some ways \nthe \"right thing\" to do, if only because it means that the index can \nfinally express something that the tree objects themselves have always \nbeen able to validly encode.\n\n\t\t\tLinus\n"},{"id":"48002","messageId":"85hcnytuq8.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707201712150.27249@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-21T01:23:59Z","receivedAt":"2007-07-21T01:23:59Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> The really sad part about this discussion is that the \".gitignore\n> trick\" is really technically no different at all from the one that\n> David Kastrup has been advocating a few times, except he calls his\n> \".gitignore\" just \".\", and seems to think that it's somehow\n> different.\n\nOh no, I don't think at all that it is somehow different: actually\nthis is _exactly_ the reason why I think that the implementation will\nbe doable even by an idiot like myself, and that is because at least\nin my first iteration, \".\"  will appear as an empty regular file to\ngit, just like \".gitignore\".  The main worry I had was that putting\n\".\" inside of a gitignore entry might stop \"git add .\" from working\nlike previously.  But I tried it, and it works just like it would with\n\".gitignore\".  Or rather like it would with \".notignore\" since\n\".gitignore\" _is_ specially treated by git, after all.\n\n> It is true that \".gitignore\" and \".\" _are_ different.\n>\n> But they are actually different in the sense that the \".gitignore\"\n> thing is something you can control, while the \".\" thing is something\n> that is in all directories on UNIX, which is exactly why it\n> _must_not_ be used by git to mark existence.\n\nBut I don't plan to have it used by git to mark existence.  The\n_existence_ can be taken for granted.  But what can't be taken for\ngranted, like with any other file, is that the file is actually being\ntracked by git.  To have it tracked, you need to add it, and it must\nnot be covered by gitignore.\n\n> Exactly because it has thus lost its ability to be something you can\n> tune per-directory in the working tree!\n\nBut it should not let the user lose his ability to let or let not git\ntrack the file.\n\n> That said, I actually like my patch, because the git tree structures\n> actually lend themselves very naturally to the \"empty tree\", and I\n> know people have even built up those kinds of trees on purpose, even\n> if the index doesn't support that notion.\n\nAnd that is the reason I will be working with the \"empty file .\"\nmetaphor: it would be way above my head to make the index support new\nfile types or even structures, and change the evaporate-when-empty\nsemantics of trees and so on, while catching all special cases.\n\nI have no chance in hell to implement a new feature with a reasonable\namount of time and work.  That's a task for people with a larger brain\nthan mine who have my full admiration and respect.  The best I can\nhope to achieve is a clever hack.\n\nAnd if that works, people can still pile exceptions on it and redo it\nas a \"proper feature\".\n\nYou are _perfectly_ correct that my proposal is _not_ a jot different\nfrom registering a regular empty file \".notignore\", and it is on\n_purpose_, because I could not handle the complications if it were.\n\nThe only difference is that I am calling the file \".\".  Which is in\n_all_ respects nothing more than a naming convention.\n\nHowever, this convention has distinct advantages over \".notignore\":\n\na) I don't have to depart as far from reality.  Whenever I try\nregistering \".\", I can rely on the work directory actually _having_\n\".\" as a _real_, not a pseudofile.  It will not actually be a\n_regular_ file as I'll tell git: that's a wart of my prototype\nimplementation which will, no doubt, eventually be fixed by others\n_if_ the code does its job fine apart from being ugly to look at.  It\nmay not be even necessary internally to think of \".\" other than as an\nempty regular file, but git should probably not talk too loud about it\nlest people laugh at it.\n\nb) it already means something to people.  Now this is a two-edged\nsword, since \"almost, but not quite, entirely unlike\" concepts are not\nnecessarily helpful in computing.  In this case, however, I think the\nmatch is close enough to help people understand what is going on\nrather than the other way round.  \".\" was introduced because people\nwanted to have a good way to refer to a directory as an element of\nitself.  So using \".\" as a self-reference for a directory is quite in\nthe spirit of that name.\n\n> So in that sense, teaching the index about an empty tree is in some\n> ways the \"right thing\" to do, if only because it means that the\n> index can finally express something that the tree objects themselves\n> have always been able to validly encode.\n\nIf you define the tree objects by the physical in-memory or\nin-repository data structures encoding them, then you are correct.  I\nam somewhat reluctant to parade around another red cape, but in this\nparticular case, the size of the wet spot in my pants does not as much\nrelate to the physical layout of the data structure (big deal,\nprobably 30 lines of code all around), but rather to the extent and\nassumptions of functions accessing it.  Namely, data layout and\naccessor functions _together_ constitute a tree object.  So for me the\n\"evaporate-when-empty\" property, while not inherent in the physical\nlayout of the object, is still an inherent part of its structure which\nI would not want to touch: finding and fixing and debugging all code\nelements which explicitly or implicitly rely on that assumption is\nsomething I would not entrust myself with.\n\nI might have been more inclined to dabble with that approach if the\ntree stuff were written in something more object-oriented, say, clean\nand concise C++, except that clean and concise C++ code in the wild is\neven more of a mythical beast than clean and concise TeX code, and C++\nitself is such a mindboggingly complex contraption...  I digress.\n\nAll the best,\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48005","messageId":"85d4ymtnrx.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"85hcnytuq8.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-21T03:54:10Z","receivedAt":"2007-07-21T03:54:10Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"David Kastrup <dak@gnu.org> writes:\n\n> The only difference is that I am calling the file \".\".  Which is in\n> _all_ respects nothing more than a naming convention.\n>\n> However, this convention has distinct advantages over \".notignore\":\n>\n> a) I don't have to depart as far from reality.  Whenever I try\n> registering \".\", I can rely on the work directory actually _having_\n> \".\" as a _real_, not a pseudofile.  It will not actually be a\n> _regular_ file as I'll tell git: that's a wart of my prototype\n> implementation which will, no doubt, eventually be fixed by others\n> _if_ the code does its job fine apart from being ugly to look at.\n\nUpdate: well, I am still digging through the code, but this is all so\nwell factored that it might be perfectly feasible to have S_ISDIR\nentries after all without too much of a hassle.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48006","messageId":"851wf2bcqy.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707181557270.27353@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-21T04:29:41Z","receivedAt":"2007-07-21T04:29:41Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> This really updates three different areas, which are nicely\n> separated into three different files, so while it's one single\n> patch, you can actually follow along the changes by just looking at\n> the differences in each file, which directly translate to separate\n> conceptual changes:\n\nOk, I have now acquired enough passing familiarity with the code that\nI find part of my way around it.  Most of your patch looks like it\ncaters for the S_ISDIR type not previously in use in the index (how\nabout the repository?).  So that makes for quite a bit of nicer looks.\nThe disadvantage is that it introduces a new data type and thus one\nhas to check all the code paths to see how older versions of git will\ncater with newer data.\n\nMy idea of a fake zero-length file would have had predictable side\neffects:\n\nFor checking out, git would have created the directory it needed to\nplace the \"file\", then try to write an empty file called \".\" and\nfailing.  Apart from an error message (if we aren't root on Solaris),\nthis would have worked exactly as intended.\n\nFor deletion on checking out, git would have tried deleting \".\" and\nfailed.  I have not checked the code to see whether git takes this as\na clue not to attempt deleting the containing directory.  If not,\nagain stuff would have worked as intended.  If yes, well, the user\nneeds to clean up manually.\n\nI am not sure what code paths are executed when using S_ISDIR now in\nunmodified git.  As a theoretical question for now: do git\nrepositories carry some versioning inside them?  Something like \"don't\ntouch me if you are not at least version x\"?\n\nAnyway, the code becomes quite less of a dirty hack by using that data\ntype, so I am pretty much taking your code (which has no overlap to\nthe work I have done already) as is.  Seems like it should play\ntogether quite nicely with my own stuff.\n\nSo thanks for doing the heavy lifting in a difficult area.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48007","messageId":"alpine.LFD.0.999.0707202135450.27249@woody.linux-foundation.org","threadId":"9086","inReplyTo":"851wf2bcqy.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-21T04:51:13Z","receivedAt":"2007-07-21T04:51:13Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 21 Jul 2007, David Kastrup wrote:\n> \n> Ok, I have now acquired enough passing familiarity with the code that\n> I find part of my way around it.  Most of your patch looks like it\n> caters for the S_ISDIR type not previously in use in the index (how\n> about the repository?).\n\nThe object database has always had S_ISDIR (well, \"always\" is since very \nearly on, when I realized that flat trees didn't cut it).\n\n> The disadvantage is that it introduces a new data type and thus one\n> has to check all the code paths to see how older versions of git will\n> cater with newer data.\n\nTake a look at the \"subproject\" patches - those did the same (adding the \nntion of a gitlink to the index), except those also changed how the tree \nobject looked, since now a tree could contain pointers to commits too. \n\n> My idea of a fake zero-length file would have had predictable side\n> effects:\n\nAs far as I can tell, it would have been exactly the same thing as the \nS_IFDIR, just instead of the S_IFDIR check, you'd have had to check the \nend of the filename for being '/'.\n\nOtherwise? Exactly the same.\n\nExcept for the fact that we already supported S_IFGITLINK for subprojects \n(and there it matches the \"struct tree\" entry, so it really *does* make \nmore sense that way), so supporting S_IFDIR was actually easier.\n\nBut hey, that's an implementation detail. I don't actually care all that \nmuch. In many ways, the \"long-term\" data structures are much more \nimportant than the index, the index is a purely temporary - and even more \nimportantly - a purely local datastructure.\n\nThe more important thing is in many ways the object storage, and that's \nalso the reason for doing the index the way I did - it more closely \nmatches what the object storage does (ie the \"index\" ends up mirroring a \nlinearized and unpacked \"tree\" object).\n\n\t\t\tLinus\n"},{"id":"48008","messageId":"alpine.LFD.0.999.0707202154220.27249@woody.linux-foundation.org","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707202135450.27249@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-21T05:08:26Z","receivedAt":"2007-07-21T05:08:26Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 20 Jul 2007, Linus Torvalds wrote:\n> \n> As far as I can tell, it would have been exactly the same thing as the \n> S_IFDIR, just instead of the S_IFDIR check, you'd have had to check the \n> end of the filename for being '/'.\n\nBTW, there is actually one big difference, and the '/' at the end actually \nhas one huge advantage.\n\nWhy? Because my preliminary patches sort the index entries wrong. A \ndirectory should always sort *as*if* it had the '/' at the end.\n\nSee base_name_compare() for details.\n\nAnd we've never done that for the index, because the index has never had \nthis issue (since it never contained directories). So sit down and compare \nbase_name_compare (for tree entries) with cache_name_compare() (for index \nentries), and see how the latter doesn't care about the type of names.\n\nThis was actually something that I hit already with subproject support, \nand one of my very first patches even had some (aborted) code to start \nsorting subprojects in the index the way we sort directories.\n\nAnd I *should* have done it that way, but I never did. It now makes the \nS_ISDIR handling harder, because directories really do have to be sorted \nas if they had the '/' at the end, or \"git-fsck\" will complain about bad \nsorting.\n\nSad, sad, sad. It effectively means that S_IFGITLINK is *not* quite the \nsame as S_IFDIR, because they sort differently. Duh.\n\nOf course, it seldom matters, but basically, you should test a directory \nstructure that has the files\n\n\tdir.c\n\tdir/test\n\nin it, and the \"dir\" directory should always sort _after_ \"dir.c\".\n\nAnd yes, having the index entry with a '/' at the end would handle that \nautomatically.\n\nAs it is, with the \"mode\" difference, it instead needs to fix up \n\"cache_name_compare()\". Admittedly, that would actually be a cleanup \n(since it would now match base_name_compare() in logic, and could actually \nuse that to do the name comparison!), but it's a damn painful cleanup \nbecause we don't even pass in the mode to \"cache_name_compare()\", since we \nnever needed it.\n\nGaah.\n\ncache_name_compare itself isn't used in that many places, but it's used \nby \"index_name_pos()/cache_name_pos()\", which *is* used in many places. \nAnd again, that one doesn't even have the mode, so it cannot pass it down.\n\nSo it probably *is* easier to add the '/' at the end of the name instead, \nto make directories sort the right way in the index. I'd still suggest you \n*also* make the mode be S_IFDIR, though (and preferably make git-fsck \nactually verify that the mode and the last character of the name \nmatches!).\n\n\t\tLinus\n"},{"id":"48010","messageId":"85sl7i9w1j.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"alpine.LFD.0.999.07072\u000402135450.27249@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-21T05:15:52Z","receivedAt":"2007-07-21T05:15:52Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Sat, 21 Jul 2007, David Kastrup wrote:\n>> \n>> Ok, I have now acquired enough passing familiarity with the code\n>> that I find part of my way around it.  Most of your patch looks\n>> like it caters for the S_ISDIR type not previously in use in the\n>> index (how about the repository?).\n>\n> The object database has always had S_ISDIR (well, \"always\" is since\n> very early on, when I realized that flat trees didn't cut it).\n\nThen I think I have a bit of a problem: I should think that S_ISDIR in\nthe repository presumably marks a tree object (still very fuzzy around\nthe concepts here).  An explicitly checked-in directory (under my\nscheme always named \".\" inside of its tree) would presumably also have\nS_ISDIR in the repository but behave quite differently.\n\n> As far as I can tell, it would have been exactly the same thing as the \n> S_IFDIR, just instead of the S_IFDIR check, you'd have had to check the \n> end of the filename for being '/'.\n\nRelative file name of \".\", more or less.  Both names satisfy S_IFDIR\nin the filesystem, though.\n\n> Otherwise? Exactly the same.\n\n> The more important thing is in many ways the object storage, and\n> that's also the reason for doing the index the way I did - it more\n> closely matches what the object storage does (ie the \"index\" ends up\n> mirroring a linearized and unpacked \"tree\" object).\n\nI still have to get enough of a clue about the object store to see how\nthis pans out.  I would not want to have the \".\" objects marked as\ntype \"tree\" and empty if I can avoid it.  It seems unclean, would need\nextra case separations all over the place, violate the \"empty trees\nevaporate\" property and also waste a good place for tracking\npermissions or other attributes in future.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48011","messageId":"85odi69vgt.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707202154220.27249@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-21T05:28:18Z","receivedAt":"2007-07-21T05:28:18Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Fri, 20 Jul 2007, Linus Torvalds wrote:\n>> \n>> As far as I can tell, it would have been exactly the same thing as the \n>> S_IFDIR, just instead of the S_IFDIR check, you'd have had to check the \n>> end of the filename for being '/'.\n>\n> BTW, there is actually one big difference, and the '/' at the end actually \n> has one huge advantage.\n>\n> Why? Because my preliminary patches sort the index entries wrong. A \n> directory should always sort *as*if* it had the '/' at the end.\n\nHm, that's bad.  The thing is that the directory names I am tracking\nare called \".\" (that's what I was currently trying to reconcile your\ncode with).\n\n> And I *should* have done it that way, but I never did. It now makes\n> the S_ISDIR handling harder, because directories really do have to\n> be sorted as if they had the '/' at the end, or \"git-fsck\" will\n> complain about bad sorting.\n\nHm, I'll have to check what git-fsck does.\n\n> Of course, it seldom matters, but basically, you should test a directory \n> structure that has the files\n>\n> \tdir.c\n> \tdir/test\n>\n> in it, and the \"dir\" directory should always sort _after_ \"dir.c\".\n>\n> And yes, having the index entry with a '/' at the end would handle\n> that automatically.\n\nYou completely lost me here.  I guess I'll be able to pick this up\nonly after investing considerable more time into the data structures.\nAnd I have to goto bed right now.\n\n> As it is, with the \"mode\" difference, it instead needs to fix up\n> \"cache_name_compare()\". Admittedly, that would actually be a cleanup\n> (since it would now match base_name_compare() in logic, and could\n> actually use that to do the name comparison!), but it's a damn\n> painful cleanup because we don't even pass in the mode to\n> \"cache_name_compare()\", since we never needed it.\n>\n> Gaah.\n>\n> cache_name_compare itself isn't used in that many places, but it's\n> used by \"index_name_pos()/cache_name_pos()\", which *is* used in many\n> places.  And again, that one doesn't even have the mode, so it\n> cannot pass it down.\n>\n> So it probably *is* easier to add the '/' at the end of the name instead, \n> to make directories sort the right way in the index. I'd still suggest you \n> *also* make the mode be S_IFDIR, though (and preferably make git-fsck \n> actually verify that the mode and the last character of the name \n> matches!).\n\nThe _flattened_ directory name would end in /. in my scheme.  I would\nnot want to use \"xxx/\" for a directory name, and \"xxx\" for a tree:\nthat would be completely backwards.  And I also don't like the\nduplication of xxx when listing objects.\n\nSure, that's an implementation detail, but I don't like\nimplementations hurting my eyes...\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48042","messageId":"alpine.LFD.0.999.0707210832180.27249@woody.linux-foundation.org","threadId":"9086","inReplyTo":"85odi69vgt.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-21T15:53:39Z","receivedAt":"2007-07-21T15:53:39Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 21 Jul 2007, David Kastrup wrote:\n\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n> \n> > Of course, it seldom matters, but basically, you should test a directory \n> > structure that has the files\n> >\n> > \tdir.c\n> > \tdir/test\n> >\n> > in it, and the \"dir\" directory should always sort _after_ \"dir.c\".\n> >\n> > And yes, having the index entry with a '/' at the end would handle\n> > that automatically.\n> \n> You completely lost me here.  I guess I'll be able to pick this up\n> only after investing considerable more time into the data structures.\n\nSo the basic issue is that not only does git obviously think that only \ncontent matters, but it describes it with a single SHA1. \n\nThat's not an issue at all for a single file, but if you want to describe \n*multiple* files with a single SHA1 (which git obviously very much wants \nto do), the way you generate the SHA1 matters a lot.\n\nIn particular, the order.\n\nSo git is very very strict about the ordering of tree structures. A tree \nstructure is not just a random list of\n\n\t<ASCII mode> + <space> + <filename> + <NUL> + <SHA1>\n\nit's very much an _ordered_ list of those things, because we want the SHA1 \nof the tree to be well-specified by the contents, and that means that the \ncontents of a tree object has have absolutely _zero_ ambiguity.\n\nThis means, for example, that git is very fundamentally case sensitive. \nThere's no sane way *not* to be, because if you're case insensitive in any \nway at all, you'll end up having two trees that are \"the same\", but end up \nhaving different SHA1's.\n\nIt also means that git objects have absolutely zero \"localization\". There \nis no locale at all, and there very fundamnetally *must*not* be. Again, \nfor the same reason: if you can describe the same filename with two \ndifferent encodings, you'd have two different SHA1's for the same content.\n\nSo git filenames are very much a \"stream of bytes\", not anything else. And \nthey need to sort 100% reliably, always the same way, and never with any \nlocalized meaning.\n\nAnd, partly because it seemed most natural, and partly for historical \nreasons, the way git sorts filenames is by sorting by *pathname*. So if \nyou have three files named\n\n\ta.c\n\ta/c\n\tabc\n\nthen they sort in that exact order, and no other! They sort as a \"memcmp\" \nin the full pathname, and that's really nice when you see whole \ncollections of files, and you know the list is globally sorted.\n\nSo that \"global pathname sorting\" has nice properties, and it seems \n\"obvious\", but it means that because git actually *encodes* those three \nfiles hierarchically as two different trees (because there's a \nsubdirectory there), the tree objects themselves sort a bit oddly. The \ntree obejcts themselves will look like\n\n top-level tree:\n\t100644 a.c -> blob1\n\t040000 a   -> tree2\n\t100644 abc -> blob3\n\n sub-tree:\n\t100644 c    -> blob2\n\nand notice how the *tree* is not sorted alphabetically at all. It has a \nsubtly different sort, where the entry \"a\" sorts *after* the entry \"a.c\", \nbecause we know that it's a tree entry, and thus will (in the *global* \norder) sort as if it had a \"/\" at the end!\n\nTraditionally, when we have the index, the index sorting has been very \nsimple: you just sort the names as memcmp() would sort them. But note how \nthat changes, if \"a\" is an empty directory. Now the index needs to sort as\n\n\tfile a.c\n\tdir  a\n\tfile abc\n\nbecause when we create the tree entry, it needs to be sorted the same way \nall tree entries are always sorted - as if \"a\" had a slash at the end!\n\n[ Yeah, yeah, we could make a special case and just say \"the empty tree \n  sorts differently\", but that actually results in huge problems when \n  doing a \"diff\" between two trees: our diff machinery very much depends \n  on the fact that the index and the trees always sort the same way, and \n  if we sorted the \"a\" entry (when it is an empty directory) differently \n  from the \"a\" entry (when it has entries in it), that would just be \n  insane and cause no end of trouble for comparing two trees - one with an \n  empty directory and one with content added to that directory.\n\n  So the sorting is doubly important: it's what makes \"one content\" always \n  have the same SHA1, but it is also much easier and efficient to compare \n  directories when we know they are sorted the same way. ]\n\nIn other words, introducing tree entries in the index ended up also \nintroducing all the issues that we already had with the tree objects since \nthey got split up hierarchically, but that the code didn't use to have to \ncare about.\n\nThe easiest way to solve this really does seem to be to add the rule that \nthe index entry for an empty directory has to have the \"/\" at the end of \nthe name - then the \"sort mindlessly by name\" will just continue to work.\n\nBut that was what I said was broken: my patches I sent out didn't actually \ndo that.\n\nIt's *probably* just a few lines of code, and it actually would result in \nsome nice changes (\"git ls-files\" would show a '/' at the end of an empty \ndirectory entry, for example), so this is not a big deal, but it's an \nexample of how subtly different a directory is from a file when it comes \nto git.\n\n\t\t\tLinus\n"},{"id":"48047","messageId":"85tzrxslms.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707210832180.27249@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-21T17:38:03Z","receivedAt":"2007-07-21T17:38:03Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Sat, 21 Jul 2007, David Kastrup wrote:\n>\n>> Linus Torvalds <torvalds@linux-foundation.org> writes:\n>> \n>> > Of course, it seldom matters, but basically, you should test a directory \n>> > structure that has the files\n>> >\n>> > \tdir.c\n>> > \tdir/test\n>> >\n>> > in it, and the \"dir\" directory should always sort _after_ \"dir.c\".\n>> >\n>> > And yes, having the index entry with a '/' at the end would handle\n>> > that automatically.\n>> \n>> You completely lost me here.  I guess I'll be able to pick this up\n>> only after investing considerable more time into the data structures.\n\n[Basic explanation about git sort order and trees sorting as tree/ in\norder to be in the right sort order for a prefix]\n\nOk, I could not have figured this out on my own.  Are there any design\ndocuments or does one just have to pester the list?\n\n> So the basic issue is that not only does git obviously think that only \n> content matters, but it describes it with a single SHA1. \n>\n> That's not an issue at all for a single file, but if you want to describe \n> *multiple* files with a single SHA1 (which git obviously very much wants \n> to do), the way you generate the SHA1 matters a lot.\n>\n> In particular, the order.\n>\n> So git is very very strict about the ordering of tree structures. A tree \n> structure is not just a random list of\n>\n> \t<ASCII mode> + <space> + <filename> + <NUL> + <SHA1>\n\nOk.\n\n> So git filenames are very much a \"stream of bytes\", not anything\n> else. And they need to sort 100% reliably, always the same way, and\n> never with any localized meaning.\n\nThere is some utf-8/Unicode trouble to be expected in connection with\nthat eventually: some, but not all operating and/or file systems\ncanonicalize file names, replacing accented letters by a combining\naccent and the letter.  But that's beside the point.\n\n> And, partly because it seemed most natural, and partly for\n> historical reasons, the way git sorts filenames is by sorting by\n> *pathname*. So if you have three files named\n>\n> \ta.c\n> \ta/c\n> \tabc\n>\n> then they sort in that exact order, and no other! They sort as a\n> \"memcmp\" in the full pathname, and that's really nice when you see\n> whole collections of files, and you know the list is globally\n> sorted.\n\nIt is amusing that my description of git having no external concept of\ndirectories except as an expedience for representing slashes in\nfilenames was much closer to the mark that I would have expected.\n\n> So that \"global pathname sorting\" has nice properties, and it seems \n> \"obvious\", but it means that because git actually *encodes* those three \n> files hierarchically as two different trees (because there's a \n> subdirectory there), the tree objects themselves sort a bit oddly. The \n> tree obejcts themselves will look like\n>\n>  top-level tree:\n> \t100644 a.c -> blob1\n> \t040000 a   -> tree2\n> \t100644 abc -> blob3\n>\n>  sub-tree:\n> \t100644 c    -> blob2\n>\n> and notice how the *tree* is not sorted alphabetically at all. It has a \n> subtly different sort, where the entry \"a\" sorts *after* the entry \"a.c\", \n> because we know that it's a tree entry, and thus will (in the *global* \n> order) sort as if it had a \"/\" at the end!\n>\n> Traditionally, when we have the index, the index sorting has been very \n> simple: you just sort the names as memcmp() would sort them. But note how \n> that changes, if \"a\" is an empty directory. Now the index needs to sort as\n>\n> \tfile a.c\n> \tdir  a\n> \tfile abc\n>\n> because when we create the tree entry, it needs to be sorted the same way \n> all tree entries are always sorted - as if \"a\" had a slash at the end!\n\nHere is the layout as I would scheme it:\n\ntree1:\n     0?0000 .   -> dir1\n     100644 a.c -> blob1\n     040000 a   -> tree2\n     100644 abc -> blob3\n\nsub-tree:\n     0?0000 .    -> dir2\n     100644 c    -> blob2\n\nRemember that a tree evaporates when it is empty, and if we don't want\nto mess with that (which appears like a good idea to me), the \"don't\ndelete this\" indication belongs in the subtree where its natural name\nis \".\".  Since the dir entries are _leaves_ in the tree, there is no\nnecessity for sorting them specially.  They will usually appear first,\nbut people to all sorts of things, so filenames starting with \"!\"\nmight still come before them.\n\nSo the sorted flat file list for the above would be\n.    [dir]\na.c  [file]\na/   [tree]\na/.  [dir]\na/c  [file]\nabc  [file]\n\nNote that a tree is basically just a string arrangement tool which\ngets only incidentally mapped to directories when checking out.\n\nSo I am quite unhappy that 040000 is already taken by it.  I can't\neven say, \"ok, let . look like an empty tree\" because there should not\nbe something like an empty tree!  I find the correlation empty->gone\nvery important.\n\n> [ Yeah, yeah, we could make a special case and just say \"the empty\n> tree sorts differently\", but that actually results in huge problems\n> when doing a \"diff\" between two trees: our diff machinery very much\n> depends on the fact that the index and the trees always sort the\n> same way, and if we sorted the \"a\" entry (when it is an empty\n> directory) differently from the \"a\" entry (when it has entries in\n> it), that would just be insane and cause no end of trouble for\n> comparing two trees - one with an empty directory and one with\n> content added to that directory.\n\nIt appears to me like our ideas are still out of sync: a directory\nunder my scheme is _not_ at all an empty tree, rather it is an entry\n_inside_ of a tree, making the tree non-empty (which means that git\nwill not be tempted to delete the corresponing real-world directory\n_until_ one deletes the directory entry keeping the tree alive).\n\n>   So the sorting is doubly important: it's what makes \"one content\"\n>   always have the same SHA1, but it is also much easier and\n>   efficient to compare directories when we know they are sorted the\n>   same way. ]\n>\n> It's *probably* just a few lines of code, and it actually would\n> result in some nice changes (\"git ls-files\" would show a '/' at the\n> end of an empty directory entry, for example), so this is not a big\n> deal, but it's an example of how subtly different a directory is\n> from a file when it comes to git.\n\nLinus, a directory is simply non-existent inside of git.  Trees are an\nindexing mechanism solely determined by their content.  That is not a\nsubtle difference.  Git _uses_ directories when exporting in order to\nsimulate a flat namespace.  But it is internally oblivious to their\nexistence.  And that is a perfectly elegant and reasonable approach\nand I like it very much and don't want to mess with it at all.\n\nBut I also want to have directories represented within git, because\nnot doing so leads to awkward problems.  And the proper way as I see\nit is _not_ to mess with trees and stick them with \"stay when empty\"\nflags or similar.  This messes up the whole elegance of git's flat\nname space.  The proper way is to create a distinct object that\nrepresents a physical directory.  We don't need to represent the\ncontents of it: those are already tracked in the flat namespace fine,\nwith trees serving as an implementation detail.\n\nAll we need to represent is \".\".\n\nSo git-ls-files on\n.    [dir]\na.c  [file]\na/   [tree]\na/.  [dir]\na/c  [file]\nabc  [file]\n\nshould likely list\n\n.\na.c\na/.\na/c\nabc\n\nIf one wants to see the _tree_ because of its SHA1, it may also be\nlisted.  The SHA1 of a _directory_ like a/., in contrast, is\nuninteresting: it will be the same for every directory.\n\nWhether the _tree_ is listed as \"a\" or \"a/\" is probably a matter of\ntaste.  Personally, I think \"a/\" is better for bringing across the\nnotion that it is a structuring device not really related to the\nphysical _directory_ a which is _identical_ (meaning inode-identical,\nwhich is what counts in the physical world) to \"a/.\" even though it is\nanother name of it.\n\nAnd using \"a/\" puts it closer to its natural sort order.\n\nI'd write up a philosophy paper about git's relation between trees,\nfiles, directories if that were not utterly preposterous.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48049","messageId":"46A247C1.4000902@fs.ei.tum.de","threadId":"9086","inReplyTo":"85tzrxslms.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Simon 'corecode' Schubert","fromEmail":"corecode@fs.ei.tum.de","sentAt":"2007-07-21T17:52:01Z","receivedAt":"2007-07-21T17:52:01Z","isPatch":true,"sender":{"key":"corecode@fs.ei.tum.de","avatar":"https://gravatar.com/avatar/eff9dbf0cdac0d1e6a6cd7ed0e50763edcb376b493b5253a35ff167918ad79e1?d=mp&s=160"},"body":"David Kastrup wrote:\n> But I also want to have directories represented within git, because\n> not doing so leads to awkward problems.  And the proper way as I see\n> it is _not_ to mess with trees and stick them with \"stay when empty\"\n> flags or similar.  This messes up the whole elegance of git's flat\n> name space.  The proper way is to create a distinct object that\n> represents a physical directory.  We don't need to represent the\n> contents of it: those are already tracked in the flat namespace fine,\n> with trees serving as an implementation detail.\n> \n> All we need to represent is \".\".\n\nWhat I still don't get is:  How do you carry this information about \"this directory should not be removed\" from one checkout to the next commit?  When creating a .gitignore, this file exists in the workdir.  Of course you add some data to the index to stage it.  But how does this work with your \".\" \"file\"?  You can't put that in the filesystem.\n\ncheers\n  simon\n\n-- \nServe - BSD     +++  RENT this banner advert  +++    ASCII Ribbon   /\"\\\nWork - Mac      +++  space for low €€€ NOW!1  +++      Campaign     \\ /\nParty Enjoy Relax   |   http://dragonflybsd.org      Against  HTML   \\\nDude 2c 2 the max   !   http://golden-apple.biz       Mail + News   / \\\n"},{"id":"48050","messageId":"85lkd9sk7z.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"46A247C1.4000902@fs.ei.tum.de","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-21T18:08:32Z","receivedAt":"2007-07-21T18:08:32Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Simon 'corecode' Schubert <corecode@fs.ei.tum.de> writes:\n\n> David Kastrup wrote:\n>> But I also want to have directories represented within git, because\n>> not doing so leads to awkward problems.  And the proper way as I see\n>> it is _not_ to mess with trees and stick them with \"stay when empty\"\n>> flags or similar.  This messes up the whole elegance of git's flat\n>> name space.  The proper way is to create a distinct object that\n>> represents a physical directory.  We don't need to represent the\n>> contents of it: those are already tracked in the flat namespace fine,\n>> with trees serving as an implementation detail.\n>>\n>> All we need to represent is \".\".\n>\n> What I still don't get is: How do you carry this information about\n> \"this directory should not be removed\" from one checkout to the next\n> commit?\n\nI don't.  The only information in the file system is whether a\ndirectory exists or not.  \"Should not removed\" is not a property that\nis tracked.\n\n> When creating a .gitignore, this file exists in the workdir.  Of\n> course you add some data to the index to stage it.  But how does\n> this work with your \".\" \"file\"?  You can't put that in the\n> filesystem.\n\nEither the directory is in the file system or it is not.  Like with\nevery other file.  And either git tracks the directory, in which case\nit will notice its addition (when doing git-add) and removal (when\ndoing git-rm or git-commit -a) or git doesn't track the directory.\n\nWhen git tracks the directory (a matter of gitignore settings for\nimplicit tracking, and git-add for explicit tracking), and considers\nit existent, it will not touch it.  If it tracks it but considers it\nremoved in particular commit, it will attempt to remove it.\n\n    Fineprint: actually, things are more involved here: git does not\n    actually attempt to remove directories at the time it deletes them\n    from the tree: this is sort of pointless since the sort order\n    means that there might still be files it needs to take out from\n    the physical directory).  Instead, like before, git attempts to\n    remove a physical directory whenever the corresponding tree in git\n    becomes empty, and it is a prerequisite to delete a possibly\n    tracked directory from it.\n\nAfter it has attempted to remove it, it will leave it alone since it\nis now no longer tracking it.  If you add and remove a contained file,\nit will again try to remove the directory.  If you add _both_\ndirectory and a contained file, just removing the contained file will\nnot make git attempt to delete the directory.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48070","messageId":"alpine.LFD.0.999.0707211650190.3607@woody.linux-foundation.org","threadId":"9086","inReplyTo":"85tzrxslms.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-21T23:50:42Z","receivedAt":"2007-07-21T23:50:42Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 21 Jul 2007, David Kastrup wrote:\n> \n> tree1:\n>      0?0000 .   -> dir1\n>      100644 a.c -> blob1\n>      040000 a   -> tree2\n>      100644 abc -> blob3\n\nNo. Totally broken. That \".\" entry not only doesn't buy you anything, it \nis *impossible*. You  cannot make an object point to itself. Not possible.\n\nTell me how to calculate the SHA1 for the result. Also, tell me what the \n*point*  is. There is none.\n\n> Linus, a directory is simply non-existent inside of git. \n\nYou need to learn git first.\n\nA directory doesn't exist IN THE INDEX (until my patches). But you need to \nlearn about the object database and the SHA1's. That's the real meat of \ngit, and it sure as hell knows about directories.\n\n\t\tLinus\n"},{"id":"48072","messageId":"85644dqoig.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707211650190.3607@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-22T00:18:47Z","receivedAt":"2007-07-22T00:18:47Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Sat, 21 Jul 2007, David Kastrup wrote:\n>> \n>> tree1:\n>>      0?0000 .   -> dir1\n>>      100644 a.c -> blob1\n>>      040000 a   -> tree2\n>>      100644 abc -> blob3\n>\n> No. Totally broken. That \".\" entry not only doesn't buy you\n> anything, it is *impossible*. You cannot make an object point to\n> itself. Not possible.\n\nIt does not point to itself.  The name \".\" points to an entry of type\n\"dir\", no content is involved.  trees in the repository have content,\nand _only_ content.  directories in the repository imply existence,\nand _only_ existence.\n\n> Tell me how to calculate the SHA1 for the result.\n\nSince \".\" has no content (as long as we don't decide to track any file\npermissions at one point of time), _all_ entries \".\" will have the\nsame SHA1.\n\n> Also, tell me what the *point* is. There is none.\n\nThe point is to have a reflection of the physical existence of a\ndirectory.  Not just as a manner of accommodating slashes in a flat\nfilespace, allowing certain slash-related operations to be carried out\nefficiently.\n\n>> Linus, a directory is simply non-existent inside of git.\n>\n> You need to learn git first.\n>\n> A directory doesn't exist IN THE INDEX (until my patches). But you\n> need to learn about the object database and the SHA1's. That's the\n> real meat of git, and it sure as hell knows about directories.\n\nI have written up a complete explanation about the underlying concept\nin a separate thread, maybe it would make sense reading that before\ninvesting too much time meddling over details that don't fit the large\npicture.  The point is that the object database and the SHA1 values\ntrack _trees_, not _directories_.  And a _tree_ is just a hashing\nmechanism in the repository for files.  Its existence is solely\ndependent on the existence of its contents.  The only synchronization\nwith directories is that when a tree becomes empty, git attempts to do\nan rmdir on the corresponding directory.  And of course, if git needs\nto check out a file, it creates the necessary parent directories.\n\nNow since the physical _contents_ of a directory are already tracked\nin _trees_ by git, the only missing part is the _existence_ of the\ndirectory itself: a directory must exist as long as there is a tree\n(and thus content) connected with it, but the reverse does not hold:\nwithout a tree, the directory can still exist.  Which we can represent\nby a repository entry named \".\" without content (the content is\nalready catered for by the _tree_).  This must _not_ be represented by\na _tree_ node since there is no content, and a tree without content by\n_definition_ does not exist.\n\nI must be really bad at explaining things, or I am losing a fight\nagainst preconceptions fixed beyond my imagination.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48074","messageId":"851wf1qnst.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707211650190.3607@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-22T00:34:10Z","receivedAt":"2007-07-22T00:34:10Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Sat, 21 Jul 2007, David Kastrup wrote:\n>\n>> Linus, a directory is simply non-existent inside of git.\n>\n> You need to learn git first.\n>\n> A directory doesn't exist IN THE INDEX (until my patches). But you\n> need to learn about the object database and the SHA1's. That's the\n> real meat of git, and it sure as hell knows about directories.\n\nTo put it in another way: what would happen if trees were removed from\ngit's repository completely?  Instead we would just stipulate that git\nshould only track files, not trees, and that it would remove an\noutside directory when removing the last file from the repository that\ncan't be accomodated without such a directory.\n\nNow the effect would be that git would become quite inefficient.  But\nit would not change its behavior in any other way.  Because it knows\n_zilch_ about directories.  It knows about the hierarchy of the\n_contents_, but the directories, the physical entities in the work\ntree?  It deduces a convenient point of time to try deleting them\n(when a tree collapses), and it deduces that they are there as long as\nit is tracking their content, but no information about a _directory_\nother than its _contents_ ever enter the repository or index.  About\nits _existence_, git only keeps circumstantial evidence.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48077","messageId":"alpine.LFD.0.999.0707211737090.3607@woody.linux-foundation.org","threadId":"9086","inReplyTo":"85644dqoig.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-22T00:37:51Z","receivedAt":"2007-07-22T00:37:51Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 22 Jul 2007, David Kastrup wrote:\n> \n> I must be really bad at explaining things, or I am losing a fight\n> against preconceptions fixed beyond my imagination.\n\nI really dont' see the point. But hey, code talks. \n\n\t\tLinus\n"},{"id":"48080","messageId":"85r6n1p7sb.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707211737090.3607@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-22T01:05:24Z","receivedAt":"2007-07-22T01:05:24Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Sun, 22 Jul 2007, David Kastrup wrote:\n>> \n>> I must be really bad at explaining things, or I am losing a fight\n>> against preconceptions fixed beyond my imagination.\n>\n> I really dont' see the point. But hey, code talks. \n\nYes, I am working on that.  It would have been nice if IS_DIR was not\nalready taken by trees, but one can't have everything.  So I need to\ndecide how to represent the node, and it would appear that I need to\nangle for \"file\" after all.  Since it is really quite closer to a file\nor symlink than to a tree or project.  Hm, perhaps a symlink might be\nmore expedient.  Make it have an empty reference, and it is unique.\nAnd there will be fewer places in the code manipulating symlinks than\nfiles.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48081","messageId":"f7uap7$eo1$1@sea.gmane.org","threadId":"9086","inReplyTo":"85644dqoig.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-07-22T01:16:47Z","receivedAt":"2007-07-22T01:16:47Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"David Kastrup wrote:\n\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n> \n>> On Sat, 21 Jul 2007, David Kastrup wrote:\n\n>>> Linus, a directory is simply non-existent inside of git.\n>>\n>> You need to learn git first.\n>>\n>> A directory doesn't exist IN THE INDEX (until my patches). But you\n>> need to learn about the object database and the SHA1's. That's the\n>> real meat of git, and it sure as hell knows about directories.\n> \n> I have written up a complete explanation about the underlying concept\n> in a separate thread, maybe it would make sense reading that before\n> investing too much time meddling over details that don't fit the large\n> picture.  The point is that the object database and the SHA1 values\n> track _trees_, not _directories_.  And a _tree_ is just a hashing\n> mechanism in the repository for files.  Its existence is solely\n> dependent on the existence of its contents.  The only synchronization\n> with directories is that when a tree becomes empty, git attempts to do\n> an rmdir on the corresponding directory.  And of course, if git needs\n> to check out a file, it creates the necessary parent directories.\n> \n> Now since the physical _contents_ of a directory are already tracked\n> in _trees_ by git, the only missing part is the _existence_ of the\n> directory itself: a directory must exist as long as there is a tree\n> (and thus content) connected with it, but the reverse does not hold:\n> without a tree, the directory can still exist.  Which we can represent\n> by a repository entry named \".\" without content (the content is\n> already catered for by the _tree_).  This must _not_ be represented by\n> a _tree_ node since there is no content, and a tree without content by\n> _definition_ does not exist.\n> \n> I must be really bad at explaining things, or I am losing a fight\n> against preconceptions fixed beyond my imagination.\n\nI don't understand you, or you don't understand git. \"Tree\" object\nin object database (in repository) represents a directory in the\nworking area. There was never any problem with having empty trees\nin object database, or having links to empty directory in the superdir.\nWe don't have to change anything about object database.\n\nThe problems with git problems with empty directories stems from the\nfact that index didn't have directories. Index is flattened version\nof root tree, and before subproject support it contained _only_ info\nabout blobs (file contents). At least till Linus patch...\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"48083","messageId":"85myxpp67k.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"f7uap7$eo1$1@sea.gmane.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-22T01:39:27Z","receivedAt":"2007-07-22T01:39:27Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> writes:\n\n> David Kastrup wrote:\n>\n>> I must be really bad at explaining things, or I am losing a fight\n>> against preconceptions fixed beyond my imagination.\n>\n> I don't understand you, or you don't understand git. \"Tree\" object\n> in object database (in repository) represents a directory in the\n> working area. There was never any problem with having empty trees in\n> object database, or having links to empty directory in the superdir.\n> We don't have to change anything about object database.\n\nI disagree here.  The object database _can_ represent an _empty_\ndirectory that has been added explicitly, because up to now no\noperations existed that actually left an empty tree.  But it can't\ndistinguish a _non_-empty directory that has been added explicitly\nfrom non-empty directory that has not been added explicitly.\n\nTo wit: after the sequence\n\nmkdir a\ntouch a/b\ngit-add a\ngit-commit -m x\ngit-rm a/b\ngit-commit -m x\n\nI expect git to retain an empty directory a.  But the _tree_ now can't\nbe different from the tree in the situation\n\nmkdir a\ntouch a/b\ngit-add a/b\ngit-commit -m x\ngit-rm a/b\ngit-commit -m x\n\nbecause after step 1, the trees have identical contents, and so there\nis nothing at the _identical_ step 2 that could cause different\nbehavior.\n\nBut in the second case, git must _not_ retain a.  So we need to record\nthe information that in the first case, a was added explicitly.  And\nthis can't be done with the current repository layout.  It doesn't buy\nus anything that we _have_ a representation available for an _empty_\ntree added explicitly.  We need this \"added explicitly\" information\nfor _every_ tree, not just empty ones.\n\nAnd a perfectly consistent way is to make those trees with an\nexplicitly added directory _non-empty_, by virtue of putting a file\n\".\" in them.  This file, of course, exists in every physical\ndirectory, but we may or may not decide to let it be tracked by git,\nusing the gitignore mechanism on the pattern \".\".  Perfectly\nexpedient.\n\n> The problems with git problems with empty directories stems from the\n> fact that index didn't have directories.\n\nThat basically implies that no information about directories could be\ntracked in the repository.  And yes, we need appropriate information\nin the index.  Again, the information whether a directory was added\nexplicitly.\n\n> Index is flattened version of root tree, and before subproject\n> support it contained _only_ info about blobs (file contents).\n\nAnd the repository is a versioned and hierarchically hashed version of\nthe index, but its trees contain _no_ information that is not already\ninherently represented by the files alone.  Permitting empty trees\nwould change that fundamental property, and it would not buy us the\nability to actually track directories: see above.  So it is not worth\nthe trouble to assign any meaningful concept to persisting empty trees\nrather than make them a case for git-fsck.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48084","messageId":"alpine.LFD.0.999.0707211840000.3607@woody.linux-foundation.org","threadId":"9086","inReplyTo":"85r6n1p7sb.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-22T01:41:27Z","receivedAt":"2007-07-22T01:41:27Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 22 Jul 2007, David Kastrup wrote:\n>  Make it have an empty reference, and it is unique.\n\nI *really* don't see the point.\n\nAnd you seem to have igored totally my treatise on \"content\" and how the \nstuff git tracks must be stuff that is visible and detectable in the \ntrees. And if I understand you correctly, you also wouldn't be backwards \ncompatible. \n\nIOW, there's a lot of \"why's\" at all levels.\n\nI don't see the *point*. What's the problem you're trying to solve?\n\n\t\tLinus\n"},{"id":"48085","messageId":"85fy3hp3f2.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707211840000.3607@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-22T02:39:45Z","receivedAt":"2007-07-22T02:39:45Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Sun, 22 Jul 2007, David Kastrup wrote:\n>>  Make it have an empty reference, and it is unique.\n>\n> I *really* don't see the point.\n>\n> And you seem to have igored totally my treatise on \"content\" and how\n> the stuff git tracks must be stuff that is visible and detectable in\n> the trees.\n\nOh please.  Just because you refuse to read a point-to-point reply\ndoes not mean it has not been made.\n\n\".\" _is_ visible and detectable in every tree.  But that does not mean\nit is automatically tracked by git unless it gets added explicitly, or\nimplicitly (as long as the gitignore mechanism does not kick in) by\nadding a higher level directory.\n\nIf a file does not get added explicitly or implicitly, it does not end\nup in the repository and git behaves like it knows nothing about it.\n\nAnd that's just the way it is going to be with directories.  Nothing\nmore, nothing less, nothing new.\n\n> And if I understand you correctly, you also wouldn't be backwards\n> compatible.\n\nDefine backwards compatible.  Anyway, you are the repository wizard:\nhere are the semantics I need supported for backwards compatibility:\n\nI need an entry type in the index and in the repository with the\nfollowing features:\n\na) if part of a tree, the tree is not considered empty.  Should be\n   easy.\nb) it has the name \".\".  This is not absolutely necessary, but it\n   means that the gitignore mechanism can be used for dealing with it,\n   and that's intuitive and has exactly the expressive power required\n   for the job.  Now the gitignore mechanism is isolated very locally\n   in dir.c: whether one makes the actual representation in the\n   repository based on an attribute like \"filemode\" rather than on a\n   separate entry does not actually complicate the code all too much.\n   There is, however, some level of complication since the consulted\n   .gitignore file for ignoring \".\" must, of course, be the .gitignore\n   file situated _in_ the directory.  So making \".\" sit _in_ the tree\n   rather than _on_ the tree simplifies the code considerably.  It is\n   a small amount of code, nevertheless, so it is not a major\n   strategic decision.\n\n   One conceivable implementation would be indeed similar to what the\n   \"filemode\" thing does: let us keep open the option to track, at one\n   time, permissions.  The current format has, as far as I understand,\n   all zeros in the permissions field of trees (I have not checked,\n   though).  Now if we stipulate that this is the kind of directory\n   permissions we will in all eternity _not_ support outside of git,\n   we are all set with regard to backwards compatibility: a tree with\n   permissions all zero will behave as previously: it will get removed\n   when it becomes empty (taking the corresponding work tree directory\n   with it, if possible).  And that's it.  But a tree with nonzero\n   permissions (whether they correspond to outward permissions or are\n   just a placeholder) will _not_ evaporate when becoming empty.  It\n   will be possible to explicitly or implicitly delete it: that will\n   just set its permissions all to zero so that it has the chance to\n   evaporate next time it becomes empty.\n\n> IOW, there's a lot of \"why's\" at all levels.\n>\n> I don't see the *point*. What's the problem you're trying to solve?\n\nrm -rf ./*\ngit-commit -m \"all empty\" -a\nunzip /tmp/something-with-empty-dirs.zip\ngit-add .\ngit-commit -m \"something-with-empty-dirs\"\ngit-checkout HEAD~1\n# Now I don't want empty directories and their parents lying around.\ngit-checkout master\n# Now the state after unzip should be restored faithfully\nrm -rf ./*\nunzip /tmp/something-else-with-empty-dirs\ngit-commit -a -m \"something-else\"\n# Now I want to have the state of something-else registered faithfully\n# even if it contains top-level files and directories not present in\n# something-with-empty-dirs, because supposedly . is being tracked,\n# not just every file element in it.\n\nActually, oops.  This last criterion is not met when .'s relation to\nthe tree is such that it is only considered _part_ of tree.\n\nLooks like it might be prudent to focus on the permissions-coupled\nrepresentation.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48087","messageId":"alpine.LFD.0.999.0707212040340.3607@woody.linux-foundation.org","threadId":"9086","inReplyTo":"85fy3hp3f2.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-22T03:43:11Z","receivedAt":"2007-07-22T03:43:11Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 22 Jul 2007, David Kastrup wrote:\n>\n> \".\" _is_ visible and detectable in every tree.\n\nI'm going to add you to my \"clueless\" filter, because it's not worth my \ntime to answr you any more.\n\nI told you. Several times. That \".\" is pointless exactly because it's in \n_every_ tree, and as such is no longer \"content\". It's not something that \nthe user can care about, because it has no meaning. There's no point in \ntracking it, because even if we do *not* track it, it's there, and we \ncannot do anything about it.\n\nThat was the whole difference between \".\" and \".gitignore\", and I \nexplicitly pointed out that that was the difference (and the _only_ one), \nand why it mattered.\n\nAnd you didn't listen. And now you claim that I don't read your emails. I \ndo. They just don't make any sense.\n\nConsider this discussion ended. I simply don't care any more.\n\n\t\tLinus\n"},{"id":"48088","messageId":"1DB4C18A-A21B-4176-85AD-A86C92178F54@silverinsanity.com","threadId":"9086","inReplyTo":"85tzrxslms.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Brian Gernhardt","fromEmail":"benji@silverinsanity.com","sentAt":"2007-07-22T04:00:30Z","receivedAt":"2007-07-22T04:00:30Z","isPatch":true,"sender":{"key":"benji@silverinsanity.com","avatar":"https://gravatar.com/avatar/e06c101dbc25c68114d859b4a9ec7cf8a2c52fd2b0270ef0eac0e2e63ff22311?d=mp&s=160"},"body":"\nOn Jul 21, 2007, at 1:38 PM, David Kastrup wrote:\n\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n>\n>> So git filenames are very much a \"stream of bytes\", not anything\n>> else. And they need to sort 100% reliably, always the same way, and\n>> never with any localized meaning.\n>\n> There is some utf-8/Unicode trouble to be expected in connection with\n> that eventually: some, but not all operating and/or file systems\n> canonicalize file names, replacing accented letters by a combining\n> accent and the letter.  But that's beside the point.\n\nThis issue exists today.  OS X does a number of things to filenames,  \none of which is normalizing all UTF.  The resulting error is wholly  \nnon-intuitive, but easy to solve.  Git thinks both that the file  \nexists under the name it expects and that the file is being ignored  \nas the name OS X uses.  The solution is to put the OS X normalized  \nform into .git/info/exclude.  Any other solution involves platform- \ndependent hackery and inclusion of Unicode libraries.  I perused this  \nfor a short while some months ago, but was convinced to leave it be.\n\n~~ Brian\n"},{"id":"48089","messageId":"85abtpoydg.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707212040340.3607@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-22T04:28:43Z","receivedAt":"2007-07-22T04:28:43Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Sun, 22 Jul 2007, David Kastrup wrote:\n>>\n>> \".\" _is_ visible and detectable in every tree.\n>\n> I'm going to add you to my \"clueless\" filter, because it's not worth\n> my time to answr you any more.\n\nToo bad I can't do the same.\n\n> I told you. Several times. That \".\" is pointless exactly because\n> it's in _every_ tree, and as such is no longer \"content\".\n\n\".\" is in every _non-empty_ directory tree.  But we are talking about\npermitting _empty_ trees in the repository.  And for an empty tree in\nthe repository, \".\" may or may not be in the corresponding work\ndirectory tree, depending on whether the directory exists or not.  So\nwhen we are talking about a repository tree _becoming_ empty, we need\nthe information whether or whether not we should remove it upon\nbecoming empty.  _That_ is the information content of \".\" being or not\nbeing considered part of the trackable material.  And the information\nis no longer available at the time the repository tree becomes empty\n_unless_ we already store it there when the tree is still populated.\n\n> It's not something that the user can care about, because it has no\n> meaning. There's no point in tracking it, because even if we do\n> *not* track it, it's there, and we cannot do anything about it.\n\nOk, here we go _again_.  Test case 1:\n\nmkdir a\ntouch a/b\ngit-add a/b\ngit-commit -m x\ngit-rm a/b\ngit-commit -m x\n\nNow we want to have the directory a _removed_.\n\nTest case 2:\n\nmkdir a\ntouch a/b\ngit-add a\ngit-commit -m x\ngit-rm a/b\ngit-commit -m x\n\nNow we want to have the directory a _retained_.\n\nAfter the first commit in _both_ test cases, the only file in the\ntrees / and /a is a/b.  The working directory state is _identical_ at\nthis point, and we do identical commands afterwards.\n\nThe end result is not identical, so there must be some information\ndifferent in the repository after the first commit.  This information\n_can't_ be encoded in a remaining empty tree, because both the trees /\nand /a are _non_-empty yet.\n\nSo we _must_ encode the evaporate-or-not-when-empty information\n_otherwise_ into the repository.  And we do that by _not_ having\n/a/. in the set of tracked files in test case 1, and by _having_ it in\nthe set of tracked files in test case 2.\n\n> That was the whole difference between \".\" and \".gitignore\", and I\n> explicitly pointed out that that was the difference (and the _only_\n> one), and why it mattered.\n\nYou are underestimating the power of \".gitignore\": while it is true\nthat its _physical_ presence will reliably keep git from removing the\ndirectory, its physical presence is not _actually_ required.\n\nIt is sufficient that git _believes_ in its continuing physical\nexistence.  And if we tell it \"it is still there\" whenever it takes a\nlook, then git will keep the record of .gitignore in its tree, and\nconsequently won't remove the tree and not try deleting the directory.\nHowever, once we explicitly tell it \"remove the record of .gitignore\nfrom the repository\", it will do so, and in the course of doing so\nremove the directory in the work directory together with the tree in\nthe repository.\n\n>From a user interface and logical standpoint, adding or not adding \".\"\nto the tracked content is a perfectly consistent and convenient way of\nhaving the directory kept around or not.\n\n>From the viewpoint of the internal data structures, I'll likely go\nwith tampering with (pseudo-)permissions.\n\n> And you didn't listen. And now you claim that I don't read your\n> emails. I do. They just don't make any sense.\n>\n> Consider this discussion ended. I simply don't care any more.\n\nIt is painfully clear that I could invest a few weeks of time in\ncoding better than in explaining stuff.  And I guess that's what I'll\nhave to do.  And afterwards it will be your job to wrack your head\nabout why something does all the right things for the wrong reasons\nand come up with a different explanation how and why the code works.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48095","messageId":"Pine.LNX.4.64.0707212332530.6350@asgard.lang.hm","threadId":"9086","inReplyTo":"85abtpoydg.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"","fromEmail":"david@lang.hm","sentAt":"2007-07-22T06:38:41Z","receivedAt":"2007-07-22T06:38:41Z","isPatch":true,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Sun, 22 Jul 2007, David Kastrup wrote:\n\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n>\n>> On Sun, 22 Jul 2007, David Kastrup wrote:\n>>>\n>>> \".\" _is_ visible and detectable in every tree.\n>>\n>> I'm going to add you to my \"clueless\" filter, because it's not worth\n>> my time to answr you any more.\n>\n> Too bad I can't do the same.\n>\n>> I told you. Several times. That \".\" is pointless exactly because\n>> it's in _every_ tree, and as such is no longer \"content\".\n>\n> \".\" is in every _non-empty_ directory tree.  But we are talking about\n> permitting _empty_ trees in the repository.  And for an empty tree in\n> the repository, \".\" may or may not be in the corresponding work\n> directory tree, depending on whether the directory exists or not.  So\n> when we are talking about a repository tree _becoming_ empty, we need\n> the information whether or whether not we should remove it upon\n> becoming empty.  _That_ is the information content of \".\" being or not\n> being considered part of the trackable material.  And the information\n> is no longer available at the time the repository tree becomes empty\n> _unless_ we already store it there when the tree is still populated.\n\nDavid, the point where you and Linus are talking past each other is that \nLinus is assuming that you only want to track some specific directories, \nand for that tracking \".\" doesn't work becouse it's in every directory\n\nyou apparently consider every directory equal and therefor the fact that \n\".\" exists in every directory doesn't bother you becouse you want to track \nevery directory.\n\nwhat you are not hearing is that while Linus and the other git developers \ncan see reasons to track directories sometimes, they definantly don't \nagree that you want to track directories all the time.\n\nsometimes the fact that a directory exists is significant, most of the \ntime it's not. and the difference between what is and what isn't \nsignificant isn't a per-repository or per-project thing, it's a \nper-directory thing.\n\nin one repository you will have some directories that only exist becouse \nfiles are in them, and you may have some directories that exist becouse \nyou explicitly want them to exist.\n\nboth types have the \".\" file in them (or appear to, some OS's/filesystems \ndon't actually have a \".\" on disk, they add it when needed when reporting \nto userspace), so git has no way to tell which ones you explicitly want \ntracked.\n\ncreating .gitignore in the directories that you want tracked lets the \nother directories not be trackes.\n\nDavid Lang\n\n>> It's not something that the user can care about, because it has no\n>> meaning. There's no point in tracking it, because even if we do\n>> *not* track it, it's there, and we cannot do anything about it.\n>\n> Ok, here we go _again_.  Test case 1:\n>\n> mkdir a\n> touch a/b\n> git-add a/b\n> git-commit -m x\n> git-rm a/b\n> git-commit -m x\n>\n> Now we want to have the directory a _removed_.\n>\n> Test case 2:\n>\n> mkdir a\n> touch a/b\n> git-add a\n> git-commit -m x\n> git-rm a/b\n> git-commit -m x\n>\n> Now we want to have the directory a _retained_.\n>\n> After the first commit in _both_ test cases, the only file in the\n> trees / and /a is a/b.  The working directory state is _identical_ at\n> this point, and we do identical commands afterwards.\n>\n> The end result is not identical, so there must be some information\n> different in the repository after the first commit.  This information\n> _can't_ be encoded in a remaining empty tree, because both the trees /\n> and /a are _non_-empty yet.\n>\n> So we _must_ encode the evaporate-or-not-when-empty information\n> _otherwise_ into the repository.  And we do that by _not_ having\n> /a/. in the set of tracked files in test case 1, and by _having_ it in\n> the set of tracked files in test case 2.\n>\n>> That was the whole difference between \".\" and \".gitignore\", and I\n>> explicitly pointed out that that was the difference (and the _only_\n>> one), and why it mattered.\n>\n> You are underestimating the power of \".gitignore\": while it is true\n> that its _physical_ presence will reliably keep git from removing the\n> directory, its physical presence is not _actually_ required.\n>\n> It is sufficient that git _believes_ in its continuing physical\n> existence.  And if we tell it \"it is still there\" whenever it takes a\n> look, then git will keep the record of .gitignore in its tree, and\n> consequently won't remove the tree and not try deleting the directory.\n> However, once we explicitly tell it \"remove the record of .gitignore\n> from the repository\", it will do so, and in the course of doing so\n> remove the directory in the work directory together with the tree in\n> the repository.\n>\n> From a user interface and logical standpoint, adding or not adding \".\"\n> to the tracked content is a perfectly consistent and convenient way of\n> having the directory kept around or not.\n>\n> From the viewpoint of the internal data structures, I'll likely go\n> with tampering with (pseudo-)permissions.\n>\n>> And you didn't listen. And now you claim that I don't read your\n>> emails. I do. They just don't make any sense.\n>>\n>> Consider this discussion ended. I simply don't care any more.\n>\n> It is painfully clear that I could invest a few weeks of time in\n> coding better than in explaining stuff.  And I guess that's what I'll\n> have to do.  And afterwards it will be your job to wrack your head\n> about why something does all the right things for the wrong reasons\n> and come up with a different explanation how and why the code works.\n>\n>\n"},{"id":"48111","messageId":"851wf0pzyt.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"Pine.LNX.4.64.0707212332530.6350@asgard.lang.hm","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-22T09:08:58Z","receivedAt":"2007-07-22T09:08:58Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"david@lang.hm writes:\n\n> On Sun, 22 Jul 2007, David Kastrup wrote:\n>\n>> Linus Torvalds <torvalds@linux-foundation.org> writes:\n>>\n>>> I told you. Several times. That \".\" is pointless exactly because\n>>> it's in _every_ tree, and as such is no longer \"content\".\n>>\n>> \".\" is in every _non-empty_ directory tree.  But we are talking\n>> about permitting _empty_ trees in the repository.  And for an empty\n>> tree in the repository, \".\" may or may not be in the corresponding\n>> work directory tree, depending on whether the directory exists or\n>> not.  So when we are talking about a repository tree _becoming_\n>> empty, we need the information whether or whether not we should\n>> remove it upon becoming empty.  _That_ is the information content\n>> of \".\" being or not being considered part of the trackable\n>> material.  And the information is no longer available at the time\n>> the repository tree becomes empty _unless_ we already store it\n>> there when the tree is still populated.\n>\n> David, the point where you and Linus are talking past each other is\n> that Linus is assuming that you only want to track some specific\n> directories, and for that tracking \".\" doesn't work becouse it's in\n> every directory\n>\n> you apparently consider every directory equal and therefor the fact\n> that \".\" exists in every directory doesn't bother you becouse you\n> want to track every directory.\n\nSigh.  No, I don't want to track every directory.  I want to have\nevery directory _trackable_.  Whether it is _tracked_ depends on\nwhether you _add_ it to the index.  And that depends, among other\nthings, on the gitignore patterns, and those can be specified on a\nper-directory, per-project, per-user preference.\n\n> what you are not hearing is that while Linus and the other git\n> developers can see reasons to track directories sometimes, they\n> definantly don't agree that you want to track directories all the\n> time.\n\nAnd that is why one can use per-directory, per-project and per-user\nsettings to turn the tracking off, _and_ one can decide at what level\none adds information to the index.  If you always make it a habit to\nonly ever use git-add -f and git-rm -f on _files_ and never on\ndirectories, you won't _ever_ see a difference on whether directories\nare tracked, and the contents of .gitignore won't make a difference,\neither.\n\nBut if you use git-add and git-rm on directories, then for the\nspecified directory and its children, .gitignore gets consulted.\n\n> sometimes the fact that a directory exists is significant, most of\n> the time it's not. and the difference between what is and what isn't\n> significant isn't a per-repository or per-project thing, it's a\n> per-directory thing.\n\nWhich is why one can control it per-directory using either the\n.gitignore mechanism _or_ by including the directory level in question\nin the git-add and git-rm commands or not.\n\n> in one repository you will have some directories that only exist\n> becouse files are in them, and you may have some directories that\n> exist becouse you explicitly want them to exist.\n>\n> both types have the \".\" file in them (or appear to, some\n> OS's/filesystems don't actually have a \".\" on disk, they add it when\n> needed when reporting to userspace), so git has no way to tell which\n> ones you explicitly want tracked.\n\nLike with any other file, git _has_ a way to tell.  If I don't git-add\nor git-rm the directory or one of its parents to the index, I don't\nwant to have it tracked.  And if I add the directory or one of its\nparents to the index recursively, but it is covered by .gitignore, I\ndon't want to have it tracked.\n\nIt is a pity that you have seemingly not read on, because there\nfollows a simple example:\n\n>> Ok, here we go _again_.  Test case 1:\n>>\n>> mkdir a\n>> touch a/b\n>> git-add a/b\n>> git-commit -m x\n>> git-rm a/b\n>> git-commit -m x\n>>\n>> Now we want to have the directory a _removed_.\n>>\n>> Test case 2:\n>>\n>> mkdir a\n>> touch a/b\n>> git-add a\n>> git-commit -m x\n>> git-rm a/b\n>> git-commit -m x\n>>\n>> Now we want to have the directory a _retained_.\n>>\n>> After the first commit in _both_ test cases, the only file in the\n>> trees / and /a is a/b.  The working directory state is _identical_ at\n>> this point, and we do identical commands afterwards.\n>>\n>> The end result is not identical, so there must be some information\n>> different in the repository after the first commit.  This information\n>> _can't_ be encoded in a remaining empty tree, because both the trees /\n>> and /a are _non_-empty yet.\n>>\n>> So we _must_ encode the evaporate-or-not-when-empty information\n>> _otherwise_ into the repository.  And we do that by _not_ having\n>> /a/. in the set of tracked files in test case 1, and by _having_ it in\n>> the set of tracked files in test case 2.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48125","messageId":"200707221406.25541.jnareb@gmail.com","threadId":"9086","inReplyTo":"85myxpp67k.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-07-22T12:06:24Z","receivedAt":"2007-07-22T12:06:24Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Sun, 22 July 2007, David Kastrup wrote:\n> Jakub Narebski <jnareb@gmail.com> writes:\n>> David Kastrup wrote:\n>>\n>>> I must be really bad at explaining things, or I am losing a fight\n>>> against preconceptions fixed beyond my imagination.\n\nOr you are wrong...\n\n>> I don't understand you, or you don't understand git. \"Tree\" object\n>> in object database (in repository) represents a directory in the\n>> working area. There was never any problem with having empty trees in\n>> object database, or having links to empty directory in the superdir.\n>> We don't have to change anything about object database.\n> \n> I disagree here.  The object database _can_ represent an _empty_\n> directory that has been added explicitly, because up to now no\n> operations existed that actually left an empty tree.  But it can't\n> distinguish a _non_-empty directory that has been added explicitly\n> from non-empty directory that has not been added explicitly.\n\nTrue. I forgot about that.\n\nAlthough I'd rather say that we want distinguish between automatically \ncleaned up directory (directory which will be deleted if all files in \nit would be deleted, and would be untracked if all tracked files in it \nwould be deleted), and \"sticky\" directory, which is explicitely tracked \nand have to be explicitely deleted.\n\nThe fact that it was added explicitely or non explicitely is orthogonal \nto that.\n\nIMHO it would be best to first provide plumbing infrastructure (as e.g. \nit was the case of submodule support), then add option to \ngit-update-index to change the \"stickiness\"/\"autoremoval\" status of a \ndirectory (of a tree), and _last_ think about how to change the \nporcelain (git-add and git-rm).\n\n[...]\n> But in the second case, git must _not_ retain a.  So we need to record\n> the information that in the first case, a was added explicitly.  And\n> this can't be done with the current repository layout.  It doesn't buy\n> us anything that we _have_ a representation available for an _empty_\n> tree added explicitly.  We need this \"added explicitly\" information\n> for _every_ tree, not just empty ones.\n> \n> And a perfectly consistent way is to make those trees with an\n> explicitly added directory _non-empty_, by virtue of putting a file\n> \".\" in them.  This file, of course, exists in every physical\n> directory, but we may or may not decide to let it be tracked by git,\n> using the gitignore mechanism on the pattern \".\".  Perfectly\n> expedient.\n\nHere we disagree. I think putting \".\" in a tree as marker of having it \nnot be automatically deleted when empty, as opposed to marking tree \nusing filemode in the parent, is not a good idea.\n\nThe only advantage to the \".\" idea is that it can use gitignore \nmechanism (both in-tree .gitignore, tracked or not, and info/exclude \nfile). But I also think that the fact that gitignore mechanism is \nrecursive is more of disadvantage than advantage.\n\nFirst, it is _not_ consistent. Working directory trees _always_ have '.' \nin them, while trees would have or would have not it, depending if they \nwould be \"sticky\" or \"autoremoved\".\n\nSecond, the \"easy implementation\" is anything but easy. \"git add .\" as\na way to mark directory as \"sticky\" is not backward compatibile: \ncurrently it mean to add _all contents_ of current directory. \nImplementation is tricky: as we have seen trying to unlink '.' or \ncreate '.' can unfortunately succeed on [some Sun OS, and UFS \nfilesystem] (which follows POSIX stupidly to the letter) f**king\nup the filesystem. The alternative proposal of adding \"magic mode\" to \nmark directory as \"not remove when empty\" is largely tested; it is very \nsimilar to the subproject support.\n\nThird, is contrary to the git philosophy of tracking contents. \n\"Stickiness\" is an attribute; the fact that directory is explicitely \ntracked or not does not change contents of a directory. Compare to \n'blob' which contains only contents of a file: not a filename, not a \npathname, not [subset of] filemode.\n\nFourth, is very artificial. What would you put for filemode for '.'?\n040000 (i.e. directory)? What would you put for sha1? Sha1 of an empty \ndirectory? Of an empty blob? 0{40} (which is bad idea because \ngit-diff-tree uses 0{40} to represent 'not existance')?\n\n>> The problems with git problems with empty directories stems from the\n>> fact that index didn't have directories.\n> \n> That basically implies that no information about directories could be\n> tracked in the repository.  And yes, we need appropriate information\n> in the index.  Again, the information whether a directory was added\n> explicitly.\n\nWhether directory is automatically managed by git (automatically removed \nor untracked). But we need directory entry in index for git-diff, for \nexample to recognize if there is or there is not empty directory, or if \na directory is automanaged or not.\n \n>> Index is flattened version of root tree, and before subproject\n>> support it contained _only_ info about blobs (file contents).\n> \n> And the repository is a versioned and hierarchically hashed version of\n> the index, but its trees contain _no_ information that is not already\n> inherently represented by the files alone. [...]\n\nThe above sentence is nonsensical. Index is helper for repository,\nand can be derived from repository. Not vice versa.\n\nTrees do contain information which is not inherently present by the \nblobs.\n-- \nJakub Narebski\nPoland\n"},{"id":"48142","messageId":"857iosmto0.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"200707221406.25541.jnareb@gmail.com","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-22T13:53:19Z","receivedAt":"2007-07-22T13:53:19Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"\nJakub, this mail is too long already, and it does not make sense to\ntack a changed proposal to its end since then the readers will be\nexhausted at the time they come there.  So I'll instead tack a\nfollowup to the \"big picture\" mail instead where I outline a modified\napproach which is presumably easier to understand and completely\nbackwards-compatible, incorporating your feedback.\n\nThere is probably little sense in wasting your time on a detailed\nresponse: feel free to point out where you don't see myself making\nsense.  I have no problem with people coming to different conclusions\nthat I do, but I would prefer it if it is not because they consider\nmyself a raving lunatic, but because they have different opinions\nregarding the details.\n\n\"I can follow you, but I disagree with your conclusion\" is perfectly\nfine for now since I am going to propose something else, anyway.\n\nThanks for the feedback.  It gave me some good ideas.\n\nJakub Narebski <jnareb@gmail.com> writes:\n\n> On Sun, 22 July 2007, David Kastrup wrote:\n>> Jakub Narebski <jnareb@gmail.com> writes:\n>>> David Kastrup wrote:\n>>>\n>>>> I must be really bad at explaining things, or I am losing a fight\n>>>> against preconceptions fixed beyond my imagination.\n>\n> Or you are wrong...\n\nWell, there is little reason for you to take my word on it, but I\nhappen to have a history of designing and implementing systems where I\nhave been responsible for every single byte, bootloader, firmware,\napplications, target compiler, assembler, whatever.  I have been\nexposed to Unix and working with it several years before Linux even\nexisted.  I also have a track record of being not exactly stupid.\n\nSo I pretty much can rule out that I am wrong on the factual side.\n\nBut where I may be wrong is in estimating the how obvious the design\ncan appear to others, and how useful and maintainable for others it\nmay be in the long run.  Linus says \"code talks\", but that's actually\nnot half the story.  If my code says that it works and the evidence is\nthere, but nobody is able to understand _why_ it works, it has no\nplace in a project where I am not permanently around.\n\nIf smart people don't get what I am talking about, it does not matter\nthat the patch is surprisingly well-contained: it will be a\nmaintenance nightmare because people will never figure out why\nsomething stopped working after some particular change.\n\n>> I disagree here.  The object database _can_ represent an _empty_\n>> directory that has been added explicitly, because up to now no\n>> operations existed that actually left an empty tree.  But it can't\n>> distinguish a _non_-empty directory that has been added explicitly\n>> from non-empty directory that has not been added explicitly.\n>\n> True. I forgot about that.\n\nThanks.  It is almost a revelation that anybody can agree on any point\nwith me at the moment.\n\n> IMHO it would be best to first provide plumbing infrastructure (as\n> e.g.  it was the case of submodule support), then add option to\n> git-update-index to change the \"stickiness\"/\"autoremoval\" status of\n> a directory (of a tree), and _last_ think about how to change the\n> porcelain (git-add and git-rm).\n\nSure.  It does no harm to think about reducing the amount of breaking\nporcelain, though.\n\n> [...]\n>\n>> And a perfectly consistent way is to make those trees with an\n>> explicitly added directory _non-empty_, by virtue of putting a file\n>> \".\" in them.  This file, of course, exists in every physical\n>> directory, but we may or may not decide to let it be tracked by\n>> git, using the gitignore mechanism on the pattern \".\".  Perfectly\n>> expedient.\n>\n> Here we disagree. I think putting \".\" in a tree as marker of having\n> it not be automatically deleted when empty, as opposed to marking\n> tree using filemode in the parent, is not a good idea.\n\nWell, \"not a good idea\" is a far step forward from \"stupid idiot\nbabbling nonsense\", so we may make progress towards actually being\nable to _weigh_ different options.  I can actually associate with \"not\na good idea\", not least because nobody else seems to get the idea, and\nthat makes it infeasible for maintenance.\n\nSo I'll address some points and then propose a different way of\nimplementing what will in the end amount to rather similar semantics,\nbut with a different view of looking at those semantics, one that\ncorresponds well with the implementation.\n\n> The only advantage to the \".\" idea is that it can use gitignore\n> mechanism (both in-tree .gitignore, tracked or not, and info/exclude\n> file). But I also think that the fact that gitignore mechanism is\n> recursive is more of disadvantage than advantage.\n>\n> First, it is _not_ consistent. Working directory trees _always_ have\n> '.'  in them, while trees would have or would have not it, depending\n> if they would be \"sticky\" or \"autoremoved\".\n\nLet me point out again that this inconsistency is already present in\nthe difference of tracked and untracked _files_: they are always in\nthe working directory, while trees have or not have them, depending on\nwhether they are \"registered\" or \"not\".\n\nThere is no inconsistency involved here, but it seems to make people\n_very_ uncomfortable to factor out the \"stays around even if empty\"\nfunctionality and call it \"dir/.\" from the \"can hold content\"\nfunctionality which is in effect called \"dir/\", and basically\nassociate tracked physical existence just with the former.\n\nThe recursiveness of the gitignore mechanism has the advantage that\nwhen maintaining a large repository with actual or logical\nsubprojects, one does not need to pick a single policy for all\nsubprojects.  I think that is quite important.  It could possibly be\nachieved with some other method of having per-subproject\nconfiguration, but I see little wrong in using what is there and\ndocumented already.\n\n> Second, the \"easy implementation\" is anything but easy. \"git add .\"\n> as a way to mark directory as \"sticky\" is not backward compatibile:\n> currently it mean to add _all contents_ of current directory.\n> Implementation is tricky: as we have seen trying to unlink '.' or\n> create '.' can unfortunately succeed on [some Sun OS, and UFS\n> filesystem] (which follows POSIX stupidly to the letter) f**king up\n> the filesystem.\n\nI was not suggesting actually leaving any such calls in place: after\nall, they would presumably lead to error messages.  But I agree that\nthis could lead to nasty surprises when somebody with a legacy version\nof git worked with a repository containing \".\" as explicit entries of\nsome file type.\n\n> The alternative proposal of adding \"magic mode\" to mark directory as\n> \"not remove when empty\" is largely tested; it is very similar to the\n> subproject support.\n\nGood.  Because it is what I converged to last night.\n\n> Third, is contrary to the git philosophy of tracking contents.\n> \"Stickiness\" is an attribute; the fact that directory is explicitely\n> tracked or not does not change contents of a directory. Compare to\n> 'blob' which contains only contents of a file: not a filename, not a\n> pathname, not [subset of] filemode.\n>\n> Fourth, is very artificial. What would you put for filemode for '.'?\n> 040000 (i.e. directory)?\n\nTaken already.  By something very artificial, namely a tree...  Yes,\nthis was a wart in my proposal.\n\n> What would you put for sha1?  Sha1 of an empty directory?\n\nSome fixed value.  Everywhere the same.  Not really relevant.\n\n>> That basically implies that no information about directories could\n>> be tracked in the repository.  And yes, we need appropriate\n>> information in the index.  Again, the information whether a\n>> directory was added explicitly.\n>\n> Whether directory is automatically managed by git (automatically\n> removed or untracked). But we need directory entry in index for\n> git-diff, for example to recognize if there is or there is not empty\n> directory, or if a directory is automanaged or not.\n\nOne conclusion that I have come to (and I think I am in agreement with\nLinus here) is that the information \"empty or not\" is actually useless\nseparately: when I add files below a directory to the repository, the\ndirectory _can't_ be empty.  And git has no way of knowing whether it\nis non-empty because I wanted the directory to be there, or whether it\nis non-empty because I could not have checked in the files into the\ntree below it otherwise.\n\n>> And the repository is a versioned and hierarchically hashed version\n>> of the index, but its trees contain _no_ information that is not\n>> already inherently represented by the files alone. [...]\n>\n> The above sentence is nonsensical. Index is helper for repository,\n> and can be derived from repository. Not vice versa.\n>\n> Trees do contain information which is not inherently present by the \n> blobs.\n\nCould you give examples for such information?  As long as we are not\ntalking about _history_, I am at a loss at what else you mean.  File\nnames and permissions?\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48150","messageId":"alpine.LFD.0.999.0707221023530.3607@woody.linux-foundation.org","threadId":"9086","inReplyTo":"85abtpoydg.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-22T17:28:10Z","receivedAt":"2007-07-22T17:28:10Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 22 Jul 2007, David Kastrup wrote:\n> \n> > I told you. Several times. That \".\" is pointless exactly because\n> > it's in _every_ tree, and as such is no longer \"content\".\n> \n> \".\" is in every _non-empty_ directory tree.\n\nYou're pointless.\n\nWe have no problems at all with non-empty trees. We know exactly what they \nare. We keep track of them fine, and we do not need a totally pointless \n\".\" entry for them.\n\n>  But we are talking about\n> permitting _empty_ trees in the repository.\n\nAnd WE ALREADY DO.\n\nThe empty tree looks like this: \"\". It has a SHA1 of \n4b825dc642cb6eb9a060e54bf8d69288fbee4904. It works today, and in fact, git \nuses it already. \n\nTry this:\n\n\tgit ls-tree 4b825dc642cb6eb9a060e54bf8d69288fbee4904\n\nin the git repository. What do you think that is?\n\nYour \".\" is *pointless*.\n\nAnd it's _worse_ than pointless: it's not \"content\". It doesn't add any \ninformation. It's not something you can match up  against the working tree \nmeaningfully, exactly because *every* working tree has it. As such, it's \ntotal non-information.\n\n\t\tLinus\n"},{"id":"48151","messageId":"alpine.LFD.0.999.0707221029160.3607@woody.linux-foundation.org","threadId":"9086","inReplyTo":"851wf0pzyt.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-22T17:30:06Z","receivedAt":"2007-07-22T17:30:06Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 22 Jul 2007, David Kastrup wrote:\n>\n> Sigh.  No, I don't want to track every directory.  I want to have\n> every directory _trackable_.\n\nAnd they already are. \n\nYour point is pointless. You don't understand the git data structures, and \nyou are trying to do something that makes no sense.\n\n\t\tLinus\n"},{"id":"48152","messageId":"alpine.LFD.0.999.0707221031050.3607@woody.linux-foundation.org","threadId":"9086","inReplyTo":"85abtpoydg.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-22T17:33:30Z","receivedAt":"2007-07-22T17:33:30Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 22 Jul 2007, David Kastrup wrote:\n>\n> So  when we are talking about a repository tree _becoming_ empty, we \n> need the information whether or whether not we should remove it upon\n> becoming empty.\n\nYou don't seem to realize - although I've told you now abotu a million \ntimes - that what you are talking about is:\n\n - technically exactly the same as \".gitignore\", which for some \n   unfathomable reason you cannot seem to accept.\n\n - except your use of \".\" is 100% INFERIOR exactly because the \".\" entry \n   has no meaning in the target filesystem, so it means that the bit of \n   information is no longer something that is trackable in the working \n   tree.\n\nQuite frankly, Junio would be a total idiot to take any patches that do \nwhat you want to do. Happily, he is anything but.\n\n\t\tLinus\n"},{"id":"48154","messageId":"85tzrwiakt.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707221029160.3607@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-22T17:59:14Z","receivedAt":"2007-07-22T17:59:14Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Sun, 22 Jul 2007, David Kastrup wrote:\n>>\n>> Sigh.  No, I don't want to track every directory.  I want to have\n>> every directory _trackable_.\n>\n> And they already are.\n\nTheir contents are.\n\n> Your point is pointless. You don't understand the git data\n> structures, and you are trying to do something that makes no sense.\n\nThat makes no sense to you and apparently quite a few other people,\nafter a lot of explaining.  That does not mean that it wouldn't work,\nbut it does mean that it is going nowhere: it is irrelevant whether I\nconsider the concept easy to understand and explain when nobody else\ndoes: that makes it unmaintainable.\n\nFortunately, a few other participants, notably Junio and Jakub, have\nfocused a bit more on technical details rather than my sanity in their\nsomewhat more nuanced feedback, and thus I have (in a separate thread)\nmade a new proposal that addresses a few technical shortcomings and\nthat does no longer require splitting tree-ness/directory-ness into\nseparate concepts and records, something which I considered elegant\nand others gibberish.\n\nIt boils down to encoding the \"don't-evaporate-when-empty\" or \"I told\nyou to keep track of it\" property in the directory access permissions:\nif those are zero, git does not track the corresponding directory and\nwill attempt a remove-on-empty.  If they are non-zero (probably 755 as\nlong as git stores only a sanitized version of the actual state\nthere), this means that git has been told to track the directory and\nwill not attempt to delete it until it is told to stop tracking it\nagain.\n\nThe proposal of allowing \".\" \"!.\" as a gitignore pattern to specify\nthe tracking/non-tracking indicator does still stand, but its\nsemantics are now so much decoupled from that of\n\"don't-evaporate-when-empty\" that the code would not actually overlap\nwith that of the tracking, and so discussing it is orthogonal to the\nactual proposal and can be postponed separately, and an implementation\nproferred separately once the rest is in place.\n\nSo do both of us a favor and skip the rest of the mail queue with\n\"Empty directories...\" in its title.\n\nActually, the code (and later comments for it) you produced matches\nthe areas of work and what I think needs to be done quite closer now\nthan with my original proposal.\n\nSo while the discussion with you has not really been much of a help\nexcept to show without reasonable doubt that my original approach\nwould have been unmaintainable by other persons, the code _is_ very\nhelpful.\n\nThanks,\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48165","messageId":"85ir8ci7td.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"85abtpoydg.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-22T18:58:54Z","receivedAt":"2007-07-22T18:58:54Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Sun, 22 Jul 2007, David Kastrup wrote:\n>>\n>> So  when we are talking about a repository tree _becoming_ empty, we \n>> need the information whether or whether not we should remove it upon\n>> becoming empty.\n>\n> You don't seem to realize - although I've told you now abotu a million \n> times - that what you are talking about is:\n>\n>  - technically exactly the same as \".gitignore\", which for some \n>    unfathomable reason you cannot seem to accept.\n\nLinus?  Do both of us a favor and forget about the \".\" proposal.\nSince I already dropped it, we can save time if you rant about the\nproposal I have replaced it with and call me an idiot for a different\nreason.\n\n> Quite frankly, Junio would be a total idiot to take any patches that do \n> what you want to do. Happily, he is anything but.\n\nAnd he does not come across as one.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48217","messageId":"200707222226.30788.jnareb@gmail.com","threadId":"9086","inReplyTo":"857iosmto0.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-07-22T20:26:30Z","receivedAt":"2007-07-22T20:26:30Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"David Kastrup wrote:\n \n> \"I can follow you, but I disagree with your conclusion\" is perfectly\n> fine for now since I am going to propose something else, anyway.\n> \n> Thanks for the feedback.  It gave me some good ideas.\n\nYou are welcome.\n \n> Jakub Narebski <jnareb@gmail.com> writes: \n>> On Sun, 22 July 2007, David Kastrup wrote:\n>>> Jakub Narebski <jnareb@gmail.com> writes:\n>>>> David Kastrup wrote:\n>>>>\n>>>>> I must be really bad at explaining things, or I am losing a fight\n>>>>> against preconceptions fixed beyond my imagination.\n>>\n>> Or you are wrong...\n> \n> Well, there is little reason for you to take my word on it, but I\n> happen to have a history of designing and implementing systems where I\n> have been responsible for every single byte, bootloader, firmware,\n> applications, target compiler, assembler, whatever.  I have been\n> exposed to Unix and working with it several years before Linux even\n> existed.  I also have a track record of being not exactly stupid.\n> \n> So I pretty much can rule out that I am wrong on the factual side.\n\nBig words.\n\nFirst, there is little matter of something like area of competence.\nYou might be systems master, but your idea about snapshot based \ndistributed revision control systems can be wrong because DSCM are \noutside the area you know most about.\n\nSecond, even if you are a master at given topic, you can still be wrong.\n\nMind you, I was not saying you are wrong. I was saying you could be.\n\n\n[...] \n>> The only advantage to the \".\" idea is that it can use gitignore\n>> mechanism (both in-tree .gitignore, tracked or not, and info/exclude\n>> file). But I also think that the fact that gitignore mechanism is\n>> recursive is more of disadvantage than advantage.\n[...]\n> The recursiveness of the gitignore mechanism has the advantage that\n> when maintaining a large repository with actual or logical\n> subprojects, one does not need to pick a single policy for all\n> subprojects.  I think that is quite important.  It could possibly be\n> achieved with some other method of having per-subproject\n> configuration, but I see little wrong in using what is there and\n> documented already.\n\nI think it would be best implemented by repository config, e.g. \ncore.dirManagement or something like that, which could be set to\n 1. \"autoremove\" or something like that, which gives old behavior\n    of untracking directory if it doesn't have any tracked files\n    in it, and removing directory if it doesn't have any files\n    in it.\n 2. \"noremove\" or something like that, which changes the behaviour\n    to _never_ untrack directory automatically. This can be done\n    without any changes to 'tree' object nor index. It could be useful\n    for git-svn repositories.\n 3. \"marked\" or something like that, for which you have to explicitely\n    mark directories which are not to be removed when empty.\n 4. \"recursive\" or something like that, which would automatically mark\n    as \"sticky\" all subdirectories added in a \"sticky\" repository.\n    OR directory is not removed when empty if it is marked as such,\n    or one of its parents is marked as such.\n \n>> Second, the \"easy implementation\" is anything but easy. \"git add .\"\n>> as a way to mark directory as \"sticky\" is not backward compatibile:\n>> currently it mean to add _all contents_ of current directory.\n>> Implementation is tricky: as we have seen trying to unlink '.' or\n>> create '.' can unfortunately succeed on [some Sun OS, and UFS\n>> filesystem] (which follows POSIX stupidly to the letter) f**king up\n>> the filesystem.\n> \n> I was not suggesting actually leaving any such calls in place: after\n> all, they would presumably lead to error messages.  But I agree that\n> this could lead to nasty surprises when somebody with a legacy version\n> of git worked with a repository containing \".\" as explicit entries of\n> some file type.\n\nThe \"magic mode\" solution _should_ work also with older git, I think.\n \n\n>> Fourth, is very artificial. What would you put for filemode for '.'?\n>> 040000 (i.e. directory)?\n[...]\n>> What would you put for sha1?  Sha1 of an empty directory?\n> \n> Some fixed value.  Everywhere the same.  Not really relevant.\n\nRelevant because it has to work with legacy git on strange operating \nsystems. Because git has to fsck it (and adding special casing this \n\"some fixed value\" to git-fsck is bad, bad idea).\n\nNote that sha1 cannot be sha1 of the tree. In working area '.' is self \nlink. You cannot create self link in git repository object.\n\n[...]\n>>> And the repository is a versioned and hierarchically hashed version\n>>> of the index, but its trees contain _no_ information that is not\n>>> already inherently represented by the files alone. [...]\n[...]\n>> Trees do contain information which is not inherently present by the \n>> blobs.\n> \n> Could you give examples for such information?  As long as we are not\n> talking about _history_, I am at a loss at what else you mean.  File\n> names and permissions?\n\nFile names and permissions. And they bind blobs and trees together.\nTrees do not contain any info about history.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"48193","messageId":"85wswsf8o4.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"alpine.LFD.0.999.070718\u00041710271.27353@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-22T21:08:43Z","receivedAt":"2007-07-22T21:08:43Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"\nWell, coming back to this posting in order to focus on some points\nthat were at a level more relevant to the implementation.  And I'll go\nthrough the questions assuming my permissions-based proposal.\n\nLinus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Thu, 19 Jul 2007, David Kastrup wrote:\n>> \n>> Well, kudos.  Together with the analysis from Junio, this seems like a\n>> good start.  Would you have any recommendations about what stuff one\n>> should really read in order to get up to scratch about git internals?\n>\n> Well, you do need to understand the index. That's where all the new \n> subtlety happens.\n>\n> The data structures themselves are trivial, and we've supported\n> empty trees (at the top level) from the beginning, so that part is\n> not anything new.\n>\n> However, now having a new entry type in the index (S_IFDIR) means\n> that anything that interacts with the index needs to think\n> twice. But a lot of that is just testing what happens, and so the\n> first thing to do is to have a test-suite.\n\nYes.\n\n> There's also the question about how to show an empty tree in a\n> diff.\n\nWell, there are two possibilities involved here, a more and a less\nchatty one.  Assuming that we want to do as little work as possible,\nthe transition between a tracked and a non-tracked directory will be\ngiven in one of the following manners:\n\nEither:\na) xxx: old mode 000000\n   xxx: new mode 040755\n\nwhen a directory gets tracked and\n\n   xxx: new mode 040755\n   xxx: old mode 000000\n\nwhen it gets untracked again.\n\nor\nb)\n   xxx: new directory mode 040755\n\nwhen a directory gets tracked and\n\n   xxx: deleted directory mode 040755\n\nwhen it gets untracked again.  Note that \"new\" does not mean that git\ndid not previously have had files that absolutely have required a\ndirectory for placing.  It just means that it has now actively gained\nknowledge about the directory.\n\nIn a similar vein, \"deleted\" means that git is just deleting its\nknowledge about the directory, _scheduling_ it for a single deletion\nattempt at the earliest (and actually also latest) opportunity: when\ngit happens to know about no more files that require keeping the\ndirectory around.  So perhaps the following would be more readable:\n\n   xxx: tracking directory mode 040755\n\n   xxx: forgetting directory mode 040755\n\nNow in order to cut down on the verbiage, it might be an option to\ntransmit those strings only when something happens that can't be\ndeduced from other data.  Because _if_ it can be deduced from other\ndata (like a directory being present when files in it are), then at\nleast the working copies are identical as long as both persons don't\nstart deleting files from the repository.  If they do so, when a\ndirectory becomes empty, the other side needs to know whether the\ndirectory is being tracked or not if it still wants to maintain the\nsame state in the working tree.  But if we really want to have not\njust the working tree but also the repositories in SHA1-lockstep, we\ncan't delay transmitting this information.\n\n> We've never had that: the only time we had empty trees was when we\n> compared a totally empty \"root\" tree against another tree, and then\n> it was obvious.  But what if the empty tree is a subdirectory of\n> another tree - how do you express that in a diff? Do you care? Right\n> now, since we always recurse into the tree (and then not find\n> anything), empty trees will simply not show up _at_all_ in any\n> diffs.\n\nOne would still recurse.\n\n> And what about usability issues elsewhere? With my patch, doing something \n> like a\n>\n> \tgit add directory/\n>\n> still won't do anything, because the behaviour of \"git add\" has always \n> been to recurse into directories.\n\nThis will remain the same, but the directory itself will be added if\nand only if the corresponding preference variable is set, regardless\nof whether the directory is empty.\n\n> So to add a new empty directory, you'd have to do\n>\n> \tgit update-index --add directory\n>\n> and that's not exactly user-friendly.\n\nPresumably one could, if one really wanted an explicit way, have\ngit add --directory directory\nin analogy to the --directory option of the ls command.  But I think\nthat in most cases one would not want to treat one directory different\nfrom the whole tree, so the implicit behavior regulated by a\nproject-wide preference should be sufficient in general.\n\n> So do you add a \"-n\" flag to \"git add\" to tell it to not recurse? Or\n> do you always recurse, but then if you notice that the end result is\n> empty, you add it as a directory?\n\nI always recurse (unless there is a --directory option and I have some\nstrange desire to actually use it).  I add it as a directory,\nregardless of whether it is empty or not, if my preference setting (or\ngitignore or whatever) is set to tracking directories.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48196","messageId":"85sl7gf7g1.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"7vhco28aoq.fsf@assigned-by-dhcp.cox.net","subject":"Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-22T21:35:10Z","receivedAt":"2007-07-22T21:35:10Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"\nComing full circle...\n\nJunio C Hamano <gitster@pobox.com> writes:\n\n> The right approach to take probably would be to allow entries of\n> mode 040000 in the index.  Traditionally, we allowed only 100644\n> (blobs as regular files) and 120000 (blobs as symlinks).  We\n> recently added 160000 (commit from outer space, aka subproject).\n>\n> And we do that for all directories, not just empty ones.  So if\n> you have fileA, empty/, sub/fileB tracked, your index would\n> probably have these four entries, immediately after read-tree\n> of an existing tree object:\n>\n> \t100644 15db6f1f27ef7a... 0\tfileA\n> \t040000 4b825dc642cb6e... 0\tempty\n> \t040000 e125e11d3b63e3... 0\tsub\n> \t100644 52054201c2a872... 0\tsub/fileB\n\nThis would be very much what I am proposing now, except that instead\nof 040000 we would have 040755 usually, so that when the index makes\nit into the repository where 040000 already has a meaning (a\ndisappear-when-empty tree) we get the right information.  Also note\nthat the above comes about when doing\ngit-add *\nbut not when doing\ngit-add fileA empty sub/fileB (in the latter case, the entry for sub\n                               would be missing)\n\n> If you add sub/fileC, with \"update-index\" (and \"add\"), you\n> invalidate the SHA-1 object name you stored for \"sub\" (because\n> there is no point recomputing the tree object until you know you\n> need a subtree for \"sub\" part, which does not happen until the\n> next \"write-tree\"), and end up with something like:\n>\n> \t100644 15db6f1f27ef7a... 0\tfileA\n> \t040000 4b825dc642cb6e... 0\tempty\n> \t040000 00000000000000... 0\tsub\n> \t100644 52054201c2a872... 0\tsub/fileB\n> \t100644 705bf16c546f32... 0\tsub/fileC\n>\n> These \"missing\" SHA-1 would need to be recomputed on-demand.\n\nAh, ok.  Does it even make sense to compute the SHA-1 values in the\nindex in advance?  What would they be useful for?\n\n> We have had necessary infrastructure to do this \"keeping\n> untouched tree object names in the index\" for quite some time,\n> but it is not a part of the index proper (it is stored in an\n> extension section in the index file, to keep the index\n> compatible with older versions of git).\n\nWhat is the application for which this is being used?\n\n> Having made it sound so easy, here are the issues I would expect\n> to be nontrivial (but probably not rocket surgery either).\n>\n>  * unpack-trees, which is the workhorse for twoway merge (aka\n>    \"switching branches\") and threeway merge, has a convoluted\n>    logic to avoid D/F conflicts; it can probably be cleaned up\n>    once we do the above conversion so that the index starts\n>    saying \"Hey, I have a directory here\" more explicitly.  The\n>    end result would probably be a code easier to follow.\n\nI am afraid that this is unlikely to happen, and that is because\ndirectory tracking remains optional at a fundamental level as long as\nwe want to support the current behavior as an option.  However, one\ncould conceivably add 040000 entries (rather than 040755) for\ndirectories that have not been passed into tracking but are required\nby git, if this simplifies matters.  But it sounds like something that\nmight complicate working with several different git versions on the\nsame index.\n\n>  * status, update-index --refresh, and diff-files cares about\n>    the information cached in the index from the last time\n>    lstat(2) is run on each entry.  What we should store there\n>    for \"tree\" entries is very unclear to me, but probably we\n>    should teach them to ignore the stat-matching logic for\n>    these entries.\n\nAt the current point of time, git tracks just the u+x bit for normal\nfiles, and for directories, there is really nothing worth tracking as\nlong as no attempt of restoring more mode bits is done.  Modification\ntimes are probably a bit too risky to pay attention to.\n\n>  * diff-index walks the index and a tree in parallel but does\n>    not currently expect to see a tree object in the index.  It\n>    needs to be taught to ignore these \"tree\" entries.\n\nOr do something sensible when comparing.  Understood.\n\n>  * merge-recursive and merge-index walk the index, coming up\n>    with the merge results one path at a time.  They also need to\n>    be taught to ignore these \"tree\" entries.\n\nSame here.\n\n>  * diff-index and \"read-tree -m\" should be taught to take\n>    advantage of the \"tree\" entries in the index.  For example,\n>    if diff-index finds the \"tree\" entry in the index and the\n>    subtree found from the tree object exactly match, it does not\n>    even have to descend into the tree, which would be a huge\n>    performance win (because you do not have to open the subtree\n>    and its subtrees from the tree side; you already have read\n>    everything on the index side, and still have to skip the\n>    entries in the directory).  \"read-tree -m\" also should be\n>    able to optimize two identical subtrees in the 2 or 3 trees\n>    involved.\n>\n>    Even if we follow the \"lazy invalidate\" strategy to maintain\n>    the \"tree\" entries in the normal codepath, we could have a\n>    special operation that says \"now update all the tree entries\n>    by recomputing the tree object names as needed\".  Perhaps we\n>    might want to initiate such an operation before \"read-tree\n>    -m\" automatically.\n\nOver my head, but it would appear that it can safely left for later.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48232","messageId":"85644cf3mf.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"200707222226.30788.jnareb@gmail.com","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-22T22:57:44Z","receivedAt":"2007-07-22T22:57:44Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> writes:\n\n> David Kastrup wrote:\n>\n>> So I pretty much can rule out that I am wrong on the factual side.\n>\n> Big words.\n\nSure.  It is not relevant, however.\n\n> First, there is little matter of something like area of competence.\n> You might be systems master, but your idea about snapshot based\n> distributed revision control systems can be wrong because DSCM are\n> outside the area you know most about.\n\nSlicing the concept of directory and tree into two separate things and\nthinking separately about them and their relation in working tree and\nrepository is not exactly concerned with the internals.  It obviously\nwas too artificial a concept to be understandable, and likely a worse\nidea than necessary (whether one wants to call it too smart or too\nstupid for its own good may be a matter of taste).\n\nAnyway, it would be more productive if we managed to focus on the\ntechnical aspects again.  I accept that my previous proposal was not\nfit for inclusion.\n\n> Second, even if you are a master at given topic, you can still be\n> wrong.\n>\n> Mind you, I was not saying you are wrong. I was saying you could be.\n\nWe can leave that open since no code is going to come of the first\nproposal.\n\n> [...]\n>> The recursiveness of the gitignore mechanism has the advantage that\n>> when maintaining a large repository with actual or logical\n>> subprojects, one does not need to pick a single policy for all\n>> subprojects.\n>\n> I think it would be best implemented by repository config, e.g. \n> core.dirManagement or something like that, which could be set to\n>  1. \"autoremove\" or something like that, which gives old behavior\n>     of untracking directory if it doesn't have any tracked files\n>     in it, and removing directory if it doesn't have any files\n>     in it.\n\nThat's actually not _tracking_ a directory at all, but rather\nmaintaining an independent directory in the parallel repository\nuniverse.  No information specific to directories passes the index.\n\n>  2. \"noremove\" or something like that, which changes the behaviour\n>     to _never_ untrack directory automatically. This can be done\n>     without any changes to 'tree' object nor index. It could be useful\n>     for git-svn repositories.\n\nI don't see how this could occur.  Automatic _untracking_ would happen\nwhen one untracks (aka removes) a parent directory.  But one would not\ndo this while keeping the child.\n\n>  3. \"marked\" or something like that, for which you have to explicitely\n>     mark directories which are not to be removed when empty.\n\nEquivalent to 1 in my scheme.\n\n>  4. \"recursive\" or something like that, which would automatically mark\n>     as \"sticky\" all subdirectories added in a \"sticky\" repository.\n\nIf they are covered by the add and not just implied by childs.  That is,\ngit-add a/b\nwill not make \"a\" sticky while\ngit-add a\nwill make a/b sticky.\n\n>     OR directory is not removed when empty if it is marked as such,\n>     or one of its parents is marked as such.\n\nI'd not throw too much inheritance into the equation, or things become\nintractable too easily.\n\n> The \"magic mode\" solution _should_ work also with older git, I\n> think.\n\nI think so, too, for the repository.  But of course what happens in\nthe index with old code when new data types get added is a case for\nreview, testing and praying.\n\n>>> Fourth, is very artificial. What would you put for filemode for '.'?\n>>> 040000 (i.e. directory)?\n> [...]\n>>> What would you put for sha1?  Sha1 of an empty directory?\n>> \n>> Some fixed value.  Everywhere the same.  Not really relevant.\n>\n> Relevant because it has to work with legacy git on strange operating \n> systems. Because git has to fsck it (and adding special casing this \n> \"some fixed value\" to git-fsck is bad, bad idea).\n\nI did not mean \"arbitrary value\", but the value would be computed in a\nstandard way from the node, and since the node would be the same\neverywhere, the hash would be too.\n\n> Note that sha1 cannot be sha1 of the tree. In working area '.' is\n> self link. You cannot create self link in git repository object.\n\nCertainly.  And the idea was to have \".\" be isolated from the contents\nof the tree, basically treating it as a sibling of the other entries.\nWhich is, in a way, how \".\" shared one namespace in Unix with what\namounts to _children_ of the corresponding tree.\n\nSo that was some inspiration here, probably too much so.\n\n> [...]\n>>>> And the repository is a versioned and hierarchically hashed version\n>>>> of the index, but its trees contain _no_ information that is not\n>>>> already inherently represented by the files alone. [...]\n> [...]\n>>> Trees do contain information which is not inherently present by the \n>>> blobs.\n>> \n>> Could you give examples for such information?  As long as we are not\n>> talking about _history_, I am at a loss at what else you mean.  File\n>> names and permissions?\n>\n> File names and permissions. And they bind blobs and trees together.\n\nTrees bind blobs and trees together?  Anyway, I consider the names and\npermissions properties of the files and their identity.  Stripping out\nthe blobs from under them does not actually add any information: the\ntrees still don't contain any information that would have necessitated\nlooking at directories rather than just files, their names,\npermissions and content in the work space.\n\nBut you are right in that the tree can't be replaced by the blobs.  It\nactually needs the files (namely their full names and permissions) to\nreconstruct it.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48266","messageId":"85ps2jeju9.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"85644cf3mf.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-23T06:05:02Z","receivedAt":"2007-07-23T06:05:02Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"David Kastrup <dak@gnu.org> writes:\n\n> Jakub Narebski <jnareb@gmail.com> writes:\n>\n>> I think it would be best implemented by repository config, e.g.\n\nI got sidetracked here: the gitignore stuff has in the dirmod scheme\nactually no code or concept overlap with the actual scheme, so it can\nbe considered a distraction for now and its implementation and\ndiscussion tabled.  It has one disadvantage: in order to get\n_recursive_ behavior in one tree, one needs to use the \".\" pattern in\nthe .gitignore file of the respective directory, and having a\n.gitignore file in that directory sort of defeats the idea of not\nhaving .gitignore directories around...  Of course, a single\n.gitignore file is better than ones one has to distribute through the\ntree.\n\n>> core.dirManagement or something like that, which could be set to\n>>  1. \"autoremove\" or something like that, which gives old behavior\n>>     of untracking directory if it doesn't have any tracked files\n>>     in it, and removing directory if it doesn't have any files\n>>     in it.\n>\n> That's actually not _tracking_ a directory at all, but rather\n> maintaining an independent directory in the parallel repository\n> universe.  No information specific to directories passes the index.\n\nNote: that was merely a comment on semantics, not on the matter.\n\n>>  2. \"noremove\" or something like that, which changes the behaviour\n>>     to _never_ untrack directory automatically. This can be done\n>>     without any changes to 'tree' object nor index. It could be useful\n>>     for git-svn repositories.\n>\n> I don't see how this could occur.  Automatic _untracking_ would happen\n> when one untracks (aka removes) a parent directory.  But one would not\n> do this while keeping the child.\n\nCorrection: if there was a --directory option and one used it for\ngit-rm (or no -r was given, so just one directory level was effected),\none _could_ untrack stuff on the git side accidentally.  And for\nsomething like git-svn, this might be a bad idea.  So there is\nconceivably a market for an option that never untracks a non-empty\ntree.\n\n>>  3. \"marked\" or something like that, for which you have to explicitely\n>>     mark directories which are not to be removed when empty.\n>\n> Equivalent to 1 in my scheme.\n\nAt least if scheme 1 does not forbid some _explicit_ way of saying\n\"track this and I really mean it\".\n\n>>  4. \"recursive\" or something like that, which would automatically mark\n>>     as \"sticky\" all subdirectories added in a \"sticky\" repository.\n>\n> If they are covered by the add and not just implied by childs.  That is,\n> git-add a/b\n> will not make \"a\" sticky while\n> git-add a\n> will make a/b sticky.\n\nAddition: I was thinking so much of my implementation and its\nsemantics that I did not consider one possibility that you might mean\nhere:\n\nWhen adding a/b, always also add a (and the whole hierarchy above it)\nautomatically as sticky.  Namely disallow unsticky directories in the\nrepository at all.  That would mean that\n\n  git-add a/b;git-commit -m x;git-rm a/b;git-commit -m x\n\nmight not be a noop if a was not in the repository previously: it\nwould cause a to stay around sticky until removed.  With all other\nschemes, however, it would cause a to be removed \"on behalf of the\nuser\" even if the user intended it to stay around.\n\nIndeed, this scheme might by far be the easiest to understand.  Having\nno autoremoval at all in levels higher than the deleted level is\nsomething that people might easily understand: delayed removal just\ndoes not happen anymore, and git never deletes a directory unless told\nto.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48276","messageId":"86bqe3va08.fsf@lola.quinscape.zz","threadId":"9086","inReplyTo":"85ps2jeju9.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-23T07:45:27Z","receivedAt":"2007-07-23T07:45:27Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"David Kastrup <dak@gnu.org> writes:\n\n> Addition: I was thinking so much of my implementation and its\n> semantics that I did not consider one possibility that you might mean\n> here:\n>\n> When adding a/b, always also add a (and the whole hierarchy above it)\n> automatically as sticky.  Namely disallow unsticky directories in the\n> repository at all.  That would mean that\n>\n>   git-add a/b;git-commit -m x;git-rm a/b;git-commit -m x\n>\n> might not be a noop if a was not in the repository previously: it\n> would cause a to stay around sticky until removed.  With all other\n> schemes, however, it would cause a to be removed \"on behalf of the\n> user\" even if the user intended it to stay around.\n>\n> Indeed, this scheme might by far be the easiest to understand.\n> Having no autoremoval at all in levels higher than the deleted level\n> is something that people might easily understand: delayed removal\n> just does not happen anymore, and git never deletes a directory\n> unless told to.\n\nAnd of course, it would be a nuisance for people managing a\npatch-based workflow.  But those can actually easily set the\nrepository preferences differently, and even\nfind -type d -empty -delete\nis not too hard to do.  So it would even be feasible as default.\n\nBut I think that in practice, the \"track only what has been added\nrecursively\" approach is a good default.  And since patches without\ndir information never add anything recursively, it would mostly keep\nthe directories clean.\n\n-- \nDavid Kastrup\n"},{"id":"48342","messageId":"87ps2inab5.fsf@hades.wkstn.nix","threadId":"9086","inReplyTo":"863azka7d4.fsf@lola.quinscape.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Nix","fromEmail":"nix@esperi.org.uk","sentAt":"2007-07-23T20:18:22Z","receivedAt":"2007-07-23T20:18:22Z","isPatch":true,"sender":{"key":"nix@esperi.org.uk","avatar":"https://avatars.githubusercontent.com/u/6503005?v=4"},"body":"On 19 Jul 2007, David Kastrup stated:\n> Tomash Brechko <tomash.brechko@gmail.com> writes:\n>> Please consider this: I myself use Git to track my own local\n>> projects, and for this usage you proposal have no value for me,\n>> i.e. as a _Source_ Code Management system Git is rather complete.\n>> But I also track /etc and ~/ in Git, and for this I'd love to have\n>> directories, permissions, ownership, other attributes, to be\n>> tracked.  I have Perl script wrapping Git that allows me to filter\n>> tracked paths by full regexps instead of Git's file globs, and also\n>> to filter out too big files assuming that they are binary anyway.\n>\n> Look, git _tracks_ contents.  Your permissions managements needs to be\n> told explicitly when and how things change.  So you end up with git\n> _tracking_ material and your permissions/directory management needing\n> the level of manual handholding Subversion demands.\n\nActually, if we had a post-checkout hook, we could use a pre-commit hook\nto keep track of directory existence, permissions, et seq, and a post-\ncheckout hook to restore them.\n\n(But we don't, at least not yet. Adding one is probably quite easy.)\n"},{"id":"48354","messageId":"85y7h6dewp.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"87ps2inab5.fsf@hades.wkstn.nix","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-23T20:49:10Z","receivedAt":"2007-07-23T20:49:10Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Nix <nix@esperi.org.uk> writes:\n\n> On 19 Jul 2007, David Kastrup stated:\n>> Tomash Brechko <tomash.brechko@gmail.com> writes:\n>>> Please consider this: I myself use Git to track my own local\n>>> projects, and for this usage you proposal have no value for me,\n>>> i.e. as a _Source_ Code Management system Git is rather complete.\n>>> But I also track /etc and ~/ in Git, and for this I'd love to have\n>>> directories, permissions, ownership, other attributes, to be\n>>> tracked.  I have Perl script wrapping Git that allows me to filter\n>>> tracked paths by full regexps instead of Git's file globs, and also\n>>> to filter out too big files assuming that they are binary anyway.\n>>\n>> Look, git _tracks_ contents.  Your permissions managements needs to\n>> be told explicitly when and how things change.  So you end up with\n>> git _tracking_ material and your permissions/directory management\n>> needing the level of manual handholding Subversion demands.\n>\n> Actually, if we had a post-checkout hook, we could use a pre-commit\n> hook to keep track of directory existence, permissions, et seq, and\n> a post- checkout hook to restore them.\n\nActually, tracking permissions would be cheap: one just needs to\nreplace the permission-munging macros in git with identity.  Ownership\n-- well, that's harder.\n\nBut my sentiment remains: git _tracks_ stuff: it notices when things\nmove around and follows them.  Statically snapshotting permissions\ncreates a layer that is quite less flexible.  The information gets\ndetached.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48361","messageId":"87lkd6n62i.fsf@hades.wkstn.nix","threadId":"9086","inReplyTo":"85y7h6dewp.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Nix","fromEmail":"nix@esperi.org.uk","sentAt":"2007-07-23T21:49:57Z","receivedAt":"2007-07-23T21:49:57Z","isPatch":true,"sender":{"key":"nix@esperi.org.uk","avatar":"https://avatars.githubusercontent.com/u/6503005?v=4"},"body":"On 23 Jul 2007, David Kastrup uttered the following:\n> Nix <nix@esperi.org.uk> writes:\n>> Actually, if we had a post-checkout hook, we could use a pre-commit\n>> hook to keep track of directory existence, permissions, et seq, and\n>> a post- checkout hook to restore them.\n>\n> Actually, tracking permissions would be cheap: one just needs to\n> replace the permission-munging macros in git with identity.  Ownership\n> -- well, that's harder.\n>\n> But my sentiment remains: git _tracks_ stuff: it notices when things\n> move around and follows them.  Statically snapshotting permissions\n> creates a layer that is quite less flexible.  The information gets\n> detached.\n\nNot if you record it in a file which is checked in in the same commit\nthat is tracked, it isn't (that's what the pre-commit hook is for). It's\ntrue that git won't natively have any knowledge of that data, but Linus\nhas fairly effectively shown that it shouldn't have any such knowledge\nand doesn't need it.\n\n(You might want to give git-diff knowledge of it, just so it can skip\nit unless a new flag is given. Give the file a nice format, and bingo,\nreadable permission/ownership diffs!)\n\n(I'd recommend storing the names of user/group file owners as well as\nthe uids, so you can --- given suitable permissions --- chown to the\nright username in preference to uid if that user exists at checkout\ntime.)\n\n\nDoing this *efficiently* is another matter: probably a pair of hooks are\nneeded, run on pre-checkout and post-checkout: they can communicate so\nas only to fiddle permissions on things which are newly appeared or\nwhose permissions have changed.\n\nObviously because the permissions, ownerships et al aren't recorded in\nthe index this will slow committing down, but given that\ngit-update-index will already have sucked the entire tree's inodes into\nthe page cache anyway, I don't think a second pass over the working tree\nsnarfing permissions would slow it down much.\n\n\nAs I need this anyway (I'm backing up a filesystem via git, yes, I'm\ninsane but I need version control and it's horrifically redundant so\npacking it will save heaps of space), I guess I'd better get off my\nrear and write the code.\n\n(The recent commit-as-a-builtin's introduction of a run_hook() function\nwill be pretty damn useful: good timing, I guess.)\n"},{"id":"48362","messageId":"87hcnun5dc.fsf@hades.wkstn.nix","threadId":"9086","inReplyTo":"87lkd6n62i.fsf@hades.wkstn.nix","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Nix","fromEmail":"nix@esperi.org.uk","sentAt":"2007-07-23T22:05:03Z","receivedAt":"2007-07-23T22:05:03Z","isPatch":true,"sender":{"key":"nix@esperi.org.uk","avatar":"https://avatars.githubusercontent.com/u/6503005?v=4"},"body":"On 23 Jul 2007, nix@esperi.org.uk outgrape:\n> (I'd recommend storing the names of user/group file owners as well as\n> the uids, so you can --- given suitable permissions --- chown to the\n> right username in preference to uid if that user exists at checkout\n> time.)\n\nSuddenly this gets more complex. git-merge-file(1) has to understand the\ncontents of this file, so as not to consider merges conflicting unless\ntwo files actually have different permissions (i.e. doing a line by line\ndiff, and combining the two such that at most one file with a given name\nexists in the result), and so as not to consider lines with differing\nownerships conflicting unless we're running under a uid in which we can\nchange ownerships at all. (I'd like to track ownership but it's looking\nlike a bit of a nest of snakes.)\n\nAnd the problem is that while git has a lot of strategies for merging\n*trees*, its file merge system is totally unpluggable: it just falls\nback to xdiff's merging system. I guess I'll have to add that feature :)\n\n(How does this cope with binary files, I wonder? I seem to recall\nsomething about that flying past back before the volume of the git list\noverwhelmed me...)\n"},{"id":"48364","messageId":"85k5sqdavo.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"87lkd6n62i.fsf@hades.wkstn.nix","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-23T22:16:11Z","receivedAt":"2007-07-23T22:16:11Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Nix <nix@esperi.org.uk> writes:\n\n> On 23 Jul 2007, David Kastrup uttered the following:\n>> Nix <nix@esperi.org.uk> writes:\n>>> Actually, if we had a post-checkout hook, we could use a pre-commit\n>>> hook to keep track of directory existence, permissions, et seq, and\n>>> a post- checkout hook to restore them.\n>>\n>> Actually, tracking permissions would be cheap: one just needs to\n>> replace the permission-munging macros in git with identity.  Ownership\n>> -- well, that's harder.\n>>\n>> But my sentiment remains: git _tracks_ stuff: it notices when things\n>> move around and follows them.  Statically snapshotting permissions\n>> creates a layer that is quite less flexible.  The information gets\n>> detached.\n>\n> Not if you record it in a file which is checked in in the same\n> commit that is tracked, it isn't (that's what the pre-commit hook is\n> for).\n\nI have my doubts that anybody but git actually has a clue what to\nsnapshot when, and where to place it: don't forget that index\nmanipulation and committing are done at different times, and you need\nnot even commit all of the index.\n\n> It's true that git won't natively have any knowledge of that data,\n> but Linus has fairly effectively shown that it shouldn't have any\n> such knowledge and doesn't need it.\n\nLast time I looked, git tracked the executable bit.  For kernel\ndevelopment, this is pretty much what it takes, and with colloborative\nwork, tracking anything but the owner permissions is going to lead to\nannoying and verbose merge behavior quite a lot.  And of the owner\npermissions, r and w complicate proper handling when unset.\n\nBut being able to specify other masks for applications other than\nmulti-site colloborative development would likely not hurt.\n\n> Doing this *efficiently* is another matter: probably a pair of hooks\n> are needed, run on pre-checkout and post-checkout: they can\n> communicate so as only to fiddle permissions on things which are\n> newly appeared or whose permissions have changed.\n>\n> Obviously because the permissions, ownerships et al aren't recorded\n> in the index this will slow committing down,\n\nIt will also detach the time where the file contents and the\npermissions get recorded.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48366","messageId":"alpine.LFD.0.999.0707231527050.3607@woody.linux-foundation.org","threadId":"9086","inReplyTo":"85k5sqdavo.fsf@lola.goethe.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-23T22:31:46Z","receivedAt":"2007-07-23T22:31:46Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 24 Jul 2007, David Kastrup wrote:\n\n> Nix <nix@esperi.org.uk> writes:\n> \n> > It's true that git won't natively have any knowledge of that data,\n> > but Linus has fairly effectively shown that it shouldn't have any\n> > such knowledge and doesn't need it.\n> \n> Last time I looked, git tracked the executable bit.\n\nActually, originally it tracked the whole mode word.\n\nIt was a total disaster. People who had different umasks etc got mode \nclashes all the time, and you ended up having silly and unnecessary \nconflicts.\n\nThe same would be true (to an even higher degree) if we tracked owner and \ngroup information etc.\n\nSo practically speaking, you want to track the *minimal* possible state, \nnot the maximal one. \n\nThis is one of those \"in theory\" vs \"in practice\" things. In *theory*, it \nwould be nice for an SCM to track everything that is known about a file. \nIn *practice*, that sucks.\n\nSo this does mean that if you want to explicitly track certain things \n(ownership and more complete file permissions, or ACL's, or \"resource \nforks\", or any number of other things that a file *could* have on various \nsystems), you end up havign to track them in something else than git, or \nyou end up having to track them as a separate \"metadata file\".\n\nOne such metadata file is, for example, the \".gitattributes\" file. It \n*could* be used to contain things like path-based rules for ownership, \nnot just things like whether to check out with CRLF etc.\n\n\t\t\tLinus\n"},{"id":"48369","messageId":"f83bfv$95g$1@sea.gmane.org","threadId":"9086","inReplyTo":"87hcnun5dc.fsf@hades.wkstn.nix","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-07-23T22:52:47Z","receivedAt":"2007-07-23T22:52:47Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Nix wrote:\n\n> And the problem is that while git has a lot of strategies for merging\n> *trees*, its file merge system is totally unpluggable: it just falls\n> back to xdiff's merging system. I guess I'll have to add that feature :)\n\nNot true. You can add custom diff driver for files using gitattributes\nsystem.\n \n> (How does this cope with binary files, I wonder? I seem to recall\n> something about that flying past back before the volume of the git list\n> overwhelmed me...)\n\nxdiff has binary diff, and git has some kind of \"ascii-armored\" binary diff\noutput. As to how to merge binary files: I suspect that they always\nconflict, unless the merge is trivial.\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"48375","messageId":"873azen1c7.fsf@hades.wkstn.nix","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707231527050.3607@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Nix","fromEmail":"nix@esperi.org.uk","sentAt":"2007-07-23T23:32:08Z","receivedAt":"2007-07-23T23:32:08Z","isPatch":true,"sender":{"key":"nix@esperi.org.uk","avatar":"https://avatars.githubusercontent.com/u/6503005?v=4"},"body":"On 23 Jul 2007, Linus Torvalds spake thusly:\n> So practically speaking, you want to track the *minimal* possible state, \n> not the maximal one. \n\nI think it depends on your use case. For source code and indeed anything\nwith heavy merges, this is true: but I'm increasingly using git as a\nsort of `merged historical tar' to store images of entire random\nfilesystem trees across time, and gaining the benefit of the packer's\nlovely space-efficiency as well (doing this with svn would be a lost\ncause, twice the space usage before you even think about the\nrepository). And in that case, preserving everything you can makes\nsense.\n\n(Perhaps what I should be doing is tarring the directory tree up and\nstoring the *tarball* in git. I'll try that and see what it does to pack\nsizes. These are version-controlled backups of my mother's magnum opus\nin progress so you can understand that I don't want to destroy them\naccidentally: I'd never hear the end of it! ;) )\n\n> So this does mean that if you want to explicitly track certain things \n> (ownership and more complete file permissions, or ACL's, or \"resource \n> forks\", or any number of other things that a file *could* have on various \n> systems), you end up havign to track them in something else than git, or \n> you end up having to track them as a separate \"metadata file\".\n\nYes indeed: that's why I proposed doing this using a couple of new hooks\ndriving entirely optional permissions-preservation stuff. Most use cases\nreally won't want to track this, so this sort of stuff shouldn't impose\nupon the git core or upon anyone who doesn't want it. (However, the\nability to have alternative file merging strategies *may* be useful\nelsewhere, perhaps.)\n"},{"id":"48381","messageId":"alpine.LFD.0.999.0707231638020.3607@woody.linux-foundation.org","threadId":"9086","inReplyTo":"873azen1c7.fsf@hades.wkstn.nix","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-23T23:57:44Z","receivedAt":"2007-07-23T23:57:44Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 24 Jul 2007, Nix wrote:\n>\n> On 23 Jul 2007, Linus Torvalds spake thusly:\n> > So practically speaking, you want to track the *minimal* possible state, \n> > not the maximal one. \n> \n> I think it depends on your use case. For source code and indeed anything\n> with heavy merges, this is true\n\nYes, very obviously. Git is targeted towards source code and working in a \ndistributed manner across a very wide variety of users and setups, while \nsomething that would be more targeted towards a special scenario and much \nstricter usage would find that the \"minimum\" set is much bigger, and might \nwell include ACL's and usr information.\n\n> but I'm increasingly using git as a sort of `merged historical tar' to \n> store images of entire random filesystem trees across time, and gaining \n> the benefit of the packer's lovely space-efficiency as well (doing this \n> with svn would be a lost cause, twice the space usage before you even \n> think about the repository). And in that case, preserving everything you \n> can makes sense.\n\nOn the other hand, almost all the space-efficiency comes from things that \ndelta well, and change quickly. That includes the file data itself (and \nvery much the tree contents), but it doesn't necessarily include things \nlike permissions and user information - mainly because that doesn't \nactually delta at all (not because it can't, but because it hardly ever \nchanges, and when it does change, it often changes all over the map).\n\nTo make an example of your \"tar\" situation: if you want to be space- \nefficient in a tar-like setting, you should *not* make user information be \nsomething that is per-file at all! Why? Because in 99% of all tar-files, \nthere is a single user name.\n\nSo even your usage *may* actually be much better off using git as a \"data \nbackend\", and using something totally different for \"user/group\" \ninformation. Yes, you'd have to make a \"shim layer\" on top of git to hide \nthe fact that the user information is handled separately, but that \nshouldn't be that hard per se.\n\n> (Perhaps what I should be doing is tarring the directory tree up and\n> storing the *tarball* in git. I'll try that and see what it does to pack\n> sizes. These are version-controlled backups of my mother's magnum opus\n> in progress so you can understand that I don't want to destroy them\n> accidentally: I'd never hear the end of it! ;) )\n\nYou don't want to do this. \n\nThere's a few reasons, but the two big ones are:\n\n - the git delta logic is strictly a \"single delta base\" thing.\n\n   Yes, git would be able to find the delta's between two tar-files (as \n   long as you don't compress them), and express one tar-file in terms of \n   the other, and it would probably save a fair amount of disk.\n\n   But it would not be able to do _nearly_ as well as it can if you store \n   individual files, and let git just find the best delta per-file (and \n   not just \"one delta base for the whole tar-ball\")\n\n - git is very much optimized for \"many small files\". Yes, you can check \n   in large files, and it works fine, but quite frankly, all the design \n   and heavy optimizations have been about having trees with tens of \n   thousands of files, but the files individually reasonably small.\n\n   A lot of the speed advantages of git come from efficiently pruning away \n   whole sub-directory structures, for example, and not even touching the \n   data at all!\n\n   So if you track just one file that changes in every version, all the \n   things that make git fly are basically disabled, and you won't take \n   full advantage of what git does.\n\n> Yes indeed: that's why I proposed doing this using a couple of new hooks\n> driving entirely optional permissions-preservation stuff. Most use cases\n> really won't want to track this, so this sort of stuff shouldn't impose\n> upon the git core or upon anyone who doesn't want it. (However, the\n> ability to have alternative file merging strategies *may* be useful\n> elsewhere, perhaps.)\n\nThe \".gitattributes\" file really could be used for some of that. Using it \nto track ownership and full permissions would not be impossible, and it \ncould have interesting semantics (especially as .gitattibutes is path \npattern based - so you could literally do a \"user\" attribute, and say that \neverything in a particular subdirectory is owned by a particular user).\n\nThat wouldn't be UNIX-like semantics, of course, but it can be very useful \nfor certain things. \n\nTaking an example of something totally independent of git, look at how \n\"udev\" handles permissions, for example. In situations like that, static \nuser information is useless, and it actually ends up setting up modes and \nownership based on name-based patterns rather than having each file have a \npermission/user (because individual files appear and disappear, the \nname-based patterns are the things that matter).\n\nSo if you *just* want to track a regular filesystem layout, that's not the \nright thing, but \"udev\" does show an example of a totally different way of \ndescribing ownership and permissions, and one which wouldn't actually be \nat all foreign to git.\n\n\t\tLinus\n"},{"id":"48399","messageId":"87y7h6l27z.fsf@hades.wkstn.nix","threadId":"9086","inReplyTo":"86ps2ithyl.fsf@lola.quinscape.zz","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Nix","fromEmail":"nix@esperi.org.uk","sentAt":"2007-07-24T06:56:00Z","receivedAt":"2007-07-24T06:56:00Z","isPatch":true,"sender":{"key":"nix@esperi.org.uk","avatar":"https://avatars.githubusercontent.com/u/6503005?v=4"},"body":"On 24 Jul 2007, David Kastrup spake thusly:\n> But merging will become nicer if the permissions actually stay\n> associated with the file rather than the file name.  Even in things\n> like /etc backups, blobs not infrequently relocate from one place to\n> another when the system gets updated.\n\nEven without that we'd need to merge without context, i.e. with totally\nindependent lines, for such a file. So it's not the standard git file\nmerge.\n"},{"id":"48601","messageId":"877iooxfxb.fsf@hades.wkstn.nix","threadId":"9086","inReplyTo":"f83bfv$95g$1@sea.gmane.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"Nix","fromEmail":"nix@esperi.org.uk","sentAt":"2007-07-25T22:43:44Z","receivedAt":"2007-07-25T22:43:44Z","isPatch":true,"sender":{"key":"nix@esperi.org.uk","avatar":"https://avatars.githubusercontent.com/u/6503005?v=4"},"body":"On 23 Jul 2007, Jakub Narebski spake thusly:\n\n> Nix wrote:\n>\n>> And the problem is that while git has a lot of strategies for merging\n>> *trees*, its file merge system is totally unpluggable: it just falls\n>> back to xdiff's merging system. I guess I'll have to add that feature :)\n>\n> Not true. You can add custom diff driver for files using gitattributes\n> system.\n\nOo. Excellent, I didn't notice that. Thank you.\n"},{"id":"48747","messageId":"200707270133.25221.robin.rosenberg.lists@dewire.com","threadId":"9086","inReplyTo":"85lkdezi08.fsf@lola.goethe.zz","subject":"Re: Empty directories...","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg.lists@dewire.com","sentAt":"2007-07-26T23:33:24Z","receivedAt":"2007-07-26T23:33:24Z","isPatch":false,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"\n(\n\tI don't know which mail is the best to reply to and I probably missed \n\tsomething in the thread, so bear with me if I'm repeating anything.\n)\n\nDavid. Reconsider \"tracking\" all directories and what that would give, \ncompared to explicitly tracking specific ones and the requires magic entries.\n\nSay we have a config setting that tells git never to remove empty trees. Linus \npatches could be a start for representing trees in the index. As an \noptimization the index could prune trees from the index if they contain \nthings as long as the index *effectively* remembers all trees.\n\nUsing the patches again we could add empty directories to the index and remove \nthem. No directory would be removed automatically, except maybe by a merge.\n\nWe would probably have only a few empty directories and new unexpected ones\nwould only pop up when we remove all blobs from one. Git status could tell us\nabout them so we will not forget them. It could even tell us about \"new\" empty\ndirectories, which is probably the most important thing you'd want to know. \n\nForgetting to untrack an empty directory would not be a big deal.\n\nWhether to retain empty trees or not should be a repository policy, but an all \nor nothing setting.\n\n-- robin\n"},{"id":"48761","messageId":"854pjq775r.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"200707270133.25221.robin.rosenberg.lists@dewire.com","subject":"Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-27T05:22:08Z","receivedAt":"2007-07-27T05:22:08Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Robin Rosenberg <robin.rosenberg.lists@dewire.com> writes:\n\n> (\n> \tI don't know which mail is the best to reply to and I probably missed \n> \tsomething in the thread, so bear with me if I'm repeating anything.\n> )\n>\n> David. Reconsider \"tracking\" all directories and what that would\n> give, compared to explicitly tracking specific ones and the requires\n> magic entries.\n\nIt would be quite a nuisance for a patch-based workflow, since patches\ndon't talk about the creation and deletion of directories.\n\nThe \"track only when entered approach\" has the advantage that\ndirectories that were only created to accommodate patches will be\nremoved again when becoming empty.\n\nOf course, once doing \"git-add top-level\" will level the difference.\n\n> Say we have a config setting that tells git never to remove empty\n> trees.\n\nWhy wouldn't I have tree/zap removed when doing git-rm tree?\n\n> Linus patches could be a start for representing trees in the\n> index. As an optimization the index could prune trees from the index\n> if they contain things as long as the index *effectively* remembers\n> all trees.\n\nBut it doesn't.  If you do git-add tree, optimizing the dir entry away\nsince tree/zap exists, then subsequently do git-rm tree/zap, of course\nthere is nothing to do except remove tree/zap, and the tree is gone.\n\nOne can't start tracking trees explicitly only when they become empty,\nbecause one can't know whether to track them then.\n\n> Using the patches again we could add empty directories to the index\n> and remove them. No directory would be removed automatically, except\n> maybe by a merge.\n\nI currently have the problem that\n\nrm -rf *\nunzip some-archive\ngit-add some-archive\ngit-commit -a -m whatever\ngit-checkout something else\n\nleaves empty directory skeletons lying around.\n\n> We would probably have only a few empty directories and new\n> unexpected ones would only pop up when we remove all blobs from\n> one. Git status could tell us about them so we will not forget\n> them.\n\nI don't want a source management system to tell me whenever it is\ngoing to annoy me.\n\n> It could even tell us about \"new\" empty directories, which is\n> probably the most important thing you'd want to know.\n>\n> Forgetting to untrack an empty directory would not be a big deal.\n>\n> Whether to retain empty trees or not should be a repository policy,\n> but an all or nothing setting.\n\nWith that approach idea the workflow\n\n\"Apply a patch creating something/hello\"\n\"Undo the patch creating something/hello\"\n\nwill leave something lying around.  For somebody managing hundreds of\ndirectories, that would be a nuisance.\n\nI don't say that a \"track all parents automatically\" approach would\nnot have its merits: it would likely prevent some mistakes and be\neasily understandable to most users.  But for managing a patch\nworkflow, it would appear to get in the way.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48880","messageId":"854pjo533r.fsf@lola.goethe.zz","threadId":"9086","inReplyTo":"alpine.LFD.0.999.0707202154220.27249@woody.linux-foundation.org","subject":"Re: [RFC PATCH] Re: Empty directories...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-28T08:44:56Z","receivedAt":"2007-07-28T08:44:56Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> Why? Because my preliminary patches sort the index entries wrong. A \n> directory should always sort *as*if* it had the '/' at the end.\n>\n> See base_name_compare() for details.\n>\n> And we've never done that for the index, because the index has never had \n> this issue (since it never contained directories). So sit down and compare \n> base_name_compare (for tree entries) with cache_name_compare() (for index \n> entries), and see how the latter doesn't care about the type of names.\n>\n> This was actually something that I hit already with subproject support, \n> and one of my very first patches even had some (aborted) code to start \n> sorting subprojects in the index the way we sort directories.\n>\n> And I *should* have done it that way, but I never did. It now makes the \n> S_ISDIR handling harder, because directories really do have to be sorted \n> as if they had the '/' at the end, or \"git-fsck\" will complain about bad \n> sorting.\n>\n> Sad, sad, sad. It effectively means that S_IFGITLINK is *not* quite the \n> same as S_IFDIR, because they sort differently. Duh.\n>\n> Of course, it seldom matters, but basically, you should test a directory \n> structure that has the files\n>\n> \tdir.c\n> \tdir/test\n>\n> in it, and the \"dir\" directory should always sort _after_ \"dir.c\".\n>\n> And yes, having the index entry with a '/' at the end would handle that \n> automatically.\n\nPersonally, I am not much in favor of using different names in index\nand repository.\n\n> As it is, with the \"mode\" difference, it instead needs to fix up \n> \"cache_name_compare()\". Admittedly, that would actually be a cleanup \n> (since it would now match base_name_compare() in logic, and could actually \n> use that to do the name comparison!), but it's a damn painful cleanup \n> because we don't even pass in the mode to \"cache_name_compare()\", since we \n> never needed it.\n>\n> Gaah.\n>\n> cache_name_compare itself isn't used in that many places,\n\ndir.c and readcache.c\n\n> but it's used by \"index_name_pos()/cache_name_pos()\", which *is*\n> used in many places.\n\ncache_name_pos:\nbuiltin-apply.c\nbuiltin-blame.c\nbuiltin-checkout-index.c\nbuiltin-ls-files.c\nbuiltin-mv.c\nbuiltin-read-tree.c\nbuiltin-rm.c\nbuiltin-update-index.c\ndiff.c\ndiff-lib.c\ndir.c\nmerge-index.c\nsha1_name.c\nunpack-trees.c\nwt-status.c\n\nindex_name_pos:\nread-cache.c\n\n> And again, that one doesn't even have the mode, so it cannot pass it\n> down.\n\n> So it probably *is* easier to add the '/' at the end of the name\n> instead, to make directories sort the right way in the index. I'd\n> still suggest you *also* make the mode be S_IFDIR, though (and\n> preferably make git-fsck actually verify that the mode and the last\n> character of the name matches!).\n\nActually, pretty much all of the above files are likely to get touched\nby directory support one way or another anyway.  One really should aim\nfor the cleanest solution in the long run, and this for me more or\nless means that it makes no sense to have different names in index and\nrepository.  Putting that slash in always would probably simplify some\nlogic in the repository as well, but I don't really like something as\nmarker-like as \"/\" in the data structures.  Putting a slash there\nwould involve a three-phase plan:\n\na) make fsck and the other code deal gracefully with either slash or\n   no slash.\nWait until everybody uses this code.\n\nb) make the code actually _put_ slashes there.\nWait until everybody has used this code.\n\nc) deal with it for all eternity, oops: since rewriting the\n   cryptographic history of existing repositories is pretty much out\n   as far as I understand (which might be insufficient), one has to\n   navigate around slash/noslash all the time when accessing\n   repositories, including the sorting.  The index, however, can at\n   one point of time phase out the slash-specific sorting.  There is\n   no such thing as prehistoric indexes we would need to mind.\n\nI guess that looks like not being worth the pain.  Double the code or\nno money back.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"}]}