{"thread":{"id":"24457","subject":"Avery Pennarun's git-subtree?","startedAt":"2010-07-21T17:15:52Z","lastAt":"2010-07-28T22:27:06Z","messageCount":58,"participants":["Bryan Larsen","Ævar Arnfjörð Bjarmason","Avery Pennarun","Jens Lehmann","Jonathan Nieder","Elijah Newren","Chris Webb","Marc Branchaud","skillzero@gmail.com","Sverre Rabbelier","Jakub Narebski","Nguyen Thai Ngoc Duy","Linus Torvalds","Eugene Sajine","Junio C Hamano"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"145946","messageId":"4C472B48.8050101@gmail.com","threadId":"24457","inReplyTo":null,"subject":"Avery Pennarun's git-subtree?","fromName":"Bryan Larsen","fromEmail":"bryan.larsen@gmail.com","sentAt":"2010-07-21T17:15:52Z","receivedAt":"2010-07-21T17:15:52Z","isPatch":false,"sender":{"key":"bryan@larsen.st","avatar":"https://avatars.githubusercontent.com/u/32073?v=4"},"body":"I've been using Avery Pennarun's git-subtree \n(http://github.com/apenwarr/git-subtree) for a while now and have been \nfinding it very useful and problem-free.\n\nGit submodules have been particularly problematic for me on a project \nwhich contains submodules which contain submodules.  git-subtree \"just \nworks\", without any futzing.\n\nWe've also had problems with less git savvy users dropping patches \nbecause they've occurred inside of a module.\n\nIt would be really nice if git-subtree became an part of git.    Avery \nhas submitted git-subtree in the past and has indicated a willingness to \ndo so again if there was a good chance of acceptance.\n\nAvery's announcment of v0.3 is also informative: \nhttp://kerneltrap.org/mailarchive/git/2010/2/4/22366\n\nthank you,\nBryan\n"},{"id":"145951","messageId":"AANLkTilivtS4TccZXHz2N_n_2RpY6q_5sw7zwdWKdnYE@mail.gmail.com","threadId":"24457","inReplyTo":"4C472B48.8050101@gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2010-07-21T19:43:12Z","receivedAt":"2010-07-21T19:43:12Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"On Wed, Jul 21, 2010 at 17:15, Bryan Larsen <bryan.larsen@gmail.com> wrote:\n> I've been using Avery Pennarun's git-subtree\n> (http://github.com/apenwarr/git-subtree) for a while now and have been\n> finding it very useful and problem-free.\n>\n> Git submodules have been particularly problematic for me on a project which\n> contains submodules which contain submodules.  git-subtree \"just works\",\n> without any futzing.\n>\n> We've also had problems with less git savvy users dropping patches because\n> they've occurred inside of a module.\n\nWhat sort of workflows do you find bad with git-submodule that are\nbetter with git-subtree?\n\nThe submodule concept is simple, but a lot of the implementation is\nbad IMO. It doesn't integrate well, e.g. you have to remember to do\ngit clone --recursive, or git clone and git submodule update --init\nafter that, submodules don't remember what branch you wanted, so git\nsubmodule foreach 'git pull' doesn't DWYM (although I have a hack for\nthat) etc.\n\nI've also wondered if we couldn't just store all the heads .gitmodules\npoint to inside the main .git repository, and just git gc them when\nsubmodules are removed.\n\nI'd planned to maybe submit patches to fix some of these UI issues,\nknowing about more of them would help. I also haven't tried\ngit-subtree.\n"},{"id":"145953","messageId":"AANLkTinl1SB1x1bEObLIo-LWjvxM-Yf1PfdUp4DNJda3@mail.gmail.com","threadId":"24457","inReplyTo":"AANLkTilivtS4TccZXHz2N_n_2RpY6q_5sw7zwdWKdnYE@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-21T19:56:54Z","receivedAt":"2010-07-21T19:56:54Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Wed, Jul 21, 2010 at 3:43 PM, Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n> What sort of workflows do you find bad with git-submodule that are\n> better with git-subtree?\n>\n> The submodule concept is simple, but a lot of the implementation is\n> bad IMO. It doesn't integrate well, e.g. you have to remember to do\n> git clone --recursive, or git clone and git submodule update --init\n> after that, submodules don't remember what branch you wanted, so git\n> submodule foreach 'git pull' doesn't DWYM (although I have a hack for\n> that) etc.\n\nIn my experience, there is exactly one killer problem with submodules\nthat people are looking to solve with git-subtree:\n\nBranching.\n\nIf you have a random developer in your office and they need to make a\npatch to one of your subprojects in the course of making their main\nproject work, with submodules this requires incredibly error-prone\ncontortions involving branching both projects, making sure you have\npush access to both projects, learning how to use git-submodule, etc.\nAnd then merging that branch into someone else's branch is\ncomplicated, particularly if they've also applied their own changes to\nthe subproject.\n\nWith git-subtree, that developer just commits the changes to the\nmerged project - and that's it.  Then you or someone else, who knows\nhow git-subtree works, at any point in the future, can submit the\nsubproject changes upstream, or not, as appropriate.\n\nNo amount of bugfixing in git submodule can fix this workflow, because\nit's not a result of bugs.  (The bugs, particularly the\ndisconnected-by-default HEADs on submodule checkouts, do make it a bit\nworse :( )  It would require a fundamental redesign to make this work\nnicely with submodules.\n\ngit-subtree is certainly a fundamental redesign.  Arguably there might\nbe even better ways to design it, of course.  And submodules are good\nfor certain other situations that git-subtree isn't, so it's obviously\nnot a one-for-one replacement.\n\nIf we can get some kind of consensus in principle that git-subtree is\na good idea to merge into git core, I can prepare some patches and we\ncan talk about the details.\n\nHave fun,\n\nAvery\n"},{"id":"145955","messageId":"AANLkTikl2zKcie3YGhBHrGbYbX3yB9QCtuJTKjsAfK07@mail.gmail.com","threadId":"24457","inReplyTo":"AANLkTinl1SB1x1bEObLIo-LWjvxM-Yf1PfdUp4DNJda3@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2010-07-21T20:36:04Z","receivedAt":"2010-07-21T20:36:04Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"On Wed, Jul 21, 2010 at 19:56, Avery Pennarun <apenwarr@gmail.com> wrote:\n\n> No amount of bugfixing in git submodule can fix this workflow, because\n> it's not a result of bugs.  (The bugs, particularly the\n> disconnected-by-default HEADs on submodule checkouts, do make it a bit\n> worse :( )  It would require a fundamental redesign to make this work\n> nicely with submodules.\n\nI think most of those can be fixed, actually. The only requirement\nthat the git plumbing imposes on git-submodules is that a \"commit\"\nentry exist in your tree, the rest is just (ugly plumbing).\n\nThus, we could:\n\n   * Hack git-submodule (or its replacement) to check import the tree\n     that contains that \"commit\" into one central .git\n\n   * Fix git status / git commit so that you could commit into\n     submodules, i.e.:\n\n     for each submodule in this-commit:\n         chdir $submodule && commit\n     done && cd $root && commit -m\"bumping sumbodules\"\n\n   * Make git-push push the submodule contents and the\n     superprojects. You'd just need to have commit access to the url\n     listed in .gitmodules.\n\nWhat's missing from that (which would be nice) is the ability to check\nout a subdirectory from another repository. That could (I think) be\ndone by just adding a normal \"tree\" entry, and then specifying that\nthat tree can be found in git://... instead of the main tree.\n\n> If we can get some kind of consensus in principle that git-subtree is\n> a good idea to merge into git core, I can prepare some patches and we\n> can talk about the details.\n\nFrom having looked at it briefly it looks very nice. But it looks to\nme as if the main differences between git-submodule and git-subtree\nare in the porcelain, not the plumbing.\n\nIt would be a lot less confusing to users of Git in the long term if\nwe would at least try to unify these two approaches instead of having\ntwo mutually incompatible ways of doing essentially the same thing.\n"},{"id":"145956","messageId":"AANLkTimiROxqf7KcRKTZvMvsFdd4w3jK_GLeZR8n7tdA@mail.gmail.com","threadId":"24457","inReplyTo":"AANLkTikl2zKcie3YGhBHrGbYbX3yB9QCtuJTKjsAfK07@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-21T21:09:17Z","receivedAt":"2010-07-21T21:09:17Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Wed, Jul 21, 2010 at 4:36 PM, Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n> On Wed, Jul 21, 2010 at 19:56, Avery Pennarun <apenwarr@gmail.com> wrote:\n>> No amount of bugfixing in git submodule can fix this workflow, because\n>> it's not a result of bugs.  (The bugs, particularly the\n>> disconnected-by-default HEADs on submodule checkouts, do make it a bit\n>> worse :( )  It would require a fundamental redesign to make this work\n>> nicely with submodules.\n> [...]\n> I think most of those can be fixed, actually. The only requirement\n> that the git plumbing imposes on git-submodules is that a \"commit\"\n> entry exist in your tree, the rest is just (ugly plumbing).\n\nSure.  But this commit object (and the objects it points to) are never\nautomatically pushed, fetched, or fsck'd.  They're second class\ncitizens.  As it turns out, this was a major design mistake in\nimplementing the submodule commit objects.\n\nAll the behaviour people *currently* get from submodules could have\nbeen obtained without using a new 'commit' object type at all.  Just\nadd a commitid to the horrible junk (including repo URLs, argh) that\nalready needs to get pasted into .gitmodules, and have git-commit at\nthe top level update .gitmodules automatically (as it currently\nupdates the 'commit' tree entries).  Problem solved (at least, solved\nto exactly the extent that it is today).\n\nWhat we *really* want is a way to have git actually recurse through\ncommit objects when doing *any* operation, as if they were tree\nobjects.  If we had that, submodules could be beautiful (because you'd\npush them to the same repo, etc and users would see none of the\ncomplexity).  But this doesn't exist.  And for backward compatibility\nat this point, we'd probably need to introduce an entirely new kind of\ntree entry to support such a thing.\n\n> Thus, we could:\n>\n>   * Hack git-submodule (or its replacement) to check import the tree\n>     that contains that \"commit\" into one central .git\n\nThis part is relatively easy, I think - at least in concept, although\nI bet there would be widespread implementation tweaks - and would\nclean up a lot of the mess.  However it would require a change to the\n.git/index file format to remember when a subdir is a commit and not a\n\"normal\" tree so that it doesn't silently commit the next thing as a\ntree instead.\n\n>   * Fix git status / git commit so that you could commit into\n>     submodules, i.e.:\n>\n>     for each submodule in this-commit:\n>         chdir $submodule && commit\n>     done && cd $root && commit -m\"bumping submodules\"\n\nAfter making the earlier change to get rid of the extra .git subdirs,\nthis next requirement would actually be considerably more work,\nbecause 'git commit' would need to know how to update a subcommit\nwithout changing HEAD.  You certainly couldn't just code it up as a\nrecursive \"git commit\" as you imply (and as you could do right now).\n\n>   * Make git-push push the submodule contents and the\n>     superprojects. You'd just need to have commit access to the url\n>     listed in .gitmodules.\n\nThis is really a *killer* problem, and you're making it sound easy.\nLet's imagine that my app has 25 different submodules - not\nunreasonable at all in a world with dozens of ever-changing ruby gems\nand suchlike.\n\nNow, if I want to branch my project, I might have to branch 25\nprojects just so I can push my changes?  It's totally awful.  And the\nawfulness is multiplied many times over if .gitmodules has hard-coded\nrepo paths, because then I have to update the repo path in my branch\nbut not the other branch, and merging will have conflicts.  You might\nthink that my .git/config could just override .gitmodules, but then\nsome guy trying to fetch my branch will fail to fetch the submodules\nfrom my branch and get errors and have no idea what's going on.\n\nAnd you might think that using relative repo paths in .gitmodules\nwould work, but that's only if I branched all 25 submodules in the\n*first* place.  In real life, most subprojects point at the original\nproject's home repo by default (because nobody thinks they'll be\npatching 25 subprojects when they start, and they're probably right),\nbut then you have to individually change the URLs when you decide you\nneed to patch them, and life gets complicated and ugly, especially\nwhen the next guy goes to fork your project and now needs to fork some\nsubprojects but not others.\n\nThere is no good solution to the submodule problem if each submodule\nhas to go in its own repo.  I've been thinking about this for years\nnow, and watching lots of discussions about it on the git mailing\nlist, and I just can't see any other option.  All the submodules have\nto get pushed to and fetched from the same repo by default.  Anything\nelse is insane.\n\nOne option might be to store the submodule commit refs as refs in your\nsuperproject.  That wouldn't actually be so bad, except for the\naforementioned problem that fetch/push/clone/etc don't actually trace\nthrough commit objects when deciding what objects to send you, so\nfetching the ref of the superproject wouldn't autofetch the subproject\nrefs.  Also, you could accidentally delete one of the subproject refs\nand lose tons of history without ever realizing it.  That's error\nprone and confusing... and clutters up your repo refs list with\nadministrative stuff you didn't actually want in the first place.\n\n> What's missing from that (which would be nice) is the ability to check\n> out a subdirectory from another repository. That could (I think) be\n> done by just adding a normal \"tree\" entry, and then specifying that\n> that tree can be found in git://... instead of the main tree.\n\nActually that's already easy with submodules (and git-subtree makes it\neasy too, though slightly different).  Just fetch the commit from the\nother repo, and do:\n\n   git checkout FETCH_HEAD -- subdirname\n\n>> If we can get some kind of consensus in principle that git-subtree is\n>> a good idea to merge into git core, I can prepare some patches and we\n>> can talk about the details.\n>\n> From having looked at it briefly it looks very nice. But it looks to\n> me as if the main differences between git-submodule and git-subtree\n> are in the porcelain, not the plumbing.\n\nNo.  The fundamental difference is exactly one: git-subtree uses\nnormal 'tree' entries (rather than commits) in its trees, so that all\nthe git tools recurse through them like any other tree.  Thus you\ndon't need any extra refs, extra .git dirs, etc.  That allows you to\nbypass all the useless behaviour git has around 'commit' entries.\nThis is very much a plumbing difference.\n\nThe git-submodule porcelain happens to independently be kind of\nannoying and inconvenient, but that would be much easier to fix if it\nweren't for the plumbing-related problems.\n\n> It would be a lot less confusing to users of Git in the long term if\n> we would at least try to unify these two approaches instead of having\n> two mutually incompatible ways of doing essentially the same thing.\n\nTrue.  But I don't have the time, and implementing the new 'commit'\nentry semantics sounds like a lot of work (as opposed to arguing about\nthem, which I guess I'm good at but which seems unproductive).\n\nIn productive terms: git-subtree is solving problems for real users\nright now.  It might solve more problems for more users if it were\nintegrated into the core and thus made \"official.\"  Nothing precludes\nmaking submodules better later.\n\nHave fun,\n\nAvery\n"},{"id":"145958","messageId":"AANLkTik0gqWjvHhULey89JuatjSSYsCwFegfr6E4bUKx@mail.gmail.com","threadId":"24457","inReplyTo":"AANLkTimiROxqf7KcRKTZvMvsFdd4w3jK_GLeZR8n7tdA@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-21T21:20:33Z","receivedAt":"2010-07-21T21:20:33Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Wed, Jul 21, 2010 at 5:09 PM, Avery Pennarun <apenwarr@gmail.com> wrote:\n> All the submodules have\n> to get pushed to and fetched from the same repo by default.  Anything\n> else is insane.\n\n...and just to clarify, by far the least insane option here is to have\nthe whole thing all under a single ref, which is currently impossible\nwith submodules.\n\n>> What's missing from that (which would be nice) is the ability to check\n>> out a subdirectory from another repository. That could (I think) be\n>> done by just adding a normal \"tree\" entry, and then specifying that\n>> that tree can be found in git://... instead of the main tree.\n>\n> Actually that's already easy with submodules (and git-subtree makes it\n> easy too, though slightly different).  Just fetch the commit from the\n> other repo, and do:\n>\n>   git checkout FETCH_HEAD -- subdirname\n\nSorry, that's not right.  You can use this instead for roughly the\neffect you want:\n\n    git read-tree --prefix subdirname FETCH_HEAD: && git checkout subdirname\n\nHave fun,\n\nAvery\n"},{"id":"145969","messageId":"4C4778DE.9090905@web.de","threadId":"24457","inReplyTo":"AANLkTimiROxqf7KcRKTZvMvsFdd4w3jK_GLeZR8n7tdA@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Jens Lehmann","fromEmail":"jens.lehmann@web.de","sentAt":"2010-07-21T22:46:54Z","receivedAt":"2010-07-21T22:46:54Z","isPatch":false,"sender":{"key":"jens.lehmann@web.de","avatar":"https://avatars.githubusercontent.com/u/135220?v=4"},"body":"Am 21.07.2010 23:09, schrieb Avery Pennarun:\n> What we *really* want is a way to have git actually recurse through\n> commit objects when doing *any* operation, as if they were tree\n> objects.\n\nThis would not be useful for every work flow (or to put it in other\nwords: this is not what I *really* want ;-). And as you pointed\nout, that only works when you have a single repo you are working\nagainst (like you do in your subtree model).\n\nBut unless I got something wrong (which might very well be the\ncase, as I never have used subtree myself), all changes to the\nsubtree will only show up in that single repo, unless you actively\npush them somewhere else. And that, it seems to me, is as easy to\nforget as you can right now forget to push a submodules commit you\nalready recorded and pushed in the superproject). So am I wrong\nassuming that subtree is more focused on a single repo containing\nall commits which /might/ then be shared, while submodules are\nabout /always/ sharing code via their own repo?\n\n\n> There is no good solution to the submodule problem if each submodule\n> has to go in its own repo.  I've been thinking about this for years\n> now, and watching lots of discussions about it on the git mailing\n> list, and I just can't see any other option.  All the submodules have\n> to get pushed to and fetched from the same repo by default.  Anything\n> else is insane.\n\nI have to object here. Your insanity is someone else's work flow ;-)\nAnd I am the last one not to admit that there are some severe\nusability warts still to be fixed for submodules (I put up a - not\nnecessarily complete - list at\nhttp://wiki.github.com/jlehmann/git-submod-enhancements/ ). And\nmyself and others are actively working on them (the next bigger\nthing after a new config option about when to consider a submodule\nmodified are recursive checkouts, so that \"git submodule update\"\nwill hopefully be almost obsolete in the near future).\n"},{"id":"145971","messageId":"AANLkTin2a_HoXLuUjq5oryMA4UhZrJ6VNyD3uOZhz5nP@mail.gmail.com","threadId":"24457","inReplyTo":"AANLkTimiROxqf7KcRKTZvMvsFdd4w3jK_GLeZR8n7tdA@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2010-07-21T23:46:26Z","receivedAt":"2010-07-21T23:46:26Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"On Wed, Jul 21, 2010 at 21:09, Avery Pennarun <apenwarr@gmail.com> wrote:\n> On Wed, Jul 21, 2010 at 4:36 PM, Ævar Arnfjörð Bjarmason\n> <avarab@gmail.com> wrote:\n>> On Wed, Jul 21, 2010 at 19:56, Avery Pennarun <apenwarr@gmail.com> wrote:\n>>> No amount of bugfixing in git submodule can fix this workflow, because\n>>> it's not a result of bugs.  (The bugs, particularly the\n>>> disconnected-by-default HEADs on submodule checkouts, do make it a bit\n>>> worse :( )  It would require a fundamental redesign to make this work\n>>> nicely with submodules.\n>> [...]\n>> I think most of those can be fixed, actually. The only requirement\n>> that the git plumbing imposes on git-submodules is that a \"commit\"\n>> entry exist in your tree, the rest is just (ugly plumbing).\n>\n> Sure.  But this commit object (and the objects it points to) are never\n> automatically pushed, fetched, or fsck'd.  They're second class\n> citizens.  As it turns out, this was a major design mistake in\n> implementing the submodule commit objects.\n>\n> All the behaviour people *currently* get from submodules could have\n> been obtained without using a new 'commit' object type at all.  Just\n> add a commitid to the horrible junk (including repo URLs, argh) that\n> already needs to get pasted into .gitmodules, and have git-commit at\n> the top level update .gitmodules automatically (as it currently\n> updates the 'commit' tree entries).  Problem solved (at least, solved\n> to exactly the extent that it is today).\n\nYeah, that does sound better than the current mess.\n\n> What we *really* want is a way to have git actually recurse through\n> commit objects when doing *any* operation, as if they were tree\n> objects.  If we had that, submodules could be beautiful (because you'd\n> push them to the same repo, etc and users would see none of the\n> complexity).  But this doesn't exist.  And for backward compatibility\n> at this point, we'd probably need to introduce an entirely new kind of\n> tree entry to support such a thing.\n>\n>> Thus, we could:\n>>\n>>   * Hack git-submodule (or its replacement) to check import the tree\n>>     that contains that \"commit\" into one central .git\n>\n> This part is relatively easy, I think - at least in concept, although\n> I bet there would be widespread implementation tweaks - and would\n> clean up a lot of the mess.  However it would require a change to the\n> .git/index file format to remember when a subdir is a commit and not a\n> \"normal\" tree so that it doesn't silently commit the next thing as a\n> tree instead.\n>\n>>   * Fix git status / git commit so that you could commit into\n>>     submodules, i.e.:\n>>\n>>     for each submodule in this-commit:\n>>         chdir $submodule && commit\n>>     done && cd $root && commit -m\"bumping submodules\"\n>\n> After making the earlier change to get rid of the extra .git subdirs,\n> this next requirement would actually be considerably more work,\n> because 'git commit' would need to know how to update a subcommit\n> without changing HEAD.  You certainly couldn't just code it up as a\n> recursive \"git commit\" as you imply (and as you could do right now).\n>\n>>   * Make git-push push the submodule contents and the\n>>     superprojects. You'd just need to have commit access to the url\n>>     listed in .gitmodules.\n>\n> This is really a *killer* problem, and you're making it sound easy.\n> Let's imagine that my app has 25 different submodules - not\n> unreasonable at all in a world with dozens of ever-changing ruby gems\n> and suchlike.\n>\n> Now, if I want to branch my project, I might have to branch 25\n> projects just so I can push my changes?  It's totally awful.  And the\n> awfulness is multiplied many times over if .gitmodules has hard-coded\n> repo paths, because then I have to update the repo path in my branch\n> but not the other branch, and merging will have conflicts.  You might\n> think that my .git/config could just override .gitmodules, but then\n> some guy trying to fetch my branch will fail to fetch the submodules\n> from my branch and get errors and have no idea what's going on.\n>\n> And you might think that using relative repo paths in .gitmodules\n> would work, but that's only if I branched all 25 submodules in the\n> *first* place.  In real life, most subprojects point at the original\n> project's home repo by default (because nobody thinks they'll be\n> patching 25 subprojects when they start, and they're probably right),\n> but then you have to individually change the URLs when you decide you\n> need to patch them, and life gets complicated and ugly, especially\n> when the next guy goes to fork your project and now needs to fork some\n> subprojects but not others.\n>\n> There is no good solution to the submodule problem if each submodule\n> has to go in its own repo.  I've been thinking about this for years\n> now, and watching lots of discussions about it on the git mailing\n> list, and I just can't see any other option.  All the submodules have\n> to get pushed to and fetched from the same repo by default.  Anything\n> else is insane.\n\nYeah, bundling the submodules in the upstream repo so only one person\never has to worry about gathering them up and pushing them to the\ncentral repo sounds better for most uses than the current submodule\nimplementation.\n\nOTOH, I have some submodules that I track on GitHub that would really\ninflate the size of the repo that's tracking them. So there are\ndefinitely use cases for having the tree somewhere remotely as well,\nespecially for large submodules like game art, which some people have\nreported submodules for.\n\n> One option might be to store the submodule commit refs as refs in your\n> superproject.  That wouldn't actually be so bad, except for the\n> aforementioned problem that fetch/push/clone/etc don't actually trace\n> through commit objects when deciding what objects to send you, so\n> fetching the ref of the superproject wouldn't autofetch the subproject\n> refs.  Also, you could accidentally delete one of the subproject refs\n> and lose tons of history without ever realizing it.  That's error\n> prone and confusing... and clutters up your repo refs list with\n> administrative stuff you didn't actually want in the first place.\n>\n>> What's missing from that (which would be nice) is the ability to check\n>> out a subdirectory from another repository. That could (I think) be\n>> done by just adding a normal \"tree\" entry, and then specifying that\n>> that tree can be found in git://... instead of the main tree.\n>\n> Actually that's already easy with submodules (and git-subtree makes it\n> easy too, though slightly different).  Just fetch the commit from the\n> other repo, and do:\n>\n>   git checkout FETCH_HEAD -- subdirname\n>\n>>> If we can get some kind of consensus in principle that git-subtree is\n>>> a good idea to merge into git core, I can prepare some patches and we\n>>> can talk about the details.\n>>\n>> From having looked at it briefly it looks very nice. But it looks to\n>> me as if the main differences between git-submodule and git-subtree\n>> are in the porcelain, not the plumbing.\n>\n> No.  The fundamental difference is exactly one: git-subtree uses\n> normal 'tree' entries (rather than commits) in its trees, so that all\n> the git tools recurse through them like any other tree.  Thus you\n> don't need any extra refs, extra .git dirs, etc.  That allows you to\n> bypass all the useless behaviour git has around 'commit' entries.\n> This is very much a plumbing difference.\n>\n> The git-submodule porcelain happens to independently be kind of\n> annoying and inconvenient, but that would be much easier to fix if it\n> weren't for the plumbing-related problems.\n>\n>> It would be a lot less confusing to users of Git in the long term if\n>> we would at least try to unify these two approaches instead of having\n>> two mutually incompatible ways of doing essentially the same thing.\n>\n> True.  But I don't have the time, and implementing the new 'commit'\n> entry semantics sounds like a lot of work (as opposed to arguing about\n> them, which I guess I'm good at but which seems unproductive).\n>\n> In productive terms: git-subtree is solving problems for real users\n> right now.  It might solve more problems for more users if it were\n> integrated into the core and thus made \"official.\"  Nothing precludes\n> making submodules better later.\n\nSure, don't get me wrong. git-subtree looks very useful, and I have no\nobjection to having it in git.git, and even if it's not optimal for\neverything good working software now shouldn't be held up by some\ntheoretical pie-in-the-sky system.\n"},{"id":"145977","messageId":"AANLkTim9nfRGjhpn2Mj-1GntLsDX7xeyL2pegB84aZX8@mail.gmail.com","threadId":"24457","inReplyTo":"4C4778DE.9090905@web.de","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-22T01:09:06Z","receivedAt":"2010-07-22T01:09:06Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Wed, Jul 21, 2010 at 6:46 PM, Jens Lehmann <Jens.Lehmann@web.de> wrote:\n> Am 21.07.2010 23:09, schrieb Avery Pennarun:\n>> What we *really* want is a way to have git actually recurse through\n>> commit objects when doing *any* operation, as if they were tree\n>> objects.\n>\n> This would not be useful for every work flow (or to put it in other\n> words: this is not what I *really* want ;-). And as you pointed\n> out, that only works when you have a single repo you are working\n> against (like you do in your subtree model).\n\nBut you see, the utter failure of the way git-submodule works is that\nit required a change to the git repository format, but that repository\nformat change resulted in absolutely *zero* improvement.\n\nThe tree object of the parent points at 'commit xxxx'.  But everything\nin git has been *specially modified* to *just ignore* that 'commit\nxxxx'.  It would have given exactly the same functionality - and much\nless confusingly - if .gitmodules would just include the desired\ncommitid of the child project.  You could still have the same 'git\nsubmodule' command with the same syntax and semantics.  And it\nwouldn't have bastardized the git repo format.\n\nIt would have been just as good to just dump something into your\nMakefile to go 'git clone' the subprojects from somewhere before\nbuilding.  Seriously, it would be one or two lines of code; all of\ngit-submodule replaces about one or two lines of code in your\nMakefile.  And you know what?  If I just used that one or two lines of\ncode, I'd have all sorts of flexibility in where the subprojects get\ncloned from, which I currently don't have, and which is the insanity\nthat drove me to write git-subtree in the first place.\n\nHOWEVER\n\nI'm not saying we can change that now.  I'm not suggesting that this\nfeature can be safely removed or changed at all.  Furthermore, I\ntotally agree that having large subprojects *not* be in your repo is\nsometimes a good idea.  I just think it was actually a bad idea to\nintrusively add support to git to implement this when it could have\nbeen done without modifying git at all.\n\nI also believe that the vast majority of people who use git-submodules\nwould rather have it work differently.  (Again, this is not to\nsubtract functionality.  The existing functionality is useful\nsometimes.)\n\n> But unless I got something wrong (which might very well be the\n> case, as I never have used subtree myself), all changes to the\n> subtree will only show up in that single repo, unless you actively\n> push them somewhere else. And that, it seems to me, is as easy to\n> forget as you can right now forget to push a submodules commit you\n> already recorded and pushed in the superproject). So am I wrong\n> assuming that subtree is more focused on a single repo containing\n> all commits which /might/ then be shared, while submodules are\n> about /always/ sharing code via their own repo?\n\nYes, this is absolutely intentional.  It also matches exactly with\neverything else in the git repo philosophy!\n\nI make my own clone.  I mess with it, I fiddle with it, I make 17\nclones on my local machine, I throw away what I don't like, I pull\nmerge, I rebase, and then *eventually* I submit *some* of my patches\nupstream.  git-subtree lets you do all those things.  git-submodule\nstomps on you repeatedly if you try.\n\nTo wit:\n\n- cloning a local supermodule on my local machine to another copy:\nevery call to 'git submodule update' re-downloads submodule repos from\nthe remote machine, because the submodule path is hardcoded to point\nat a remote machine.  Better still, if I've modified any of my\nsubprojects without pushing changes upstream, the clone will fail,\nbecause the new copy of the superproject will have no access to my\nsubproject's patches.  (If .gitmodules supplies a relative path, it's\neven worse, because my 'origin' in the new copy is now pointing to a\nlocal folder, not a remote one, and all the submodules don't exist\nthere.)\n\n- branching a local supermodule on my local machine: fails to branch\nthe submodule automatically and makes it super easy to lose patches\naltogether (since by default, they're committed to a detached HEAD).\n\n- pulling/merging: always causes a conflict if local and remote have\nmodified the same submodule.\n\n- rebasing: always causes a conflict if local and remote have modified\nthe same submodule.  Also requires you to rebase submodules separately\nfrom the supermodule.  (Yes, this happens often in real life.)\n\n- submitting upstream: requires me to have a separate repo that's a\ncopy of the upstream repo, and to manage at least one subrepo branch\nfor every superproject branch, just to track my submissions.  With\ngit-subtree, no extra repos are necessary.\n\nIt's very clear that git-submodule's current behaviour totally\nmismatches the entire git philosophy.  That's why it's so impossible\nto make the git-submodule command usable.\n\nAnother mental exercise: try to think of any other part of git where\nit would be considered remotely acceptable to put the absolute or\nrelative URL of one repo inside another repo.  git URLs are an\nimplementation detail of clone/fetch/push/pull.  The *content* that\ngit manages should not have to deal with that stuff.  With\ngit-submodule, it has to.  With git-subtree, it doesn't.\n\n>> There is no good solution to the submodule problem if each submodule\n>> has to go in its own repo.  I've been thinking about this for years\n>> now, and watching lots of discussions about it on the git mailing\n>> list, and I just can't see any other option.  All the submodules have\n>> to get pushed to and fetched from the same repo by default.  Anything\n>> else is insane.\n>\n> I have to object here. Your insanity is someone else's work flow ;-)\n\nSorry.  I was being a little hyperbolic.  Some people might want to do\nuse multiple repos for certain things - but I believe those people are\nmuch more rare than the kind who want to do it my way.  And\nfurthermore, even those people would probably actually like it better\nif *most* of their subprojects - the smallish ones - could be all in\none repo.\n\nEven if you like multiple repos, I'm sure you don't like being\n*forced* to manually fork multiple repos just to fork a single\nsuperproject.  I'm sure you don't like updating .gitmodules to change\nthe absolute URL of a submodule, and then getting merge conflicts when\nsomeone else had to do the same thing.  There's no way you like that.\nIf you like that, then you really are insane. :)\n\n> And I am the last one not to admit that there are some severe\n> usability warts still to be fixed for submodules (I put up a - not\n> necessarily complete - list at\n> http://wiki.github.com/jlehmann/git-submod-enhancements/ ). And\n> myself and others are actively working on them (the next bigger\n> thing after a new config option about when to consider a submodule\n> modified are recursive checkouts, so that \"git submodule update\"\n> will hopefully be almost obsolete in the near future).\n\nI don't believe you can fix git-submodule by fixing surface warts.\nIt's fundamentally broken.  Since we're stuck with supporting the\ncurrent behaviour at the end of time, fixing the surface warts might\nbe necessary and even mildly helpful.  It will also be soul sucking\nsince no matter how hard you try, people will still hate the result.\n\nHave fun,\n\nAvery\n"},{"id":"146007","messageId":"4C488C9D.60606@gmail.com","threadId":"24457","inReplyTo":"m31vavn8la.fsf@localhost.localdomain","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Bryan Larsen","fromEmail":"bryan.larsen@gmail.com","sentAt":"2010-07-22T18:23:25Z","receivedAt":"2010-07-22T18:23:25Z","isPatch":false,"sender":{"key":"bryan@larsen.st","avatar":"https://avatars.githubusercontent.com/u/32073?v=4"},"body":">\n> Using git-subtree has its warts too: I don't think for example that there is\n> a way to get a log _automatically excluding_ history subtree-merged\n> subprojects.  Or is it there?\n>\n\nIt works exactly right for me when I used git-subtree in \"squashed\" \nmode.  Changes which were done in tree show up separately in the log, \nchanges which were pulled in via git-subtree pull show up as a single \nsummary entry in the log.\n\nThis discussion has been about how to improve git submodules, which is \nsorely needed.   However, it's quite clear that git submodules will \nnever work as well as git subtrees in certain quite common situations. \n  If fixed, git submodules will be more appropriate in other situations. \n   However, I'm not asking to remove git submodules or prevent anybody \nfrom fixing them, I'm just asking that git subtree be merged.\n\nDoes anybody actually oppose the merger of git-subtree, which has (at \nleast) hundreds of users despite its out-of-tree status?\n\nthanks,\nBryan\n"},{"id":"146011","messageId":"AANLkTimOb2VjYI21wQsC64lm4HsVPwpRWd1twIUBnbJ3@mail.gmail.com","threadId":"24457","inReplyTo":"m31vavn8la.fsf@localhost.localdomain","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-22T19:41:52Z","receivedAt":"2010-07-22T19:41:52Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Thu, Jul 22, 2010 at 5:57 AM, Jakub Narebski <jnareb@gmail.com> wrote:\n> Avery Pennarun <apenwarr@gmail.com> writes:\n>> The tree object of the parent points at 'commit xxxx'.  But everything\n>> in git has been *specially modified* to *just ignore* that 'commit\n>> xxxx'.  It would have given exactly the same functionality - and much\n>> less confusingly - if .gitmodules would just include the desired\n>> commitid of the child project.  You could still have the same 'git\n>> submodule' command with the same syntax and semantics.  And it\n>> wouldn't have bastardized the git repo format.\n>\n> Actually the prototype implementation by Martin Waitz worked in such way,\n> i.e. with special file in top directory holding SHA-1 of submodule commits,\n> what you can read on https://git.wiki.kernel.org/index.php/SubprojectSupport\n> page.\n>\n> The low level plumbing with 'commit' entries in the 'tree' object was\n> created by Linus Torvalds (CC-ed).  I don't remember discussion about why\n> this solution was chosen, though.  But please read about differences between\n> git-subtree and git-submodule below.\n\nI actually think Linus's contribution - the particular change to the\nrepo format to have trees link to commits - was exactly right.  If we\nwant to talk about failings of git-subtree, they all precisely come\ndown to the fact that, because it has tree->tree links instead of\ntree->commit links, it has to stash commitid information in the commit\nmessage, which is gross and error prone.\n\ngit-subtree would have benefitted from tree->commit links, but because\ngit's implementation of them is broken, that wasn't an option.\n\nUnfortunately everything built *on top of* Linus's file format\ncontribution has turned out to be a disaster.  Actually making the\nsubprojects have their own local .git repositories was a disaster, for\nexactly the same reasons that having every subdir in svn have its own\n.svn directory (or in cvs, every directory has its own CVS directory)\nis a disaster.  When you split things up that way, you can't easily do\nglobal atomic operations across the entire set of content.  And you\ncan accidentally have a subdir pointing at a totally different place\nthan the parent thinks it is.  And you have CVS/.svn/.git directories\ncluttering stuff up everywhere.\n\nThe tree->commit links do not preclude you doing wonderful global\natomic operations across the entire set of content.  The separate\nrepository garbage absolutely does.\n\n>> To wit:\n>>\n>> - cloning a local supermodule on my local machine to another copy:\n>> every call to 'git submodule update' re-downloads submodule repos from\n>> the remote machine, because the submodule path is hardcoded to point\n>> at a remote machine.\n>\n> Errrr... the URL to submodule repository (I guess it is what you meant here\n> by \"submodule path\") in the config file overrides URL to submodule\n> repository in '.gitmodules' for a reason.  So the plumbing support is here,\n> it is only failing of an UI that we don't have '--recursive-local' or\n> '--convert-submodules' (like '--convert-links' in wget) in \"git clone\".\n\nLet me be more specific.\n\nI create an app named myapp on github:\n\n   git://github.com/apenwarr/myapp\n\nIt uses 17 different ruby gems, which I import as subprojects.  I have\ntwo choices:\n\n[1] .gitmodules can use absolute paths to the original gem locations:\n\n   git://github.com/rubygems/gem[1..n]\n\n[2] Or else I can fork them all and use relative paths in .gitmodules:\n\n   ../gem[1..n]\n   translates to --> git://github.com/apenwarr/gem[1..n]\n\nAt this phase, both options are okay (though option #2 is obviously\nmuch more work).  My next step will be to clone myapp onto my local\nmachine:\n\n   git clone --recursive git://github.com/apenwarr/myapp\n\nAnd it will grab all the submodules just fine.\n\nNow let's say I want to change gem13.  If I used option #1, I have to\nnow go fork gem13 on github.  Then do one of the following:\n\n[1A] Re-point my .git/config file to point at the new submodule\nlocation, git://github.com/apenwarr/gem13 but leave .gitmodules alone\n\n[1B] or update both .git/config and .gitmodules\n\nIf I do #1A, then when I push my changes, the *next* guy who clones\ngit://github.com/apenwarr/myapp will fail; the gem13 link in myapp\npoints at a commit that is *only* in apenwarr/gem13, not\nrubygems/gem13.\n\nIf I do #1B, then if someone else does something similar in their own\ncopy and pulls from me, we will have a conflict in .gitmodules.\n\nIn both cases, if two people need to patch gem13 during their changes\nto myapp, merges will fail because there is no submodule-recursive\nmerge (and trying to write one would be incredibly hard since it would\nhave to communicate across sub-repositories).\n\nSo if you do #1, then I don't know of any options other than #1A and\n#1B, and neither one works.\n\nNow, if I had done #2 instead, things are a little better, because\nwe're using relative paths in .gitmodules so when the second guy\nclones a copy of myapp, he can also clone a copy of all 17 gems, and\nall the paths will still work.\n\nWhen the second guy does 'git pull apenwarr myapp' it will still fail,\nthough; it will try to get the latest gem13 from ../gem13 -->\nsecondguy/gem13, when actually the required commits are in\napenwarr/gem13.\n\nFurthermore, 'git clone --recursive myapp myapp2' will totally fail,\nbecause it will then expect gem[1..n] to all be in separate local\ndirectories at the same level as myapp, which they aren't.  (You might\nbe saying: what do you need that for?  Well, I rarely do.  But\nsometimes.  And as long as I don't use git-submodule, it works fine.)\n\nYou can fix warts all day long.  You can't make it work, because it's\nnot just warts; the insides are rotten.\n\n>> - branching a local supermodule on my local machine: fails to branch\n>> the submodule automatically and makes it super easy to lose patches\n>> altogether (since by default, they're committed to a detached HEAD).\n>\n> That's UI problem, too.  Theough I guess that using detached HEAD was\n> choosen because it is simplest solution.\n\nI've seen the discussion about submodule branch names go by on the git\nlist a few times, and I participated once or twice.  The current\noption was certainly chosen because it's the simplest; unfortunately,\nit's also non-functional, and all the other options are also awful.\n\nHere it is in a nutshell: if I'm branching myapp, I already have a\nbranch that I want to store all my changes under; it's the branch I'm\nworking on in myapp.  That's not to say I want that same myapp branch\nname *in my gem13 repository*; my branchname is probably something\nlike add-feature-to-myapp, which has nothing to do with gem13.  The\nchanges required to gem13 to implement add-feature-to-myapp are\nprobably just a tiny bugfix or config option.  gem13 doesn't know\nanything about myapp.  The upstream gem13 maintainers certainly don't\ncare about myapp.  As a guy who *just wants to get work done on myapp\nright now*, thinking about what to name my trivial one-patch temporary\nbranch gem13 is a *waste of time*.\n\nI don't *want* my gem13 changes to have a branchname.\n\nSo the disconnected HEAD is the right answer then, right?  No!  The\ndefault disconnected HEAD makes it *far* too easy to lose my changes.\nI don't want to name my branch, but I *have* to, because I *have* to\npush it somewhere separately, because if I don't, then my changes to\nmyapp will be useless to everyone who tries to pull from me.\n\nThe question of what to name the submodule branch is unanswerable\nbecause it's the wrong question.\n\n> Otherwise you would have either\n> put submodule branch name in '.gitmodules' (but that's contrary to git\n> philosophy that branches are ephemeral and branch names are local matter),\n\nSurely including *repository URLs* inside the *repository content* is\nat least as bad as including branch names.  If we're going to do one,\nwe might as well do the other.  But it won't help, because the stored\nbranch name will probably be 'master', and my personal hacked-up copy\nof gem13 shouldn't be on a branch named master anyway.\n\n>> - pulling/merging: always causes a conflict if local and remote have\n>> modified the same submodule.\n>>\n>> - rebasing: always causes a conflict if local and remote have modified\n>> the same submodule.  Also requires you to rebase submodules separately\n>> from the supermodule.  (Yes, this happens often in real life.)\n>\n> That's a matter of UI, and lack of merge strategy that can merge\n> submodules... although if I remember correctly there was some preliminary or\n> proof of concept work on submodule-aware merge strategy.\n>\n> \"git merge\" and \"git rebase\" would have to acquire '--recursive' option.\n> Currently you probably need to use 'git submodule foreach ...', I guess.\n\nMerge and rebase are actually very different here.  Merging is\nsomething I might expect to work across submodules eventually;\nrebasing is much less obvious, because successive versions of myapp\nmight actually be jumping back and forth between versions of gem13.\nThen what does it mean to auto-rebase gem13 when you're rebasing\nmyapp?\n\nYou should check out git-subtree --squash here; it's quite interesting\nand makes rebasing easy, even if the subtree version is alternating\nback and forth.  I'm not sure how you'd map it onto git-submodule,\nthough, even if git-submodule weren't broken.\n\n>> - submitting upstream: requires me to have a separate repo that's a\n>> copy of the upstream repo, and to manage at least one subrepo branch\n>> for every superproject branch, just to track my submissions.  With\n>> git-subtree, no extra repos are necessary.\n>\n> NOTE that it is important design decision to have by default separate object\n> storage for submodules.\n\nI certainly won't deny that :)  This discussion is about whether it\nwas the right decision.\n\n> First, this allow to not clone submodule, and do not download its objects.\n> This is *impossible* with git-subtree (with using 'subtree' merge strategy).\n> I'm not sure how commonly this feature is used in real life, but somebody\n> here in this thread gave example of submodule with arts, which is large\n> because it contains large / many binary files, while being required to have\n> only for some.\n>\n> Second, from what I remember this was implemented also for perfomance\n> reasons... though I don't remember reasoning used.\n\nI think this ended up being a terrible mistake.  The problems you\nidentify come down to this:\n\n1) Sometimes I want to clone only some subdirs of a project\n2) Sometimes I don't want the entire history because it's too big.\n3) Super huge git repositories start to degrade in performance.\n\n(Actually #3 isn't really a problem as far as I've ever seen, and bup\nstores hundreds of gigs, including trees that reference millions of\nblobs, in a single git repo without dying.  But okay, maybe this is a\nproblem sometimes for some types of operations.)\n\nThese problems come up regardless of whether you're using submodules.\nThe hard truth of the matter is that people are using submodules to\ntry to solve these problems, but they were never caused by the lack of\nsubmodules in the first place.\n\nWhen I clone the Linux kernel, sometimes I just don't want the entire\nhistory.  That's why people invented shallow clones (although last\ntime I checked, they were still a little half-assed).\n\nWhen I clone KDE, sometimes I don't want all the subprograms;\nsometimes I do. That's why people invented sparse checkouts, and why\n(I think) it would be nice to have sparse clones as well (where you\ndon't even download the objects for subtrees you don't care about).\n\nThere is simply not a clear path from \"my repo is too big\" to \"all my\nproblems will be solved if git-submodule is implemented correctly.\"\n\nThe truth is, problems 1-3 are easily solvable by improving the git\nimplementation, without any change in architecture and without\nrequiring people to layout their projects differently.\n\nThe *real* need for submodules - the need you can't fix without\nsubmodules - has nothing to do with these requirements.  It's about\neach submodule wanting to have its own lifecycle, owner, changelog,\nand release process, and - perhaps this is actually the killer\nrequirement - each supermodule wanting to be able to cleanly rewind a\nsubmodule if they don't like the new version.\n\n>> It's very clear that git-submodule's current behaviour totally\n>> mismatches the entire git philosophy.  That's why it's so impossible\n>> to make the git-submodule command usable.\n>\n> That's very strong accusation.\n\nAgreed... but that doesn't make it wrong :)\n\n> Using git-subtree has its warts too: I don't think for example that there is\n> a way to get a log _automatically excluding_ history subtree-merged\n> subprojects.  Or is it there?\n\nThere's git-subtree merge --squash.  It's pretty cool.  Also insane\nand not as good as real tree->commit links.  I will gladly admit to\ngit-subtree's warts.\n\n>\n> Sumodule                           | Subtree\n> -----------------------------------+----------------------------------\n> must clone recursively submodules; | automatically gets all subtrees\n\nYup.\n\n> can not clone some submodules      | cannot leave out some subtree, but\n>                                   | nowadays can not checkout it\n\nI don't understand what you mean on the right-hand side here.  FWIW,\nsubtree forces you to always checkout the entire thing (unless you use\ngit sparse checkouts, I guess; maybe that's what you mean).\n\n> rebase and merge needs separate    | rebase and merge works normally\n> work in submodule currently        |\n\nTrue.\n\n> easy to send updates upstream      | need not to worry about submodule\n> to submodule repo                  | repository\n\nIt's actually easy to send subtree updates upstream with the new 'git\nsubtree push' command, which was contributed recently.  Or you can\nsend them via format-patch if you use 'git subtree split'.  It's one\nline more of typing than doing it on a submodule repo, and that one\nline is greatly offset by the hugely reduced typing by not using\nsubmodules.\n\nHave fun,\n\nAvery\n"},{"id":"146013","messageId":"20100722195653.GC4439@burratino","threadId":"24457","inReplyTo":"AANLkTimOb2VjYI21wQsC64lm4HsVPwpRWd1twIUBnbJ3@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2010-07-22T19:56:53Z","receivedAt":"2010-07-22T19:56:53Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Avery Pennarun wrote:\n\n> Unfortunately everything built *on top of* Linus's file format\n> contribution has turned out to be a disaster.\n\nAside: this kind of statement might make it unlikely for exactly\nthose who would benefit most from your opinions to read them.\n\nWell, that is my guess, anyway.  I know that I have not found the time\nto read your email (though I would like to) because I suspect based on\nsuch sweeping statements that it would take a while to separate the\nuseful part from the rest.\n\nOf course I am glad to see people thinking about these issues.\nMy comment is only about how the results get presented.\n\nJonathan\n"},{"id":"146015","messageId":"AANLkTinI6uOQyJcJvbNLhNde8yURyQMSua438TZQKyXv@mail.gmail.com","threadId":"24457","inReplyTo":"20100722195653.GC4439@burratino","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-22T20:06:02Z","receivedAt":"2010-07-22T20:06:02Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Thu, Jul 22, 2010 at 3:56 PM, Jonathan Nieder <jrnieder@gmail.com> wrote:\n> Avery Pennarun wrote:\n>> Unfortunately everything built *on top of* Linus's file format\n>> contribution has turned out to be a disaster.\n>\n> Aside: this kind of statement might make it unlikely for exactly\n> those who would benefit most from your opinions to read them.\n>\n> Well, that is my guess, anyway.  I know that I have not found the time\n> to read your email (though I would like to) because I suspect based on\n> such sweeping statements that it would take a while to separate the\n> useful part from the rest.\n\nUnfortunately you will find that the rest of my email more or less\njust expands in detail on those sweeping statements.\n\nSorry.\n\nAvery\n"},{"id":"146017","messageId":"AANLkTilhKR5wuJPPIF1SiRcTJ0fmz1oqp_NfuSSuKMOn@mail.gmail.com","threadId":"24457","inReplyTo":"20100722195653.GC4439@burratino","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2010-07-22T20:17:08Z","receivedAt":"2010-07-22T20:17:08Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"On Thu, Jul 22, 2010 at 19:56, Jonathan Nieder <jrnieder@gmail.com> wrote:\n> Avery Pennarun wrote:\n>\n>> Unfortunately everything built *on top of* Linus's file format\n>> contribution has turned out to be a disaster.\n>\n> Aside: this kind of statement might make it unlikely for exactly\n> those who would benefit most from your opinions to read them.\n>\n> Well, that is my guess, anyway.  I know that I have not found the time\n> to read your email (though I would like to) because I suspect based on\n> such sweeping statements that it would take a while to separate the\n> useful part from the rest.\n>\n> Of course I am glad to see people thinking about these issues.\n> My comment is only about how the results get presented.\n\nWell, it's not like Linus is the image of calmness when attacking\nsomething he perceives as crap design either >:)\n\nAnyway, to answer Bryan's question. My comments in previous messages\nshouldn't be interpreted as opposition to git-subtree being merged at\nall. It's clearly very useful, especially for cases where\ngit-submodule is wanting. I'd be happy to review a patch that\nintegrated it into the Git tree.\n\nBut it's also clear that we have a lot of tribal knowledge about the\nlackings of git submodule / git subtree. It would be *really* useful\nif people like Avery and Jens which have obviously thought hard about\nthe submodule/subtree issues would draft up some (calmly written) docs\nabout how the two differ (with comparison tables etc.).\n\nThat'd be a very helpful resource for Git users in deciding which one\nto use.\n"},{"id":"146020","messageId":"AANLkTikUqbAGPXcAmGsx_oML0tZHpMWHOFl7CCkjBca0@mail.gmail.com","threadId":"24457","inReplyTo":"20100722195653.GC4439@burratino","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2010-07-22T20:43:08Z","receivedAt":"2010-07-22T20:43:08Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"Hi,\n\nOn Thu, Jul 22, 2010 at 1:56 PM, Jonathan Nieder <jrnieder@gmail.com> wrote:\n> Avery Pennarun wrote:\n>\n>> Unfortunately everything built *on top of* Linus's file format\n>> contribution has turned out to be a disaster.\n>\n> Aside: this kind of statement might make it unlikely for exactly\n> those who would benefit most from your opinions to read them.\n>\n> Well, that is my guess, anyway.  I know that I have not found the time\n> to read your email (though I would like to) because I suspect based on\n> such sweeping statements that it would take a while to separate the\n> useful part from the rest.\n\nI'd usually agree with such a sentiment, but I don't think it's\naccurate in this case.  Having read Avery's emails in this thread, I\nthink he does a really good job explaining why submodules don't (and\nwon't) work for a lot of people.  I think he provided a better\nexplanation than I could have for why I've never had much luck with\nsubmodules (and further convinced me that not only do they not work\nfor me now, but they aren't ever going to fulfill the usecases I had).\n\nI can't really add much other than that we've been relatively happy\nwith git-subtree and would like to see it or something like it merged.\n Our problems with it so far have turned out to be issues in other\nareas of git (e.g. the known issue about --prefix being ignored with\nthe code being merged under a different directory due to\nrename-detection, and the bugs in merge-recursive's handling of D/F\nchanges).\n\n\nElijah\n"},{"id":"146024","messageId":"AANLkTinLOQEXk3AIKxC_prMO_j7hcFVh-zW6hqFfuhhI@mail.gmail.com","threadId":"24457","inReplyTo":"AANLkTikUqbAGPXcAmGsx_oML0tZHpMWHOFl7CCkjBca0@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-22T21:32:34Z","receivedAt":"2010-07-22T21:32:34Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Thu, Jul 22, 2010 at 4:43 PM, Elijah Newren <newren@gmail.com> wrote:\n> (e.g. the known issue about --prefix being ignored with\n> the code being merged under a different directory due to\n> rename-detection, [...])\n\nAside: if this is the bug I think it is, then it's is fixed by the git\nmerge -Xsubtree feature, which has since been merged into git.  (I\nthink Elijah knew that, I just wanted to make sure it's clear to\nanyone else reading.)\n\nHave fun,\n\nAvery\n"},{"id":"146025","messageId":"AANLkTil2ZHsV-PH46wfeqqvB27akRw3egDWRtJbIPLXb@mail.gmail.com","threadId":"24457","inReplyTo":"AANLkTilhKR5wuJPPIF1SiRcTJ0fmz1oqp_NfuSSuKMOn@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-22T21:33:57Z","receivedAt":"2010-07-22T21:33:57Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Thu, Jul 22, 2010 at 4:17 PM, Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n> But it's also clear that we have a lot of tribal knowledge about the\n> lackings of git submodule / git subtree. It would be *really* useful\n> if people like Avery and Jens which have obviously thought hard about\n> the submodule/subtree issues would draft up some (calmly written) docs\n> about how the two differ (with comparison tables etc.).\n>\n> That'd be a very helpful resource for Git users in deciding which one\n> to use.\n\nI think I'm too biased to write that, but if someone else wants to\ntake the lead, I could certainly contribute.\n\nHave fun,\n\nAvery\n"},{"id":"146045","messageId":"20100723083149.GD27082@arachsys.com","threadId":"24457","inReplyTo":"AANLkTimOb2VjYI21wQsC64lm4HsVPwpRWd1twIUBnbJ3@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Chris Webb","fromEmail":"chris@arachsys.com","sentAt":"2010-07-23T08:31:50Z","receivedAt":"2010-07-23T08:31:50Z","isPatch":false,"sender":{"key":"chris@arachsys.com","avatar":"https://avatars.githubusercontent.com/u/299056?v=4"},"body":"Avery Pennarun <apenwarr@gmail.com> writes:\n\n> I actually think Linus's contribution - the particular change to the\n> repo format to have trees link to commits - was exactly right.  If we\n> want to talk about failings of git-subtree, they all precisely come\n> down to the fact that, because it has tree->tree links instead of\n> tree->commit links, it has to stash commitid information in the commit\n> message, which is gross and error prone.\n> \n> git-subtree would have benefitted from tree->commit links, but because\n> git's implementation of them is broken, that wasn't an option.\n\nI considered using submodules for one of my projects, and decided against\nfor some of the usability reasons with multiple repositories which you\nhighlight. (I didn't know about subtree.)\n\nYou've surely considered this already, but reading your description in this\nthread, my first thought is that commits within trees could mean different\nthings depending on whether they're at paths listed in .gitmodules or not.\nIf the path is listed, the commit is in an external repository. If it isn't,\nit's a reference to a local commit, allowing submodules to live in the same\nrepo as their parent and share some of the advantages you describe for\nsub-tree.\n\nOver time, git could then become smarter about recursing through commits in\ntrees, although I can see a potential problem with needing to know about a\n.gitmodules blob in the top-level tree when we're examining a deeper level\ntree.\n\nCheers,\n\nChris.\n"},{"id":"146044","messageId":"AANLkTinBtXmMei2Q6MZrXWxa3t+_quGdzpcq46EZvgvG@mail.gmail.com","threadId":"24457","inReplyTo":"20100723083149.GD27082@arachsys.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-23T08:40:44Z","receivedAt":"2010-07-23T08:40:44Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Fri, Jul 23, 2010 at 4:31 AM, Chris Webb <chris@arachsys.com> wrote:\n> You've surely considered this already, but reading your description in this\n> thread, my first thought is that commits within trees could mean different\n> things depending on whether they're at paths listed in .gitmodules or not.\n> If the path is listed, the commit is in an external repository. If it isn't,\n> it's a reference to a local commit, allowing submodules to live in the same\n> repo as their parent and share some of the advantages you describe for\n> sub-tree.\n\nI think it would be better if we could abandon .gitmodules entirely;\nit's really only useful for listing repository URLs, and listing\nrepository URLs is a major part of the problem.\n\nSomething that would be neat, and at least vaguely backward-compatible\nwould be to simply *try* fetching the linked commit objects from a\nremote repo, and checking them out from the local repo.  If the\nobjects exists, fetch/checkout of them will just work; if they don't,\nthen it can (for backwards compatibility) revert to the current\nbehaviour.  Push would, if the objects exist, send them to the remote\nrepo.\n\nThen there could be a .gitconfig option that flips this new behaviour\non and off, ie. auto-checkouts subprojects that *can* be checked out\nwithout any extra knowledge, or not.  If not, then you have to use the\nold-style git submodule stuff.\n\n(This proposal is not as easy as it sounds; to do it *right* would\ninvolve not having a separate .git repo for each subproject.  That\nmeans changes to the index file format and a bunch of related stuff.\nThough I guess you could keep the sub-repo stuff and it would still be\nbetter than what we have now.)\n\nHave fun,\n\nAvery\n"},{"id":"146076","messageId":"4C49B0E9.1090300@web.de","threadId":"24457","inReplyTo":"AANLkTimOb2VjYI21wQsC64lm4HsVPwpRWd1twIUBnbJ3@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Jens Lehmann","fromEmail":"jens.lehmann@web.de","sentAt":"2010-07-23T15:10:33Z","receivedAt":"2010-07-23T15:10:33Z","isPatch":false,"sender":{"key":"jens.lehmann@web.de","avatar":"https://avatars.githubusercontent.com/u/135220?v=4"},"body":"Am 22.07.2010 21:41, schrieb Avery Pennarun:\n> I create an app named myapp on github:\n> \n>    git://github.com/apenwarr/myapp\n> \n> It uses 17 different ruby gems, which I import as subprojects.  I have\n> two choices:\n> \n> [1] .gitmodules can use absolute paths to the original gem locations:\n> \n>    git://github.com/rubygems/gem[1..n]\n> \n> [2] Or else I can fork them all and use relative paths in .gitmodules:\n> \n>    ../gem[1..n]\n>    translates to --> git://github.com/apenwarr/gem[1..n]\n\nYou forgot what we do as best practice at work:\n\n[3] Fork the gem repos on github (or another server reachable by your\n    co-workers) and use those, so you don't have to change the URL\n    later:\n\n    git://github.com/apenwarrrubygems/gem[1..n]\n\nYour problems go away, setup has to be done only once on project\nstart and not for every developer, you can use your own branchnames\nand you have a staging repo from where you can push patches upstream\nif necessary.\n\n\n> Surely including *repository URLs* inside the *repository content* is\n> at least as bad as including branch names.  If we're going to do one,\n> we might as well do the other.  But it won't help, because the stored\n> branch name will probably be 'master', and my personal hacked-up copy\n> of gem13 shouldn't be on a branch named master anyway.\n\nYou sure are aware that having a branch name associated with a\nsubmodule checkout is a request repeatedly made?\n\n\n> The *real* need for submodules - the need you can't fix without\n> submodules - has nothing to do with these requirements.  It's about\n> each submodule wanting to have its own lifecycle, owner, changelog,\n> and release process, and - perhaps this is actually the killer\n> requirement - each supermodule wanting to be able to cleanly rewind a\n> submodule if they don't like the new version.\n\nThat is just one example. Another one is code shared between\ndifferent repos (think: libraries) where you want to make sure that\na bugfix in the library made in project A will make it to the shared\ncode repo and thus doesn't have to be fixed again by projects B to X.\nThis was one of the reasons we preferred submodules over subtrees\nin our evaluation, because there is no incentive to push fixes inside\nthe subtree back to its own repo like there is when using submodules.\n\n\n>>> It's very clear that git-submodule's current behaviour totally\n>>> mismatches the entire git philosophy.  That's why it's so impossible\n>>> to make the git-submodule command usable.\n>>\n>> That's very strong accusation.\n> \n> Agreed... but that doesn't make it wrong :)\n\nBut calling a feature \"impossible to make ... usable\" is an\ninteresting thing to say about a feature lots of people are\nusing productively in their daily work, no? ;-)\n\n\n>> rebase and merge needs separate    | rebase and merge works normally\n>> work in submodule currently        |\n> \n> True.\n\nNope, there is a patch in pu doing\nthat when it is a simple fast forward\nand giving you advice when both sides\nare already merged inside the submodule\n(CCed Heiko, because he is the author\nof that feature)\n\nIt is the /commits/ that have to be\ndone twice, once in the submodule and\nthen in the superproject. (But that is\nnot necessarily bad, imagine having git\ngui as a submodule: you would be\nautomagically reminded that stuff for\ngit gui should be sent somewhere else\nthan to Junio).\n"},{"id":"146077","messageId":"4C49B103.2010002@web.de","threadId":"24457","inReplyTo":"AANLkTil2ZHsV-PH46wfeqqvB27akRw3egDWRtJbIPLXb@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Jens Lehmann","fromEmail":"jens.lehmann@web.de","sentAt":"2010-07-23T15:10:59Z","receivedAt":"2010-07-23T15:10:59Z","isPatch":false,"sender":{"key":"jens.lehmann@web.de","avatar":"https://avatars.githubusercontent.com/u/135220?v=4"},"body":"Am 22.07.2010 23:33, schrieb Avery Pennarun:\n> On Thu, Jul 22, 2010 at 4:17 PM, Ævar Arnfjörð Bjarmason\n> <avarab@gmail.com> wrote:\n>> But it's also clear that we have a lot of tribal knowledge about the\n>> lackings of git submodule / git subtree. It would be *really* useful\n>> if people like Avery and Jens which have obviously thought hard about\n>> the submodule/subtree issues would draft up some (calmly written) docs\n>> about how the two differ (with comparison tables etc.).\n>>\n>> That'd be a very helpful resource for Git users in deciding which one\n>> to use.\n> \n> I think I'm too biased to write that, but if someone else wants to\n> take the lead, I could certainly contribute.\n\nWhile I don't consider myself biased, I just don't know enough about\nthe details of the subtree approach to write that.\n\nBut I would certainly contribute to the submodule side of such a\ndocument too.\n"},{"id":"146078","messageId":"4C49B10A.8060900@web.de","threadId":"24457","inReplyTo":"AANLkTinBtXmMei2Q6MZrXWxa3t+_quGdzpcq46EZvgvG@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Jens Lehmann","fromEmail":"jens.lehmann@web.de","sentAt":"2010-07-23T15:11:06Z","receivedAt":"2010-07-23T15:11:06Z","isPatch":false,"sender":{"key":"jens.lehmann@web.de","avatar":"https://avatars.githubusercontent.com/u/135220?v=4"},"body":"Am 23.07.2010 10:40, schrieb Avery Pennarun:\n> I think it would be better if we could abandon .gitmodules entirely;\n> it's really only useful for listing repository URLs, and listing\n> repository URLs is a major part of the problem.\n\nThen where do you get the URL to clone the submodule from on \"git\nclone --recursive\"?\n"},{"id":"146079","messageId":"4C49B195.2090101@web.de","threadId":"24457","inReplyTo":"AANLkTinBtXmMei2Q6MZrXWxa3t+_quGdzpcq46EZvgvG@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Jens Lehmann","fromEmail":"jens.lehmann@web.de","sentAt":"2010-07-23T15:13:25Z","receivedAt":"2010-07-23T15:13:25Z","isPatch":false,"sender":{"key":"jens.lehmann@web.de","avatar":"https://avatars.githubusercontent.com/u/135220?v=4"},"body":"Am 23.07.2010 10:40, schrieb Avery Pennarun:\n> I think it would be better if we could abandon .gitmodules entirely;\n> it's really only useful for listing repository URLs, and listing\n> repository URLs is a major part of the problem.\n\nThen where do you get the URL to clone the submodule from on \"git\nclone --recursive\"?\n"},{"id":"146081","messageId":"4C49B31F.8000102@xiplink.com","threadId":"24457","inReplyTo":"AANLkTimOb2VjYI21wQsC64lm4HsVPwpRWd1twIUBnbJ3@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Marc Branchaud","fromEmail":"marcnarc@xiplink.com","sentAt":"2010-07-23T15:19:59Z","receivedAt":"2010-07-23T15:19:59Z","isPatch":false,"sender":{"key":"marcnarc@xiplink.com","avatar":"https://avatars.githubusercontent.com/u/14980203?v=4"},"body":"On 10-07-22 03:41 PM, Avery Pennarun wrote:\n> \n> 1) Sometimes I want to clone only some subdirs of a project\n> 2) Sometimes I don't want the entire history because it's too big.\n> 3) Super huge git repositories start to degrade in performance.\n\nThe reason we turned to submodules is precisely to deal with repository size.\n Our code base encompasses the entire FreeBSD tree plus different versions of\nthe Linux kernel, along with various third-party libraries & apps.  You don't\nneed everything to build a given product (a FreeBSD product doesn't use any\nLinux kernels, for example) but because all the products share common code we\nneed to be able to branch and tag the common code along with the uncommon code.\n\nSo a straight \"git clone\" that would need to fetch all of FreeBSD plus 4\ndifferent Linux kernels and check all that out is a major problem, especially\nfor our automated build system (which could definitely be implemented better,\nbut still).  In truth it's the checkout that takes the most time by far,\nthough commands like git-status also take inconveniently long.\n\nWe chose git-submodule over git-subtree mainly because git-submodule lets us\nselectively checkout different parts of our code.  (AFAIK sparse checkouts\naren't yet an option.)  We didn't really consider git-subtree because it's\nnot an official part of git, and we didn't want to have to teach (and nag)\nall our developers to install and maintain it in addition to keeping up with\ngit itself.  Besides, git-submodule's collection-of-independent-repos model\nworks fairly well in our situation, though the implementation could\ndefinitely be improved (and Jens's list is a really good start).\n\nNeither submodule nor subtree really solves our situation, but right now\ngit-submodule is the only thing \"official\" git offers to manage\nloosely-coupled code.  It would be nice to see git-submodule added to the\ntoolkit, but it would be even nicer if git had better ways to deal with\n\"vast\" repositories.\n\nAnother tool folks should keep in mind in this discussion is 'repo' which\nGoogle built for the Android project.  Android's code base is also too vast\nto work well in a single git repository, and I don't think subtrees or\nsubmodules would be a good match for them either.\n\n\t\tM.\n"},{"id":"146083","messageId":"4C49BDB7.3080805@gmail.com","threadId":"24457","inReplyTo":"4C49B0E9.1090300@web.de","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Bryan Larsen","fromEmail":"bryan.larsen@gmail.com","sentAt":"2010-07-23T16:05:11Z","receivedAt":"2010-07-23T16:05:11Z","isPatch":false,"sender":{"key":"bryan@larsen.st","avatar":"https://avatars.githubusercontent.com/u/32073?v=4"},"body":"On 10-07-23 11:10 AM, Jens Lehmann wrote:\n> Am 22.07.2010 21:41, schrieb Avery Pennarun:\n>> I create an app named myapp on github:\n>>\n>>     git://github.com/apenwarr/myapp\n>>\n>> It uses 17 different ruby gems, which I import as subprojects.  I have\n>> two choices:\n>>\n>> [1] .gitmodules can use absolute paths to the original gem locations:\n>>\n>>     git://github.com/rubygems/gem[1..n]\n>>\n>> [2] Or else I can fork them all and use relative paths in .gitmodules:\n>>\n>>     ../gem[1..n]\n>>     translates to -->  git://github.com/apenwarr/gem[1..n]\n>\n> You forgot what we do as best practice at work:\n>\n> [3] Fork the gem repos on github (or another server reachable by your\n>      co-workers) and use those, so you don't have to change the URL\n>      later:\n>\n>      git://github.com/apenwarrrubygems/gem[1..n]\n>\n> Your problems go away, setup has to be done only once on project\n> start and not for every developer, you can use your own branchnames\n> and you have a staging repo from where you can push patches upstream\n> if necessary.\n\nWhat's best practice for open source projects?   I do this, but nobody \nexcept my coworkers can push to my forks, so it's a huge rigamarole just \nto get a fix into a submodule.\n\n>\n> That is just one example. Another one is code shared between\n> different repos (think: libraries) where you want to make sure that\n> a bugfix in the library made in project A will make it to the shared\n> code repo and thus doesn't have to be fixed again by projects B to X.\n> This was one of the reasons we preferred submodules over subtrees\n> in our evaluation, because there is no incentive to push fixes inside\n> the subtree back to its own repo like there is when using submodules.\n\nBut you stated above that each project has its own fork of the library. \n   So there's no special incentive to push changes from the fork back to \nits master repo.\n\n>\n>\n>>>> It's very clear that git-submodule's current behaviour totally\n>>>> mismatches the entire git philosophy.  That's why it's so impossible\n>>>> to make the git-submodule command usable.\n>>>\n>>> That's very strong accusation.\n>>\n>> Agreed... but that doesn't make it wrong :)\n>\n> But calling a feature \"impossible to make ... usable\" is an\n> interesting thing to say about a feature lots of people are\n> using productively in their daily work, no? ;-)\n\nIn my experience, it's possible to make it usable if and only if:\n\n1.  you have a small team\n2.  all of whom are very comfortable with git\n3.  changes inside submodules are either infrequent or only happen in a \nsingle direction\n4.  the project is not public/open source\n\nI think #4 is the killer reason why submodules don't work.  It works \nfine if the submodule is fairly independent, but if you have a patch to \nthe submodule that was created for and in the context of the \nsuperproject, things get really annoying really quickly.\n\nBryan\n"},{"id":"146092","messageId":"4C49CD49.4010101@web.de","threadId":"24457","inReplyTo":"4C49BDB7.3080805@gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Jens Lehmann","fromEmail":"jens.lehmann@web.de","sentAt":"2010-07-23T17:11:37Z","receivedAt":"2010-07-23T17:11:37Z","isPatch":false,"sender":{"key":"jens.lehmann@web.de","avatar":"https://avatars.githubusercontent.com/u/135220?v=4"},"body":"Am 23.07.2010 18:05, schrieb Bryan Larsen:\n> On 10-07-23 11:10 AM, Jens Lehmann wrote:\n>> That is just one example. Another one is code shared between\n>> different repos (think: libraries) where you want to make sure that\n>> a bugfix in the library made in project A will make it to the shared\n>> code repo and thus doesn't have to be fixed again by projects B to X.\n>> This was one of the reasons we preferred submodules over subtrees\n>> in our evaluation, because there is no incentive to push fixes inside\n>> the subtree back to its own repo like there is when using submodules.\n> \n> But you stated above that each project has its own fork of the library.   So there's no special incentive to push changes from the fork back to its master repo.\n\nWhen you are not working on your own, it is preferable to be able to\nget changes upstream into a submodules repo to share them.\nSo if you can do that (either via push or patches sent by email or\nwhatever), then use it's URL directly (and then you have the incentive\nthat fixes get pushed, which is nice).\nOr you can't, then use a fork reachable by the people you work with\n(then you still can see all fixes made by your group in the forked\nrepo and can decide to push them upstream). Then pushing fixes back\nto the original repo is a matter of courtesy, as it is with every\nother work flow I know.\nAnd I think that is just the same thing we all do with plain git\nrepos when working with others: If you can push, you use it directly\nto clone from, if you can't, you fork it.\n\n\n> In my experience, it's possible to make it usable if and only if:\n> \n> 1.  you have a small team\n> 2.  all of whom are very comfortable with git\n> 3.  changes inside submodules are either infrequent or only happen in a single direction\n> 4.  the project is not public/open source\n>\n> I think #4 is the killer reason why submodules don't work.  It works fine if the submodule is fairly independent, but if you have a patch to the submodule that was created for and in the context of the superproject, things get really annoying really quickly.\n\nWhat is the problem with the \"forked repo\" solution for #4?\n"},{"id":"146104","messageId":"4C49E720.20207@gmail.com","threadId":"24457","inReplyTo":"4C49CD49.4010101@web.de","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Bryan Larsen","fromEmail":"bryan.larsen@gmail.com","sentAt":"2010-07-23T19:01:52Z","receivedAt":"2010-07-23T19:01:52Z","isPatch":false,"sender":{"key":"bryan@larsen.st","avatar":"https://avatars.githubusercontent.com/u/32073?v=4"},"body":"On 10-07-23 01:11 PM, Jens Lehmann wrote:\n> Am 23.07.2010 18:05, schrieb Bryan Larsen:\n>> On 10-07-23 11:10 AM, Jens Lehmann wrote:\n>>> That is just one example. Another one is code shared between\n>>> different repos (think: libraries) where you want to make sure that\n>>> a bugfix in the library made in project A will make it to the shared\n>>> code repo and thus doesn't have to be fixed again by projects B to X.\n>>> This was one of the reasons we preferred submodules over subtrees\n>>> in our evaluation, because there is no incentive to push fixes inside\n>>> the subtree back to its own repo like there is when using submodules.\n>>\n>> But you stated above that each project has its own fork of the library.   So there's no special incentive to push changes from the fork back to its master repo.\n>\n> When you are not working on your own, it is preferable to be able to\n> get changes upstream into a submodules repo to share them.\n> So if you can do that (either via push or patches sent by email or\n> whatever), then use it's URL directly (and then you have the incentive\n> that fixes get pushed, which is nice).\n> Or you can't, then use a fork reachable by the people you work with\n> (then you still can see all fixes made by your group in the forked\n> repo and can decide to push them upstream). Then pushing fixes back\n> to the original repo is a matter of courtesy, as it is with every\n> other work flow I know.\n> And I think that is just the same thing we all do with plain git\n> repos when working with others: If you can push, you use it directly\n> to clone from, if you can't, you fork it.\n\nSo basically you're saying: sometimes you can use a non-forked \nrepository, which has a whole bunch of disadvantages, but has the minor \nadvantage that you're \"forced\" to push your changes upstream.\n\nWhich I see as a disadvantage because that means you're pushing untested \nchanges.\n\nOr else you use a forked repo, which is basically the same as using \ngit-subtree, except for a lot of additional admin hassle.\n\n>\n>\n>> In my experience, it's possible to make it usable if and only if:\n>>\n>> 1.  you have a small team\n>> 2.  all of whom are very comfortable with git\n>> 3.  changes inside submodules are either infrequent or only happen in a single direction\n>> 4.  the project is not public/open source\n>>\n>> I think #4 is the killer reason why submodules don't work.  It works fine if the submodule is fairly independent, but if you have a patch to the submodule that was created for and in the context of the superproject, things get really annoying really quickly.\n>\n> What is the problem with the \"forked repo\" solution for #4?\n>\n\nPlease tell me how I can set up a public project on github where project \nA contains module X, so that Joe Average User can clone A, make a change \nin the module X and send a simple pull request to get that change into \nA.   The change is one that's inappropriate to push upstream to X \nwithout additional work, but is appropriate for A at this point in time. \n  Joe's a beginning git user.\n\nThat's actually a simple use case compared to others I've run into.\n\nBryan\n"},{"id":"146117","messageId":"AANLkTimSoe9iqu4cJCH1d4rVsWHpFn3+8pbrCxsnVM1D@mail.gmail.com","threadId":"24457","inReplyTo":"4C49B0E9.1090300@web.de","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-23T22:32:18Z","receivedAt":"2010-07-23T22:32:18Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Fri, Jul 23, 2010 at 11:10 AM, Jens Lehmann <Jens.Lehmann@web.de> wrote:\n> You forgot what we do as best practice at work:\n>\n> [3] Fork the gem repos on github (or another server reachable by your\n>    co-workers) and use those, so you don't have to change the URL\n>    later:\n>\n>    git://github.com/apenwarrrubygems/gem[1..n]\n>\n> Your problems go away, setup has to be done only once on project\n> start and not for every developer, you can use your own branchnames\n> and you have a staging repo from where you can push patches upstream\n> if necessary.\n\nNow all your fellow developers have to push their submodule code to a\nsingle upstream repo?  That's rather centralized and un-git-like.\n\nFor the rest, Brian Larsen answered this one well, and I agree with him.\n\n>> Surely including *repository URLs* inside the *repository content* is\n>> at least as bad as including branch names.  If we're going to do one,\n>> we might as well do the other.  But it won't help, because the stored\n>> branch name will probably be 'master', and my personal hacked-up copy\n>> of gem13 shouldn't be on a branch named master anyway.\n>\n> You sure are aware that having a branch name associated with a\n> submodule checkout is a request repeatedly made?\n\nOf course it is; I requested it myself.  Then, two years later after\nthinking about the problem a lot and writing git-subtree out of\nfrustration, I realized that even if this feature existed, it wouldn't\nhelp at all.\n\nIf you use git-submodule, you must push your submodule commits\nseparately or the supermodule is broken for everybody but you.  To\npush a submodule, you need a) an upstream to push to and b) a branch\nname.  It's easy to forget to create a branch name, so of course\npeople request that feature.\n\nHowever, the real problem is \"you must push your submodule commits\nseparately.\"  Fix that, and I can guarantee that the request for\nsubmodule branch naming will disappear.\n\n> That is just one example. Another one is code shared between\n> different repos (think: libraries) where you want to make sure that\n> a bugfix in the library made in project A will make it to the shared\n> code repo and thus doesn't have to be fixed again by projects B to X.\n> This was one of the reasons we preferred submodules over subtrees\n> in our evaluation, because there is no incentive to push fixes inside\n> the subtree back to its own repo like there is when using submodules.\n\nI think you'd like svn; it's pretty cool.  All changes made to a\nproject need to get pushed to a central upstream repo so you never\nforget to share them.\n\n>>> rebase and merge needs separate    | rebase and merge works normally\n>>> work in submodule currently        |\n>>\n>> True.\n>\n> Nope, there is a patch in pu doing\n> that when it is a simple fast forward\n> and giving you advice when both sides\n> are already merged inside the submodule\n> (CCed Heiko, because he is the author\n> of that feature)\n\nFast forwards are not merges, and pu is not now.\n\n> It is the /commits/ that have to be\n> done twice, once in the submodule and\n> then in the superproject. (But that is\n> not necessarily bad, imagine having git\n> gui as a submodule: you would be\n> automagically reminded that stuff for\n> git gui should be sent somewhere else\n> than to Junio).\n\nYup, I agree that requiring a separate commit to the submodule repo is\nnot a bad idea.  I always do this anyway even when using git-subtree,\nbecause I'm thinking ahead to the day when I'll push my submodule\nchanges upstream and I want my commit message to make sense.  But\nthat's because I think ahead like that.  Having the tool force me to\ndo it would be harmless and help people avoid mistakes.\n\nThe syntax for it ought to be nice though.  I should be able to do:\n\n    git commit -- path/to/submodule\n\nAnd have it commit everything in the submodule tree as a new commit in\nthe submodule.  I don't want to have to think about cd'ing to\npath/to/submodule just so I can commit the files I changed in there.\n\nHave fun,\n\nAvery\n"},{"id":"146118","messageId":"AANLkTinb8YqBesa5Uh7+uuKAF0BhNjAGmvai5bNO3hV4@mail.gmail.com","threadId":"24457","inReplyTo":"4C49B10A.8060900@web.de","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-23T22:33:31Z","receivedAt":"2010-07-23T22:33:31Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Fri, Jul 23, 2010 at 11:11 AM, Jens Lehmann <Jens.Lehmann@web.de> wrote:\n> Am 23.07.2010 10:40, schrieb Avery Pennarun:\n>> I think it would be better if we could abandon .gitmodules entirely;\n>> it's really only useful for listing repository URLs, and listing\n>> repository URLs is a major part of the problem.\n>\n> Then where do you get the URL to clone the submodule from on \"git\n> clone --recursive\"?\n\nIf you're asking that question, you're missing my point entirely.  In\nmy proposed model, the submodule objects are all in the same repo as\nthe superproject, so there *is* no separate URL.  And thus there is no\nmore need for .gitmodules.\n\nHave fun,\n\nAvery\n"},{"id":"146120","messageId":"AANLkTi=LHYDhY=424YZpO3yGqGGsxpY2Sj8=ULNKvAQX@mail.gmail.com","threadId":"24457","inReplyTo":"4C49B31F.8000102@xiplink.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-23T22:50:49Z","receivedAt":"2010-07-23T22:50:49Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Fri, Jul 23, 2010 at 11:19 AM, Marc Branchaud <marcnarc@xiplink.com> wrote:\n> On 10-07-22 03:41 PM, Avery Pennarun wrote:\n>> 1) Sometimes I want to clone only some subdirs of a project\n>> 2) Sometimes I don't want the entire history because it's too big.\n>> 3) Super huge git repositories start to degrade in performance.\n>\n> The reason we turned to submodules is precisely to deal with repository size.\n\nI believe that's very common.\n\nHowever, I wonder whether that's actually a good reason for git to\ndevelop better submodules, or actually just a good reason for git to\nget better support for handling huge repositories.\n\nMy bup project (http://github.com/apenwarr/bup) is all about huge\nrepositories.  It handles repositories with hundreds of gigabytes, and\ntrees containing millions of files (entire filesystems), quite nicely.\n Of course, it's not a version control system, so it won't solve your\nproblems.  It's just evidence that large repositories are actually\nquite manageable without changing the fundamentals of git.\n\n>  Our code base encompasses the entire FreeBSD tree plus different versions of\n> the Linux kernel, along with various third-party libraries & apps.  You don't\n> need everything to build a given product (a FreeBSD product doesn't use any\n> Linux kernels, for example) but because all the products share common code we\n> need to be able to branch and tag the common code along with the uncommon code.\n\nHonest question: do you care about the wasted disk space and download\ntime for these extra files?  Or just the fact that git gets slow when\nyou have them?\n\nHow people answer that question very much affects the way git should\nbe designed.\n\n> So a straight \"git clone\" that would need to fetch all of FreeBSD plus 4\n> different Linux kernels and check all that out is a major problem, especially\n> for our automated build system (which could definitely be implemented better,\n> but still).\n\nTo be absolutely pedantic, the four linux kernels likely share most of\ntheir objects and so you're only paying the cost (at least during\nfetch) of including it once :)\n\n(If you're actually using git-submodule and each copy of the kernel is\nits own module, then it might be cloning the kernel four times\nseparately, in which case the objects *don't* get shared, so this ends\nup being much more expensive than it should be.  That could be fixed\nby slightly improving git-submodule to share some objects rather than\nrearchitecting it though.)\n\n> In truth it's the checkout that takes the most time by far,\n> though commands like git-status also take inconveniently long.\n\nYeah, git could stand to be optimized a bit here.  And since Windows\nstats files about 10x slower than Linux, this problem occurs about 10x\nsooner on Windows, which makes using git on Windows (which sadly I\nhave to do sometimes) extremely painful compared to Linux.\n\nIMHO, the correct answer here is to have an inotify-based daemon prod\nat the .git/index automatically when files get updated, so that git\nitself doesn't have to stat/readdir through the entire tree in order\nto do any of its operations.  (Windows also has something like inotify\nthat would work.)  If you had this, then git\nstatus/diff/checkout/commit would be just as fast with zillions of\nfiles as with 10 files.  Sooner or later, if nobody implements this, I\npromise I'll get around to it since inotify is actually easy to code\nfor :)\n\nAlso note that the only reason submodules are faster here is that\nthey're ignoring possibly important changes.  Notably, when you do\n'git status' from the top level, it won't warn you if you have any\nnot-yet-committed files in any of your submodules.  Personally, I\nconsider that to be really important information, but to obtain it\nwould make 'git status' take just as long as without submodules, so\nyou wouldn't get any benefit.  (I think nowadays there's a way to get\nthis recursive status information if you want it, but it'll be slow of\ncourse.)\n\n> We chose git-submodule over git-subtree mainly because git-submodule lets us\n> selectively checkout different parts of our code.  (AFAIK sparse checkouts\n> aren't yet an option.)\n\nFair enough.  If you could confirm or deny my theory that this is\n*entirely* a performance related concern (as opposed to disk space /\ndownload time), that would be helpful.\n\n> We didn't really consider git-subtree because it's\n> not an official part of git, and we didn't want to have to teach (and nag)\n> all our developers to install and maintain it in addition to keeping up with\n> git itself.\n\nArguably, this is a vote for including git-subtree into the core\n(which was Bryan's point when he started this thread); it obviously is\nbeing rejected sometimes by git users simply because it's not in the\ncore, even though it could help them.\n\nHave fun,\n\nAvery\n"},{"id":"146127","messageId":"AANLkTinhd2DYh7WXzMvhMkqp98fYtTWWuQi0RSL9Rome@mail.gmail.com","threadId":"24457","inReplyTo":"AANLkTi=LHYDhY=424YZpO3yGqGGsxpY2Sj8=ULNKvAQX@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"","fromEmail":"skillzero@gmail.com","sentAt":"2010-07-24T00:58:31Z","receivedAt":"2010-07-24T00:58:31Z","isPatch":false,"sender":{"key":"skillzero@gmail.com","avatar":null},"body":"On Fri, Jul 23, 2010 at 3:50 PM, Avery Pennarun <apenwarr@gmail.com> wrote:\n\n> Honest question: do you care about the wasted disk space and download\n> time for these extra files?  Or just the fact that git gets slow when\n> you have them?\n\nI have the similar situation to the original poster (huge trees) and\nfor me it's all three: disk space, download time, and performance. My\ntree has a few relatively small (< 20 MB) shared directories of common\ncode, a few large (2-6 GB) directories of code for OS's, and then\nseveral medium size (< 500 MB) directories for application code. The\napplication developers only care about the app+shared directories (and\nare very annoyed by the massive space and performance impact of the OS\ndirectories). The firmware-only developers only care about OS+shared\nand are mildly annoyed by the medium space and performance impact of\nthe app directories. I work on all of the pieces, but even I would\nprefer to have things separated so when I work on the apps, git\nstatus/etc doesn't take a big hit for close to a million files in the\nOS directories (particularly when doing git status on Windows). Even\nwhen using the -uno option to git status, it's still pretty slow (over\na minute).\n\ngit-submodule might be technically possible in this situation, but\nhaving to commit and push each submodule and then commit and push the\nsuper module makes it slightly worse than just dealing with the\nspace/download/performance issues of one huge repository.\n\ngit-subtree could also possibly help, but there's still extra work to\nsplit and merge each repository. And I'm not sure how it handles\ncommit IDs across the repositories because I want to be able to say \"I\nfixed that bug in shared/code.c in commit abc123\" and have both the\nOS+shared and the apps+shared people be able git log abc123 and see\nthe same change (and merge/cherry-pick/etc.).\n\nI think what I want is a way to do a sparse checkout where some sort\nof module is maintained in the git repository (probably just an\nINI-style file with paths) so I can clone directly from the server and\nit figures out the objects I need for the full history of only\napps+shared (or firmware+shared, etc.) on the server side and only\nsends those objects. I still want to be able to branch, tag, and refer\nto commit IDs. So I only take the space/download/performance hit of\ndirectories included in the module, but I don't have to manually\nmaintain that view of the repository (as I do with git-submodule and\ngit-subtree).\n\nThe closest thing to that so far for me has been the sparse checkout\nsupport added in git 1.7 combined with a convenience script I wrote.\nEveryone still has a huge download and .git directory, but at least\nthe working copy is limited to the paths specified in the module so\ngit status isn't super slow (although just having all those objects in\nthe .git directory still slows it down quite a bit).\n"},{"id":"146133","messageId":"AANLkTimLayG_HFxGdq+Tt8hU_MApBpSdHHiYPxcakpRJ@mail.gmail.com","threadId":"24457","inReplyTo":"AANLkTinhd2DYh7WXzMvhMkqp98fYtTWWuQi0RSL9Rome@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-24T01:20:07Z","receivedAt":"2010-07-24T01:20:07Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Fri, Jul 23, 2010 at 8:58 PM,  <skillzero@gmail.com> wrote:\n> On Fri, Jul 23, 2010 at 3:50 PM, Avery Pennarun <apenwarr@gmail.com> wrote:\n>> Honest question: do you care about the wasted disk space and download\n>> time for these extra files?  Or just the fact that git gets slow when\n>> you have them?\n>\n> I have the similar situation to the original poster (huge trees) and\n> for me it's all three: disk space, download time, and performance. My\n> tree has a few relatively small (< 20 MB) shared directories of common\n> code, a few large (2-6 GB) directories of code for OS's, and then\n> several medium size (< 500 MB) directories for application code. The\n> application developers only care about the app+shared directories (and\n> are very annoyed by the massive space and performance impact of the OS\n> directories).\n\nGiven how cheap disk space is nowadays, I'm curious about this.  Are\nthey really just annoyed by the performance problem, and they complain\nabout the extra size because they blame the performance on the extra\nfiles?  Or are they honestly short of disk space?\n\nSimilarly, are all your developers located at the same office?  If so,\nthen bandwidth ought not be an issue.\n\nI'm pushing extra hard on this because I believe there are lots of\nopportunities to just improve git performance on huge repositories.\nAnd if the only *real* reason people need to split repositories is\nthat performance goes down, then that's fixable, and you may need\nneither git-submodule nor git-subtree.\n\n> I work on all of the pieces, but even I would\n> prefer to have things separated so when I work on the apps, git\n> status/etc doesn't take a big hit for close to a million files in the\n> OS directories (particularly when doing git status on Windows). Even\n> when using the -uno option to git status, it's still pretty slow (over\n> a minute).\n\nThis is indeed a problem with large repositories.  Of course,\nsplitting them with git-submodule is kind of cheating, because it just\nmakes git-status *not look* to see if those files are dirty or not.\nIf they are dirty and you forget to commit them, you'll never know\nuntil someone tells you later.  It would be functionally equivalent to\njust have git-status not look inside certain subdirs of a single\nrepository.\n\nIn any case, this is a pretty clear optimization target (especially\nsince Windows is so amazingly slow at statting files): just have a\ndaemon running inotify (or the Windows equivalent) that tracks whether\nfiles are up-to-date or not.  Then git would never need to recurse\nthrough the entire tree, and operations like status, diff, checkout,\nand commit could be fast even with a million-file repository.\n\n> git-subtree could also possibly help, but there's still extra work to\n> split and merge each repository. And I'm not sure how it handles\n> commit IDs across the repositories because I want to be able to say \"I\n> fixed that bug in shared/code.c in commit abc123\" and have both the\n> OS+shared and the apps+shared people be able git log abc123 and see\n> the same change (and merge/cherry-pick/etc.).\n\ngit-subtree (if you don't use --squash) keeps all the commit IDs.  It\nis extra work to split and merge between repositories, though.  It\ndoesn't solve your repository-is-too-large problem.\n\n> I think what I want is a way to do a sparse checkout where some sort\n> of module is maintained in the git repository (probably just an\n> INI-style file with paths) so I can clone directly from the server and\n> it figures out the objects I need for the full history of only\n> apps+shared (or firmware+shared, etc.) on the server side and only\n> sends those objects. I still want to be able to branch, tag, and refer\n> to commit IDs. So I only take the space/download/performance hit of\n> directories included in the module, but I don't have to manually\n> maintain that view of the repository (as I do with git-submodule and\n> git-subtree).\n\nYes, better sparse checkout and sparse fetch would be very valuable\nhere and would eliminate a lot of the reasons people have for misusing\nsubmodules.\n\n> (although just having all those objects in\n> the .git directory still slows it down quite a bit).\n\nYou're the second person who has mentioned this today (the first one\nwas to me in a private email).  I'd like to understand this better.\n\nIn my bup project (http://github.com/apenwarr/bup) we regularly create\ngit repositories with hundreds of gigabytes of packs, comprising tens\nor hundreds of millions of objects, and the repository doesn't get\nslow.  (Obviously this is a separate issue from having a huge work\ntree with a million files in it.)  In repositories this thoroughly\nhuge, we did find a way to improve memory usage versus git's pack .idx\nfiles (bup has '.midx' files that combine multiple indexes into one,\nthus reducing the binary search steps).  But this only matters when\nyou get well over 10 gigabytes of stuff and you're wading through it\nusing crappy python code (as bup does) and frequently inserting a\nmillion objects at a time (as bup does).  The git usage pattern is\nmuch simpler and therefore faster.\n\nHow big is your .git directory and what performance problems do you\nsee?  I assume you've done 'git gc' to clean up all the loose objects,\nright?\n\nHave fun,\n\nAvery\n"},{"id":"146202","messageId":"AANLkTikx5EtQ0yvdkqN1Q1QAudFZfbd+_jpoa9ztLrz1@mail.gmail.com","threadId":"24457","inReplyTo":"AANLkTimLayG_HFxGdq+Tt8hU_MApBpSdHHiYPxcakpRJ@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"","fromEmail":"skillzero@gmail.com","sentAt":"2010-07-24T19:40:11Z","receivedAt":"2010-07-24T19:40:11Z","isPatch":false,"sender":{"key":"skillzero@gmail.com","avatar":null},"body":"On Fri, Jul 23, 2010 at 6:20 PM, Avery Pennarun <apenwarr@gmail.com> wrote:\n> On Fri, Jul 23, 2010 at 8:58 PM,  <skillzero@gmail.com> wrote:\n>> On Fri, Jul 23, 2010 at 3:50 PM, Avery Pennarun <apenwarr@gmail.com> wrote:\n>>> Honest question: do you care about the wasted disk space and download\n>>> time for these extra files?  Or just the fact that git gets slow when\n>>> you have them?\n>>\n>> I have the similar situation to the original poster (huge trees) and\n>> for me it's all three: disk space, download time, and performance. My\n>> tree has a few relatively small (< 20 MB) shared directories of common\n>> code, a few large (2-6 GB) directories of code for OS's, and then\n>> several medium size (< 500 MB) directories for application code. The\n>> application developers only care about the app+shared directories (and\n>> are very annoyed by the massive space and performance impact of the OS\n>> directories).\n>\n> Given how cheap disk space is nowadays, I'm curious about this.  Are\n> they really just annoyed by the performance problem, and they complain\n> about the extra size because they blame the performance on the extra\n> files?  Or are they honestly short of disk space?\n\nI think it's both space and performance. When you're using SSD drives,\nstorage still pretty expensive. A 128 GB or less SSD is pretty common\nin a laptop so you can run out pretty quick, especially when you're\nworking concurrently on a few different branches at the same time.\nIt's useful to keep multiple working copies (e.g. git-new-workdir)\nbecause rebuild time can be significant when switching branches.\n\n> Similarly, are all your developers located at the same office?  If so,\n> then bandwidth ought not be an issue.\n\nBandwidth isn't a big problem because you don't need to re-download\nthe repo very often. However, people work at home a lot where\nbandwidth is more limited. The biggest complaint I hear about\nbandwidth is that people tend to re-download when something goes wrong\n(i.e. inexperience with git resulting in a repository they can't\nrecover due to git resets, etc).\n\n> I'm pushing extra hard on this because I believe there are lots of\n> opportunities to just improve git performance on huge repositories.\n> And if the only *real* reason people need to split repositories is\n> that performance goes down, then that's fixable, and you may need\n> neither git-submodule nor git-subtree.\n\nPerformance degradation is my biggest complaint with large\nrepositories. Your inotify/FSEvents/etc daemon idea sounds interesting\nto deal with the stat issue.\n\n> This is indeed a problem with large repositories.  Of course,\n> splitting them with git-submodule is kind of cheating, because it just\n> makes git-status *not look* to see if those files are dirty or not.\n> If they are dirty and you forget to commit them, you'll never know\n> until someone tells you later.  It would be functionally equivalent to\n> just have git-status not look inside certain subdirs of a single\n> repository.\n\nI think it's only cheating if you're using all of the submodules. The\nmain purpose of submodules for me (although I don't currently use\nsubmodules) would be so I don't need to keep modules on disk that I\ndon't care about. If a developer is working on an app, they don't need\nthe OS directories/modules so they get much faster git status/etc and\nthere wouldn't be other directories to have dirty files in. That said,\nif I was using git submodule, I'd want git status to show me all the\nsubmodules that were checked out.\n\n>> (although just having all those objects in\n>> the .git directory still slows it down quite a bit).\n>\n> You're the second person who has mentioned this today (the first one\n> was to me in a private email).  I'd like to understand this better.\n\nWhat I'm basing this on is that even when I'm using a sparse checkout\nsuch that I have only a small subset of the files in my working\ndirectory, git status seems singifncantly slower for me than an\nequivalent git repository that only has that subset of files. That's\nnot very scientific, but that's what made me think just having a large\n.git directory with lots of objects/history slows down git status even\nif the working copy doesn't have a lot of files.\n\nI will try to experiment and see if I can narrow it down with some real numbers.\n\nBTW...what's the policy on CC'ing people on git mailing list replies?\nShould it be trimmed or not? I've received complaints in the past, but\nI was never really clear what the recommended policy is.\n"},{"id":"146205","messageId":"AANLkTi=Qp5CNCe=V2LCH2_EcTkxSpJ6+EHkk_BmUr9+B@mail.gmail.com","threadId":"24457","inReplyTo":"AANLkTi=LHYDhY=424YZpO3yGqGGsxpY2Sj8=ULNKvAQX@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Sverre Rabbelier","fromEmail":"srabbelier@gmail.com","sentAt":"2010-07-24T20:07:13Z","receivedAt":"2010-07-24T20:07:13Z","isPatch":false,"sender":{"key":"srabbelier@gmail.com","avatar":"https://avatars.githubusercontent.com/u/3098?v=4"},"body":"Heya,\n\nOn Fri, Jul 23, 2010 at 17:50, Avery Pennarun <apenwarr@gmail.com> wrote:\n> IMHO, the correct answer here is to have an inotify-based daemon prod\n> at the .git/index automatically when files get updated, so that git\n> itself doesn't have to stat/readdir through the entire tree in order\n> to do any of its operations.  (Windows also has something like inotify\n> that would work.)  If you had this, then git\n> status/diff/checkout/commit would be just as fast with zillions of\n> files as with 10 files.  Sooner or later, if nobody implements this, I\n> promise I'll get around to it since inotify is actually easy to code\n> for :)\n\nFrom what I've heard both SVN and Mercurial have something like that\nand it's incredible unstable and icky and nasty and bad and will eat\nyour babies. Then again, I don't have any experience with inotify, so\nif you say that it's all good and awesome, who am I to doubt that :).\n\n-- \nCheers,\n\nSverre Rabbelier\n"},{"id":"146218","messageId":"201007250036.55811.jnareb@gmail.com","threadId":"24457","inReplyTo":"4C488C9D.60606@gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2010-07-24T22:36:49Z","receivedAt":"2010-07-24T22:36:49Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Dnia czwartek 22. lipca 2010 20:23, Bryan Larsen napisał:\n> >\n> > Using git-subtree has its warts too: I don't think for example that there is\n> > a way to get a log _automatically excluding_ history subtree-merged\n> > subprojects.  Or is it there?\n> >\n> \n> It works exactly right for me when I used git-subtree in \"squashed\" \n> mode.  Changes which were done in tree show up separately in the log, \n> changes which were pulled in via git-subtree pull show up as a single \n> summary entry in the log.\n> \n> This discussion has been about how to improve git submodules, which is \n> sorely needed.   However, it's quite clear that git submodules will \n> never work as well as git subtrees in certain quite common situations. \n>   If fixed, git submodules will be more appropriate in other situations. \n>    However, I'm not asking to remove git submodules or prevent anybody \n> from fixing them, I'm just asking that git subtree be merged.\n> \n> Does anybody actually oppose the merger of git-subtree, which has (at \n> least) hundreds of users despite its out-of-tree status?\n\nI am very much *for* merging git-subtree into git core.  It is not that\nmuch different from e.g. \"git submodule\" or \"git remote\" porcelain\ncommands.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"146239","messageId":"AANLkTikEo=Qw56WCxkFdmGqQcQoiTsnBy+Dt6zHpkOii@mail.gmail.com","threadId":"24457","inReplyTo":"AANLkTikx5EtQ0yvdkqN1Q1QAudFZfbd+_jpoa9ztLrz1@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2010-07-25T01:47:47Z","receivedAt":"2010-07-25T01:47:47Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sun, Jul 25, 2010 at 5:40 AM,  <skillzero@gmail.com> wrote:\n>>> (although just having all those objects in\n>>> the .git directory still slows it down quite a bit).\n>>\n>> You're the second person who has mentioned this today (the first one\n>> was to me in a private email).  I'd like to understand this better.\n>\n> What I'm basing this on is that even when I'm using a sparse checkout\n> such that I have only a small subset of the files in my working\n> directory, git status seems singifncantly slower for me than an\n> equivalent git repository that only has that subset of files. That's\n> not very scientific, but that's what made me think just having a large\n> .git directory with lots of objects/history slows down git status even\n> if the working copy doesn't have a lot of files.\n\nHmm... I recall I experienced some slower operations on webkit with\nsparse checkout too.\n\n>\n> I will try to experiment and see if I can narrow it down with some real numbers.\n\nYes, I'd appreciate that.\n\nBy the way, how hard is it to use git-replace to implement narrow clone?\n-- \nDuy\n"},{"id":"146327","messageId":"4C4C9743.9080902@web.de","threadId":"24457","inReplyTo":"AANLkTimSoe9iqu4cJCH1d4rVsWHpFn3+8pbrCxsnVM1D@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Jens Lehmann","fromEmail":"jens.lehmann@web.de","sentAt":"2010-07-25T19:57:55Z","receivedAt":"2010-07-25T19:57:55Z","isPatch":false,"sender":{"key":"jens.lehmann@web.de","avatar":"https://avatars.githubusercontent.com/u/135220?v=4"},"body":"Am 24.07.2010 00:32, schrieb Avery Pennarun:\n> On Fri, Jul 23, 2010 at 11:10 AM, Jens Lehmann <Jens.Lehmann@web.de> wrote:\n>> You forgot what we do as best practice at work:\n>>\n>> [3] Fork the gem repos on github (or another server reachable by your\n>>    co-workers) and use those, so you don't have to change the URL\n>>    later:\n>>\n>>    git://github.com/apenwarrrubygems/gem[1..n]\n>>\n>> Your problems go away, setup has to be done only once on project\n>> start and not for every developer, you can use your own branchnames\n>> and you have a staging repo from where you can push patches upstream\n>> if necessary.\n> \n> Now all your fellow developers have to push their submodule code to a\n> single upstream repo?  That's rather centralized and un-git-like.\n\nBut isn't that exactly the same thing you would have to do for your\nsuperproject too to be able to push your changes for your fellows?\n\n\n>> It is the /commits/ that have to be\n>> done twice, once in the submodule and\n>> then in the superproject. (But that is\n>> not necessarily bad, imagine having git\n>> gui as a submodule: you would be\n>> automagically reminded that stuff for\n>> git gui should be sent somewhere else\n>> than to Junio).\n> \n> Yup, I agree that requiring a separate commit to the submodule repo is\n> not a bad idea.  I always do this anyway even when using git-subtree,\n> because I'm thinking ahead to the day when I'll push my submodule\n> changes upstream and I want my commit message to make sense.  But\n> that's because I think ahead like that.  Having the tool force me to\n> do it would be harmless and help people avoid mistakes.\n\nAnd submodules force you to do that.\n\n\n> The syntax for it ought to be nice though.  I should be able to do:\n> \n>     git commit -- path/to/submodule\n> \n> And have it commit everything in the submodule tree as a new commit in\n> the submodule.  I don't want to have to think about cd'ing to\n> path/to/submodule just so I can commit the files I changed in there.\n\nYes, that would be a nice feature (assuming you have a branch in the\nsubmodule to commit these changes to ;-).\n"},{"id":"146356","messageId":"201007261051.41663.jnareb@gmail.com","threadId":"24457","inReplyTo":"AANLkTi=LHYDhY=424YZpO3yGqGGsxpY2Sj8=ULNKvAQX@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2010-07-26T08:51:38Z","receivedAt":"2010-07-26T08:51:38Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Sat, 24 Jul 2010 00:50, Avery Pennarun wrote:\n> On Fri, Jul 23, 2010 at 11:19 AM, Marc Branchaud <marcnarc@xiplink.com> wrote:\n>> On 10-07-22 03:41 PM, Avery Pennarun wrote:\n>>> 1) Sometimes I want to clone only some subdirs of a project\n>>> 2) Sometimes I don't want the entire history because it's too big.\n>>> 3) Super huge git repositories start to degrade in performance.\n>>\n>> The reason we turned to submodules is precisely to deal with repository size.\n> \n> I believe that's very common.\n> \n> However, I wonder whether that's actually a good reason for git to\n> develop better submodules, or actually just a good reason for git to\n> get better support for handling huge repositories.\n> \n> My bup project (http://github.com/apenwarr/bup) is all about huge\n> repositories.  It handles repositories with hundreds of gigabytes, and\n> trees containing millions of files (entire filesystems), quite nicely.\n>  Of course, it's not a version control system, so it won't solve your\n> problems.  It's just evidence that large repositories are actually\n> quite manageable without changing the fundamentals of git.\n\nThere is also git-bigfiles project, although it is more about large\n[binary] files than large repositories per se (many files, long history).\n\nNote that with 'bup' you might not see problems with large repositories\nbecause it does not examine code paths that are slow in large repositories\n(gc, log, path-delimited log).\n\n>>  Our code base encompasses the entire FreeBSD tree plus different versions of\n>> the Linux kernel, along with various third-party libraries & apps.  You don't\n>> need everything to build a given product (a FreeBSD product doesn't use any\n>> Linux kernels, for example) but because all the products share common code we\n>> need to be able to branch and tag the common code along with the uncommon code.\n\nSidenote: I have noticed there very important ability of submodules, which\ngit-subtree lacks, or at least doesn't have it directly, namely ability\nto tag in submodule separately of tagging superproject as whole (so e.g.\nsuperproject v1.6.2 includes subproject 'foo' v0.99 which is foo/v0.99\ntag in superproject).\n \n>> So a straight \"git clone\" that would need to fetch all of FreeBSD plus 4\n>> different Linux kernels and check all that out is a major problem, especially\n>> for our automated build system (which could definitely be implemented better,\n>> but still).\n> \n> To be absolutely pedantic, the four linux kernels likely share most of\n> their objects and so you're only paying the cost (at least during\n> fetch) of including it once :)\n> \n> (If you're actually using git-submodule and each copy of the kernel is\n> its own module, then it might be cloning the kernel four times\n> separately, in which case the objects *don't* get shared, so this ends\n> up being much more expensive than it should be.  That could be fixed\n> by slightly improving git-submodule to share some objects rather than\n> rearchitecting it though.)\n\nThis issue is orthogonal to the fact of using submodules, it is a matter\nof setting up alternates to share object storage.\n \n>> In truth it's the checkout that takes the most time by far,\n>> though commands like git-status also take inconveniently long.\n> \n> Yeah, git could stand to be optimized a bit here.  And since Windows\n> stats files about 10x slower than Linux, this problem occurs about 10x\n> sooner on Windows, which makes using git on Windows (which sadly I\n> have to do sometimes) extremely painful compared to Linux.\n> \n> IMHO, the correct answer here is to have an inotify-based daemon prod\n> at the .git/index automatically when files get updated, so that git\n> itself doesn't have to stat/readdir through the entire tree in order\n> to do any of its operations.  (Windows also has something like inotify\n> that would work.)  If you had this, then git\n> status/diff/checkout/commit would be just as fast with zillions of\n> files as with 10 files.  Sooner or later, if nobody implements this, I\n> promise I'll get around to it since inotify is actually easy to code\n> for :)\n\nIIUC the problem is that inotify is not automatically recursive, so\ndaemon would have to take care of adding inotify trigger to each newly\ncreated subdirectory.\n\n> Also note that the only reason submodules are faster here is that\n> they're ignoring possibly important changes.  Notably, when you do\n> 'git status' from the top level, it won't warn you if you have any\n> not-yet-committed files in any of your submodules.  Personally, I\n> consider that to be really important information, but to obtain it\n> would make 'git status' take just as long as without submodules, so\n> you wouldn't get any benefit.  (I think nowadays there's a way to get\n> this recursive status information if you want it, but it'll be slow of\n> course.)\n\nErrr... didn't it got improved in recent git?  I think git-status now\nincludes information about submodules if configured so / unless configured\notherwise.  Isn't it?\n\n>> We chose git-submodule over git-subtree mainly because git-submodule lets us\n>> selectively checkout different parts of our code.  (AFAIK sparse checkouts\n>> aren't yet an option.)\n\nSparse checkouts are here, IIRC, but they do not solve problem of disk\nspace (they are still in repository, even if not checked out), and speed\n(they still need to be fetched, even if not checked out).\n\n> Fair enough.  If you could confirm or deny my theory that this is\n> *entirely* a performance related concern (as opposed to disk space /\n> download time), that would be helpful.\n> \n>> We didn't really consider git-subtree because it's\n>> not an official part of git, and we didn't want to have to teach (and nag)\n>> all our developers to install and maintain it in addition to keeping up with\n>> git itself.\n> \n> Arguably, this is a vote for including git-subtree into the core\n> (which was Bryan's point when he started this thread); it obviously is\n> being rejected sometimes by git users simply because it's not in the\n> core, even though it could help them.\n\nWell, patch management interfaces such as StGIT, Guilt and TopGit are\nalso outside git code (and should be), same with GUI tools such as qgit.\nThat shouldn't prevent people from using them ;-)\n\nBut I am all for having git-subtree in core: we have git-remote, haven't\nwe?  Besides git-subtree fits some workflows better than git-submodule\n(and vice versa).\n\n-- \nJakub Narebski\nPoland\n"},{"id":"146357","messageId":"201007261056.58985.jnareb@gmail.com","threadId":"24457","inReplyTo":"AANLkTinhd2DYh7WXzMvhMkqp98fYtTWWuQi0RSL9Rome@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2010-07-26T08:56:58Z","receivedAt":"2010-07-26T08:56:58Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Sat, Jul 24, 2010, skillzero@gmail.com napisał:\n> On Fri, Jul 23, 2010 at 3:50 PM, Avery Pennarun <apenwarr@gmail.com> wrote:\n> \n> > Honest question: do you care about the wasted disk space and download\n> > time for these extra files?  Or just the fact that git gets slow when\n> > you have them?\n> \n> I have the similar situation to the original poster (huge trees) and\n> for me it's all three: disk space, download time, and performance. My\n> tree has a few relatively small (< 20 MB) shared directories of common\n> code, a few large (2-6 GB) directories of code for OS's, and then\n> several medium size (< 500 MB) directories for application code. The\n> application developers only care about the app+shared directories (and\n> are very annoyed by the massive space and performance impact of the OS\n> directories). The firmware-only developers only care about OS+shared\n> and are mildly annoyed by the medium space and performance impact of\n> the app directories. I work on all of the pieces, but even I would\n> prefer to have things separated so when I work on the apps, git\n> status/etc doesn't take a big hit for close to a million files in the\n> OS directories (particularly when doing git status on Windows). Even\n> when using the -uno option to git status, it's still pretty slow (over\n> a minute).\n> \n> git-submodule might be technically possible in this situation, but\n> having to commit and push each submodule and then commit and push the\n> super module makes it slightly worse than just dealing with the\n> space/download/performance issues of one huge repository.\n\nBut this is just a matter for improving UI for dealing with submodules,\nisn't it.   For example having \"git commit --recursive\" would help\nwith 'having to commit each submodule', though how you would write commit\nmessages then: perhaps supermodule commit message could be by default\ncomposed out of submodules commits (if any).  \"git push --recursive\"\n(or some support for push in \"git remote\") would help with 'having to\npush each submodule'.\n\nIsn't it?\n-- \nJakub Narebski\nPoland\n"},{"id":"146371","messageId":"201007261513.57045.jnareb@gmail.com","threadId":"24457","inReplyTo":"AANLkTikx5EtQ0yvdkqN1Q1QAudFZfbd+_jpoa9ztLrz1@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2010-07-26T13:13:53Z","receivedAt":"2010-07-26T13:13:53Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Sat, Jul 24, 2010, skillzero@gmail.com wrote:\n> On Fri, Jul 23, 2010 at 6:20 PM, Avery Pennarun <apenwarr@gmail.com> wrote:\n>> On Fri, Jul 23, 2010 at 8:58 PM,  <skillzero@gmail.com> wrote:\n>>> On Fri, Jul 23, 2010 at 3:50 PM, Avery Pennarun <apenwarr@gmail.com> wrote:\n\n>> This is indeed a problem with large repositories.  Of course,\n>> splitting them with git-submodule is kind of cheating, because it just\n>> makes git-status *not look* to see if those files are dirty or not.\n>> If they are dirty and you forget to commit them, you'll never know\n>> until someone tells you later.  It would be functionally equivalent to\n>> just have git-status not look inside certain subdirs of a single\n>> repository.\n> \n> I think it's only cheating if you're using all of the submodules. The\n> main purpose of submodules for me (although I don't currently use\n> submodules) would be so I don't need to keep modules on disk that I\n> don't care about. If a developer is working on an app, they don't need\n> the OS directories/modules so they get much faster git status/etc and\n> there wouldn't be other directories to have dirty files in. [...]\n\nThere are two issues that make submodules or git-subtree a better\nsolution.  If you work with subprojects via upstream subproject \nrepository, and you don't always need / want all subprojects, \ngit-submodule is better.  If you always have checked out all subprojects,\nand you edit them in superproject, git-subtree is better.\n \n\n-- \nJakub Narebski\nPoland\n"},{"id":"146381","messageId":"4C4DA683.9020102@xiplink.com","threadId":"24457","inReplyTo":"AANLkTi=LHYDhY=424YZpO3yGqGGsxpY2Sj8=ULNKvAQX@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Marc Branchaud","fromEmail":"marcnarc@xiplink.com","sentAt":"2010-07-26T15:15:15Z","receivedAt":"2010-07-26T15:15:15Z","isPatch":false,"sender":{"key":"marcnarc@xiplink.com","avatar":"https://avatars.githubusercontent.com/u/14980203?v=4"},"body":"On 10-07-23 06:50 PM, Avery Pennarun wrote:\n> On Fri, Jul 23, 2010 at 11:19 AM, Marc Branchaud <marcnarc@xiplink.com> wrote:\n>> On 10-07-22 03:41 PM, Avery Pennarun wrote:\n>>> 1) Sometimes I want to clone only some subdirs of a project\n>>> 2) Sometimes I don't want the entire history because it's too big.\n>>> 3) Super huge git repositories start to degrade in performance.\n>>\n>> The reason we turned to submodules is precisely to deal with repository size.\n> \n> I believe that's very common.\n> \n> However, I wonder whether that's actually a good reason for git to\n> develop better submodules, or actually just a good reason for git to\n> get better support for handling huge repositories.\n\nI think that's a fundamental question, but part of the problem in coming up\nwith an answer is that there's no agreed-upon definition of how to handle\nhuge repos.  People have provided tools that answer the question in ways they\nlike, but I think the fact that these issues keep coming up is proof that git\nisn't there yet.\n\n>>  Our code base encompasses the entire FreeBSD tree plus different versions of\n>> the Linux kernel, along with various third-party libraries & apps.  You don't\n>> need everything to build a given product (a FreeBSD product doesn't use any\n>> Linux kernels, for example) but because all the products share common code we\n>> need to be able to branch and tag the common code along with the uncommon code.\n> \n> Honest question: do you care about the wasted disk space and download\n> time for these extra files?  Or just the fact that git gets slow when\n> you have them?\n\nIt's not the disk space or the extra download time.  It's how long takes to\ncheckout all those files, and how long it takes to \"git status\" in a unified\nrepo.\n\n>> So a straight \"git clone\" that would need to fetch all of FreeBSD plus 4\n>> different Linux kernels and check all that out is a major problem, especially\n>> for our automated build system (which could definitely be implemented better,\n>> but still).\n> \n> To be absolutely pedantic, the four linux kernels likely share most of\n> their objects and so you're only paying the cost (at least during\n> fetch) of including it once :)\n\nThat is true, but like I said the problem is the checkout.  Our different\nproducts use different kernels (or FreeBSD):\n\n\tProduct 1 -- Linux vX\n\tProduct 2 -- Linux vY\n\tProduct 3 -- FreeBSD\n\n(Luckily we're currently only using one version of FreeBSD...)\n\nAll the products use common code.  When we release, we need to tag the common\ncode and the particular Linux kernel (or FreeBSD) we built the product with.\n We can't stuff all the Linux kernels into a single submodule, because then\nthe repo will be \"dirty\" if we checkout a different Linux kernel to build a\ndifferent product.  Even in a unified repo we'd need the kernels to live in\ntheir own trees.\n\nSo we've ended up with individual submodules for each Linux kernel, and we've\ntaught our automated build to only clone/checkout the kernel it needs to\nbuild the target product.  Otherwise the checkout I/O overshadows the actual\nbuild time, especially when we try to run several builds in parallel on one\nslave machine.\n\n> (If you're actually using git-submodule and each copy of the kernel is\n> its own module, then it might be cloning the kernel four times\n> separately, in which case the objects *don't* get shared, so this ends\n> up being much more expensive than it should be.  That could be fixed\n> by slightly improving git-submodule to share some objects rather than\n> rearchitecting it though.)\n\nEven with the --reference parameter, it's still a problem.\n\n>>  In truth it's the checkout that takes the most time by far,\n>> though commands like git-status also take inconveniently long.\n> \n> Yeah, git could stand to be optimized a bit here.  And since Windows\n> stats files about 10x slower than Linux, this problem occurs about 10x\n> sooner on Windows, which makes using git on Windows (which sadly I\n> have to do sometimes) extremely painful compared to Linux.\n> \n> IMHO, the correct answer here is to have an inotify-based daemon prod\n> at the .git/index automatically when files get updated, so that git\n> itself doesn't have to stat/readdir through the entire tree in order\n> to do any of its operations.  (Windows also has something like inotify\n> that would work.)  If you had this, then git\n> status/diff/checkout/commit would be just as fast with zillions of\n> files as with 10 files.  Sooner or later, if nobody implements this, I\n> promise I'll get around to it since inotify is actually easy to code\n> for :)\n> \n> Also note that the only reason submodules are faster here is that\n> they're ignoring possibly important changes.  Notably, when you do\n> 'git status' from the top level, it won't warn you if you have any\n> not-yet-committed files in any of your submodules.  Personally, I\n> consider that to be really important information, but to obtain it\n> would make 'git status' take just as long as without submodules, so\n> you wouldn't get any benefit.  (I think nowadays there's a way to get\n> this recursive status information if you want it, but it'll be slow of\n> course.)\n\nI'm happy with a \"git status\" that can ignore uninitialized submodules and\nstill probe into initialized/cloned ones.  I agree that it's important for\n\"git status\" to be correct.\n\n>> We chose git-submodule over git-subtree mainly because git-submodule lets us\n>> selectively checkout different parts of our code.  (AFAIK sparse checkouts\n>> aren't yet an option.)\n> \n> Fair enough.  If you could confirm or deny my theory that this is\n> *entirely* a performance related concern (as opposed to disk space /\n> download time), that would be helpful.\n\nConsider it confirmed.  Honestly, disk space is a complete non-issue.  It's\nalways nice to have faster download times, but it hasn't been an issue for us\nand there are already several ways to work around it anyway.\n\n>>  We didn't really consider git-subtree because it's\n>> not an official part of git, and we didn't want to have to teach (and nag)\n>> all our developers to install and maintain it in addition to keeping up with\n>> git itself.\n> \n> Arguably, this is a vote for including git-subtree into the core\n> (which was Bryan's point when he started this thread); it obviously is\n> being rejected sometimes by git users simply because it's not in the\n> core, even though it could help them.\n\nYes, I have no objection to seeing git-subtree becoming an official part of\ngit.  My only complaint would be that it doesn't really help git deal with\nhuge repos.\n\n\t\tM.\n"},{"id":"146392","messageId":"4C4DB9AC.9000306@xiplink.com","threadId":"24457","inReplyTo":"AANLkTimLayG_HFxGdq+Tt8hU_MApBpSdHHiYPxcakpRJ@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Marc Branchaud","fromEmail":"marcnarc@xiplink.com","sentAt":"2010-07-26T16:37:00Z","receivedAt":"2010-07-26T16:37:00Z","isPatch":false,"sender":{"key":"marcnarc@xiplink.com","avatar":"https://avatars.githubusercontent.com/u/14980203?v=4"},"body":"On 10-07-23 09:20 PM, Avery Pennarun wrote:\n> \n> I'm pushing extra hard on this because I believe there are lots of\n> opportunities to just improve git performance on huge repositories.\n> And if the only *real* reason people need to split repositories is\n> that performance goes down, then that's fixable, and you may need\n> neither git-submodule nor git-subtree.\n\nI think I should mention one aspect of what we're doing, which is that a lot\nof our submodules are based on external code, and that we occasionally need\nto modify or customize some of that code.  So it's quite nice for us to\nmaintain private git mirrors of the external repos, with our own private\nbranches that contain our modifications.  Although we want to get much of our\nchanges incorporated into the upstream code bases, upstream release cycles\nare rarely in sync with ours.\n\nSo it's very convenient for use to have our external-code modifications\ncontained in private branches in our private mirrors, and to rebase those\nbranches to keep up with upstream releases.  We also often use these private\nbranches to maintain the code that integrates the external code bases into\nour overall build system.\n\nI mention this purely because this pattern is so convenient that I don't want\nto see it get lost in whatever may arise from this discussion.\n\n\t\tM.\n"},{"id":"146394","messageId":"AANLkTimQywtn-0Fcr-ceLeHGeSBNROt+T=K+TowF_u5h@mail.gmail.com","threadId":"24457","inReplyTo":"4C4DB9AC.9000306@xiplink.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2010-07-26T16:41:42Z","receivedAt":"2010-07-26T16:41:42Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Mon, Jul 26, 2010 at 9:37 AM, Marc Branchaud <marcnarc@xiplink.com> wrote:\n>\n> I think I should mention one aspect of what we're doing, which is that a lot\n> of our submodules are based on external code, and that we occasionally need\n> to modify or customize some of that code.  So it's quite nice for us to\n> maintain private git mirrors of the external repos, with our own private\n> branches that contain our modifications.  Although we want to get much of our\n> changes incorporated into the upstream code bases, upstream release cycles\n> are rarely in sync with ours.\n\nTHIS.\n\nThis is why I always thought that submodules absolutely have to be\ncommits, not trees. It's why the git submodule data structures are\ndone the way they are. Anything that makes the submodule just a tree\nis fundamentally broken, I think.\n\nThat said, I'm not competent to comment on the actual user interface\nissues. I can well believe that git-subtree has a nicer interface.\n\n             Linus\n"},{"id":"146397","messageId":"AANLkTimAkfgsMLUY5Lhi=-Rd=v5ZkT4_SBd8wxHhG-LR@mail.gmail.com","threadId":"24457","inReplyTo":"AANLkTil2ZHsV-PH46wfeqqvB27akRw3egDWRtJbIPLXb@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Eugene Sajine","fromEmail":"euguess@gmail.com","sentAt":"2010-07-26T17:34:57Z","receivedAt":"2010-07-26T17:34:57Z","isPatch":false,"sender":{"key":"euguess@gmail.com","avatar":null},"body":"On Thu, Jul 22, 2010 at 5:33 PM, Avery Pennarun <apenwarr@gmail.com> wrote:\n> On Thu, Jul 22, 2010 at 4:17 PM, Ęvar Arnfjörš Bjarmason\n> <avarab@gmail.com> wrote:\n>> But it's also clear that we have a lot of tribal knowledge about the\n>> lackings of git submodule / git subtree. It would be *really* useful\n>> if people like Avery and Jens which have obviously thought hard about\n>> the submodule/subtree issues would draft up some (calmly written) docs\n>> about how the two differ (with comparison tables etc.).\n>>\n>> That'd be a very helpful resource for Git users in deciding which one\n>> to use.\n>\n> I think I'm too biased to write that, but if someone else wants to\n> take the lead, I could certainly contribute.\n>\n> Have fun,\n>\n> Avery\n\n\nI personally tried to understand submodules, but my attempts to find\neasy way to use them have failed miserably;) probably i have to spend\neven more time in order to understand if i can benefit from them or\nnot. So, i think this kind of comparison would be very beneficial for\n\"mere mortals\"\n\nI would like to share an idea how it can be organized:\n\nWe could create a file in doc section of git.git or in Avery's repo\nnamed git_submodule_vs_git_subtree or just use a separate topic of the\nlist.\n\nThe file would look like this:\n\ngit-submodule |           feature                  | git-subtree\n______________________________________________________________________\n     +        | ability to tag submodule without   |     -\n   (comments) | tagging the whole tree             |  (comments)\n______________________________________________________________________\n\n\nAvery and Jens could add features they think are beneficial for one\nproject or another and answer to each other this way. They could mark\njust presence or abscence of the feature by +/- like above or specify\nkey approaches how to do different things.\nFor example, how to configure new submodule (main sequence of commands\nto create, add ), how to do that with sub-tree...\n\nI think this simple feature matrix will answer a lot of questions.\n\njust my 2 cents...\n\nThanks,\nEugene\n"},{"id":"146398","messageId":"4C4DC799.6070702@gmail.com","threadId":"24457","inReplyTo":"AANLkTimQywtn-0Fcr-ceLeHGeSBNROt+T=K+TowF_u5h@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Bryan Larsen","fromEmail":"bryan.larsen@gmail.com","sentAt":"2010-07-26T17:36:25Z","receivedAt":"2010-07-26T17:36:25Z","isPatch":false,"sender":{"key":"bryan@larsen.st","avatar":"https://avatars.githubusercontent.com/u/32073?v=4"},"body":"On 10-07-26 12:41 PM, Linus Torvalds wrote:\n> On Mon, Jul 26, 2010 at 9:37 AM, Marc Branchaud<marcnarc@xiplink.com>  wrote:\n>>\n>> I think I should mention one aspect of what we're doing, which is that a lot\n>> of our submodules are based on external code, and that we occasionally need\n>> to modify or customize some of that code.  So it's quite nice for us to\n>> maintain private git mirrors of the external repos, with our own private\n>> branches that contain our modifications.  Although we want to get much of our\n>> changes incorporated into the upstream code bases, upstream release cycles\n>> are rarely in sync with ours.\n>\n> THIS.\n>\n> This is why I always thought that submodules absolutely have to be\n> commits, not trees. It's why the git submodule data structures are\n> done the way they are. Anything that makes the submodule just a tree\n> is fundamentally broken, I think.\n>\n> That said, I'm not competent to comment on the actual user interface\n> issues. I can well believe that git-subtree has a nicer interface.\n>\n>               Linus\n>\n\nTo me, that's what git-subtree is: an internal private mirror of an \nexternal repo.   Using git submodule moves that into a separately \nmanaged repo, which is just unnecessary hassle.  Why maintain repo \ncalled \"clone of library X for project A\" when you can just stick it \ninside of project A without any downsides?\n\nFor us, changes are made in the superproject and tested in the \nsuperproject.  Once they're tested, a git subtree push or a git subtree \nsplit pushes the patches to the subproject.   Once the subproject has \naccepted the patches, a git subtree pull merges them.   Same workflow as \nthe \"private git mirror of external repo\" listed above, just without the \nhassle of having another repo to manage.\n\nBryan\n"},{"id":"146401","messageId":"AANLkTi=XrAwSe9Lfr8FDT00VS5+PZDx3pvh6+hC8wy2Z@mail.gmail.com","threadId":"24457","inReplyTo":"4C4DC799.6070702@gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2010-07-26T17:48:54Z","receivedAt":"2010-07-26T17:48:54Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Mon, Jul 26, 2010 at 10:36 AM, Bryan Larsen <bryan.larsen@gmail.com> wrote:\n>\n> To me, that's what git-subtree is: an internal private mirror of an external\n> repo.   Using git submodule moves that into a separately managed repo, which\n> is just unnecessary hassle.  Why maintain repo called \"clone of library X\n> for project A\" when you can just stick it inside of project A without any\n> downsides?\n\nWithout any downsides?\n\nWhat about merging? What about complex history? IOW, what about\n_anything_ but a few extra one-liner patches?\n\nBackground: the only time I ever used CVS modules, we had submodules\nfor things like gcc, binutils, etc. And maintained them separately\nfrom upstream for _years_. Not with some simple one-liner fixes, but\nwith big fundamental changes that couldn't be sent upstream (and\nwouldn't have been accepted anyway) etc.\n\nTHAT is the problem space. Not \"just a mirror of another project\".\n\n                   Linus\n"},{"id":"146502","messageId":"20100727182841.GA25124@worldvisions.ca","threadId":"24457","inReplyTo":"AANLkTimQywtn-0Fcr-ceLeHGeSBNROt+T=K+TowF_u5h@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-27T18:28:41Z","receivedAt":"2010-07-27T18:28:41Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Mon, Jul 26, 2010 at 09:41:42AM -0700, Linus Torvalds wrote:\n\n> On Mon, Jul 26, 2010 at 9:37 AM, Marc Branchaud <marcnarc@xiplink.com> wrote:\n> >\n> > I think I should mention one aspect of what we're doing, which is that a lot\n> > of our submodules are based on external code, and that we occasionally need\n> > to modify or customize some of that code.  So it's quite nice for us to\n> > maintain private git mirrors of the external repos, with our own private\n> > branches that contain our modifications.  Although we want to get much of our\n> > changes incorporated into the upstream code bases, upstream release cycles\n> > are rarely in sync with ours.\n> \n> THIS.\n> \n> This is why I always thought that submodules absolutely have to be\n> commits, not trees. It's why the git submodule data structures are\n> done the way they are. Anything that makes the submodule just a tree\n> is fundamentally broken, I think.\n\nI agree completely.  The major failing of git-subtree is that it uses\ntree->tree links instead of tree->commit links.\n\nThis was necessary only because git fundamentally *mistreats* tree->commit\nlinks: it refuses to push or fetch through them automatically.  That is,\nwhen I fetch a superproject that has a tree->commit link in it, git won't\nfetch the subproject's history starting at the targeted commit, even if the\nremote repo *has* that history.  And if I make a patch to the subproject,\npushing the superproject won't push that patch.\n\nHave fun,\n\nAvery\n"},{"id":"146506","messageId":"20100727183658.GB25124@worldvisions.ca","threadId":"24457","inReplyTo":"201007261056.58985.jnareb@gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-27T18:36:58Z","receivedAt":"2010-07-27T18:36:58Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Mon, Jul 26, 2010 at 10:56:58AM +0200, Jakub Narebski wrote:\n> On Sat, Jul 24, 2010, skillzero@gmail.com napisał:\n> > git-submodule might be technically possible in this situation, but\n> > having to commit and push each submodule and then commit and push the\n> > super module makes it slightly worse than just dealing with the\n> > space/download/performance issues of one huge repository.\n> \n> But this is just a matter for improving UI for dealing with submodules,\n> isn't it.   For example having \"git commit --recursive\" would help\n> with 'having to commit each submodule', though how you would write commit\n> messages then: perhaps supermodule commit message could be by default\n> composed out of submodules commits (if any).  \"git push --recursive\"\n> (or some support for push in \"git remote\") would help with 'having to\n> push each submodule'.\n\nFor \"recursive\" commit, for my own workflow, I would rather have it work\nlike this: from the toplevel, I can 'git commit' any set of files, as long\nas they all fall inside a particular submodule.  That is, if I do\n\n\tgit commit mod1/*.c mod2/*.c\n\t\nit should reject it (with a helpful message), because the commit would cross\nsubmodule boundaries.  But if I do\n\n\tgit commit mod1/*.c\n\t\nI think it should create a new commit in mod1, leave my superproject\npointing at that new commit, and stop (ie. without the superproject having\ncommitted the new commit pointer).\n\nWhy?  Because my normal workflow is:\n\n  - make a bunch of superproject/submodule changes until they work.\n  - commit the submodule changes with a submodule-relevant message\n  - commit the superproject change with a supermodule-relevant message\n  \nI wouldn't want to share commit messages between the two, so actually having\na single commit process be \"recursive\" would not do me any good.\n\nHowever, pushing is a separate issue entirely.  Having push be recursive\nwould be easy, but it doesn't solve the *real* problem with pushing: git\ndoesn't know what branch to push to in the submodule, and the submodule most\nlikely isn't pointing at a pushable repo at all, even if the supermodule is. \nThis is why I keep coming back to the idea that I really want to push all\nthe submodule objects into the superproject's repo.\n\nHave fun,\n\nAvery\n"},{"id":"146508","messageId":"20100727184047.GC25124@worldvisions.ca","threadId":"24457","inReplyTo":"4C4C9743.9080902@web.de","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-27T18:40:47Z","receivedAt":"2010-07-27T18:40:47Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Sun, Jul 25, 2010 at 09:57:55PM +0200, Jens Lehmann wrote:\n\n> Am 24.07.2010 00:32, schrieb Avery Pennarun:\n> > On Fri, Jul 23, 2010 at 11:10 AM, Jens Lehmann <Jens.Lehmann@web.de> wrote:\n> >> You forgot what we do as best practice at work:\n> >>\n> >> [3] Fork the gem repos on github (or another server reachable by your\n> >>    co-workers) and use those, so you don't have to change the URL\n> >>    later:\n> >>\n> >>    git://github.com/apenwarrrubygems/gem[1..n]\n> >>\n> >> Your problems go away, setup has to be done only once on project\n> >> start and not for every developer, you can use your own branchnames\n> >> and you have a staging repo from where you can push patches upstream\n> >> if necessary.\n> > \n> > Now all your fellow developers have to push their submodule code to a\n> > single upstream repo?  That's rather centralized and un-git-like.\n> \n> But isn't that exactly the same thing you would have to do for your\n> superproject too to be able to push your changes for your fellows?\n\nNo.  On github, only I can push to my superproject's history, and yet\neveryone can still pull from me.\n\nWith what you're proposing, for all my submodules, we can't each have our\nown project; we all have to push to the shared one.\n\n(Just to be clear: I don't want to fork *every submodule by hand every\ntime*.  I just want *my* stuff to be in *my* repo.  The easiest way to do\nthis would be to have all my changes in a single repo, ie. my fork of the\nsuperproject.)\n\n> >> It is the /commits/ that have to be\n> >> done twice, once in the submodule and\n> >> then in the superproject. (But that is\n> >> not necessarily bad, imagine having git\n> >> gui as a submodule: you would be\n> >> automagically reminded that stuff for\n> >> git gui should be sent somewhere else\n> >> than to Junio).\n> > \n> > Yup, I agree that requiring a separate commit to the submodule repo is\n> > not a bad idea.  I always do this anyway even when using git-subtree,\n> > because I'm thinking ahead to the day when I'll push my submodule\n> > changes upstream and I want my commit message to make sense.  But\n> > that's because I think ahead like that.  Having the tool force me to\n> > do it would be harmless and help people avoid mistakes.\n> \n> And submodules force you to do that.\n\nYes.  This is a limitation of submodules, but not one that bothers me.  And\nit encourages good behaviour.\n\n> > The syntax for it ought to be nice though.  I should be able to do:\n> > \n> >     git commit -- path/to/submodule\n> > \n> > And have it commit everything in the submodule tree as a new commit in\n> > the submodule.  I don't want to have to think about cd'ing to\n> > path/to/submodule just so I can commit the files I changed in there.\n> \n> Yes, that would be a nice feature (assuming you have a branch in the\n> submodule to commit these changes to ;-).\n\nNo, I explicitly *don't* want to have to have a branch in the submodule;\nthat's too much extra thinking at that stage.\n\nHave fun,\n\nAvery\n"},{"id":"146511","messageId":"AANLkTi=6SDQ2A0Zxf8DiSSNzSfUS43M7wmCkKKraOd8w@mail.gmail.com","threadId":"24457","inReplyTo":"201007261051.41663.jnareb@gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-27T19:15:08Z","receivedAt":"2010-07-27T19:15:08Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Mon, Jul 26, 2010 at 4:51 AM, Jakub Narebski <jnareb@gmail.com> wrote:\n> On Sat, 24 Jul 2010 00:50, Avery Pennarun wrote:\n>> My bup project (http://github.com/apenwarr/bup) is all about huge\n>> repositories.  It handles repositories with hundreds of gigabytes, and\n>> trees containing millions of files (entire filesystems), quite nicely.\n>>  Of course, it's not a version control system, so it won't solve your\n>> problems.  It's just evidence that large repositories are actually\n>> quite manageable without changing the fundamentals of git.\n>\n> There is also git-bigfiles project, although it is more about large\n> [binary] files than large repositories per se (many files, long history).\n\nRight.  git-bigfiles is valuable, but it's valuable with or without\nsubmodules.  (If you have large blobs, submodules won't save you.)\n\nbup happens to have its own way of dealing with large files too, but\nit may not be applicable to git.  It does result in lots and lots of\nsmaller objects, though, which is why I know git repositories are\nfundamentally capable of handling lots and lots of smaller objects :)\n\n> Note that with 'bup' you might not see problems with large repositories\n> because it does not examine code paths that are slow in large repositories\n> (gc, log, path-delimited log).\n\ngc is a huge problem.  bup avoids it entirely (it foregoes delta\ncompression); git gc fails completely on such large repositories (100+\nGB).  There's no reason this has to be true forever, but yes, to\nsupport really big repos, git gc would need to be improved somewhat.\nFor most reasonably sane repos (a few GB) you can get reasonable\nperformance by just making your biggest packfiles .keep so they don't\nkeep getting repacked all the time.\n\nCompared to that, log feels like not a problem at all :)  At least\nperformance-wise.  The thing that sucks about log using git-subtree,\nof course, is that you get all these log messages from multiple\nprojects jammed together into a single repo, which is rarely what you\nwant, even if it's fast.  I think the \"best\" solution is a single repo\nwith all your objects, but still keeping the histories of each\nsubmodule separate.\n\n>> IMHO, the correct answer here is to have an inotify-based daemon prod\n>> at the .git/index automatically when files get updated, so that git\n>> itself doesn't have to stat/readdir through the entire tree in order\n>> to do any of its operations.  (Windows also has something like inotify\n>> that would work.)  If you had this, then git\n>> status/diff/checkout/commit would be just as fast with zillions of\n>> files as with 10 files.  Sooner or later, if nobody implements this, I\n>> promise I'll get around to it since inotify is actually easy to code\n>> for :)\n>\n> IIUC the problem is that inotify is not automatically recursive, so\n> daemon would have to take care of adding inotify trigger to each newly\n> created subdirectory.\n\nYeah, the inotify API is kind of gross that way.  But it can be done,\nand people do.  (eg. the beagle project)\n\n>> Also note that the only reason submodules are faster here is that\n>> they're ignoring possibly important changes.  Notably, when you do\n>> 'git status' from the top level, it won't warn you if you have any\n>> not-yet-committed files in any of your submodules.  Personally, I\n>> consider that to be really important information, but to obtain it\n>> would make 'git status' take just as long as without submodules, so\n>> you wouldn't get any benefit.  (I think nowadays there's a way to get\n>> this recursive status information if you want it, but it'll be slow of\n>> course.)\n>\n> Errr... didn't it got improved in recent git?  I think git-status now\n> includes information about submodules if configured so / unless configured\n> otherwise.  Isn't it?\n\nYes, but you're still left with the choice between slow (checks all\nfiles in all submodules) and not slow (might miss stuff).  This isn't\na submodule question, really, it's an overall performance question\nwith huge checkouts with or without submodules.\n\n>>> We chose git-submodule over git-subtree mainly because git-submodule lets us\n>>> selectively checkout different parts of our code.  (AFAIK sparse checkouts\n>>> aren't yet an option.)\n>\n> Sparse checkouts are here, IIRC, but they do not solve problem of disk\n> space (they are still in repository, even if not checked out), and speed\n> (they still need to be fetched, even if not checked out).\n\nHmm, don't mix bandwidth usage (and thus the slowness of fetch) with\nslowness during everyday usage.  I don't mind a slow fetch now and\nthen, but 'git status' should be fast. AFAIK, sparse checkouts\n*should* make git status faster.  If they don't, it's probably just a\nbug.\n\nHave fun,\n\nAvery\n"},{"id":"146520","messageId":"7vaapc7jv8.fsf@alter.siamese.dyndns.org","threadId":"24457","inReplyTo":"20100727182841.GA25124@worldvisions.ca","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2010-07-27T20:25:15Z","receivedAt":"2010-07-27T20:25:15Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Avery Pennarun <apenwarr@gmail.com> writes:\n\n> On Mon, Jul 26, 2010 at 09:41:42AM -0700, Linus Torvalds wrote:\n>\n>> On Mon, Jul 26, 2010 at 9:37 AM, Marc Branchaud <marcnarc@xiplink.com> wrote:\n>> >\n>> > I think I should mention one aspect of what we're doing, which is that a lot\n>> > of our submodules are based on external code, and that we occasionally need\n>> > to modify or customize some of that code.  So it's quite nice for us to\n>> > maintain private git mirrors of the external repos, with our own private\n>> > branches that contain our modifications.  Although we want to get much of our\n>> > changes incorporated into the upstream code bases, upstream release cycles\n>> > are rarely in sync with ours.\n>> \n>> THIS.\n>> \n>> This is why I always thought that submodules absolutely have to be\n>> commits, not trees. It's why the git submodule data structures are\n>> done the way they are. Anything that makes the submodule just a tree\n>> is fundamentally broken, I think.\n>\n> I agree completely.  The major failing of git-subtree is that it uses\n> tree->tree links instead of tree->commit links.\n>\n> This was necessary only because git fundamentally *mistreats* tree->commit\n> links: it refuses to push or fetch through them automatically.\n\nI do not think that is so \"fundamental\" as you seem to think.\n\nIsn't it just the matter of how the default UI of object transfer commands\n(like push and fetch) are set up?\n\nAdmittedly, the way the default UI is set up is to strongly favor the\nearly design decision we made back when Linus did his initial \"gitlink\"\nimplementation, which is \"separate project lives in a separate repository,\nand not having to check out any subproject should be the norm for using a\nsuperproject\".  \n\nSome \"recursive\" operations have been added to commands for which it makes\nsense (e.g. \"clone --recursive\") by people who cared enough.  Even though\nthere are a few other commands that shouldn't ever learn the recursive\nmode (e.g. \"commit --recursive -m $msg\" would not make sense), there still\nare some commands where a similar \"--recursive\" option would make sense\nbut haven't learned it (e.g. \"push --recursive\").\n\nI also consider it merely a lack of UI enhancement that you have to clone\nthe submodule again (or cannot switch to a clean slate very easily) when\nswitching between revisions of superproject before and after you add a\nsubmodule, and nothing fundamental.  \n\nWhen switching back in history to lose a recent submodule, the user\nexperience should be like switching to a revision that didn't have a\ndirectory.  You shouldn't be able to lose your change in that directory,\nbut if the directory is clean, you should be able to lose it.  And when\nyou switch to a more recent revision that has the submodule, you should be\nable to get it back (again, if you have a precious file there, the\ncheckout should barf).\n\nWe have added support for having \"gitdir: $dir\" in a regular file .git\nexactly because we wanted to be able to stash away the submodule's .git\ndirectory somewhere inside .git (e.g. .git/modules/<submodulename>) in the\nsuperproject when we do that kind of branch switching, so that we can get\nit back when switching back to a revision with the submodule without\nhaving to re-clone (also this presumably would help when you move the\nsubmodule in the superproject tree), but there haven't been further work\nto make use of this in \"git submodule update\" (it probably needs to start\nby teaching \"git clone\" how to make use of \"gitdir: $dir\", if anybody is\ninterested).\n\nBy the way, I also do not think it is such a bad thing that git-subtree\ndoes not bind commit into its superproject tree while it is working\n\"natively\" (in a \"git-subtree\" workflow), but allows users to easily split\nthe history into an exportable shape to upstreams of its submodules when\nsuch an operqation is needed.  If you rarely push back to upstreams but\nconstantly consume their changes, that sounds like a reasonable way to go.\n"},{"id":"146521","messageId":"AANLkTim0A0MAmpgAiaYSgYO=YbZ2gc4Upx3MQQopx6DG@mail.gmail.com","threadId":"24457","inReplyTo":"7vaapc7jv8.fsf@alter.siamese.dyndns.org","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-27T20:57:24Z","receivedAt":"2010-07-27T20:57:24Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Tue, Jul 27, 2010 at 4:25 PM, Junio C Hamano <gitster@pobox.com> wrote:\n> Avery Pennarun <apenwarr@gmail.com> writes:\n>> I agree completely.  The major failing of git-subtree is that it uses\n>> tree->tree links instead of tree->commit links.\n>>\n>> This was necessary only because git fundamentally *mistreats* tree->commit\n>> links: it refuses to push or fetch through them automatically.\n>\n> I do not think that is so \"fundamental\" as you seem to think.\n>\n> Isn't it just the matter of how the default UI of object transfer commands\n> (like push and fetch) are set up?\n\nWell, I call it fundamental because there's currently no way to get\nthe git UI to do otherwise.  It's not really just a \"default.\"  To\ndepend on this changing would have prevented me from writing\ngit-subtree, which is why I didn't depend on it.  However, I agree\nthat it's fixable.\n\nNote that the way git treats a checked-out submodule (as you describe\nbelow) is also very fundamental to how this works.  git-subtree\nwouldn't have the usability that it does if 'git checkout branchname'\ndidn't work perfectly will all the subtrees, which it currently does,\nbut which it wouldn't if I had relied on tree->commit links.\n\n> Some \"recursive\" operations have been added to commands for which it makes\n> sense (e.g. \"clone --recursive\") by people who cared enough.  Even though\n> there are a few other commands that shouldn't ever learn the recursive\n> mode (e.g. \"commit --recursive -m $msg\" would not make sense), there still\n> are some commands where a similar \"--recursive\" option would make sense\n> but haven't learned it (e.g. \"push --recursive\").\n\nOne problem with this line of reasoning is that \"--recursive\" is\nalways an option.  But if submodules are ever to be easy to use, I\nthink it should be the default (or settable as a default using git\nconfig).  This would take us a *long* way towards usability (of\ncourse, in addition to adding the missing features, as you mention).\n\nAlso, I haven't tried it, but I think 'git gc' will prune away objects\nif the only reference to them is a 'commit' link from a tree.  This\nwould be undesirable too.\n\n> I also consider it merely a lack of UI enhancement that you have to clone\n> the submodule again (or cannot switch to a clean slate very easily) when\n> switching between revisions of superproject before and after you add a\n> submodule, and nothing fundamental.\n\nI mostly agree with this.  There is one problem I don't know how to\nsolve with this idea, though: what happens when commit A adds a\nsubmodule in modules/mod1, commit B removes it, and then commit C\nre-adds the same submodules in modules/mod1-again?  Will it reuse the\nsame submodule .git directory or a new one?  Share objects or not?\nShare branch names or not?  Share .git/config or not?\n\nUnless you have some kind of \"unique id\" scheme for submodules, this\ngets impossible to handle correctly.  And the git objects themselves\n(trees that link to commits) have nowhere to put such things.\n\nBy comparison, simply putting all the stuff related to all the\nsubmodules into the supermodule's repo creates none of these confusing\nproblems.  You could even still choose not to checkout individual\nsubmodules' trees if you wanted.\n\n> When switching back in history to lose a recent submodule, the user\n> experience should be like switching to a revision that didn't have a\n> directory.  You shouldn't be able to lose your change in that directory,\n> but if the directory is clean, you should be able to lose it.  And when\n> you switch to a more recent revision that has the submodule, you should be\n> able to get it back (again, if you have a precious file there, the\n> checkout should barf).\n\nIt sounds like you're proposing that we delete the entire submodule's\ndirectory hierarchy when the submodule commit link goes away.  Note\nthat this isn't what happens in the non-submodule case: all the *.o\nfiles, for example, in a deleted subdirectory are not automatically\ndeleted by git.  And I think this is the behaviour we should expect.\n\nWith that in mind, the situations where checkout barfs because of a\n\"precious\" file should be the same as they are in normal git: it\nshould only be a problem if the files in question differ between the\noriginally-checked-out tree and the newly-checked-out tree.\n\nApologies if that's what you meant in the first place.\n\n> We have added support for having \"gitdir: $dir\" in a regular file .git\n> exactly because we wanted to be able to stash away the submodule's .git\n> directory somewhere inside .git (e.g. .git/modules/<submodulename>) in the\n> superproject when we do that kind of branch switching, so that we can get\n> it back when switching back to a revision with the submodule without\n> having to re-clone (also this presumably would help when you move the\n> submodule in the superproject tree), but there haven't been further work\n> to make use of this in \"git submodule update\" (it probably needs to start\n> by teaching \"git clone\" how to make use of \"gitdir: $dir\", if anybody is\n> interested).\n\nI guess the real question is: just how much of a \"real\" repository do\nwe want a submodule to act like?\n\nThoughts:\n\n- object store: I think this should just always be shared with the\nsuperproject.  There's no reason to separate them that I can see.\n\n- branches: should be a way to simply not worry about branches and\njust use what's in the superproject.  Other people seem to want to be\nable to have a set of branches/tags for their submodule.\n\n- .git/config: entirely shared?  entirely separate?\n\n- remotes: I would want my submodules to never do their own\npushing/pulling, and leave that to the supermodule; other people seem\nto disagree.\n\nFor the particular model I'm proposing, I'm just not sure that *any*\nof the features of a separate repo are warranted... and having them\nadds a lot of complication.  (In the most basic level, you suddenly\nneed to track .git directories as submodules are added/deleted/moved\naround when you checkout different revisions of the superproject, and\nthere seems to be no way to do that elegantly.)\n\nHave fun,\n\nAvery\n"},{"id":"146533","messageId":"7v39v47hkv.fsf@alter.siamese.dyndns.org","threadId":"24457","inReplyTo":"AANLkTim0A0MAmpgAiaYSgYO=YbZ2gc4Upx3MQQopx6DG@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2010-07-27T21:14:40Z","receivedAt":"2010-07-27T21:14:40Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Avery Pennarun <apenwarr@gmail.com> writes:\n\n> ...  There is one problem I don't know how to\n> solve with this idea, though: what happens when commit A adds a\n> submodule in modules/mod1, commit B removes it, and then commit C\n> re-adds the same submodules in modules/mod1-again?  Will it reuse the\n> same submodule .git directory or a new one?  Share objects or not?\n> Share branch names or not?  Share .git/config or not?\n>\n> Unless you have some kind of \"unique id\" scheme for submodules, this\n> gets impossible to handle correctly.  And the git objects themselves\n> (trees that link to commits) have nowhere to put such things.\n\nI vaguely recall that we already had discussed and more or less resolved\nit at the design level at some point.  Looking for \"three-level thing\" in\nthe gmane archive might be beneficial, although all I recall these three\nwords as search keywords and do not have a detailed recollection of actual\ndiscussion ;-)\n"},{"id":"146532","messageId":"4C4F4C4D.608@web.de","threadId":"24457","inReplyTo":"20100727184047.GC25124@worldvisions.ca","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Jens Lehmann","fromEmail":"jens.lehmann@web.de","sentAt":"2010-07-27T21:14:53Z","receivedAt":"2010-07-27T21:14:53Z","isPatch":false,"sender":{"key":"jens.lehmann@web.de","avatar":"https://avatars.githubusercontent.com/u/135220?v=4"},"body":"Am 27.07.2010 20:40, schrieb Avery Pennarun:\n> With what you're proposing, for all my submodules, we can't each have our\n> own project; we all have to push to the shared one.\n> \n> (Just to be clear: I don't want to fork *every submodule by hand every\n> time*.  I just want *my* stuff to be in *my* repo.  The easiest way to do\n> this would be to have all my changes in a single repo, ie. my fork of the\n> superproject.)\n\nFair enough, but that would not be the Right Thing for my use cases.\n(E.g. I am using submodules to have a single upstream repo for a library\nwhich I use in almost all my projects. And fixes to that library I do in\none of these projects shall be fetchable in all other projects right\nafter I pushed them to the submodules repo, without having to push them\nout of the superprojects repo into the shared one /again/. The situation\nat dayjob is the same and I assume a lot of people are using submodules\nthis way).\n\nSo I would vote for not breaking the *feature* submodules currently have:\nto use a different repo than that used for the superproject. Because that\nenables you to have shared content. I am not against having the /choice/\nto have the submodules objects in the same repo as the superproject, but\nthat should be an option and not mandatory.\n"},{"id":"146542","messageId":"4C4F508A.8050200@web.de","threadId":"24457","inReplyTo":"AANLkTim0A0MAmpgAiaYSgYO=YbZ2gc4Upx3MQQopx6DG@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Jens Lehmann","fromEmail":"jens.lehmann@web.de","sentAt":"2010-07-27T21:32:58Z","receivedAt":"2010-07-27T21:32:58Z","isPatch":false,"sender":{"key":"jens.lehmann@web.de","avatar":"https://avatars.githubusercontent.com/u/135220?v=4"},"body":"Am 27.07.2010 22:57, schrieb Avery Pennarun:\n> One problem with this line of reasoning is that \"--recursive\" is\n> always an option.  But if submodules are ever to be easy to use, I\n> think it should be the default (or settable as a default using git\n> config).  This would take us a *long* way towards usability (of\n> course, in addition to adding the missing features, as you mention).\n\nAnd that is exactly what I am currently doing:\n\n- I already teached diff and status to always recurse (and just\n  sent a patch to add a config option for that behavior, as some\n  users either can't pay the performance costs or don't want to\n  see submodules show up as modified just because they contain\n  untracked files).\n\n- I posted a WIP patch doing recursive checkouts (that is basically\n  working but I still have to put in the safety checks so that no\n  modifications to submodules are accidentally discarded unless -f\n  is used).\n\n- I am working on a recursive fetch too.\n\nAnd then there is other stuff on my list to be tackled; I try to\nfix these issues so that the most annoying problems get solved\nfirst.\n\nUnfortunately that does not proceed as fast as i wished, but\nhopefully I can show some progress in the near future. Of course\nany help would greatly be appreciated ;-)\n"},{"id":"146621","messageId":"4C503262.80702@xiplink.com","threadId":"24457","inReplyTo":"20100727183658.GB25124@worldvisions.ca","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Marc Branchaud","fromEmail":"marcnarc@xiplink.com","sentAt":"2010-07-28T13:36:34Z","receivedAt":"2010-07-28T13:36:34Z","isPatch":false,"sender":{"key":"marcnarc@xiplink.com","avatar":"https://avatars.githubusercontent.com/u/14980203?v=4"},"body":"On 10-07-27 02:36 PM, Avery Pennarun wrote:\n> \n> For \"recursive\" commit, for my own workflow, I would rather have it work\n> like this: from the toplevel, I can 'git commit' any set of files, as long\n> as they all fall inside a particular submodule.  That is, if I do\n> \n> \tgit commit mod1/*.c mod2/*.c\n> \t\n> it should reject it (with a helpful message), because the commit would cross\n> submodule boundaries.  But if I do\n> \n> \tgit commit mod1/*.c\n> \t\n> I think it should create a new commit in mod1, leave my superproject\n> pointing at that new commit, and stop (ie. without the superproject having\n> committed the new commit pointer).\n\nI think that makes perfect sense.  I'd also want the updated pointer to be\nunstaged.\n\n> Why?  Because my normal workflow is:\n> \n>   - make a bunch of superproject/submodule changes until they work.\n>   - commit the submodule changes with a submodule-relevant message\n>   - commit the superproject change with a supermodule-relevant message\n>   \n> I wouldn't want to share commit messages between the two, so actually having\n> a single commit process be \"recursive\" would not do me any good.\n\nThat's the workflow I'd like to follow as well.\n\nIn terms of achieving this workflow with submodules and branching, what's\nrequired is that branching in the superproject takes the submodules off of\nthe detached HEAD and onto something that won't get automatically\ngarbage-collected in a few weeks.\n\nThat could be done simply by applying the superproject's branch to all the\nsubmodules.  A command like\n\n\tsuperproject/$ git branch foo origin/master\n\nwould create the submodule branches on the commits identified for the\nsubmodules in the superproject's origin/master commit.  To make that work\nsmoothly I think requires all the submodules' .git directories, so the branch\nname can be recorded in all of them.\n\nAnd so I think that either \"git fetch\" has to recursively obtain (and update)\nall submodule repos, or there needs to be some kind of on-demand retrieval\nmechanism.  Other ideas for grand-unified object stores (which I haven't been\nfollowing too closely) could work as well.\n\nSo with unified branching and available .git directories, I think a recursive\ncheckout is doable and makes sense.  I'd still like to control which\nsubmodules a checkout might recurse through, but I think the sparse-checkout\nsystem is the way to handle that.\n\nI also suspect that non-fast-forward submodule merges could be workable,\nwhere regular merges are performed individually in the submodules before\nmerging in the superproject.\n\nOne final, somewhat orthogonal thought:  I think that \"git commit\nsubmodule-dir\" should require -f if the remote associated with the submodule\ndoesn't have the commit ID you're trying to commit.\n\n> However, pushing is a separate issue entirely.  Having push be recursive\n> would be easy, but it doesn't solve the *real* problem with pushing: git\n> doesn't know what branch to push to in the submodule, and the submodule most\n> likely isn't pointing at a pushable repo at all, even if the supermodule is. \n> This is why I keep coming back to the idea that I really want to push all\n> the submodule objects into the superproject's repo.\n\nI agree that recursive pushing doesn't make much sense, so there's no need to\ntry to implement it.  I think having \"git commit\" reject unpushed submodule\nupdates in the superproject goes a long way to alleviating misordered pushing.\n\n\t\tM.\n"},{"id":"146648","messageId":"201007282032.42106.jnareb@gmail.com","threadId":"24457","inReplyTo":"20100727183658.GB25124@worldvisions.ca","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2010-07-28T18:32:37Z","receivedAt":"2010-07-28T18:32:37Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Tue, Jul 27, 2010, Avery Pennarun wrote:\n> On Mon, Jul 26, 2010 at 10:56:58AM +0200, Jakub Narebski wrote:\n> > On Sat, Jul 24, 2010, skillzero@gmail.com napisał:\n> > >\n> > > git-submodule might be technically possible in this situation, but\n> > > having to commit and push each submodule and then commit and push the\n> > > super module makes it slightly worse than just dealing with the\n> > > space/download/performance issues of one huge repository.\n> > \n> > But this is just a matter for improving UI for dealing with submodules,\n> > isn't it.   For example having \"git commit --recursive\" would help\n> > with 'having to commit each submodule', though how you would write commit\n> > messages then: perhaps supermodule commit message could be by default\n> > composed out of submodules commits (if any).  \"git push --recursive\"\n> > (or some support for push in \"git remote\") would help with 'having to\n> > push each submodule'.\n> \n> For \"recursive\" commit, for my own workflow, I would rather have it work\n> like this: from the toplevel, I can 'git commit' any set of files, as long\n> as they all fall inside a particular submodule.  That is, if I do\n> \n> \tgit commit mod1/*.c mod2/*.c\n> \t\n> it should reject it (with a helpful message), because the commit would cross\n> submodule boundaries.  But if I do\n> \n> \tgit commit mod1/*.c\n> \t\n> I think it should create a new commit in mod1, leave my superproject\n> pointing at that new commit, and stop (ie. without the superproject having\n> committed the new commit pointer).\n> \n> Why?  Because my normal workflow is:\n> \n>   - make a bunch of superproject/submodule changes until they work.\n>   - commit the submodule changes with a submodule-relevant message\n>   - commit the superproject change with a supermodule-relevant message\n>   \n> I wouldn't want to share commit messages between the two, so actually having\n> a single commit process be \"recursive\" would not do me any good.\n\nI think it is quite good idea, but it covers only one of the three most\ncommon (I think) used versions of git-commit:\n * git commit <files>        # your proposal covers this\n * git commit -a             # but I think either this\n * git commit                # or this is actually more common\n\nAlso \"git commit .\" in a submodule cannot be done in this proposal,\nbecause it is indistinguishable from \"git commit <submodule>\" committing\nstate of submodule in supermodule.\n\nPerhaps it would be matter of porting \"--relative=<path>\" or adding\n\"--submodule=<name>\" option to git-commit?\n\n> However, pushing is a separate issue entirely.  Having push be recursive\n> would be easy, but it doesn't solve the *real* problem with pushing: git\n> doesn't know what branch to push to in the submodule, and the submodule most\n> likely isn't pointing at a pushable repo at all, even if the supermodule is. \n> This is why I keep coming back to the idea that I really want to push all\n> the submodule objects into the superproject's repo.\n\nI think there should be two easy to obtain variants of recursive clone:\n\n1. Current one, where each submodule gets its own repository in the place\n   it is checked out in working area (in worktree) of supermodule.\n\n2. New one, where submodule repositories are in .git/submodules/<name>\n   in supermodule GIT_DIR, and submodules use gitfiles (probably with\n   some notation that path is relative to supermodule, like e.g. //<path>\n   or .../<path>).\n\nI'm not sure though how it would translate into pushing...\n\n-- \nJakub Narebski\nPoland\n"},{"id":"146659","messageId":"201007290027.12443.jnareb@gmail.com","threadId":"24457","inReplyTo":"AANLkTikEo=Qw56WCxkFdmGqQcQoiTsnBy+Dt6zHpkOii@mail.gmail.com","subject":"Re: Avery Pennarun's git-subtree?","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2010-07-28T22:27:06Z","receivedAt":"2010-07-28T22:27:06Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Dnia niedziela 25. lipca 2010 03:47, Nguyen Thai Ngoc Duy napisał:\n> \n> By the way, how hard is it to use git-replace to implement narrow clone?\n\nI don't think that git-replace should be used to implement narrow clone,\nalthough it could probable be abused to do so.  The refs/replaces \nmechanism is about static replacements...\n\n-- \nJakub Narebski\nPoland\n"}]}