{"thread":{"id":"43065","subject":"Re: [RFC] Submodules in GIT","startedAt":"2006-11-28T10:29:09Z","lastAt":"2006-12-17T00:21:10Z","messageCount":160,"participants":["Martin Waitz","Josef Weidendorfer","Torgil Svensson","sf","Shawn Pearce","Linus Torvalds","Michael K. Edwards","Andreas Ericsson","Jakub Narebski","Sam Vilain","R. Steve McKown","Sven Verdoolaege","Uwe Kleine-Koenig","Andy Parkins","Stephan Feder","Jon Loeliger","Daniel Barkalow","Junio C Hamano","Alan Chandler","Steven Grimm"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"294936","messageId":"200611281029.11918.andyparkins@gmail.com","threadId":"43065","inReplyTo":"456C0313.3020308@op5.se","subject":"Re: [RFC] Submodules in GIT","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2006-11-28T10:29:09Z","receivedAt":"2006-11-28T10:29:09Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Tuesday 2006 November 28 09:36, Andreas Ericsson wrote:\n\n> I'd actually prefer the second solution here and let git print a list of\n> submodules with dirty state and ask for some sort of user-response\n> before creating the actual commit. As non-interactive commits should\n> always be clean, requiring user intervention on non-clean state should\n> be a safe thing to do.\n\nI'd agree.  However, is there a need to require user intervention?  Can we not \nmake the following analogies to normal git operation:\n\nfile in working directory -> submodule working directory\nfile in index -> submodule repository\n\nIt's perfectly possible to make a commit with different contents in the index \nand the working directory - it shows up in the git-status output very nicely.  \nWhy not deal with submodules in the same way?\n\nNow imagine the following repository:\n  file1\n  file2\n  submodule1/file3\nMake changes to file2 and file3, but don't update-index or commit.  git-status \nwould show:\n\n# Changed but not updated:\n#   (use git-update-index to mark for commit)\n#\n#       modified:   file2\n#       dirty:      submodule1\n#\nnothing to commit\n\nNow \"git-update-index file2\" and git-status\n\n# Updated but not checked in:\n#   (will commit)\n#\n#       modified:   file2\n#\n# Changed but not updated:\n#   (use git-update-index to mark for commit)\n#\n#       dirty:      submodule1/\n#\n\nNow do a commit in submodule1/ and git-status in the supermodule.\n\n# Updated but not checked in:\n#   (will commit)\n#\n#       modified:   file2\n#       submodule:  submodule1/\n#\n\nObviously the detail would be different, but you get the idea.  There is \nalmost no difference between git-with-submodules and git-as-normal.\n\nI suppose there would actually need to be an extra step were the submodule is \nadded to the supermodule index.  So really there would be three states from \ngit-status:\n\n# Updated but not checked in:\n#   (will commit)\n#\n#       modified:   file2\n#\n# Changed but not updated:\n#   (use git-update-index to mark for commit)\n#\n#       modified:   file1\n#       submodule:  submodule1/\n#\n# Dirty submodules:\n#   (commit changes in the submodule to clean)\n#\n#       dirty:      submodule1/\n#\n\nWhich means: since the last supermodule commit there has been\n * a change to file2, which is in the index and would be committed.\n * a change to file1, which is not in the index and won't be committed.\n * a commit to submodule1, which won't be committed\n * changes to the submodule working directory\n\nThis really reinforces Linus's interpretation that submodules are \ndirectories - they would presumably just get a new object type and be \nreferenced in the tree object.  git-update-index would be blind to dirty \nsubmodules with no new commit, just as git-update-index on an unchanged file \nhas no effect.\n\nHas this question been answered yet?  How does the supermodule know which \nbranch to track in the submodule?  Does it simply track HEAD or when the \nsubmodule is added to the supermodule is it told which branch to track?  I \nsuppose it's got to be HEAD really hasn't it?\n\n\nAndy\n\n-- \nDr Andy Parkins, M Eng (hons), MIEE\n"},{"id":"296114","messageId":"ekh45n$rfc$1@sea.gmane.org","threadId":"43065","inReplyTo":"200611281029.11918.andyparkins@gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-11-28T10:50:10Z","receivedAt":"2006-11-28T10:50:10Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Andy Parkins wrote:\n\n>                                  How does the supermodule know which \n> branch to track in the submodule?  Does it simply track HEAD or when the \n> submodule is added to the supermodule is it told which branch to track?  I \n> suppose it's got to be HEAD really hasn't it?\n\nI think that the proper place for that would be supermodule _index_.\nThe supermodule tree would have commit entry, and the index would have\nsymbolic branch (and perhaps some infor about where to find refs for\nsubmodule).\n\nThis I guess breaks index abstraction slightly, but on the other hand\nallows for tracking non-HEAD branch of submodule, and for submodule to\nnot know about supermodule at all...\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n\n"},{"id":"296601","messageId":"200611281335.38728.andyparkins@gmail.com","threadId":"43065","inReplyTo":"ekh45n$rfc$1@sea.gmane.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2006-11-28T13:35:37Z","receivedAt":"2006-11-28T13:35:37Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Tuesday 2006 November 28 10:50, Jakub Narebski wrote:\n\n> I think that the proper place for that would be supermodule _index_.\n> The supermodule tree would have commit entry, and the index would have\n> symbolic branch (and perhaps some infor about where to find refs for\n> submodule).\n>\n> This I guess breaks index abstraction slightly, but on the other hand\n> allows for tracking non-HEAD branch of submodule, and for submodule to\n> not know about supermodule at all...\n\nThe reason I thought it would have to be HEAD at all times, is to prevent \nsituations where the supermodule commit doesn't reflect the state of the \ncurrent tree.\n\nLet's imagine that we're doing non-HEAD tracking in the supermodule.\n  supermodule\n   +-------- libsubmodule1\n   +-------- libsubmodule2\nSo, you do a \"make\" in supermodule; this of course will call make in each of \nthe submodules.  You test the output and find that it's all working nicely.  \nTime for a supermodule commit.  We want to freeze this working state.  You \ncommit and tag \"supermodule-rc1\"\n\nUnfortunately, during development, you've switched libsubmodule1 to \nbranch \"development\", but supermodule isn't tracking libsubmodule1/HEAD it's \ntracking libsubmodule1/master.  Your supermodule commit doesn't capture a \nsnapshot of the tree you're using.\n\nNow you say to the mailing list \"hey guys, can you test \"supermodule-rc1\"?  \nThey check it out, and find that everything is broken.  Why?  Because what \nyou wanted to check in was libsubmodule@development, but what actually went \nin was libsubmodule@master.\n\nI think I've talked myself into the position where it definitely has to be \nHEAD being tracked in the submodules; anything else is a disaster waiting to \nhappen because commit doesn't check in your current tree.\n\n\n\nAndy\n\n-- \nDr Andy Parkins, M Eng (hons), MIEE\n"},{"id":"295620","messageId":"20061128154434.GD28337@spearce.org","threadId":"43065","inReplyTo":"200611281335.38728.andyparkins@gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-11-28T15:44:34Z","receivedAt":"2006-11-28T15:44:34Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Andy Parkins <andyparkins@gmail.com> wrote:\n> I think I've talked myself into the position where it definitely has to be \n> HEAD being tracked in the submodules; anything else is a disaster waiting to \n> happen because commit doesn't check in your current tree.\n\nYes, but not only that, HEAD is the only thing that fits with the\nrest of the git repository/index/working directory model.\n\nLets review...\n\nWhat's HEAD?  Its the commit which matches the index state as\nclosely as possible, with the only differences being the changes in\nprogress that are being prepared for the next commit (whose parent\nwill be HEAD).  If the index and working directory are both clean\n(no changes) then its also the current content of this directory,\nright?\n\nWhat's the index?  Its what you are about to commit.\n\nWhat's the working directory?  Its the current content, which may\nalso be partially checked out or dirty.\n\nSo HEAD in a submodule is the current content of that submodule.\nTherefore any update-index call on a submodule should load HEAD\n(totally ignoring whatever branch it refers to) into the supermodule\nindex.\n\n-- \n"},{"id":"297329","messageId":"200611281629.08636.andyparkins@gmail.com","threadId":"43065","inReplyTo":"20061128154434.GD28337@spearce.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2006-11-28T16:29:05Z","receivedAt":"2006-11-28T16:29:05Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Tuesday 2006 November 28 15:44, Shawn Pearce wrote:\n\n> So HEAD in a submodule is the current content of that submodule.\n> Therefore any update-index call on a submodule should load HEAD\n> (totally ignoring whatever branch it refers to) into the supermodule\n> index.\n\nI was with you right up until here.\n\nWhy should a submodule do anything to the supermodule?  This is like saying, \nwhen I edit a working tree file, it should automatically call update-index.  \nThe supermodule index should only be updated in response to a manual \nupdate-index (or commit -a I suppose).\n\nWorse, if you allow that to happen, the supermodule can commit a state that \ncannot be retrieved from the submodule's repository.  The ONLY thing a \nsupermodule can record about a submodule is a commit.  Changing the index \ndoesn't create a commit, so it can't change anything in the supermodule.\n\nIf you change the submodule index then that submodule is \"dirty\", this state \nhas no parallel with normal git operation.  The nearest thing is that you've \nchanged a file but not saved it.  Apart from showing the \"dirty\" state in the \nsupermodule's git-status, I don't see that there is anything that the \nsupermodule can do - it can't go around committing in a repository that it \nnot itself.\n\nIMO, it should always be possible to take a submodule and work on it in \nisolation - in an extreme case, by moving it out of the supermodule tree \nentirely.\n\nIn summary, from the supermodule's point of view:\n * A submodule with changed working directory is \"dirty-wd\"\n * A submodule with changed index is \"dirty-idx\" from the supermodule's\n * A submodule with changed HEAD (since the last supermodule commit) \n   is \"changed but not updated\" and can hence be \"update-index\"ed into the\n   supermodule\n * A submodule with changed HEAD that has been added to the supermodule index\n   is \"updated but not checked in\"\n * A submodule with changed HEAD (since the last supermodule update-index) is\n   both \"changed but not updated\" _and_ \"updated but not checked in\", just \n   like any normal file.\n\nWhat's needed then:\n * A way of telling git to treat a particular directory as a submodule instead\n   of a directory\n * git-status gets knowledge of how to check for \"dirty\" submodules\n * git-commit-tree learns about how to store \"submodule\" object types in\n   trees.  The submodule object type will be nothing more than the hash of the\n   current HEAD commit.  (This might be my ignorance, perhaps it's just \n   update-index that needs to know this)\n\nI don't know enough about the plumbing to know if my description above is \nusing the right nomenclature - I'm sure someone will correct me.\n\nIn my head, it would look something like this:\n\n$ mkdir supermodule; cd supermodule\n$ git init-db\n$ git clone proto://host/submodule.git\n$ git add --submodule submodule\n$ git update-index submodule\n$ git commit -m \"Added submodule to supermodule\"\n[ edit submodule ]\n$ git status\nsubmodule is dirty, the working directory has changed\n[ update-index in submodule ]\n$ git status\nsubmodule is dirty, the index has changed\n[ commit in submodule ]\n$ git status\nsubmodule is changed but not updated\n$ git update-index submodule\n$ git status\nsubmodule is updated but not checked in\n$ git commit -m \"Record submodule change in supermodule\"\n\nAm I crazy?\n\n\n\nAndy\n\n-- \nDr Andy Parkins, M Eng (hons), MIEE\n"},{"id":"298680","messageId":"20061128163651.GG28337@spearce.org","threadId":"43065","inReplyTo":"200611281629.08636.andyparkins@gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-11-28T16:36:51Z","receivedAt":"2006-11-28T16:36:51Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Andy Parkins <andyparkins@gmail.com> wrote:\n> On Tuesday 2006 November 28 15:44, Shawn Pearce wrote:\n> \n> > So HEAD in a submodule is the current content of that submodule.\n> > Therefore any update-index call on a submodule should load HEAD\n> > (totally ignoring whatever branch it refers to) into the supermodule\n> > index.\n> \n> I was with you right up until here.\n> \n> Why should a submodule do anything to the supermodule?  This is like saying, \n> when I edit a working tree file, it should automatically call update-index.  \n> The supermodule index should only be updated in response to a manual \n> update-index (or commit -a I suppose).\n\nYou misread my poorly written statement.  :-)\n\nWhat I meant to say was that update-index run in the supermodule\nwould load the submodule content into the supermodule index; much\nas an update-index on a file would load the content of that file\ninto the index.\n \n> IMO, it should always be possible to take a submodule and work on it in \n> isolation - in an extreme case, by moving it out of the supermodule tree \n> entirely.\n\nAside from sharing object directories, yes.\n \n> In summary, from the supermodule's point of view:\n>  * A submodule with changed working directory is \"dirty-wd\"\n>  * A submodule with changed index is \"dirty-idx\" from the supermodule's\n>  * A submodule with changed HEAD (since the last supermodule commit) \n>    is \"changed but not updated\" and can hence be \"update-index\"ed into the\n>    supermodule\n>  * A submodule with changed HEAD that has been added to the supermodule index\n>    is \"updated but not checked in\"\n>  * A submodule with changed HEAD (since the last supermodule update-index) is\n>    both \"changed but not updated\" _and_ \"updated but not checked in\", just \n>    like any normal file.\n> \n> What's needed then:\n>  * A way of telling git to treat a particular directory as a submodule instead\n>    of a directory\n>  * git-status gets knowledge of how to check for \"dirty\" submodules\n>  * git-commit-tree learns about how to store \"submodule\" object types in\n>    trees.  The submodule object type will be nothing more than the hash of the\n>    current HEAD commit.  (This might be my ignorance, perhaps it's just \n>    update-index that needs to know this)\n\nErr, uhm, more like git-write-tree.  git-commit-tree doesn't\ncare about the tree content.  And all of the tree reading code.\nAnd all object traversal code (e.g. rev-list --objects).  Martin\nWaitz's submodule prototype has been working on those details.\nIts non-trivial due to the number of locations affected.\n\n> In my head, it would look something like this:\n> \n> $ mkdir supermodule; cd supermodule\n> $ git init-db\n> $ git clone proto://host/submodule.git\n> $ git add --submodule submodule\n> $ git update-index submodule\n> $ git commit -m \"Added submodule to supermodule\"\n> [ edit submodule ]\n> $ git status\n> submodule is dirty, the working directory has changed\n> [ update-index in submodule ]\n> $ git status\n> submodule is dirty, the index has changed\n> [ commit in submodule ]\n> $ git status\n> submodule is changed but not updated\n> $ git update-index submodule\n> $ git status\n> submodule is updated but not checked in\n> $ git commit -m \"Record submodule change in supermodule\"\n\nYes, exactly my thoughts on the matter.\n \n> Am I crazy?\n\nMaybe, but I'm not a shrink.  Your email looked sane.  :-)\n\n-- \n"},{"id":"296726","messageId":"1164735520.4724.28.camel@cashmere.sps.mot.com","threadId":"43065","inReplyTo":"200611281629.08636.andyparkins@gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Jon Loeliger","fromEmail":"jdl@freescale.com","sentAt":"2006-11-28T17:38:41Z","receivedAt":"2006-11-28T17:38:41Z","isPatch":false,"sender":{"key":"jdl@jdl.com","avatar":"https://gravatar.com/avatar/75ce9a10b151acd2c28ec4ab2136dba7b2ff1634530bd04b155981a749d08a64?d=mp&s=160"},"body":"On Tue, 2006-11-28 at 10:29, Andy Parkins wrote:\n\n> IMO, it should always be possible to take a submodule and work on it in \n> isolation - in an extreme case, by moving it out of the supermodule tree \n> entirely.\n\nThis seems to me to be tantamount to saying something like:\n\n    We need a \"recursively defined git repository\" that\n    is representable as a git repository.\n\nThat is, can the tree object be changed from containing\njust \"blob\" and \"tree\" references to also having a new\n\"git\" reference as well?\n\njdl\n\n"},{"id":"297548","messageId":"456C94E2.6010708@midwinter.com","threadId":"43065","inReplyTo":"200611281335.38728.andyparkins@gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Steven Grimm","fromEmail":"koreth@midwinter.com","sentAt":"2006-11-28T19:58:26Z","receivedAt":"2006-11-28T19:58:26Z","isPatch":false,"sender":{"key":"koreth@midwinter.com","avatar":"https://gravatar.com/avatar/71b4d2e8b62f168bdc9e9205341159e3567003b4f9e2127c617c5fa0a1f5bad2?d=mp&s=160"},"body":"Andy Parkins wrote:\n> Unfortunately, during development, you've switched libsubmodule1 to \n> branch \"development\", but supermodule isn't tracking libsubmodule1/HEAD it's \n> tracking libsubmodule1/master.  Your supermodule commit doesn't capture a \n> snapshot of the tree you're using.\n>   \n\nHow about if the supermodule commit errors out by default if you commit \na different submodule branch than the one you committed the previous \ntime? Require the user to explicitly acknowledge that yes, they want to \ncheck in the contents of \"development\" now, even though the supermodule \nwas tracking \"master\" before.\n\nOtherwise I think you could easily end up with just the opposite \nsituation, where you forget you've checked out \"development\" for a \nmoment to look at something, and end up inadvertently committing a bunch \nof stuff that's not ready for prime time yet. In a standalone git \nsetting, that's no big deal since the commit only updates the current \nbranch and doesn't touch the master branch, but (as I understand the \nproposal) in a supermodule setting you'd actually end up essentially \ndoing a merge between your development branch and the previously \ncommitted master. Or maybe not a merge, but worse, you'd *replace* the \npreviously committed master with what's in your dev branch.\n\nI think wanting to commit a submodule on a different branch than last \ntime is probably not a typical day-to-day use case, so we should make \nsure the user really wants to do it (but allow it if so.)\n\nOn a related note, it would be great from a usability point of view if \nthere were a way to say \"I always want to be on the same branch in all \nsubmodules and the supermodule.\" I think a common scenario will be that \nyou are doing development that touches a couple of different \napplications and your development effort is really a single set of \nchanges even though it happens to cross submodule boundaries. If this \nbranches-in-sync option is turned on, I'd want \"git checkout \ndevelopment\" to check out the development branch in the entire set of \nrepositories.\n\nMore generally, while I 100% agree that it's very useful to be able to \noperate independently on each submodule, I think it's also going to be \ncommon to use submodules to selectively clone different pieces of a \nlarger project. Say your current development effort needs server A, \nlibrary B, and documentation C, and you want to have *just* those pieces \nin your environment. You don't particularly care about the details of \nhow the system has assembled the pieces you want; you want to be able to \nmake your changes and push them when you're done. They are really just \npieces of a larger code base, not independent entities that happen to be \npulled together into a composite workspace temporarily.\n\nFor that use case, I don't want the system to act differently depending \non whether server A and library B are in the same submodule or separate \nones; I want to treat the supermodule as the repository, and the system \nshould take care of the details of managing the submodules. When I do \n\"git commit -a\" I want it to give me one editor to write one commit \ncomment that covers all of the changes I've made, and when I do \"git \ncheckout -b\" I want a new branch to apply across all the files I'm \nworking with.\n\nIt is entirely possible that the above is a matter best left to the \nporcelain layer, and that's fine with me. But I think the Perforce-style \n\"compose a single workspace out of different bits of a larger project\" \nmodel is hugely useful and whatever submodule system Git ends up with, \nit should be able to emulate as much of that feature as possible.\n\n"},{"id":"294052","messageId":"20061128210211.GI28337@spearce.org","threadId":"43065","inReplyTo":"456C94E2.6010708@midwinter.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-11-28T21:02:11Z","receivedAt":"2006-11-28T21:02:11Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Steven Grimm <koreth@midwinter.com> wrote:\n> Andy Parkins wrote:\n> >Unfortunately, during development, you've switched libsubmodule1 to \n> >branch \"development\", but supermodule isn't tracking libsubmodule1/HEAD \n> >it's tracking libsubmodule1/master.  Your supermodule commit doesn't \n> >capture a snapshot of the tree you're using.\n> >  \n>\n> Or maybe not a merge, but worse, you'd *replace* the \n> previously committed master with what's in your dev branch.\n\nRight, you would be replacing the prior branch of that submodule with\nthe new submodule branch.\n\nI think the safety valve you are looking for here is two things:\n\n  * don't automatically update the submodule's HEAD into the\n    supermodule's index.\n\n  * make sure the submodule's HEAD is a fast-forward of the\n    supermodule's index, with a --force option to force it\n\tanyway.\n\nOtherwise the developer just has to know what he/she is doing.\nToday you can put stuff that isn't ready for prime-time into a\nrepository on the wrong branch just by applying the wrong patch,\nor cherry-picking the wrong commit, etc...  the user can (and\nwill) make mistakes.  But they can also easily recover from them\nby rewinding history and redoing it.\n\n> On a related note, it would be great from a usability point of view if \n> there were a way to say \"I always want to be on the same branch in all \n> submodules and the supermodule.\"\n\nThat's not really an issue.\n\nA branch doesn't exist just because you checked-out the branch, or\nbecause you created it.  A branch exists because there were two or\nmore commits (B and C) which use the same parent (A) and two or more\nof those commits survive, e.g. they have refs which point to them\n(directly or indirectly) or they were merged into another commit\nwhich itself survives.\n\nTherefore if the supermodule is on the \"development branch\" the\nsubmodules are also immediately on the same branch, because their\nHEADs are derived from whatever is stored in the supermodule's tree.\nAnd that tree is derived from whatever \"development branch\" means.\n\nReally what you want/need is a special head in the submodule\nwhich acts as the \"branch that corresponds to the supermodule\".\nThis probably should just be a naked SHA1 stored in HEAD, which\nis committable only because a supermodule exists in a higher level\ndirectory.\n\nThe fact that the submodule project has branches *at all* is\ntotally irrelevant once you start to speak about that submodule\nwithin the supermodule, as its the supermodule which determines\nthe branch of the submodule.\n\n> But I think the Perforce-style \n> \"compose a single workspace out of different bits of a larger project\" \n> model is hugely useful\n\nThat's a mess.\n\nYou start to get into weird cases where the directory structure\nexpected by the build process is no longer intact, because the user\nhas sliced it apart in weird ways.  And there's no single version\nwhich corresponds to that workspace as (if I recall correctly)\nyou can pick different tags or branches at will.  I believe that\nClearCase has the same bug.\n\nYou also can't version that now spliced workspace, aside from taking\nthe configuration file and putting that under version control too.\n\nHowever I think the proposal on the table will support that to some\ndegree, in that you can take any version of any repository and embed\nit at any directory of any other repository.  This means you can\nfor example embed the Linux kernel, glibc and gcc projects into\na larger \"embedded device\" repository, but you cannot alter the\nstructure of any of those three projects without making your own\nlocally developed branch of them.  Which is actually the correct\nthing to do as any subslicing of a repository is exactly that:\na locally developed branch of that repository.\n\n-- \n"},{"id":"296861","messageId":"20061129160355.GF18810@admingilde.org","threadId":"43065","inReplyTo":"200611281335.38728.andyparkins@gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-11-29T16:03:56Z","receivedAt":"2006-11-29T16:03:56Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Tue, Nov 28, 2006 at 01:35:37PM +0000, Andy Parkins wrote:\n> The reason I thought it would have to be HEAD at all times, is to prevent \n> situations where the supermodule commit doesn't reflect the state of the \n> current tree.\n\nThe way I wanted to address this is to show in the supermodule\ngit-status that the submodule is using another branch.\nThat way you are warned and can decide not to commit the supermodule.\n\nI implemented tracking of refs/heads/master (not HEAD) without much\nthinking, and only recently began to think about possible problems with\nthis approach.\n\nBut I think it is an important design decision to take, so I'd like to\nhave consensus here.\n\nPro HEAD:\n - update-index on submodule really updates the supermodule index with\n   a commit that resembles the working directory.\nContra HEAD:\n - HEAD is not garanteed to be equal to the working directory anyway,\n   you may have uncommitted changes.\n - when updating the supermodule, you have to take care that your\n   submodules are on the right branch.\n   You might for example have some testing-throwawy branch in one\n   submodule and don't want to merge it with other changes yet.\n\nPro refs/heads/master:\n - the supermodule really tracks one defined branch of development.\n - you can easily overwrite one submodule by changing to another branch,\n   without fearing that changes in the supermodule change anything\n   there.\nContra refs/heads/master:\n - after updating the supermodule, you may not have the correct working\n   directory checked out everywhere, because some submodules may be on a\n   different branch.\n - there is one branch in the submodule which is special to all the other.\n\nI think that most of the disadvantages of refs/heads/master can be\nsolved by printing the above-mentioned warning in git-status when the\nsubmodule is using another branch (similiar to the\nplanned-but-not-implemented warn if the submodule has uncommited\nchanges).\n\nI don't yet know how to cope with tracking HEAD directly, so I'm still\nin favor of tracking refs/heads/master, as already implemented.\n\n-- \nMartin Waitz\n"},{"id":"297694","messageId":"20061129161543.GG18810@admingilde.org","threadId":"43065","inReplyTo":"200611281629.08636.andyparkins@gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-11-29T16:15:43Z","receivedAt":"2006-11-29T16:15:43Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Tue, Nov 28, 2006 at 04:29:05PM +0000, Andy Parkins wrote:\n> In summary, from the supermodule's point of view:\n>  * A submodule with changed working directory is \"dirty-wd\"\n>  * A submodule with changed index is \"dirty-idx\" from the supermodule's\n>  * A submodule with changed HEAD (since the last supermodule commit) \n>    is \"changed but not updated\" and can hence be \"update-index\"ed into the\n>    supermodule\n>  * A submodule with changed HEAD that has been added to the supermodule index\n>    is \"updated but not checked in\"\n>  * A submodule with changed HEAD (since the last supermodule update-index) is\n>    both \"changed but not updated\" _and_ \"updated but not checked in\", just \n>    like any normal file.\n\nwhen tracking refs/heads/master instead of HEAD, you also get:\n   * A submodule where HEAD is not pointing to refs/heads/master is\n     \"dirty-branch\" or something.\n\n\n> What's needed then:\n>  * A way of telling git to treat a particular directory as a submodule instead\n>    of a directory\nThis is handled by creating a GIT repository in this directory.\nMy current implementation needs some more magic by the user to add it to\nthe index, but I plan to change this to the way that GIT repositories\nwill be recognized as possible submodules.\n\n>  * git-status gets knowledge of how to check for \"dirty\" submodules\nThis is on top of my TODO.\n\n>  * git-commit-tree learns about how to store \"submodule\" object types in\n>    trees.  The submodule object type will be nothing more than the hash of the\n>    current HEAD commit.  (This might be my ignorance, perhaps it's just \n>    update-index that needs to know this)\nit's only update-index that has to know this.\nOtherwise it would be implicitly updated and you would never get your\n\"changed but not updated\" status as above.\n\n\n-- \nMartin Waitz\n"},{"id":"296092","messageId":"200611292000.23778.andyparkins@gmail.com","threadId":"43065","inReplyTo":"20061129160355.GF18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2006-11-29T20:00:22Z","receivedAt":"2006-11-29T20:00:22Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Wednesday 2006, November 29 16:03, Martin Waitz wrote:\n\n> The way I wanted to address this is to show in the supermodule\n> git-status that the submodule is using another branch.\n> That way you are warned and can decide not to commit the supermodule.\n\nThe problem I see with tracking a particular branch is that it makes it less \nconvenient to use git's quick-branching features in the submodules.  Let's \nsay I want to try something out quickly in a submodule, I make a branch, \ncommit, commit, \"hmm, looks good, let's snapshot it in the supermodule\", make \na supermodule branch, \"oh no, I've got to tell the supermodule to track the \nnew (but temporary) branch in the submodule do a commit, switch the submodule \nbranch back to master, delete the temporary branch, remember that the \nsupermodule is tracking that branch and tell the supermodule to track \nsomething else instead...  It all seems too complicated to me.\n\n> Pro HEAD:\n>  - update-index on submodule really updates the supermodule index with\n>    a commit that resembles the working directory.\n\nOuch.  Why does the submodule need to update the supermodule index?  That \nshould be done by update-index in the supermodule.   Further, how is the \nsupermodule index going to represent working directory changes in the \nsubmodule?  The only link between the two is a commit hash.  It has to be \nlike that otherwise you haven't made a supermodule-submodule, you've just \nmade one super-repository.  Also, if you don't store submodule commit hashes, \nthen there is no way to guarantee that you're going to be able get back the \nstate of the submodule again.\n\n> Contra HEAD:\n>  - HEAD is not garanteed to be equal to the working directory anyway,\n>    you may have uncommitted changes.\n\nThat's the case for every file in a repository, so isn't really a worry.  It's \nthe equivalent of changing a file and not updating the index - who cares?  As \nlong as update-index tells you that the submodule is dirty and what to do to \nclean it, everything is great.\n\n>  - when updating the supermodule, you have to take care that your\n>    submodules are on the right branch.\n>    You might for example have some testing-throwawy branch in one\n>    submodule and don't want to merge it with other changes yet.\n\nWhat is the \"right\" branch though?  As I said above, if you're tracking one \nbranch in the submodule then you've effectively locked that submodule to that \nbranch for all supermodule uses.  Or you've made yourself a big rod to beat \nyourself with everytime you want to do some development on an \"off\" branch on \nthe submodule.\n\n> Pro refs/heads/master:\n>  - the supermodule really tracks one defined branch of development.\n\nWhy is this a pro?\n\n>  - you can easily overwrite one submodule by changing to another branch,\n>    without fearing that changes in the supermodule change anything\n>    there.\n\nYou can always do that anyway by simply not running update-index for the \nsubmodule in the supermodule.\n\n> Contra refs/heads/master:\n>  - after updating the supermodule, you may not have the correct working\n>    directory checked out everywhere, because some submodules may be on a\n>    different branch.\n\nThis seems like the biggest problem to me - doesn't this negate all the \nadvantages of a submodule system?  After a check in, you have no idea if what \nyou checked in was what was in your working tree.\n\n\nAndy\n\n-- \nDr Andrew Parkins, M Eng (Hons), AMIEE\n"},{"id":"294248","messageId":"456EC738.6000103@b-i-t.de","threadId":"43065","inReplyTo":"200611281629.08636.andyparkins@gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"sf","fromEmail":"sf@b-i-t.de","sentAt":"2006-11-30T11:57:44Z","receivedAt":"2006-11-30T11:57:44Z","isPatch":false,"sender":{"key":"sf@b-i-t.de","avatar":null},"body":"Andy Parkins wrote:\n...\n> Worse, if you allow that to happen, the supermodule can commit a state that \n> cannot be retrieved from the submodule's repository.  The ONLY thing a \n> supermodule can record about a submodule is a commit.\n\nSo what? You have a submodule commit that only exists in the \nsupermodule. I fail to see the problem. The changes you made to the \nsubmodule _in the supermodule_ can later be pulled from wherever you want.\n\nRegards\n\nStephan\n"},{"id":"298339","messageId":"456ECBA5.7010409@op5.se","threadId":"43065","inReplyTo":"200611292000.23778.andyparkins@gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-11-30T12:16:37Z","receivedAt":"2006-11-30T12:16:37Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Andy Parkins wrote:\n> On Wednesday 2006, November 29 16:03, Martin Waitz wrote:\n> \n  >>  - when updating the supermodule, you have to take care that your\n>>    submodules are on the right branch.\n>>    You might for example have some testing-throwawy branch in one\n>>    submodule and don't want to merge it with other changes yet.\n> \n> What is the \"right\" branch though?  As I said above, if you're tracking one \n> branch in the submodule then you've effectively locked that submodule to that \n> branch for all supermodule uses.  Or you've made yourself a big rod to beat \n> yourself with everytime you want to do some development on an \"off\" branch on \n> the submodule.\n> \n\nPerhaps I'm just daft, but I fail to see how you can safely track a \ntopic-branch that might get rewinded or rebased in the submodule without \ncrippling the supermodule. Wasn't the intention that the supermodule has \na new tree object (called \"submodule\") that points to a commit in the \nsubmodule from where it gets its tree and stuff? Is the intention that \nthe supermodule pulls all of the submodules history into its own ODB? If \nso, what's the difference between just having one large repository. If \nnot, how can you make it not break in case the commit it references in \nthe submodule is pruned away?\n\nOne possible way would ofcourse be to add something like this to the \nsupermodule commit:\nsubmodule directory/commit-sha1\ntree submodule-tree-sha1\n\nbut then you're in trouble because the supermodule will have the same \nfiles as all the submodules stored in its own tree. I'm confused. Could \nsomeone shed some light on how this sub-/super-module connection is \nsupposed to work in the supermodule's commit objects?\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\n"},{"id":"295857","messageId":"200611301240.22938.andyparkins@gmail.com","threadId":"43065","inReplyTo":"456ECBA5.7010409@op5.se","subject":"Re: [RFC] Submodules in GIT","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2006-11-30T12:40:17Z","receivedAt":"2006-11-30T12:40:17Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Thursday 2006 November 30 12:16, Andreas Ericsson wrote:\n\n> > What is the \"right\" branch though?  As I said above, if you're tracking\n> > one branch in the submodule then you've effectively locked that submodule\n> > to that branch for all supermodule uses.  Or you've made yourself a big\n> > rod to beat yourself with everytime you want to do some development on an\n> > \"off\" branch on the submodule.\n>\n> Perhaps I'm just daft, but I fail to see how you can safely track a\n> topic-branch that might get rewinded or rebased in the submodule without\n> crippling the supermodule. Wasn't the intention that the supermodule has\n\nWho said anything but rebase/rewind?  As it happens though, I don't see why \nyou can't (it wouldn't be pleasant though).  A rebase or rewind still leaves \nthe original commit in the object database, so provided no one runs \ngit-prune, there is no catastrophic failure.\n\n> a new tree object (called \"submodule\") that points to a commit in the\n> submodule from where it gets its tree and stuff? Is the intention that\n> the supermodule pulls all of the submodules history into its own ODB? If\n> so, what's the difference between just having one large repository. If\n> not, how can you make it not break in case the commit it references in\n> the submodule is pruned away?\n\nI certainly never suggested anything /but/ storing a submodule type that \nstores the commit.  The current debate is about whether the supermodule \nshould track HEAD or some defined branch in the submodule.\n\n> but then you're in trouble because the supermodule will have the same\n> files as all the submodules stored in its own tree. I'm confused. Could\n> someone shed some light on how this sub-/super-module connection is\n> supposed to work in the supermodule's commit objects?\n\nI don't really know, I only joined in to stand up against commit in the \nsupermodule triggering commits in the submodule.  That lead to me trying to \nget an understanding of how it would work.\n\nAs far as I can see, the only way a submodule is any use is if it is always a \nsubmodule-commit-hash that is noted in the supermodule tree object.  That \nmeans that the supermodule will only commit clean submodules.  The rest is \njust UI to show something useful in the difficult cases when the submodule \ntree is dirty.\n\n\nAndy\n-- \nDr Andy Parkins, M Eng (hons), MIEE\n"},{"id":"297582","messageId":"20061130170625.GH18810@admingilde.org","threadId":"43065","inReplyTo":"200611292000.23778.andyparkins@gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-11-30T17:06:25Z","receivedAt":"2006-11-30T17:06:25Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Wed, Nov 29, 2006 at 08:00:22PM +0000, Andy Parkins wrote:\n> On Wednesday 2006, November 29 16:03, Martin Waitz wrote:\n> \n> > The way I wanted to address this is to show in the supermodule\n> > git-status that the submodule is using another branch.\n> > That way you are warned and can decide not to commit the supermodule.\n> \n> The problem I see with tracking a particular branch is that it makes it less \n> convenient to use git's quick-branching features in the submodules.  Let's \n> say I want to try something out quickly in a submodule, I make a branch, \n> commit, commit, \"hmm, looks good, let's snapshot it in the supermodule\", make \n> a supermodule branch, \"oh no, I've got to tell the supermodule to track the \n> new (but temporary) branch in the submodule do a commit, switch the submodule \n> branch back to master, delete the temporary branch, remember that the \n> supermodule is tracking that branch and tell the supermodule to track \n> something else instead...  It all seems too complicated to me.\n\nWhat about:\nYou decide to try something out quickly and create a new branch in the\nsubmodule. After you have verified that it works, you merge it to the\nsubmodules master branch and commit that to the supermodule.\nNot that complicated, isn't it?\nIn fact, my current implementation does not even allow to change the\nbranch name of the submodule which is tracked by the supermodule ;-).\n\n> > Pro HEAD:\n> >  - update-index on submodule really updates the supermodule index with\n> >    a commit that resembles the working directory.\n> \n> Ouch.  Why does the submodule need to update the supermodule index?\n\nPlease excuse that I am not an native english speaker and I may have\ncaused some confusion here.\n\n> That should be done by update-index in the supermodule.\n\nThat is exactly what I wanted to say. In the supermoduel you call\nupdate-index (with the submodule path as argument) to update the index\nof the supermodule. Just like normal files. Nothing new.\n\n> Further, how is the supermodule index going to represent working\n> directory changes in the submodule?  The only link between the two is\n> a commit hash.  It has to be like that otherwise you haven't made a\n> supermodule-submodule, you've just made one super-repository.  Also,\n> if you don't store submodule commit hashes, then there is no way to\n> guarantee that you're going to be able get back the state of the\n> submodule again.\n\nThis is handled in the next paragraph.\nThe argument really is: HEAD always points to the checked out branch,\nso it really has a relationship to the working directory.\n\n> > Contra HEAD:\n> >  - HEAD is not garanteed to be equal to the working directory anyway,\n> >    you may have uncommitted changes.\n> \n> That's the case for every file in a repository, so isn't really a\n> worry.  It's the equivalent of changing a file and not updating the\n> index - who cares?  As long as update-index tells you that the\n> submodule is dirty and what to do to clean it, everything is great.\n\nYes, it's not a real counter-argument, but it relativates the previous\npro-argument.\n\n> >  - when updating the supermodule, you have to take care that your\n> >    submodules are on the right branch.\n> >    You might for example have some testing-throwawy branch in one\n> >    submodule and don't want to merge it with other changes yet.\n> \n> What is the \"right\" branch though?  As I said above, if you're tracking one \n> branch in the submodule then you've effectively locked that submodule to that \n> branch for all supermodule uses.\n\nyes, but luckily GIT branches are very flexible.\n\n> Or you've made yourself a big rod to beat yourself with everytime you\n> want to do some development on an \"off\" branch on the submodule.\n\nI don't think it is that bad.\n\n> > Pro refs/heads/master:\n> >  - the supermodule really tracks one defined branch of development.\n> \n> Why is this a pro?\n\nYou always know which branch in the submodule is the \"upstream\" branch\nwhich is managed by the supermodule.\nYou can easily have several topic-branches and merge updates from the\nmaster branch.\notherwise you always have to remember which branch holds your current\ncontents from the supermodule.\n\nWhen viewed from the supermodule, you are storing one branch per\nsubmodule in your tree.\n\n> >  - you can easily overwrite one submodule by changing to another branch,\n> >    without fearing that changes in the supermodule change anything\n> >    there.\n> \n> You can always do that anyway by simply not running update-index for the \n> submodule in the supermodule.\n\nSuppose you are working on a complicated feature in one submodule.\nYou create your own branch for that feature and work on it.\nNow you want to update your project, so you pull a new supermodule\nversion. Now this pull also included one (for you unimportant) change\nin the submodule.\nI think it is more clear to update the master branch with the new\nversion coming from the supermodule, while leaving your work intact\n(you haven't commited it to the supermodule yet, so the supermodule\nshould not care about your changes, it's just some dirty tree).\nThen you can freely merge between your branch and master as you like and\nare not forced to merge at once. And perhaps you even do not want to\nmerge at all, because you are on an experimental branch which really is\nmutually exclusive with the current supermodule contents.\n\n> > Contra refs/heads/master:\n> >  - after updating the supermodule, you may not have the correct working\n> >    directory checked out everywhere, because some submodules may be on a\n> >    different branch.\n> \n> This seems like the biggest problem to me - doesn't this negate all the \n> advantages of a submodule system?  After a check in, you have no idea if what \n> you checked in was what was in your working tree.\n\nOf course you know: git-status will tell it.\nThis is no different to today, where you can commit while still leaving\na part of the tree dirty.\n\n-- \nMartin Waitz\n"},{"id":"297799","messageId":"456F29A2.1050205@op5.se","threadId":"43065","inReplyTo":"20061130170625.GH18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-11-30T18:57:38Z","receivedAt":"2006-11-30T18:57:38Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Martin Waitz wrote:\n> hoi :)\n> \n> On Wed, Nov 29, 2006 at 08:00:22PM +0000, Andy Parkins wrote:\n>> On Wednesday 2006, November 29 16:03, Martin Waitz wrote:\n>>\n>> Further, how is the supermodule index going to represent working\n>> directory changes in the submodule?  The only link between the two is\n>> a commit hash.  It has to be like that otherwise you haven't made a\n>> supermodule-submodule, you've just made one super-repository.  Also,\n>> if you don't store submodule commit hashes, then there is no way to\n>> guarantee that you're going to be able get back the state of the\n>> submodule again.\n> \n> This is handled in the next paragraph.\n> The argument really is: HEAD always points to the checked out branch,\n> so it really has a relationship to the working directory.\n> \n>>> Contra HEAD:\n>>>  - HEAD is not garanteed to be equal to the working directory anyway,\n>>>    you may have uncommitted changes.\n>> That's the case for every file in a repository, so isn't really a\n>> worry.  It's the equivalent of changing a file and not updating the\n>> index - who cares?  As long as update-index tells you that the\n>> submodule is dirty and what to do to clean it, everything is great.\n> \n> Yes, it's not a real counter-argument, but it relativates the previous\n> pro-argument.\n> \n>>>  - when updating the supermodule, you have to take care that your\n>>>    submodules are on the right branch.\n>>>    You might for example have some testing-throwawy branch in one\n>>>    submodule and don't want to merge it with other changes yet.\n>> What is the \"right\" branch though?  As I said above, if you're tracking one \n>> branch in the submodule then you've effectively locked that submodule to that \n>> branch for all supermodule uses.\n> \n> yes, but luckily GIT branches are very flexible.\n> \n\nThere's no real technical reason for locking it to a single branch \nthough, and in case of a fork in the upstream submodule project, you \nmight suddenly decide that \"the other team\" is heading in a much more \ninteresting direction and you want to use their work in your module \ninstead. Will you now have to maintain a separate branch just to keep \nthe same name as the branch the original team used?\n\n>> Or you've made yourself a big rod to beat yourself with everytime you\n>> want to do some development on an \"off\" branch on the submodule.\n> \n> I don't think it is that bad.\n> \n\nIt could be, and as has already been stated, there's no real reason to \nlimit this to a particular branch, so I don't see why we would want to \nimpose such non-real restrictions.\n\n>>> Pro refs/heads/master:\n>>>  - the supermodule really tracks one defined branch of development.\n>> Why is this a pro?\n> \n> You always know which branch in the submodule is the \"upstream\" branch\n> which is managed by the supermodule.\n\nNo you don't. The branch-name might be moved to some other tip of the \nDAG, and that's exactly the same as changing the branch you're tracking.\n\n> You can easily have several topic-branches and merge updates from the\n> master branch.\n> otherwise you always have to remember which branch holds your current\n> contents from the supermodule.\n> \n\nNo you don't. The only thing you need is the commit-sha.\n\n> When viewed from the supermodule, you are storing one branch per\n> submodule in your tree.\n> \n\nWrong again. You're storing one particular point in the revision history.\n\n>>>  - you can easily overwrite one submodule by changing to another branch,\n>>>    without fearing that changes in the supermodule change anything\n>>>    there.\n>> You can always do that anyway by simply not running update-index for the \n>> submodule in the supermodule.\n> \n> Suppose you are working on a complicated feature in one submodule.\n> You create your own branch for that feature and work on it.\n> Now you want to update your project, so you pull a new supermodule\n> version. Now this pull also included one (for you unimportant) change\n> in the submodule.\n\ngit reset to the rescue.\n\n> I think it is more clear to update the master branch with the new\n> version coming from the supermodule, while leaving your work intact\n> (you haven't commited it to the supermodule yet, so the supermodule\n> should not care about your changes, it's just some dirty tree).\n> Then you can freely merge between your branch and master as you like and\n> are not forced to merge at once. And perhaps you even do not want to\n> merge at all, because you are on an experimental branch which really is\n> mutually exclusive with the current supermodule contents.\n> \n\nThis is all just policy though. Tools that enforce a certain policy are \nnot good tools.\n\nThe only problem I'm seeing atm is that the supermodule somehow has to \nmark whatever commits it's using from the submodule inside the submodule \nrepo so that they effectively become un-prunable, otherwise the \nsupermodule may some day find itself with a history that it can't restore.\n\nThe really major problem with this is that now you'll have one \nrepository of the submodule that is actually special, so it's not \ncertain you can go and use any repository at all of the submodule code, \nsince the upstream repo most likely won't be all that interested in \nhaving all of that meta-data inside it. In reality, I'm sure this will \nbe a small problem though, as submodules that are in reality projects \nwhich the supermodule's maintainer isn't the owner of will most likely \nnever rewind their history beyond the supermodules stored commit. It's \nsomething fsck will have to be taught to watch for though.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"296926","messageId":"200612010902.51264.andyparkins@gmail.com","threadId":"43065","inReplyTo":"20061130170625.GH18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2006-12-01T09:02:48Z","receivedAt":"2006-12-01T09:02:48Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Thursday 2006 November 30 17:06, Martin Waitz wrote:\n\n> You can easily have several topic-branches and merge updates from the\n> master branch.\n> otherwise you always have to remember which branch holds your current\n> contents from the supermodule.\n\nWHAT?  I've got to make merges (that I don't necessarily want) in order to \ncommit in the supermodule?  This completely negates any useful functioning of \nbranches in the submodule.  I want to be able to make a quick development \nbranch in the submodule and NOT merge that code into master and then be able \nto still commit that in the supermodule.\n\nI think you're imagining the binding between the super and sub is very much \ntighter than it should be.  What if I'm working on a development version of \nthe supermodule, which includes a stable version of the submodule?  Vice \nversa?\n\n> When viewed from the supermodule, you are storing one branch per\n> submodule in your tree.\n\nThat prevents me \"trying something out\" on a topic branch in the submodule.  \nHere's a scenario using my suggested \"supermodule tracks submodule HEAD\" \nmethod.\n\n * You're developerA\n * Make a development branch in the supermodule\n * In the submodule, make a whole load of topic branches\n * Make a development branch in the submodule\n * Merge the topic branches into the development branch of the submodule\n * Commit in the supermodule.  This capture\n * Tag that commit \"my-tested-arrangement-of-submodule-features\"\n * Push that tag to the central repository - tell the world.\n * DeveloperB checks out that tag and tries it.  Great stuff.\n\nNow: here's the secret fact that I didn't tell you that will break \nyour \"supermodule tracks submodule branch\" method.  DeveloperB has decided to \nhave this in his remote:\n  Pull: refs/heads/master:refs/heads/upstream/master\nOops. The supermodule, which has been told to track the \"master\" branch in the \nsubmodule is tracking different things in developerA's repository from \ndeveloperB's repository.  Worse, what if developerB did this:\n  Pull: refs/heads/master:refs/heads/development\n  Pull: refs/heads/development:refs/heads/master\n\nBranches are completely arbitrary per-repository.  You cannot rely on them \nbeing consistent between different repositories.  If you store the name of a \nsubmodule branch in a supermodule - that supermodule is only valid for that \none special case of your particular version of the submodule.\n\n\nAndy\n-- \nDr Andy Parkins, M Eng (hons), MIEE\n"},{"id":"297288","messageId":"20061201110032.GL18810@admingilde.org","threadId":"43065","inReplyTo":"200612010902.51264.andyparkins@gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-01T11:00:32Z","receivedAt":"2006-12-01T11:00:32Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Fri, Dec 01, 2006 at 09:02:48AM +0000, Andy Parkins wrote:\n> On Thursday 2006 November 30 17:06, Martin Waitz wrote:\n> \n> > You can easily have several topic-branches and merge updates from the\n> > master branch.\n> > otherwise you always have to remember which branch holds your current\n> > contents from the supermodule.\n> \n> WHAT?  I've got to make merges (that I don't necessarily want) in\n> order to commit in the supermodule?  This completely negates any\n> useful functioning of branches in the submodule.  I want to be able to\n> make a quick development branch in the submodule and NOT merge that\n> code into master and then be able to still commit that in the\n> supermodule.\n\nexactly!\n\nPlease think about it.\n\nIf you track HEAD, then this means that you track HEAD.\nIn _both_ directions!\n\nSo you not only store your submodule HEAD commit in the supermodule when you\ndo commit to the supermodule, it also means that your submodule HEAD\nwill be updated when you update your supermodule.\nAnd what happens if you already commited something to HEAD in the mean\ntime? Exactly: a merge is needed.\n\nAnd you are right: you might not want to do this now, because you\nbranched off, because you _wanted_ to have some development which is\n_independent_ to the current supermodule work.\n\nSo tracking HEAD really makes branching in the submodule hard to work\nwith.\n\nWhat does the supermodule provide to the submodule? It stores one\nreference to a commit sha1. Just like a reference inside refs/heads\ninside the submodule. There really is not much difference between the\nsha1 stored inside the supermodules tree and one stored inside refs/.\nSo from the submodules point of view, the supermodule is not much more\nthen one special branch.\nBut it is not possible to use the supermodule index directly as one\n\"magic\" branch for several reasons.\nSo we need synchronization methods between the index entry for the\nsubmodule which is stored in the supermodule and the references in the\nsubmodule. These are git-update-index/git-commit and git-checkout, both\ncalled explicitly or implicitly in the supermodule.\nAnd I really think it makes sense to have a one-to-one relationship\nbetween the submodule \"branch\" stored in the supermodule and the\nbranchname used in the submodule.\n\n> I think you're imagining the binding between the super and sub is very much \n> tighter than it should be.  What if I'm working on a development version of \n> the supermodule, which includes a stable version of the submodule?  Vice \n> versa?\n\nI don't see your problem here.\n\n> > When viewed from the supermodule, you are storing one branch per\n> > submodule in your tree.\n> \n> That prevents me \"trying something out\" on a topic branch in the submodule.  \n> Here's a scenario using my suggested \"supermodule tracks submodule HEAD\" \n> method.\n> \n>  * You're developerA\n>  * Make a development branch in the supermodule\n>  * In the submodule, make a whole load of topic branches\n>  * Make a development branch in the submodule\n>  * Merge the topic branches into the development branch of the submodule\n>  * Commit in the supermodule.  This capture\n>  * Tag that commit \"my-tested-arrangement-of-submodule-features\"\n>  * Push that tag to the central repository - tell the world.\n>  * DeveloperB checks out that tag and tries it.  Great stuff.\n\nThis is still supposed to be a distributed system.\nDeveloperB does not only check out the whole project including several\nmodules. He is also supposed to _work_ with it.\n\nWhat if DeveloperB also has several topic branches?\nWhen he checks out the new supermodule, only his current HEAD in the\nsubmodule will be updated.\nSo he first has to change to some supermodule-tracking branch inside the\nsubmodule, then pull the supermodule updates, then eventually merge the\nnew contents of his supermodule-tracking branch into his topic branches.\nSo why not make this \"let's update one supermodule-tracking-branch\"\nautomatic?\n\n> Now: here's the secret fact that I didn't tell you that will break\n> your \"supermodule tracks submodule branch\" method.  DeveloperB has\n> decided to have this in his remote:\n>   Pull: refs/heads/master:refs/heads/upstream/master\n> Oops. The supermodule, which has been told to track the \"master\"\n> branch in the submodule is tracking different things in developerA's\n> repository from developerB's repository.\n\nSo what? He can do to the repository whatever he wants?\nHe wants to change one submodule to a different branch?\nHe can do so!\nBut please do not expect the system to magically be able to resolve\nproblems. If you _by intent_ changed the submodule to another branch\nwhich is incompatible to the one used in the submodule you can't expect\nthat this is magically merged.\nThis is the same as with normal files.\nSure you can replace one file with new contents that are different to\nthe one used by someone else.  Don't expect this can be merged\nautomatically. So now you have two forks/branches of the project.\nSo what?\n\nSame for a system including submodules:\nIf you change one submodule to a totally different branch, then you\neffectivley forked/branched the entire project.\n(Nomenclature is a bit difficult here: what I mean by totally different\nbranch is: the submodule commit tracked by the supermodule is not\ndirectly connected to the one tracked by an old version of the\nsupermodule).\n\nSo whenever you introduce conflicting changes somewhere in the project\n(be it in a submodule or in a file) you _always_ fork/branch the entire\nproject (i.e. the topmost supermodule).\nYou can't circumvent that.\n\nSo what are submodule branches good for then?\nTo store other lines of development which are not yet / not any more\ntracked by the supermodule.  Perhaps you store references to branches\nstored in another supermodule, or another standalone repository.\nOr a temporary branch which is only used for testing.\nThere are really many possiblilities.\nBut they all have one thing in common: they are not meant to be tracked\nby the supermodule.\n\n> Branches are completely arbitrary per-repository.\n\nYes, but a submodule is special here: it really has one special branch.\nThe module is not independent any more.  That is the _nature_ of a\nsubmodule.\n\n> You cannot rely on them being consistent between different\n> repositories.\n\nSure, we are in a distributed system.\nBut the supermodule always has to know which branch in the submodule has\nto be tracked.  The easiest thing is to always use the default\nrefs/heads/master.  Surely this could be changed if there is a need.\n\n-- \nMartin Waitz\n"},{"id":"296102","messageId":"45701A24.5060500@b-i-t.de","threadId":"43065","inReplyTo":"456F29A2.1050205@op5.se","subject":"Re: [RFC] Submodules in GIT","fromName":"sf","fromEmail":"sf@b-i-t.de","sentAt":"2006-12-01T12:03:48Z","receivedAt":"2006-12-01T12:03:48Z","isPatch":false,"sender":{"key":"sf@b-i-t.de","avatar":null},"body":"Andreas Ericsson wrote:\n...\n> The only problem I'm seeing atm is that the supermodule somehow has to \n> mark whatever commits it's using from the submodule inside the submodule \n> repo so that they effectively become un-prunable, otherwise the \n> supermodule may some day find itself with a history that it can't restore.\n\nThat has nothing to do with submodules. What you state here is the \nproblem of alternate repositories.\n\nThere are two solutions:\n\n1. Do not use alternates.\n\n2. Do not prune a repository that is used as an alternate repository by \nother repositories.\n\nFor the submodule discussion that would mean:\n\n1. Only fetch and work on branches of submodules you are interested in. \nIt does not matter that the origin repository contains (probably orders \nof magnitude) more data. You will never touch that.\n\n2. You can never prune the main (the supermodule's) repository, at least \nnot with what git provides today.\n\nThat is why the sanest approach to subprojects is to put commits into \ntree objects, define a way to name these commits and make git understand \nthese new commit names. Done. Works.\n\nRegards\n\nStephan\n"},{"id":"295290","messageId":"45701B8D.1030508@b-i-t.de","threadId":"43065","inReplyTo":"20061201110032.GL18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"sf","fromEmail":"sf@b-i-t.de","sentAt":"2006-12-01T12:09:49Z","receivedAt":"2006-12-01T12:09:49Z","isPatch":false,"sender":{"key":"sf@b-i-t.de","avatar":null},"body":"Martin Waitz wrote:\n...\n> So you not only store your submodule HEAD commit in the supermodule when you\n> do commit to the supermodule, it also means that your submodule HEAD\n> will be updated when you update your supermodule.\n\nWhy the magic? The typical workflow in git is\n\n1. You work on a branch, i.e. edit and commit and so on.\n2. At some point, you decide to share the work you did on that branch \n(e-mail a patch, merge into another branch, push upstream or let it by \npulled by upstream)\n\nI fail to understand why these two steps have to be mixed up. Someone \ncare to explain?\n\nRegards\n\nStephan\n"},{"id":"293984","messageId":"20061201121110.GP18810@admingilde.org","threadId":"43065","inReplyTo":"45701A24.5060500@b-i-t.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-01T12:11:10Z","receivedAt":"2006-12-01T12:11:10Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Fri, Dec 01, 2006 at 01:03:48PM +0100, sf wrote:\n> Andreas Ericsson wrote:\n> 2. You can never prune the main (the supermodule's) repository, at least \n> not with what git provides today.\n\nIt even already works (well, not with what git provides today, but with\nmy implementation). git-prune simply walks all the submodules, too, when\ndoing it's reachability analysis.\n\nWhat does not work is a prune inside the submodule, because it does not\nknow about all the commits used by the supermodule.\n\n-- \nMartin Waitz\n"},{"id":"297532","messageId":"20061201121234.GQ18810@admingilde.org","threadId":"43065","inReplyTo":"45701B8D.1030508@b-i-t.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-01T12:12:34Z","receivedAt":"2006-12-01T12:12:34Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Fri, Dec 01, 2006 at 01:09:49PM +0100, sf wrote:\n> Martin Waitz wrote:\n> ...\n> >So you not only store your submodule HEAD commit in the supermodule when \n> >you\n> >do commit to the supermodule, it also means that your submodule HEAD\n> >will be updated when you update your supermodule.\n> \n> Why the magic? The typical workflow in git is\n> \n> 1. You work on a branch, i.e. edit and commit and so on.\n> 2. At some point, you decide to share the work you did on that branch \n> (e-mail a patch, merge into another branch, push upstream or let it by \n> pulled by upstream)\n\n3. Other people want to use your new work.\n\n-- \nMartin Waitz\n"},{"id":"296656","messageId":"4570289D.9050802@b-i-t.de","threadId":"43065","inReplyTo":"20061201121234.GQ18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"sf","fromEmail":"sf@b-i-t.de","sentAt":"2006-12-01T13:05:33Z","receivedAt":"2006-12-01T13:05:33Z","isPatch":false,"sender":{"key":"sf@b-i-t.de","avatar":null},"body":"Martin Waitz wrote:\n> hoi :)\n> \n> On Fri, Dec 01, 2006 at 01:09:49PM +0100, sf wrote:\n>> Martin Waitz wrote:\n>> ...\n>> >So you not only store your submodule HEAD commit in the supermodule when \n>> >you\n>> >do commit to the supermodule, it also means that your submodule HEAD\n>> >will be updated when you update your supermodule.\n>> \n>> Why the magic? The typical workflow in git is\n>> \n>> 1. You work on a branch, i.e. edit and commit and so on.\n>> 2. At some point, you decide to share the work you did on that branch \n>> (e-mail a patch, merge into another branch, push upstream or let it by \n>> pulled by upstream)\n> \n> 3. Other people want to use your new work.\n\nSorry, if that was not obvious: You actually procceed with one of the \noptions I listed in Step 2. What I wanted to state is that with git you \ndo not mix up committing (which is local to your repository and your \nbranch) and publishing.\n\nRegards\n\nStephan\n"},{"id":"297716","messageId":"45702C50.9050307@b-i-t.de","threadId":"43065","inReplyTo":"20061201121110.GP18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"sf","fromEmail":"sf@b-i-t.de","sentAt":"2006-12-01T13:21:20Z","receivedAt":"2006-12-01T13:21:20Z","isPatch":false,"sender":{"key":"sf@b-i-t.de","avatar":null},"body":"Martin Waitz wrote:\n> hoi :)\n> \n> On Fri, Dec 01, 2006 at 01:03:48PM +0100, sf wrote:\n>> Andreas Ericsson wrote:\n>> 2. You can never prune the main (the supermodule's) repository, at least \n>> not with what git provides today.\n> \n> It even already works (well, not with what git provides today, but with\n> my implementation). git-prune simply walks all the submodules, too, when\n> doing it's reachability analysis.\n> \n> What does not work is a prune inside the submodule, because it does not\n> know about all the commits used by the supermodule.\n\nI just had a short (really short) look at your work. My impression is \nthat your repository setup is much too complicated.\n\nAs I proposed elsewhere: For submodules to work you only need to allow \ncommits in tree objects (that is what your implementation requires as \nwell). Everything else is in the tools. Much simpler.\n\nRegards\n\nStephan\n"},{"id":"295190","messageId":"20061201133558.GU18810@admingilde.org","threadId":"43065","inReplyTo":"4570289D.9050802@b-i-t.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-01T13:35:58Z","receivedAt":"2006-12-01T13:35:58Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Fri, Dec 01, 2006 at 02:05:33PM +0100, sf wrote:\n> >On Fri, Dec 01, 2006 at 01:09:49PM +0100, sf wrote:\n> >>Martin Waitz wrote:\n> >>>So you not only store your submodule HEAD commit in the supermodule\n> >>>when you do commit to the supermodule, it also means that your\n> >>>submodule HEAD will be updated when you update your supermodule.\n> >>\n> >>Why the magic? The typical workflow in git is\n> >>\n> >>1. You work on a branch, i.e. edit and commit and so on.\n> >>2. At some point, you decide to share the work you did on that branch \n> >>(e-mail a patch, merge into another branch, push upstream or let it by \n> >>pulled by upstream)\n> >\n> >3. Other people want to use your new work.\n> \n> Sorry, if that was not obvious: You actually procceed with one of the \n> options I listed in Step 2. What I wanted to state is that with git you \n> do not mix up committing (which is local to your repository and your \n> branch) and publishing.\n\nI guess you are refering to not mix up committing to the submodule and\nupdating the supermodule index.\nThese are really two separate steps, I just combined them above because I\nwanted to put emphasis on the other part: it is not a one-way flow, it\nis bidirectional, so your HEAD would have to changed if the supermodule\ngets updated.\nAnd I consider changing HEAD, without looking at the branch it points\nto, to be a bad thing.\n\n-- \nMartin Waitz\n"},{"id":"294800","messageId":"20061201134311.GV18810@admingilde.org","threadId":"43065","inReplyTo":"45702C50.9050307@b-i-t.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-01T13:43:11Z","receivedAt":"2006-12-01T13:43:11Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Fri, Dec 01, 2006 at 02:21:20PM +0100, sf wrote:\n> I just had a short (really short) look at your work. My impression is \n> that your repository setup is much too complicated.\n\nWell, I'm not really satisfied with the UI part.\nWhat exactly do you find complicated?\n\n> As I proposed elsewhere: For submodules to work you only need to allow \n> commits in tree objects (that is what your implementation requires as \n> well). Everything else is in the tools. Much simpler.\n\nI do not quite get your point.\nThe core of my work allows to put commits into tree objects.\nThen there is some more (but not quite finished) work to make the tools\nwork together with submodules.  So no, not everything is there yet.\n\n-- \nMartin Waitz\n"},{"id":"295662","messageId":"45703174.8000609@op5.se","threadId":"43065","inReplyTo":"20061201133558.GU18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-12-01T13:43:16Z","receivedAt":"2006-12-01T13:43:16Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Martin Waitz wrote:\n> hoi :)\n> \n> On Fri, Dec 01, 2006 at 02:05:33PM +0100, sf wrote:\n>>> On Fri, Dec 01, 2006 at 01:09:49PM +0100, sf wrote:\n>>>> Martin Waitz wrote:\n>>>>> So you not only store your submodule HEAD commit in the supermodule\n>>>>> when you do commit to the supermodule, it also means that your\n>>>>> submodule HEAD will be updated when you update your supermodule.\n>>>> Why the magic? The typical workflow in git is\n>>>>\n>>>> 1. You work on a branch, i.e. edit and commit and so on.\n>>>> 2. At some point, you decide to share the work you did on that branch \n>>>> (e-mail a patch, merge into another branch, push upstream or let it by \n>>>> pulled by upstream)\n>>> 3. Other people want to use your new work.\n>> Sorry, if that was not obvious: You actually procceed with one of the \n>> options I listed in Step 2. What I wanted to state is that with git you \n>> do not mix up committing (which is local to your repository and your \n>> branch) and publishing.\n> \n> I guess you are refering to not mix up committing to the submodule and\n> updating the supermodule index.\n> These are really two separate steps, I just combined them above because I\n> wanted to put emphasis on the other part: it is not a one-way flow, it\n> is bidirectional, so your HEAD would have to changed if the supermodule\n> gets updated.\n> And I consider changing HEAD, without looking at the branch it points\n> to, to be a bad thing.\n> \n\nSo a commit in the supermodule turns into a commit in the submodule? \nThat's just plain wrong. If it doesn't, why would the submodule HEAD \nhave to change?\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\n"},{"id":"296862","messageId":"20061201134610.GW18810@admingilde.org","threadId":"43065","inReplyTo":"45703174.8000609@op5.se","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-01T13:46:11Z","receivedAt":"2006-12-01T13:46:11Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Fri, Dec 01, 2006 at 02:43:16PM +0100, Andreas Ericsson wrote:\n> So a commit in the supermodule turns into a commit in the submodule? \n\nno.\n\n> If it doesn't, why would the submodule HEAD have to change?\n\nSo how do you update your submodule?\n\nRemember: if you git-pull in the supermodule, you want to update the\nwhole thing, including all submodules.\n\n-- \nMartin Waitz\n"},{"id":"295224","messageId":"45703375.4050500@b-i-t.de","threadId":"43065","inReplyTo":"20061201133558.GU18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Stephan Feder","fromEmail":"sf@b-i-t.de","sentAt":"2006-12-01T13:51:49Z","receivedAt":"2006-12-01T13:51:49Z","isPatch":false,"sender":{"key":"sf@b-i-t.de","avatar":null},"body":"Martin Waitz wrote:\n> hoi :)\n> \n> On Fri, Dec 01, 2006 at 02:05:33PM +0100, sf wrote:\n>> >On Fri, Dec 01, 2006 at 01:09:49PM +0100, sf wrote:\n>> >>Martin Waitz wrote:\n>> >>>So you not only store your submodule HEAD commit in the supermodule\n>> >>>when you do commit to the supermodule, it also means that your\n>> >>>submodule HEAD will be updated when you update your supermodule.\n>> >>\n>> >>Why the magic? The typical workflow in git is\n>> >>\n>> >>1. You work on a branch, i.e. edit and commit and so on.\n>> >>2. At some point, you decide to share the work you did on that branch \n>> >>(e-mail a patch, merge into another branch, push upstream or let it by \n>> >>pulled by upstream)\n>> >\n>> >3. Other people want to use your new work.\n>> \n>> Sorry, if that was not obvious: You actually procceed with one of the \n>> options I listed in Step 2. What I wanted to state is that with git you \n>> do not mix up committing (which is local to your repository and your \n>> branch) and publishing.\n> \n> I guess you are refering to not mix up committing to the submodule and\n> updating the supermodule index.\n\nThe opposite: If you work in the supermodule, even if it is in the code \nof the submodule, you only commit to the supermodule. The submodule does \nnot \"know\" about these changes after step 1.\n\n> These are really two separate steps, I just combined them above because I\n> wanted to put emphasis on the other part: it is not a one-way flow, it\n> is bidirectional, so your HEAD would have to changed if the supermodule\n> gets updated.\n\nWhy do you mix up supermodule and submodule? The way I see your proposal \nyou cannot change submodule and supermodule independently. That is a \nhuge drawback.\n\nRegards\n\n"},{"id":"297220","messageId":"45703ACB.6050007@b-i-t.de","threadId":"43065","inReplyTo":"20061201134311.GV18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Stephan Feder","fromEmail":"sf@b-i-t.de","sentAt":"2006-12-01T14:23:07Z","receivedAt":"2006-12-01T14:23:07Z","isPatch":false,"sender":{"key":"sf@b-i-t.de","avatar":null},"body":"Martin Waitz wrote:\n> hoi :)\n> \n> On Fri, Dec 01, 2006 at 02:21:20PM +0100, sf wrote:\n>> I just had a short (really short) look at your work. My impression is \n>> that your repository setup is much too complicated.\n> \n> Well, I'm not really satisfied with the UI part.\n> What exactly do you find complicated?\n\nI am not talking about the UI. On the contrary, I am talking about the \nrepository, i.e. the directories and files on your disk and their \ncontents. You add a lot of additional information to the repository that \nis not needed at all.\n\n>> As I proposed elsewhere: For submodules to work you only need to allow \n>> commits in tree objects (that is what your implementation requires as \n>> well). Everything else is in the tools. Much simpler.\n> \n> I do not quite get your point.\n> The core of my work allows to put commits into tree objects.\n\nThat is fine.\n\n> Then there is some more (but not quite finished) work to make the tools\n> work together with submodules.  So no, not everything is there yet.\n\nAnd what is already there is a lot of meta information (see above). You \ndo not need that.\n\nFor example, in the index, if it is a commit (i.e. a subproject), store \nthe commit id (not the commit's tree id ). Make the tools handle this \ncase (as yet, all code expects only trees and blobs when they parse \ntrees). Especially, extend update-index to be able to store a commit \ninstead of the tree.\n\nOr else, do not change what is recorded in the index. Then, at commit \ntime, you not only commit the superproject but also all subprojects.\n\nOr allow both.\n\nAnyway, you can create commits in tree objects. See, you did not need to \n  store any additional information in the repository.\n\nTo push and pull you have to extend the tools as well. That is the next \nstep.\n\nRegards\n\nStephan\n\n-- \nb.i.t.\nberatungsgesellschaft für informations-technologie mbh\nStephan Feder\nelisabethenstr. 62   fon: +49(0)6151/827575\n64283 darmstadt      fax: +49(0)6151/827576\nmailto:sf@b-i-t.de   www: http://www.b-i-t.de\n"},{"id":"298466","messageId":"457041AD.4010601@op5.se","threadId":"43065","inReplyTo":"20061201134610.GW18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-12-01T14:52:29Z","receivedAt":"2006-12-01T14:52:29Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Martin Waitz wrote:\n> hoi :)\n> \n> On Fri, Dec 01, 2006 at 02:43:16PM +0100, Andreas Ericsson wrote:\n>> So a commit in the supermodule turns into a commit in the submodule? \n> \n> no.\n> \n>> If it doesn't, why would the submodule HEAD have to change?\n> \n> So how do you update your submodule?\n> \n\nBy committing to it separately, or by getting changes from the upstream \nproject (openssl, libcurl, ...).\n\n> Remember: if you git-pull in the supermodule, you want to update the\n> whole thing, including all submodules.\n> \n\nOnly if the new commits I pull into the supermodule DAG has commits \nwhich includes a new shapshot of the submodule, otherwise it wouldn't be \nnecessary.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\n"},{"id":"293960","messageId":"20061201145817.GY18810@admingilde.org","threadId":"43065","inReplyTo":"45703375.4050500@b-i-t.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-01T14:58:17Z","receivedAt":"2006-12-01T14:58:17Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Fri, Dec 01, 2006 at 02:51:49PM +0100, Stephan Feder wrote:\n> If you work in the supermodule, even if it is in the code of the\n> submodule, you only commit to the supermodule. The submodule does not\n> \"know\" about these changes after step 1.\n\nI think we are using totally different definitions of \"submodule\".\n\nFor me a submodule is responsible for everything in or below a certain\ndirectory.  So by definition when you change something in this\ndirectory, you have to change it in the submodule.\nYou can't change the submodule contents in the supermodule without also\nchanging the submodule.\nThis is just like you can't commit a change to a file without also\nchanging the file.\n\nThen the supermodule just records the current content of the entire\ntree.  The only new thing is that instead of simple files there are now\nsubmodules and that are also recorded.\n\n> Why do you mix up supermodule and submodule? The way I see your proposal \n> you cannot change submodule and supermodule independently. That is a \n> huge drawback.\n\nNo, this is the benefit you get by introducing submodules.\nWhy would you want to introduce a submodule when it is not linked to the\nsupermodule?\n\n-- \nMartin Waitz\n"},{"id":"297886","messageId":"20061201150045.GZ18810@admingilde.org","threadId":"43065","inReplyTo":"457041AD.4010601@op5.se","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-01T15:00:46Z","receivedAt":"2006-12-01T15:00:46Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Fri, Dec 01, 2006 at 03:52:29PM +0100, Andreas Ericsson wrote:\n> >Remember: if you git-pull in the supermodule, you want to update the\n> >whole thing, including all submodules.\n> >\n> \n> Only if the new commits I pull into the supermodule DAG has commits \n> which includes a new shapshot of the submodule, otherwise it wouldn't be \n> necessary.\n\nOf course.\n\nBut if the supermodule contains changes to the submodule, you still\nhave to change the submodule.  And this implies changing the submodule\nHEAD or some branch.\n\n-- \nMartin Waitz\n"},{"id":"297207","messageId":"20061201150746.GA18810@admingilde.org","threadId":"43065","inReplyTo":"45703ACB.6050007@b-i-t.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-01T15:07:46Z","receivedAt":"2006-12-01T15:07:46Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Fri, Dec 01, 2006 at 03:23:07PM +0100, Stephan Feder wrote:\n> And what is already there is a lot of meta information (see above). You \n> do not need that.\n\nWhat information are you refering to?\nPerhaps you have looked into my old branch?\nThe current implementation is in \"module2\".\n\n> For example, in the index, if it is a commit (i.e. a subproject), store \n> the commit id (not the commit's tree id ).\n\nThis is exactly what I have done.\n\n> Especially, extend update-index to be able to store a commit \n> instead of the tree.\n\nDone, except that update-index never stores trees ;-)\n\n> Or else, do not change what is recorded in the index. Then, at commit \n> time, you not only commit the superproject but also all subprojects.\n\nBut then submodules would be handled differently to files which I wanted\nto avoid.\n\n> To push and pull you have to extend the tools as well. That is the next \n> step.\n\nAlso done.\n\n-- \nMartin Waitz\n"},{"id":"296514","messageId":"45704EA3.40203@b-i-t.de","threadId":"43065","inReplyTo":"20061201145817.GY18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Stephan Feder","fromEmail":"sf@b-i-t.de","sentAt":"2006-12-01T15:47:47Z","receivedAt":"2006-12-01T15:47:47Z","isPatch":false,"sender":{"key":"sf@b-i-t.de","avatar":null},"body":"Martin Waitz wrote:\n> hoi :)\n> \n> On Fri, Dec 01, 2006 at 02:51:49PM +0100, Stephan Feder wrote:\n>> If you work in the supermodule, even if it is in the code of the\n>> submodule, you only commit to the supermodule. The submodule does not\n>> \"know\" about these changes after step 1.\n> \n> I think we are using totally different definitions of \"submodule\".\n\nNo so different. The way I see it is that \"I\" (meaning with submodules \nimplemented as I proposed) could pull regularly from \"your\" repositories \n(implemented as you proposed) and work with the result (including \nsubmodules). Could you do the same?\n\n> For me a submodule is responsible for everything in or below a certain\n> directory.  So by definition when you change something in this\n> directory, you have to change it in the submodule.\n\nBut you do not consider the case where you cannot change the submodule \nbecause you do not own it.\n\nFor example, git has the subproject xdiff. If git had been able to work \nwith subprojects as I envision, and if xdiff had been published as a git \nrepository (not necessarily subproject enabled), it could have been \npulled in git's subdirectory xdiff as a subproject. There would not have \nbeen a separate branch or even repository for xdiff in the git repository.\n\nAll changes to xdiff in git could have been committed to the git \nrepository only. Independently, they could have been published to \nupstream and be put into the xdiff repository by its author. But the \nlast part is what only the owner of the xdiff repository is able to decide.\n\n(Ok, ok... the example sucks badly because xdiff has been massively \nchanged for its usage in git so the changes would not be integrated by \nupstream. But you can imagine where you use a library essentially as is, \nonly if you discover bugs you fix them immediately in your repository \nand keep those fixes in your version of the library, even on upgrade, \nuntil the bugs have been fixed by upstream.)\n\n> You can't change the submodule contents in the supermodule without also\n> changing the submodule.\n> This is just like you can't commit a change to a file without also\n> changing the file.\n\nThere is a difference. I would say: If you commit a change to a file in \none branch, it need not be changed in all branches.\n\n> Then the supermodule just records the current content of the entire\n> tree.  The only new thing is that instead of simple files there are now\n> submodules and that are also recorded.\n\nYes, and that is all you need. If the changes are to be part of a branch \nof the submodule, they have to be pulled. That is an independent operation.\n\n>> Why do you mix up supermodule and submodule? The way I see your proposal \n>> you cannot change submodule and supermodule independently. That is a \n>> huge drawback.\n> \n> No, this is the benefit you get by introducing submodules.\n> Why would you want to introduce a submodule when it is not linked to the\n> supermodule?\n\nBecause the submodule must be independent of the supermodule.\n\nI see where you are coming from. You have one project that is divided \ninto subprojects but the subprojects themselves are not independent.\n\nWhat I would like to solve is the followng: You have a project X, an \nthis project is made part of two other projects Y and Z (as a submodule \nor subproject or whatever you want to call it). The project X need not, \nmust not or cannot care that it was made a subproject. But in projects Y \nand Z, you must be able to bugfix or extend or modify the code of \nprojectX, and you must be able to push and pull changes between all \nthree projects (of course we are only talking about the code part of \nproject X).\n\nDo you see where your solution makes that impossible, and that with more \nchanges to the repository layout?\n\nRegards\n\nStephan\n\n"},{"id":"296317","messageId":"45705283.5080900@b-i-t.de","threadId":"43065","inReplyTo":"20061201150746.GA18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Stephan Feder","fromEmail":"sf@b-i-t.de","sentAt":"2006-12-01T16:04:19Z","receivedAt":"2006-12-01T16:04:19Z","isPatch":false,"sender":{"key":"sf@b-i-t.de","avatar":null},"body":"Martin Waitz wrote:\n> hoi :)\n> \n> On Fri, Dec 01, 2006 at 03:23:07PM +0100, Stephan Feder wrote:\n>> And what is already there is a lot of meta information (see above). You \n>> do not need that.\n> \n> What information are you refering to?\n> Perhaps you have looked into my old branch?\n> The current implementation is in \"module2\".\n\nI was looking into git-init-module.sh (branch module2). There you set up \na separate git repository for the submodule and store references to it \ninto the supermodules's repository.\n\n> \n>> For example, in the index, if it is a commit (i.e. a subproject), store \n>> the commit id (not the commit's tree id ).\n> \n> This is exactly what I have done.\n> \n>> Especially, extend update-index to be able to store a commit \n>> instead of the tree.\n> \n> Done, except that update-index never stores trees ;-)\n\nYes, I forgot.\n\n>> Or else, do not change what is recorded in the index. Then, at commit \n>> time, you not only commit the superproject but also all subprojects.\n> \n> But then submodules would be handled differently to files which I wanted\n> to avoid.\n\nOn the other hand, it feels more naturally to only commit at the end of \nyour work. So both alternatives have their merits.\n\n> \n>> To push and pull you have to extend the tools as well. That is the next \n>> step.\n> \n> Also done.\n\nI hope I have time to give your solution a try.\n\nRegards\n\nStephan\n\n-- \nb.i.t.\nberatungsgesellschaft für informations-technologie mbh\nStephan Feder\nelisabethenstr. 62   fon: +49(0)6151/827575\n64283 darmstadt      fax: +49(0)6151/827576\nmailto:sf@b-i-t.de   www: http://www.b-i-t.de\n"},{"id":"298214","messageId":"20061201161558.GC18810@admingilde.org","threadId":"43065","inReplyTo":"45705283.5080900@b-i-t.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-01T16:15:58Z","receivedAt":"2006-12-01T16:15:58Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Fri, Dec 01, 2006 at 05:04:19PM +0100, Stephan Feder wrote:\n> I was looking into git-init-module.sh (branch module2). There you set up \n> a separate git repository for the submodule and store references to it \n> into the supermodules's repository.\n\nyes.\n\nThis is to be able to call git-fsck-objects and git-prune in the\ntoplevel supermodule.  When traversing the object tree, it already knows\nabout all submodule, but only about those versions that are really part\nof their supermodule.\nSo I have to teach git about separate submodule branches which may be\nused, in order not to prune them away.\n\n-- \nMartin Waitz\n"},{"id":"297706","messageId":"45705A94.2070509@op5.se","threadId":"43065","inReplyTo":"20061201150045.GZ18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-12-01T16:38:44Z","receivedAt":"2006-12-01T16:38:44Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Martin Waitz wrote:\n> hoi :)\n> \n> On Fri, Dec 01, 2006 at 03:52:29PM +0100, Andreas Ericsson wrote:\n>>> Remember: if you git-pull in the supermodule, you want to update the\n>>> whole thing, including all submodules.\n>>>\n>> Only if the new commits I pull into the supermodule DAG has commits \n>> which includes a new shapshot of the submodule, otherwise it wouldn't be \n>> necessary.\n> \n> Of course.\n> \n> But if the supermodule contains changes to the submodule, you still\n> have to change the submodule.  And this implies changing the submodule\n> HEAD or some branch.\n> \n\nNot really. I fail to see why HEAD needs to be changed so long as the \ncommit is in the submodule's odb.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\n"},{"id":"295576","messageId":"Pine.LNX.4.64.0612010844380.3695@woody.osdl.org","threadId":"43065","inReplyTo":"45705A94.2070509@op5.se","subject":"Re: [RFC] Submodules in GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-01T16:49:20Z","receivedAt":"2006-12-01T16:49:20Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 1 Dec 2006, Andreas Ericsson wrote:\n\n> Martin Waitz wrote:\n> > \n> > But if the supermodule contains changes to the submodule, you still\n> > have to change the submodule.  And this implies changing the submodule\n> > HEAD or some branch.\n> > \n> \n> Not really. I fail to see why HEAD needs to be changed so long as the commit\n> is in the submodule's odb.\n\nRight. A commit in the supermodule should _not_ imply a commit in the \nsubmodule.\n\nMaybe I should take a look at the code, but it sounds like people are \nstill trying to \"mix\" submodules too much. \n\nThink of it this way: one common use for submodules is really to just \n(occasionally) track somebody elses code. The submodule should be a \ntotally pristine copy from somebody else (ie it might be the \"intel driver \nfor X.org\" submodule, maintained within intel), and the supermodule just \nrefers to it indirectly (ie the supermodule might be the \"Fedora Core X \ngroup\" which contains all the different drivers from different people).\n\nSo anything that mixes super-modules and sub-modules too much will always \nbreak this kind of model.\n\nA supermodule can never \"contain changes\" to a submodule. A supermodule \nwould always just point to the submodule, and not have any changes \nwhat-so-ever of its own. The submodule is self-sufficient, and always \ncontains all its _own_ changes.\n\n"},{"id":"296264","messageId":"20061201165418.GD18810@admingilde.org","threadId":"43065","inReplyTo":"45704EA3.40203@b-i-t.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-01T16:54:19Z","receivedAt":"2006-12-01T16:54:19Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Fri, Dec 01, 2006 at 04:47:47PM +0100, Stephan Feder wrote:\n> No so different. The way I see it is that \"I\" (meaning with submodules \n> implemented as I proposed) could pull regularly from \"your\" repositories \n> (implemented as you proposed) and work with the result (including \n> submodules). Could you do the same?\n\nSorry, but with all that many people proposing things I am a bit lost\nnow.  Sometimes I thought you want exactly the same thing as I do,\nsometimes I think we are talking in totally different directions.\n\n> >For me a submodule is responsible for everything in or below a certain\n> >directory.  So by definition when you change something in this\n> >directory, you have to change it in the submodule.\n> \n> But you do not consider the case where you cannot change the submodule \n> because you do not own it.\n\nI do not understand you here.\nThe submodule is part of the supermodule, and the one who sets up the\nrepository owns the whole thing, including all submodules, just like all\nthe files which are part of the project.\n\nIf you mean the upstream repository of the submodule, then yes, this is\nof course completely separated from the submodule and may be owned by\nsomeone else.  Consequently, this upstream repository of course does not\nneed to change when someone introduces changes in the supermodule.\n\n> For example, git has the subproject xdiff. If git had been able to work \n> with subprojects as I envision, and if xdiff had been published as a git \n> repository (not necessarily subproject enabled), it could have been \n> pulled in git's subdirectory xdiff as a subproject.\n\nThis could have been done if submodule support would have been available\nat the time xdiff was introduced, yes.\n\n> There would not have been a separate branch or even repository for\n> xdiff in the git repository.\n\nWhat separate branch or repository are you talking about?\n\n> All changes to xdiff in git could have been committed to the git \n> repository only.\n\nYes, but if it would have been integrated as a submodule it obviously\nwould have been committed to the xdiff submodule inside the git\nrepository.\nSo the changes are really part of the git repository, but you could go\nto the \"git/xdiff\" directory and only see the changes in the submodule,\nwithout the normal supermodule history.\n\n> Independently, they could have been published to upstream and be put\n> into the xdiff repository by its author.  But the last part is what\n> only the owner of the xdiff repository is able to decide.\n\nOf course, everything still works like normal git repositories.\n\n> >You can't change the submodule contents in the supermodule without also\n> >changing the submodule.\n> >This is just like you can't commit a change to a file without also\n> >changing the file.\n> \n> There is a difference. I would say: If you commit a change to a file in \n> one branch, it need not be changed in all branches.\n\nBut you need to change _at_least_ one branch.\nOtherwise you cannot commit to a branch.\n\nSo if you change something in a submodule, you have to change one branch\nin the submodule.\nIf you call git-checkout in the supermodule this will result in\nsomething like a git-reset in the submodule.\n\n> >No, this is the benefit you get by introducing submodules.\n> >Why would you want to introduce a submodule when it is not linked to the\n> >supermodule?\n> \n> Because the submodule must be independent of the supermodule.\n> \n> I see where you are coming from. You have one project that is divided \n> into subprojects but the subprojects themselves are not independent.\n> \n> What I would like to solve is the followng: You have a project X, an \n> this project is made part of two other projects Y and Z (as a submodule \n> or subproject or whatever you want to call it). The project X need not, \n> must not or cannot care that it was made a subproject. But in projects Y \n> and Z, you must be able to bugfix or extend or modify the code of \n> projectX, and you must be able to push and pull changes between all \n> three projects (of course we are only talking about the code part of \n> project X).\n\nOf course.\n\nSo if you wanted to check out everything, you could have something like\n~/src/X, ~/src/Y/X, and ~/src/Z/X.\nAll of these would be GIT repositories, all of them have their\nindependent branches.\n\nWhat I am saying is just that if you update Y, and the new Y contains an\nupdated version of X, then ~/src/Y/X/.git/refs/heads/master will be\nchanged by the pull, resulting in the new version of X being checked out\nin ~/src/Y/X (alongside all the other updates inside ~/src/Y).\nThis of course is independend from ~/src/X or  ~/src/Z/X.\n\n> Do you see where your solution makes that impossible, and that with more \n> changes to the repository layout?\n\nNo ;-)\n\n-- \nMartin Waitz\n"},{"id":"295442","messageId":"20061201165708.GE18810@admingilde.org","threadId":"43065","inReplyTo":"45705A94.2070509@op5.se","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-01T16:57:08Z","receivedAt":"2006-12-01T16:57:08Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"On Fri, Dec 01, 2006 at 05:38:44PM +0100, Andreas Ericsson wrote:\n> >But if the supermodule contains changes to the submodule, you still\n> >have to change the submodule.  And this implies changing the submodule\n> >HEAD or some branch.\n> >\n> \n> Not really. I fail to see why HEAD needs to be changed so long as the \n> commit is in the submodule's odb.\n\nBecause I want the submodule to act as a normal git repository.\nPlease note that I also voted against changing HEAD directly, but that\nthe new commit which came from the supermodule is just stored in one\nbranch of the submodule, as part of the supermodule checkout.\n\n-- \nMartin Waitz\n"},{"id":"295765","messageId":"457061A7.2000102@b-i-t.de","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612010844380.3695@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"sf","fromEmail":"sf@b-i-t.de","sentAt":"2006-12-01T17:08:55Z","receivedAt":"2006-12-01T17:08:55Z","isPatch":false,"sender":{"key":"sf@b-i-t.de","avatar":null},"body":"Linus Torvalds wrote:\n...\n > Think of it this way: one common use for submodules is really to just\n > (occasionally) track somebody elses code. The submodule should be a\n > totally pristine copy from somebody else (ie it might be the \"intel \ndriver\n > for X.org\" submodule, maintained within intel), and the supermodule just\n > refers to it indirectly (ie the supermodule might be the \"Fedora Core X\n > group\" which contains all the different drivers from different people).\n\nCould you please be a little bit more specific about how you would store \nthe \"pristine copy\". There seems to be some agreement to store the \ncommit id of the submodule instead of a plain tree id in the \nsupermodules tree object, and that all objects that are reachable from \nthis commit are made part of the supermodule repository (either fetched \nor via alternates). Do you agree?\n\n...\n > A supermodule can never \"contain changes\" to a submodule. A supermodule\n > would always just point to the submodule, and not have any changes\n > what-so-ever of its own. The submodule is self-sufficient, and always\n > contains all its _own_ changes.\n\nThat is one of the points Martin Waitz and I are discussing.\n\nIf I understand you correctly you cannot make any changes to the \nsubmodules code _in the supermodule's repository_, no bugfixes, no \nextensions, no adaptions, nothing. Do you mean that?\n\nThat would be a third alternative. In my opinion the usefulness of \nsubmodules would be unnecessarily restricted if it comes to the choice \nof either using the code from upstream as is or do not use submodules at \nall. What is the point of the restriction?\n\nRegards\n\n"},{"id":"294309","messageId":"20061201171420.GF18810@admingilde.org","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612010844380.3695@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-01T17:14:20Z","receivedAt":"2006-12-01T17:14:20Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Fri, Dec 01, 2006 at 08:49:20AM -0800, Linus Torvalds wrote:\n> Think of it this way: one common use for submodules is really to just \n> (occasionally) track somebody elses code. The submodule should be a \n> totally pristine copy from somebody else (ie it might be the \"intel driver \n> for X.org\" submodule, maintained within intel), and the supermodule just \n> refers to it indirectly (ie the supermodule might be the \"Fedora Core X \n> group\" which contains all the different drivers from different people).\n\nYes, but it is not only about tracking, also about distributing\nsubmodules.\n\nOne Fedora X developer fixes a bug in the intel driver, commits that to\nthe submodule and then updates the supermodule to the new version (by\ncalling \"git-update-index drivers/intel && git-commit\" or something).  Then\nanother Feora X developer updates his X repository.  By pulling the\nsupermodule he also gets a new version of the submodule.\nAnd this new version of the submodule is stored in a branch which can be\naccessed by the submodule.\n\n> A supermodule can never \"contain changes\" to a submodule.\n\nThe supermodule always contains _the_entire_ submodule with its complete\nhistory, so it also does contain changes.  But it does not per-se\ncontain changes, only indirectly (i.e. the commits in the submodule are\nnot part of the supermodule commit chain).\n\n> A supermodule would always just point to the submodule, and not have\n> any changes what-so-ever of its own. The submodule is self-sufficient,\n> and always contains all its _own_ changes.\n\nYes.\n\n-- \nMartin Waitz\n"},{"id":"297724","messageId":"45706758.2020907@b-i-t.de","threadId":"43065","inReplyTo":"20061201165418.GD18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Stephan Feder","fromEmail":"sf@b-i-t.de","sentAt":"2006-12-01T17:33:12Z","receivedAt":"2006-12-01T17:33:12Z","isPatch":false,"sender":{"key":"sf@b-i-t.de","avatar":null},"body":"Martin Waitz wrote:\n> hoi :)\n> \n> On Fri, Dec 01, 2006 at 04:47:47PM +0100, Stephan Feder wrote:\n>> No so different. The way I see it is that \"I\" (meaning with submodules \n>> implemented as I proposed) could pull regularly from \"your\" repositories \n>> (implemented as you proposed) and work with the result (including \n>> submodules). Could you do the same?\n> \n> Sorry, but with all that many people proposing things I am a bit lost\n> now.  Sometimes I thought you want exactly the same thing as I do,\n> sometimes I think we are talking in totally different directions.\n\nWe are in agreement about two fundamental parts of the implementation \nand their meaning:\n\n1. A submodule is stored as a commit id in a tree object.\n\n2. Every object that is reachable from the submodule's commit are \nreachable from the supermodule's repository.\n\nPlease confirm.\n\nWe now argue about how to work with that repository _object_ model.\n\n>> >For me a submodule is responsible for everything in or below a certain\n>> >directory.  So by definition when you change something in this\n>> >directory, you have to change it in the submodule.\n>> \n>> But you do not consider the case where you cannot change the submodule \n>> because you do not own it.\n> \n> I do not understand you here.\n> The submodule is part of the supermodule, and the one who sets up the\n> repository owns the whole thing, including all submodules, just like all\n> the files which are part of the project.\n\nIf you mean by \"owns the whole thing\" what I stated above in 2. the we \nagree.\n\n> If you mean the upstream repository of the submodule, then yes, this is\n> of course completely separated from the submodule and may be owned by\n> someone else.  Consequently, this upstream repository of course does not\n> need to change when someone introduces changes in the supermodule.\n\nI think we still agree.\n\n>> For example, git has the subproject xdiff. If git had been able to work \n>> with subprojects as I envision, and if xdiff had been published as a git \n>> repository (not necessarily subproject enabled), it could have been \n>> pulled in git's subdirectory xdiff as a subproject.\n> \n> This could have been done if submodule support would have been available\n> at the time xdiff was introduced, yes.\n> \n>> There would not have been a separate branch or even repository for\n>> xdiff in the git repository.\n> \n> What separate branch or repository are you talking about?\n\nThat's it: There is no need for a separate branch or repository. If you \nhave the subproject's commit in the superproject's object database (and \nwe really have that, see 1. and 2. above), why do you _have to_ store it \nelsewhere?\n\n>> All changes to xdiff in git could have been committed to the git \n>> repository only.\n> \n> Yes, but if it would have been integrated as a submodule it obviously\n> would have been committed to the xdiff submodule inside the git\n> repository.\n\nNo. The xdiff submodule would only exist as part of the git repository. \nYou could, f.e., access the xdiff commit in git HEAD as HEAD:xdiff// \n(again my proposed syntax). HEAD:xdiff//~2:xemit.c would give you the \ngrandparent of xemit.c in the xdiff submodule. And so on. You can even \nhave submodules that have themselves submodules.\n\n> So the changes are really part of the git repository, but you could go\n> to the \"git/xdiff\" directory and only see the changes in the submodule,\n> without the normal supermodule history.\n\nSee above.\n\n>> Independently, they could have been published to upstream and be put\n>> into the xdiff repository by its author.  But the last part is what\n>> only the owner of the xdiff repository is able to decide.\n> \n> Of course, everything still works like normal git repositories.\n\nOK.\n\n>> >You can't change the submodule contents in the supermodule without also\n>> >changing the submodule.\n>> >This is just like you can't commit a change to a file without also\n>> >changing the file.\n>> \n>> There is a difference. I would say: If you commit a change to a file in \n>> one branch, it need not be changed in all branches.\n> \n> But you need to change _at_least_ one branch.\n> Otherwise you cannot commit to a branch.\n\nBut only the supermodule's branch.\n\n> So if you change something in a submodule, you have to change one branch\n> in the submodule.\n\nNo.\n\n> If you call git-checkout in the supermodule this will result in\n> something like a git-reset in the submodule.\n\nIf you mean the submodule repository created by init-module I \nunderstand. But why create this \"helper repository at all\"?\n\n>> >No, this is the benefit you get by introducing submodules.\n>> >Why would you want to introduce a submodule when it is not linked to the\n>> >supermodule?\n>> \n>> Because the submodule must be independent of the supermodule.\n>> \n>> I see where you are coming from. You have one project that is divided \n>> into subprojects but the subprojects themselves are not independent.\n>> \n>> What I would like to solve is the followng: You have a project X, an \n>> this project is made part of two other projects Y and Z (as a submodule \n>> or subproject or whatever you want to call it). The project X need not, \n>> must not or cannot care that it was made a subproject. But in projects Y \n>> and Z, you must be able to bugfix or extend or modify the code of \n>> projectX, and you must be able to push and pull changes between all \n>> three projects (of course we are only talking about the code part of \n>> project X).\n> \n> Of course.\n> \n> So if you wanted to check out everything, you could have something like\n> ~/src/X, ~/src/Y/X, and ~/src/Z/X.\n> All of these would be GIT repositories, all of them have their\n> independent branches.\n> \n> What I am saying is just that if you update Y, and the new Y contains an\n> updated version of X, then ~/src/Y/X/.git/refs/heads/master will be\n> changed by the pull, resulting in the new version of X being checked out\n> in ~/src/Y/X (alongside all the other updates inside ~/src/Y).\n> This of course is independend from ~/src/X or  ~/src/Z/X.\n> \n>> Do you see where your solution makes that impossible, and that with more \n>> changes to the repository layout?\n> \n> No ;-)\n> \n\nSorry, have to leave for home so I must leave that uncommented. \nHopefully I can join in during the weekend.\n\nRegards\n\n"},{"id":"294219","messageId":"45706F19.6090609@op5.se","threadId":"43065","inReplyTo":"457061A7.2000102@b-i-t.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-12-01T18:06:17Z","receivedAt":"2006-12-01T18:06:17Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"sf wrote:\n> \n> That is one of the points Martin Waitz and I are discussing.\n> \n> If I understand you correctly you cannot make any changes to the \n> submodules code _in the supermodule's repository_, no bugfixes, no \n> extensions, no adaptions, nothing. Do you mean that?\n> \n> That would be a third alternative. In my opinion the usefulness of \n> submodules would be unnecessarily restricted if it comes to the choice \n> of either using the code from upstream as is or do not use submodules at \n> all. What is the point of the restriction?\n> \n\nThat depends on your definition of submodule. In my eyes, a submodule is \na separate repo that can be committed to separately (and generally also \nbuilt separately), although it's usually built into something else. I'm \nimagining most submodules will contain only library code and its testing \nroutines.\n\nInsofar as I've envisioned submodules, it's a separate git repo where \nyou simply record a certain snapshot of the sub-repo with a commit in \nthe super-module, like so:\n\n$ git commit ssl-functions/*.[ch] openssl -m \"Upgraded openssl with \nnecessary changes to core code\"\n\n(yes, I know it's horrid to use -m to commit, and I daily advocate \nagainst it where I work, but you get the idea, I'm sure)\n\nIsn't this how it's supposed to work? Enlighten me, and please remember \nthat I'm drunk atm, so make it obvious ;-)\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\n"},{"id":"298639","messageId":"45706F8C.3070000@op5.se","threadId":"43065","inReplyTo":"20061201165708.GE18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-12-01T18:08:12Z","receivedAt":"2006-12-01T18:08:12Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"\n\nMartin Waitz wrote:\n> On Fri, Dec 01, 2006 at 05:38:44PM +0100, Andreas Ericsson wrote:\n>>> But if the supermodule contains changes to the submodule, you still\n>>> have to change the submodule.  And this implies changing the submodule\n>>> HEAD or some branch.\n>>>\n>> Not really. I fail to see why HEAD needs to be changed so long as the \n>> commit is in the submodule's odb.\n> \n> Because I want the submodule to act as a normal git repository.\n> Please note that I also voted against changing HEAD directly, but that\n> the new commit which came from the supermodule is just stored in one\n> branch of the submodule, as part of the supermodule checkout.\n> \n\nYou're assuming the super- and sub-module will share HEAD, or at least \nODB, I think. I'm not convinced this is necessary. Convince me. I'll go \ndrink bear and get some dancing done while you're at it ;-)\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\n"},{"id":"296084","messageId":"20061201184801.GG18810@admingilde.org","threadId":"43065","inReplyTo":"45706758.2020907@b-i-t.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-01T18:48:02Z","receivedAt":"2006-12-01T18:48:02Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"On Fri, Dec 01, 2006 at 06:33:12PM +0100, Stephan Feder wrote:\n> We are in agreement about two fundamental parts of the implementation \n> and their meaning:\n> \n> 1. A submodule is stored as a commit id in a tree object.\n> \n> 2. Every object that is reachable from the submodule's commit are \n> reachable from the supermodule's repository.\n\nCorrect.\n\n> >>For example, git has the subproject xdiff. If git had been able to work \n> >>with subprojects as I envision, and if xdiff had been published as a git \n> >>repository (not necessarily subproject enabled), it could have been \n> >>pulled in git's subdirectory xdiff as a subproject.\n> >\n> >This could have been done if submodule support would have been available\n> >at the time xdiff was introduced, yes.\n> >\n> >>There would not have been a separate branch or even repository for\n> >>xdiff in the git repository.\n> >\n> >What separate branch or repository are you talking about?\n> \n> That's it: There is no need for a separate branch or repository. If you \n> have the subproject's commit in the superproject's object database (and \n> we really have that, see 1. and 2. above), why do you _have to_ store it \n> elsewhere?\n\nLet's see if I understand you correctly:\n\nYou don't want to create an additional .git directory for the submodule\nand just handle everything with one toplevel .git repository for the\nwhole project.\nWithout the .git directory, you of course do not have refs/heads inside\nthe submodule.\n\nSo this is a different user-interface approach to submodules when\ncompared to my approach.  But the basis is the same and both could\ninter-operate.\n\nNow your submodule is no longer seen as an independent git repository\nand I think this would cause problems when you want to push/pull between\nthe submodule and its upstream repository.\nNo technical problems, but UI-problems because now your submodule is\nhandled completly different to a \"normal\" repository.\n\n\n> >Yes, but if it would have been integrated as a submodule it obviously\n> >would have been committed to the xdiff submodule inside the git\n> >repository.\n> \n> No. The xdiff submodule would only exist as part of the git repository. \n\nBut you could still call the \"xdiff\" part of the git repository a\nsubmodule.  And then changes to the xdiff directory result in a new\nsubmodule commit, even when there is no direct reference to it.\nSo you'd still \"commit to the xdiff submodule\".\n\n> You could, f.e., access the xdiff commit in git HEAD as HEAD:xdiff// \n> (again my proposed syntax). HEAD:xdiff//~2:xemit.c would give you the \n> grandparent of xemit.c in the xdiff submodule.\n\ngit-cat-file commit HEAD:xdiff already works out of the box (even\ncat-file tree to get the submodule tree).  But up to now revision\nparsing follows the file name only once.\n\nWhat about just separating things with \"/\"?\n\ncommit HEAD\ntree   HEAD/\nblob   HEAD/Makefile\ncommit HEAD/xdiff\ntree   HEAD/xdiff/\nblob   HEAD/xdiff~2/xemit.c\n\nthis may add some confusion when used with hierarchical branches, but\nit's still unique:\n\n\trefs/heads/master/xdiff/xemit.c\n\nJust use as many path components until a matching reference is found,\nthen start peeling.\nOr just use / between super and submodule:\n\n\trefs/heads/master:xdiff/xemit.c\n\nI think this is easier to read then\n\n\trefs/heads/master:xdiff//:xemit.c\n\n\n> If you mean the submodule repository created by init-module I \n> understand. But why create this \"helper repository at all\"?\n\nBecause it helps \"normal\" git operations ;-)\n\n-- \nMartin Waitz\n"},{"id":"294329","messageId":"20061201185100.GH18810@admingilde.org","threadId":"43065","inReplyTo":"45706F8C.3070000@op5.se","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-01T18:51:00Z","receivedAt":"2006-12-01T18:51:00Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Fri, Dec 01, 2006 at 07:08:12PM +0100, Andreas Ericsson wrote:\n> You're assuming the super- and sub-module will share HEAD, or at least \n> ODB, I think.\n\nnot HEAD, only ODB.\n\n> I'm not convinced this is necessary. Convince me. I'll go \n> drink bear and get some dancing done while you're at it ;-)\n\nGet me a beer and I will convince you :-)\n\n-- \nMartin Waitz\n"},{"id":"296694","messageId":"200612011917.19252.andyparkins@gmail.com","threadId":"43065","inReplyTo":"45706758.2020907@b-i-t.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2006-12-01T19:17:17Z","receivedAt":"2006-12-01T19:17:17Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Friday 2006, December 01 17:33, Stephan Feder wrote:\n\n> 1. A submodule is stored as a commit id in a tree object.\n>\n> 2. Every object that is reachable from the submodule's commit are\n> reachable from the supermodule's repository.\n\nI'm still not convinced about 2.  Why should any of the submodule commits be \nin the supermodule repository?  I know that is what you've implemented, but \nit still feels like too much of a blending of the submodule into the \nsupermodule.\n\nIn fact, why should the submodule commits be even visible in the supermodule?  \nThat tree->submodule commit is sufficient; there isn't any need to view \nsubmodule history in the supermodule.\n\n\n\nAndy\n\n-- \nDr Andrew Parkins, M Eng (Hons), AMIEE\n"},{"id":"298108","messageId":"20061201193802.GI18810@admingilde.org","threadId":"43065","inReplyTo":"200612011917.19252.andyparkins@gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-01T19:38:02Z","receivedAt":"2006-12-01T19:38:02Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Fri, Dec 01, 2006 at 07:17:17PM +0000, Andy Parkins wrote:\n> In fact, why should the submodule commits be even visible in the\n> supermodule?  That tree->submodule commit is sufficient; there isn't\n> any need to view submodule history in the supermodule.\n\nWell, but there is a need for a common object traversal.\nYou need that when sending all objects between two supermodule versions\nand also when you determine which objects are still reachable.\n\nThe easiest way to implement the common object traversal is to have all\nobjects in one object repository.\n\nIt may be possible to use two object stores and still do the common\nobject traversal but I do not think that gives you any benefits.\nYou still don't have a totally separated repository then, because\nyou can't do a reachability analysis in the submodule repository alone.\n\n-- \nMartin Waitz\n"},{"id":"295652","messageId":"Pine.LNX.4.64.0612011134080.3695@woody.osdl.org","threadId":"43065","inReplyTo":"457061A7.2000102@b-i-t.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-01T20:13:09Z","receivedAt":"2006-12-01T20:13:09Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 1 Dec 2006, sf wrote:\n>\n> Linus Torvalds wrote:\n> ...\n> > Think of it this way: one common use for submodules is really to just\n> > (occasionally) track somebody elses code. The submodule should be a\n> > totally pristine copy from somebody else (ie it might be the \"intel driver\n> > for X.org\" submodule, maintained within intel), and the supermodule just\n> > refers to it indirectly (ie the supermodule might be the \"Fedora Core X\n> > group\" which contains all the different drivers from different people).\n> \n> Could you please be a little bit more specific about how you would store the\n> \"pristine copy\".\n\nNote that it's not necessarily \"pristine\", since the submodule clearly is \na local git repository in its own right. So like _any_ git repository, you \ncan (and may well end up) having your own local branches in the submodule, \nwith your own local modifications.\n\nSo I'm not claiming that a submodule must always match some external git \ntree 100%, and that it must be read-only or anything like that. I'm just \nsaying that I suspect that quite often, one of the MOST IMPORTANT parts is \nthat the submodule is really something that somebody else technically \nmaintains, and that this is actually one of the _reasons_ why it is a \nsubmodule in the first place. \n\nFor example, a lot of projects end up having some kind of \"library \ncomponent\" as a submodule. Take something like a video player project, \nwhich would have something like ffmpeg as a submodule, not because you'd \nmaintain ffmpeg yourself, but simply because (let's say) the library \ninterface changes enough, or you need a specific version with some of your \nown fixes that haven't been released widely yet, so you want to carry all \nthe libraries you need _with_ you, even though you don't really maintain \nthat submodule. You at most have some small extensions of your own.\n\nNow, in this situation, it's relaly really _important_ that the submodule \nreally is totally independent of the supermodule, for several reasons.\n\nFor example, since you don't \"really\" own that project, carrying around \nyour own fixes is really really painful. We know it happens all the time, \nand a lot of projects end up needing their own version, but the _last_ \nthing you want is to be in merge hell all the time. So as a supermodule \nmaintainer, the best possible thing for you is to be able to push back \nthose local changes to the original project maintainer, so that you \n_don't_ have to maintain your own changes.\n\nBut you need to realize that the real maintainer of the submodule is \nTOTALLY UNINTERESTED in your supermodule. He's not going to maintain it, \nand in fact, if you have anything in the submodule that ends up talking \nabout your supermodule, that's just going to make it a lot less likely \nthat the upstream maintainer will ever pull your changes. He might take a \ndiff from you, but in a perfect world, you'd actually be able to tell him: \n\n \"Hey, I've got a git repository with a few fixes to your ffmpeg git tree, \n  please pull from git://myhost.com/submodule.git to get these fixes:\n\n\t... explanation of fixes and commits that are relevant to\n\tffmpeg, and have nothing to do with the supermodule, except\n\tthat you need those bug-fixes because you _use_ ffmpeg ...\n\n  Thanks\"\n\nSee?\n\nSo this is why it's really important that the submodule really is a git \nrepository in its own right, and why committing stuff in the supermodule \nNEVER affect the submodule itself directly (it might _cause_ you to also \ndo a commit in the submodule indirectly, but the submodule commit MUST be \ntotally independent, and stand on its own).\n\nNow, you don't _have_ to push things upstream, of course. You can always \njust maintain your own submodule branch, and every once in a while, inside \nthe submodule, you do\n\n\t# fetch the development in the origin/master branch\n\tgit fetch submodule-origin origin/master\n\n\t# rebase our own special magic sauce on top of that\n\tgit rebase origin/master\n\nto update your submodule, and _then_ you do a commit in the supermodule \n(after testing that the update is all ok, of course) which will update the \n\"commit\" pointer in the supermodule.\n\nNotice? In this example, we really maintained the submodule AS a \nsubmodule. It was independent, but tied into the supermodule, so that when \nwe clone the supermodule, or do things like bisection on a supermodule, we \nalways end up cloning the submodule too (and in the case of bisection, we \nreally only bisect the supermodule, but the submodule always gets \n\"tracked\" in the sense that we would always check out the state of the \nsubmodule that was appropriate for that particular commit in the \nsupermodule).\n\n> There seems to be some agreement to store the commit id of\n> the submodule instead of a plain tree id in the supermodules tree object, and\n> that all objects that are reachable from this commit are made part of the\n> supermodule repository (either fetched or via alternates). Do you agree?\n\nWell, I would actually argue that you may often want to have a supermodule \nand then at least have the _option_ to decide to not fetch all the \nsubmodules.\n\nFor an example of this kind of usage, let me tell you how we operated at \nTransmeta a few years ago, which I'm not saying is the _only_ way to \noperate, but it's ONE way to do it, and I'll also explain _why_ we did it, \nand why we had submodules.\n\nIn the case of transmeta, we had our own tools, our own programs, and we \n\"owned\" all of those. We _also_ used a lot of external tools, like gcc \netc. However, different people worked on different parts, and if you \nworked on the actual x86 JIT part, you probably didn't want to have all of \nthe gcc stuff in your tree _too_. That just took a lot of space, and you \nreally didn't want to compile the whole toolchain (which took hours), \nsince there were precompiled binaries readily available.\n\nStill, from a _release_ standpoint, when we released a new binary, that \nbinary very much depended not just on the actual JIT sources, but on the \nwhole toolchain. So if you wanted to be able to re-create a release, you \nreally needed _everything_. You couldn't just take the \"current version\" \nof the toolchain, you needed to have the toolchain that was used AT THE \nTIME OF THE RELEASE.\n\nAnd this is a _classic_ example of when you'd want to use submodules. \nNotice how everybody wanted _some_ of the submodules, but really only the \nrelease people wanted them _all_. The higher up the chain you were, the \nless likely you were to really want to muck around with the compiler and \nthe linker, for example. \n\nAnd nobody really owned all modules. \n\nSo what you really want is:\n\n - a supermodule maintainer that is not really the maintainer of _any_ of \n   the submodules, but that does the main \"build world\" infrastructure \n   (and generally would tend to also maintain the source control \n   infrastructure itself)\n\n - submodules that had their own maintainers, and where the maintainers \n   may or may not have wanted the supermodule, but even when they wanted \n   the supermodule, they might not want _all_ of the submodules, simply \n   because they just didn't care.\n\n - some of the submodules then have _upstream_ sources that were totally \n   independent, and that you would want to track, but you had zero power \n   AT ALL over them, and yet you migt well want to push back at least some \n   of the fixes you did - at least the ones that made sense even outside \n   your own project - just to avoid having to maintain a _huge_ set of \n   internal patches.\n\nSo no, I don't think the supermodule should even _force_ people to always \nget all the submodules. It migth be the default case, but at the same \ntime, it's just being polite to let users decide on their own whether they \nreally want _all_ of the build infrastructure sources.\n\n> If I understand you correctly you cannot make any changes to the submodules\n> code _in the supermodule's repository_, no bugfixes, no extensions, no\n> adaptions, nothing. Do you mean that?\n\nYes. I think you should make all changes _within_ the submodule, because \nthe submodule should still be an independent git tree in its own right.\n\nBut obviously, you'd often use a private _branch_ in the submodule beause \nyou end up having whatever private extensions. That's always true: we \nalways have the \"master\" branch that is kind of the default \"private \nbranch\" for any repository, but obviously that is often extended upon, and \nyou may have several private branches. \n\nFor example, after you've done a big update (from some external upstream \nsource) in the submodule that you are using, you migth decide that you do \nall the work on that new big update in a _new_ private branch within the \nsubmodule - and get the submodule changes all squared away on its own \n_before_ you then decide to commit the end result (the tip of that new \nprivate branch) within the supermodule.\n\nIe, you very much should be able to to do\n\n\tgit clone supermodule/that/one/submodule my-own-version-of-submodule\n\nto clone a submodule _without_ getting anything else (but still get all \nthe work you did within he submodule - very much including your own \nprivate branch work).\n\nAnd the importance of keeping the submodule independent is partly just \nstability and sanity, but partly also scalability. For example, the \n\"index\" in a supermodule should NOT include the indexes of all the \nsubmodules. That's really important, because the index doesn't really \nscale. Things do slow down with large indexes. \n\nFor example, git can handle tens of thousands of files easily. I suspect \nit scales well to hundreds of thousands of filenames. But with \nsupermodules, you really can end up in the situation where you have _tens_ \nof these submodules, maybe even hundreds. And if you try to maintain one \nunified index for the _whole_ thing, I guarantee you that you'll start \nfeeling the pain. Indexing millions of files is just not going to be \npretty.\n\nSo just from a git stability and scalability point, it's important to keep \nsubprojects _separate_. There is obviously integration stuff, but they \nshould still be seen as truly independent projects. Even the supermodule \nshould have clearly its own life even _regardless_ of submodules, because \n(as I said) quite often you may want the supermodule, but you don't want \nto have _all_ of the submodules.\n\nBut it's more than that stability and scalability thing too - keeping them \nseparate is what allows you to do pulls and pushes on an individual \nsubproject basis, and have people really work at that level. For example, \nif you're the compiler guy at a company, you really do want to work with \nother compiler people _outside_ the company, but you sure as hell may not \nbe able to give them access to your supermodule. But you may want to work \non _just_ the compiler parts (or at least share some branches in public), \nwhich means that the subproject really has to be able to work \n_independently_ of the supermodule.\n\nSo \"independent\" here is really key, for several reasons. And that all \nmeans, for example, that here must NEVER be any \"backpointers\". A \nsubproject really can _never_ have backpointers to the superproject, \nbecause that fundamentally means that the above kind of \"compiler guy \nworks on the compiler subproject in public\" cannot work, if your \nsupermodule isn't public.\n\n"},{"id":"294044","messageId":"20061201203002.GJ18810@admingilde.org","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612011134080.3695@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-01T20:30:02Z","receivedAt":"2006-12-01T20:30:02Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nLinus, you are a lot better in describing all my thoughts than I myself.\n;-)\n\n-- \nMartin Waitz\n"},{"id":"295873","messageId":"200612012104.39897.andyparkins@gmail.com","threadId":"43065","inReplyTo":"20061201193802.GI18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2006-12-01T21:04:37Z","receivedAt":"2006-12-01T21:04:37Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Friday 2006, December 01 19:38, Martin Waitz wrote:\n\n> On Fri, Dec 01, 2006 at 07:17:17PM +0000, Andy Parkins wrote:\n> > In fact, why should the submodule commits be even visible in the\n> > supermodule?  That tree->submodule commit is sufficient; there isn't\n> > any need to view submodule history in the supermodule.\n>\n> Well, but there is a need for a common object traversal.\n> You need that when sending all objects between two supermodule versions\n> and also when you determine which objects are still reachable.\n\nNo you don't; when traversing the supermodule history you will come across \ntrees that have submodule commit hashes in them, that is all the other end \nneeds to know.  If it wants it can then connect to the submodule and clone \nsubmodule to submodule.  The whole operation doesn't have to be done in the \nsupermodule though.\n\n> The easiest way to implement the common object traversal is to have all\n> objects in one object repository.\n\nThat's true; but is it the right way?  I really really think the submodule \nobjects should be in the submodule itself.\n\n> It may be possible to use two object stores and still do the common\n> object traversal but I do not think that gives you any benefits.\n\nThere is one benefit - you can git-clone the submodule just as you would if it \nwere not a submodule.  In fact, from the submodule's point of view it knows \nnothing about the supermodule.\n\n> You still don't have a totally separated repository then, because\n> you can't do a reachability analysis in the submodule repository alone.\n\nI'm going to guess by reachability analysis, you mean that the submodule \ndoesn't know that some of it's commits are referenced by the supermodule.  As \nI suggested elsewhere in the thread, that's easily fixed by making a \nrefs/supermodule/commitXXXX file for each supermodule commit that references \nas particular submodule commit.  Then you can git-prune, git-fsck whenever \nyou want.\n\n\nAndy\n-- \nDr Andrew Parkins, M Eng (Hons), AMIEE\n"},{"id":"298333","messageId":"20061201213722.GK18810@admingilde.org","threadId":"43065","inReplyTo":"200612012104.39897.andyparkins@gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-01T21:37:23Z","receivedAt":"2006-12-01T21:37:23Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Fri, Dec 01, 2006 at 09:04:37PM +0000, Andy Parkins wrote:\n> > It may be possible to use two object stores and still do the common\n> > object traversal but I do not think that gives you any benefits.\n> \n> There is one benefit - you can git-clone the submodule just as you\n> would if it were not a submodule.  In fact, from the submodule's point\n> of view it knows nothing about the supermodule.\n\nThe submodule repository obviously has to able to reach all its objects.\nThis is easily doable with the shared object database.\n\nSo you can already clone the submodule standalone.\n\n> I'm going to guess by reachability analysis, you mean that the\n> submodule doesn't know that some of it's commits are referenced by the\n> supermodule.  As I suggested elsewhere in the thread, that's easily\n> fixed by making a refs/supermodule/commitXXXX file for each\n> supermodule commit that references as particular submodule commit.\n\nI wouldn't call this \"easily\".\n\n-- \nMartin Waitz\n"},{"id":"297391","messageId":"200612012154.33834.andyparkins@gmail.com","threadId":"43065","inReplyTo":"20061201213722.GK18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2006-12-01T21:54:32Z","receivedAt":"2006-12-01T21:54:32Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Friday 2006, December 01 21:37, Martin Waitz wrote:\n\n> > I'm going to guess by reachability analysis, you mean that the\n> > submodule doesn't know that some of it's commits are referenced by the\n> > supermodule.  As I suggested elsewhere in the thread, that's easily\n> > fixed by making a refs/supermodule/commitXXXX file for each\n> > supermodule commit that references as particular submodule commit.\n>\n> I wouldn't call this \"easily\".\n\nOf course it is; when you write a supermodule commit you have it's hash, \n$SUPERMODULE_HASH, you have the commit-hash of the submodule commit you're \nreferencing, $SUBMODULE_HASH.  It's not really hard to do\n\necho $SUBMODULE_HASH > \nsubmodule/.git/refs/supermodules/commit$SUPERMODULE_HASH\n\nIs it?\n\n\nAndy\n\n-- \nDr Andrew Parkins, M Eng (Hons), AMIEE\n"},{"id":"296070","messageId":"200612012306.41410.Josef.Weidendorfer@gmx.de","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612011134080.3695@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-01T22:06:40Z","receivedAt":"2006-12-01T22:06:40Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Friday 01 December 2006 21:13, Linus Torvalds wrote: \n> > There seems to be some agreement to store the commit id of\n> > the submodule instead of a plain tree id in the supermodules tree object, and\n> > that all objects that are reachable from this commit are made part of the\n> > supermodule repository (either fetched or via alternates). Do you agree?\n> \n> Well, I would actually argue that you may often want to have a supermodule \n> and then at least have the _option_ to decide to not fetch all the \n> submodules.\n\nIf you want to allow this, you have to be able to cut off fetching the\nobjects of the supermodule at borders to given submodules, the ones you\ndo not want to track. With \"border\" I mean the submodule commit in some\ntree of the supermodule.\n\nThis looks a little bit like a shallow clone, where you introduce\ngraft points at the border to some of the submodule's object DAGs.\nBut I am not sure that this is scalable: for supermodules with\na large number of submodules you are not interested in,\nyour graft file would grow very fast, as there will be new borders\nwith every change in some submodule, which happens to be tracked\nin the supermodule.\n\nSo IMHO, instead of a huge graft file, you want to have a fast way\nto check at a submodule border which submodule this given border is\ngoing into. Then, at fetch time, you easily can decide that you do\nnot want to fetch any object from the submodule.\nOtherwise, you would have to ask the remote end at cloning time:\n\"Is this commit from some submodule I am locally not interested in?\"\n\nSo I think we should introduce a submodule namespaces in supermodules.\nAnd at every border from super- to submodules, the name of the\nsubmodule we are going into should be specified.\nWhich actually means that we need to introduce a \"submodule\" object,\nand trees of a supermodule can have such submodule objects as borders\ninto a submodule. In a submodule object, of course we have the\nSHA1 of the commit into the submodule DAG, and there would be the global\nunique name we have choosen for this submodule in this supermodule.\nSomething like\n\n submodule: gcc\n commit: 6287376...\n\nBefore cloning a supermodule, you should be able to list the names of\nthe submodules available, and select the submodules you want to have\ncloned together with the supermodule.\n\n> Ie, you very much should be able to to do\n>\n>         git clone supermodule/that/one/submodule\n> my-own-version-of-submodule\n>\n> to clone a submodule _without_ getting anything else (but still get all\n> the work you did within he submodule - very much including your own\n> private branch work).\n\nSo in the example, \"that/one/submodule\" is _not_ the path of the working\ntree which happens to be the root of the submodule at current supermodule\nHEAD, but the unique name from the submodule namespace.\n\nThis is important, as you should be able to move the root of a submodule\ninside of your supermodule like moving any other file or directory.\nI.e. for every supermodule commit, the path to the root directory of a\ngiven submodule can change, making it useless as a name for a submodule\nselection at clone time.\n\n"},{"id":"296543","messageId":"20061201220821.GL18810@admingilde.org","threadId":"43065","inReplyTo":"200612012154.33834.andyparkins@gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-01T22:08:21Z","receivedAt":"2006-12-01T22:08:21Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Fri, Dec 01, 2006 at 09:54:32PM +0000, Andy Parkins wrote:\n> On Friday 2006, December 01 21:37, Martin Waitz wrote:\n> \n> > > I'm going to guess by reachability analysis, you mean that the\n> > > submodule doesn't know that some of it's commits are referenced by the\n> > > supermodule.  As I suggested elsewhere in the thread, that's easily\n> > > fixed by making a refs/supermodule/commitXXXX file for each\n> > > supermodule commit that references as particular submodule commit.\n> >\n> > I wouldn't call this \"easily\".\n> \n> Of course it is; when you write a supermodule commit you have it's hash, \n> $SUPERMODULE_HASH, you have the commit-hash of the submodule commit you're \n> referencing, $SUBMODULE_HASH.  It's not really hard to do\n> \n> echo $SUBMODULE_HASH > \n> submodule/.git/refs/supermodules/commit$SUPERMODULE_HASH\n\nI guess you are aware that you have to scan _all_ trees inside _all_\nsupermodule commits for possible references.\n\nSo what do you do with deleted submodules?\nYou wouldn't want them to still sit around in your working directory,\nbut you still have to preserve them.\n\n-- \nMartin Waitz\n"},{"id":"296773","messageId":"20061201221230.GM18810@admingilde.org","threadId":"43065","inReplyTo":"200612012306.41410.Josef.Weidendorfer@gmx.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-01T22:12:30Z","receivedAt":"2006-12-01T22:12:30Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Fri, Dec 01, 2006 at 11:06:40PM +0100, Josef Weidendorfer wrote:\n> > Well, I would actually argue that you may often want to have a\n> > supermodule and then at least have the _option_ to decide to not\n> > fetch all the submodules.\n> \n> If you want to allow this, you have to be able to cut off fetching the\n> objects of the supermodule at borders to given submodules, the ones you\n> do not want to track. With \"border\" I mean the submodule commit in some\n> tree of the supermodule.\n\nI don't think this is something special to submodules.  There has been\ninterest in checking out only a part of the tree even before talking\nabout submodules and I really think this feature should be independent\nto submodules.\n\n-- \nMartin Waitz\n"},{"id":"293820","messageId":"200612012326.22779.Josef.Weidendorfer@gmx.de","threadId":"43065","inReplyTo":"20061201221230.GM18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-01T22:26:22Z","receivedAt":"2006-12-01T22:26:22Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Friday 01 December 2006 23:12, Martin Waitz wrote:\n> hoi :)\n> \n> On Fri, Dec 01, 2006 at 11:06:40PM +0100, Josef Weidendorfer wrote:\n> > > Well, I would actually argue that you may often want to have a\n> > > supermodule and then at least have the _option_ to decide to not\n> > > fetch all the submodules.\n> > \n> > If you want to allow this, you have to be able to cut off fetching the\n> > objects of the supermodule at borders to given submodules, the ones you\n> > do not want to track. With \"border\" I mean the submodule commit in some\n> > tree of the supermodule.\n> \n> I don't think this is something special to submodules.  There has been\n> interest in checking out only a part of the tree even before talking\n> about submodules and I really think this feature should be independent\n> to submodules.\n\nIt's not about checking out part of the tree, it's about fetching only\npart of the objects: If you have a slow modem and want to clone a\nsupermodule, you are not interested in fetching all the objects from\nsome submodules.\n\nSo it is more like a shallow clone. But even here, submodules are special\nas you have defined borders between supermodule and submodules. This gives\nyou the freedom the introduce a submodule namespace, and allows you to\npoint to a submodule: \"I do not want you!\".\n\nWith shallow clone, there you do not have this option, so there, you need\nto use something like grafting.\n\nBTW: In your submodule implementation, is the user allowed to change the\nrelative path of the root of some submodule, e.g. with \"git-mv\" ?\n\nJosef\n"},{"id":"294821","messageId":"Pine.LNX.4.64.0612011423100.3695@woody.osdl.org","threadId":"43065","inReplyTo":"200612012306.41410.Josef.Weidendorfer@gmx.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-01T22:26:51Z","receivedAt":"2006-12-01T22:26:51Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 1 Dec 2006, Josef Weidendorfer wrote:\n> > \n> > Well, I would actually argue that you may often want to have a supermodule \n> > and then at least have the _option_ to decide to not fetch all the \n> > submodules.\n> \n> If you want to allow this, you have to be able to cut off fetching the\n> objects of the supermodule at borders to given submodules, the ones you\n> do not want to track. With \"border\" I mean the submodule commit in some\n> tree of the supermodule.\n>\n> This looks a little bit like a shallow clone\n\nNo. \n\nI would say that it looks more like a \"partial checkout\" than a shallow \nclone.\n\nA shallow clone limits the data in \"time\" - we have _some_ data, but we \ndon't have all of the history of that data.\n\nIn contrast, a submodule that we don't fetch is an all-or-nothing \nsituation: we simply don't have the data at all, and it's really a matter \nof simply not recursing into that submodule at all - much more like not \nchecking out a particular part of the tree.\n\nSo if a shallow clone is a \"limit in time\", a lack of a module (or a lack \nof a checkout for a subtree in general - you could certainly imagine doing \nthe same thing even _within_ a git repository, and indeed, we did discuss \nexactly that at one point in time) is more of a \"limit in space\".\n\n"},{"id":"293942","messageId":"4570AE4E.8000600@stephan-feder.de","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612011134080.3695@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"sf","fromEmail":"sf-gmane@stephan-feder.de","sentAt":"2006-12-01T22:35:58Z","receivedAt":"2006-12-01T22:35:58Z","isPatch":false,"sender":{"key":"sf-gmane@stephan-feder.de","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Fri, 1 Dec 2006, sf wrote:\n>> Linus Torvalds wrote:\n>> ...\n>>> Think of it this way: one common use for submodules is really to just\n>>> (occasionally) track somebody elses code. The submodule should be a\n>>> totally pristine copy from somebody else (ie it might be the \"intel driver\n>>> for X.org\" submodule, maintained within intel), and the supermodule just\n>>> refers to it indirectly (ie the supermodule might be the \"Fedora Core X\n>>> group\" which contains all the different drivers from different people).\n>> Could you please be a little bit more specific about how you would store the\n>> \"pristine copy\".\n> \n> Note that it's not necessarily \"pristine\", since the submodule clearly is \n> a local git repository in its own right. So like _any_ git repository, you \n> can (and may well end up) having your own local branches in the submodule, \n> with your own local modifications.\n> \n> So I'm not claiming that a submodule must always match some external git \n> tree 100%, and that it must be read-only or anything like that. I'm just \n> saying that I suspect that quite often, one of the MOST IMPORTANT parts is \n> that the submodule is really something that somebody else technically \n> maintains, and that this is actually one of the _reasons_ why it is a \n> submodule in the first place. \n> \n> For example, a lot of projects end up having some kind of \"library \n> component\" as a submodule. Take something like a video player project, \n> which would have something like ffmpeg as a submodule, not because you'd \n> maintain ffmpeg yourself, but simply because (let's say) the library \n> interface changes enough, or you need a specific version with some of your \n> own fixes that haven't been released widely yet, so you want to carry all \n> the libraries you need _with_ you, even though you don't really maintain \n> that submodule. You at most have some small extensions of your own.\n> \n> Now, in this situation, it's relaly really _important_ that the submodule \n> really is totally independent of the supermodule, for several reasons.\n> \n> For example, since you don't \"really\" own that project, carrying around \n> your own fixes is really really painful. We know it happens all the time, \n> and a lot of projects end up needing their own version, but the _last_ \n> thing you want is to be in merge hell all the time. So as a supermodule \n> maintainer, the best possible thing for you is to be able to push back \n> those local changes to the original project maintainer, so that you \n> _don't_ have to maintain your own changes.\n\nTrue. But if you need the changes to the submodule for your supermodule\nto function, and upstream either does not want to merge your changes or\nthe merge will be available only after a long time, then what is the\nalternative? You must be able to keep local changes, and you must be\nable to keep pulling from upstream. Of course, what you describe is the\nideal case: You find a bug, push the fix upstream, and in no time at all\nyour fix is merged and you can just pull a new version into your\nsuperproject, but that might be wishful thinking.\n\n> But you need to realize that the real maintainer of the submodule is \n> TOTALLY UNINTERESTED in your supermodule. He's not going to maintain it, \n> and in fact, if you have anything in the submodule that ends up talking \n> about your supermodule, that's just going to make it a lot less likely \n> that the upstream maintainer will ever pull your changes. He might take a \n> diff from you, but in a perfect world, you'd actually be able to tell him: \n> \n>  \"Hey, I've got a git repository with a few fixes to your ffmpeg git tree, \n>   please pull from git://myhost.com/submodule.git to get these fixes:\n> \n> \t... explanation of fixes and commits that are relevant to\n> \tffmpeg, and have nothing to do with the supermodule, except\n> \tthat you need those bug-fixes because you _use_ ffmpeg ...\n> \n>   Thanks\"\n> \n> See?\n\nNo! All you need is a naming scheme to address the commit of the\nsubproject that should be pulled. The extreme case would be to just\naddress it with its id (well, currently you cannot do that with git\npull, but that is fixable). But I already proposed a syntax for naming\ncommits which are \"hidden\" in a superproject: Just name the path as\ndescribed in git-rev-parse and append double slashes (to indicate that\nyou mean the commit, not the tree it contains). So no manual work needs\nbe done by upstream.\n\n[snipped: about independence of submodule branches]\n\n>> There seems to be some agreement to store the commit id of\n>> the submodule instead of a plain tree id in the supermodules tree object, and\n>> that all objects that are reachable from this commit are made part of the\n>> supermodule repository (either fetched or via alternates). Do you agree?\n> \n> Well, I would actually argue that you may often want to have a supermodule \n> and then at least have the _option_ to decide to not fetch all the \n> submodules.\n\n[transmeta example snipped]\n\n> So no, I don't think the supermodule should even _force_ people to always \n> get all the submodules. It migth be the default case, but at the same \n> time, it's just being polite to let users decide on their own whether they \n> really want _all_ of the build infrastructure sources.\n\nIf you want to track some chosen submodules there are two easy solutions:\n\n1. If you want to track their state as it appears from the supermodule's\nview, pull from master:<submodule>//\n2. If you want to track their state from their own development branches,\n pull from <submodule>/master\n\nCan you see the difference?\n\n>> If I understand you correctly you cannot make any changes to the submodules\n>> code _in the supermodule's repository_, no bugfixes, no extensions, no\n>> adaptions, nothing. Do you mean that?\n> \n> Yes. I think you should make all changes _within_ the submodule, because \n> the submodule should still be an independent git tree in its own right.\n\nEvery commit is a git tree in its own right, is it not?\n\n[description of independent submodule development snipped]\n\n> And the importance of keeping the submodule independent is partly just \n> stability and sanity, but partly also scalability. For example, the \n> \"index\" in a supermodule should NOT include the indexes of all the \n> submodules. That's really important, because the index doesn't really \n> scale. Things do slow down with large indexes. \n> \n> For example, git can handle tens of thousands of files easily. I suspect \n> it scales well to hundreds of thousands of filenames. But with \n> supermodules, you really can end up in the situation where you have _tens_ \n> of these submodules, maybe even hundreds. And if you try to maintain one \n> unified index for the _whole_ thing, I guarantee you that you'll start \n> feeling the pain. Indexing millions of files is just not going to be \n> pretty.\n\nI am not sure I understand what you say.\n\n1. If you are working on a submodule, then the supermodule never enters\nthe picture. You are working independently. So far, so good.\n\n2. If you are working on the supermodule, git will not be able to\nfunction? How would you work without submodules, in which case you would\n have simply one large project?\n\n> So just from a git stability and scalability point, it's important to keep \n> subprojects _separate_. There is obviously integration stuff, but they \n> should still be seen as truly independent projects. Even the supermodule \n> should have clearly its own life even _regardless_ of submodules, because \n> (as I said) quite often you may want the supermodule, but you don't want \n> to have _all_ of the submodules.\n> \n> But it's more than that stability and scalability thing too - keeping them \n> separate is what allows you to do pulls and pushes on an individual \n> subproject basis, and have people really work at that level. For example, \n> if you're the compiler guy at a company, you really do want to work with \n> other compiler people _outside_ the company, but you sure as hell may not \n> be able to give them access to your supermodule. But you may want to work \n> on _just_ the compiler parts (or at least share some branches in public), \n> which means that the subproject really has to be able to work \n> _independently_ of the supermodule.\n\nI totally agree. When I try to explain why submodules work that only\nexist as part of one or more supermodules, I do not mean to say that you\ncannot or should not have independent branches or repositories for the\nsubmodules' code.\n\n> So \"independent\" here is really key, for several reasons. And that all \n> means, for example, that here must NEVER be any \"backpointers\". A \n> subproject really can _never_ have backpointers to the superproject, \n> because that fundamentally means that the above kind of \"compiler guy \n> works on the compiler subproject in public\" cannot work, if your \n> supermodule isn't public.\n\nI took that for granted: from a commit you only ever look backwards (in\ntime/history dimension) or downwards (in content dimension).\n\nRegards\n\n"},{"id":"295336","messageId":"20061201224021.GN18810@admingilde.org","threadId":"43065","inReplyTo":"200612012326.22779.Josef.Weidendorfer@gmx.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-01T22:40:22Z","receivedAt":"2006-12-01T22:40:22Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Fri, Dec 01, 2006 at 11:26:22PM +0100, Josef Weidendorfer wrote:\n> It's not about checking out part of the tree, it's about fetching only\n> part of the objects: If you have a slow modem and want to clone a\n> supermodule, you are not interested in fetching all the objects from\n> some submodules.\n\nSo when you want to suppress one submodule, how is this not about only\nchecking out part of the tree?\nOk, you also want to avoid downloading the submodule, but you first have\nto solve the partial checkout.\n\n> BTW: In your submodule implementation, is the user allowed to change the\n> relative path of the root of some submodule, e.g. with \"git-mv\" ?\n\nIn principle: yes.\nHowever there are some links between both repositories that have to be\nupdated manually (for the shared object repository and for ignoring\nsubmodule files in the supermodule).\nBut I expect that much of this configuration stuff will vanish when\nsubmodules are better integrated in git.\n\nRename detection for submodules would be another interesting thing to\nhave. It should be much easier as for files because we can simply check\nfor common ancestors and do not have to guess based on the diff.\n\n-- \nMartin Waitz\n"},{"id":"298499","messageId":"4570AF8F.1000801@stephan-feder.de","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612011423100.3695@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"sf","fromEmail":"sf-gmane@stephan-feder.de","sentAt":"2006-12-01T22:41:19Z","receivedAt":"2006-12-01T22:41:19Z","isPatch":false,"sender":{"key":"sf-gmane@stephan-feder.de","avatar":null},"body":"Linus Torvalds wrote:\n...\n> In contrast, a submodule that we don't fetch is an all-or-nothing \n> situation: we simply don't have the data at all, and it's really a matter \n> of simply not recursing into that submodule at all - much more like not \n> checking out a particular part of the tree.\n\nIf you do not want to fetch all of the supermodule then do not fetch the\nsupermodule. Instead fetch only the submodules you are interested in.\nYou do not have to fetch the whole repository.\n\nRegards\n\n"},{"id":"295884","messageId":"200612012355.03493.Josef.Weidendorfer@gmx.de","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612011423100.3695@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-01T22:55:03Z","receivedAt":"2006-12-01T22:55:03Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Friday 01 December 2006 23:26, Linus Torvalds wrote:\n> \n> On Fri, 1 Dec 2006, Josef Weidendorfer wrote:\n> > > \n> > > Well, I would actually argue that you may often want to have a supermodule \n> > > and then at least have the _option_ to decide to not fetch all the \n> > > submodules.\n> > \n> > If you want to allow this, you have to be able to cut off fetching the\n> > objects of the supermodule at borders to given submodules, the ones you\n> > do not want to track. With \"border\" I mean the submodule commit in some\n> > tree of the supermodule.\n> >\n> > This looks a little bit like a shallow clone\n> \n> No. \n> \n> I would say that it looks more like a \"partial checkout\" than a shallow \n> clone.\n> \n> A shallow clone limits the data in \"time\" - we have _some_ data, but we \n> don't have all of the history of that data.\n> \n> In contrast, a submodule that we don't fetch is an all-or-nothing \n> situation: we simply don't have the data at all, and it's really a matter \n> of simply not recursing into that submodule at all - much more like not \n> checking out a particular part of the tree.\n\nOK.\n\nI still think it should be about \"limit in space\" regarding the\nobjects in the local repository.\n\nFor a project containing \"gcc\" as submodule, and I am not\ninterested in this submodule, there should be a way to not need\nto fetch all the objects from the gcc submodule at clone time.\n\n\nWhat about my other argument for a submodule namespace:\nYou want to be able to move the relative root path of a submodule\ninside of your supermodule, but yet want to have a unique name\nfor the submodule:\n- to be able to just clone a submodule without having to know\nthe current position in HEAD\n- more practically, e.g. to be able to name a submodule\nindependent from any current commit you are on in the supermodule,\ne.g. to be able to store some meta information about a submodule:\n- \"Where is the official upstream of this submodule?\"\n- \"Should git allow to commit rewind actions of this submodule\n   in the supermodule?\" (which, AFAICS, exactly has the same\n   problems as publishing a rewound branch: you will get into\n   merge hell when you want to pull upstream changes into the\n   supermodule)\n- \"Should this submodule be checked out?\"\nand so on.\n\n"},{"id":"295370","messageId":"200612020003.14377.Josef.Weidendorfer@gmx.de","threadId":"43065","inReplyTo":"4570AF8F.1000801@stephan-feder.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-01T23:03:14Z","receivedAt":"2006-12-01T23:03:14Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Friday 01 December 2006 23:41, sf wrote:\n> Linus Torvalds wrote:\n> ...\n> > In contrast, a submodule that we don't fetch is an all-or-nothing \n> > situation: we simply don't have the data at all, and it's really a matter \n> > of simply not recursing into that submodule at all - much more like not \n> > checking out a particular part of the tree.\n> \n> If you do not want to fetch all of the supermodule then do not fetch the\n> supermodule. Instead fetch only the submodules you are interested in.\n> You do not have to fetch the whole repository.\n\nBut what, when I *want* to fetch the supermodule because of the source\ntree which is only available in the supermodule?\n\nOf course, you can argue that the only objects in trees of a supermodule\nshould be submodule commits, but this is quite restricting the usage of\nsupermodules.\n\nSee further arguments for a submodule namespace in my other mail.\nYou probably want to specify policies about submodule handling. This\ninformation has to be indexed by some name independent from a \nsupermodule commit.\n\nJosef\n"},{"id":"295700","messageId":"20061201230732.GO18810@admingilde.org","threadId":"43065","inReplyTo":"200612012355.03493.Josef.Weidendorfer@gmx.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-01T23:07:32Z","receivedAt":"2006-12-01T23:07:32Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Fri, Dec 01, 2006 at 11:55:03PM +0100, Josef Weidendorfer wrote:\n> What about my other argument for a submodule namespace:\n> You want to be able to move the relative root path of a submodule\n> inside of your supermodule, but yet want to have a unique name\n> for the submodule:\n> - to be able to just clone a submodule without having to know\n> the current position in HEAD\n> - more practically, e.g. to be able to name a submodule\n> independent from any current commit you are on in the supermodule,\n> e.g. to be able to store some meta information about a submodule:\n> - \"Where is the official upstream of this submodule?\"\n\nyou can always have a bare repository for all used modules lying around\nin some defined location.  There is no need for a unique submodule-name.\n\n-- \nMartin Waitz\n"},{"id":"297708","messageId":"Pine.LNX.4.64.0612011505190.3695@woody.osdl.org","threadId":"43065","inReplyTo":"4570AF8F.1000801@stephan-feder.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-01T23:09:40Z","receivedAt":"2006-12-01T23:09:40Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 1 Dec 2006, sf wrote:\n> Linus Torvalds wrote:\n> ...\n> > In contrast, a submodule that we don't fetch is an all-or-nothing \n> > situation: we simply don't have the data at all, and it's really a matter \n> > of simply not recursing into that submodule at all - much more like not \n> > checking out a particular part of the tree.\n> \n> If you do not want to fetch all of the supermodule then do not fetch the\n> supermodule.\n\nSo why do you want to limit it? There's absolutely no cost to saying \"I \nwant to see all the common shared infrastructure, but I'm actually only \ninterested in this one submodule that I work with\".\n\nAlso, anybody who works on just the build infrastructure simply may not \ncare about all the submodules. The submodules may add up to hundreds of \ngigs of stuff. Not everybody wants them. But you may still want to get the \ncommon build infrastructure.\n\nIn other words, your \"all or nothing\" approach is\n (a) not friendly\nand\n (b) has no real advantages anyway, since modules have to be independent \n     enough that you _can_ split them off for other reasons anyway.\n\nSo forcing that \"you have to take everything\" mentality onyl has \nnegatives, and no positives. Why do it?\n\n"},{"id":"298088","messageId":"200612020017.44275.Josef.Weidendorfer@gmx.de","threadId":"43065","inReplyTo":"20061201221230.GM18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-01T23:17:44Z","receivedAt":"2006-12-01T23:17:44Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Friday 01 December 2006 23:12, Martin Waitz wrote:\n> hoi :)\n> \n> On Fri, Dec 01, 2006 at 11:06:40PM +0100, Josef Weidendorfer wrote:\n> > > Well, I would actually argue that you may often want to have a\n> > > supermodule and then at least have the _option_ to decide to not\n> > > fetch all the submodules.\n> > \n> > If you want to allow this, you have to be able to cut off fetching the\n> > objects of the supermodule at borders to given submodules, the ones you\n> > do not want to track. With \"border\" I mean the submodule commit in some\n> > tree of the supermodule.\n> \n> I don't think this is something special to submodules.  There has been\n> interest in checking out only a part of the tree even before talking\n> about submodules and I really think this feature should be independent\n> to submodules.\n\nAfter some thinking, a submodule namespace even is important for checking\nout only parts of a supermodule, exactly because the root of a submodule\npotentially can change at every commit.\n\nWhen checking out some arbitrary supermodule commit, how do you check that\nat some submodule border, the user did not want to check out the submodule\nat all? You need a way to check the DAG identity you are diving\ninto at this border: lets say by going to the root commit of this DAG (!).\nAnd via this identity, you have to check whether the user had\nspecified that he wants the submodule to be check out.\nWithout any further meta information (indexed by a submodule name!), this\ninformation is only available from the checkout the user switched from,\nas there would be no file in the working tree from this submodule?\n\nQuite a pain.\n\n"},{"id":"297059","messageId":"200612012323.54403.alan@chandlerfamily.org.uk","threadId":"43065","inReplyTo":"20061201203002.GJ18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Alan Chandler","fromEmail":"alan@chandlerfamily.org.uk","sentAt":"2006-12-01T23:23:54Z","receivedAt":"2006-12-01T23:23:54Z","isPatch":false,"sender":{"key":"alan@chandlerfamily.org.uk","avatar":"https://gravatar.com/avatar/1862247e5ea8eac114c842f9dc3a5db6253754e24ef7171757cf97eedce48b8c?d=mp&s=160"},"body":"On Friday 01 December 2006 20:30, Martin Waitz wrote:\n> hoi :)\n>\n> Linus, you are a lot better in describing all my thoughts than I myself.\n> ;-)\n\nYes\n\nSome of the clearest explanations described from a strategic point of view \ncome from posts such as these. \n\nAnd he keeps saying he can't write documentation.\n\n \n-- \nAlan Chandler\n"},{"id":"294210","messageId":"Pine.LNX.4.64.0612011510290.3695@woody.osdl.org","threadId":"43065","inReplyTo":"200612012355.03493.Josef.Weidendorfer@gmx.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-01T23:30:32Z","receivedAt":"2006-12-01T23:30:32Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 1 Dec 2006, Josef Weidendorfer wrote:\n> \n> What about my other argument for a submodule namespace:\n> You want to be able to move the relative root path of a submodule\n> inside of your supermodule, but yet want to have a unique name\n> for the submodule:\n> - to be able to just clone a submodule without having to know\n> the current position in HEAD\n\nUmm? I don't get the issue. A submodule is a git repo in its own right, \nand you clone it exactly like you'd clone any other repo. It _does_ have a \nHEAD. It has it's own branches. It has everything.\n\nSo when you clone a submodule, you always get all those branches. The \nsupermodule will not _point_ to them all (the branches are local to the \nsubmodule, and _will_ depend on things like \"which upstreams module am I \ntracking\"), but they'll have to be there, exactly _because_ the submodule \nhas an existence and is tracked on its own.\n\nIn the trivial case where the submodule doesn't even _have_ any external \nexistence at all (ie it's always maintained as _just_ a submodule, it \nwould probably tend to have just one branch, and a clone would get \nwhatever that branch is), but that's just a degenerate special case of the \nmuch richer \"this submodule actually has a life of its own\" case.\n\n> - more practically, e.g. to be able to name a submodule\n> independent from any current commit you are on in the supermodule,\n> e.g. to be able to store some meta information about a submodule:\n\nThe current commit within the supermodule would be _totally_ invisible to \nthe submodule.\n\nOf course, if HEAD _differs_ from that commit within the supermodule, then \na \"git diff\" (when done from within the supermodule) should show that, but \nagain, that's actually only as seen from the _supermodule_. \n\n> - \"Where is the official upstream of this submodule?\"\n\nThat's entirely a question for the submodule. You cannot ask that question \nwithin the confines of the supermodule, because it's not even a relevant \nquestion in that context. Two different supermodule repositories may well \ndecide to get their submodules from difference places, just because they \ngot cloned from different places (or even just for practical reasons like \n\"that other site is closer to me\").\n\nSo the official upstream of a submodule must NOT be encoded inside the \nsupermodule (or at least not within its _objects_). Exactly because the \nupstream location is not a \"global\" thing - it's per-repository, and thus \nmust not be encoded in the global data (ie the objects).\n\nIt should be be encoded in some _ephemeral_ place, eg in the \".git/config\" \nfile or in a \".git/remotes/origin\"-like file (either in the supermodule or \nthe submodule, and I would seriously suggest you do it within in the \nsubmodule itself, because you'll want it exactly when you decide to work \non the submodule and upgrade _that_).\n\n> - \"Should git allow to commit rewind actions of this submodule\n>    in the supermodule?\" (which, AFAICS, exactly has the same\n>    problems as publishing a rewound branch: you will get into\n>    merge hell when you want to pull upstream changes into the\n>    supermodule)\n\nThe only thing that a submodule must NOT be allowed to do on its own is \npruning (and it's distant cousin \"git repack -d\"). You must always prune \nfrom the supermodule, because the submodule cannot really know on its own \nwhat references point into it.\n\n(There are alternatives. One alternative is to never allow rewinding - or \ndeletion - of branches in a submodule, and thus solve the problem that \nway. That is the easier solution, because it also means that a \"clone\" of \na supermodule can just recursively clone the submodules independently \n_without_ having to worry about reachability, but it's really _really_ \ndraconian).\n\n> - \"Should this submodule be checked out?\"\n\nThis, I think, requires too much configuration to say separately for every \npossible submodule, so I would suggest that the way to make that decision \nis:\n\n - \"git clone\" by default will fetch and check out all submodules (and \n   obviously they have to be described some way outside of the object \n   database, just so that you don't have to parse the _whole_ history of \n   the _whole_ supermodule just to find all possible submodules. So the \n   supermodule _will_ need some \"list of submodules and where to get them\" \n   in a config file or other).\n\n - add a flag (possibly just re-use the current \"-n\" flag) that disables \n   that recursive fetching of submodules entirely.\n\n - have a way to fetch individual submodules one-by-one (that capacity \n   obviously has to be there anyway, since the \"recursive\" git clone has \n   to be able to do it, so this is likely just \"git clone\" again, with \n   just logic added to say \"when you clone something and are _already_ \n   within a superproject, the clonee becomes a subproject automatically\"\n\nI dunno. And I'd also like to point out that things don't have to all work \nfully before we can do at least some cases of this. For example, if the \ninitial version just always clones everything, big deal. I'm not saying \nthat we have to have support for things like this on \"Day 1\", I'm just \nsaying that I think people will want to be able to not fetch and check out \neverything, so the design should _allow_ for it.\n\n(But I also think that as long as submodules are independent enough, the \n\"design\" part should fall out on its own, and it just becomes a \"small \nmatter of programming\" to actually get it to work).\n\n"},{"id":"295717","messageId":"4570BC07.4080203@stephan-feder.de","threadId":"43065","inReplyTo":"20061201184801.GG18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"sf","fromEmail":"sf-gmane@stephan-feder.de","sentAt":"2006-12-01T23:34:31Z","receivedAt":"2006-12-01T23:34:31Z","isPatch":false,"sender":{"key":"sf-gmane@stephan-feder.de","avatar":null},"body":"Martin Waitz wrote:\n> On Fri, Dec 01, 2006 at 06:33:12PM +0100, Stephan Feder wrote:\n>> We are in agreement about two fundamental parts of the implementation \n>> and their meaning:\n>>\n>> 1. A submodule is stored as a commit id in a tree object.\n>>\n>> 2. Every object that is reachable from the submodule's commit are \n>> reachable from the supermodule's repository.\n> \n> Correct.\n\nGood. For me that is the main point. As I said before the user interface\nis not so important because it can be changed anytime, but to change the\nobject database later is close to impossible.\n\n...\n> Let's see if I understand you correctly:\n> \n> You don't want to create an additional .git directory for the submodule\n> and just handle everything with one toplevel .git repository for the\n> whole project.\n\nYes.\n\n> Without the .git directory, you of course do not have refs/heads inside\n> the submodule.\n\nCorrect..\n\n> So this is a different user-interface approach to submodules when\n> compared to my approach.  But the basis is the same and both could\n> inter-operate.\n\nBig YES.\n\n> Now your submodule is no longer seen as an independent git repository\n> and I think this would cause problems when you want to push/pull between\n> the submodule and its upstream repository.\n\nYou can always pick a single commit or several commits out of a larger\nrepository and have a complete git repository.\n\nAnd I already explained how to push and pull even from within superprojects.\n\n> No technical problems, but UI-problems because now your submodule is\n> handled completly different to a \"normal\" repository.\n\nYes and no. You can always have branches that are only concerned with\nsubmodules' code, say, in refs/heads/submodules/<submodule>/.\n\"submodules\" here is simply an example and has not deeper meaning. You\ncould call it foo or whatever you like. Or you could use\nrefs/heads/<submodule>/ if it suits you.\n\nBut if you mean the submodule as seen from the supermodule, then there\nis a difference. Naturally, because the concept of submodules is new to git.\n\n>>> Yes, but if it would have been integrated as a submodule it obviously\n>>> would have been committed to the xdiff submodule inside the git\n>>> repository.\n>> No. The xdiff submodule would only exist as part of the git repository. \n> \n> But you could still call the \"xdiff\" part of the git repository a\n> submodule.  And then changes to the xdiff directory result in a new\n> submodule commit, even when there is no direct reference to it.\n> So you'd still \"commit to the xdiff submodule\".\n\nLet's make certain that we understand each other. I see a clear\ndistinction between the submodule code in a supermodule branch (commits\nin the supermodule's tree and nothing else) and submodule branches which\nare independent of the superproject. Supermodule branches and submodule\nbranches do not interact, only if I want them to.\n\n> \n>> You could, f.e., access the xdiff commit in git HEAD as HEAD:xdiff// \n>> (again my proposed syntax). HEAD:xdiff//~2:xemit.c would give you the \n>> grandparent of xemit.c in the xdiff submodule.\n> \n> git-cat-file commit HEAD:xdiff already works out of the box (even\n> cat-file tree to get the submodule tree).  But up to now revision\n> parsing follows the file name only once.\n> \n> What about just separating things with \"/\"?\n> \n> commit HEAD\n> tree   HEAD/\n> blob   HEAD/Makefile\n> commit HEAD/xdiff\n> tree   HEAD/xdiff/\n> blob   HEAD/xdiff~2/xemit.c\n> \n> this may add some confusion when used with hierarchical branches, but\n> it's still unique:\n> \n> \trefs/heads/master/xdiff/xemit.c\n> \n> Just use as many path components until a matching reference is found,\n> then start peeling.\n> Or just use / between super and submodule:\n> \n> \trefs/heads/master:xdiff/xemit.c\n> \n> I think this is easier to read then\n> \n> \trefs/heads/master:xdiff//:xemit.c\n\nThe double slashes is the only way I can think of that clearly indicates\nthat I do not mean the contents named by the path, but the commit that\nyou find there. Once you have named a commit in that way, you can\ncontinue to apply other revision naming suffixes, paths, and so on.\n\nLet's try. What does git cat-file -p\nmaster:dir/sub//^^^:sub/dir/sub//^:dir/file mean?\n\nExplanation: Take branch master and go to path dir/sub. There you will\nfind a commit. Take its grand-grandparent and go to path sub/dir/sub\n(the first sub is a subproject as well but we do not care). There you\nwill, again, find a commit. Take its parent and go to path dir/file\nwhich happens to be a blob the contents of which you want to cat.\n\nIn reality you will never see these kinds of complex paths. Have you\never seen something like git cat-file -p\nbd2c39f58f915af532b488c5bda753314f0db603~12^{commit}^2^5~8^2~308:README ?\n\n>> If you mean the submodule repository created by init-module I \n>> understand. But why create this \"helper repository at all\"?\n> \n> Because it helps \"normal\" git operations ;-)\n\n\nLet's see. I still have to try.\n\nRegards\n\n"},{"id":"298679","messageId":"200612020036.08826.Josef.Weidendorfer@gmx.de","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612011505190.3695@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-01T23:36:08Z","receivedAt":"2006-12-01T23:36:08Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Saturday 02 December 2006 00:09, Linus Torvalds wrote:\n> \n> On Fri, 1 Dec 2006, sf wrote:\n> > Linus Torvalds wrote:\n> > ...\n> > > In contrast, a submodule that we don't fetch is an all-or-nothing \n> > > situation: we simply don't have the data at all, and it's really a matter \n> > > of simply not recursing into that submodule at all - much more like not \n> > > checking out a particular part of the tree.\n> > \n> > If you do not want to fetch all of the supermodule then do not fetch the\n> > supermodule.\n> \n> So why do you want to limit it? There's absolutely no cost to saying \"I \n> want to see all the common shared infrastructure, but I'm actually only \n> interested in this one submodule that I work with\".\n\nSo you are for a global submodule namespace in supermodule repositories,\ndo I understand correctly?\n\nOtherwise, how would you specify the submodules at clone time given the\nability that submodule roots can have relative path changed arbitrarily\nbetween commits?\n\n"},{"id":"295693","messageId":"4570BFA4.8070903@stephan-feder.de","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612011505190.3695@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"sf","fromEmail":"sf-gmane@stephan-feder.de","sentAt":"2006-12-01T23:49:56Z","receivedAt":"2006-12-01T23:49:56Z","isPatch":false,"sender":{"key":"sf-gmane@stephan-feder.de","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Fri, 1 Dec 2006, sf wrote:\n>> Linus Torvalds wrote:\n>> ...\n>>> In contrast, a submodule that we don't fetch is an all-or-nothing \n>>> situation: we simply don't have the data at all, and it's really a matter \n>>> of simply not recursing into that submodule at all - much more like not \n>>> checking out a particular part of the tree.\n>> If you do not want to fetch all of the supermodule then do not fetch the\n>> supermodule.\n> \n> So why do you want to limit it? There's absolutely no cost to saying \"I \n> want to see all the common shared infrastructure, but I'm actually only \n> interested in this one submodule that I work with\".\n\nIf you need a common infrastructure to be able to work with the\nsubmodule, then the submodule is not independent of of the supermodule.\nI see a contradiction in your requirements.\n\n> Also, anybody who works on just the build infrastructure simply may not \n> care about all the submodules. The submodules may add up to hundreds of \n> gigs of stuff. Not everybody wants them. But you may still want to get the \n> common build infrastructure.\n\nSee above.\n\n> In other words, your \"all or nothing\" approach is\n>  (a) not friendly\n> and\n>  (b) has no real advantages anyway, since modules have to be independent \n>      enough that you _can_ split them off for other reasons anyway.\n> \n> So forcing that \"you have to take everything\" mentality onyl has \n> negatives, and no positives. Why do it?\n\n(There have been lots of use cases for shallow clones but for a long\ntime git did not support them).\n\nIf you can extend this partial fetch feature to the non-subproject case\nI would agree with your reasoning. What makes the subprojects so special\nin this regard. Do I have to turn a plain tree into a subproject to be\nable to ignore it? Once you can restrict fetches to parts of the\ncontents you get the ability to restrict fetches to the \"common\ninfrastructure\" and selected submodules for free.\n\nRegards\n\nStephan\n"},{"id":"296331","messageId":"Pine.LNX.4.64.0612011540010.3695@woody.osdl.org","threadId":"43065","inReplyTo":"200612020036.08826.Josef.Weidendorfer@gmx.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-02T00:12:10Z","receivedAt":"2006-12-02T00:12:10Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 2 Dec 2006, Josef Weidendorfer wrote:\n> \n> So you are for a global submodule namespace in supermodule repositories,\n> do I understand correctly?\n> \n> Otherwise, how would you specify the submodules at clone time given the\n> ability that submodule roots can have relative path changed arbitrarily\n> between commits?\n\nThe only _true_ namespace would be the SHA1 of the commit (and maybe allow \na pointer to a tag too, but the namespace ends up being the same).\n\nHow to _find_ a repository that contains that SHA1 must be left to higher \nlevels. After all, repositories move around, and the place you found them \noriginally is not a stable name.\n\nSo within the supermodule, on a \"git object\" level, a submodule should \njust be named by the SHA1 that was it's HEAD when it was committed within \nthe supermodule. So in the \"tree object\", you'd see something like the \nfollowing when you go \"git ls-tree HEAD\" on the superproject:\n\n\t...\n\t100644 blob 08602f522183dc43787616f37cba9b8af4e3dade\txdiff-interface.c\n\t100644 blob 1346908bea31319aabeabdfd955e2ea9aab37456\txdiff-interface.h\n\t040000 tree 959dd5d97e665998eb26c764d3a889ae7903d9c2\txdiff\n\t050000 link 0215ffb08ce99e2bb59eca114a99499a4d06e704\txyzzy\n\nwhere that 050000 is the new magic type (I picked one out of my *ss: it's \nnot a valid type for a file mode, so it's a godo choice, but it could be \nanythign that cannot conflict with a real file), which just specifies the \n\"link\" part. The SHA1 is the SHA1 of the commit, and the \"xyzzy\" is \nobviously just the name within the directory of the submodule.\n\nThat's all that is actually required for a lot of git commands that \nalready expect all objects to be available (ie \"git checkout\", \"git diff\" \netc).\n\nIt only gets interesting for commands that fetch new objects, ie do a \n\"pull/fetch\" op, and you'd need to know where/how to fetch new objects for \nthe xyzzy subproject, so that's a \"naming\" issue. You have a few choices:\n\n - get all the objects directly from the subproject as if it was one big \n   project.\n\n   I actually think this sucks. Why? Because it puts an insane load on the \n   server side, which basically needs to traverse the object list of the \n   _sum_ of all projects. An initial clone (or a really big pull, which \n   comes to the same thing) would be absolutely horrendous\n\nSo I'd strongly argue against that approach, for scalability reasons. So \ninstead, you should really try to do pulls etc one git repo at a time:\n\n - take the \"list of subprojects\" from the supermodule, and pull them all \n   one by one.\n\n   This again makes subprojects \"less seamless\", and makes each subproject \n   more of a separate thing, with the project list gotten from the \n   superproject and parsed separately. But it means you have none of the \n   scalability problems, since you never see things as one huge project \n   with millions of files and even more objects.\n\nThe second approach also means that you can see the \"supermodule\" support \nin git as less of a \"plumbing\" thing, and it's largely just a thin veneer \naround the core plumbing that really doesn't understand about multiple \nrepositories at all (apart from the single \"link\" extension in the tree \nobject), and it's really just scripting to get the subprojects to \"look\" \nlike one thing, when they really are pretty much independent.\n\n"},{"id":"297988","messageId":"200612020114.42858.Josef.Weidendorfer@gmx.de","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612011510290.3695@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-02T00:14:42Z","receivedAt":"2006-12-02T00:14:42Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Saturday 02 December 2006 00:30, Linus Torvalds wrote:\n> On Fri, 1 Dec 2006, Josef Weidendorfer wrote:\n> > \n> > What about my other argument for a submodule namespace:\n> > You want to be able to move the relative root path of a submodule\n> > inside of your supermodule, but yet want to have a unique name\n> > for the submodule:\n> > - to be able to just clone a submodule without having to know\n> > the current position in HEAD\n> \n> Umm? I don't get the issue. A submodule is a git repo in its own right, \n> and you clone it exactly like you'd clone any other repo. It _does_ have a \n> HEAD. It has it's own branches. It has everything.\n\nI just thought about the case when you want to clone a submodule directly\nout of the supermodule repository, at a given realive path. And that can\nbe changing.\n\nOf course, every project which happens to be submodule of some supermodule,\nalso can have its own repository, as it is fully independent. And then,\nyou of course can clone from without any knowledge of its relative position\nin the supermodule.\n\n> In the trivial case where the submodule doesn't even _have_ any external \n> existence at all (ie it's always maintained as _just_ a submodule, it \n> would probably tend to have just one branch, and a clone would get \n> whatever that branch is), but that's just a degenerate special case of the \n> much richer \"this submodule actually has a life of its own\" case.\n\nYes.\n\n> > - more practically, e.g. to be able to name a submodule\n> > independent from any current commit you are on in the supermodule,\n> > e.g. to be able to store some meta information about a submodule:\n> \n> The current commit within the supermodule would be _totally_ invisible to \n> the submodule.\n\nOf course.\n\nYet, you need some name to store meta information of submodules\ninto some config file of the supermodule, like whether you want to have\nit checked out (see below).\n\nIn that case, such a name for a submodule does not have to be global in\nthe supermodule project...\n\n> > - \"Should git allow to commit rewind actions of this submodule\n> >    in the supermodule?\" (which, AFAICS, exactly has the same\n> >    problems as publishing a rewound branch: you will get into\n> >    merge hell when you want to pull upstream changes into the\n> >    supermodule)\n> \n> The only thing that a submodule must NOT be allowed to do on its own is \n> pruning (and it's distant cousin \"git repack -d\"). You must always prune \n> >from the supermodule, because the submodule cannot really know on its own \n> what references point into it.\n\nYes. I just gave an example of a policy some project may want for submodule\nhandling.\n\n> > - \"Should this submodule be checked out?\"\n> \n> This, I think, requires too much configuration to say separately for every \n> possible submodule, so I would suggest that the way to make that decision \n> is:\n> \n>  - \"git clone\" by default will fetch and check out all submodules (and \n>    obviously they have to be described some way outside of the object \n>    database, just so that you don't have to parse the _whole_ history of \n>    the _whole_ supermodule just to find all possible submodules. So the \n>    supermodule _will_ need some \"list of submodules and where to get them\" \n>    in a config file or other).\n\nExactly. And in this list, you have to specify names.\n\nThe thing I wanted to discuss is whether such names would need to be globally\nunique in the project containing submodles, or not.\n\nIf yes, it IMHO makes a lot of sense to introduce \"submodule objects\" which contain\nthese submodule names, and which are used as pointers to submodule commits in\nsupermodule trees.\n\n"},{"id":"297680","messageId":"Pine.LNX.4.64.0612011621380.3695@woody.osdl.org","threadId":"43065","inReplyTo":"200612020114.42858.Josef.Weidendorfer@gmx.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-02T00:33:52Z","receivedAt":"2006-12-02T00:33:52Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 2 Dec 2006, Josef Weidendorfer wrote:\n> > \n> > The current commit within the supermodule would be _totally_ invisible to \n> > the submodule.\n> \n> Of course.\n> \n> Yet, you need some name to store meta information of submodules\n> into some config file of the supermodule, like whether you want to have\n> it checked out (see below).\n\nYes, you do need to have a list of submodules somewhere, and you'd need to \nmaintain that separately. One of the results of having the submodules be \nindependent from the supermodule is that it's not all \"automatically \nintegrated\", and thus the supermodule does end up having to have things \nlike that maintained separately. \n\nAnd yes, if you screw that up, you wouldn't be able to fetch submodules \nproperly etc, even if you see the supermodule, and yes, this sounds more \nlike the CVS \"Entries\" kind of file that is more \"tacked on\" than really \ndeeply integrated. But I think the separation is _more_ than worth the \nfact that you can see things being separate.\n\nIn fact, I'm very much arguing for keeping things as separate as possible, \nwhile just integrating to the smallest possible degree (just _barely_ \nenough that you can do things like \"git clone\" and it will fetch multiple \nrepositories and put them all in the right places, and \"git diff\" and \nfriends will do reasonably sane things).\n\nKeep it simple, stupid. \n\n> >  - \"git clone\" by default will fetch and check out all submodules (and \n> >    obviously they have to be described some way outside of the object \n> >    database, just so that you don't have to parse the _whole_ history of \n> >    the _whole_ supermodule just to find all possible submodules. So the \n> >    supermodule _will_ need some \"list of submodules and where to get them\" \n> >    in a config file or other).\n> \n> Exactly. And in this list, you have to specify names.\n\nYes. \n\n> The thing I wanted to discuss is whether such names would need to be globally\n> unique in the project containing submodles, or not.\n\nMy preference would be for it to be \"local\", just because (as I \nmentioned), with mirroring etc, it might well be that you want to fetch \nthings from the _closest_ repository. That's really not a global decision, \nit's a local one.\n\n> If yes, it IMHO makes a lot of sense to introduce \"submodule objects\" which contain\n> these submodule names, and which are used as pointers to submodule commits in\n> supermodule trees.\n\nYou could do it that way, and then it would be global. It would work, and \nin many ways it would probably be \"simpler\" on a supermodule level.\n\nThe advantage of a global namespace is that you can much more easily \nupdate it - \"git fetch\" will just fetch the new file(s) that describe the \nsubprojects very naturally if they are all global. Putting them in a local \n.git/config file has it's advantages (see above), but it also makes it \nvery hard to version them, and to update the list - it would have to \nbecome manual.\n\nThere are possibly combinations of the two approaches: have a \"global \nnamespace\" that describes the canonical place to get the subprojects, but \nhave some way to add local \"translation\" of the canonical names into \nlocally preferred versions (eg you could just have a way to say \"this is \nthe local mirror for that global canonical place\")\n\nMaybe that would work?\n\n"},{"id":"296976","messageId":"200612020922.43832.andyparkins@gmail.com","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612011540010.3695@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2006-12-02T09:22:40Z","receivedAt":"2006-12-02T09:22:40Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Saturday 2006, December 02 00:12, Linus Torvalds wrote:\n\n> \t100644 blob 08602f522183dc43787616f37cba9b8af4e3dade\txdiff-interface.c\n> \t100644 blob 1346908bea31319aabeabdfd955e2ea9aab37456\txdiff-interface.h\n> \t040000 tree 959dd5d97e665998eb26c764d3a889ae7903d9c2\txdiff\n> \t050000 link 0215ffb08ce99e2bb59eca114a99499a4d06e704\txyzzy\n>\n> where that 050000 is the new magic type (I picked one out of my *ss: it's\n> not a valid type for a file mode, so it's a godo choice, but it could be\n> anythign that cannot conflict with a real file), which just specifies the\n> \"link\" part. The SHA1 is the SHA1 of the commit, and the \"xyzzy\" is\n> obviously just the name within the directory of the submodule.\n\nCan I argue that the hash in that object should actually be to a real object \nin the supermodule repository rather than a link?  Then THAT object would \ncontain the hash?  So in your above example:\n\n  100644 blob 08602f522183dc43787616f37cba9b8af4e3dade\txdiff-interface.c\n  100644 blob 1346908bea31319aabeabdfd955e2ea9aab37456\txdiff-interface.h\n  040000 tree 959dd5d97e665998eb26c764d3a889ae7903d9c2\txdiff\n  050000 link a7f26495b7b7e32bf949efbd91ee32267b792cba\txyzzy\n\nAnd then the local object a7f26495b7b7e32bf949efbd91ee32267b792cba would \ncontain your original hash 0215ffb08ce99e2bb59eca114a99499a4d06e704.\n\nThe reason I suggest this as without out it the \"link\" object is the only hash \nin the tree that doesn't point to a valid object.  The contents of objects is \nentirely arbitrary so it's perfectly okay for that to contain a hash that \nwon't dereference to a real object in the supermodule.\n\nThe main advantage of this is (I think) that git-prune, git-fsck, and whatever \nelse relies on tree objects all being real, don't need to be modified at all.\n\nIt also gives you scope to later add fields to the \"link\" object if you \nwanted.\n\n\nAndy\n-- \nDr Andrew Parkins, M Eng (Hons), AMIEE\n"},{"id":"296419","messageId":"200612020927.50787.andyparkins@gmail.com","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612011621380.3695@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2006-12-02T09:27:49Z","receivedAt":"2006-12-02T09:27:49Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Saturday 2006, December 02 00:33, Linus Torvalds wrote:\n\n> Yes, you do need to have a list of submodules somewhere, and you'd need to\n> maintain that separately. One of the results of having the submodules be\n\nWhy?  You just recursively search for every \"link\" object in the supermodule.  \nThat tells you which submodules you need and where they should be.\n\nDuring a supermodule clone, it can tell the client end to start a new clone \nwith the correct path because it knows what the local path is at that moment.\n\n\n\nAndy\n-- \nDr Andrew Parkins, M Eng (Hons), AMIEE\n"},{"id":"298434","messageId":"200612021004.22236.andyparkins@gmail.com","threadId":"43065","inReplyTo":"20061201220821.GL18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2006-12-02T10:04:20Z","receivedAt":"2006-12-02T10:04:20Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Friday 2006, December 01 22:08, Martin Waitz wrote:\n\n> > echo $SUBMODULE_HASH >\n> > submodule/.git/refs/supermodules/commit$SUPERMODULE_HASH\n>\n> I guess you are aware that you have to scan _all_ trees inside _all_\n> supermodule commits for possible references.\n\nNo you don't; you do it as part of the appropriate normal operations.\n\n * supermodule commit - scan the current tree for \"link\" objects in the\n   tree.  If you find one write the reference in the submodule.\n * adding a new submodule - if this is a new submodule there can't be any\n   references in the supermodule already.\n * cloning a supermodule, every new commit that gets written in the \n   supermodule gets checked from \"link\" objects.\n\n> So what do you do with deleted submodules?\n> You wouldn't want them to still sit around in your working directory,\n> but you still have to preserve them.\n\nNow that is a tricky one.  Mind you, I think that problem exists for any \nimplementation.  I haven't got a good answer for that.\n\n\nAndy\n\n-- \nDr Andrew Parkins, M Eng (Hons), AMIEE\n"},{"id":"295862","messageId":"200612021232.08699.Josef.Weidendorfer@gmx.de","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612011540010.3695@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-02T11:32:08Z","receivedAt":"2006-12-02T11:32:08Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Saturday 02 December 2006 01:12, Linus Torvalds wrote:\n> On Sat, 2 Dec 2006, Josef Weidendorfer wrote:\n> > \n> > So you are for a global submodule namespace in supermodule repositories,\n> > do I understand correctly?\n> > \n> > Otherwise, how would you specify the submodules at clone time given the\n> > ability that submodule roots can have relative path changed arbitrarily\n> > between commits?\n> \n> The only _true_ namespace would be the SHA1 of the commit (and maybe allow \n> a pointer to a tag too, but the namespace ends up being the same).\n\nI am not so sure about this.\nPerhaps we want the namespace to be more than the space of commit ids.\n\nSuppose you have some superproject which uses two compiler major versions\n(GCC 3 and GCC 4) as submodules because you want to have your regression\ntest suite run with both major versions.\nSo you would have a submodule at path \"gcc3/\" and \"gcc4/\" in your supermodule.\nAs both the gcc 3 and gcc 4 are branches from the same project, the submodule\nlinks will go into a connected DAG (suppose GCC uses git).\n\nAlone from the commit link, it is not easy to see to what submodule it belongs\nto (at least from a practical point of view).\n\nSo it actually _is_ more information if the proposed link objects in the supermodule\ncontain some submodule ID they belong to. They only need be unique in the scope of\nthe superproject (not really globally unique).\n\nSo another argument for submodule names: Merging. Otherwise, how do you decide\nto which submodule a link belongs to, especially in the scope of above example?\n\n> How to _find_ a repository that contains that SHA1 must be left to higher \n> levels. After all, repositories move around, and the place you found them \n> originally is not a stable name.\n\nI did not talk about a special format for these submodule IDs yet.\nWe could use an URL, but with such a value a user automatically associates\nsome semantic which can be confusing, as repository URLs can change. \n\nWe can use some symbolic name which has some meaning in the scope of\nthe superproject, and is specified at submodule creation, like \"gcc3\" or\n\"gcc4\". However, this is a local decision of the person which is importing\nthe submodule. So if two developers of the same project using supermodules\nindependently decide that the import of \"gcc\" as submodule is the right\nthing, but use slightly different submodule IDs, you will get 2 different\nsubmodules when merging.\n\nI argue that this is even the correct thing, and they should decide about the\nname before both are doing the import, or only one imports and the other\npulls.\n\nAnother option for a submodule ID could be the root commit of the submodule\ncommit DAG. This looks nice as such an ID really is globally unique for\nprojects (more or less: the first commit always contains the time stamp at\ncreation, and the author/commiter email address, even if the tree happens\nto be the same because you start with the same dummy file).\n\nBut my example above (with the 2 different submodules from the\nsame GCC project) shows that this is not working. A superproject never\ncould create different submodules from the same (e.g. GCC) project.\n\nSo I just vote for a symbolic name choosen at submodule creation time.\n\n"},{"id":"295292","messageId":"ekrsio$1qe$1@sea.gmane.org","threadId":"43065","inReplyTo":"20061201110032.GL18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-02T12:48:17Z","receivedAt":"2006-12-02T12:48:17Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"[cut]\n\nFrom this discussion I think it follows that supermodule should track HEAD\nversion of submodule. Perhaps the supermodule index should have sha1 of\nsubmodule commit, so (as usual) you have to update-index in supermodule to\nrecord changes in submodule; the difference being that you update to HEAD\nversion, not to working directory version. Or you can just git-commit -a\nin supermodule which would take working directory version of files, and HEAD\nversion of submodules.\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n\n"},{"id":"297737","messageId":"ekrtph$6nh$1@sea.gmane.org","threadId":"43065","inReplyTo":"45706758.2020907@b-i-t.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-02T13:08:58Z","receivedAt":"2006-12-02T13:08:58Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Stephan Feder wrote:\n\n> That's it: There is no need for a separate branch or repository. If you \n> have the subproject's commit in the superproject's object database (and \n> we really have that, see 1. and 2. above), why do you _have to_ store it \n> elsewhere?\n\nIt would be much simpler to have subproject's commit in subproject object\ndatabase, and have it available in superproject's object database by the\nway of alternates.\n\nOtherwise when commiting new submodule state in supermodule you would have\nto fetch all the needed objects (submodule mighe have evolved few commits\nin history inbetween) into superproject's object database.\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n\n"},{"id":"297263","messageId":"ekru4l$6nh$2@sea.gmane.org","threadId":"43065","inReplyTo":"200612012104.39897.andyparkins@gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-02T13:14:54Z","receivedAt":"2006-12-02T13:14:54Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Andy Parkins wrote:\n\n>> You still don't have a totally separated repository then, because\n>> you can't do a reachability analysis in the submodule repository alone.\n> \n> I'm going to guess by reachability analysis, you mean that the submodule \n> doesn't know that some of it's commits are referenced by the supermodule.  As \n> I suggested elsewhere in the thread, that's easily fixed by making a \n> refs/supermodule/commitXXXX file for each supermodule commit that references \n> as particular submodule commit.  Then you can git-prune, git-fsck whenever \n> you want.\n\nI think it would be better resolve this in universal way by adding\nto git repository layout the optional \"borrowers\" file, which would\nprotect against pruning objects that are referenced by repositories\nwhich have given repository as one of the \"alternates\".\n\nBy the way, how to slurp all the objects from alternates into repo\nobject repository?\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n\n"},{"id":"294513","messageId":"200612021450.46005.Josef.Weidendorfer@gmx.de","threadId":"43065","inReplyTo":"200612021004.22236.andyparkins@gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-02T13:50:45Z","receivedAt":"2006-12-02T13:50:45Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Saturday 02 December 2006 11:04, Andy Parkins wrote:\n> > So what do you do with deleted submodules?\n> > You wouldn't want them to still sit around in your working directory,\n> > but you still have to preserve them.\n> \n> Now that is a tricky one.  Mind you, I think that problem exists for any \n> implementation.  I haven't got a good answer for that.\n\nThat suggests that it is probably better to separate submodule repositories\nfrom their checked out working trees. Why not put the GITDIRs of the submodules\nin subdirectories of the supermodules GITDIR instead?\n\n"},{"id":"295789","messageId":"e7bda7770612021057mc9f3eb9q7fc047dd1b5c235f@mail.gmail.com","threadId":"43065","inReplyTo":"4570BFA4.8070903@stephan-feder.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Torgil Svensson","fromEmail":"torgil.svensson@gmail.com","sentAt":"2006-12-02T18:57:46Z","receivedAt":"2006-12-02T18:57:46Z","isPatch":false,"sender":{"key":"torgil.svensson@gmail.com","avatar":null},"body":"> If you need a common infrastructure to be able to work with the\n> submodule, then the submodule is not independent of of the supermodule.\n> I see a contradiction in your requirements.\n\nHere's an real-world example that doesn't contradict:\n\nhttp://amarok.kde.org/wiki/Installation_HowTo#From_Anonymous_SVN\n\n\"svn co -N svn://anonsvn.kde.org/home/kde/trunk/extragear/multimedia\ncd multimedia\nsvn co svn://anonsvn.kde.org/home/kde/branches/KDE/3.5/kde-common/admin\nsvn up amarok\n\nTo compile the sources (from the multimedia directory):\"\n\nand there's probably very few people that want to clone the entire KDE\nmultimedia sub&super-module in this case.\n\n//Torgil\n\n\nOn 12/2/06, sf <sf-gmane@stephan-feder.de> wrote:\n> Linus Torvalds wrote:\n> >\n> > On Fri, 1 Dec 2006, sf wrote:\n> >> Linus Torvalds wrote:\n> >> ...\n> >>> In contrast, a submodule that we don't fetch is an all-or-nothing\n> >>> situation: we simply don't have the data at all, and it's really a matter\n> >>> of simply not recursing into that submodule at all - much more like not\n> >>> checking out a particular part of the tree.\n> >> If you do not want to fetch all of the supermodule then do not fetch the\n> >> supermodule.\n> >\n> > So why do you want to limit it? There's absolutely no cost to saying \"I\n> > want to see all the common shared infrastructure, but I'm actually only\n> > interested in this one submodule that I work with\".\n>\n> If you need a common infrastructure to be able to work with the\n> submodule, then the submodule is not independent of of the supermodule.\n> I see a contradiction in your requirements.\n>\n> > Also, anybody who works on just the build infrastructure simply may not\n> > care about all the submodules. The submodules may add up to hundreds of\n> > gigs of stuff. Not everybody wants them. But you may still want to get the\n> > common build infrastructure.\n>\n> See above.\n>\n> > In other words, your \"all or nothing\" approach is\n> >  (a) not friendly\n> > and\n> >  (b) has no real advantages anyway, since modules have to be independent\n> >      enough that you _can_ split them off for other reasons anyway.\n> >\n> > So forcing that \"you have to take everything\" mentality onyl has\n> > negatives, and no positives. Why do it?\n>\n> (There have been lots of use cases for shallow clones but for a long\n> time git did not support them).\n>\n> If you can extend this partial fetch feature to the non-subproject case\n> I would agree with your reasoning. What makes the subprojects so special\n> in this regard. Do I have to turn a plain tree into a subproject to be\n> able to ignore it? Once you can restrict fetches to parts of the\n> contents you get the ability to restrict fetches to the \"common\n> infrastructure\" and selected submodules for free.\n>\n> Regards\n>\n> Stephan\n>\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"294103","messageId":"Pine.LNX.4.64.0612021114270.3476@woody.osdl.org","threadId":"43065","inReplyTo":"e7bda7770612021057mc9f3eb9q7fc047dd1b5c235f@mail.gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-02T19:41:38Z","receivedAt":"2006-12-02T19:41:38Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 2 Dec 2006, Torgil Svensson wrote:\n> \n> Here's an real-world example that doesn't contradict:\n\nAnd I'll add the note that people who do things like submodules aren't \ngenerally even _used_ to them being \"seamless\", and most of the time \nprobably don't even want complete seamlessness.\n\nAs the example that Torgil points to shows, people are quite used to \nactually even naming the submodules separately, and things like having the \n\"default\" set of submodules not equal the \"complete\" set. \n\nIn other words, I don't think people expect or want something hugely more \ncomplicated than the CVS/modules kind of file. \n\nWhat people _do_ want (and that CVS in general is horribly bad at, and \nthis is not a module-specific issue) is to have the _versioning_ work \nwell. When you check out a specific version of a module, you want any \n_linked_ modules to follow along too.\n\nThis is the same reason why CVS users use tags a lot: because even \n_within_ a single project (no modules, no nothing), it's often hard to \nre-create the exact state of a version any other way. So you tag every \nsingle file and do insane things like that, because CVS just isn't very \ngood at guaranteeing consistency across the whole project.\n\nThe exact same thing is true about subprojects. I don't think that people \nwho have used CVS subprojects a lot really mind the CVS/modules file \nitself (but hey, maybe I'm wrong - I've seen _other_ people maintain \nmodules in CVS, but I've never done it myself), but they do mind the fact \nthat it's hard as hell to do something as simple as \"get all modules back \nto version X\" without lots and lots of careful crud (ie tagging every \nsingl emodule, things like that).\n\nNow, I'm not exactly sure who wants to use git modules, so this is the \ntime to ask: did you hate the CVS/modules file? Or was it something you \nset up once, and then basically forgot about? People clearly use the \nability to mark certain modules as depending on each other, and aliases to \nsay \"if you ask for this module, you actually get a set of _these_ \nmodules\".\n\n_I_ suspect that that isn't the problem people had, and isn't what they \nhave any problems with. What CVS didn't do very well (or at all, afaik) is \nto say \"I want supermodule version XYZ\", and then got all the submodules \nautomatically to that (reliable) state. And THAT is something I think is \nreally important for submodules, and it's why I think the most important \npart isn't actually all the veneer to make \"git clone\" and \"git pull\" work \n(which is really about the CVS/modules kind of wrapper parsing), but \nactually about the supermodule \"tree\" object pointing to a very specific \nversion, so that you get the exact same \"atomic snapshotting\" of multiple \ntrees that you get within a single git tree.\n\nIn other words, I _suspect_ that that is really what module users are all \nabout. They want the ability to specify an arbitrary collection of these \natomic snapshots (for releases etc), and just want a way to copy and move \nthose things around, and are less interested in making everything else \nvery seamless (because most people are happy to do the actual \n_development_ entirely within the submodules, so the \"development\" part \nis actually not that important for the supermodule, the supermodule is \nmostly for aggregation and snapshots, and tying different versions of \ndifferent submodules together).\n\nSo that's where I come from. And maybe I'm totally wrong. I'd like to hear \nwhat people who actually _use_ submodules think.\n\n"},{"id":"295350","messageId":"20061202194602.GP18810@admingilde.org","threadId":"43065","inReplyTo":"4570BC07.4080203@stephan-feder.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-02T19:46:03Z","receivedAt":"2006-12-02T19:46:03Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Sat, Dec 02, 2006 at 12:34:31AM +0100, sf wrote:\n> > Now your submodule is no longer seen as an independent git repository\n> > and I think this would cause problems when you want to push/pull between\n> > the submodule and its upstream repository.\n> \n> You can always pick a single commit or several commits out of a larger\n> repository and have a complete git repository.\n> \n> And I already explained how to push and pull even from within superprojects.\n\nSure it you are able to make it work, but it needs more work on the UI part.\nHow do you handle the index? How do you allow to clone only the\nsubmodule?\n\nI really thought about such a setup too, but then decided that it is\nmuch easier to work with submodules when you can really see it as a\nrepository of its own.\n\n> > But you could still call the \"xdiff\" part of the git repository a\n> > submodule.  And then changes to the xdiff directory result in a new\n> > submodule commit, even when there is no direct reference to it.\n> > So you'd still \"commit to the xdiff submodule\".\n> \n> Let's make certain that we understand each other. I see a clear\n> distinction between the submodule code in a supermodule branch (commits\n> in the supermodule's tree and nothing else) and submodule branches which\n> are independent of the superproject. Supermodule branches and submodule\n> branches do not interact, only if I want them to.\n\nAgreed.\nI think the thing which caused some discussion is that I make the\ncurrent submodule commit which is used by the supermodule available in a\nrefs/head in the submodule.\nSo there is one \"branch\" in the submodule which corresponds to the\nversion used by the supermodule, but this is just for user interface.\nIt's most important purpose is to give this special commit a name, so\nthat it can be used in merges, etc.\n\nBy selecting another refs/heads \"branch\" in the submodule you can also\neasily detach the submodule from the supermodule.\nIt is really important to understand that you can't branch the submodule\nalone and still have it connected to the supermodule, because the\nsupermodule always tracks only one commit for each submodule.\nSo every branch that affects the project has to be done on project\n(topmost supermodule) level.\nBut of course the submodule can have other branches which are not\ntracked by the supermodule.\nSo by checking out refs/heads/master (as it is used in my\nimplementation) you can attach the submodule to the supermodule (attach\nas in: bring the working directory in sync with the whole project), and\nyou can detach it by selecting another refs/heads (the submodule is\nstill part of the supermodule, but not in the state which is currently\nvisible in the working directory).\nThis may sound confusing, but it really is the only semantic for\nsubmodule branches that makes sense.\nThere are fears that you may commit something that does not match your\ncurrent working directory.  Sure, but you explicitly asked for it and I\nthink it won't be a problem if git-status tells about this fact.\n\n\n> The double slashes is the only way I can think of that clearly indicates\n> that I do not mean the contents named by the path, but the commit that\n> you find there. Once you have named a commit in that way, you can\n> continue to apply other revision naming suffixes, paths, and so on.\n\nWith the current semantics, you can already get to the submodule commit\n(just leave out your double slashes), but what is missing is simply to\napply all the modifiers again on this submodule commit.\nSo I think we can do without the double slashes.\n\n-- \nMartin Waitz\n"},{"id":"298383","messageId":"Pine.LNX.4.64.0612021144520.3476@woody.osdl.org","threadId":"43065","inReplyTo":"200612021232.08699.Josef.Weidendorfer@gmx.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-02T19:52:13Z","receivedAt":"2006-12-02T19:52:13Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 2 Dec 2006, Josef Weidendorfer wrote:\n> > \n> > The only _true_ namespace would be the SHA1 of the commit (and maybe allow \n> > a pointer to a tag too, but the namespace ends up being the same).\n> \n> I am not so sure about this.\n> Perhaps we want the namespace to be more than the space of commit ids.\n\nI don't think it would be wrong at all to have a \"link object\" type, and \nhave the \"link\" tree entry actually point to that \"link object\" instead of \npointing directly to the commit in the submodule.\n\nAnd yes, that extra indirection would allow for more flexibility (the \n\"link object\" can contain comments about the particular version used, \npointers to where you can get it - whether human-readable or strictly \nmeant for automation - etc etc).\n\nSo I agree with Andy Parkins' comment about the link object allowing not \nonly extended namespaces, but also allowing a certain amount of \nflexibility (ie there's some built-in extensibility and ability to perhaps \nadd future fields if there's a new object type).\n\nI just want the naming of the links themselves to use all the same SHA1 \nhashes etc, so that you always have a very explicit, and very trustworthy \nversion - and never end up in the situation that you know which repository \nyou want at that position, but you don't know exactly which commit in that \nrepo was supposed to be checked out with that particular version of the \nsuper-module.\n\n"},{"id":"293802","messageId":"20061202201242.GQ18810@admingilde.org","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612011505190.3695@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-02T20:12:42Z","receivedAt":"2006-12-02T20:12:42Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Fri, Dec 01, 2006 at 03:09:40PM -0800, Linus Torvalds wrote:\n> On Fri, 1 Dec 2006, sf wrote:\n> > If you do not want to fetch all of the supermodule then do not fetch the\n> > supermodule.\n> \n> So why do you want to limit it? There's absolutely no cost to saying \"I \n> want to see all the common shared infrastructure, but I'm actually only \n> interested in this one submodule that I work with\".\n\nAn interesting way to support this \"only fetch some modules\" use-case is\nto use several supermodules.\n\nSo you could have one supermodule which is geared towards developers and\nonly contains the modules they use.  Another supermodule contails all\nthe toolchain sources.  And then there is the supermodule used for\nreleases which is just a merge of all the other supermodules.\n\nThe concept is so flexible that you don't have to introduce lots of\nother things as module namespaces.  Just use the tools you have in a\ncreative way ;-)\n\n-- \nMartin Waitz\n"},{"id":"295479","messageId":"20061202201826.GR18810@admingilde.org","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612011540010.3695@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-02T20:18:26Z","receivedAt":"2006-12-02T20:18:26Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Fri, Dec 01, 2006 at 04:12:10PM -0800, Linus Torvalds wrote:\n> It only gets interesting for commands that fetch new objects, ie do a \n> \"pull/fetch\" op, and you'd need to know where/how to fetch new objects for \n> the xyzzy subproject, so that's a \"naming\" issue. You have a few choices:\n> \n>  - get all the objects directly from the subproject as if it was one big \n>    project.\n> \n>    I actually think this sucks. Why? Because it puts an insane load on the \n>    server side, which basically needs to traverse the object list of the \n>    _sum_ of all projects. An initial clone (or a really big pull, which \n>    comes to the same thing) would be absolutely horrendous\n\nI don't buy your scalability argument.\nBy dividing the object traversal in separate steps you do not win\nanything.  The complexity of the operation still stays the same, as you\nstill have to traverse the exact same amount of objects.\n\nBy separating the repositories you just make reachability analyis be\ntotally awkward, without winning anything.\n\n-- \nMartin Waitz\n"},{"id":"296406","messageId":"20061202202103.GS18810@admingilde.org","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612021144520.3476@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-02T20:21:03Z","receivedAt":"2006-12-02T20:21:03Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Sat, Dec 02, 2006 at 11:52:13AM -0800, Linus Torvalds wrote:\n> I don't think it would be wrong at all to have a \"link object\" type, and \n> have the \"link\" tree entry actually point to that \"link object\" instead of \n> pointing directly to the commit in the submodule.\n> \n> And yes, that extra indirection would allow for more flexibility (the \n> \"link object\" can contain comments about the particular version used, \n> pointers to where you can get it - whether human-readable or strictly \n> meant for automation - etc etc).\n\nWhat makes a submodule so special that now we suddenly have to store\nthose stuff in the object database?\n\nStoring a fetch location would grossly contradict the distributed nature\nof git.  I really do not see _any_ reason to store more information than\nthe commit sha1 of the submodule.\n\n-- \nMartin Waitz\n"},{"id":"295721","messageId":"20061202202428.GT18810@admingilde.org","threadId":"43065","inReplyTo":"200612020017.44275.Josef.Weidendorfer@gmx.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-02T20:24:28Z","receivedAt":"2006-12-02T20:24:28Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Sat, Dec 02, 2006 at 12:17:44AM +0100, Josef Weidendorfer wrote:\n> After some thinking, a submodule namespace even is important for checking\n> out only parts of a supermodule, exactly because the root of a submodule\n> potentially can change at every commit.\n\nhave you ever thought about the idea that the location may be an\nimportant thing to consider for your decision.\n\nPerhaps the submodule is now used for something else (this is why it was\nmoved) and that now you'd like to keep it?\n\nAnyway, you can just create several supermodules or implement generic\npartial tree support for git.  I do not see any reason to special case\nsubmodules here.\n\n-- \nMartin Waitz\n"},{"id":"298549","messageId":"20061202204012.GU18810@admingilde.org","threadId":"43065","inReplyTo":"200612021004.22236.andyparkins@gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-02T20:40:12Z","receivedAt":"2006-12-02T20:40:12Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Sat, Dec 02, 2006 at 10:04:20AM +0000, Andy Parkins wrote:\n> On Friday 2006, December 01 22:08, Martin Waitz wrote:\n> \n> > > echo $SUBMODULE_HASH >\n> > > submodule/.git/refs/supermodules/commit$SUPERMODULE_HASH\n> >\n> > I guess you are aware that you have to scan _all_ trees inside _all_\n> > supermodule commits for possible references.\n> \n> No you don't; you do it as part of the appropriate normal operations.\n> \n>  * supermodule commit - scan the current tree for \"link\" objects in the\n>    tree.  If you find one write the reference in the submodule.\n>  * adding a new submodule - if this is a new submodule there can't be any\n>    references in the supermodule already.\n>  * cloning a supermodule, every new commit that gets written in the \n>    supermodule gets checked from \"link\" objects.\n\n * removing a branch from the supermodule.\n   OK, this is an infrequent operation and it can be handled by redoing\n   everything.\n\nI just don't like to duplicate information which is already available\neasily.  We'd need much to many special cases, just to correctly support\nreachablility analysis.\nKISS.\n\n> > So what do you do with deleted submodules?\n> > You wouldn't want them to still sit around in your working directory,\n> > but you still have to preserve them.\n> \n> Now that is a tricky one.  Mind you, I think that problem exists for any \n> implementation.  I haven't got a good answer for that.\n\nIf you just keep it in a shared object repository you don't have any\nproblems.\n\nPlease note that it is not required to keep it in one physical location.\nYou can still use alternates/whatever to store some objects in another\nrepository, but you need to be able to access all objects from the\nsupermodule.\n\n-- \nMartin Waitz\n"},{"id":"297888","messageId":"20061202204350.GV18810@admingilde.org","threadId":"43065","inReplyTo":"200612021450.46005.Josef.Weidendorfer@gmx.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-02T20:43:50Z","receivedAt":"2006-12-02T20:43:50Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Sat, Dec 02, 2006 at 02:50:45PM +0100, Josef Weidendorfer wrote:\n> On Saturday 02 December 2006 11:04, Andy Parkins wrote:\n> > > So what do you do with deleted submodules?\n> > > You wouldn't want them to still sit around in your working directory,\n> > > but you still have to preserve them.\n> > \n> > Now that is a tricky one.  Mind you, I think that problem exists for any \n> > implementation.  I haven't got a good answer for that.\n> \n> That suggests that it is probably better to separate submodule repositories\n> from their checked out working trees. Why not put the GITDIRs of the submodules\n> in subdirectories of the supermodules GITDIR instead?\n\nWhy not simply use a shared object database instead?\n\nYou can still have an alternative to some standalone bare repository of\nthe submodule if you do not like to store submodule objects in the\nsupermodule repository.\n\n-- \nMartin Waitz\n"},{"id":"298561","messageId":"Pine.LNX.4.64.0612021242080.3476@woody.osdl.org","threadId":"43065","inReplyTo":"20061202201826.GR18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-02T20:44:20Z","receivedAt":"2006-12-02T20:44:20Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 2 Dec 2006, Martin Waitz wrote:\n> \n> I don't buy your scalability argument.\n\nTry it.\n\nReally. Get the mozilla import (450MB project), and clone it on a machine \nwith half a gig of RAM or less.\n\nThen, clone a couple of smaller archives that end up being 450MB \n_combined_, but clone them separately.\n\nAnd watch the memory usage.\n\n> By separating the repositories you just make reachability analyis be\n> totally awkward, without winning anything.\n\nTrust me. Try it out.\n\n"},{"id":"294653","messageId":"Pine.LNX.4.64.0612021245081.3476@woody.osdl.org","threadId":"43065","inReplyTo":"20061202202103.GS18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-02T20:46:01Z","receivedAt":"2006-12-02T20:46:01Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 2 Dec 2006, Martin Waitz wrote:\n> \n> What makes a submodule so special that now we suddenly have to store\n> those stuff in the object database?\n\nI'm not sure it is. I suspect a pure commit link with just a CVS-style \n\"modules\" file is sufficient. I'm just saying that I don't think it is \n_wrong_ to possibly want to expand it.\n\n"},{"id":"294579","messageId":"20061202205853.GW18810@admingilde.org","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612021245081.3476@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-02T20:58:53Z","receivedAt":"2006-12-02T20:58:53Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Sat, Dec 02, 2006 at 12:46:01PM -0800, Linus Torvalds wrote:\n> On Sat, 2 Dec 2006, Martin Waitz wrote:\n> > \n> > What makes a submodule so special that now we suddenly have to store\n> > those stuff in the object database?\n> \n> I'm not sure it is. I suspect a pure commit link with just a CVS-style \n> \"modules\" file is sufficient. I'm just saying that I don't think it is \n> _wrong_ to possibly want to expand it.\n\nIf we later see that we really want to have it we can always introduce\nit later.  I don't think we should do it now if we don't see clear\nbenefits _now_.\n\nSo I was not against the link object itself (initially I wanted to do it\nthis way, too), only agains the information which was proposed to be\nstored there.  Up to now I haven't found anything which makes sense to\nstore next to the submodule commit to define the identity of the\nsubmodule.\n\n-- \nMartin Waitz\n"},{"id":"296600","messageId":"20061202210640.GX18810@admingilde.org","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612021242080.3476@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-02T21:06:40Z","receivedAt":"2006-12-02T21:06:40Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Sat, Dec 02, 2006 at 12:44:20PM -0800, Linus Torvalds wrote:\n> On Sat, 2 Dec 2006, Martin Waitz wrote:\n> > \n> > I don't buy your scalability argument.\n> \n> Try it.\n> \n> Really. Get the mozilla import (450MB project), and clone it on a machine \n> with half a gig of RAM or less.\n> \n> Then, clone a couple of smaller archives that end up being 450MB \n> _combined_, but clone them separately.\n> \n> And watch the memory usage.\n\nDo I understand you correctly that the problem is not the algorithmic\ncomplexity but that you have to map the objects at once instead of map\nthem in small parts one after the other?\n\n-- \nMartin Waitz\n"},{"id":"297570","messageId":"Pine.LNX.4.64.0612021252380.3476@woody.osdl.org","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612021242080.3476@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-02T21:22:22Z","receivedAt":"2006-12-02T21:22:22Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 2 Dec 2006, Linus Torvalds wrote:\n> \n> And watch the memory usage.\n\nBtw, just in case you don't understand _why_ this is true, the fact is, in \na git repository, quite fundamentally, because we don't have \"backlinks\" \nat any stage at all, we don't know - and fundamentally _cannot_ know - \nwhether we're goign to see the same object in the future.\n\nSo operations like \"git-rev-list --objects\" (or, these days, more commonly \nanything that just does the equivalent of that internally using the \nlibrary interfaces - ie \"git pack-objects\" and friends) VERY FUNDAMENTALLY \nhave to hold on to the object flags for the whole lifetime of the whole \noperation.\n\nAnd you should realize that this is really really fundamental. You can't \nfix it with \"smarter memory management\". You can't fix it with \"garbage \ncollection\". This is _not_ a result of the fact that we use C and malloc, \nand we don't free those objects, like some people sometimes seem to \nbelieve.\n\nSo garbage collection will never help this kind of situation. It flows \n_directly_ from the fact that our objects are immutable: because they are \nimmutable, they don't have any backpointers, because we cannot (and must \nnot) add backpointers to an old existing object when a new object is \ncreated that points to it.\n\nSo this really isn't a memory management issue. You could somewhat work \naround it by adding a \"caching layer\" on top of git, and allow that \ncaching layer to modify their cache of old objects (so that they can \ncontain back-pointers), but for 99% of all users that would actually make \nperformance MUCH WORSE, and it would also be a serious problem for \ncoherency issues (one of the things that immutable objects cause is that \nthere are basically never any race conditions, while a \"caching layer\" \nlike this would have some serious issues about serialization).\n\nSo: the very fundamental nature and choices that were made in git also \nmeans that when you have something like git-pack-objects that wants to \nwalk the whole repo, you will end up with something that remembers EVERY \nSINGLE OBJECT it walked. \n\nAnd while I've worked very hard to make the memory footprint of individual \nobjects as small as possible, and this means that this all works fine even \nfor fairly large databases (especially since very few operations actually \ndo this \"traverse the whole friggin tree\" thing), it does mean that \nthere's a very fundamental limit to scalability. You can't just make a \nwhole repository a hundred times bigger - because the operations that \ntraverse the whole thing will require a hundred times more memory!\n\nNow, in \"real\" projects, this is not a problem. I can pretty much \n_guarantee_ that memory sizes and hardware will grow faster than projects \ngrow. I'm not AT ALL worried about the fact that in ten years, the linux \nkernel repository will likely be two or three times the size it is now. \nBecause I'm absolutely convinced that in ten years, the machines we have \nnow will be obsolete.\n\nSo on any \"individual project\" basis, the fact that memory requirements \nscale roughly as O(n) in the total repository size is simply not a \nproblem. In fact, O(n) is pretty damn good, especially since the constant \nis pretty small (basically 28 bytes per object - and 20 of those bytes \nare the SHA1 that you simply cannot avoid).\n\nBut it does mean that supermodules really should NOT be so seamless that \ndoing a \"git clone\" on a supermodule does one _large_ clone. Because it's \nsimply going to be better to:\n\n - when you clone the supermodule, track the commits you need on all \n   submodules (this _may_ be a reason in itself for the \"link\" object, \n   just so that you can traverse the supermodule object dependencies and \n   know what subobject you are looking at even _without_ having to look at \n   the path you got there from)\n\n - clone submodules one-by-one, using the list of objects you gathered.\n\nMaybe there are other solutions, but quite frankly, I doubt it. Yes, \nyou'll end up \"traversing\" exactly as many objects either way, but the \n\"globe subobjects one by one\" is going to be a _hell_ of a lot more \nmemory-efficient, and quite frankly, \"memory usage\" == \"performance\"  \nunder many loads (notably, any load that uses too much memory will _suck_ \nperformance-wise, either because of swapping or simply because it will \nthrow out caches that \"many small invocations\" would not have thrown out).\n\nSo I guarantee that it's going to be better to do five clones of five \nsmall repositories over one clone of one big one. If only because you need \nless memory to do the five smaller clones.\n\n"},{"id":"296781","messageId":"Pine.LNX.4.64.0612021322560.3476@woody.osdl.org","threadId":"43065","inReplyTo":"20061202210640.GX18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-02T21:29:48Z","receivedAt":"2006-12-02T21:29:48Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 2 Dec 2006, Martin Waitz wrote:\n> \n> Do I understand you correctly that the problem is not the algorithmic\n> complexity but that you have to map the objects at once instead of map\n> them in small parts one after the other?\n\nNot map them, but track their \"used\" flag. Yes. You can unmap objects any \ntime at all (since you can just always re-create them at any time very \neasily and cheaply), but the one thing you CANNOT recreate is the object \nflags. See \"struct object\", and the \"used\" and FLAG_BITS in particular.\n\nAlmost all git programs need the FLAG_BITS. Something as simple as just \ntraversing the commit history needs at a minimum one _single_ bit for each \nobject: \"Have I already seen this\". In reality, you tend to need two or \nthree more (ie the UNINTERESTING bit ends up being as important as the \nSEEN bit, because it's what determines whether it's reachable from some \ncommit we're _not_ interested in, and in the end that's what allows us to \nnot traverse the whole history).\n\nSo you need at a MINIMUM to track the bits\n\n\t#define SEEN            (1u<<0)\n\t#define UNINTERESTING   (1u<<1)\n\nand in practice almost everything needs\n\n\t#define SHOWN           (1u<<3)\n\ntoo (SEEN is for deciding whether to _traverse_ something, SHOWN is for \ndeciding whether you've already output the data for this, and the \ndifference is crucial for any depth-first DAG algorithm, since you need \nto test-and-set the one bit when you first encounter the object, and \ntest-and-set the other bit when you \"leave\" the object).\n\nSo three bits are minimal to _any_ git traversal algorithm. Many specific \nissues want more bits (eg the TREECHANGE bit may not be quite as \nfundamnetal, but it sure ends up being critical for the \"track subtree\" \ncase).\n\n"},{"id":"296026","messageId":"200612030155.09630.Josef.Weidendorfer@gmx.de","threadId":"43065","inReplyTo":"20061202202428.GT18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-03T00:55:08Z","receivedAt":"2006-12-03T00:55:08Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Saturday 02 December 2006 21:24, Martin Waitz wrote:\n> On Sat, Dec 02, 2006 at 12:17:44AM +0100, Josef Weidendorfer wrote:\n> > After some thinking, a submodule namespace even is important for checking\n> > out only parts of a supermodule, exactly because the root of a submodule\n> > potentially can change at every commit.\n> \n> have you ever thought about the idea that the location may be an\n> important thing to consider for your decision.\n\nWhich decision, for what? Sorry, I do not understand.\n\nDo you want to say that relative submodule root paths should be kept fix\nthe whole lifetime of a supermodule?\nIe. a submodule \"identity\" is bound to its relative path, and when we\nmove it, it should be seen as deleting at and creating a totally new,\ndifferent submodule?\n\nThat's fine.\nBut you have to handle submodule creation/deletion neverless. And while\nyou are at a commit which has a given submodule deleted, you have to\nkeep the submodule data somewhere - referencing it with a name.\nI do not speak here about the object database, that could be combined;\nbut about all the other files in .git/ of the currently not checked out\nsubmodule.\n\n> Perhaps the submodule is now used for something else (this is why it was\n> moved) and that now you'd like to keep it?\n\nCan you give a usage szenario? What do you mean here?\n\n\n> Anyway, you can just create several supermodules or implement generic\n> partial tree support for git.  I do not see any reason to special case\n> submodules here.\n\nWhat should such a general partial tree support look like? I suppose you\nwant to configure paths which should not be checked out. As long as you\nsay that a given submodule always has to exist at a given path, you are\nright: then, you can say: \"Please, do not check out this submodule\" which\nis the same as \"Do not check out this path\". \n\nBut I think it is quite restrictive to not allow to move submodules around.\nWhen the supermodule upstream decides to move a submodule, your partial\ntree config to not check out a submodule will be lost.\nBut more important, if you made changes to a given submodule, and pull from\nupstream which changed the submodule position in-between, your changes will\nbe not taken over to the new position, as the move is seen as creation of\na totally independent submodule.\n\n"},{"id":"295543","messageId":"200612030202.34990.Josef.Weidendorfer@gmx.de","threadId":"43065","inReplyTo":"20061202204350.GV18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-03T01:02:34Z","receivedAt":"2006-12-03T01:02:34Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Saturday 02 December 2006 21:43, Martin Waitz wrote:\n> On Sat, Dec 02, 2006 at 02:50:45PM +0100, Josef Weidendorfer wrote:\n> > On Saturday 02 December 2006 11:04, Andy Parkins wrote:\n> > > > So what do you do with deleted submodules?\n> > > > You wouldn't want them to still sit around in your working directory,\n> > > > but you still have to preserve them.\n> > > \n> > > Now that is a tricky one.  Mind you, I think that problem exists for any \n> > > implementation.  I haven't got a good answer for that.\n> > \n> > That suggests that it is probably better to separate submodule repositories\n> > from their checked out working trees. Why not put the GITDIRs of the submodules\n> > in subdirectories of the supermodules GITDIR instead?\n> \n> Why not simply use a shared object database instead?\n\nSure. I have no problem with this.\n\nBut can we go one step further?\nAFAICS your submodules store the .git/ directories of submodules directly\nat submodule position in the working tree - but you have a link .git/objects\ninto the object database of the supermodule.\nWhen the user wants to delete the submodule, he would remove this .git/ directory,\ntoo. So you loose the .git/refs of the submodule etc. I would suggest to put\nthe submodule .git dirs into the .git dir of the supermodule.\n\n"},{"id":"294833","messageId":"200612030211.17159.Josef.Weidendorfer@gmx.de","threadId":"43065","inReplyTo":"20061202205853.GW18810@admingilde.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-03T01:11:16Z","receivedAt":"2006-12-03T01:11:16Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Saturday 02 December 2006 21:58, Martin Waitz wrote:\n> So I was not against the link object itself (initially I wanted to do it\n> this way, too), only agains the information which was proposed to be\n> stored there.  Up to now I haven't found anything which makes sense to\n> store next to the submodule commit to define the identity of the\n> submodule.\n\nIsn't it enough reason that a porcelain probably wants to store meta\ninformation for a given submodule, giving the need to put a name/identity\nto it?\n\nJosef\n"},{"id":"296969","messageId":"200612030307.26429.Josef.Weidendorfer@gmx.de","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612021252380.3476@woody.osdl.org","subject":"Thoughts about memory requirements in traversals [Was: Re: [RFC] Submodules in GIT]","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-03T02:07:26Z","receivedAt":"2006-12-03T02:07:26Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Saturday 02 December 2006 22:22, Linus Torvalds wrote:\n> So operations like \"git-rev-list --objects\" (or, these days, more commonly \n> anything that just does the equivalent of that internally using the \n> library interfaces - ie \"git pack-objects\" and friends) VERY FUNDAMENTALLY \n> have to hold on to the object flags for the whole lifetime of the whole \n> operation.\n>\n> [...]\n> \n> So this really isn't a memory management issue. You could somewhat work \n> around it by adding a \"caching layer\" on top of git, and allow that \n> caching layer to modify their cache of old objects (so that they can \n> contain back-pointers), but for 99% of all users that would actually make \n> performance MUCH WORSE, and it would also be a serious problem for \n> coherency issues (one of the things that immutable objects cause is that \n> there are basically never any race conditions, while a \"caching layer\" \n> like this would have some serious issues about serialization).\n\nThinking about this...\nYou have to make very sure to always update the caching layer containing\nthe backlinks on every addition of a further object. You can do this\nbecause you always reached this new object by some other object, which\nexactly is the backpointer.\n\nNow let us suppose we are able to do this.\nWhat does this give us?\n\nTake a look at object traversal:\nWe have to store the flag \"already visited\" for objects we could reach\nagain in the traversal. But with the backlinks, we can see that most\nof the objects can only be reached via one path, and therefore, there\nis no need to store the flag, as it never will be queried in the\nfurther traversal.\n(Similar for objects with two paths: When you have visited the object\ntwo times, you can throw away the flag, as it is not queried any more).\n\nRegarding the caching layer and object traversal, it would have been\nenough to only store \"is this object reachable via more than 1 path?\".\nFor this, the \"cache\" could be the set of objects reachable with\nmore than one path.\nAnd such a set stored in a file should be quite managable, and be\nquite small, relative to the size of the object database.\n\nIn fact, this \"cache\" can be created with a usual object traversal\n(which has the original memory requirement), but as long as we do\nnot add objects to the database, further traversals would only need\na fraction of memory.\n\nWhen only adding a small number of objects, it should be easy to\nupdate the cache; while with big actions like fetching/pulling,\nwe simply should remove the file with the backlink information.\n\n> problem. In fact, O(n) is pretty damn good, especially since the constant \n> is pretty small (basically 28 bytes per object - and 20 of those bytes \n> are the SHA1 that you simply cannot avoid).\n\nAgain only some thoughts...\nPack files are fully self-contained object stores, yes?\nSo in the scope of a single pack file, the offset of this object is enough\nas object identification.\nIf we could make sure that in any given algorithm touching objects, like\ncommit traversal, we always have the offset available when we need to do\nan object lookup, then, it should be enough to store object flags only\nindexed by the offset of this object in the pack.\nThe translation SHA1 -> offset can be done with the pack index.\nAs you usually have multiple packs, a (pack number / offset) tuple should\nbe enough as object ID.\n\nThinking even one step further:\nWould it make sense to define an encoding format for the content of\ncommit and tree objects inside of packs, where the SHA1 is replaced by the\noffset of the object in this pack?\nAs exactly the SHA1 is the least compressable thing, this could promise\nquite a benefit.\nAFAIK, we currently only use these offsets for referencing objects in\ndelta chains.\n\nMore about the original topic of this thread (and off-topic to the\nnew subject):\n\n> But it does mean that supermodules really should NOT be so seamless that \n> doing a \"git clone\" on a supermodule does one _large_ clone. Because it's \n> simply going to be better to:\n> \n>  - when you clone the supermodule, track the commits you need on all \n>    submodules (this _may_ be a reason in itself for the \"link\" object, \n>    just so that you can traverse the supermodule object dependencies and \n>    know what subobject you are looking at even _without_ having to look at \n>    the path you got there from)\n> \n>  - clone submodules one-by-one, using the list of objects you gathered.\n\nWithout submodule identities, we would have to clone path-by-path, as\nwe can not distinguish different submodules apart from there location.\n\n"},{"id":"294269","messageId":"Pine.LNX.4.64.0612021814530.3476@woody.osdl.org","threadId":"43065","inReplyTo":"200612030307.26429.Josef.Weidendorfer@gmx.de","subject":"Re: Thoughts about memory requirements in traversals [Was: Re: [RFC] Submodules in GIT]","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-03T02:25:29Z","receivedAt":"2006-12-03T02:25:29Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 3 Dec 2006, Josef Weidendorfer wrote:\n> \n> Thinking about this...\n> You have to make very sure to always update the caching layer containing\n> the backlinks on every addition of a further object. You can do this\n> because you always reached this new object by some other object, which\n> exactly is the backpointer.\n\nYou're missing the big issue.\n\nThe issue is that a cache like that would ABSOLUTELY SUCK.\n\nYou could speed up the non-common operations with it, but:\n\n - any changes would become a LOT more expensive to do, because they all \n   need to update every single object they add (ie a \"commit\" would now \n   have to add backpointers TO EVERY SINGLE BLOB).\n\n   Imagine what this does to something like the kernel, where a commit \n   reaches 22,000 files!\n\n   You can do it at a finer granularity (ie do just the direct backlinks \n   and only do the \"tree->blob\" and \"tree->tree\" things rather than the \n   full commit reachability, but it's still going to be MUCH more painful \n   than what we do now.\n\n - the cache would be a lot bigger than the current pack-files, and it \n   would be fragile as hell to boot. Because it needs to get rewritten for \n   every operation, it gets corrupted much more easily, and that's \n   ignoring things like race conditions, so it would now need a ton of \n   locking that git simply doesn't do at all.\n\n - everything would basically slow down.\n\n - you couldn't do shared object databases AT ALL, because backpointers \n   wouldn't work. The whole _reason_ you can share object databases is the \n   same reason we can't have backpointers: objects are immutable and never \n   change depending on circustances.\n\nThe _only_ downside of the current situation is literally the 24 or 28 \nbytes per object that we look at. For most operations, we don't even look \nat that many objects, so it's really the worst-case things.\n\n> In fact, this \"cache\" can be created with a usual object traversal\n> (which has the original memory requirement), but as long as we do\n> not add objects to the database, further traversals would only need\n> a fraction of memory.\n\nRight. If the project is totally read-only, the cache would work well.\n\nFor real development, it would SUCK. It would make things like \"git reset\" \nvery expensive indeed, for example (you'd have to unwind the whole cache: \neither regenerating it - which would take minutes - or being very careful \nindeed and being able to always remove objects properly and keeping track \nof them 100%).\n\nIOW, it's nasty nasty nasty. And it doesn't really even help anything but \na case that we actually already handle really well (I spent a lot of \neffort on making the memory footprint minimal).\n\nBut it does mean that you do NOT want to traverse a hundred different \nproject \"as if\" they were one. That's really the only thing it means.\n\nAnd since you can do submodules as independent projects, and you SHOULD do \nthem that way for tons of other reasons _anyway_, even that isn't a reason \nto screw up all the _wonderful_ properties of the git object database.\n\nSo what I'm trying to say is that the immutable non-backpointer nature of \nthe git database is what makes it so WONDERFUL. It's efficient, it's \ndense, it's stable, and it allows us all the clever things we do. But it \nmeans that we do end up alway spending 28 bytes per object, and we can \nnever throw those 28 bytes away during a single \"traversal\" run.\n\n"},{"id":"296344","messageId":"20061203024655.GD26668@spearce.org","threadId":"43065","inReplyTo":"200612030307.26429.Josef.Weidendorfer@gmx.de","subject":"Re: Thoughts about memory requirements in traversals [Was: Re: [RFC] Submodules in GIT]","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-12-03T02:46:55Z","receivedAt":"2006-12-03T02:46:55Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Josef Weidendorfer <Josef.Weidendorfer@gmx.de> wrote:\n> Thinking even one step further:\n> Would it make sense to define an encoding format for the content of\n> commit and tree objects inside of packs, where the SHA1 is replaced by the\n> offset of the object in this pack?\n> As exactly the SHA1 is the least compressable thing, this could promise\n> quite a benefit.\n\nI actually had the same idea the other day.  I discarded it after\nthinking about it for a minute.  Here's the problem:\n\nLets say we do this for the tree and parent IDs in a commit, because\nthese are the most commonly needed part of a commit during revision\ntraversal.  So we want to put the offset to the tree and the offset\nto each parent at the front of the commit somehow to make them very\ncheap to access.\n\nThis means that when we start to write out a commit we need to know\nthe offset to the tree that commit references.  But git-pack-objects\nsorts object by type: commit, tree, blob (I forget where tags go,\nbut they aren't important in this context).  So generally *all*\ncommits appear before the first tree.  So when we write out the first\ncommit we need to know exactly how many bytes every commit will need\n(compressed mind you) in this pack so we can determine the position\nof the first tree.  Now do this for every commit and every tree\nthat those commits use...  yes, its a lot of work to precompute\nand store all offsets before you even write out the first byte.\n\nIts even worse with parent commits because ancestors tend to appear\nbehind the commit (newest->oldest) so that \"git log\" can benefit\nfrom OS read-ahead.  So you also have to keep track of your parent\ncommmit offsets.  Not pretty.\n\nExtending that idea to tree objects (store the offset of the entry)\nmakes the issue even uglier.\n\nOh, and packs aren't entirely self-contained.  A pack is only self\ncontained in the sense that no object in the pack deltafies against\nan object outside of the pack[1].  However by design an object\n(e.g. a commit or a tree) can reference an object which is either\nloose or which is in another pack.  This is especially important\nfor every large projects where not every commit/tree/tag/blob will\nfit into one 4 giB file.\n\n**1** Except in the case of thin packs, which are used only on the\nnetwork and only to save bandwidth.\n\n> AFAIK, we currently only use these offsets for referencing objects in\n> delta chains.\n\nYes, that's a recent feature to reference a delta base.\n\n-- \n"},{"id":"298127","messageId":"200612030421.18662.Josef.Weidendorfer@gmx.de","threadId":"43065","inReplyTo":"20061203024655.GD26668@spearce.org","subject":"Re: Thoughts about memory requirements in traversals [Was: Re: [RFC] Submodules in GIT]","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-03T03:21:18Z","receivedAt":"2006-12-03T03:21:18Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Sunday 03 December 2006 03:46, Shawn Pearce wrote:\n> Josef Weidendorfer <Josef.Weidendorfer@gmx.de> wrote:\n> > Thinking even one step further:\n> > Would it make sense to define an encoding format for the content of\n> > commit and tree objects inside of packs, where the SHA1 is replaced by the\n> > offset of the object in this pack?\n> > As exactly the SHA1 is the least compressable thing, this could promise\n> > quite a benefit.\n> [...]\n> \n> This means that when we start to write out a commit we need to know\n> the offset to the tree that commit references.  But git-pack-objects\n> sorts object by type: commit, tree, blob (I forget where tags go,\n> but they aren't important in this context).  So generally *all*\n> commits appear before the first tree.  So when we write out the first\n> commit we need to know exactly how many bytes every commit will need\n> (compressed mind you) in this pack so we can determine the position\n> of the first tree.  Now do this for every commit and every tree\n> that those commits use...  yes, its a lot of work to precompute\n> and store all offsets before you even write out the first byte.\n\nYes, it looks like a hen-and-egg problem, but IMHO you can\nhandle it nicely with another redirection, i.e. a table you build\nup while repacking the file, and storing this table at the end.\n\nYou simply sequentially renumber any object SHA, starting from 0\nin the order you see them. You can do two renumberings, one for\nthe objects contained in the original pack (1), and one for the\nexternal ones (2). Put these new numbers (with a bit distinguishing\n(1) and (2)) as replacement into commit/tree objects.\nAt the end, you have the new offsets for objects in (1). Put\nredirection tables for (1) [new number -> new offset]\nand (2) [other new number->SHA1 of external object] at the end\nof the new pack.\nThis way, you effectivly have removed all incompressable SHAs from\nthe pack file aside from one entry in the redirection tables for\neach external object.\n\nThe only problem I see is how to decode the objects, i.e. how to\nget the original SHA1 from an offset: we can not recalculate the\nSHA1 from the object content as we changed the content itself.\nBut there should be a way to store the SHA1 in front of the object\nsomehow, perhaps it is already given by the current format? \n\nAm I missing something here?\n\n"},{"id":"298696","messageId":"20061203062916.GY18810@admingilde.org","threadId":"43065","inReplyTo":"200612030155.09630.Josef.Weidendorfer@gmx.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-03T06:29:16Z","receivedAt":"2006-12-03T06:29:16Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Sun, Dec 03, 2006 at 01:55:08AM +0100, Josef Weidendorfer wrote:\n> On Saturday 02 December 2006 21:24, Martin Waitz wrote:\n> > On Sat, Dec 02, 2006 at 12:17:44AM +0100, Josef Weidendorfer wrote:\n> > > After some thinking, a submodule namespace even is important for checking\n> > > out only parts of a supermodule, exactly because the root of a submodule\n> > > potentially can change at every commit.\n> > \n> > have you ever thought about the idea that the location may be an\n> > important thing to consider for your decision.\n> \n> Which decision, for what? Sorry, I do not understand.\n\nto check out, or not to check out.\n\n> What should such a general partial tree support look like? I suppose you\n> want to configure paths which should not be checked out. As long as you\n> say that a given submodule always has to exist at a given path, you are\n> right: then, you can say: \"Please, do not check out this submodule\" which\n> is the same as \"Do not check out this path\". \n\nYou could say something like \"do not check out anything below \"test/\".\nIf then some submodule moves from \"test/foo\" to \"build/foo\", it will be\nchecked out, because this module is now not only used for testing, but\nis needed for building in the new version of the supermodule.\n\n-- \nMartin Waitz\n"},{"id":"297749","messageId":"e7bda7770612030119v197cbc95h6b3fa9e22b78c058@mail.gmail.com","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612021114270.3476@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Torgil Svensson","fromEmail":"torgil.svensson@gmail.com","sentAt":"2006-12-03T09:19:13Z","receivedAt":"2006-12-03T09:19:13Z","isPatch":false,"sender":{"key":"torgil.svensson@gmail.com","avatar":null},"body":"On 12/2/06, Linus Torvalds <torvalds@osdl.org> wrote:\n>\n> In other words, I don't think people expect or want something hugely more\n> complicated than the CVS/modules kind of file.\n\nWhat about the case when you want _everything_, do you then have to\nknow the names of all submodules, present and past?\n\nIf you have an old irrelevant submodule in the history that happens to\nhave the same name as one of them you are interested in, do you get\nthis as well?\n\nDuring a debugging session it might be convenient to do a \"all but X\"\nkind of fetch if you have a project dependent on several small modules\nand one of them is the big black sheep.\n\nFor simple cases, I think it's sufficient to have the \"everyone or\nno-one\" option. If git enforces sending submodules one by one and\nrequires the fetching side to specify links explicitly couldn't the\nselection be left to the user to decide with \"hooks\" or plumbing?\nDefault hook could implement a simple white- or black-list.\n\n"},{"id":"294961","messageId":"200612030942.20946.andyparkins@gmail.com","threadId":"43065","inReplyTo":"200612021255.59972.Josef.Weidendorfer@gmx.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2006-12-03T09:42:19Z","receivedAt":"2006-12-03T09:42:19Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Saturday 2006, December 02 11:55, Josef Weidendorfer wrote:\n\n> > > \t100644 blob 08602f522183dc43787616f37cba9b8af4e3dade\txdiff-interface.c\n> > > \t100644 blob 1346908bea31319aabeabdfd955e2ea9aab37456\txdiff-interface.h\n> > > \t040000 tree 959dd5d97e665998eb26c764d3a889ae7903d9c2\txdiff\n> > > \t050000 link 0215ffb08ce99e2bb59eca114a99499a4d06e704\txyzzy\n> > >\n> > > where that 050000 is the new magic type (I picked one out of my *ss:\n> > > it's not a valid type for a file mode, so it's a godo choice, but it\n> > > could be anythign that cannot conflict with a real file), which just\n> > > specifies the \"link\" part. The SHA1 is the SHA1 of the commit, and the\n> > > \"xyzzy\" is obviously just the name within the directory of the\n> > > submodule.\n> >\n> > Can I argue that the hash in that object should actually be to a real\n> > object in the supermodule repository rather than a link?\n>\n> That is the thing we already are discussing here :-)\n> IMHO, submodule IDs make a lot of sense, and this needs to specify the\n> submodule ID at every link. Which would force us to use seperate objects.\n\nI wasn't really going as deep as a submodule ID.  Just moving the submodule \ncommit hash from the supermodule tree, to a supermodule \"link\" object.  What \ngoes in that object is a separate problem I believe.\n\nThe primary reason I think it's a good idea is that it is consistent with \nevery other hash in the tree.  It seems to be inconsistent to say\n\nblob objects have a hash that points to an object in this repo\ntree objects have a hash that points to an object in this repo\nlink objects have a hash the points to an object in a different repo\n\n> However, I am not speaking about some separation issue, but more about a\n> design decision. For fetching/pulling/merging, you want be able to\n> distinguish submodules not only by the commit id into the submodule:\n> multiple\n> submodules could link into the same DAG (but different branches) of another\n> repository which would make unique fetching/pulling/merge decisions\n> difficult, especially when you think about the possibility that the\n> relative root path of a submodule in a supermodule should be able to change\n> at any supermodule commit.\n\nI can't say I've understood what you mean here.  There is no difference in \nfacilities if there is a link object in the local repository as well.  It's \nmerely an extra layer of indirection.  Apart from the tiny cost of \ndereferencing that link object, there is no disadvantage.\n\n\nAndy\n\n-- \nDr Andrew Parkins, M Eng (Hons), AMIEE\n"},{"id":"297403","messageId":"ekubag$6km$2@sea.gmane.org","threadId":"43065","inReplyTo":"200612030421.18662.Josef.Weidendorfer@gmx.de","subject":"Re: Thoughts about memory requirements in traversals [Was: Re: [RFC] Submodules in GIT]","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-03T11:10:08Z","receivedAt":"2006-12-03T11:10:08Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Josef Weidendorfer wrote:\n\n> On Sunday 03 December 2006 03:46, Shawn Pearce wrote:\n>> Josef Weidendorfer <Josef.Weidendorfer@gmx.de> wrote:\n>>> Thinking even one step further:\n>>> Would it make sense to define an encoding format for the content of\n>>> commit and tree objects inside of packs, where the SHA1 is replaced by\n>>> the offset of the object in this pack?\n>>> As exactly the SHA1 is the least compressable thing, this could promise\n>>> quite a benefit.\n>> [...]\n>> \n>> This means that when we start to write out a commit we need to know\n>> the offset to the tree that commit references.  But git-pack-objects\n>> sorts object by type: commit, tree, blob (I forget where tags go,\n>> but they aren't important in this context).  So generally *all*\n>> commits appear before the first tree.  So when we write out the first\n>> commit we need to know exactly how many bytes every commit will need\n>> (compressed mind you) in this pack so we can determine the position\n>> of the first tree.  Now do this for every commit and every tree\n>> that those commits use...  yes, its a lot of work to precompute\n>> and store all offsets before you even write out the first byte.\n> \n> Yes, it looks like a hen-and-egg problem, but IMHO you can\n> handle it nicely with another redirection, i.e. a table you build\n> up while repacking the file, and storing this table at the end.\n> \n> You simply sequentially renumber any object SHA, starting from 0\n> in the order you see them. You can do two renumberings, one for\n> the objects contained in the original pack (1), and one for the\n> external ones (2). Put these new numbers (with a bit distinguishing\n> (1) and (2)) as replacement into commit/tree objects.\n> At the end, you have the new offsets for objects in (1). Put\n> redirection tables for (1) [new number -> new offset]\n> and (2) [other new number->SHA1 of external object] at the end\n> of the new pack.\n> This way, you effectivly have removed all incompressable SHAs from\n> the pack file aside from one entry in the redirection tables for\n> each external object.\n> \n> The only problem I see is how to decode the objects, i.e. how to\n> get the original SHA1 from an offset: we can not recalculate the\n> SHA1 from the object content as we changed the content itself.\n> But there should be a way to store the SHA1 in front of the object\n> somehow, perhaps it is already given by the current format? \n> \n> Am I missing something here?\n\nDoesn't this idea clash with the object and delta reusing for repack?\nHmmm... perhaps with the two indirect tables it wouldn't, only\nthe tables would need to be recalculated... or perhaps it would because\nof offset clashes.\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n\n"},{"id":"297140","messageId":"200612031247.12957.Josef.Weidendorfer@gmx.de","threadId":"43065","inReplyTo":"ekubag$6km$2@sea.gmane.org","subject":"Re: Thoughts about memory requirements in traversals [Was: Re: [RFC] Submodules in GIT]","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-03T11:47:12Z","receivedAt":"2006-12-03T11:47:12Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Sunday 03 December 2006 12:10, Jakub Narebski wrote:\n> > You simply sequentially renumber any object SHA, starting from 0\n> > in the order you see them. You can do two renumberings, one for\n> > the objects contained in the original pack (1), and one for the\n> > external ones (2). Put these new numbers (with a bit distinguishing\n> > (1) and (2)) as replacement into commit/tree objects.\n> > At the end, you have the new offsets for objects in (1). Put\n> > redirection tables for (1) [new number -> new offset]\n> > and (2) [other new number->SHA1 of external object] at the end\n> > of the new pack.\n> \n> Doesn't this idea clash with the object and delta reusing for repack?\n\nIn general, yes: you modify object content by encoding, and if you want\nto reuse the objects without decompression, and need to keep the info\nto be able to decode, ie. the redirection table.\n\nThis gets problematic if you want to join multiple packs, or fetch objects\nand want to reuse the compressed representation, as the object renumbering\nis only local to one pack, and numbers can be reused between packs.\n\nSo this idea is probably for archival packs only.\n\n"},{"id":"298517","messageId":"Pine.LNX.4.64.0612030946150.3476@woody.osdl.org","threadId":"43065","inReplyTo":"e7bda7770612030119v197cbc95h6b3fa9e22b78c058@mail.gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-03T17:54:35Z","receivedAt":"2006-12-03T17:54:35Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 3 Dec 2006, Torgil Svensson wrote:\n>\n> On 12/2/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> > \n> > In other words, I don't think people expect or want something hugely more\n> > complicated than the CVS/modules kind of file.\n> \n> What about the case when you want _everything_, do you then have to\n> know the names of all submodules, present and past?\n\nAfaik, the way people do this historically is simply:\n\n - often have an alias for \"everything\" (eg \"all\" or \"src\" or \"world\"), \n   and if you want everything, you basically ask for it by checking out \n   the \"src\" module.\n\n   Ie this is the \"upstream\" way to let downstream check out everything.\n\n - if you're downstream, and you have a partial repo, and you realize that \n   you want everything else, you just look at gitweb (assuming it is \n   extended to show module information, of course ;) or the .gitmodules \n   (or whatever it would be called) file to get the other pieces manually.\n\nBut hey, I also think it would be fine to have \"git clone --allmodules\" or \nsomething (\"fetch\" too). I think this whole question will depend more on \nhow people end up _using_ module support than on any technical issues per \nse. Again, I suspect the people who now set up modules in CVS are likely \nto have a better idea than I do about how they usually do it (and why).\n\n> If you have an old irrelevant submodule in the history that happens to\n> have the same name as one of them you are interested in, do you get\n> this as well?\n\nI dunno. Details, details. I'm also not sure this is hugely important.\n\nIt could be \"solved\" by simply having the requirement that all modules \nneed to be named differently (notice that \"module name\" is _not_ the same \nthing as \"the directory name where the module shows up\". That's not the \ncase even in CVS modules, and with a \"link\" type in the git tree object, \nthe directory where a module shows up would basically be totally \nindependent of the \"name\" of the module).\n\n> During a debugging session it might be convenient to do a \"all but X\"\n> kind of fetch if you have a project dependent on several small modules\n> and one of them is the big black sheep.\n\nI suspect it's more common to name the modules you want to fetch \nexplicitly, rather than make it a \"negative\" choice, but that sounds \nlargely like just an interface issue.\n\n"},{"id":"298691","messageId":"200612031933.42968.andyparkins@gmail.com","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612021114270.3476@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2006-12-03T19:33:41Z","receivedAt":"2006-12-03T19:33:41Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Saturday 2006, December 02 19:41, Linus Torvalds wrote:\n\n> Now, I'm not exactly sure who wants to use git modules, so this is the\n> time to ask: did you hate the CVS/modules file? Or was it something you\n> set up once, and then basically forgot about? People clearly use the\n> ability to mark certain modules as depending on each other, and aliases to\n> say \"if you ask for this module, you actually get a set of _these_\n> modules\".\n\nNever used CVS/modules, but I used svn:externals.  I have a few projects that \nare libraries that I use in many other projects.  So, my directory tree looks \nlike this:\n\n projects/\n   libX/\n   projectP/\n    libX/\n   projectQ/\n    libX/\n\nThe nightmare I had was that I would add a feature to projectP/libX, and \ncommit it.  Great.  Then later I'd do \"svn update\" in projectQ - HAVING MADE \nNO CHANGES TO IT - and libX would update to the latest version, which turns \nout to be incompatible with projectQ, and I can no longer even build \nprojectQ.  If only libX would stay where it was put.  The worst of it is if \nyou check out an older version, say \"stable-release\" that you tagged last \nyear, the svn:external would always just check out the latest version, so \nyou'd have to go back through the logs to find out what approximate submodule \nrevision you should really check out, check it out and then remember not to \ndo svn update, because that would just reset the external to the latest \nversion.  AHHHHHHH!  Maddening to say the least.\n\nThis fits exactly with what you have described as the primary reason for \nwanting submodules.  I didn't want seamless integration, I was happy to \nchange into projectP/libX to make libX commits.  All I actually wanted was \nthe particular checkout of libX for a particular checkin of projectP to be \nremembered.  That's it.  Anything else is just gravy.\n\nI'm doing exactly the same sort of thing now but with git.  git hasn't fixed \nthe problem (yet) but certainly hasn't made it any worse than it was.  \nsvn:externals were nothing more than a way of storing a URL in the \nrepository - who cares, I wish now I'd never bothered, they serve no version \ncontrol purpose and are merely a UI convenience.\n\n\n\n\nAndy\n-- \nDr Andrew Parkins, M Eng (Hons), AMIEE\n"},{"id":"294388","messageId":"20061203204644.GZ18810@admingilde.org","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612021242080.3476@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-03T20:46:44Z","receivedAt":"2006-12-03T20:46:44Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Sat, Dec 02, 2006 at 12:44:20PM -0800, Linus Torvalds wrote:\n> And watch the memory usage.\n\nhmm, really sad, it was such a nice concept until now...\nYou are right, I have to think more about scalability. O(N) anywhere is\nreally bad for submodules.  They really should be able to bundle the kernel,\nmozilla, qt and whatnot into one project and that will get huge.\n\n-- \nMartin Waitz\n"},{"id":"295896","messageId":"20061203221630.GA940MdfPADPa@greensroom.kotnet.org","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612011510290.3695@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Sven Verdoolaege","fromEmail":"skimo@kotnet.org","sentAt":"2006-12-03T22:16:30Z","receivedAt":"2006-12-03T22:16:30Z","isPatch":false,"sender":{"key":"skimo@kotnet.org","avatar":null},"body":"On Fri, Dec 01, 2006 at 04:12:10PM -0800, Linus Torvalds wrote:\n> So within the supermodule, on a \"git object\" level, a submodule should \n> just be named by the SHA1 that was it's HEAD when it was committed within \n> the supermodule. So in the \"tree object\", you'd see something like the \n> following when you go \"git ls-tree HEAD\" on the superproject:\n> \n> \t...\n> \t100644 blob 08602f522183dc43787616f37cba9b8af4e3dade\txdiff-interface.c\n> \t100644 blob 1346908bea31319aabeabdfd955e2ea9aab37456\txdiff-interface.h\n> \t040000 tree 959dd5d97e665998eb26c764d3a889ae7903d9c2\txdiff\n> \t050000 link 0215ffb08ce99e2bb59eca114a99499a4d06e704\txyzzy\n> \n> where that 050000 is the new magic type (I picked one out of my *ss: it's \n> not a valid type for a file mode, so it's a godo choice, but it could be \n> anythign that cannot conflict with a real file), which just specifies the \n> \"link\" part. The SHA1 is the SHA1 of the commit, and the \"xyzzy\" is \n> obviously just the name within the directory of the submodule.\n> \n> That's all that is actually required for a lot of git commands that \n> already expect all objects to be available (ie \"git checkout\", \"git diff\" \n> etc).\n\nBut is this object (and all the objects it points to) going to be\navailable (in the superproject) ?\nThe following seems to suggest that you think they shouldn't.\nHow is fsck-objects then going to check that such an object is\nvalid ?  Is it going to call fsck-objects recursively on the\n(available) submodules ?\n\nOn Fri, Dec 01, 2006 at 03:30:32PM -0800, Linus Torvalds wrote:\n> The only thing that a submodule must NOT be allowed to do on its own is \n> pruning (and it's distant cousin \"git repack -d\").\n\nHow are you going to enforce this if the submodule isn't supposed\nto know that it is being used as a submodule ?\n\n> You must always prune \n> from the supermodule, because the submodule cannot really know on its own \n> what references point into it.\n\nHow is one of the supermodules going to know what references from other\nsupermodules containing the submodule point into the submodule ?\n\n"},{"id":"298766","messageId":"Pine.LNX.4.64.0612031421030.3476@woody.osdl.org","threadId":"43065","inReplyTo":"20061203221630.GA940MdfPADPa@greensroom.kotnet.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-03T22:32:02Z","receivedAt":"2006-12-03T22:32:02Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 3 Dec 2006, Sven Verdoolaege wrote:\n> \n> On Fri, Dec 01, 2006 at 03:30:32PM -0800, Linus Torvalds wrote:\n> > The only thing that a submodule must NOT be allowed to do on its own is \n> > pruning (and it's distant cousin \"git repack -d\").\n> \n> How are you going to enforce this if the submodule isn't supposed\n> to know that it is being used as a submodule ?\n\nNote that there's actually two \"submodules\":\n\n - there's the submodule \"project\" itself.\n\n   This one must be totally unaware of the supermodule, because this one \n   might be cloned and copied _independently_ of the supermodule.\n\n - there's the PARTICULAR CHECKED-OUT COPY of the submodule that is \n   actually checked out in a supermodule.\n\n   This is just a specific _instance_ of the particular submodule.\n\nSo a particular instance of a submodule might be \"aware\" of the fact that \nit's a submodule of a supermodule. For example, the \"awareness\" migth be \nas simple as just a magic flag file inside it's .git/ directory. And that \nawareness would be what simply disabled pruning or \"repack -d\" within that \nparticular instance.\n\nBut this magic flag doesn't affect the bigger-picture git repository. It's \na _private_ flag. So it doesn't affect the git part, any more than it \nreally affects the git repository that you may have a\n\n\t[user]\n\t\tname = Myname\n\t\temail = myemail\n\nin your .git/config file.\n\nSee? You can have private data in a git repository, but that doesn't mean \nthat it's visible as _repository_ data. But it can still affect how git \ncommands act (eg the \"user\" definitions above will affect the default user \ninformation that \"git commit\" uses, of course, without actually affecting \nthe git archive in any other way)\n\n> How is one of the supermodules going to know what references from other\n> supermodules containing the submodule point into the submodule ?\n\nWhy would it care? They are other supermodules. It doesn't matter, the \nsame way it doesn't matter that _my_ \"git\" tree may not have all the same \nreferences that _your_ \"git\" repo has. If I want to get the same \nreferences, I'd need to fetch them from you, and at that point, I'd need \nto get all the objects that are pointed to by those refs too. But only on \n\"git fetch\" do you actually start caring.\n\n"},{"id":"295385","messageId":"ekvk8s$c3d$1@sea.gmane.org","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612031421030.3476@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-03T22:49:03Z","receivedAt":"2006-12-03T22:49:03Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Linus Torvalds wrote:\n\n> On Sun, 3 Dec 2006, Sven Verdoolaege wrote:\n>> \n>> On Fri, Dec 01, 2006 at 03:30:32PM -0800, Linus Torvalds wrote:\n>>> The only thing that a submodule must NOT be allowed to do on its own is \n>>> pruning (and it's distant cousin \"git repack -d\").\n>> \n>> How are you going to enforce this if the submodule isn't supposed\n>> to know that it is being used as a submodule ?\n> \n> Note that there's actually two \"submodules\":\n> \n>  - there's the submodule \"project\" itself.\n> \n>    This one must be totally unaware of the supermodule, because this one \n>    might be cloned and copied _independently_ of the supermodule.\n> \n>  - there's the PARTICULAR CHECKED-OUT COPY of the submodule that is \n>    actually checked out in a supermodule.\n> \n>    This is just a specific _instance_ of the particular submodule.\n> \n> So a particular instance of a submodule might be \"aware\" of the fact that \n> it's a submodule of a supermodule. For example, the \"awareness\" migth be \n> as simple as just a magic flag file inside it's .git/ directory. And that \n> awareness would be what simply disabled pruning or \"repack -d\" within that \n> particular instance.\n\nIf we use objects/info/alternates (or equivalent, e.g. objects/info/modules,\nor modules file) in superproject to refer to submodule repository object\ndatabase (so superproject has access to all the objects including\nsubmodule), I'd prefer to have in submodule objects/info/borrowers file,\nwhich would point to superproject (and to other repositories which have\nsubmodule as one of alternate object databases) for git-prune and friends\nto check which parts are truly unreachable.\n\nThis would be generic solution to the problem with alternates, not only\nspecific to submodule support.\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n\n"},{"id":"297154","messageId":"200612041212.55911.Josef.Weidendorfer@gmx.de","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612031421030.3476@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-04T11:12:55Z","receivedAt":"2006-12-04T11:12:55Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Sunday 03 December 2006 23:32, Linus Torvalds wrote:\n> So a particular instance of a submodule might be \"aware\" of the fact that \n> it's a submodule of a supermodule. For example, the \"awareness\" migth be \n> as simple as just a magic flag file inside it's .git/ directory. And that \n> awareness would be what simply disabled pruning or \"repack -d\" within that \n> particular instance.\n\nThat prohibits the problem in your supermodule and your instance of the\ngiven submodule.\n\nBut IMHO, using a submodule commit which could be removed by pruning in\nanother instance of the submodule is really not the thing you ever want.\nIf you start your own branch in a submodule, and start to rely on it in\nthe supermodule, you _will_ want to push this to the submodule upstream.\n\nAnd if you find that you have to rebase in the submodule, you simply\nhave to rewrite your branch commits in the supermodule too. Otherwise,\nyou effectively fork the submodule project purely for your superproject.\n\nSo I suppose that in practical use, pruning in submodules probably\nwould not have any negative effect. If it has, you made something\nwrong. So you probably should only a submodule commit if it has\n\"publishing quality\" (unless being on a temporary supermodule branch).\n\nIe. any \"borrowers\" file should be empty.\n\n"},{"id":"294115","messageId":"f2b55d220612041056w68db5891t105054c0d35efe45@mail.gmail.com","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612011621380.3695@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Michael K. Edwards","fromEmail":"medwards.linux@gmail.com","sentAt":"2006-12-04T18:56:22Z","receivedAt":"2006-12-04T18:56:22Z","isPatch":false,"sender":{"key":"medwards.linux@gmail.com","avatar":null},"body":"(I wrote most of this a couple of days ago, so it's not at the tip of\nthe conversational tree, so to speak.  But it's effectively a response\nto Linus's \"what do you want to do with submodules\" question, with\nsome thoughts on implementation.  Sorry it's so long; like Blaise\nPascal, \"I would have written a shorter letter, but I did not have the\ntime.\")\n\nThe supermodule concept, implemented right, could really improve\ncooperation among embedded platform integrators, boutique distro\npublishers, and other editorial contributors to sprawling metaprojects\nwho don't want to run kernel.org-scale mirrors.  To make this work,\nyou need sparse repositories (conserving resources when fetching, by\nomitting the bulk of currently un-needed submodules that can reliably\nbe obtained later from elsewhere) and shallow cloning (conserving\nresources when publishing, by referring cloners to a third-party\nrepository for universally available content).\n\nFor instance, it would be a wonderful thing if the pile-o-patches\nnightmare that is PTXdist (and crosstool and buildtool and every other\napproach I have seen for ongoing maintenance of embedded toolchains\nand userlands) were obsoleted by a git supermodule.  Its submodules\nwould mostly track external projects, but would also logically contain\nthe fix-up patches worked out during platform integration, checked in\nto branches anchored at each upstream release point.  The supermodule\nwould contain all of the build automation, log auditing, and remote\nunit testing stuff, as well as the metadata for each submodule\ninvolved in this platform build cycle.\n\nAt a content level, the sparsely populated / shallowly published\nsupermodule wouldn't be much different from today's PTXdist.  But the\npay-off comes when you merge forward to a new release of some base\ncomponent (compiler, library, etc.) and discover that some of your\nfix-ups have been adopted or obsoleted upstream, and new fix-ups are\nneeded for components that depend on the updated bit, and the set of\nconfigurables has changed (for which you need to compensate in the\nmeta-configurator).  Instead of piling up versioned patch directories,\nyou commit fix-ups to the sub-modules, which other integration\nbranches can ignore (if they aren't affected), merge, or cherry-pick.\n\nAs I understand it, in today's git, every content object is a patch to\nthe _data_ of one and only one git repository, containing the label of\nthe preceding _data_ state plus a diff of file contents and\nattributes.  Assuming this model is retained, any clean state of a\n\"leaf\" module (one with no submodules) can be reached by replaying a\nseries of patches, starting from the repository's root node (an empty\ndirectory with the hopefully unique label generated by init-db).  The\nlabel (SHA1) of the last patch is therefore a perfectly good label for\nthis _data_ state.\n\nIf all we were trying to do with supermodules was to capture and track\nvarious states of the submodules' data, we could extend the format of\ncontent objects to include \"state X of submodule with init-db label\nY\".  That would have the effect of capturing submodule states as\n_data_ in non-\"leaf\" modules.  We would have to help cloners find a\nplace from which to pull these states, of course; and it's easy to get\nsidetracked onto that part of the problem.  But that's not where the\nbang for the buck is in supermodules.\n\nThe whole model of distributed supermodules, with references to\nslightly diverging submodules whose content should mostly be fetched\nfrom external sources, smells to me just like LVM.  The external\nsources (like an LVM volume of which you have taken a \"snapshot\") make\nup the bulk of the content pool.  They also give you a window into\ndevelopments on the submodule's own branches (like being able to peek\nforward and merge changes from the original volume).  The supermodule\n(the snapshot volume) provides most of the interesting refs (submodule\ncommits referenced by supermodule tags and branch heads), along with\nenough \"journaled\" content to replay forward from some checkpoint\nguaranteed to be available in each external source to any of these\nrefs.\n\nThe implication here is that submodule states are not just SHA1 labels\nto be embedded within supermodule data diffs.  One ought to be able to\nclone a supermodule without immediately cloning full copies of any of\nits submodules.  This ought to populate the clone's content database\nwith all of the quanta of submodule content that aren't guaranteed to\nbe available from any not-too-stale submodule mirror.  When cloning,\nyou don't want to have to inspect every supermodule state for\nsubmodule states that are outside the global subset.  So the\nsupermodule needs to maintain a set of supplemental refs from which\nall referenced submodule states can be reached.  This allows you to\ntraverse the portion of the pool of submodule content that can't be\nreached from true submodule branch heads.\n\nOn 12/1/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> Yes, you do need to have a list of submodules somewhere, and you'd need to\n> maintain that separately. One of the results of having the submodules be\n> independent from the supermodule is that it's not all \"automatically\n> integrated\", and thus the supermodule does end up having to have things\n> like that maintained separately.\n\nThis is not a defect; it's a virtue.  It's important for every commit\nto the supermodule to contain the information of which submodule\nbranches you're currently on and how far along them you've crawled.\nAny particular supermodule commit point is likely to reflect an\nintegration milestone visible only to the person working at the\nsupermodule level.  No content object should ever cross a submodule\nboundary, because then you wouldn't be able to apply it to the\nsubmodule in isolation (or in another supermodule state) or identify\nit when it is applied upstream and propagates back to you in a pull.\n But the supermodule can also contain supplemental refs (heads and\ntags) that don't exist in the submodule (and shouldn't necessarily be\npushed to it); the commits they refer to are localized to the\nsubmodule but may not be reachable from any of the submodule's branch\nheads.\n\n> And yes, if you screw that up, you wouldn't be able to fetch submodules\n> properly etc, even if you see the supermodule, and yes, this sounds more\n> like the CVS \"Entries\" kind of file that is more \"tacked on\" than really\n> deeply integrated. But I think the separation is _more_ than worth the\n> fact that you can see things being separate.\n\nThere is an opportunity for useful deep integration here.  The same\nalgorithm that does reachability analysis for \"git prune\" can dig from\nsupermodule down to submodules, copying objects into the supermodule\ndatabase until it hits a commit that is advertised as \"global\" by the\nsubmodule.  \"git clone\" of the supermodule can then pull the bulk of\nthe submodules (a superset of the \"global\" subset) from (a mirror of)\nthe canonical place for each, and use the supermodule object database\nas an alternate source for commits that don't exist in the \"canonical\"\nsubmodule.\n\n> In fact, I'm very much arguing for keeping things as separate as possible,\n> while just integrating to the smallest possible degree (just _barely_\n> enough that you can do things like \"git clone\" and it will fetch multiple\n> repositories and put them all in the right places, and \"git diff\" and\n> friends will do reasonably sane things).\n>\n> Keep it simple, stupid.\n\nAs simple as possible; but no simpler.  The \"alternates\" / \"git clone\n--reference\" model is already almost powerful enough for the\nsupermodule to contain a \"journal\" of submodule commits that haven't\nyet been retired to the canonical subset (guaranteed present in each\nmirror).  The only difference is that the supermodule should be\nconsidered a \"weak alternates\" source.  Commit objects in the\nsupermodule's database should be visible to submodule-level operations\n(so that commits which are accepted upstream get flowed in nicely\nduring \"git pull\").\n\nBut if a commit becomes reachable from a ref that is really in the\nsubmodule (not just one of the supermodule's \"supplemental refs\",\nwhich should _not_ be visible to submodule operations), then it should\nbe copied into the submodule's object database.  (The refs internal to\nthe submodule should retain their integrity even if the supermodule is\ninaccessible.)  The existing \"strong alternates\" mechanism should be\nreserved for repos which are at least as public and persistent as the\nsubmodule, and supermodules don't qualify (e. g., Linus's transmeta\nscenario).\n\n> On Sat, 2 Dec 2006, Josef Weidendorfer wrote:\n> > The thing I wanted to discuss is whether such names would need to be globally\n> > unique in the project containing submodles, or not.\n>\n> My preference would be for it to be \"local\", just because (as I\n> mentioned), with mirroring etc, it might well be that you want to fetch\n> things from the _closest_ repository. That's really not a global decision,\n> it's a local one.\n\nI think \"global resource, local provider\" is the way to go, with each\nprovider advertising what checkpoints of what resources it can supply.\n When I clone or pull, I should be able to consult a local mapping of\nsubmodule URIs to \"mirrors\" (which may well be local repositories\ncontaining content and branches that aren't in the \"official\"\nupstream).  The only thing that may need \"global\" agreement is the\nboundaries of the \"global\" subset for each submodule, i. e., the set\nof commit objects that can reliably be obtained from any mirror of the\n\"official\" upstream repository.  That doesn't need to be terribly\nclever; \"at least three days old on a globally published branch\" would\nprobably be a perfectly good heuristic.\n\n> > If yes, it IMHO makes a lot of sense to introduce \"submodule objects\" which contain\n> > these submodule names, and which are used as pointers to submodule commits in\n> > supermodule trees.\n>\n> You could do it that way, and then it would be global. It would work, and\n> in many ways it would probably be \"simpler\" on a supermodule level.\n\nI think the implication of \"submodule objects\" is that supermodule\ndiffs would say \"roll submodule X from commit-id A to commit-id B\".  I\ndon't think that would work very well for pulls/merges in the sparsely\npopulated scenario, because you want to be able to pull the\nnon-canonical subset of the individual diffs between states A and B\ninto the supermodule's object pool.  When you decide later to flesh\nout submodule X, you should only have to clone some canonical mirror\nand then fast-forward to state B using objects you already have in the\nsupermodule pool.\n\nThe merge case is even clearer.  Suppose I pull updates from two\nremote branches of the supermodule onto my master branch.  Each remote\nbranch has added the same submodule, cloned from third-party\nrepositories whose clone history goes back to the same origin.  (The\nexample I have in mind is when some project switches to git from some\nother SCM, and the maintainers of the remote branches port their\nintegration patches over from their git-svn tracker submodule to a\nclone of upstream's new git repo.)  I should be able to postpone the\nmerge effort, come back later and clone the upstream repo, then merge\nthe non-canonical commits that were pulled earlier.\n\nI might want to decide at supermodule pull time to postpone pulling\nthe bodies of the submodule commits; but I want the full sequence of\nsubmodule commit IDs in the supermodule commit object.  So it's not so\nmuch the supermodule _state_ that has a hierarchical structure; it's\nthe supermodule _diffs_ and _object_pool_ that become hierarchical.\n\n> The advantage of a global namespace is that you can much more easily\n> update it - \"git fetch\" will just fetch the new file(s) that describe the\n> subprojects very naturally if they are all global. Putting them in a local\n> .git/config file has it's advantages (see above), but it also makes it\n> very hard to version them, and to update the list - it would have to\n> become manual.\n\nI think the only global-to-local-namespace mapping applies to the\ndifferent labels for the \"empty repository\" state generated at init-db\ntime.  Given the init-db SHA1 of the linux kernel repository, I should\nbe able to choose any mirror or clone of that repository as a source\nfor objects in its \"global set\".  I expect this provider not to\nscribble on globally published branches, but that isn't even all that\ncritical; anything outside the canonical set is kept in the\nsupermodule's object pool, so I can always blow the submodule away and\nregenerate it from a different mirror.\n\n> There are possibly combinations of the two approaches: have a \"global\n> namespace\" that describes the canonical place to get the subprojects, but\n> have some way to add local \"translation\" of the canonical names into\n> locally preferred versions (eg you could just have a way to say \"this is\n> the local mirror for that global canonical place\")\n>\n> Maybe that would work?\n\nSure.  But all you really need from the canonical place is its init-db\nSHA1 (permanent) and its list of globally published branches\n(monotonically expanding).  A URL for it is a convenient shorthand but\ndoesn't have to be persistent.\n\nCheers,\n"},{"id":"296596","messageId":"e7bda7770612041226j4d4a5584m279afa9a2d7dfe74@mail.gmail.com","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612030946150.3476@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Torgil Svensson","fromEmail":"torgil.svensson@gmail.com","sentAt":"2006-12-04T20:26:09Z","receivedAt":"2006-12-04T20:26:09Z","isPatch":false,"sender":{"key":"torgil.svensson@gmail.com","avatar":null},"body":"On 12/3/06, Linus Torvalds <torvalds@osdl.org> wrote:\n>\n> > If you have an old irrelevant submodule in the history that happens to\n> > have the same name as one of them you are interested in, do you get\n> > this as well?\n>\n> It could be \"solved\" by simply having the requirement that all modules\n> need to be named differently (notice that \"module name\" is _not_ the same\n> thing as \"the directory name where the module shows up\".\n\nOkay, missed that part.  I wasn't familiar with contents of the CVS\nmodules files and misinterpreted your suggestion.\n\nMODULE [OPTIONS] [&OTHERMODULE...] [DIR] [FILES]\n\nSo all this is UI only and the \"normal\" operations on the supermodule\n"},{"id":"297677","messageId":"Pine.LNX.4.64.0612041234390.3476@woody.osdl.org","threadId":"43065","inReplyTo":"e7bda7770612041226j4d4a5584m279afa9a2d7dfe74@mail.gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-04T20:41:23Z","receivedAt":"2006-12-04T20:41:23Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 4 Dec 2006, Torgil Svensson wrote:\n> \n> Okay, missed that part.  I wasn't familiar with contents of the CVS\n> modules files and misinterpreted your suggestion.\n> \n> MODULE [OPTIONS] [&OTHERMODULE...] [DIR] [FILES]\n> \n> So all this is UI only and the \"normal\" operations on the supermodule\n> will just ignore what's behind the commit-links?\n\nRight. That's how CVS modules work (although in the case of CVS modules, \nthe \"dir\" thing is obviously there in the \"modules\" file, so it's not \n_purely_ UI in CVS - this would likely be different in a git \nimplementation, because the \"tree\" object ends up telling not just the \nexact version, but the location too).\n\nSo my suggestion basically boils down to:\n\n - \"fetch\" and \"clone\" etc will just look at the \"modules\" file, and \n   recursively fetch/clone whatever the module files talks about. This is \n   the \"thin veneer to make it _look_ like git actually understands \n   submodules\" part. It woudln't really - they're very much tacked on.\n\n - the tree entries are what makes the \"once you have all the submodule \n   objects, this is how you can do 'diff' and 'checkout' on them, and this \n   is what tells you the exact version that goes along with a particular \n   supermodule version\".\n\nIn other words, the simple and stupid way to do this is to just consider \nthese two things two totally independent issues, and have different \nmechanisms for telling different operations what to do.\n\nIs it \"pretty\"? No. The whole sub-module thing wouldn't be a tightly \nintegrated low-level thing, it would very much be all about tracking \nmultiple _separate_ git repositories, and just make them work well \ntogether. They'd very much still be separate, with just some simple \ninfrastructure glue to make them look somewhat integrated.\n\nSo yeah, it's a bit hacky, but for the reasons I've tried to outline, I \nactually think that users _want_ hacky. Exactly because \"deep integration\" \nends up having so many _bad_ features, so it's better to have a thin and \nsimple layer that you can actually see past if you want to.\n\n"},{"id":"294097","messageId":"e7bda7770612041336s73e677ebh758b030f9f75c1d8@mail.gmail.com","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612041234390.3476@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Torgil Svensson","fromEmail":"torgil.svensson@gmail.com","sentAt":"2006-12-04T21:36:48Z","receivedAt":"2006-12-04T21:36:48Z","isPatch":false,"sender":{"key":"torgil.svensson@gmail.com","avatar":null},"body":"On 12/4/06, Linus Torvalds <torvalds@osdl.org> wrote:\n>\n> So yeah, it's a bit hacky, but for the reasons I've tried to outline, I\n> actually think that users _want_ hacky. Exactly because \"deep integration\"\n> ends up having so many _bad_ features, so it's better to have a thin and\n> simple layer that you can actually see past if you want to.\n\nThin and simple sounds very good. Let's try it with an example. Lets\nsay we have one apllication App1 and three librarys (Lib1, Lib2, Lib3)\nwith the following dependency-graph:\n\n        App1\n          /\\\n         /  \\\n   Lib1   Lib2\n       \\     /\n        \\   /\n        Lib3 (don't really needed for this example but looks nice)\n\nAll components can be used individually and have their own upstream,\nmaintainer etc.\n\nTo compile App1 however, I need some files from both Lib1 and Lib2\nspecifying it's API. To satisfy these dependencies, It sounds\nreasonable to link Lib2 and Lib3 submodules from App1. In your\nconcept, can I construct a modules file to fetch the API files and\n"},{"id":"294660","messageId":"4574CBFF.8040708@vilain.net","threadId":"43065","inReplyTo":"f2b55d220612041056w68db5891t105054c0d35efe45@mail.gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2006-12-05T01:31:43Z","receivedAt":"2006-12-05T01:31:43Z","isPatch":false,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"Michael K. Edwards wrote:\n> who don't want to run kernel.org-scale mirrors.  To make this work,\n> you need sparse repositories (conserving resources when fetching, by\n> omitting the bulk of currently un-needed submodules that can reliably\n> be obtained later from elsewhere) and shallow cloning (conserving\n> resources when publishing, by referring cloners to a third-party\n> repository for universally available content).\n\nDid you see GitTorrent?  http://gittorrent.utsl.gen.nz/  A lot of\nsimilar ideas to what you mention.  Sorry, still no prototype :)\n\nI'd see the submodules thing as a good way to glue together a whole\nbunch of repositories, so that the core mirror servers only have to\nmirror a small-ish number of repositories.\n\nSam.\n"},{"id":"296123","messageId":"Pine.LNX.4.64.0612042053080.20138@iabervon.org","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612021114270.3476@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2006-12-05T02:33:12Z","receivedAt":"2006-12-05T02:33:12Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Sat, 2 Dec 2006, Linus Torvalds wrote:\n\n> So that's where I come from. And maybe I'm totally wrong. I'd like to hear \n> what people who actually _use_ submodules think.\n\nI think you'd rather hear from people who _would_ use submodules; I've \nworked on a number of projects that would have benefitted from that \ngeneral functionality, but nobody trusted the implementation enough to \nactually use it.\n\nAt my work, we're doing a bunch of stuff with microcontrollers. We've got \nabout a dozen different boards with microcontrollers, and each of them has \ndifferent firmware. We also have a bunch of code that can go on any of the \nboards.\n\nThe way things are organized currently is that each board has its own \nproject, and there's a \"common-micro\" project with the common code. This \nsort of works, but it means that when you change things in common-micro, \nyou never know what effect this will have on boards other than the one \nyou're actually working on. What I'd like to have is that each project has \na \"common-micro\" subdirectory, and changes to each of these can be merged \ninto each other, but that doesn't happen automaticly, and each board's \nrevisions include the common-micro revision they were created with.\n\nA few notes: \n\nI'd never work on common-micro in isolation. Nothing in there even \ncompiles by itself, because the compiler needs to know the target \nmicrocontroller type, which depends on the board it's for. It only makes \nsense to prepare a new revision of \"common-micro\" in the context of some \nparticular board, at least if you want to test it at all.\n\nI'd sometimes want to include temporary hacks in the common-micro for a \nparticular board, when things are late and I need to change some library \nbehavior in a way that I know works for the board I'm working on, but I \ndon't have time to think about all of the other boards (each of which is \nspecial in some way).\n\nI'd often make some change that I know improves the cleanliness of \ncommon-micro, but which requires changes to every board to compensate. I \ndon't want to make all of the changes at once; I'll update each board \nappropriately the next time I work on it. But, of course, until I update \neach board, that board needs to keep using the version of common-micro \nwithout the changes.\n\nI don't want to have repository states where stuff doesn't work, which \nmeans I can't do it as one big tree; I need to be able to make a commit \nwith board1 and a common-micro change without having a version of board2 \nthat would use the changed common-micro, because I haven't come up with a \nboard2 version that works with it yet.\n\n\t-Daniel\n"},{"id":"294813","messageId":"20061205090125.GA2428@cepheus","threadId":"43065","inReplyTo":"456F29A2.1050205@op5.se","subject":"Re: [RFC] Submodules in GIT","fromName":"Uwe Kleine-Koenig","fromEmail":"zeisberg@informatik.uni-freiburg.de","sentAt":"2006-12-05T09:01:25Z","receivedAt":"2006-12-05T09:01:25Z","isPatch":false,"sender":{"key":"u.kleine-koenig@pengutronix.de","avatar":"https://gravatar.com/avatar/354b5e3ceb2806a2f1e1e382ac29ddbdad18288654da62b61eb13583a857eee7?d=mp&s=160"},"body":"Hello,\n\nAndreas Ericsson wrote:\n> The only problem I'm seeing atm is that the supermodule somehow has to \n> mark whatever commits it's using from the submodule inside the submodule \n> repo so that they effectively become un-prunable, otherwise the \n> supermodule may some day find itself with a history that it can't restore.\nOne could circumvent that by creating a separate repo for the submodule\nat checkout time and pull the needed objects in the supermodule's odb\nwhen commiting the supermodule.  This way prune in the submodule cannot\ndo any harm, because in it's odb are no objects that are important for\nthe supermodule.\n\nUwe\n\n-- \nUwe Kleine-Koenig\n\n"},{"id":"295283","messageId":"45754AFE.1070207@op5.se","threadId":"43065","inReplyTo":"20061205090125.GA2428@cepheus","subject":"Re: [RFC] Submodules in GIT","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-12-05T10:33:34Z","receivedAt":"2006-12-05T10:33:34Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Uwe Kleine-Koenig wrote:\n> Hello,\n> \n> Andreas Ericsson wrote:\n>> The only problem I'm seeing atm is that the supermodule somehow has to \n>> mark whatever commits it's using from the submodule inside the submodule \n>> repo so that they effectively become un-prunable, otherwise the \n>> supermodule may some day find itself with a history that it can't restore.\n> One could circumvent that by creating a separate repo for the submodule\n> at checkout time and pull the needed objects in the supermodule's odb\n> when commiting the supermodule.  This way prune in the submodule cannot\n> do any harm, because in it's odb are no objects that are important for\n> the supermodule.\n> \n\nYes, but then you'd lose history connectivity (I'm assuming you'd only \npull in the tree and blob objects from the submodule, and prefix the \ntree-entrys with whatever directory you're storing the submodul in).\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\n"},{"id":"297447","messageId":"45754C0E.3070904@op5.se","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612041234390.3476@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-12-05T10:38:06Z","receivedAt":"2006-12-05T10:38:06Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Linus Torvalds wrote:\n> \n> On Mon, 4 Dec 2006, Torgil Svensson wrote:\n>> Okay, missed that part.  I wasn't familiar with contents of the CVS\n>> modules files and misinterpreted your suggestion.\n>>\n>> MODULE [OPTIONS] [&OTHERMODULE...] [DIR] [FILES]\n>>\n>> So all this is UI only and the \"normal\" operations on the supermodule\n>> will just ignore what's behind the commit-links?\n> \n> Right. That's how CVS modules work (although in the case of CVS modules, \n> the \"dir\" thing is obviously there in the \"modules\" file, so it's not \n> _purely_ UI in CVS - this would likely be different in a git \n> implementation, because the \"tree\" object ends up telling not just the \n> exact version, but the location too).\n> \n> So my suggestion basically boils down to:\n> \n>  - \"fetch\" and \"clone\" etc will just look at the \"modules\" file, and \n>    recursively fetch/clone whatever the module files talks about. This is \n>    the \"thin veneer to make it _look_ like git actually understands \n>    submodules\" part. It woudln't really - they're very much tacked on.\n> \n>  - the tree entries are what makes the \"once you have all the submodule \n>    objects, this is how you can do 'diff' and 'checkout' on them, and this \n>    is what tells you the exact version that goes along with a particular \n>    supermodule version\".\n> \n> In other words, the simple and stupid way to do this is to just consider \n> these two things two totally independent issues, and have different \n> mechanisms for telling different operations what to do.\n> \n> Is it \"pretty\"? No. The whole sub-module thing wouldn't be a tightly \n> integrated low-level thing, it would very much be all about tracking \n> multiple _separate_ git repositories, and just make them work well \n> together. They'd very much still be separate, with just some simple \n> infrastructure glue to make them look somewhat integrated.\n> \n> So yeah, it's a bit hacky, but for the reasons I've tried to outline, I \n> actually think that users _want_ hacky. Exactly because \"deep integration\" \n> ends up having so many _bad_ features, so it's better to have a thin and \n> simple layer that you can actually see past if you want to.\n> \n\nIndeed. With the \"tight\" integration option we'd also have to have the \nmechanism to rewrite the tree-entries with the location where the \nsubmodule is located in the working tree. This might be needed anyways, \nbut it sure as hell seems a lot easier to just tack that part on when \ndoing a checkout and actually creating all the files.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\n"},{"id":"295325","messageId":"45754D27.9070701@op5.se","threadId":"43065","inReplyTo":"e7bda7770612041336s73e677ebh758b030f9f75c1d8@mail.gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-12-05T10:42:47Z","receivedAt":"2006-12-05T10:42:47Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Torgil Svensson wrote:\n> On 12/4/06, Linus Torvalds <torvalds@osdl.org> wrote:\n>>\n>> So yeah, it's a bit hacky, but for the reasons I've tried to outline, I\n>> actually think that users _want_ hacky. Exactly because \"deep \n>> integration\"\n>> ends up having so many _bad_ features, so it's better to have a thin and\n>> simple layer that you can actually see past if you want to.\n> \n> Thin and simple sounds very good. Let's try it with an example. Lets\n> say we have one apllication App1 and three librarys (Lib1, Lib2, Lib3)\n> with the following dependency-graph:\n> \n>        App1\n>          /\\\n>         /  \\\n>   Lib1   Lib2\n>       \\     /\n>        \\   /\n>        Lib3 (don't really needed for this example but looks nice)\n> \n> All components can be used individually and have their own upstream,\n> maintainer etc.\n> \n> To compile App1 however, I need some files from both Lib1 and Lib2\n> specifying it's API. To satisfy these dependencies, It sounds\n> reasonable to link Lib2 and Lib3 submodules from App1. In your\n> concept, can I construct a modules file to fetch the API files and\n> their history without checking out the whole Lib1 and Lib2 source?\n\nI think not. Then it wouldn't be a submodule anymore, but just some \nrandom sources from an upstream project. Not that it's an uncommon \nworkflow or anything, but it's sort of akin to just importing the SHA1 \nimplementation (a few source-files with no real interest in the history \nof those source-files) from openssl into a different project rather than \nactually using the entire openssl lib (which would be nice to have as a \nsubmodule).\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\n"},{"id":"294438","messageId":"el3jeb$gkr$1@sea.gmane.org","threadId":"43065","inReplyTo":"45754C0E.3070904@op5.se","subject":"Re: [RFC] Submodules in GIT","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-05T11:01:28Z","receivedAt":"2006-12-05T11:01:28Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Andreas Ericsson wrote:\n\n> Indeed. With the \"tight\" integration option we'd also have to have the \n> mechanism to rewrite the tree-entries with the location where the \n> submodule is located in the working tree. This might be needed anyways, \n> but it sure as hell seems a lot easier to just tack that part on when \n> doing a checkout and actually creating all the files.\n\nExcellent idea! This way most of the concerns for \"separate repositories for\nsubmodules\" layout about ability to rename directory the submodule resides\nin, or move submodule are resolved. The other part would be to use\nsubmodule-aware git-mv to move submodule(s).\n\nPerhaps the following solution would work best:\n * refs/submodules/<module> holds sha1 of top commit in submodule\n * objects/info/submodules is a file which can be automatically generated\n   (or at least automatically updated) on checkout, with the following\n   contents:\n\n   <module> TAB or SPC <path to submodule, or GIT_DIR of submodule, or\n                        GIT_OBJECT_DIRECTORY of submodule>\n\n   with the usual rule that # and ; means comment, \\ at end of line is used\n   for continuations, empty lines doesn't matter etc.\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n\n"},{"id":"298544","messageId":"el3jsh$i62$1@sea.gmane.org","threadId":"43065","inReplyTo":"45754D27.9070701@op5.se","subject":"Re: [RFC] Submodules in GIT","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-05T11:09:02Z","receivedAt":"2006-12-05T11:09:02Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Andreas Ericsson wrote:\n\n> Torgil Svensson wrote:\n>> On 12/4/06, Linus Torvalds <torvalds@osdl.org> wrote:\n>>>\n>>> So yeah, it's a bit hacky, but for the reasons I've tried to outline, I\n>>> actually think that users _want_ hacky. Exactly because \"deep \n>>> integration\"\n>>> ends up having so many _bad_ features, so it's better to have a thin and\n>>> simple layer that you can actually see past if you want to.\n>> \n>> Thin and simple sounds very good. Let's try it with an example. Lets\n>> say we have one apllication App1 and three librarys (Lib1, Lib2, Lib3)\n>> with the following dependency-graph:\n>> \n>>        App1\n>>        /  \\\n>>       /    \\\n>>   Lib1   Lib2\n>>       \\    /\n>>        \\  /\n>>        Lib3 (don't really needed for this example but looks nice)\n>> \n>> All components can be used individually and have their own upstream,\n>> maintainer etc.\n>> \n>> To compile App1 however, I need some files from both Lib1 and Lib2\n>> specifying it's API. To satisfy these dependencies, It sounds\n>> reasonable to link Lib2 and Lib3 submodules from App1. In your\n>> concept, can I construct a modules file to fetch the API files and\n>> their history without checking out the whole Lib1 and Lib2 source?\n> \n> I think not. Then it wouldn't be a submodule anymore, but just some \n> random sources from an upstream project. Not that it's an uncommon \n> workflow or anything, but it's sort of akin to just importing the SHA1 \n> implementation (a few source-files with no real interest in the history \n> of those source-files) from openssl into a different project rather than \n> actually using the entire openssl lib (which would be nice to have as a \n> submodule).\n\nNote that this is what partial checkouts (another great idea nobody\nimplemented yet[*1*]; you can do partial checkout but there is no UI for\nthis, and working with partial checkouts is bit hard) is about, although it\nwould buy you only working area space, and not repository (object database\nstorage) space.\n\nFor now, you can imitate this by having in in Lib1 and Lib2 the 'includes'\nbranch which would contain only the API (and which you would have to keep\nup to date with 'master', but it should be fairly easy: just merge changes\ninto 'includes', perhaps with help of git-rerere, or [nonexisting]\ngit-rerere2).\n\n[*1*] Although with our track[*2*] I guess it is reasonable to think it\nwould get implemented soon.\n[*2*] Out of four \"great ideas\": shallow clone / sparse clone, submodules\nsupport, lazy clone / remote alternates, two are in example-implementation\n(submodules support) and beta work (shallow clone is in 'next').\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n\n"},{"id":"297821","messageId":"el3k0a$i62$2@sea.gmane.org","threadId":"43065","inReplyTo":"45754AFE.1070207@op5.se","subject":"Re: [RFC] Submodules in GIT","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-05T11:11:04Z","receivedAt":"2006-12-05T11:11:04Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Andreas Ericsson wrote:\n\n> Uwe Kleine-Koenig wrote:\n>\n>> Andreas Ericsson wrote:\n>>> The only problem I'm seeing atm is that the supermodule somehow has to \n>>> mark whatever commits it's using from the submodule inside the submodule \n>>> repo so that they effectively become un-prunable, otherwise the \n>>> supermodule may some day find itself with a history that it can't restore.\n>>>\n>> One could circumvent that by creating a separate repo for the submodule\n>> at checkout time and pull the needed objects in the supermodule's odb\n>> when commiting the supermodule.  This way prune in the submodule cannot\n>> do any harm, because in it's odb are no objects that are important for\n>> the supermodule.\n> \n> Yes, but then you'd lose history connectivity (I'm assuming you'd only \n> pull in the tree and blob objects from the submodule, and prefix the \n> tree-entrys with whatever directory you're storing the submodul in).\n\nI thought that Uwe meant pulling (getting) _all_ the needed objects from\nsubmodule object repository into supermodule object repository: commits,\ntrees and blobs, full history.\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n\n"},{"id":"298420","messageId":"20061205150217.GA5573@cepheus","threadId":"43065","inReplyTo":"45754AFE.1070207@op5.se","subject":"Re: [RFC] Submodules in GIT","fromName":"Uwe Kleine-Koenig","fromEmail":"zeisberg@informatik.uni-freiburg.de","sentAt":"2006-12-05T15:02:17Z","receivedAt":"2006-12-05T15:02:17Z","isPatch":false,"sender":{"key":"u.kleine-koenig@pengutronix.de","avatar":"https://gravatar.com/avatar/354b5e3ceb2806a2f1e1e382ac29ddbdad18288654da62b61eb13583a857eee7?d=mp&s=160"},"body":"Hella Andreas,\n\nAndreas Ericsson wrote:\n> >>The only problem I'm seeing atm is that the supermodule somehow has to \n> >>mark whatever commits it's using from the submodule inside the submodule \n> >>repo so that they effectively become un-prunable, otherwise the \n> >>supermodule may some day find itself with a history that it can't restore.\n> >One could circumvent that by creating a separate repo for the submodule\n> >at checkout time and pull the needed objects in the supermodule's odb\n> >when commiting the supermodule.  This way prune in the submodule cannot\n> >do any harm, because in it's odb are no objects that are important for\n> >the supermodule.\n> \n> Yes, but then you'd lose history connectivity (I'm assuming you'd only \n> pull in the tree and blob objects from the submodule, and prefix the \n> tree-entrys with whatever directory you're storing the submodul in).\nThat's the reason for me prefering to pull in the complete commit.\n\nI don't understand what you mean with \"prefix the tree-entrys with\nwhatever directory you're storing the submodul in\".\nMaybe one of us doesn't understand tree objects correctly.  AFAICT they\ndon't store the location where they occur, so there is no need to store\na prefix.  E.g. \n\n\tzeisberg@cepheus:/tmp$ mkdir test-repo\n\tzeisberg@cepheus:/tmp$ cd test-repo/\n\tzeisberg@cepheus:/tmp/test-repo$ git-init-db \n\tdefaulting to local storage area\n\tzeisberg@cepheus:/tmp/test-repo$ echo LD_FLAGS=-ltest > Makefile\n\tzeisberg@cepheus:/tmp/test-repo$ git add Makefile\n\tzeisberg@cepheus:/tmp/test-repo$ git commit -m 'test1'\n\tCommitting initial tree 754eadab39642175748bb02155d2959176bcf014\n\tzeisberg@cepheus:/tmp/test-repo$ mkdir subdir\n\tzeisberg@cepheus:/tmp/test-repo$ cp Makefile subdir/\n\tzeisberg@cepheus:/tmp/test-repo$ git add subdir/\n\tzeisberg@cepheus:/tmp/test-repo$ git commit -m 'test2'\n\tzeisberg@cepheus:/tmp/test-repo$ git ls-tree HEAD\n\t100644 blob 610bafd79f92c7e546b104d5b22795df1f099723    Makefile\n\t040000 tree 754eadab39642175748bb02155d2959176bcf014    subdir\n\nSo the tree that only contains the Makefile specifing LD_FLAGS has the\nsha1id 754eadab39642175748bb02155d2959176bcf014 independent of being the\nroot of my project or a subtree.\n\nBut maybe I misunderstood you?\n\nBest regards\nUwe\n\n-- \nUwe Kleine-Koenig\n\nIf a lawyer and an IRS agent were both drowning, and you could only save\n"},{"id":"296720","messageId":"457590AD.4000806@op5.se","threadId":"43065","inReplyTo":"20061205150217.GA5573@cepheus","subject":"Re: [RFC] Submodules in GIT","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-12-05T15:30:53Z","receivedAt":"2006-12-05T15:30:53Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Uwe Kleine-Koenig wrote:\n> Hella Andreas,\n> \n> Andreas Ericsson wrote:\n>>>> The only problem I'm seeing atm is that the supermodule somehow has to \n>>>> mark whatever commits it's using from the submodule inside the submodule \n>>>> repo so that they effectively become un-prunable, otherwise the \n>>>> supermodule may some day find itself with a history that it can't restore.\n>>> One could circumvent that by creating a separate repo for the submodule\n>>> at checkout time and pull the needed objects in the supermodule's odb\n>>> when commiting the supermodule.  This way prune in the submodule cannot\n>>> do any harm, because in it's odb are no objects that are important for\n>>> the supermodule.\n>> Yes, but then you'd lose history connectivity (I'm assuming you'd only \n>> pull in the tree and blob objects from the submodule, and prefix the \n>> tree-entrys with whatever directory you're storing the submodul in).\n> That's the reason for me prefering to pull in the complete commit.\n> \n> I don't understand what you mean with \"prefix the tree-entrys with\n> whatever directory you're storing the submodul in\".\n> Maybe one of us doesn't understand tree objects correctly.  AFAICT they\n> don't store the location where they occur, so there is no need to store\n> a prefix.  E.g. \n> \n> \t100644 blob 610bafd79f92c7e546b104d5b22795df1f099723    Makefile\n> \t040000 tree 754eadab39642175748bb02155d2959176bcf014    subdir\n> \n> So the tree that only contains the Makefile specifing LD_FLAGS has the\n> sha1id 754eadab39642175748bb02155d2959176bcf014 independent of being the\n> root of my project or a subtree.\n> \n> But maybe I misunderstood you?\n> \n\nNopes. I just didn't think of the fact that subtrees are trees and never \nstore any path-info no matter what. So basically the supermodule can \nstore all trees of all submodules for each commit adding a new submodule \nrevision (which is neat, since \"casuals\" never have to bother with \ngetting all the submodules if they want to see all the code used in any \nparticular revision), while we invent the new tree object \"subm\" that \npoints to a commit in the submodule repo. We then teach the tools to \nrecognize when the *real* submodule repo is present and just don't check \nout trees from the supermodule odb that lead us to directories where \nsubmodules reside. Simple and beautiful. Me likes.\n\n*IF* we teach the history viewers about submodules is a different matter \nthough. I'm not sure it would make much sense to have simple text-mode \nbrowsers show the submodule history, although I can imagine qgit and \ngitk wanting to take advantage of their nice side-by-side DAG displaying \ncode to show all the repos in parallell, or link between them in some \npoint-and-click kind of way.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\n"},{"id":"298132","messageId":"20061205160019.GR940MdfPADPa@greensroom.kotnet.org","threadId":"43065","inReplyTo":"20061205090125.GA2428@cepheus","subject":"Re: [RFC] Submodules in GIT","fromName":"Sven Verdoolaege","fromEmail":"skimo@kotnet.org","sentAt":"2006-12-05T16:00:19Z","receivedAt":"2006-12-05T16:00:19Z","isPatch":false,"sender":{"key":"skimo@kotnet.org","avatar":null},"body":"On Tue, Dec 05, 2006 at 10:01:25AM +0100, Uwe Kleine-Koenig wrote:\n> Hello,\n> \n> Andreas Ericsson wrote:\n> > The only problem I'm seeing atm is that the supermodule somehow has to \n> > mark whatever commits it's using from the submodule inside the submodule \n> > repo so that they effectively become un-prunable, otherwise the \n> > supermodule may some day find itself with a history that it can't restore.\n> One could circumvent that by creating a separate repo for the submodule\n> at checkout time and pull the needed objects in the supermodule's odb\n> when commiting the supermodule.  This way prune in the submodule cannot\n> do any harm, because in it's odb are no objects that are important for\n> the supermodule.\n\nI _think_ Linus argued against doing this (for scalability reasons),\nalthough he didn't actually answer my question when I asked him directly.\nIn his proposal you wouldn't need to do this, because the particular\nchecked-out copy of the submodule that is located in a subdirectory\nof a superproject would not be allowed to be pruned and it seems that\nMartin has also implemented it like this.\n\n"},{"id":"295040","messageId":"4575EDA7.2090705@stephan-feder.de","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612042053080.20138@iabervon.org","subject":"Re: [RFC] Submodules in GIT","fromName":"sf","fromEmail":"sf-gmane@stephan-feder.de","sentAt":"2006-12-05T22:07:35Z","receivedAt":"2006-12-05T22:07:35Z","isPatch":false,"sender":{"key":"sf-gmane@stephan-feder.de","avatar":null},"body":"Daniel Barkalow wrote:\n> On Sat, 2 Dec 2006, Linus Torvalds wrote:\n> \n>> So that's where I come from. And maybe I'm totally wrong. I'd like to hear \n>> what people who actually _use_ submodules think.\n> \n> I think you'd rather hear from people who _would_ use submodules; I've \n> worked on a number of projects that would have benefitted from that \n> general functionality, but nobody trusted the implementation enough to \n> actually use it.\n> \n> At my work, we're doing a bunch of stuff with microcontrollers. We've got \n> about a dozen different boards with microcontrollers, and each of them has \n> different firmware. We also have a bunch of code that can go on any of the \n> boards.\n> \n> The way things are organized currently is that each board has its own \n> project, and there's a \"common-micro\" project with the common code. This \n> sort of works, but it means that when you change things in common-micro, \n> you never know what effect this will have on boards other than the one \n> you're actually working on. What I'd like to have is that each project has \n> a \"common-micro\" subdirectory, and changes to each of these can be merged \n> into each other, but that doesn't happen automaticly, and each board's \n> revisions include the common-micro revision they were created with.\n\nOur setup and requirements at work are exactly the same: We have a few\nmain projects that are developed independently and we have one \"helper\"\nproject for code that is general enough to be reused. So work on the\nhelper project is only done while working on one of the main projects.\nWhen we switch to another main project we integrate the changes to the\n\"helper\" project.\n\nThat's the theory, at least.\n\nRegards\n\nStephan\n"},{"id":"295344","messageId":"1165602554.19135.309.camel@cashmere.sps.mot.com","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612011134080.3695@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Jon Loeliger","fromEmail":"jdl@freescale.com","sentAt":"2006-12-08T18:29:14Z","receivedAt":"2006-12-08T18:29:14Z","isPatch":false,"sender":{"key":"jdl@jdl.com","avatar":"https://gravatar.com/avatar/75ce9a10b151acd2c28ec4ab2136dba7b2ff1634530bd04b155981a749d08a64?d=mp&s=160"},"body":"On Fri, 2006-12-01 at 14:13, Linus Torvalds wrote:\n\n> So this is why it's really important that the submodule really is a git \n> repository in its own right, and why committing stuff in the supermodule \n> NEVER affect the submodule itself directly (it might _cause_ you to also \n> do a commit in the submodule indirectly, but the submodule commit MUST be \n> totally independent, and stand on its own).\n\nAn implication of this is that the entire administrative\nresponsibility for having some super-sub module interaction\nlies entirely with the supermodule.\n\nWhy not have a \"glue\" object at the \"stub\"-interface of\nthe supermodule tree that provides policy mappings to\nthe sub-modules.  Perhaps indicating git URL location,\nmappings of branch names between super- and sub- modules,\nspecial commit SHA1s, user policy or config choices at\nthe boundary, and things like that.\n\nIs that the sort of direction we are headed?\n\njdl\n\n"},{"id":"294750","messageId":"20061208184532.GE940MdfPADPa@greensroom.kotnet.org","threadId":"43065","inReplyTo":"1165602554.19135.309.camel@cashmere.sps.mot.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Sven Verdoolaege","fromEmail":"skimo@kotnet.org","sentAt":"2006-12-08T18:45:32Z","receivedAt":"2006-12-08T18:45:32Z","isPatch":false,"sender":{"key":"skimo@kotnet.org","avatar":null},"body":"On Fri, Dec 08, 2006 at 12:29:14PM -0600, Jon Loeliger wrote:\n> Why not have a \"glue\" object at the \"stub\"-interface of\n> the supermodule tree that provides policy mappings to\n> the sub-modules.  Perhaps indicating git URL location,\n> mappings of branch names between super- and sub- modules,\n> special commit SHA1s, user policy or config choices at\n> the boundary, and things like that.\n> \n> Is that the sort of direction we are headed?\n\nNot unless you have something useful in mind that could be put in\nthese glue objects.  URLs and branch names, in particular, should\nnot be stored in the repository itself, but in configuration files,\nsince they will be different for different copies of the repo.\n\n"},{"id":"294713","messageId":"200612091434.15001.rsmckown@yahoo.com","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612021114270.3476@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"R. Steve McKown","fromEmail":"rsmckown@yahoo.com","sentAt":"2006-12-09T21:34:14Z","receivedAt":"2006-12-09T21:34:14Z","isPatch":false,"sender":{"key":"rsmckown@yahoo.com","avatar":null},"body":"On Saturday 02 December 2006 12:41 pm, Linus Torvalds wrote:\n> In other words, I _suspect_ that that is really what module users are all\n> about. They want the ability to specify an arbitrary collection of these\n> atomic snapshots (for releases etc), and just want a way to copy and move\n> those things around, and are less interested in making everything else\n> very seamless (because most people are happy to do the actual\n> _development_ entirely within the submodules, so the \"development\" part\n> is actually not that important for the supermodule, the supermodule is\n> mostly for aggregation and snapshots, and tying different versions of\n> different submodules together).\n>\n> So that's where I come from. And maybe I'm totally wrong. I'd like to hear\n> what people who actually _use_ submodules think.\n\nHere's some thoughts on subprojects from my company's perspective.  I \napologize for the long message.\n\nAbstract: We use submodules heavily in CVS and SVN.  I like what I've read \nfrom Linus about the \"thin veneer\" approach of integrating subprojects.  It \nseems conceptually to provide the support we desire.  For us, it's important \nthat the mandated linkage between a master project and a subproject is \nminimal to maximize our flexibility in building our processes.\n\n\nWe develop and maintain a lot of embedded applications.  Both for higher level \nsystems (ex: 32MB RAM/32MB storage) running the Linux kernel and a customized \nset of libs/app support code and more deeply embedded environments (ex: 8KB \nof RAM and 32KB of storage).  Even though these two cases are very different \nin many repects, the version management issues are the same.\n\n- We (mostly) track everything needed to build historical versions of code \nwith 100% fidelity.  This includes all of the tools used to compile, build, \ntest, deploy, debug, etc. the actual build results themselves.  I initially \nlooked at Vesta several years ago.  I love their conceptual approach to this \nproblem (integrated build system that caches mid-level build results within \nthe repository itself), but it's too unwieldy, very hard to set up (lots of \nup-front effort), and lacks many useful features.\n\n- Most of our \"applications\" are a relatively small amount of app-specific \ncode with references to several/many shared modules.  Shared modules can \ncontain support tools, like build/test/debug/deploy support for a given \nembedded platform, in-house developed shared app code, or shared code \ndeveloped by third parties.\n\n- We use CVS to manage our larger system development projects.  The repo is \nabout 2GB and has several dozen application-code submodules.  We use the \n\"third party sources\" approach to tracking submodules as outlined in Ch.13 of \nthe CVS manual.  Additionally, we manage our \"buildox\" (similar to buildroot \nin concept) in another CVS repo.  All prior interesting versions of the \nbuildroot can be built from source (toolchains, everything), if necessary.  \nApplications contain metadata (a file...) in the repo so the app-level build \nsystem can ensure it is being ran under the correct version of buildbox; \nclunky but serviceable.  CVS is a nightmare because of its poor \nbranch/tagging facilities, and many of the things we *ought* to be doing with \nrevision control we don't because of the complexity.\n\n- We use SVN to manage our deeply embedded system projects.  The repo is about \n250MB in size.  Applications use the svn:externals property to reference \nneeded modules.  We aren't using a buildbox in this environment yet (bad!).  \nSVN's simple branching and svn:externals are a giant leap forward in \ncomparison to CVS's capabilities.\n\n\nBelow are some common use case scenarios that are to varying degrees unweildy \nin CVS and/or SVN.  Many of these involving non-trivial branching and merging \noperations are nearly impractical in CVS, and the lack of merge tracking (to \nsupport repeated safe merging from one branch to another) makes some of these \na bit tricky in SVN too.  Of course neither repo supports \ndisconnected/distributed operation, which would make a number of activities \nthat much simpler as well.\n\n- Round trip module management.  A specific app requires a change to a shared \nmodule, so it makes a local branch to develop the change.  The \"diff\" is \npresented to the maintainer (who may be inhouse).  The next interesting \nmaintainer version of the module gets imported into our repo (if in house, \nit's already there), where the app can reference it.  This merge process may \nleave changes not yet implemented (or never to be implemented) by the module \nmaintainer in the local branch used by the apps.  Other apps are unaffected, \nas they are linking to a prior version in the local branch.\n\n- Pragmatic development.  It's typical that in developing an application, a \ndeveloper will need to simultaneously make changes to one or more submodules.  \nIf more than trivial, he/she should branch the submodules and continually \ntracking the HEAD of those branches in the relevant app.  This is so complex \nand fraught with problems in CVS that it doesn't get done, and developers \nhouse too much change over time in their working directories.  With SVN and \nsvn:externals, the process is workable.  It is nice that an svn:external can \npoint to (the HEAD of) a branch when making changes.\n\n- An application implements a new feature internally (say support for a new \ndigital chipset in the embedded world) which later needs to be \"promoted\" to \na subproject for use by others.  Pretty easy in SVN.  A challenge in CVS; \nit's really not possible to \"convert\" app code into a \"third party source\" \nand retain an historical link.\n\n- Updating build tools.  In concept no different than updating a shared code \nmodule.  In practice, due to the buildbox strategy, it's a bit convoluted.  I \ndon't expect this to get much smoother.  Getting Vesta-like features, where \nintegrated build suport can cache lower-level build results in a version-safe \nmanner (like the binary code built when the cross toolchain was built) would \nbe killer, but that's surely OT for the submodules discussion.\n\nThanks,\n"},{"id":"296872","messageId":"e7bda7770612100347j78854d79x547084972ed14e99@mail.gmail.com","threadId":"43065","inReplyTo":"200612091434.15001.rsmckown@yahoo.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Torgil Svensson","fromEmail":"torgil.svensson@gmail.com","sentAt":"2006-12-10T11:47:53Z","receivedAt":"2006-12-10T11:47:53Z","isPatch":false,"sender":{"key":"torgil.svensson@gmail.com","avatar":null},"body":"What if we use linus \"module\" file concept and allow the link objects\nto track subtrees? An object may look like this:\n\ncommit: <SHA1>\nlink: <SHA1> /path/to/remote/tree/or/blob\n\n\nTracking upstream library:\n--------------------------\nclone as usual\n\n\nInhouse libraries/applications:\n-------------------------------\nTo satisfy versioning of build-dependencies - make links of type\n\"external/lib1_header.h\" -> \"<commit>/headers/lib1_header.h\" (blob)\n\"external/lib1_interface\" -> \"<commit>/api\" (tree)\n\nIf git supports \"sparse fetching\" of subtrees we can follow the\nhistory in the submodule only concerning the files we want without\nfetching the whole subtree. \"modules\" file could specify something\nlike \"always clone on fetch\"\n\n\nBuild environment\n------------------\nFirst make links to all tools, applications, etc ...\n\"buildtools/random_app1\" -> \"<commit>/\"\n\"buildtools/random_app2\" -> \"<commit>/\"\n\"sub_build_projects/user_interface\" -> \"<commit>/\"\n\"sub_build_projects/kernel\" -> \"<commit>/\"\n\"apps/special_app1\" -> \"<commit>/\"\n\"libs/special_lib1\" ->\n\"<commit-from-another-build-project>/special/lib/binary/path\"\n\nHere we can have a build system that for example creates a \"i386\"\nfolder and the repo itself\n\n\nDocumentation release\n----------------------\n\"Lib1/\" -> \"<lib1 commit>/docs\"\n\"Lib2/\" -> \"<lib2 commit>/docs\"\n\"App1/\" -> \"<app1 commit>/docs\"\n\n\nSpecial customer release for a specific HW platform\n---------------------------------------------------\n\"Lib1/lib1.h\" -> \"<lib1-commit>/headers/lib1.h\"\n\"Lib1/lib1.so\" -> \"<build-environment-commit>/i386/Lib1/lib1.so\"\n\"Lib1/docs\" -> \"<lib1-commit>/docs\"\n\"App1_binary\" -> \"<build-environment-commit>/i386/App1/App1_binary\"\n\"docs\" -> \"<app1-commit>/docs\"\n\ncommit&tag&bag this and send to customer. If the customer says\nsomething is broken, we can make an SHA1 of the customers tree and\nimmediately see if there's changes not belonging to us.\n\n\nNow this can be broken in so many ways that I can't even count, so I\nappreciate some feedback to correct my head.\n\n\nOn 12/9/06, R. Steve McKown <rsmckown@yahoo.com> wrote:\n> On Saturday 02 December 2006 12:41 pm, Linus Torvalds wrote:\n> > In other words, I _suspect_ that that is really what module users are all\n> > about. They want the ability to specify an arbitrary collection of these\n> > atomic snapshots (for releases etc), and just want a way to copy and move\n> > those things around, and are less interested in making everything else\n> > very seamless (because most people are happy to do the actual\n> > _development_ entirely within the submodules, so the \"development\" part\n> > is actually not that important for the supermodule, the supermodule is\n> > mostly for aggregation and snapshots, and tying different versions of\n> > different submodules together).\n> >\n> > So that's where I come from. And maybe I'm totally wrong. I'd like to hear\n> > what people who actually _use_ submodules think.\n>\n> Here's some thoughts on subprojects from my company's perspective.  I\n> apologize for the long message.\n>\n> Abstract: We use submodules heavily in CVS and SVN.  I like what I've read\n> from Linus about the \"thin veneer\" approach of integrating subprojects.  It\n> seems conceptually to provide the support we desire.  For us, it's important\n> that the mandated linkage between a master project and a subproject is\n> minimal to maximize our flexibility in building our processes.\n>\n>\n> We develop and maintain a lot of embedded applications.  Both for higher level\n> systems (ex: 32MB RAM/32MB storage) running the Linux kernel and a customized\n> set of libs/app support code and more deeply embedded environments (ex: 8KB\n> of RAM and 32KB of storage).  Even though these two cases are very different\n> in many repects, the version management issues are the same.\n>\n> - We (mostly) track everything needed to build historical versions of code\n> with 100% fidelity.  This includes all of the tools used to compile, build,\n> test, deploy, debug, etc. the actual build results themselves.  I initially\n> looked at Vesta several years ago.  I love their conceptual approach to this\n> problem (integrated build system that caches mid-level build results within\n> the repository itself), but it's too unwieldy, very hard to set up (lots of\n> up-front effort), and lacks many useful features.\n>\n> - Most of our \"applications\" are a relatively small amount of app-specific\n> code with references to several/many shared modules.  Shared modules can\n> contain support tools, like build/test/debug/deploy support for a given\n> embedded platform, in-house developed shared app code, or shared code\n> developed by third parties.\n>\n> - We use CVS to manage our larger system development projects.  The repo is\n> about 2GB and has several dozen application-code submodules.  We use the\n> \"third party sources\" approach to tracking submodules as outlined in Ch.13 of\n> the CVS manual.  Additionally, we manage our \"buildox\" (similar to buildroot\n> in concept) in another CVS repo.  All prior interesting versions of the\n> buildroot can be built from source (toolchains, everything), if necessary.\n> Applications contain metadata (a file...) in the repo so the app-level build\n> system can ensure it is being ran under the correct version of buildbox;\n> clunky but serviceable.  CVS is a nightmare because of its poor\n> branch/tagging facilities, and many of the things we *ought* to be doing with\n> revision control we don't because of the complexity.\n>\n> - We use SVN to manage our deeply embedded system projects.  The repo is about\n> 250MB in size.  Applications use the svn:externals property to reference\n> needed modules.  We aren't using a buildbox in this environment yet (bad!).\n> SVN's simple branching and svn:externals are a giant leap forward in\n> comparison to CVS's capabilities.\n>\n>\n> Below are some common use case scenarios that are to varying degrees unweildy\n> in CVS and/or SVN.  Many of these involving non-trivial branching and merging\n> operations are nearly impractical in CVS, and the lack of merge tracking (to\n> support repeated safe merging from one branch to another) makes some of these\n> a bit tricky in SVN too.  Of course neither repo supports\n> disconnected/distributed operation, which would make a number of activities\n> that much simpler as well.\n>\n> - Round trip module management.  A specific app requires a change to a shared\n> module, so it makes a local branch to develop the change.  The \"diff\" is\n> presented to the maintainer (who may be inhouse).  The next interesting\n> maintainer version of the module gets imported into our repo (if in house,\n> it's already there), where the app can reference it.  This merge process may\n> leave changes not yet implemented (or never to be implemented) by the module\n> maintainer in the local branch used by the apps.  Other apps are unaffected,\n> as they are linking to a prior version in the local branch.\n>\n> - Pragmatic development.  It's typical that in developing an application, a\n> developer will need to simultaneously make changes to one or more submodules.\n> If more than trivial, he/she should branch the submodules and continually\n> tracking the HEAD of those branches in the relevant app.  This is so complex\n> and fraught with problems in CVS that it doesn't get done, and developers\n> house too much change over time in their working directories.  With SVN and\n> svn:externals, the process is workable.  It is nice that an svn:external can\n> point to (the HEAD of) a branch when making changes.\n>\n> - An application implements a new feature internally (say support for a new\n> digital chipset in the embedded world) which later needs to be \"promoted\" to\n> a subproject for use by others.  Pretty easy in SVN.  A challenge in CVS;\n> it's really not possible to \"convert\" app code into a \"third party source\"\n> and retain an historical link.\n>\n> - Updating build tools.  In concept no different than updating a shared code\n> module.  In practice, due to the buildbox strategy, it's a bit convoluted.  I\n> don't expect this to get much smoother.  Getting Vesta-like features, where\n> integrated build suport can cache lower-level build results in a version-safe\n> manner (like the binary code built when the cross toolchain was built) would\n> be killer, but that's surely OT for the submodules discussion.\n>\n> Thanks,\n> Steve\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"295149","messageId":"457E692E.7060708@op5.se","threadId":"43065","inReplyTo":"1165602554.19135.309.camel@cashmere.sps.mot.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-12-12T08:32:46Z","receivedAt":"2006-12-12T08:32:46Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Jon Loeliger wrote:\n> On Fri, 2006-12-01 at 14:13, Linus Torvalds wrote:\n> \n>> So this is why it's really important that the submodule really is a git \n>> repository in its own right, and why committing stuff in the supermodule \n>> NEVER affect the submodule itself directly (it might _cause_ you to also \n>> do a commit in the submodule indirectly, but the submodule commit MUST be \n>> totally independent, and stand on its own).\n> \n> An implication of this is that the entire administrative\n> responsibility for having some super-sub module interaction\n> lies entirely with the supermodule.\n> \n\nThat's a good thing. I wouldn't want the openssl maintainers to have to \nbother with every project that uses their code, and I'm fairly certain \nthey feel the same.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\n"},{"id":"297386","messageId":"e7bda7770612141327r11368e4dtabe8077e96545040@mail.gmail.com","threadId":"43065","inReplyTo":"e7bda7770612100347j78854d79x547084972ed14e99@mail.gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Torgil Svensson","fromEmail":"torgil.svensson@gmail.com","sentAt":"2006-12-14T21:27:26Z","receivedAt":"2006-12-14T21:27:26Z","isPatch":false,"sender":{"key":"torgil.svensson@gmail.com","avatar":null},"body":"On 12/10/06, Torgil Svensson <torgil.svensson@gmail.com> wrote:\n> What if we use linus \"module\" file concept and allow the link objects\n> to track subtrees? An object may look like this:\n>\n> commit: <SHA1>\n> link: <SHA1> /path/to/remote/tree/or/blob\n\n> Special customer release for a specific HW platform\n> ---------------------------------------------------\n> \"Lib1/lib1.h\" -> \"<lib1-commit>/headers/lib1.h\"\n> \"Lib1/lib1.so\" -> \"<build-environment-commit>/i386/Lib1/lib1.so\"\n> \"App1_binary\" -> \"<build-environment-commit>/i386/App1/App1_binary\"\n\nThis example is somewhat complex since the build for lib1.so and the\nheader-file might not has gone through the same commit on the lib1\nsubproject.  Consider this example:\n\n\nlib1 - library project (source tracking)\n------------------------------------------\nBlob: /src/lib1.h\n\n\napp1 - application project (source tracking)\n-----------------------------------------\nLink: /headers/lib1.h -> <lib1-commit1>/src/lib1.h\n\n\nbuild1 - Build project (binary build tracking)\n------------------------------------\nLink: /src/lib1 -> <lib1-commit2>/\nLink: /src/app1 -> <app1-commit>/\nBlob: /i386/lib1/lib1.so\nBlob: /i386/app1/app1\n\n\nRelease Project (file compilation tracking)\n-----------------------------------\nLink: /headers/lib1.h -> <lib1-commit3>/src/lib1.h\nLink: /bin/lib1.so -> <build1-commit>/i386/lib1/lib1.so\nLink: /bin/app1 -> <build1-commit>/i386/app1/app1\n\n\n<lib1-commit1>, <lib1-commit2> and <lib1-commit3> should be the same,\ndictated by the app1 project. Can we enforce this in the modules file\nor should the different supermodules fix this somehow using\nscripts/hooks?\n\nHow do the super-projects in this case get access to the blobs pointed\nby the links - transparent or explicit in the build-process?\n\n"},{"id":"295065","messageId":"200612150007.44331.Josef.Weidendorfer@gmx.de","threadId":"43065","inReplyTo":"e7bda7770612141327r11368e4dtabe8077e96545040@mail.gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-14T23:07:44Z","receivedAt":"2006-12-14T23:07:44Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Thursday 14 December 2006 22:27, Torgil Svensson wrote:\n> This example is somewhat complex since the build for lib1.so and the\n> header-file might not has gone through the same commit on the lib1\n> subproject.  Consider this example:\n\nIf you want to track build results for some source,\nwhy would you ever want these builds go out of sync with the source?\nAs the built files depend on the source (and other things), the\nsource should be a submodule of the build project.\n\nHmm... I think I see a problem / wish for submodules here.\n\nWith the current submodule proposal, we force submodules to be\nsubdirectories inside of a supermodule.\n\nYour example has the folling submodule dependence\n(\"X ==> Y\" means Y being a submodule of X):\n\n  App     ==>   Lib\n   ^             ^\n   |             |\n AppBuild ==> LibBuild\n\nIf we force submodules to be subdirectories of supermodules,\nLib needlessly will have to appear two times in a checkout of\nAppBuild.\n\nHowever, there is nothing wrong with it. Yet, you perhaps want\nthe 2 Lib submodules not to go out of sync. This easily\ncan be done with symlinking the Lib checkouts. As they are submodules,\neverything should work fine.\n\nPerhaps an option you want to have is to force a checkout\nof AppBuild to make these symlinking itself when it detects\nidentical submodules links.\n\nHmmm... the only problem with a symlink is that it can go wrong\nwhen moved. Unfortunately, I do not have a good solution for\nthis. We can not make UNIX symlinks smart in any way.\nHardlinking directories would be a solution, but that is not\npossible.\n\nAnother thing:\nWith normal \"$buildroot != $srcroot\" environments, the source\ncan not be a subdirectory of the build directory.\nYet, we want to specify submodule/supermodule relation.\nThis is difficult to do with a submodule object, as it needs\nto appear in trees in the supermodule.\n\nActually, the best workaround for this is to make Lib a direct\nsubmodule of AppBuild, and specify the relationship of\nLibBuild ==> Lib only in AppBuild.\n\nBTW, build project commits probably should not depend on any\nhistory of other build commits.\nSo you actually want all build commits to be root commits, and\nhave a tag name which could include the source commit id from\nwhich the build was done. This gives some loose coupling.\n\n\n> Link: /headers/lib1.h -> <lib1-commit3>/src/lib1.h\n> Link: /bin/lib1.so -> <build1-commit>/i386/lib1/lib1.so\n> Link: /bin/app1 -> <build1-commit>/i386/app1/app1\n> \n> \n> <lib1-commit1>, <lib1-commit2> and <lib1-commit3> should be the same,\n> dictated by the app1 project.\n\nI do not see any problem here. Symlinks are stored in the git repository.\nAs the AppBuild commit depends on App and LibBuild submodule commits, the\nsymlinks always should be correct.\n\n> Can we enforce this in the modules file \n> or should the different supermodules fix this somehow using\n> scripts/hooks?\n\nI do not see any need for an hook. But of course, a checkout hook should\nbe able to generate files/links. However, IMHO this should be not\ndone with hooks but with Makefile targets.\n \n> How do the super-projects in this case get access to the blobs pointed\n> by the links - transparent or explicit in the build-process?\n\nSubmodules should automatically be checked out when checking out the\nsupermodule. So the blobs should already be there.\nOr do I miss something?\n\n"},{"id":"295653","messageId":"e7bda7770612150943j71a7362bmb509cea3b7756003@mail.gmail.com","threadId":"43065","inReplyTo":"200612150007.44331.Josef.Weidendorfer@gmx.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Torgil Svensson","fromEmail":"torgil.svensson@gmail.com","sentAt":"2006-12-15T17:43:59Z","receivedAt":"2006-12-15T17:43:59Z","isPatch":false,"sender":{"key":"torgil.svensson@gmail.com","avatar":null},"body":"On 12/15/06, Josef Weidendorfer <Josef.Weidendorfer@gmx.de> wrote:\n> If you want to track build results for some source,\n> why would you ever want these builds go out of sync with the source?\n\nI don't, bad wording by me. That was the problem I wanted to address.\n\n\n> Your example has the folling submodule dependence\n> (\"X ==> Y\" means Y being a submodule of X):\n>\n>   App     ==>   Lib\n>    ^             ^\n>    |             |\n>  AppBuild ==> LibBuild\n\nIn my example \"AppBuild\" and \"LibBuild\" were the same project but this\nscenario is relevant as well.\n\n\n> If we force submodules to be subdirectories of supermodules,\n> Lib needlessly will have to appear two times in a checkout of\n> AppBuild.\n>\n> However, there is nothing wrong with it. Yet, you perhaps want\n> the 2 Lib submodules not to go out of sync. This easily\n> can be done with symlinking the Lib checkouts. As they are submodules,\n> everything should work fine.\n\nThis is interesting. In my notation:\n\n/path/to/link/name -> <commit>/path/to/subtree\n\nmeans that there is a link named \"name\" in the tree object for\n\"path/to/link\". The link points to a \"link object\" specifying a\nsubtree or blob of the tree that is pointed to in a submodule commit.\nThis is not currently implemented but has at least the following\nadvantages:\n\n1. You can access files in a submodule without fetching the whole\nsubmodule (which may be very large). (App1 is only interested in\nlib1.h, the rest is toally irrelevant)\n2. Superproject can access referenced (linked) files in it's own\nfolder-structure without being forced a structure by the subproject.\n\nIf you do a symlink instead, doesn't you loose versioning information?\nWhat happens with the symlinks if someone clones the superproject?\n\n\n>\n> Perhaps an option you want to have is to force a checkout\n> of AppBuild to make these symlinking itself when it detects\n> identical submodules links.\n>\n> Hmmm... the only problem with a symlink is that it can go wrong\n> when moved. Unfortunately, I do not have a good solution for\n> this. We can not make UNIX symlinks smart in any way.\n> Hardlinking directories would be a solution, but that is not\n> possible.\n>\n\nWouldn't specifying the submodule path in the link object fit in well\nhere? Then each \"link object\" can represent a checked out tree from\nthe subproject in the superproject directory-structure.\n\n\n> Another thing:\n> With normal \"$buildroot != $srcroot\" environments, the source\n> can not be a subdirectory of the build directory.\n\nThis is true for symlinks and would also be corrected if we have a\n(sparse) submodule checkout there in it's place.\n\n\n> BTW, build project commits probably should not depend on any\n> history of other build commits.\n\nWhy? Can you give an example here.\n\n\n> > Link: /headers/lib1.h -> <lib1-commit3>/src/lib1.h\n> > Link: /bin/lib1.so -> <build1-commit>/i386/lib1/lib1.so\n> > Link: /bin/app1 -> <build1-commit>/i386/app1/app1\n> >\n> >\n> > <lib1-commit1>, <lib1-commit2> and <lib1-commit3> should be the same,\n> > dictated by the app1 project.\n>\n> I do not see any problem here. Symlinks are stored in the git repository.\n> As the AppBuild commit depends on App and LibBuild submodule commits, the\n> symlinks always should be correct.\n\nThe main reason for these \"links\" are for versioning purposes: the\nuniqe SHA1 of the \"link\" representing a tree/blob in a version of the\nsubmodule should be \"included\" in the supermodules commit. Symlinks\nwon't give that at all.\n\n\n> > How do the super-projects in this case get access to the blobs pointed\n> > by the links - transparent or explicit in the build-process?\n>\n> Submodules should automatically be checked out when checking out the\n> supermodule. So the blobs should already be there.\n> Or do I miss something?\n\nProbably not as that was a piece of the puzzle that I was missing.\n\n\n"},{"id":"295591","messageId":"200612152242.50472.Josef.Weidendorfer@gmx.de","threadId":"43065","inReplyTo":"e7bda7770612150943j71a7362bmb509cea3b7756003@mail.gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-15T21:42:49Z","receivedAt":"2006-12-15T21:42:49Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Friday 15 December 2006 18:43, Torgil Svensson wrote:\n> On 12/15/06, Josef Weidendorfer <Josef.Weidendorfer@gmx.de> wrote:\n> > However, there is nothing wrong with it. Yet, you perhaps want\n> > the 2 Lib submodules not to go out of sync. This easily\n> > can be done with symlinking the Lib checkouts. As they are submodules,\n> > everything should work fine.\n> \n> This is interesting. In my notation:\n> \n> /path/to/link/name -> <commit>/path/to/subtree\n> \n> means that there is a link named \"name\" in the tree object for\n> \"path/to/link\". The link points to a \"link object\" specifying a\n> subtree or blob of the tree that is pointed to in a submodule commit.\n\nAh, now I understand. I somehow missed this notation.\n\n> This is not currently implemented but has at least the following\n> advantages:\n> \n> 1. You can access files in a submodule without fetching the whole\n> submodule (which may be very large). (App1 is only interested in\n> lib1.h, the rest is toally irrelevant)\n> 2. Superproject can access referenced (linked) files in it's own\n> folder-structure without being forced a structure by the subproject.\n\nThat all sounds fine, but how do you create such symlinks in practice?\nDo you want to introduce special porcelain commands to create them?\nEspecially, what is the SCM user supposed to do to change the link\ntarget, ie. from\n <commit>/path/to/subtree\nto \n <commit>/path2/to2/subtree2\n?\nShould this do a re-checkout at the other point?\n\nBy linking a file from a submodule, such a link seems to force that\nthis file has to be at a fixed position in the submodule. Otherwise,\nsome magic has to happen when the file is moved in the submodule,\npossibly leading to a dangling link, eg. if the whole subdirectory\nspecified in the link is removed.\n\nIMHO this is getting way to complex.\nMuch simpler is to include the full submodule at some path in\nthe supermodule, and create normal symlinks from the supermodule\ninto the submodule.\n\nIf you only want to check out part of a submodule, this should be\ndone with path-limiting checkouts, which should be a feature totally\nindependent from submodules.\n\nAnd if you want to limit the number of objects transferred in cloning\nof a subproject, it is better to further split this subproject into\nmultiple subprojects itself.\n \n> If you do a symlink instead, doesn't you loose versioning information?\n\nOf course, you need the submodule fully checked out somewhere in the\nsupermodule, and the link goes into the submodule directory. The\nversioning is given by the supermodule/submodule link.\n\n> What happens with the symlinks if someone clones the superproject?\n\nAs already said: the link has to go into a submodule directory, which\nwill be checked out automatically with the clone of the supermodule.\n\n> > Perhaps an option you want to have is to force a checkout\n> > of AppBuild to make these symlinking itself when it detects\n> > identical submodules links.\n> >\n> > Hmmm... the only problem with a symlink is that it can go wrong\n> > when moved. Unfortunately, I do not have a good solution for\n> > this. We can not make UNIX symlinks smart in any way.\n> > Hardlinking directories would be a solution, but that is not\n> > possible.\n> >\n> \n> Wouldn't specifying the submodule path in the link object fit in well\n> here? Then each \"link object\" can represent a checked out tree from\n> the subproject in the superproject directory-structure.\n\nThe problem is not the representation in the git repository, but the\nchecked out module/submodule, where you need to use normal UNIX file semantics.\nTo move submodules around, the user should be able to just use\nthe normal UNIX \"mv\" commands, and git should be able to detect move\nactions after the fact.\nThe simple thing here is that currently, git does not have this problem\nas it tracks content, and does not even try to detect any moves at\ncommit time. This is different with submodules, as there, you want to\nbe able to track moves of any submodule root directories.\n\nThis now becomes a problem if you use symlinks to \"unify\" multiple checkouts\nof the same submodule at multiple places in the supermodule, and move\nthe symlink around, as it easily can get dangling this way. Thus, you would\nnot have a way to see what submodule this link was talking about.\n\nAnd for this thing, I do not see how your link object could help.\n\nSo it is better to use a simple submodule concept, and for this corner\ncases, we perhaps could expect the user to fix e.g. a dangling symlink\nto a previous submodule checkout himself, using a meaningful error message.\n\n> > BTW, build project commits probably should not depend on any\n> > history of other build commits.\n> \n> Why? Can you give an example here.\n\nIf you have a source commit chain A => B => C => D, you want\nto make any build commits totally independent: you first only\nare interested in a build commit for source versions A and D,\nand later find out that a build commit for B and C would be nice,\ntoo. If you force build commits into some history order, this\norder now would be A => D => B => C, which makes no sense.\n\nBuild commit independence can easily be achieved by making every commit\nparentless, without further history. You still have the link\nto the source version via the submodule link in the tree.\nBut to not loose any such build commits, they have to appear\nas tags or refs (unless integrated in another superproject\nbuild commit).\n\n\n> > > Link: /headers/lib1.h -> <lib1-commit3>/src/lib1.h\n> > > Link: /bin/lib1.so -> <build1-commit>/i386/lib1/lib1.so\n> > > Link: /bin/app1 -> <build1-commit>/i386/app1/app1\n> > >\n> > >\n> > > <lib1-commit1>, <lib1-commit2> and <lib1-commit3> should be the same,\n> > > dictated by the app1 project.\n> >\n> > I do not see any problem here. Symlinks are stored in the git repository.\n> > As the AppBuild commit depends on App and LibBuild submodule commits, the\n> > symlinks always should be correct.\n> \n> The main reason for these \"links\" are for versioning purposes: the\n> uniqe SHA1 of the \"link\" representing a tree/blob in a version of the\n> submodule should be \"included\" in the supermodules commit. Symlinks\n> won't give that at all.\n\nThe version coupling will be there if the whole submodule is available\nat some path in the supermodule checkout, as said above.\n\n"},{"id":"293898","messageId":"e7bda7770612151543o39c9d233q91ea643a134196d3@mail.gmail.com","threadId":"43065","inReplyTo":"200612152242.50472.Josef.Weidendorfer@gmx.de","subject":"Re: [RFC] Submodules in GIT","fromName":"Torgil Svensson","fromEmail":"torgil.svensson@gmail.com","sentAt":"2006-12-15T23:43:56Z","receivedAt":"2006-12-15T23:43:56Z","isPatch":false,"sender":{"key":"torgil.svensson@gmail.com","avatar":null},"body":"On 12/15/06, Josef Weidendorfer <Josef.Weidendorfer@gmx.de> wrote:\n> That all sounds fine, but how do you create such symlinks in practice?\n\nI'm very open to suggestions here, but the concept growing in my head\nis based around Linus 'module'-file and keep things simple. A git\nconfiguration file that specifies:\n* link name for reference\n* local path to link\n* submodule source\n* submodule path to tree/blob\n* submodule commit / HEAD / branch\n* options (depth-limit , ...)\n\nI'm reconsidering having the path-name in the link, it should be\nsufficient to have two SHA1's, one for the commit and one for the\ntree/blob. Super-module should have the tree/blob in it's database so\nthat the link part only is there for version information and reference\n(checking dirty state or history on the submodule). This way it easy\nto clone the super-project and use it without having to map up all\nsub-project sources. Sub-project sources is not important for version\ninformation and could always be specified in the project in a\nREADME-type of file.\n\n\n> Especially, what is the SCM user supposed to do to change the link\n> target, ie. from\n>  <commit>/path/to/subtree\n> to\n>  <commit>/path2/to2/subtree2\n> ?\n> Should this do a re-checkout at the other point?\n\nThat would be a change in the modules file, maybe through a command\nthat also fixes the link. The link will have to be updated in the\nindex and commited as normal.\n\n\n> By linking a file from a submodule, such a link seems to force that\n> this file has to be at a fixed position in the submodule. Otherwise,\n> some magic has to happen when the file is moved in the submodule,\n> possibly leading to a dangling link, eg. if the whole subdirectory\n> specified in the link is removed.\n\nSince we have the SHA1 (this is what we're using) and tree/blob\ninformation in the super-modules database the change itself is not a\nproblem. The problem is to track renames/moves and your remove case in\nthe submodule. The tool that tracks the submodule should probably\nwarn/exit here and we would fix up the modules file manually.\n\n\n> IMHO this is getting way to complex.\n\nOne of complex situation here as I see it is the ability to handle to\ntrack/checkout only a subset (tree/blob) of the submodule. This is\nalso quite an important feature - in my example it means the\ndifference of tracking one header file versus the whole source.\n\n\n> If you only want to check out part of a submodule, this should be\n> done with path-limiting checkouts, which should be a feature totally\n> independent from submodules.\n\nIf we can do path-limiting checkouts on a repo (module) we also can do\nit on a sub-module since they are exactly the same. This is a very\npowerful feature and it'd be a huge waste if it wasn't allowed for a\nsuper-module to do on submodules.\n\n\n> And if you want to limit the number of objects transferred in cloning\n> of a subproject, it is better to further split this subproject into\n> multiple subprojects itself.\n\nWhat if we have no control of the submodule?  This can be tracked from\nupstream, sourceforge, another company, etc. The submodule will often\nlive their own life and could be X, kernel, gcc, cairo, whatever, ...\n\n\n> The problem is not the representation in the git repository, but the\n> checked out module/submodule, where you need to use normal UNIX file semantics.\n> To move submodules around, the user should be able to just use\n> the normal UNIX \"mv\" commands, and git should be able to detect move\n> actions after the fact.\n\nIf we disregard the commit info, the link will act exactly as a normal\ntree/blob. Git can know we're moving a subproject by watching the\nmodule file. The main problem is to keep modules file up-to-date with\nreality. We could enforce module file validity by disallowing such\noperations and let the user do a \"force\" operation which also alters\nthe modules file.\n\n\n> This now becomes a problem if you use symlinks to \"unify\" multiple checkouts\n> of the same submodule at multiple places in the supermodule, and move\n> the symlink around, as it easily can get dangling this way. Thus, you would\n> not have a way to see what submodule this link was talking about.\n\nThe symlink only exists in the modules file. We only have the SHA1's\nat the tree-level and there we have everything underneath the\ntree/blob SHA1 in our database. We will only know if the modules\nsymlink file is dangling next time we fetch from the submodule - here\nwe would notify the user but our database is still consistent.\n\n\n> If you have a source commit chain A => B => C => D, you want\n> to make any build commits totally independent: you first only\n> are interested in a build commit for source versions A and D,\n> and later find out that a build commit for B and C would be nice,\n> too. If you force build commits into some history order, this\n> order now would be A => D => B => C, which makes no sense.\n\nIt makes no sense because the user seem to have act irrationally. The\ncommit-chain is completely valid as it has tracked the correct history\nof the builds. I can't see any problems here, the build-project is\nindependent of the source-project with it's own history. We can hope\nthe user has given good explanations for his/her actions in the commit\nmessages though.\n\n\n"},{"id":"294448","messageId":"e7bda7770612151713k418434e6gd8d565e49a766477@mail.gmail.com","threadId":"43065","inReplyTo":"e7bda7770612151543o39c9d233q91ea643a134196d3@mail.gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Torgil Svensson","fromEmail":"torgil.svensson@gmail.com","sentAt":"2006-12-16T01:13:35Z","receivedAt":"2006-12-16T01:13:35Z","isPatch":false,"sender":{"key":"torgil.svensson@gmail.com","avatar":null},"body":"On 12/16/06, Torgil Svensson <torgil.svensson@gmail.com> wrote:\n\n> I'm very open to suggestions here, but the concept growing in my head\n> is based around Linus 'module'-file and keep things simple. A git\n> configuration file that specifies:\n> * link name for reference\n> * local path to link\n> * submodule source\n> * submodule path to tree/blob\n> * submodule commit / HEAD / branch\n> * options (depth-limit , ...)\n>\n> I'm reconsidering having the path-name in the link, it should be\n> sufficient to have two SHA1's, one for the commit and one for the\n> tree/blob. Super-module should have the tree/blob in it's database so\n> that the link part only is there for version information and reference\n> (checking dirty state or history on the submodule). This way it easy\n> to clone the super-project and use it without having to map up all\n> sub-project sources. Sub-project sources is not important for version\n> information and could always be specified in the project in a\n> README-type of file.\n>\n\nSee it as the link only is there for the version handling between\ndifferent modules and it's the module file that give an UI to the the\nlink (which project, branch, ....). Many users will not care whats\nbehind those links, but if they want to edit the link they have to\ncreate the modules file or fetch it somewhere - it may even be\nprovided and version controlled in the project itself.\n\nexample tree object:\n\n100644 blob <sha1 of blob>    README\n100644 blob <sha1 of blob>    REPORTING-BUGS\n100644 link <sha1 of blob>     <sha1 of commit>\n040000 tree <sha1 of tree>    arch\n040000 tree <sha1 of tree>    block\n040000 link <sha1 of tree>     <sha1 of commit>\n\nNote that the links functions exactly as the blobs and trees in the\ndatabase. The difference is that they origin from _a_subproject_ (we\ndon't care which in this stage) with the specified commit SHA1. If the\nlink isn't represented in the modules file, it's no big deal, it can\nbe added later on if needed.\n\nIf the blame-game begins or if we want to check what we're using on a\nsubmodule level we can always pinpoint the exact file/tree content and\n"},{"id":"296301","messageId":"e7bda7770612151720w2e65fe83s9942e1ec1f7092a2@mail.gmail.com","threadId":"43065","inReplyTo":"e7bda7770612151713k418434e6gd8d565e49a766477@mail.gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Torgil Svensson","fromEmail":"torgil.svensson@gmail.com","sentAt":"2006-12-16T01:20:00Z","receivedAt":"2006-12-16T01:20:00Z","isPatch":false,"sender":{"key":"torgil.svensson@gmail.com","avatar":null},"body":"On 12/16/06, Torgil Svensson <torgil.svensson@gmail.com> wrote:\n>\n> example tree object:\n>\n> 100644 blob <sha1 of blob>    README\n> 100644 blob <sha1 of blob>    REPORTING-BUGS\n> 100644 link <sha1 of blob>     <sha1 of commit>\n> 040000 tree <sha1 of tree>    arch\n> 040000 tree <sha1 of tree>    block\n> 040000 link <sha1 of tree>     <sha1 of commit>\n>\n\nSorry, I was sloppy and forgot the names:\n\n100644 blob <sha1 of blob>    README\n100644 blob <sha1 of blob>    REPORTING-BUGS\n100644 link <sha1 of blob>     <sha1 of commit>   AUTHORS\n040000 tree <sha1 of tree>    arch\n040000 tree <sha1 of tree>    block\n040000 link <sha1 of tree>     <sha1 of commit>   misc\n\nNow it doesn't looks like trees/blobs anymore so maybe a link object is handy:\n\n100644 blob <sha1 of blob>    README\n100644 blob <sha1 of blob>    REPORTING-BUGS\n100644 link <sha1 of link>      AUTHORS\n040000 tree <sha1 of tree>    arch\n040000 tree <sha1 of tree>    block\n040000 link <sha1 of link>     misc\n\nlink-object:\n<sha1 of commit>\n"},{"id":"294966","messageId":"elviac$63t$1@sea.gmane.org","threadId":"43065","inReplyTo":"e7bda7770612151720w2e65fe83s9942e1ec1f7092a2@mail.gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-16T01:34:29Z","receivedAt":"2006-12-16T01:34:29Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Torgil Svensson wrote:\n\n> On 12/16/06, Torgil Svensson <torgil.svensson@gmail.com> wrote:\n>>\n>> example tree object:\n>>\n>> 100644 blob <sha1 of blob>    README\n>> 100644 blob <sha1 of blob>    REPORTING-BUGS\n>> 100644 link <sha1 of blob>     <sha1 of commit>\n>> 040000 tree <sha1 of tree>    arch\n>> 040000 tree <sha1 of tree>    block\n>> 040000 link <sha1 of tree>     <sha1 of commit>\n>>\n> \n> Sorry, I was sloppy and forgot the names:\n> \n> 100644 blob <sha1 of blob>    README\n> 100644 blob <sha1 of blob>    REPORTING-BUGS\n> 100644 link <sha1 of blob>     <sha1 of commit>   AUTHORS\n> 040000 tree <sha1 of tree>    arch\n> 040000 tree <sha1 of tree>    block\n> 040000 link <sha1 of tree>     <sha1 of commit>   misc\n> \n> Now it doesn't looks like trees/blobs anymore so maybe a link object\n> is handy: \n> \n> 100644 blob <sha1 of blob>    README\n> 100644 blob <sha1 of blob>    REPORTING-BUGS\n> 100644 link <sha1 of link>      AUTHORS\n> 040000 tree <sha1 of tree>    arch\n> 040000 tree <sha1 of tree>    block\n> 040000 link <sha1 of link>     misc\n> \n> link-object:\n> <sha1 of commit>\n> <sha1 of tree/blob>\n\nWhat do you need <sha1 of tree/blob> for in link-object? Wouldn't you\nuse usually the sha1 of top tree of a commit, which is uniquely defined\nby commit object, so you need only <ahs1 of commit>?\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n\n"},{"id":"297767","messageId":"Pine.LNX.4.64.0612151747470.3849@woody.osdl.org","threadId":"43065","inReplyTo":"e7bda7770612151713k418434e6gd8d565e49a766477@mail.gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-16T01:49:33Z","receivedAt":"2006-12-16T01:49:33Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 16 Dec 2006, Torgil Svensson wrote:\n> \n> 100644 blob <sha1 of blob>    README\n> 100644 blob <sha1 of blob>    REPORTING-BUGS\n> 100644 link <sha1 of blob>     <sha1 of commit>\n> 040000 tree <sha1 of tree>    arch\n> 040000 tree <sha1 of tree>    block\n> 040000 link <sha1 of tree>     <sha1 of commit>\n\nThat 040000 needs to be something else.\n\nIn order for something like a git-fsck-objects to know that it's a link, \nit needs to be marked as such. \n\nIn git, we never just randomly open an object by SHA1, and then figure out \nits type. We always open things by explicitly knowing both the type and \nthe SHA1, and if the object we find has the wrong type, that's a \nconsistency error in the database (or the user).\n\n"},{"id":"298678","messageId":"Pine.LNX.4.64.0612151809510.3557@woody.osdl.org","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612151747470.3849@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-16T02:12:24Z","receivedAt":"2006-12-16T02:12:24Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 15 Dec 2006, Linus Torvalds wrote:\n\n> \n> On Sat, 16 Dec 2006, Torgil Svensson wrote:\n> > \n> > 100644 blob <sha1 of blob>    README\n> > 100644 blob <sha1 of blob>    REPORTING-BUGS\n> > 100644 link <sha1 of blob>     <sha1 of commit>\n> > 040000 tree <sha1 of tree>    arch\n> > 040000 tree <sha1 of tree>    block\n> > 040000 link <sha1 of tree>     <sha1 of commit>\n> \n> That 040000 needs to be something else.\n\nSide note: that's not to say that I would really see why you'd want to \nhave both the tree and the commit SHA1's, and why you seemingly think that \nthe links don't need a filename. Hmm?\n\nIf you require the tree objects to be in the database, you might as well \nrequire that the commit object be there. But you could make rules that say \nthat subprojects don't need the whole commit history, for example (which \nis just a shallow clone in the subproject).\n\n"},{"id":"296398","messageId":"e7bda7770612160040v1a769153p909a8cd40e5ea991@mail.gmail.com","threadId":"43065","inReplyTo":"elviac$63t$1@sea.gmane.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Torgil Svensson","fromEmail":"torgil.svensson@gmail.com","sentAt":"2006-12-16T08:40:53Z","receivedAt":"2006-12-16T08:40:53Z","isPatch":false,"sender":{"key":"torgil.svensson@gmail.com","avatar":null},"body":"On 12/16/06, Jakub Narebski <jnareb@gmail.com> wrote:\n> > 100644 blob <sha1 of blob>  >\n> >>\n> >\n> > Sorry, I was sloppy and forgot the names:\n> >\n> > 100644 blob <sha1 of blob>    README\n> > 100644 blob <sha1 of blob>    REPORTING-BUGS\n> > 100644 link <sha1 of blob>     <sha1 of commit>   AUTHORS\n> > 040000 tree <sha1 of tree>    arch\n> > 040000 tree <sha1 of tree>    block\n> > 040000 link <sha1 of tree>     <sha1 of commit>   misc\n> >\n> > Now it doesn't looks like trees/blobs anymore so maybe a link object\n> > is handy:\n> >  README\n> > 100644 blob <sha1 of blob>    REPORTING-BUGS\n> > 100644 link <sha1 of link>      AUTHORS\n> > 040000 tree <sha1 of tree>    arch\n> > 040000 tree <sha1 of tree>    block\n> > 040000 link <sha1 of link>     misc\n> >\n> > link-object:\n> > <sha1 of commit>\n> > <sha1 of tree/blob>\n>\n> What do you need <sha1 of tree/blob> for in link-object? Wouldn't you\n> use usually the sha1 of top tree of a commit, which is uniquely defined\n> by commit object, so you need only <ahs1 of commit>?\n>\n\n1. \"Sparse\" repository's - In my example, I want to cherry-pick\nheader-files or binary-files from different projects without fetching\nall, potentially huge, submodules in their entirety. Imaging having X,\nkernel, gcc, gtk and libc6 as sub-projects and you really only care\nabout some header files.\n\n2. Super-module directory-hierarchy independent from submodules.\nSuper-project want to have the header-files and binaries it's own way.\nThis also gives version controlled file-collections, the \"release\ncase\" in my example - collecting different binaries and header-files\nfrom different submodules together in a new directory-structure, add\nsome documentation and configuration files and get the whole thing\nunder strong version-control down to the beginning of time for each\nlittle component.\n\n3. Super-module development independent of submodules - If we have the\ntree/blob-object with all it contents in the database many\ngit-operations can act as the link (commit) wasn't there since we have\naccess to all relevant data to work with. This makes it easy to clone\nthe super-project and work on it seamlessly without having to care\nabout submodules or mapping up submodule repository's (unless you want\nto modify the links or the data underneath it of course).\n\n"},{"id":"295061","messageId":"e7bda7770612160050j526c2b86gbabeae13f2ff114a@mail.gmail.com","threadId":"43065","inReplyTo":"Pine.LNX.4.64.0612151809510.3557@woody.osdl.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Torgil Svensson","fromEmail":"torgil.svensson@gmail.com","sentAt":"2006-12-16T08:50:29Z","receivedAt":"2006-12-16T08:50:29Z","isPatch":false,"sender":{"key":"torgil.svensson@gmail.com","avatar":null},"body":"On 12/16/06, Linus Torvalds <torvalds@osdl.org> wrote:\n\n> Side note: that's not to say that I would really see why you'd want to\n> have both the tree and the commit SHA1's, and why you seemingly think that\n> the links don't need a filename. Hmm?\n\nI really want that file-name back - we can call it a mind short-circuit.\n\n> If you require the tree objects to be in the database, you might as well\n> require that the commit object be there. But you could make rules that say\n> that subprojects don't need the whole commit history, for example (which\n> is just a shallow clone in the subproject).\n\nYou have a very good point here, this would give us the history of the\n"},{"id":"297033","messageId":"em0fpq$45b$1@sea.gmane.org","threadId":"43065","inReplyTo":"e7bda7770612160040v1a769153p909a8cd40e5ea991@mail.gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-16T09:57:36Z","receivedAt":"2006-12-16T09:57:36Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"<opublikowany i wysłany>\n\nTorgil Svensson wrote:\n\n> On 12/16/06, Jakub Narebski <jnareb@gmail.com> wrote:\n\n>>> Now it doesn't looks like trees/blobs anymore so maybe a link object\n>>> is handy:\n>>>  README\n>>> 100644 blob <sha1 of blob>    REPORTING-BUGS\n>>> 100644 link <sha1 of link>      AUTHORS\n>>> 040000 tree <sha1 of tree>    arch\n>>> 040000 tree <sha1 of tree>    block\n>>> 040000 link <sha1 of link>     misc\n\nThis would be (using the submodule original proposal)\n\n    140000 link <sha1 of link>     misc\n\n>>> link-object:\n>>> <sha1 of commit>\n>>> <sha1 of tree/blob>\n>>\n>> What do you need <sha1 of tree/blob> for in link-object? Wouldn't you\n>> use usually the sha1 of top tree of a commit, which is uniquely defined\n>> by commit object, so you need only <sha1 of commit>?\n>>\n> \n> 1. \"Sparse\" repository's - In my example, I want to cherry-pick\n> header-files or binary-files from different projects without fetching\n> all, potentially huge, submodules in their entirety. Imaging having X,\n> kernel, gcc, gtk and libc6 as sub-projects and you really only care\n> about some header files.\n> \n> 2. Super-module directory-hierarchy independent from submodules.\n> Super-project want to have the header-files and binaries it's own way.\n> This also gives version controlled file-collections, the \"release\n> case\" in my example - collecting different binaries and header-files\n> from different submodules together in a new directory-structure, add\n> some documentation and configuration files and get the whole thing\n> under strong version-control down to the beginning of time for each\n> little component.\n\nAll fine, but this does not and I think cannot protect us from the\nfact that we can have <sha1 of tree/blob> which doesn't match\n<sha1 of commit>.\n\nI think it would be better to have sparse/partial checkout first.\nBut that is just my idea. Because with <sha1 of tree/blob> which\nis not sha1 of commit tree you might loose (I think) the ability\nto merge, for example your changes to submodule with upstream.\n\n> 3. Super-module development independent of submodules - If we have the\n> tree/blob-object with all it contents in the database many\n> git-operations can act as the link (commit) wasn't there since we have\n> access to all relevant data to work with. This makes it easy to clone\n> the super-project and work on it seamlessly without having to care\n> about submodules or mapping up submodule repository's (unless you want\n> to modify the links or the data underneath it of course).\n\nThis is I think irrelevant to the fact if we have only <sha1 of commit>,\nor link object and also <sha1 of tree/blob>\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n\n"},{"id":"296633","messageId":"7vac1oi1ce.fsf@assigned-by-dhcp.cox.net","threadId":"43065","inReplyTo":"em0fpq$45b$1@sea.gmane.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-12-16T10:25:53Z","receivedAt":"2006-12-16T10:25:53Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"[jc: adding back people on CC list while feeling sick of having\nto do so...]\n\nJakub Narebski <jnareb@gmail.com> writes:\n\n>>>> Now it doesn't looks like trees/blobs anymore so maybe a link object\n>>>> is handy:\n>>>>  README\n>>>> 100644 blob <sha1 of blob>    REPORTING-BUGS\n>>>> 100644 link <sha1 of link>      AUTHORS\n>>>> 040000 tree <sha1 of tree>    arch\n>>>> 040000 tree <sha1 of tree>    block\n>>>> 040000 link <sha1 of link>     misc\n>\n> This would be (using the submodule original proposal)\n>\n>     140000 link <sha1 of link>     misc\n\nIf I recall correctly, the original original proposal used\n160000 because that is not a bitpattern used in stat.h, and a\nlink behaves like a directory and a symbolic link at the same\ntime.\n"},{"id":"297358","messageId":"e7bda7770612160705l61d1f350n70a8ba91754491c9@mail.gmail.com","threadId":"43065","inReplyTo":"em0fpq$45b$1@sea.gmane.org","subject":"Re: [RFC] Submodules in GIT","fromName":"Torgil Svensson","fromEmail":"torgil.svensson@gmail.com","sentAt":"2006-12-16T15:05:56Z","receivedAt":"2006-12-16T15:05:56Z","isPatch":false,"sender":{"key":"torgil.svensson@gmail.com","avatar":null},"body":"On 12/16/06, Jakub Narebski <jnareb@gmail.com> wrote:\n> All fine, but this does not and I think cannot protect us from the\n> fact that we can have <sha1 of tree/blob> which doesn't match\n> <sha1 of commit>.\n\nTrue, that will be a real problem. Unless we have a bug in git, do you\nsee a scenario in which this is likely to happen?\n\n\n> I think it would be better to have sparse/partial checkout first.\n> But that is just my idea. Because with <sha1 of tree/blob> which\n> is not sha1 of commit tree you might loose (I think) the ability\n> to merge, for example your changes to submodule with upstream.\n\nThat's correct. I also want a sparse/partial checkout but I don't want\nthe full submodule path. I'm also perfectly fine (for my current\nuse-cases) with not being able to merge upstream unless we're tracking\nthe commit tree (here, we might not want to specify the tree SHA1).\n\nI'm not trying to impose a technically fragile solution here [I don't\nbelieve it is, but I'm not the most competent to say that either], I'm\ntrying to find solutions for my use cases and I had problems adapting\nthem to the current suggestion.\n\n\n> > 3. Super-module development independent of submodules - If we have the\n> > tree/blob-object with all it contents in the database many\n> > git-operations can act as the link (commit) wasn't there since we have\n> > access to all relevant data to work with. This makes it easy to clone\n> > the super-project and work on it seamlessly without having to care\n> > about submodules or mapping up submodule repository's (unless you want\n> > to modify the links or the data underneath it of course).\n>\n> This is I think irrelevant to the fact if we have only <sha1 of commit>,\n> or link object and also <sha1 of tree/blob>\n\n"},{"id":"296938","messageId":"e7bda7770612160738w68a47790vef922804efd56c76@mail.gmail.com","threadId":"43065","inReplyTo":"e7bda7770612160705l61d1f350n70a8ba91754491c9@mail.gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Torgil Svensson","fromEmail":"torgil.svensson@gmail.com","sentAt":"2006-12-16T15:38:50Z","receivedAt":"2006-12-16T15:38:50Z","isPatch":false,"sender":{"key":"torgil.svensson@gmail.com","avatar":null},"body":"On 12/16/06, Torgil Svensson <torgil.svensson@gmail.com> wrote:\n> On 12/16/06, Jakub Narebski <jnareb@gmail.com> wrote:\n> > All fine, but this does not and I think cannot protect us from the\n> > fact that we can have <sha1 of tree/blob> which doesn't match\n> > <sha1 of commit>.\n>\n> True, that will be a real problem. Unless we have a bug in git, do you\n> see a scenario in which this is likely to happen?\n>\n> I also want a sparse/partial checkout but I don't want\n> the full submodule path.\n\nThis might not be as problematic as we think. If we do the same\nsparse/partial checkout (what's the definition here?) with the <sha1\nof tree/blob> as we do with the only <sha1 of commit> case and\nconsider the <sha1 of tree/blob> to be a _local_ (to the\nsuper-project) shortcut. Then we only track the submodules using the\ncommit - local conflicts are easier to handle, git would refuse to\ncommit a <sha1 of tree/blob> not present in the commit tree.\n\nWe might even consider two object types:\nmodule: <sha1 of commit> name\nlink: <sha1 of commit> <sha1 of tree/blob> name\n\n"},{"id":"295524","messageId":"200612161732.11746.jnareb@gmail.com","threadId":"43065","inReplyTo":"e7bda7770612160705l61d1f350n70a8ba91754491c9@mail.gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-16T16:32:10Z","receivedAt":"2006-12-16T16:32:10Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Torgil Svensson wrote:\n> On 12/16/06, Jakub Narebski <jnareb@gmail.com> wrote:\n>> All fine, but this does not and I think cannot protect us from the\n>> fact that we can have <sha1 of tree/blob> which doesn't match\n>> <sha1 of commit>.\n> \n> True, that will be a real problem. Unless we have a bug in git, do you\n> see a scenario in which this is likely to happen?\n\nWell, I just rather have than <sha1 of tree/blob> the definition\nof sparse checkout (for example subdirectory name, or file name,\nor glob pattern).\n\nBesides you need the name of directory (for tree) or file (for blob),\notherwise you would have no way to update it when submodule advances\nversion, and you want to use new submodule version. And if you have\nthat, you don't need <sha1 of tree/blob> in repository, in link object.\nYou might want it in the index, for performance reasons, though.\n \n>> I think it would be better to have sparse/partial checkout first.\n>> But that is just my idea. Because with <sha1 of tree/blob> which\n>> is not sha1 of commit tree you might loose (I think) the ability\n>> to merge, for example your changes to submodule with upstream.\n> \n> That's correct. I also want a sparse/partial checkout but I don't want\n> the full submodule path. I'm also perfectly fine (for my current\n> use-cases) with not being able to merge upstream unless we're tracking\n> the commit tree (here, we might not want to specify the tree SHA1).\n\nWith sparse (for example defined by 'src/*.h') or partial (for example\ndefined by 'Documentation/') checkout you should be able to merge\nupstream... unless conflicts are in the not checked out part.\n\n> I'm not trying to impose a technically fragile solution here [I don't\n> believe it is, but I'm not the most competent to say that either], I'm\n> trying to find solutions for my use cases and I had problems adapting\n> them to the current suggestion.\n\nHave you read  http://git.or.cz/gitwiki/SubprojectSupport on GitWiki?\nHave you tested the experimental submodule support (proof of concept)\n  http://git.admingilde.org/tali/git.git/module2\nby Martin Waitz?\n\n-- \nJakub Narebski\n"},{"id":"294252","messageId":"e7bda7770612161621p324ba357x883ed28c46597750@mail.gmail.com","threadId":"43065","inReplyTo":"200612161732.11746.jnareb@gmail.com","subject":"Re: [RFC] Submodules in GIT","fromName":"Torgil Svensson","fromEmail":"torgil.svensson@gmail.com","sentAt":"2006-12-17T00:21:10Z","receivedAt":"2006-12-17T00:21:10Z","isPatch":false,"sender":{"key":"torgil.svensson@gmail.com","avatar":null},"body":"On 12/16/06, Jakub Narebski <jnareb@gmail.com> wrote:\n> Well, I just rather have than <sha1 of tree/blob> the definition\n> of sparse checkout (for example subdirectory name, or file name,\n> or glob pattern).\n\nThis is entirely an UI issue:\n\nOn 12/16/06, Torgil Svensson <torgil.svensson@gmail.com> wrote:\n> is based around Linus 'module'-file and keep things simple. A git\n> configuration file that specifies:\n> * link name for reference\n> * local path to link\n> * submodule source\n> * submodule path to tree/blob\n> * submodule commit / HEAD / branch\n> * options (depth-limit , ...)\n\n\n\nOn 12/16/06, Jakub Narebski <jnareb@gmail.com> wrote:\n> And if you have that, you don't need <sha1 of tree/blob> in repository, in link object.\n\nCorrect. Since the commit contains all the version information, the\nfollowing combinations should give the same information iff we keep\nthe commit in the database:\n1. <sha1 of commit> + <sha1 of tree/blob>\n2. <sha1 of commit> + <symlink to tree/blob>\n\nI used the sha1 because I wanted them to behave exactly like\ntrees/blobs in the database for operations that can disregard the\ncommit info. Now, if we keep the commit in the database as Linus\nsuggests we can reach the target from there with a symlink. This would\nbe more readable but also cost a few object lookups extra iterating\nover the symlink.\n\n\n> With sparse (for example defined by 'src/*.h') or partial (for example\n> defined by 'Documentation/') checkout you should be able to merge\n> upstream... unless conflicts are in the not checked out part.\n\nThis would be a great feature! Will this conflict with path shortcuts?\nIf so, we might consider two types of objects: \"link\" which cannot\nmerge upstream and \"module\" which can merge upstream and contains a\n.git object repository.\n\nIMHO, \"module\" is a more intuitive name for specifying a\n(functionality wise fully fledged) submodule with a repository inside.\n\"link\" could be used for just mirroring a tree/blob. I'm not sure if a\nseparation is needed on a technical level.\n\n\n> Have you read  http://git.or.cz/gitwiki/SubprojectSupport on GitWiki?\nYes\n> Have you tested the experimental submodule support (proof of concept)\nNot yet\n\n\n"}]}