{"thread":{"id":"43100","subject":"Re: Subprojects tasks","startedAt":"2006-12-16T18:32:36Z","lastAt":"2006-12-18T10:30:18Z","messageCount":24,"participants":["Sven Verdoolaege","Josef Weidendorfer","Martin Waitz","Jakub Narebski","Junio C Hamano","Alan Chandler"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"297413","messageId":"7vzm9nelob.fsf@assigned-by-dhcp.cox.net","threadId":"43100","inReplyTo":null,"subject":"Subprojects tasks","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-12-16T18:32:36Z","receivedAt":"2006-12-16T18:32:36Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Because I am primarily a plumber, I was thinking about the\nchanges that need to be done at the plumbing level.  I only\nlooked at the prototype when it was announced, and I do not know\nthe progress you made since then.  Could you tell us the current\nstatus?\n\nI am assuming that the overall design is based on what Linus\nproposed long time ago with his \"gitlink\" object.  That is,\n\n * the index and the tree object for the superproject has a\n   \"link\" object that tells there is a directory and the\n   corresponding commit object name from the subproject.  Unlike\n   my previous \"bind commit\" based prototype, index does not\n   have any blobs nor trees from the subproject.\n\n * the subproject is on its own, and can exist unaware of the\n   existence of its superproject (there is no back-link at the\n   object layer).\n\n * the subproject and the superproject are loosely coupled.  An\n   act of committing in one does not automatically make a\n   corresponding commit in the other at the plumbing level.\n\nFor now, I assume that the representation of the \"link\" object\nis (after the usual object header) 40-byte hexadecimal SHA-1 of\nthe commit object, plus a LF, but that is a minor detail.\n\nAt the object layer, obviously you would need a new object type\nallocated, \"link\", and the mode-bits assigned in the tree\nobject.  You should be able to cover cat-file, fsck-objects\n(connectivity through link->commit), pack-objects and\nunpack-objects (they need to know about the new object type) at\nthat stage.\n\nWith the index, in addition to the mode bits and \"link\" object's\nSHA-1, you would need to decide what to do with ce_match_stat(),\nto keep track of the information \"update-index --refresh\" updates.\nMy recommendation is to:\n\n (1) Look at the directory the \"link\" is at, and find .git/\n     subdirectory (that is the $GIT_DIR for the subproject) and\n     its .git/HEAD;\n\n (2) If that points at a loose ref, use the file's stat()\n     information (e.g stat(\"$sub/.git/refs/heads/master\"));\n\n (3) Otherwise, use the packed-ref file's stat() information\n     (e.g stat(\"$sub/.git/packed-refs\")).\n\nThen ce_match_stat() for a \"link\" entry can do the same\ncomputation and tell if the subproject has changed its HEAD.\n\nI think \"update-index --add $directory\" should check if .git/\nexists and looks like a valid repository, and make a \"link\"\nobject out of \"$directory/.git/HEAD\".\n\nAnother issue with the index is what to put in the cache_tree\nstructure; I think \"link\" can be treated just like blob (both\nfiles and symlinks).\n\nThen read-tree (bulk of it is in unpack-trees.c) needs to be\ntaught to read in \"link\" and put that into the index -- this\nshould be straightforward.\n\nAfter you have a working index, you should be able to do\nwrite-tree (writes the new \"link\" entry as is, without\ndescending into the subproject) trivially.\n\nIt is debatable what 'checkout-index -f' should do when the\nsubproject is already checked out and its HEAD points at a\ndifferent commit.  I am tempted to say that it should go there\nand run \"reset --hard\", but I feel uneasy about that because it\nis a blatant layering violation.  Maybe it should simply ignore\nlink entries and let the Porcelain layer take care of them.\n\nThen there are three diff- brothers at the plumbing level.  I\nthink it is reasonable not to make them recurse into \"link\",\neven with the presense of -r (recursive) option (Porcelain \"git\ndiff\" might want to recurse into the subproject, perhaps with a\nnew --recurse-harder option, though).\n\nThat means diff-files either skips a \"link\" entry if\nce_match_stat() says it is clean, or feeds \"link\" and its\nrecorded SHA-1 from the index on the left hand side, and 0{40}\non the right hand side with \"link\" type (after verifying that\n\"$sub/.git\" is still there -- otherwise you would say that the\nworking tree has lost that subproject).  \"Read from the working\ntree\" done for diff-files for a \"link\" object would grab the\ncommit SHA-1 from the tip of the current branch of the\nsubproject, format it as the value of a \"link\" object (I am\nassuming just a 40-byte commit SHA-1, plus a LF) and would\ncompare that with the result of read_sha1_file() on the link\nobject recorded in the index when producing -p (patch) output.\n\ndiff-tree would compare \"link\" entries without descending into\nthem.  \"link\" and \"blob\" would compare just like \"symlink\" and\n\"blob\" would compare.\n\ndiff-index with --cached would work like diff-tree (two concrete\nSHA-1), and without --cached would work like diff-files (one\nSHA-1 from the tree, another is either from the index if\nce_match_stat() says it is clean, 0{40} otherwise).\n\nI suspect the hardest part is \"rev-list --objects\" (now most of\nit is found in revision.c).  Theoretically, if the code can\nhandle \"tag\"s, it should be able to handle \"link\"s, but I have a\nfeeling that the ancestry traversal code that walks commits is\nnot prepared to see \"commit\" object to appear from somewhere in\nthe middle of traversal.  A commit so far can be wrapped only by\ntags zero or more times, and a tag never appears inside anything\nbut another tag, so the code can just keep peeling the tag until\nit sees a non-tag and after that it will be living in the world\nthat has only commit->tree->blob hierarchy, and can afford to do\nthe ancestry based solely on \"commit\" and can treat reachability\nfor \"tree\" and \"blob\" as afterthought.  But I think the updated\ncode needs to know that \"link\" needs to be unwrapped and\ncontained \"commit\" needs to be injected back to the ancestry\nwalking machinery.\n\nOnce you have \"rev-list --objects\", you should be able to drive\npack-objects with its output.  I do not think there is much to\nchange in that program.\n\nMy gut feeling is that it would take about 2-3 weeks for a\ncompetent plumber working on full time to make the above changes\nto the plumbing side into presentable shape.\n\nOn the Porcelain side, you would need policy design, some of\nwhich were discussed on the list, such as what committing and\nfetching in a superproject mean and should do.  I do not have a\nguestimate of the amount of work that would involve.\n"},{"id":"295059","messageId":"em1erl$pne$1@sea.gmane.org","threadId":"43100","inReplyTo":"7vzm9nelob.fsf@assigned-by-dhcp.cox.net","subject":"Re: Subprojects tasks","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-16T18:45:11Z","receivedAt":"2006-12-16T18:45:11Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Junio C Hamano wrote:\n\n>  (1) Look at the directory the \"link\" is at, and find .git/\n>      subdirectory (that is the $GIT_DIR for the subproject) and\n>      its .git/HEAD;\n\nOr .gitlink file, if we decide to implement it (as lightweight checkout and\nsupport for submodules which one can easily move/rename).\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n\n"},{"id":"293844","messageId":"20061216203553.GA25274MdfPADPa@greensroom.kotnet.org","threadId":"43100","inReplyTo":"7vzm9nelob.fsf@assigned-by-dhcp.cox.net","subject":"Re: Subprojects tasks","fromName":"Sven Verdoolaege","fromEmail":"skimo@kotnet.org","sentAt":"2006-12-16T20:35:53Z","receivedAt":"2006-12-16T20:35:53Z","isPatch":false,"sender":{"key":"skimo@kotnet.org","avatar":null},"body":"On Sat, Dec 16, 2006 at 10:32:36AM -0800, Junio C Hamano wrote:\n> I suspect the hardest part is \"rev-list --objects\" (now most of\n> it is found in revision.c).  [..]  But I think the updated\n> code needs to know that \"link\" needs to be unwrapped and\n> contained \"commit\" needs to be injected back to the ancestry\n> walking machinery.\n\nDo we want \"link\" to be unwrapped, though ?\n\n> Once you have \"rev-list --objects\", you should be able to drive\n> pack-objects with its output.\n\nWouldn't we then run into the scalability problems Linus was\nconcerned about ?\n\n"},{"id":"295809","messageId":"7v64cbeeiv.fsf@assigned-by-dhcp.cox.net","threadId":"43100","inReplyTo":"20061216203553.GA25274MdfPADPa@greensroom.kotnet.org","subject":"Re: Subprojects tasks","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-12-16T21:07:04Z","receivedAt":"2006-12-16T21:07:04Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Sven Verdoolaege <skimo@kotnet.org> writes:\n\n> On Sat, Dec 16, 2006 at 10:32:36AM -0800, Junio C Hamano wrote:\n>> I suspect the hardest part is \"rev-list --objects\" (now most of\n>> it is found in revision.c).  [..]  But I think the updated\n>> code needs to know that \"link\" needs to be unwrapped and\n>> contained \"commit\" needs to be injected back to the ancestry\n>> walking machinery.\n>\n> Do we want \"link\" to be unwrapped, though ?\n>\n>> Once you have \"rev-list --objects\", you should be able to drive\n>> pack-objects with its output.\n>\n> Wouldn't we then run into the scalability problems Linus was\n> concerned about ?\n\nHmph.\n\nIf the plumbing layer does not have to (although I haven't\nthought it through, it does feel like it even shouldn't) unwrap\n\"link\" and let the Porcelain layer to deal with it, that would\ncertainly make rev-list/revision.c part simpler.\n\nI like it.\n\n"},{"id":"295516","messageId":"20061216225810.GD12411@admingilde.org","threadId":"43100","inReplyTo":"7v64cbeeiv.fsf@assigned-by-dhcp.cox.net","subject":"Re: Subprojects tasks","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-16T22:58:10Z","receivedAt":"2006-12-16T22:58:10Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nJunio, I'll take a more detailed look at your mail tomorrow, after I\nregenerated from all the Guinness I had tonight ;-)\n\nOn Sat, Dec 16, 2006 at 01:07:04PM -0800, Junio C Hamano wrote:\n> Sven Verdoolaege <skimo@kotnet.org> writes:\n> > On Sat, Dec 16, 2006 at 10:32:36AM -0800, Junio C Hamano wrote:\n> >> I suspect the hardest part is \"rev-list --objects\" (now most of\n> >> it is found in revision.c).  [..]  But I think the updated\n> >> code needs to know that \"link\" needs to be unwrapped and\n> >> contained \"commit\" needs to be injected back to the ancestry\n> >> walking machinery.\n\nWell, I already got to the point of using the commit directly,\ninstead of any link object.  It even worked with rev-list --objects\nin all my test cases.  That is, I could correctly clone/pack/pull\nthe complete project including all modules.\n\n> > Wouldn't we then run into the scalability problems Linus was\n> > concerned about ?\n\nThis is a real problem.\n\n> If the plumbing layer does not have to (although I haven't\n> thought it through, it does feel like it even shouldn't) unwrap\n> \"link\" and let the Porcelain layer to deal with it, that would\n> certainly make rev-list/revision.c part simpler.\n\nYes.  However, it makes other things more complicated.\nIf the plumbing does not do all the subproject stuff and you don't have\neverything in one database it is much more difficult to really get\na consistent database when cloning or fetching (you have to get even old\nsubmodule commits which are not reachable by the current supermodule\ntree anymore, perhaps even submodules which do not exist anymore).\n\nI did not have much time to think about these issues in the last day and\nam not yet convinced on how to proceed,\n\n-- \nMartin Waitz\n"},{"id":"295706","messageId":"20061216230108.GE12411@admingilde.org","threadId":"43100","inReplyTo":"em1erl$pne$1@sea.gmane.org","subject":"Re: Subprojects tasks","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-16T23:01:08Z","receivedAt":"2006-12-16T23:01:08Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Sat, Dec 16, 2006 at 07:45:11PM +0100, Jakub Narebski wrote:\n> Or .gitlink file, if we decide to implement it (as lightweight checkout and\n> support for submodules which one can easily move/rename).\n\nI still don't get the advantage of a .gitlink file over an ordinary\nrepository with alternates or a symlink.\n\n-- \nMartin Waitz\n"},{"id":"296983","messageId":"20061216231422.GE25274MdfPADPa@greensroom.kotnet.org","threadId":"43100","inReplyTo":"20061216225810.GD12411@admingilde.org","subject":"Re: Subprojects tasks","fromName":"Sven Verdoolaege","fromEmail":"skimo@kotnet.org","sentAt":"2006-12-16T23:14:22Z","receivedAt":"2006-12-16T23:14:22Z","isPatch":false,"sender":{"key":"skimo@kotnet.org","avatar":null},"body":"On Sat, Dec 16, 2006 at 11:58:10PM +0100, Martin Waitz wrote:\n> On Sat, Dec 16, 2006 at 01:07:04PM -0800, Junio C Hamano wrote:\n> > Sven Verdoolaege <skimo@kotnet.org> writes:\n> > > On Sat, Dec 16, 2006 at 10:32:36AM -0800, Junio C Hamano wrote:\n> > >> I suspect the hardest part is \"rev-list --objects\" (now most of\n> > >> it is found in revision.c).  [..]  But I think the updated\n> > >> code needs to know that \"link\" needs to be unwrapped and\n> > >> contained \"commit\" needs to be injected back to the ancestry\n> > >> walking machinery.\n> \n> Well, I already got to the point of using the commit directly,\n> instead of any link object.\n\nI think Junio is simply refering to the type of the object as represented\nin a tree and that the value would indeed just be the commit hash, as in\nyour implementation.\n\n"},{"id":"295079","messageId":"200612170015.24162.jnareb@gmail.com","threadId":"43100","inReplyTo":"20061216230108.GE12411@admingilde.org","subject":"Re: Subprojects tasks","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-16T23:15:23Z","receivedAt":"2006-12-16T23:15:23Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hi!\n\nMartin Waitz wrote:\n> On Sat, Dec 16, 2006 at 07:45:11PM +0100, Jakub Narebski wrote:\n>>\n>> Or .gitlink file, if we decide to implement it (as lightweight checkout and\n>> support for submodules which one can easily move/rename).\n> \n> I still don't get the advantage of a .gitlink file over an ordinary\n> repository with alternates or a symlink.\n\nMoving or renaming the directory with a submodule. With alternates,\nwhen you rename or move directory with a submodule, you have to add\nalternate for new place / new name, or alter existing alternate.\nWith symlinks you risk broken symlinks.\n\nWhen using alternates-like modules file, you can regenerate or\ngenerate \"alternates\" on checkout, but...\n\nWith .gitlink file you can specify GIT_DIR sor submodule as given\ndirectory relative to this directory or one of its parents, so you\ncan rename and move submodules freely.\n\n\nP.S. The second (first?) purpose of .gitlink is to be able to have\nlightweight checkout, i.e. more than one working area associated with\none repository.\n\nP.P.S. Cc to the author of current .gitlink proposal, to Josef\nWeidendorfer.\n  Message-ID: <200612082252.31245.Josef.Weidendorfer@gmx.de>\n  http://permalink.gmane.org/gmane.comp.version-control.git/33755\n-- \nJakub Narebski\n"},{"id":"298107","messageId":"200612170101.09615.Josef.Weidendorfer@gmx.de","threadId":"43100","inReplyTo":"200612170015.24162.jnareb@gmail.com","subject":"Re: Subprojects tasks","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-17T00:01:09Z","receivedAt":"2006-12-17T00:01:09Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Sunday 17 December 2006 00:15, Jakub Narebski wrote:\n> Hi!\n> \n> Martin Waitz wrote:\n> > On Sat, Dec 16, 2006 at 07:45:11PM +0100, Jakub Narebski wrote:\n> >>\n> >> Or .gitlink file, if we decide to implement it (as lightweight checkout and\n> >> support for submodules which one can easily move/rename).\n> > \n> > I still don't get the advantage of a .gitlink file over an ordinary\n> > repository with alternates or a symlink.\n> \n> Moving or renaming the directory with a submodule. With alternates,\n> when you rename or move directory with a submodule, you have to add\n> alternate for new place / new name, or alter existing alternate.\n> With symlinks you risk broken symlinks.\n\nYes.\n\nIMHO it simply is added flexibility to allow a checkout to be separate from\nthe .git/ directory, same as explicitly setting $GIT_DIR would do.\nSo this .gitlink file is on the one hand one kind of convenience for users\nwhich want to keep their repository separate, yet do not want to specify\n$GIT_DIR all the time in front of git commands.\nThe .gitlink file simply makes the linkage to the separate repository\npersistent.\n\nIn the scope of submodules, you get the benefit that you can not lose\nsubmodule repositories by doing a \"rm -rf *\" (or similar, e.g. deleting\ndirs with submodules in it) in the supermodule checkout. Actually, the\nlatter is a valid action: delete a submodule in the next commit;\nwhen going back at an earlier commit, the submodule should be there again.\nSo IMHO you allow far more possibilities by separating GITDIR from the\ncheckout of submodules.\n\nHowever. I think that this .gitlink file proposal can be seen as kind\nof independent from submodule support at first; it should be easy to\nmake this work together later on. E.g. submodule root directories\ncan be easily detected when they have a .gitlink file (instead of\n.git/ directory), and so on.\n\nThis said, I started implementing it, but do not have anything useful\nto show yet.\nSome issues:\n* Probably, it is better to go with a _file_ .git instead of a file\n.gitlink, as this way, the user is forced to either go with the\ngit repository in _directory_ .git/ or external linkage with\nthe _file_ \".git\".\n* Even when a .gitlink file is detected, we should honor a\n$GIT_DIR environment variable set by the user. Unfortunately, $GIT_DIR\nalso can be set by porcelain commands to specify \"this command only\nworks in the toplevel directory of a git checkout\", i.e. these\nporcelain commands set GIT_DIR to \".git\". IMHO this is a hack, and\nwe explicitly should tell the plumbing about these need e.g. via another\nenvironment variable (or a option) without implicitly forcing it by\nsetting $GIT_DIR. \n* In the way to make the .gitlink file as flexible as\npossible (and to use it for lightweight checkouts), it really should\nsupport $GIT_HEAD_FILE, which would replace \"HEAD\" with the content\nof $GIT_HEAD_FILE. E.g. with GIT_HEAD_FILE=MYHEAD, the command\n\"git log HEAD\" really should internally work as \"git log MYHEAD\"\n(ie. use the .git/MYHEAD file instead). It is arguable whether the\nusage of \"ORIG_HEAD\" by the user or in porcelain should map to file\n\"ORIG_MYHEAD\". Probably not. However, changing this in all places\nis some work, and I assume that therefore nobody has ever implemented\n$GIT_HEAD_FILE - which IMHO really would be useful by itself. \n\n"},{"id":"294981","messageId":"200612170108.04033.Josef.Weidendorfer@gmx.de","threadId":"43100","inReplyTo":"200612170015.24162.jnareb@gmail.com","subject":"Re: Subprojects tasks","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-17T00:08:03Z","receivedAt":"2006-12-17T00:08:03Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Sunday 17 December 2006 00:15, Jakub Narebski wrote:\n> > I still don't get the advantage of a .gitlink file over an ordinary\n> > repository with alternates or a symlink.\n\nForgot one thing:\nTo separate the repository files from a checkout, a symlink is not\nenough, as you lose the linkage when you move the checkout or the\nrepository; you could use an absolute symlink target, but that also\nhas inconveniences.\n\nSo you want some kind of smart linking. And this is another\nimportant part of the .gitlink file proposal.\n\n"},{"id":"293875","messageId":"200612170132.14276.Josef.Weidendorfer@gmx.de","threadId":"43100","inReplyTo":"20061216225810.GD12411@admingilde.org","subject":"Re: Subprojects tasks","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-17T00:32:14Z","receivedAt":"2006-12-17T00:32:14Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Saturday 16 December 2006 23:58, Martin Waitz wrote:\n> > If the plumbing layer does not have to (although I haven't\n> > thought it through, it does feel like it even shouldn't) unwrap\n> > \"link\" and let the Porcelain layer to deal with it, that would\n> > certainly make rev-list/revision.c part simpler.\n> \n> Yes.  However, it makes other things more complicated.\n> If the plumbing does not do all the subproject stuff and you don't have\n> everything in one database\n\nEven without plumbing doing the subproject stuff, we could use the\nsame, unified database for the objects. Or do I miss something?\n\nAs you said: the problem are submodule commit in superproject trees which\nare not reachable by refs of the submodule. However, we only need these\ncommits when cloning/fetching the submodule in the scope of cloning/fetching\nthe superproject; we simply can not use here a normal repository of the\nsubmodule, as these commits would be not available there.\n\nWe should add a plumbing command for \"Give me the minimal set of commits\n(from all submodules) which have all the submodule link object ids as ancestors\nwhich appear in the history of a given commit (from a superproject)\".\nWith this, building the set of objects to pack/fetch/clone into a unified\nobject database for a superproject with its submodules should be easy.\n\nIt is also needed for pruning the unified object database.\nPruning in submodules simply would print out an error \"Pruning in submodules\nnot supported. Prune in the superproject instead\".\n\nJosef\n"},{"id":"297331","messageId":"200612170848.10092.alan@chandlerfamily.org.uk","threadId":"43100","inReplyTo":"7vzm9nelob.fsf@assigned-by-dhcp.cox.net","subject":"Re: Subprojects tasks","fromName":"Alan Chandler","fromEmail":"alan@chandlerfamily.org.uk","sentAt":"2006-12-17T08:48:10Z","receivedAt":"2006-12-17T08:48:10Z","isPatch":false,"sender":{"key":"alan@chandlerfamily.org.uk","avatar":"https://gravatar.com/avatar/1862247e5ea8eac114c842f9dc3a5db6253754e24ef7171757cf97eedce48b8c?d=mp&s=160"},"body":"On Saturday 16 December 2006 18:32, Junio C Hamano wrote:\n> Because I am primarily a plumber, I was thinking about the\n> changes that need to be done at the plumbing level.  I only\n> looked at the prototype when it was announced, and I do not know\n> the progress you made since then.  Could you tell us the current\n> status?\n>\n> I am assuming that the overall design is based on what Linus\n> proposed long time ago with his \"gitlink\" object.  That is,\n>\n>  * the index and the tree object for the superproject has a\n>    \"link\" object that tells there is a directory and the\n>    corresponding commit object name from the subproject.  Unlike\n>    my previous \"bind commit\" based prototype, index does not\n>    have any blobs nor trees from the subproject.\n>\n>  * the subproject is on its own, and can exist unaware of the\n>    existence of its superproject (there is no back-link at the\n>    object layer).\n\nI have been following the submodules (subprojects - is there any \ndifference?) discussion from afar, getting lost quite frequently in \nwhat is actually being discussed and why.  I don't think the idea I \nexpress below has been mentioned, but apologies if it has.\n\nOne element I felt has been missing is a vigorous discussion of what \nsubmodules are for and what are their use cases.  The \"submodule is on \nits own\" issue seems to have crept into the discussion - but there was \none use case that was discussed, where some actually help by the \nsubmodule could be useful.\n\nThe use case was when the supermodule wanted to make use of the header \nfiles of the submodule because it was using the submodule as a library.\n\nThis did make me wonder if the submodule should not export some form \nof \"approved\" set of content (or files - and I do think care is needed \nhere as to which it is when we think about renames) which is both\n\na) a subset of the full tree that is stored at commit time, and\nb) does itself have a commit history \n\n(I am clearly thinking that would be the standard \"include\" files, but \nnot the actual source of the library - (but it might include the \nlibrary it self as a prebuilt binary library?)\n\nThis does suggest it is a tree object stored in the repository - and \nthat it is linked in time via a set of commit objects - I'll call them \nthe \"export commits\".  I am not sure whether a new commit should be \nmade everytime there is any change (via a normal commit) to this \ncontent, or (and I slightly favour this) there is a new commit made \nwhich is somewhat akin to a tag when the project wants to release a new \nversion of its interface. \n\nSupermodules, which then made use of that library would, the do some \nform of shallow clone, shallow in the sense that it only pulled in the \nexported commit content and also (possibly) shallow in the sense that \nit does not need to go back in time to get old versions of the exported \ncommit.\n\n\n-- \nAlan Chandler\n"},{"id":"294458","messageId":"em34cd$ni8$1@sea.gmane.org","threadId":"43100","inReplyTo":"200612170848.10092.alan@chandlerfamily.org.uk","subject":"Re: Subprojects tasks","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-17T10:01:10Z","receivedAt":"2006-12-17T10:01:10Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Alan Chandler wrote:\n\n> The use case was when the supermodule wanted to make use of the header \n> files of the submodule because it was using the submodule as a library.\n> \n> This did make me wonder if the submodule should not export some form \n> of \"approved\" set of content (or files - and I do think care is needed \n> here as to which it is when we think about renames) which is both\n> \n> a) a subset of the full tree that is stored at commit time, and\n> b) does itself have a commit history \n> \n> (I am clearly thinking that would be the standard \"include\" files, but \n> not the actual source of the library - (but it might include the \n> library it self as a prebuilt binary library?)\n> \n> This does suggest it is a tree object stored in the repository - and \n> that it is linked in time via a set of commit objects - I'll call them \n> the \"export commits\".  I am not sure whether a new commit should be \n> made everytime there is any change (via a normal commit) to this \n> content, or (and I slightly favour this) there is a new commit made \n> which is somewhat akin to a tag when the project wants to release a new \n> version of its interface. \n\nIn the absence of sparse/partial checkout, and it's use in submodule\nsupport, this can be solvd purely on porcelain level.\n\nYou would have to simply maintain separate 'includes' branch, similarly\nto how 'html' and 'man' (and 'todo') branches are maintained in git.git\nrepository -- it would be your 'set of commit objects'. Then the only\nthink that would be needed is some commit / post-commit hook which would\nexamine if commit touches \"include\" files and if it does, make a commit\nin the 'includes' ('inc' for short) branch.\n\nSuportmodule would then use either 'master' branch for full sources,\nor 'includes' branch for headers only.\n\nP.S. Cc: Alan Chandler <alan@chandlerfamily.org.uk>, \nJunio C Hamano <junkio@cox.net>, git@vger.kernel.org\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n\n"},{"id":"296541","messageId":"20061217111723.GF12411@admingilde.org","threadId":"43100","inReplyTo":"7vzm9nelob.fsf@assigned-by-dhcp.cox.net","subject":"Re: Subprojects tasks","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-17T11:17:23Z","receivedAt":"2006-12-17T11:17:23Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Sat, Dec 16, 2006 at 10:32:36AM -0800, Junio C Hamano wrote:\n> Because I am primarily a plumber, I was thinking about the\n> changes that need to be done at the plumbing level.  I only\n> looked at the prototype when it was announced, and I do not know\n> the progress you made since then.  Could you tell us the current\n> status?\n\nMost of the things you described are already implemented in\nhttp://git.admingilde.org/tali/git.git/module2\n\nIf there is interest in it, I can generate some nice patches out of it.\nHowever with Linus concerns about scalability I'm not sure it is ready\nyet.  But if you prefer patches for discussion I'll send them here.\n\n> I am assuming that the overall design is based on what Linus\n> proposed long time ago with his \"gitlink\" object.  That is,\n> \n>  * the index and the tree object for the superproject has a\n>    \"link\" object that tells there is a directory and the\n>    corresponding commit object name from the subproject.  Unlike\n>    my previous \"bind commit\" based prototype, index does not\n>    have any blobs nor trees from the subproject.\n\nIn contrast to your description, my implementation does not\nintroduce a new \"link\" type but instead adds the reference to the\nsubmodule commit directly to the parent tree object and to its\nindex.\n\n>  * the subproject is on its own, and can exist unaware of the\n>    existence of its superproject (there is no back-link at the\n>    object layer).\n\nyes, this is essential.\nThere may be links in this particular instance of the submodule, i.e.\nthe repository/working directory which are checked out by the\nsupermodule may be coupled to the supermodule, but it must always\nbe possible to clone/push/pull the submodule alone.\n\n>  * the subproject and the superproject are loosely coupled.  An\n>    act of committing in one does not automatically make a\n>    corresponding commit in the other at the plumbing level.\n\nthis is essential, too.\n\n> With the index, in addition to the mode bits and \"link\" object's\n> SHA-1, you would need to decide what to do with ce_match_stat(),\n> to keep track of the information \"update-index --refresh\" updates.\n> My recommendation is to:\n> \n>  (1) Look at the directory the \"link\" is at, and find .git/\n>      subdirectory (that is the $GIT_DIR for the subproject) and\n>      its .git/HEAD;\n> \n>  (2) If that points at a loose ref, use the file's stat()\n>      information (e.g stat(\"$sub/.git/refs/heads/master\"));\n> \n>  (3) Otherwise, use the packed-ref file's stat() information\n>      (e.g stat(\"$sub/.git/packed-refs\")).\n> \n> Then ce_match_stat() for a \"link\" entry can do the same\n> computation and tell if the subproject has changed its HEAD.\n\nyes.\nHowever I decicided not to read in HEAD but some specific branch.\nThis may sound arbitrary and I did not really like to make\n\"master\" (the branch I chose) even more special, but you will\nunderstand it when looking at the checkout below.\n\n> Another issue with the index is what to put in the cache_tree\n> structure; I think \"link\" can be treated just like blob (both\n> files and symlinks).\n\nHmm, I never cared about cache_tree up to now.  I guess I should learn\nabout it to understand the influence on submodules.\n\n> Then read-tree (bulk of it is in unpack-trees.c) needs to be\n> taught to read in \"link\" and put that into the index -- this\n> should be straightforward.\n> \n> After you have a working index, you should be able to do\n> write-tree (writes the new \"link\" entry as is, without\n> descending into the subproject) trivially.\n\nWhere do you want to write the link to?\nWhat I do here is update one branch (\"master\") of the submodule to\nthe new commit which was stored in the parent index.\nIf this branch is currently checked out, the working directory will\nbe updated, too.  If there is no working directory for the submodule\nyet, it will be created.\n\nUpdating one special branch instead of HEAD is because the submodule\ncommits which are stored in the supermodule really can be considered\nas a special branch which happens to not be stored in an ordinary ref.\nIn order to make it visible to the user the commit is copied to a\nnormal ref.\nThis approach also integrates better with branches in the submodule.\nWhen you want to start parallel development in a branch you eigther\nwant to do this in the complete supermodule scope -- then you have\nto branch the supermodule --, or you want to do it independent to the\nversion stored in the supermodule -- then you don't want a supermodule\ncheckout to mess with your branch.\nSo it makes sense to have one branch which is tracked by the parent\nand other branches which are independent from the parent.\n\n\n> It is debatable what 'checkout-index -f' should do when the\n> subproject is already checked out and its HEAD points at a\n> different commit.  I am tempted to say that it should go there\n> and run \"reset --hard\", but I feel uneasy about that because it\n> is a blatant layering violation.  Maybe it should simply ignore\n> link entries and let the Porcelain layer take care of them.\n\nWhere exactly do you see the layering violation?\nWell I think it makes sense to use read-tree -m <old> <new> in the\nsubmodule instead of a hard reset, but when the supermodule is checked\nout the submodule really should move to its new version.\n(At least the branch which is tracked by the parent should do so.)\n\n\n> Then there are three diff- brothers at the plumbing level.\n\nAll the diff stuff is what is still missing in my implementation.\nIf you ask for a diff in the parent, it will happily diff the\nsubmodules commit objects ;-)\n\n\n> I suspect the hardest part is \"rev-list --objects\" (now most of\n> it is found in revision.c).  Theoretically, if the code can\n> handle \"tag\"s, it should be able to handle \"link\"s, but I have a\n> feeling that the ancestry traversal code that walks commits is\n> not prepared to see \"commit\" object to appear from somewhere in\n> the middle of traversal.  A commit so far can be wrapped only by\n> tags zero or more times, and a tag never appears inside anything\n> but another tag, so the code can just keep peeling the tag until\n> it sees a non-tag and after that it will be living in the world\n> that has only commit->tree->blob hierarchy, and can afford to do\n> the ancestry based solely on \"commit\" and can treat reachability\n> for \"tree\" and \"blob\" as afterthought.  But I think the updated\n> code needs to know that \"link\" needs to be unwrapped and\n> contained \"commit\" needs to be injected back to the ancestry\n> walking machinery.\n\nWell, a simple and dump version (i.e. my current implementation) can\njust do the same for commits as it does for trees: just recursively\ndescend.  Of course this is prohibitive in anything but toy projects.\n\nA better approach is to put all the submodule commits on the pending\nlist and do the normal ancestry walk for them again.  But this would\nalso need all reachable objects from all modules to be known to one\nprocess.\n\nThis could be solved by having one pending list per submodule and\nthen flush all objects before moving to the next submodule, or\njust processing the submodule in a different process.\nBut when the SEEN information is not shared between submodules then\nrev-list could output the same object twice if a blob or tree is\nused by several submodules.  This may not be a problem if all the\ncode which processes rev-list output is idempotent, but I haven't\nlooked into this in detail.\n\nOf course, when rev-list for submodules is already split out there\nis the valid question if it really makes sense to descend into\nsubmodules when doing rev-list.\nNot doing so would natually decouple sub- from supermodule but then\na lot of operations that depend on rev-list (clone, push, pull)\nhave to be heavily modified.\n\nGetting this straight in an efficient way is the next challenge.\n\n-- \nMartin Waitz\n"},{"id":"294333","messageId":"20061217114546.GG12411@admingilde.org","threadId":"43100","inReplyTo":"200612170101.09615.Josef.Weidendorfer@gmx.de","subject":"Re: Subprojects tasks","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-17T11:45:46Z","receivedAt":"2006-12-17T11:45:46Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Sun, Dec 17, 2006 at 01:01:09AM +0100, Josef Weidendorfer wrote:\n> IMHO it simply is added flexibility to allow a checkout to be separate from\n> the .git/ directory, same as explicitly setting $GIT_DIR would do.\n> So this .gitlink file is on the one hand one kind of convenience for users\n> which want to keep their repository separate, yet do not want to specify\n> $GIT_DIR all the time in front of git commands.\n> The .gitlink file simply makes the linkage to the separate repository\n> persistent.\n\nI can see the reason for wanting to use another object database,\nbut HEAD and index should always be stored together with the\nchecked out directory.  So perhaps we just need some smart way to\nsearch for the object database, but keep the .git directory.\n\n-- \nMartin Waitz\n"},{"id":"297118","messageId":"200612171401.10585.jnareb@gmail.com","threadId":"43100","inReplyTo":"20061217114546.GG12411@admingilde.org","subject":"Re: Subprojects tasks","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-17T13:01:09Z","receivedAt":"2006-12-17T13:01:09Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Martin Waitz wrote:\n> On Sun, Dec 17, 2006 at 01:01:09AM +0100, Josef Weidendorfer wrote:\n\n>> IMHO it simply is added flexibility to allow a checkout to be separate from\n>> the .git/ directory, same as explicitly setting $GIT_DIR would do.\n>> So this .gitlink file is on the one hand one kind of convenience for users\n>> which want to keep their repository separate, yet do not want to specify\n>> $GIT_DIR all the time in front of git commands.\n>> The .gitlink file simply makes the linkage to the separate repository\n>> persistent.\n> \n> I can see the reason for wanting to use another object database,\n> but HEAD and index should always be stored together with the\n> checked out directory.  So perhaps we just need some smart way to\n> search for the object database, but keep the .git directory.\n\nWell, in the .gitlink proposal you could specify GIT_DIR for checkout,\nor separately: GIT_OBJECT_DIRECTORY, GIT_INDEX_FILE, GIT_REFS_DIRECTORY\n(does not exist yet), GIT_HEAD_FILE (does not exist yet, and I suppose\nit wouldn't be easy to implement it). By the way, that's why I'm for\n.gitlink name for the file, not .git -- this way .gitlink can \"shadow\"\nwhat's in .git, for example specifying in a smart way where to search\n(where to find) object database, but HEAD and index would be stored\ntogether with the checked out directory in .git\n\nBy the way, I'm rather partial to supermodule following HEAD in submodule,\nnot specified branch. First, I think it is easier from implementation\npoint of view: you don't have to remember which branch supermodule should\ntake submodule commits from; and this cannot be fixed branch name like\n'master'. For example 'maint' branch of supermodule could track 'maint'\nbranch of submodule, 'master' branch of supermodule track 'master'\nbranch of submodule, 'next' branch of supermodule tranck 'master' (!)\nbranch of submodule, 'pu' branch of supermodule track 'next' (!) branch\nof submodule. \n\nSecond, if you want to do some independent work on the module not related\nto work on submodule you should really clone (clone -l -s) submodule\nand work in separate checkout; the complaint that with tracking HEAD\nyou can check-in wrong version of submodule to supermodule commit\ndoesn't hold, because you still would have problem that _tree_\nof supermodule would have wrong version of submodule. And moving to\nusing single defined branch of submodule brings multitude of other\nproblems: for example you might usually track 'master' version of\nsubmodule, but for a short time need to track 'next' branch because\nit has functionality you need; and another time you need to move\nto 'maint' branch or even your own branch because 'master' version\nbreaks something in supermodule.\n\nHmmm... I wonder how planned allowing to checking out tags, non-head\nbranches (e.g. tracking/remote branches) and arbitrary commits but\nforbidding committing when HEAD is not a refs/heads/ branch would\naffect submodules / subprojects...\n\n-- \nJakub Narebski\n"},{"id":"297621","messageId":"20061217134848.GH12411@admingilde.org","threadId":"43100","inReplyTo":"200612171401.10585.jnareb@gmail.com","subject":"Re: Subprojects tasks","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-17T13:48:48Z","receivedAt":"2006-12-17T13:48:48Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Sun, Dec 17, 2006 at 02:01:09PM +0100, Jakub Narebski wrote:\n> Well, in the .gitlink proposal you could specify GIT_DIR for checkout,\n> or separately: GIT_OBJECT_DIRECTORY, GIT_INDEX_FILE, GIT_REFS_DIRECTORY\n> (does not exist yet), GIT_HEAD_FILE (does not exist yet, and I suppose\n> it wouldn't be easy to implement it). By the way, that's why I'm for\n> .gitlink name for the file, not .git -- this way .gitlink can \"shadow\"\n> what's in .git, for example specifying in a smart way where to search\n> (where to find) object database, but HEAD and index would be stored\n> together with the checked out directory in .git\n\nWhat about .git/link or something?\n(Obviously without the capability to change GIT_DIR)\n\n> By the way, I'm rather partial to supermodule following HEAD in submodule,\n> not specified branch. First, I think it is easier from implementation\n> point of view: you don't have to remember which branch supermodule should\n> take submodule commits from; and this cannot be fixed branch name like\n> 'master'. For example 'maint' branch of supermodule could track 'maint'\n> branch of submodule, 'master' branch of supermodule track 'master'\n> branch of submodule, 'next' branch of supermodule tranck 'master' (!)\n> branch of submodule, 'pu' branch of supermodule track 'next' (!) branch\n> of submodule. \n\nThe version tracked by the supermodule is completely independent from\nany branches you define in your submodule.\nIt is of course possible to use different versions of your submodule in\ndifferent branches of your supermodule.  But the supermodule does not\nknow the name of these branches.\n\nIn the setup you described a git-checkout in the supermodule would have\nto switch to a different branch in the submodule, depending on the\nbranchname which would have to be stored in the supermodule.\nThis a lot more complex.\n\nYour scenario can also be solved in this way:\n\n\tcd supermodule\n\t(cd sub && git-reset --hard origin/master)\n\tgit add sub && git commit -m \"track master of sub\"\n\tgit checkout next\n\t(cd sub && git-reset --hard origin/master)\n\tgit add sub && git commit -m \"track master of sub\"\n\tgit checkout pu\n\t(cd sub && git-reset --hard origin/next)\n\tgit add sub && git commit -m \"track next of sub\"\n\tgit checkout maint\n\t(cd sub && git-reset --hard origin/maint)\n\tgit add sub && git commit -m \"track maint of sub\"\n\nYou only store a link to the commit of the current submodule version,\njust like a normal ref.  The reference stored in the supermodule really\nis equivalent to a normal ref, just that it is stored and updated\nslightly different to a normal one.\n\nSo whenever you checkout a different version of the supermodule, the\nsubmodule ref automatically gets the correct version.  In the example\nabove, when you checkout supermodules pu, your submodules branch will be\nreset to its origin/next (to be more precise: to the commit which was at\nthe tip of origin/next at the time it was stored in the supermodule).\n\nThe fact that the reference to the current submodule commit does not\nonly exist in the supermodule tree but also as a physical ref in the\nsubmodule is very similiar to normal files: you have one version stored\nin the object database, one in the index and one as a real file in the\nworking directory (and this working file is the equivalent of the\nsubmodule ref which is stored in submodule/.git/refs/whatever)\n\nThe reference in the submodule is just a way to be able to work on\nthe submodule.  Because well, refs are the kind of thing that is changed\nby a commit.  And these submodule commits are exactly the kind of work\nyou want to store in the supermodule.  So the equivalent to a working\nfile is not the HEAD of the submodule, but the ref which gets all\nchanges which are intended for the supermodule.\n\nThe fact that the submodule repository still supports other branches has\nnothing to do with submodule support.  These branches are totally\nindependent from the supermodule.\n\n> Second, if you want to do some independent work on the module not related\n> to work on submodule you should really clone (clone -l -s) submodule\n> and work in separate checkout;\n\nYes.\nBut I really like the possibility to switch one module to a branch which\nis not tracked by the parent, because it perhaps contains some debugging\ncode which is needed to debug some other submodule.  You can't move it\nout because you need the common build infrastructure but you don't want\nto branch the entire toplevel project because you don't want your\ndebugging changes to ever become visible at that level.\n\nSo by switching to a different branch you can effectivly say: this is\ntemporary, not meant for the superproject.\nIf you change your mind later you can always merge the submodule branch\nback to master.\n\n> the complaint that with tracking HEAD you can check-in wrong version\n> of submodule to supermodule commit doesn't hold, because you still\n> would have problem that _tree_ of supermodule would have wrong version\n> of submodule.\n\nSorry, I don't understand you here.\n\n> And moving to using single defined branch of submodule brings\n> multitude of other problems: for example you might usually track\n> 'master' version of submodule, but for a short time need to track\n> 'next' branch because it has functionality you need; and another time\n> you need to move to 'maint' branch or even your own branch because\n> 'master' version breaks something in supermodule.\n\nThat is no problem.\nThe supermodule can track whatever _version_ it wants.  You can set\nit to any version which is available in the repository, including all\nthose well known external branches.\nBut the supermodule itself does not know (and should not know) about\n\"maint\" / \"next\" / whatever branch names in the submodule.\n\n> Hmmm... I wonder how planned allowing to checking out tags, non-head\n> branches (e.g. tracking/remote branches) and arbitrary commits but\n> forbidding committing when HEAD is not a refs/heads/ branch would\n> affect submodules / subprojects...\n\nIt only affects submodules if you really track HEAD directly.\n\n-- \nMartin Waitz\n"},{"id":"295374","messageId":"200612171529.03165.jnareb@gmail.com","threadId":"43100","inReplyTo":"20061217134848.GH12411@admingilde.org","subject":"Re: Subprojects tasks","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-17T14:29:02Z","receivedAt":"2006-12-17T14:29:02Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Martin Waitz wrote:\n> On Sun, Dec 17, 2006 at 02:01:09PM +0100, Jakub Narebski wrote:\n\n>> Well, in the .gitlink proposal you could specify GIT_DIR for checkout,\n>> or separately: GIT_OBJECT_DIRECTORY, GIT_INDEX_FILE, GIT_REFS_DIRECTORY\n>> (does not exist yet), GIT_HEAD_FILE (does not exist yet, and I suppose\n>> it wouldn't be easy to implement it). By the way, that's why I'm for\n>> .gitlink name for the file, not .git -- this way .gitlink can \"shadow\"\n>> what's in .git, for example specifying in a smart way where to search\n>> (where to find) object database, but HEAD and index would be stored\n>> together with the checked out directory in .git\n> \n> What about .git/link or something?\n> (Obviously without the capability to change GIT_DIR)\n\nWell, the .gitlink proposal at it is now (by Josef) serves both as a way\nto implement lightweight checkout (i.e. having additional working dir to\nsome repository, or having working dir separate from bare repository),\nand as a way to have \"smart\" submodules (which you can move and rename)\nin submodules/subproject support.\n\nBesides, I'd rather either use config file for this (core.link or\ncore.git_dir), or use .git/GIT_DIR.\n \n>> By the way, I'm rather partial to supermodule following HEAD in submodule,\n>> not specified branch. First, I think it is easier from implementation\n>> point of view: you don't have to remember which branch supermodule should\n>> take submodule commits from; and this cannot be fixed branch name like\n>> 'master'. \n[...]\n> In the setup you described a git-checkout in the supermodule would have\n> to switch to a different branch in the submodule, depending on the\n> branchname which would have to be stored in the supermodule.\n> This a lot more complex.\n\nO.K. Now I understand why you prefer specified branch to HEAD.\nI have forgot that checkout must update submodule ref, and if we track HEAD\nwe would have to remember the branch it pointed to.\n\nBy the way, should this ref be in submodule, or in supermodule, e.g. in\nrefs/modules/<name>/HEAD? And there is a problam _what_ branch should\nbe that.\n\nBoth approaches have advantages and disadvantages...\n-- \nJakub Narebski\n"},{"id":"296496","messageId":"20061217195417.GI12411@admingilde.org","threadId":"43100","inReplyTo":"200612171529.03165.jnareb@gmail.com","subject":"Re: Subprojects tasks","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-17T19:54:17Z","receivedAt":"2006-12-17T19:54:17Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Sun, Dec 17, 2006 at 03:29:02PM +0100, Jakub Narebski wrote:\n> By the way, should this ref be in submodule, or in supermodule, e.g. in\n> refs/modules/<name>/HEAD? And there is a problam _what_ branch should\n> be that.\n\nAt the moment I simply use refs/heads/master of the submodule\nrepository, just because it is the default branch anyway.\n\nIn order to make the submodule refs which are not added to the\nsupermodule available to the supermodule anyway (for fsck and prune),\nI added a symlink .git/refs/module/<submodule> -> <submodule>/.git/refs,\nso that the submodule branch is also available as\nrefs/module/<submodule>/heads/master in the supermodule.\n\nBut I expect that all this setup stuff can be greatly simplified with a\nlittle bit more knowledge of submodules in the core.  But this cleanup\nis for later, when the basis is settled.\n\n-- \nMartin Waitz\n"},{"id":"294271","messageId":"200612180023.45815.Josef.Weidendorfer@gmx.de","threadId":"43100","inReplyTo":"20061217134848.GH12411@admingilde.org","subject":"Re: Subprojects tasks","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-17T23:23:45Z","receivedAt":"2006-12-17T23:23:45Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Sunday 17 December 2006 14:48, Martin Waitz wrote:\n> The version tracked by the supermodule is completely independent from\n> any branches you define in your submodule.\n> It is of course possible to use different versions of your submodule in\n> different branches of your supermodule.  But the supermodule does not\n> know the name of these branches.\n\nI see that you always use \"refs/heads/master\" in the submodule.\nWhat happens if you do development in the submodule, create a new commit\nthere, and want to switch supermodule branch afterwards?\nWouldn't you lose your new work, as \"refs/heads/master\" has to be reset\nto another commit when you switch the supermodule branch?\n\nIMHO it would be nice to have refs in the submodule matching all the\nbranches/tags of the supermodule.\nMeaning: \"this is the commit which is used by branch/tag XYZ in the\nsupermodule\". This can be valuable information, and a \"gitk --all\" in\nthe submodule would show you all the uses of your subproject in the\nscope of the given superproject.\nWe could occupy the local refs namespace of the\nsubmodule with the same refs as there are in the supermodule. But that\nis no problem as the original branches of the subproject would be\nin \"refs/remotes/\".\n\nWhen switching branches in the supermodule, it simply would switch\nto the same name in submodules. The submodule refs would not need\nto match the submodule object in the tree of the supermodule; instead,\nit would represent the development done in the submodule while on a\ngiven branch in the supermodule. Thus, this would allow to do bug fix commits\nfor a submodule at all places where the supermodule has a branch, without\nthe need to switch supermodule branches.\nHowever, \"git commit\" in branch X in the supermodule should give a warning\nwhen submodules are not all at the same branch X, as the commit would use\nbranch X for committing.\n\n> > Second, if you want to do some independent work on the module not related\n> > to work on submodule you should really clone (clone -l -s) submodule\n> > and work in separate checkout;\n> \n> Yes.\n> But I really like the possibility to switch one module to a branch which\n> is not tracked by the parent, because it perhaps contains some debugging\n> code which is needed to debug some other submodule.  You can't move it\n> out because you need the common build infrastructure but you don't want\n> to branch the entire toplevel project because you don't want your\n> debugging changes to ever become visible at that level.\n\nIn general, I agree with not following submodule's HEAD for supermodule\ncommits. As you cannot store any submodule branch names, this really\nwould be confusing, as after switching to another supermodule branch\nand back again, the submodule branch name would reset to a given name\n(\"master\" in your current implementation).\n\nBut why wouldn't you create a temporary branch \"debug_submodule1\" in the\nsupermodule for your use case? Branches are cheap with git, even in supermodules.\nSupermodule branches also are pure local, you never have to publish\nit somewhere, and can delete it afterwards.\n\n"},{"id":"294104","messageId":"200612180027.25308.Josef.Weidendorfer@gmx.de","threadId":"43100","inReplyTo":"20061217195417.GI12411@admingilde.org","subject":"Re: Subprojects tasks","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-17T23:27:25Z","receivedAt":"2006-12-17T23:27:25Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Sunday 17 December 2006 20:54, Martin Waitz wrote:\n> I added a symlink .git/refs/module/<submodule> -> <submodule>/.git/refs,\n> so that the submodule branch is also available as\n> refs/module/<submodule>/heads/master in the supermodule.\n\nAh.\nWhat is \"<submodule>\" in your implementation?\nIs this some encoding of the path where the submodule currently lives\nin the supermodule, or are you giving the submodules unique names\nin the context of the supermodule?\n\n"},{"id":"294106","messageId":"20061218074441.GJ12411@admingilde.org","threadId":"43100","inReplyTo":"200612180023.45815.Josef.Weidendorfer@gmx.de","subject":"Re: Subprojects tasks","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-18T07:44:41Z","receivedAt":"2006-12-18T07:44:41Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"hoi :)\n\nOn Mon, Dec 18, 2006 at 12:23:45AM +0100, Josef Weidendorfer wrote:\n> I see that you always use \"refs/heads/master\" in the submodule.\n> What happens if you do development in the submodule, create a new commit\n> there, and want to switch supermodule branch afterwards?\n> Wouldn't you lose your new work, as \"refs/heads/master\" has to be reset\n> to another commit when you switch the supermodule branch?\n\nIt should behave the same as for files:\nRefuse to update the working directory if files (or the submodule here)\nare dirty.  I guess this is not yet handled correctly by my prototype,\nbut it should not be hard to do.\n\n> IMHO it would be nice to have refs in the submodule matching all the\n> branches/tags of the supermodule.\n> Meaning: \"this is the commit which is used by branch/tag XYZ in the\n> supermodule\". This can be valuable information, and a \"gitk --all\" in\n> the submodule would show you all the uses of your subproject in the\n> scope of the given superproject.\n\nI like the idea.  Perhaps make them available similiar to the remotes\ninformation in refs/tracked/{heads,tags} or something.\n\n> When switching branches in the supermodule, it simply would switch\n> to the same name in submodules.\n\nNice idea, but I don't yet know how it really works out.\nIt may be confusing to the user if he manually switches the branch in\nthe submodule to another branch of the supermodule.  Then he really is\nusing one tracked branch, but not the currently tracked branch.\n\n> Thus, this would allow to do bug fix commits for a submodule at all\n> places where the supermodule has a branch, without the need to switch\n> supermodule branches.\n\nHmm, but when switching to another supermodule branch it would try to\nupdate the submodule branch.\nAnd simply allow the current submodule branch to be a fast forward of\nthe submodule version that the parent wants to set is a bad, as you\nwould not be able to go back to an old supermodule version then.\n\n> > > Second, if you want to do some independent work on the module not related\n> > > to work on submodule you should really clone (clone -l -s) submodule\n> > > and work in separate checkout;\n> > \n> > Yes.\n> > But I really like the possibility to switch one module to a branch which\n> > is not tracked by the parent, because it perhaps contains some debugging\n> > code which is needed to debug some other submodule.  You can't move it\n> > out because you need the common build infrastructure but you don't want\n> > to branch the entire toplevel project because you don't want your\n> > debugging changes to ever become visible at that level.\n> \n> In general, I agree with not following submodule's HEAD for supermodule\n> commits. As you cannot store any submodule branch names, this really\n> would be confusing, as after switching to another supermodule branch\n> and back again, the submodule branch name would reset to a given name\n> (\"master\" in your current implementation).\n> \n> But why wouldn't you create a temporary branch \"debug_submodule1\" in the\n> supermodule for your use case? Branches are cheap with git, even in supermodules.\n> Supermodule branches also are pure local, you never have to publish\n> it somewhere, and can delete it afterwards.\n\nSure, you can of course always use supermodule branches.\nI just wanted to point out that it still is useful to have submodule\nbranches which are independent from supermodule branches.\n\n-- \nMartin Waitz\n"},{"id":"295083","messageId":"20061218074556.GK12411@admingilde.org","threadId":"43100","inReplyTo":"200612180027.25308.Josef.Weidendorfer@gmx.de","subject":"Re: Subprojects tasks","fromName":"Martin Waitz","fromEmail":"tali@admingilde.org","sentAt":"2006-12-18T07:45:56Z","receivedAt":"2006-12-18T07:45:56Z","isPatch":false,"sender":{"key":"tali@admingilde.org","avatar":"https://gravatar.com/avatar/3f89b03eee362187effabe257898735b475673a12265c398ea9161259ae91553?d=mp&s=160"},"body":"On Mon, Dec 18, 2006 at 12:27:25AM +0100, Josef Weidendorfer wrote:\n> On Sunday 17 December 2006 20:54, Martin Waitz wrote:\n> > I added a symlink .git/refs/module/<submodule> -> <submodule>/.git/refs,\n> > so that the submodule branch is also available as\n> > refs/module/<submodule>/heads/master in the supermodule.\n> \n> Ah.\n> What is \"<submodule>\" in your implementation?\n> Is this some encoding of the path where the submodule currently lives\n> in the supermodule, or are you giving the submodules unique names\n> in the context of the supermodule?\n\nAt the moment, it's just the path inside the parent.\n\n-- \nMartin Waitz\n"},{"id":"295933","messageId":"200612181130.18652.Josef.Weidendorfer@gmx.de","threadId":"43100","inReplyTo":"20061218074441.GJ12411@admingilde.org","subject":"Re: Subprojects tasks","fromName":"Josef Weidendorfer","fromEmail":"josef.weidendorfer@gmx.de","sentAt":"2006-12-18T10:30:18Z","receivedAt":"2006-12-18T10:30:18Z","isPatch":false,"sender":{"key":"josef.weidendorfer@gmx.de","avatar":null},"body":"On Monday 18 December 2006 08:44, Martin Waitz wrote:\n> On Mon, Dec 18, 2006 at 12:23:45AM +0100, Josef Weidendorfer wrote:\n> > I see that you always use \"refs/heads/master\" in the submodule.\n> > What happens if you do development in the submodule, create a new commit\n> > there, and want to switch supermodule branch afterwards?\n> > Wouldn't you lose your new work, as \"refs/heads/master\" has to be reset\n> > to another commit when you switch the supermodule branch?\n> \n> It should behave the same as for files:\n> Refuse to update the working directory if files (or the submodule here)\n> are dirty.  I guess this is not yet handled correctly by my prototype,\n> but it should not be hard to do.\n\nAh, I see.\nYes, this is consistent with other dirty files.\n\n\n> > IMHO it would be nice to have refs in the submodule matching all the\n> > branches/tags of the supermodule.\n> > Meaning: \"this is the commit which is used by branch/tag XYZ in the\n> > supermodule\". This can be valuable information, and a \"gitk --all\" in\n> > the submodule would show you all the uses of your subproject in the\n> > scope of the given superproject.\n> \n> I like the idea.  Perhaps make them available similiar to the remotes\n> information in refs/tracked/{heads,tags} or something.\n\nYes.\nHowever, you want to do development on these branches. And\n\"refs/tracked/...\" is read-only. However, taking the whole local\nrefs namespace is not good, as you perhaps want branches independent\nof the supermodule, which could give name conflicts.\nWhat about using \"refs/{heads,tags}/supermodule/...\"? This could\nbe a compromise.\n\n> > When switching branches in the supermodule, it simply would switch\n> > to the same name in submodules.\n> \n> Nice idea, but I don't yet know how it really works out.\n> It may be confusing to the user if he manually switches the branch in\n> the submodule to another branch of the supermodule.  Then he really is\n> using one tracked branch, but not the currently tracked branch.\n\nBut you already have the same problem with your current approach, don't\nyou?\n\nActually, the most expected thing for the user really would be to use\nHEAD in supermodule commits. Every other behavior can get confusing for\nthe user: (S)He simply expects the state of the checkout to be committed.\nAny branch switching in submodules should be temporary.\n\nActually, you can be on a temporary branch in a submodule and still switch\nbranches in the supermodule. It is the same as with dirty files: The\nmodifications can be carried over to other branches and back, as long as\nthere are no conflicts. \n\nHowever, I think it is important to check that you are back on the\nright branch when committing. With warning or even error.\n\n"}]}