{"thread":{"id":"3131","subject":"Notes on Subproject Support","startedAt":"2006-01-23T01:35:14Z","lastAt":"2006-01-24T04:22:41Z","messageCount":15,"participants":["Junio C Hamano","Daniel Barkalow","Alexander Litvinov","Martin Atukunda"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"15057","messageId":"7v3bjfafql.fsf@assigned-by-dhcp.cox.net","threadId":"3131","inReplyTo":null,"subject":"Notes on Subproject Support","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-01-23T01:35:14Z","receivedAt":"2006-01-23T01:35:14Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"This is still a draft/WIP, but \"release early\" is a good\ndiscipline, so...\n\n-- >8 --\n\nNotes on Subproject Support\n===========================\nJunio C Hamano <junkio@cox.net>\nv1.0 January 22, 2006\n\nScenario\n--------\n\nThe examples in the following discussion show how this proposal\nplans to help this:\n\n. A project to build an embedded Linux appliance \"gadget\" is\n  maintained with git.\n\n. The project uses linux-2.6 kernel as its subcomponent.  It\n  starts from a particular version of the mainline kernel, but\n  adds its own code and build infrastructure to fit the\n  appliance's needs.\n\n. The working tree of the project is laid out this way:\n+\n------------\n Makefile       - Builds the whole thing.\n linux-2.6/     - The kernel, perhaps modified for the project.\n appliance/     - Applications that run on the appliance, and\n                  other bits.\n------------\n\n. The project is willing to maintain its own changes out of tree\n  of the Linux kernel project, but would want to be able to feed\n  the changes upstream, and incorporate upstream changes to its\n  own tree, taking advantage of the fact that both itself and\n  the Linux kernel project are version controlled with git.\n\nThe idea here is to:\n\n. Keep `linux-2.6/` part as an independent project.  The work by\n  the project on the kernel part can be naturally exchanged with\n  the other kernel developers this way.  Specifically, a tree\n  object contained in commit objects belonging to this project\n  does *not* have linux-2.6/ directory at the top.\n\n. Keep the `appliance/` part as another independent project.\n  Applications are supposed to be more or less independent from\n  the kernel version, but some other bits might be tied to a\n  specific kernel version.  Again, a tree object contained in\n  commit objects belonging to this project does *not* have\n  appliance/ directory at the top.\n\n. Have another project that combines the whole thing together,\n  so that the project can keep track of which versions of the\n  parts are built together.\n\nWe will call the project that binds things together the\n'toplevel project'.  Other projects that hold `linux-2.6/` part\nand `appliance/` part are called 'subprojects'.\n\nNotice that `Makefile` at the top is part of the toplevel\nproject in this example, but it is not necessary.  We could\ninstead have the appliance subproject include this file.  In\nsuch a setup, the appliance subproject would have had `Makefile`\nand `appliance/` directory at the toplevel.\n\n\nSetting up\n----------\n\nLet's say we have been working on the appliance software,\nindependently version controlled with git.  Also the kernel part\nhas been version controlled separately, like this:\n------------\n$ ls -dF current/*/.git current/*\ncurrent/Makefile    current/appliance/.git/  current/linux-2.6/.git/\ncurrent/appliance/  current/linux-2.6/\n------------\n\nNow we would want to get a combined project.  First we would\nclone from these repositories (which is not strictly needed --\nwe could use `$GIT_ALTERNATE_OBJECT_DIRECTORIES` instead):\n\n------------\n$ mkdir combined && cd combined\n$ cp ../current/Makefile .\n$ git init-db\n$ mkdir -p .git/refs/subs/{kernel,gadget}/{heads,tags}\n$ git clone-pack ../current/linux-2.6/ master | read kernel_commit junk\n$ git clone-pack ../current/appliance/ master | read gadget_commit junk\n------------\n\nWe will introduce a new command to set up a combined project:\n\n------------\n$ git bind-projects \\\n\t$kernel_commit linux-2.6/ \\\n\t$gadget_commit appliance/\n------------\n\nThis would do an equivalent of:\n\n------------\n$ git read-tree --prefix=linux-2.6/ $kernel_commit\n$ git read-tree --prefix=appliance/ $gadget_commit\n------------\n[NOTE]\n============\nEarlier outlines sent to the git mailing list talked\nabout `$GIT_DIR/bind` to record what subproject are bound to\nwhich subtree in the curent working tree and index.  This\nproposal instead records that information in the index file\nwhen `--prefix=linux-2.6/` is given to `read-tree`.\n\nAlso note that in this round of proposal, there is no separate\nbranches that keep track of heads of subprojects.\n============\n\nLet's not forget to add the `Makefile`, and check the whole\nthing out from the index file.\n------------\n$ git add Makefile\n$ git checkout-index -f -u -q -a\n------------\n\nNow our directory should be identical with the `current`\ndirectory.  After making sure of that, we should be able to\ncommit the whole thing:\n\n------------\n$ diff -x .git -r ../current ../combined\n$ git commit -m 'Initial toplevel project commit'\n------------\n\nWhich should create a new commit object that records what is in\nthe index file as its tree, with `bind` lines to record which\nsubproject commit objects are bound at what subdirectory, and\nupdates the `$GIT_DIR/refs/heads/master`.  Such a commit object\nmight look like this:\n------------\ntree 04803b09c300c8325258ccf2744115acc4c57067\nbind 5b2bcc7b2d546c636f79490655b3347acc91d17f linux-2.6/\nbind 0bdd79af62e8621359af08f0afca0ce977348ac7 appliance/\nauthor Junio C Hamano <junio@kernel.org> 1137965565 -0800\ncommitter Junio C Hamano <junio@kernel.org> 1137965565 -0800\n\nInitial toplevel project commit\n------------\n\n\nMaking further commits\n----------------------\n\nThe easiest case is when you updated the Makefile without\nchanging anything in the subprojects.  In such a case, we just\nneed to create a new commmit object that records the new tree\nwith the current `HEAD` as its parent, and with the same set of\n`bind` lines.\n\nWhen we have changes to the subproject part, we would make a\nseparate commit to the subproject part and then record the whole\nthing by making a commit to the toplevel project.  The user\ninteraction might go this way:\n------------\n$ git commit\nerror: you have changes to the subproject bound at linux-2.6/.\n$ git commit --subproject linux-2.6/\n$ git commit\n------------\n\nWith the new `--subproject` option, the directory structure\nrooted at `linux-2.6/` part is written out as a tree, and a new\ncommit object that records that tree object with the commit\nbound to that portion of the tree (`5b2bcc7b` in the above\nexample) as its parent is created.  Then the final `git commit`\nwould record the whole tree with updated `bind` line for the\n`linux-2.6/` part.\n\n\nChecking out\n------------\n\nAfter cloning such a toplevel project, `git clone` without `-n`\noption would check out the working tree.  This is done by\nreading the tree object recorded in the commit object (which\nrecords the whole thing), and adding the information from the\n\"bind\" line to the index file.\n\n------------\n$ cd ..\n$ git clone -n combined cloned ;# clone the one we created earlier\n$ cd cloned\n$ git checkout\n------------\n\nThis round of proposal does not maintain separate branch heads\nfor subprojects.  The bound commits and their subdirectories\nare recorded in the index file from the commit object, so there\nis no need to do anything other than updating the index and the\nworking tree.\n\n\nSwitching branches\n------------------\n\nAlong with the traditional two-way merge by `read-tree -m -u`,\nwe would need to look at:\n\n. `bind` lines in the current `HEAD` commit.\n\n. `bind` lines in the commit we are switching to.\n\n. subproject binding information in the index file.\n\nto make sure we do sensible things.\n\nJust like until very recently we did not allow switching\nbranches when two-way merge would lose local changes, we can\nstart by refusing to switch branches when the subprojects bound\nin the index do not match what is recorded in the `HEAD` commit.\n\nBecause in this round of the proposal we do not use the\n`$GIT_DIR/bind` file nor separate branches to keep track of\nheads of the subprojects, there is nothing else other than the\nworking tree and the index file that needs to be updated when\nswitching branches.\n\n\nMerging\n-------\n\nMerging two branches of the toplevel projects can use the\ntraditional merging mechanism mostly unchanged.  The merge base\ncomputation can be done using the `parent` ancestry information\ntaken from the two toplevel project branch heads being merged,\nand merging of the whole tree can be done with a three-way merge\nof the whole tree using the merge base and two head commits.\nFor reasons described later, we would not merge the subproject\nparts of the trees during this step, though.\n\nWhen the two branch heads use different versions of subproject,\nthings get a bit tricky.  First, let's forget for a moment about\nthe case where they bind the same project at different location.\nWe would refuse if they do not have the same number of `bind`\nlines that bind something at the same subdirectories.\n\n------------\n$ git merge 'Merge in a side branch' HEAD side\nerror: the merged heads have subprojects bound at different places.\n ours:\n\tlinux-2.6/\n\tappliance/\n theirs:\n\tkernel/\n\tgadget/\n\tmanual/\n------------\n\nSuch renaming can be handled by first moving the bind points in\nour branch, and redoing the merge (this is a rare operation\nanyway).  It might go like this:\n\n------------\n$ git bind-projects \\\n\t$kernel_commit kernel/ \\\n\t$gadget_commit gadget/\n$ git commit -m 'Prepare for merge with side branch'\n$ git merge 'Merge in a side branch' HEAD side\nerror: the merged heads have subprojects bound at different places.\n ours:\n\tkernel/\n\tgadget/\n theirs:\n\tkernel/\n\tgadget/\n\tmanual/\n------------\n\nTheir branch added another subproject, so this did not work (or\nit could be the other way around -- we might have been the one\nwith `manual/` subproject while they didn't).  This suggests\nthat we may want an option to `git merge` to allow taking a\nunion of subprojects.  Again, this is a rare operation, and\nalways taking a union would have created a toplevel project that\nhad both `kernel/` and `linux-2.6/` bound to the same Linux\nkernel project from possibly different vintage, so it would be\nprudent to require the set of bound subprojects to exactly match\nand give the user an option to take a union.\n\n------------\n$ git merge --union-subprojects 'Merge in a side branch HEAD side\nerror: the subproject at `kernel/` needs to be merged first.\n------------\n\nHere, the version of the Linux kernel project in the `side`\nbranch was different from what our branch had on our `bind`\nline.  On what kind of difference should we give this error?\nInitially, I think we could require one is the fast forward of\nthe other (ours might be ahead of theirs, or the other way\naround), and take the descendant.\n\nOr we could do an independent merge of subprojects heads, using\nthe `parent` ancestry of the bound subproject heads to find\ntheir merge-base and doing a three-way merge.  This would leave\nthe merge result in the subproject part of the working tree and\nthe index.\n\n[NOTE]\nThis is the reason we did not do the whole-tree three way merge\nearlier.  The subproject commit bound to the merge base commit\nused for the toplevel project may not be the merge base between\nthe subproject commits bound to the two toplevel project\ncommits.\n\nSo let's deal with the case to merge only a subproject part into\nour tree first.\n\n\nMerging subprojects\n-------------------\n\nAn operation of more practical importance is to be able to merge\nin changes done outside to the projects bound to our toplevel\nproject.\n\n------------\n$ git pull --subproject=kernel/ git://git.kernel.org/.../linux-2.6/\n------------\n\nmight do:\n\n. fetch the current `HEAD` commit from Linus.\n. find the subproject commit bound at kernel/ subtree.\n. perform the usual three-way merge of these two commits, in\n  `kernel/` part of the working tree.\n\nAfter that, `git commit \\--subproject` option would be needed to\nmake a commit.\n\n[NOTE]\nThis suggests that we would need to have something similar to\n`MERGE_HEAD` for merging the subproject part.\n"},{"id":"15059","messageId":"Pine.LNX.4.64.0601222104120.25300@iabervon.org","threadId":"3131","inReplyTo":"7v3bjfafql.fsf@assigned-by-dhcp.cox.net","subject":"Re: Notes on Subproject Support","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2006-01-23T03:50:32Z","receivedAt":"2006-01-23T03:50:32Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Sun, 22 Jan 2006, Junio C Hamano wrote:\n\n> Also note that in this round of proposal, there is no separate\n> branches that keep track of heads of subprojects.\n\nInteresting; I think it may become useful to allow for such heads, but we \ncan deal with that when it arises. (e.g., maybe you want to use topic \nbranches in the kernel development you do in the linux-2.6/ subdirectory \nof your superproject working tree; so long as the core isn't using refs \nfor its own purposes, this is up to the user to keep straight and we can \nhelp later when we have usage notes)\n\n> ============\n> \n> Let's not forget to add the `Makefile`, and check the whole\n> thing out from the index file.\n> ------------\n> $ git add Makefile\n\nMaybe bind-projects should be \"add-projects\", to match \"add\", which has a \nsimilar effect at the user level?\n\n> $ git checkout-index -f -u -q -a\n> ------------\n> \n> Now our directory should be identical with the `current`\n> directory.  After making sure of that, we should be able to\n> commit the whole thing:\n> \n> ------------\n> $ diff -x .git -r ../current ../combined\n> $ git commit -m 'Initial toplevel project commit'\n> ------------\n> \n> Which should create a new commit object that records what is in\n> the index file as its tree, with `bind` lines to record which\n> subproject commit objects are bound at what subdirectory, and\n> updates the `$GIT_DIR/refs/heads/master`.  Such a commit object\n> might look like this:\n> ------------\n> tree 04803b09c300c8325258ccf2744115acc4c57067\n\nDoes this tree include trees for the bound projects?\n\n> bind 5b2bcc7b2d546c636f79490655b3347acc91d17f linux-2.6/\n> bind 0bdd79af62e8621359af08f0afca0ce977348ac7 appliance/\n> author Junio C Hamano <junio@kernel.org> 1137965565 -0800\n> committer Junio C Hamano <junio@kernel.org> 1137965565 -0800\n> \n> Initial toplevel project commit\n> ------------\n> \n> \n> Making further commits\n> ----------------------\n> \n> The easiest case is when you updated the Makefile without\n> changing anything in the subprojects.  In such a case, we just\n> need to create a new commmit object that records the new tree\n> with the current `HEAD` as its parent, and with the same set of\n> `bind` lines.\n> \n> When we have changes to the subproject part, we would make a\n> separate commit to the subproject part and then record the whole\n> thing by making a commit to the toplevel project.  The user\n> interaction might go this way:\n> ------------\n> $ git commit\n> error: you have changes to the subproject bound at linux-2.6/.\n> $ git commit --subproject linux-2.6/\n> $ git commit\n> ------------\n\nI think \"cd linux-2.6 && git commit\" should work for the subproject, too, \nbut that can be a later enhancement.\n\n> With the new `--subproject` option, the directory structure\n> rooted at `linux-2.6/` part is written out as a tree, and a new\n> commit object that records that tree object with the commit\n> bound to that portion of the tree (`5b2bcc7b` in the above\n> example) as its parent is created.\n\nAnd the commit is written to the index, in the special slot for the \nsubproject, replacing its parent, I assume.\n\n> Switching branches\n> ------------------\n> \n> Along with the traditional two-way merge by `read-tree -m -u`,\n> we would need to look at:\n> \n> . `bind` lines in the current `HEAD` commit.\n> \n> . `bind` lines in the commit we are switching to.\n> \n> . subproject binding information in the index file.\n> \n> to make sure we do sensible things.\n\nThis is one place I think storing the bindings in the commit is awkward. \nread-tree deals in trees (hence the name), but will need information from \nthe commit.\n\nI think it should be possible to hide the existance of subtrees in an \nadd-on to the struct tree API such that code that doesn't handle it \nspecifically doesn't see a difference, similarly to how the index file can \nbe handled. (parse_tree would fill out the structure as if the subproject \nwere a tree instead of a commit, assuming that the structure it's \npretending to be is the full tree, but there would be an additional \nfield for the commit if it's a subproject, until we've gone through \neverything to make it work with subprojects).\n\nI'm hoping to kill off the other tree object parser, which is only used by \nls-tree and diff-index at this point, but my workstation's home directory \nhard drive seems to have gotten weirdly messed up at the hardware level \n(and seems to have lost a lot of the contents of unused storage, or \nsomething), so this may take a little while. At that point, whatever \nspecial things we do in tree objects can be handled automatically with \nchanges only to a single location.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"15060","messageId":"7v7j8r7e7s.fsf@assigned-by-dhcp.cox.net","threadId":"3131","inReplyTo":"Pine.LNX.4.64.0601222104120.25300@iabervon.org","subject":"Re: Notes on Subproject Support","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-01-23T04:36:23Z","receivedAt":"2006-01-23T04:36:23Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Daniel Barkalow <barkalow@iabervon.org> writes:\n\n>> tree 04803b09c300c8325258ccf2744115acc4c57067\n>\n> Does this tree include trees for the bound projects?\n\nYes, this part has not been changed from earlier thoughts.\n\n>> bind 5b2bcc7b2d546c636f79490655b3347acc91d17f linux-2.6/\n>> bind 0bdd79af62e8621359af08f0afca0ce977348ac7 appliance/\n>> author Junio C Hamano <junio@kernel.org> 1137965565 -0800\n>> committer Junio C Hamano <junio@kernel.org> 1137965565 -0800\n\nThe tree 04803b...  tree has everything.  If you run git-ls-tree\non 04803b... would have a tree object recorded at linux-2.6, and\nit is the same as the tree associated with the commit 5b2bcc...\n\n> I think \"cd linux-2.6 && git commit\" should work for the subproject, too, \n> but that can be a later enhancement.\n\nIt's just a matter of Porcelain scripting, so that is probably\ntrue.  However I do not want people to get too used to it and\nexpect \"cd Documentation && git commit\" to work in git.git\nrepository.\n\n>> With the new `--subproject` option, the directory structure\n>> rooted at `linux-2.6/` part is written out as a tree, and a new\n>> commit object that records that tree object with the commit\n>> bound to that portion of the tree (`5b2bcc7b` in the above\n>> example) as its parent is created.\n>\n> And the commit is written to the index, in the special slot for the \n> subproject, replacing its parent, I assume.\n\nYes.  It would probably be done with `update-index --bind` to\nupdate the bound subproject commit there.\n\n>> Switching branches\n>> ------------------\n>> \n>> Along with the traditional two-way merge by `read-tree -m -u`,\n>> we would need to look at:\n>> \n>> . `bind` lines in the current `HEAD` commit.\n>> \n>> . `bind` lines in the commit we are switching to.\n>> \n>> . subproject binding information in the index file.\n>> \n>> to make sure we do sensible things.\n>\n> This is one place I think storing the bindings in the commit is awkward. \n> read-tree deals in trees (hence the name), but will need information from \n> the commit.\n\nThat's why it is 'along with'.  Dealing with binding information\ncan be done between commits and index without bothering tree\nobjects.  read-tree would not have to deal with it, and I think\nkeeping it that way is probably a good idea.\n\nIn other words, I think the design so far does not require us to\ntouch tree objects at all, and I'd be happy if we do not have to.\n\nOne reason I started the bound commit approach was exactly\nbecause I only needed to muck with commit objects and did not\nhave to touch trees and blobs; after trying to implement the\ncore level for \"gitlink\", which I ended up touching quite a lot\nand have abandoned for now.\n\nHere is an update to the still-WIP draft.\n\n-- >8 --\nSeparate role of read-tree and update-index cleaner\n\nThe previous draft prematurely merged what read-tree --prefix\ndoes with what update-index --bind would do.  Keep them separate\nfor now until we know what the common patterns would be.\n\nIntroduce 'update-index --unbind'.  We would probably need a new\ncommand that extracts bind information out of index when we\nstart writing Porcelainish, but it is not specified yet.\n\nAttempt to clarify what the \"merging into subproject part\" would\ndo a bit.  \"git pull --subproject=\" is fetch + merge, just like\nthe current subproject-unaware 'git pull' is.\n\n---\ndiff --git a/Subpro.txt b/Subpro.txt\nindex 4036e71..837cab8 100644\n--- a/Subpro.txt\n+++ b/Subpro.txt\n@@ -95,19 +95,22 @@ $ git bind-projects \\\n \t$gadget_commit appliance/\n ------------\n \n-This would do an equivalent of:\n+This would probably do an equivalent of:\n \n ------------\n+$ rm -f \"$GIT_DIR/index\"\n $ git read-tree --prefix=linux-2.6/ $kernel_commit\n $ git read-tree --prefix=appliance/ $gadget_commit\n+$ git update-index --bind linux-2.6/ $kernel_commit\n+$ git update-index --bind appliance/ $gadget_commit\n ------------\n [NOTE]\n ============\n Earlier outlines sent to the git mailing list talked\n about `$GIT_DIR/bind` to record what subproject are bound to\n-which subtree in the curent working tree and index.  This\n+which subtree in the current working tree and index.  This\n proposal instead records that information in the index file\n-when `--prefix=linux-2.6/` is given to `read-tree`.\n+with `update-index --bind` command.\n \n Also note that in this round of proposal, there is no separate\n branches that keep track of heads of subprojects.\n@@ -258,9 +261,11 @@ our branch, and redoing the merge (this \n anyway).  It might go like this:\n \n ------------\n-$ git bind-projects \\\n-\t$kernel_commit kernel/ \\\n-\t$gadget_commit gadget/\n+$ git reset\n+$ git update-index --unbind linux-2.6/\n+$ git update-index --unbind appliance/\n+$ git update-index --bind $kernel_commit kernel/\n+$ git update-index --bind $gadget_commit gadget/\n $ git commit -m 'Prepare for merge with side branch'\n $ git merge 'Merge in a side branch' HEAD side\n error: the merged heads have subprojects bound at different places.\n@@ -336,7 +341,55 @@ make a commit.\n \n [NOTE]\n This suggests that we would need to have something similar to\n-`MERGE_HEAD` for merging the subproject part.\n+`MERGE_HEAD` for merging the subproject part.  In the case of\n+merging two toplevel project commits, we probably can read the\n+`bind` lines from the `MERGE_HEAD` commit and either our `HEAD`\n+commit or our index file.  Further, we probably would require\n+that the latter two must match, just as we currently require the\n+index file matches our `HEAD` commit before `git merge`.\n \n+Just like the current `pull = fetch + merge` semantics, the\n+subproject aware version `git pull \\--subproject=frotz` would be\n+a `git fetch \\--subproject=frotz` followed by a `git merge\n+\\--subproject=frotz`.  So the above would be:\n \n+. Fetch the head.\n++\n+------------\n+$ git fetch --subproject=kernel/ git://git.kernel.org/.../linux-2.6/\n+------------\n++\n+which would do:\n+. fetch the commit chain from the remote repository.\n+. write something like this to `FETCH_HEAD`:\n++\n+------------\n+3ee68c4...\\tfor-merge-into kernel/\\tbranch 'master' of git://.../linux-2.6\n+------------\n+\n+. Run `git merge`.\n++\n+------------\n+$ git merge --subproject=kernel/ \\\n+    'Merge git://.../linux-2.6 into kernel/' HEAD 3ee68c4...\n+------------\n+\n+. In case it does not cleanly automerge, `git merge` would write\n+the necessary information for a later `git commit` to use in\n+`MERGE_HEAD`.  It may look like this:\n++\n+------------\n+3ee68c4af3fd7228c1be63254b9f884614f9ebb2\tkernel/\n+------------\n+\n+With this, a later invocation of `git commit` to record the\n+result of hand resolving would be able to notice that:\n+\n+. We should be first resolving `kernel/` subproject.\n+. The remote `HEAD` is `3ee68c4...` commit.\n+. The merge message is `Merge git://.../linux-2.6 into kernel/`.\n \n+and make a merge commit, and register that resulting commit in\n+the index file using `update-index --bind` instead of updating\n+*any* branch head (remember, we do not use separate branches to\n+keep track of subproject heads anymore).\n"},{"id":"15062","messageId":"7v64ob1omh.fsf@assigned-by-dhcp.cox.net","threadId":"3131","inReplyTo":"7v7j8r7e7s.fsf@assigned-by-dhcp.cox.net","subject":"Re: Notes on Subproject Support","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-01-23T05:48:06Z","receivedAt":"2006-01-23T05:48:06Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <junkio@cox.net> writes:\n\n> In other words, I think the design so far does not require us to\n> touch tree objects at all, and I'd be happy if we do not have to.\n>\n> One reason I started the bound commit approach was exactly\n> because I only needed to muck with commit objects and did not\n> have to touch trees and blobs; after trying to implement the\n> core level for \"gitlink\", which I ended up touching quite a lot\n> and have abandoned for now.\n\nBTW, let's digress a bit.\n\nI think recording \"commit\" in the tree objects is in line with\nthe logical organization described in README: \"blob\" and \"tree\"\nrepresent a state, and have *nothing* to do with how we came\nabout to that state.  The historyh is described in \"commit\"\nobjects.  The bound commit approach keeps that property.\n\nThe \"gitlink\" approach, as I understand how Linus outlined in\nhis original suggestion, is a bit different.  The link objects\nappear in tree objects, and when you \"git cat-file link\" one of\nthem, you would see something like this:\n\n        commit\t5b2bcc7b2d546c636f79490655b3347acc91d17f\n        name\tkernel\n\nSo in that sense, \"gitlink\" approach departs from the original\npremise of \"commit\" being the only thing that ties things\ntogether.  Tree objects with \"gitlink\" know where they are in\nthe history [*1*].\n\nBy this, I do not mean to say that \"gitlink\" approach is\ninferior because it breaks that original premise.  I am just\npointing it out as a difference between two approaches.\n\nNow, the current way index file is used is as a staging area to\ncreate a new commit on top of the tip of the current branch.\nHowever, it is interesting to note that logically, by itself\n*alone*, it cannot be used that way.  The information the index\nfile records is something that can be used to write out a tree\nobject, and not a commit, because it does not know where the\ncurrent state sits in the history.  We have two separate files,\n$GIT_DIR/HEAD that records which branch we are on, and the\nbranch head ref the HEAD points at, which records where the\ncurrent index came from, for that purpose.  The latter tells us\nwhat commit we should use as the parent commit if we create such\na commit, and the former tells us which branch head to update\nonce we create one.  So in that sense, the index file is just a\nstaging area to create a new tree, not a new commit.\n\nWe could have done things differently.  I am not advocating to\ndo the following change, but offering a possibility as a thought\nexperiment.  It just felt interesting enough to point them out.\n\nThe index file could have recorded what commit the current state\nrecorded in the index came from.  By recording the commit the\nindex was read from in the index itself, independently from the\n$GIT_DIR/refs/heads/$branch file, we could have been able to\nallow fetching into the current branch.  When the $branch file\nfor the current branch was updated by a fast-forward fetch, we\nwould notice that the commit recorded there no longer match what\nis recorded in the index.\n\nAnother interesting consequence is if the development is a\nsingle repository and linear, we did not even need any file in\n$GIT_DIR/refs/ (\"branchless git\").  The commit recorded as the\ntopmost in the index file itself would have served as the tip of\nthe development, and we would have been able to tangle the\nhistory starting from the commit in the index file.\n\nWhile we are doing a thought experiment, let's say we allow to\nrecord more than one commits the current index is based upon.\n'git merge' would record all the parent commits there, so that\nwriting out the merge result out of the index file as a tree and\nthen recording these commits as parents would have been the way\nto create a merge commit.  We would not need the auxiliary file\n$GIT_DIR/MERGE_HEAD if we did so.\n\nIn other words, if the index file recorded the commits its\ncontents were based upon, instead of being a staging area for a\nnew tree, it would have been a staging area for a new commit.\n\nNow, the latest proposal, borrowing your idea, records the\nsubproject commits bound to subdirectories in the index itself.\nThis is halfway to make the index file a staging area for the\nnext commit.  If we were to do that, we also *could* record the\ncommits the current index is based upon, so that it can truly be\nused as a staging area to create a new commit, not just a tree.\n\nOn the other hand, this could be a reason *not* to do the\n`update-index --bind` to record the subproject information in\nthe index file.  An auxiliary file such as $GIT_DIR/bind might\nbe sufficient, just like $GIT_DIR/MERGE_HEAD has been good\nenough for us so far.  One difference between MERGE_HEAD and\nbind is that the former is very transient -- only exists during\na merge while the latter is persistent while the top commit is\nchecked out and being worked on.\n\n\n[Footnote]\n\n*1* One good property of \"gitlink\" approach is that we *could*\nextend this blob-like object to store arbitrary human readable\ninformation to represent a point-in-time from an arbitrary\nforeign SCM system.  IOW, we do not necessarily have to require\n`commit` line that name a git commit to be there.  It could say\n\"Please slurp http://www.kernel.org/pub/software/.../git.tar.gz\nand extract it in git/ directory\".\n\nOf course, for such a toplevel project commit, the tool may not\nbe able to do a checkout automatically and require the user to\ncat-file the link, download a tarball and extract the subtree\nthere manually.\n\nThe bound commit approach requires you to have git commit object\nnames on the `bind` lines, and it is fundamentally much harder\nto extend it to allow interfacing with foreign (non-)SCM\nsystems.\n"},{"id":"15063","messageId":"200601231206.53466.lan@ac-sw.com","threadId":"3131","inReplyTo":"7v64ob1omh.fsf@assigned-by-dhcp.cox.net","subject":"Re: Notes on Subproject Support","fromName":"Alexander Litvinov","fromEmail":"lan@ac-sw.com","sentAt":"2006-01-23T06:06:53Z","receivedAt":"2006-01-23T06:06:53Z","isPatch":false,"sender":{"key":"lan@ac-sw.com","avatar":null},"body":"In our development we use a little bit other case (it is simplified):\n1. We have self written C++ library for linked list implementation. Lets call \nit liblist.\n2. We have project A that use liblist as separate directory and project B that \nuse this library too.\n\nCurrently we have 3 cvs projects with cvs-modules for linking liblist to A and \nB. During development of A and B we often modify liblist to fix bugs and \nthese changes are immidatly visible to all projects who use liblist.\n\nAfter full implementation of bind functionality I see one restriction: I have \nto use one repo for storing all three projects: A, B and liblist to make \nchanges of liblist visible to all projects. The solution is to make separate \nrepos and on each change of liblist in prokect A push these changes to \nliblist repo and pull them into project B again - bit hacky solution.\n"},{"id":"15064","messageId":"7vy817z83t.fsf@assigned-by-dhcp.cox.net","threadId":"3131","inReplyTo":"7v3bjfafql.fsf@assigned-by-dhcp.cox.net","subject":"Re: Notes on Subproject Support","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-01-23T08:00:54Z","receivedAt":"2006-01-23T08:00:54Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <junkio@cox.net> writes:\n\n> This is still a draft/WIP, but \"release early\" is a good\n> discipline, so...\n\nTentatively I'm placing this document in the 'todo' branch, so\nthat people interested in the changes can ask gitweb to show\ndiffs, until I can find a better way and location to manage it.\n\nI do not think it is suitable to be in the Documentation/ area,\ndue to its being an early draft and its just-one-ofthe-proposals\nstatus.\n"},{"id":"15067","messageId":"7voe23z6d6.fsf@assigned-by-dhcp.cox.net","threadId":"3131","inReplyTo":"7v64ob1omh.fsf@assigned-by-dhcp.cox.net","subject":"Re: Notes on Subproject Support","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-01-23T08:38:29Z","receivedAt":"2006-01-23T08:38:29Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <junkio@cox.net> writes:\n\n> BTW, let's digress a bit.\n\nUgh.  Serious typo.\n\n> I think recording \"commit\" in the tree objects is in line with\n> the logical organization described in README: \"blob\" and \"tree\"\n> represent a state, and have *nothing* to do with how we came\n> about to that state.  The historyh is described in \"commit\"\n> objects.  The bound commit approach keeps that property.\n\nObviously, I think \"*NOT* recording commit in tree objects\" is\nin line with \"blobs and trees are about states, commits give\nthem their points in history\".\n"},{"id":"15068","messageId":"20060123125013.GA4472@igloo.ds.co.ug","threadId":"3131","inReplyTo":"7v3bjfafql.fsf@assigned-by-dhcp.cox.net","subject":"Re: Notes on Subproject Support","fromName":"Martin Atukunda","fromEmail":"matlads@dsmagic.com","sentAt":"2006-01-23T12:50:13Z","receivedAt":"2006-01-23T12:50:13Z","isPatch":false,"sender":{"key":"matlads@dsmagic.com","avatar":null},"body":"This proposal doesn't seem to cator for the event when a directory is\nrenamed or moved to a different location, or am I missing something?\n\n- Martin -\nOn Sun, Jan 22, 2006 at 05:35:14PM -0800, Junio C Hamano wrote:\n> This is still a draft/WIP, but \"release early\" is a good\n> discipline, so...\n> \n> -- >8 --\n> \n> Notes on Subproject Support\n> ===========================\n> Junio C Hamano <junkio@cox.net>\n> v1.0 January 22, 2006\n> \n> Scenario\n> --------\n> \n> The examples in the following discussion show how this proposal\n> plans to help this:\n> \n> . A project to build an embedded Linux appliance \"gadget\" is\n>   maintained with git.\n> \n> . The project uses linux-2.6 kernel as its subcomponent.  It\n>   starts from a particular version of the mainline kernel, but\n>   adds its own code and build infrastructure to fit the\n>   appliance's needs.\n> \n> . The working tree of the project is laid out this way:\n> +\n> ------------\n>  Makefile       - Builds the whole thing.\n>  linux-2.6/     - The kernel, perhaps modified for the project.\n>  appliance/     - Applications that run on the appliance, and\n>                   other bits.\n> ------------\n> \n> . The project is willing to maintain its own changes out of tree\n>   of the Linux kernel project, but would want to be able to feed\n>   the changes upstream, and incorporate upstream changes to its\n>   own tree, taking advantage of the fact that both itself and\n>   the Linux kernel project are version controlled with git.\n> \n> The idea here is to:\n> \n> . Keep `linux-2.6/` part as an independent project.  The work by\n>   the project on the kernel part can be naturally exchanged with\n>   the other kernel developers this way.  Specifically, a tree\n>   object contained in commit objects belonging to this project\n>   does *not* have linux-2.6/ directory at the top.\n> \n> . Keep the `appliance/` part as another independent project.\n>   Applications are supposed to be more or less independent from\n>   the kernel version, but some other bits might be tied to a\n>   specific kernel version.  Again, a tree object contained in\n>   commit objects belonging to this project does *not* have\n>   appliance/ directory at the top.\n> \n> . Have another project that combines the whole thing together,\n>   so that the project can keep track of which versions of the\n>   parts are built together.\n> \n> We will call the project that binds things together the\n> 'toplevel project'.  Other projects that hold `linux-2.6/` part\n> and `appliance/` part are called 'subprojects'.\n> \n> Notice that `Makefile` at the top is part of the toplevel\n> project in this example, but it is not necessary.  We could\n> instead have the appliance subproject include this file.  In\n> such a setup, the appliance subproject would have had `Makefile`\n> and `appliance/` directory at the toplevel.\n> \n> \n> Setting up\n> ----------\n> \n> Let's say we have been working on the appliance software,\n> independently version controlled with git.  Also the kernel part\n> has been version controlled separately, like this:\n> ------------\n> $ ls -dF current/*/.git current/*\n> current/Makefile    current/appliance/.git/  current/linux-2.6/.git/\n> current/appliance/  current/linux-2.6/\n> ------------\n> \n> Now we would want to get a combined project.  First we would\n> clone from these repositories (which is not strictly needed --\n> we could use `$GIT_ALTERNATE_OBJECT_DIRECTORIES` instead):\n> \n> ------------\n> $ mkdir combined && cd combined\n> $ cp ../current/Makefile .\n> $ git init-db\n> $ mkdir -p .git/refs/subs/{kernel,gadget}/{heads,tags}\n> $ git clone-pack ../current/linux-2.6/ master | read kernel_commit junk\n> $ git clone-pack ../current/appliance/ master | read gadget_commit junk\n> ------------\n> \n> We will introduce a new command to set up a combined project:\n> \n> ------------\n> $ git bind-projects \\\n> \t$kernel_commit linux-2.6/ \\\n> \t$gadget_commit appliance/\n> ------------\n> \n> This would do an equivalent of:\n> \n> ------------\n> $ git read-tree --prefix=linux-2.6/ $kernel_commit\n> $ git read-tree --prefix=appliance/ $gadget_commit\n> ------------\n> [NOTE]\n> ============\n> Earlier outlines sent to the git mailing list talked\n> about `$GIT_DIR/bind` to record what subproject are bound to\n> which subtree in the curent working tree and index.  This\n> proposal instead records that information in the index file\n> when `--prefix=linux-2.6/` is given to `read-tree`.\n> \n> Also note that in this round of proposal, there is no separate\n> branches that keep track of heads of subprojects.\n> ============\n> \n> Let's not forget to add the `Makefile`, and check the whole\n> thing out from the index file.\n> ------------\n> $ git add Makefile\n> $ git checkout-index -f -u -q -a\n> ------------\n> \n> Now our directory should be identical with the `current`\n> directory.  After making sure of that, we should be able to\n> commit the whole thing:\n> \n> ------------\n> $ diff -x .git -r ../current ../combined\n> $ git commit -m 'Initial toplevel project commit'\n> ------------\n> \n> Which should create a new commit object that records what is in\n> the index file as its tree, with `bind` lines to record which\n> subproject commit objects are bound at what subdirectory, and\n> updates the `$GIT_DIR/refs/heads/master`.  Such a commit object\n> might look like this:\n> ------------\n> tree 04803b09c300c8325258ccf2744115acc4c57067\n> bind 5b2bcc7b2d546c636f79490655b3347acc91d17f linux-2.6/\n> bind 0bdd79af62e8621359af08f0afca0ce977348ac7 appliance/\n> author Junio C Hamano <junio@kernel.org> 1137965565 -0800\n> committer Junio C Hamano <junio@kernel.org> 1137965565 -0800\n> \n> Initial toplevel project commit\n> ------------\n> \n> \n> Making further commits\n> ----------------------\n> \n> The easiest case is when you updated the Makefile without\n> changing anything in the subprojects.  In such a case, we just\n> need to create a new commmit object that records the new tree\n> with the current `HEAD` as its parent, and with the same set of\n> `bind` lines.\n> \n> When we have changes to the subproject part, we would make a\n> separate commit to the subproject part and then record the whole\n> thing by making a commit to the toplevel project.  The user\n> interaction might go this way:\n> ------------\n> $ git commit\n> error: you have changes to the subproject bound at linux-2.6/.\n> $ git commit --subproject linux-2.6/\n> $ git commit\n> ------------\n> \n> With the new `--subproject` option, the directory structure\n> rooted at `linux-2.6/` part is written out as a tree, and a new\n> commit object that records that tree object with the commit\n> bound to that portion of the tree (`5b2bcc7b` in the above\n> example) as its parent is created.  Then the final `git commit`\n> would record the whole tree with updated `bind` line for the\n> `linux-2.6/` part.\n> \n> \n> Checking out\n> ------------\n> \n> After cloning such a toplevel project, `git clone` without `-n`\n> option would check out the working tree.  This is done by\n> reading the tree object recorded in the commit object (which\n> records the whole thing), and adding the information from the\n> \"bind\" line to the index file.\n> \n> ------------\n> $ cd ..\n> $ git clone -n combined cloned ;# clone the one we created earlier\n> $ cd cloned\n> $ git checkout\n> ------------\n> \n> This round of proposal does not maintain separate branch heads\n> for subprojects.  The bound commits and their subdirectories\n> are recorded in the index file from the commit object, so there\n> is no need to do anything other than updating the index and the\n> working tree.\n> \n> \n> Switching branches\n> ------------------\n> \n> Along with the traditional two-way merge by `read-tree -m -u`,\n> we would need to look at:\n> \n> . `bind` lines in the current `HEAD` commit.\n> \n> . `bind` lines in the commit we are switching to.\n> \n> . subproject binding information in the index file.\n> \n> to make sure we do sensible things.\n> \n> Just like until very recently we did not allow switching\n> branches when two-way merge would lose local changes, we can\n> start by refusing to switch branches when the subprojects bound\n> in the index do not match what is recorded in the `HEAD` commit.\n> \n> Because in this round of the proposal we do not use the\n> `$GIT_DIR/bind` file nor separate branches to keep track of\n> heads of the subprojects, there is nothing else other than the\n> working tree and the index file that needs to be updated when\n> switching branches.\n> \n> \n> Merging\n> -------\n> \n> Merging two branches of the toplevel projects can use the\n> traditional merging mechanism mostly unchanged.  The merge base\n> computation can be done using the `parent` ancestry information\n> taken from the two toplevel project branch heads being merged,\n> and merging of the whole tree can be done with a three-way merge\n> of the whole tree using the merge base and two head commits.\n> For reasons described later, we would not merge the subproject\n> parts of the trees during this step, though.\n> \n> When the two branch heads use different versions of subproject,\n> things get a bit tricky.  First, let's forget for a moment about\n> the case where they bind the same project at different location.\n> We would refuse if they do not have the same number of `bind`\n> lines that bind something at the same subdirectories.\n> \n> ------------\n> $ git merge 'Merge in a side branch' HEAD side\n> error: the merged heads have subprojects bound at different places.\n>  ours:\n> \tlinux-2.6/\n> \tappliance/\n>  theirs:\n> \tkernel/\n> \tgadget/\n> \tmanual/\n> ------------\n> \n> Such renaming can be handled by first moving the bind points in\n> our branch, and redoing the merge (this is a rare operation\n> anyway).  It might go like this:\n> \n> ------------\n> $ git bind-projects \\\n> \t$kernel_commit kernel/ \\\n> \t$gadget_commit gadget/\n> $ git commit -m 'Prepare for merge with side branch'\n> $ git merge 'Merge in a side branch' HEAD side\n> error: the merged heads have subprojects bound at different places.\n>  ours:\n> \tkernel/\n> \tgadget/\n>  theirs:\n> \tkernel/\n> \tgadget/\n> \tmanual/\n> ------------\n> \n> Their branch added another subproject, so this did not work (or\n> it could be the other way around -- we might have been the one\n> with `manual/` subproject while they didn't).  This suggests\n> that we may want an option to `git merge` to allow taking a\n> union of subprojects.  Again, this is a rare operation, and\n> always taking a union would have created a toplevel project that\n> had both `kernel/` and `linux-2.6/` bound to the same Linux\n> kernel project from possibly different vintage, so it would be\n> prudent to require the set of bound subprojects to exactly match\n> and give the user an option to take a union.\n> \n> ------------\n> $ git merge --union-subprojects 'Merge in a side branch HEAD side\n> error: the subproject at `kernel/` needs to be merged first.\n> ------------\n> \n> Here, the version of the Linux kernel project in the `side`\n> branch was different from what our branch had on our `bind`\n> line.  On what kind of difference should we give this error?\n> Initially, I think we could require one is the fast forward of\n> the other (ours might be ahead of theirs, or the other way\n> around), and take the descendant.\n> \n> Or we could do an independent merge of subprojects heads, using\n> the `parent` ancestry of the bound subproject heads to find\n> their merge-base and doing a three-way merge.  This would leave\n> the merge result in the subproject part of the working tree and\n> the index.\n> \n> [NOTE]\n> This is the reason we did not do the whole-tree three way merge\n> earlier.  The subproject commit bound to the merge base commit\n> used for the toplevel project may not be the merge base between\n> the subproject commits bound to the two toplevel project\n> commits.\n> \n> So let's deal with the case to merge only a subproject part into\n> our tree first.\n> \n> \n> Merging subprojects\n> -------------------\n> \n> An operation of more practical importance is to be able to merge\n> in changes done outside to the projects bound to our toplevel\n> project.\n> \n> ------------\n> $ git pull --subproject=kernel/ git://git.kernel.org/.../linux-2.6/\n> ------------\n> \n> might do:\n> \n> . fetch the current `HEAD` commit from Linus.\n> . find the subproject commit bound at kernel/ subtree.\n> . perform the usual three-way merge of these two commits, in\n>   `kernel/` part of the working tree.\n> \n> After that, `git commit \\--subproject` option would be needed to\n> make a commit.\n> \n> [NOTE]\n> This suggests that we would need to have something similar to\n> `MERGE_HEAD` for merging the subproject part.\n> \n> \n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n\n-- \nDue to a shortage of devoted followers, the production of great leaders has been discontinued.\n"},{"id":"15070","messageId":"Pine.LNX.4.64.0601231116550.25300@iabervon.org","threadId":"3131","inReplyTo":"7v7j8r7e7s.fsf@assigned-by-dhcp.cox.net","subject":"Re: Notes on Subproject Support","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2006-01-23T16:31:11Z","receivedAt":"2006-01-23T16:31:11Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Sun, 22 Jan 2006, Junio C Hamano wrote:\n\n> Daniel Barkalow <barkalow@iabervon.org> writes:\n> \n> >> Switching branches\n> >> ------------------\n> >> \n> >> Along with the traditional two-way merge by `read-tree -m -u`,\n> >> we would need to look at:\n> >> \n> >> . `bind` lines in the current `HEAD` commit.\n> >> \n> >> . `bind` lines in the commit we are switching to.\n> >> \n> >> . subproject binding information in the index file.\n> >> \n> >> to make sure we do sensible things.\n> >\n> > This is one place I think storing the bindings in the commit is awkward. \n> > read-tree deals in trees (hence the name), but will need information from \n> > the commit.\n> \n> That's why it is 'along with'.  Dealing with binding information\n> can be done between commits and index without bothering tree\n> objects.  read-tree would not have to deal with it, and I think\n> keeping it that way is probably a good idea.\n\nI think it would be a lot more fragile if switching branches requires \nmultiple programs interacting with the index file. If things get \ninterrupted after the tree is read but before the bindings are changed, \nthe user will probably generate an inconsistant commit or have to deal \nwith figuring out what's going on. It is a nice property of the current \nsystem that the index file never exists under the usual filename without \nbeing consistant.\n\n> In other words, I think the design so far does not require us to\n> touch tree objects at all, and I'd be happy if we do not have to.\n>\n> One reason I started the bound commit approach was exactly\n> because I only needed to muck with commit objects and did not\n> have to touch trees and blobs; after trying to implement the\n> core level for \"gitlink\", which I ended up touching quite a lot\n> and have abandoned for now.\n\nI think that the same thing that worked with the index file would work for \nminimizing the impact of the changes as far as the code sees. If the \nstruct tree parser reported the tree at a path when it found a commit at a \npath, it would work just as if the original tree had used trees (like in \nyour proposal).\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"15071","messageId":"Pine.LNX.4.64.0601231137250.25300@iabervon.org","threadId":"3131","inReplyTo":"200601231206.53466.lan@ac-sw.com","subject":"Re: Notes on Subproject Support","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2006-01-23T16:48:41Z","receivedAt":"2006-01-23T16:48:41Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Mon, 23 Jan 2006, Alexander Litvinov wrote:\n\n> In our development we use a little bit other case (it is simplified):\n> 1. We have self written C++ library for linked list implementation. Lets call \n> it liblist.\n> 2. We have project A that use liblist as separate directory and project B that \n> use this library too.\n> \n> Currently we have 3 cvs projects with cvs-modules for linking liblist to A and \n> B. During development of A and B we often modify liblist to fix bugs and \n> these changes are immidatly visible to all projects who use liblist.\n> \n> After full implementation of bind functionality I see one restriction: I have \n> to use one repo for storing all three projects: A, B and liblist to make \n> changes of liblist visible to all projects. The solution is to make separate \n> repos and on each change of liblist in prokect A push these changes to \n> liblist repo and pull them into project B again - bit hacky solution.\n\nWe haven't yet discussed how pushing a repository with subprojects would \nwork. We could probably have an extra line in the remotes/ file to make \nthe desired thing happen automatically, so that you can just do \"git push\" \nand have the subproject go to its repository and the main project go to \nits repository.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"15072","messageId":"Pine.LNX.4.64.0601231200380.25300@iabervon.org","threadId":"3131","inReplyTo":"7v64ob1omh.fsf@assigned-by-dhcp.cox.net","subject":"Re: Notes on Subproject Support","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2006-01-23T17:57:29Z","receivedAt":"2006-01-23T17:57:29Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Sun, 22 Jan 2006, Junio C Hamano wrote:\n\n> Junio C Hamano <junkio@cox.net> writes:\n> \n> BTW, let's digress a bit.\n> \n> I think recording \"commit\" in the tree objects is in line with\n> the logical organization described in README: \"blob\" and \"tree\"\n> represent a state, and have *nothing* to do with how we came\n> about to that state.  The historyh is described in \"commit\"\n> objects.  The bound commit approach keeps that property.\n>\n> The \"gitlink\" approach, as I understand how Linus outlined in\n> his original suggestion, is a bit different.  The link objects\n> appear in tree objects, and when you \"git cat-file link\" one of\n> them, you would see something like this:\n> \n>         commit\t5b2bcc7b2d546c636f79490655b3347acc91d17f\n>         name\tkernel\n> \n> So in that sense, \"gitlink\" approach departs from the original\n> premise of \"commit\" being the only thing that ties things\n> together.  Tree objects with \"gitlink\" know where they are in\n> the history [*1*].\n> \n> By this, I do not mean to say that \"gitlink\" approach is\n> inferior because it breaks that original premise.  I am just\n> pointing it out as a difference between two approaches.\n\nI think it's hard to say whether the history of a subproject is part of \nthe state of the superproject or part of its history. They're certainly \nnot the same history, because the superproject history may record that \nthe superproject switched from one fork of a subproject to a different \nfork, or reverted the subproject to an earlier version, or other such \nthings. (It's the whole data/metadata issue: when you take a step back, \none level's data and metadata are all data, and there's more stuff that's \nmetadata.)\n\nI'd say that with either commits in trees or gitlink objects, it's still \nonly the commits that tie things together; but some of the things that \nthey tie together are opaquely things tied to other things. Tree objects \nwith gitlink/commits don't know where *they* are in the history; they \njust happen to store things which know about a different history.\n\n> Now, the current way index file is used is as a staging area to\n> create a new commit on top of the tip of the current branch.\n> However, it is interesting to note that logically, by itself\n> *alone*, it cannot be used that way.  The information the index\n> file records is something that can be used to write out a tree\n> object, and not a commit, because it does not know where the\n> current state sits in the history.  We have two separate files,\n> $GIT_DIR/HEAD that records which branch we are on, and the\n> branch head ref the HEAD points at, which records where the\n> current index came from, for that purpose.  The latter tells us\n> what commit we should use as the parent commit if we create such\n> a commit, and the former tells us which branch head to update\n> once we create one.  So in that sense, the index file is just a\n> staging area to create a new tree, not a new commit.\n\nOf course, we don't strictly need $GIT_DIR/HEAD to create a commit; \nthat's only needed for what we generally do with the commit once we have \nit. We do need the branch head ref (or, more abstractly, the commit that \nwe read to generate the index we modified) in order to create the commit.\n\n> We could have done things differently.  I am not advocating to\n> do the following change, but offering a possibility as a thought\n> experiment.  It just felt interesting enough to point them out.\n> \n> The index file could have recorded what commit the current state\n> recorded in the index came from.  By recording the commit the\n> index was read from in the index itself, independently from the\n> $GIT_DIR/refs/heads/$branch file, we could have been able to\n> allow fetching into the current branch.  When the $branch file\n> for the current branch was updated by a fast-forward fetch, we\n> would notice that the commit recorded there no longer match what\n> is recorded in the index.\n\nI actually think this would have been a good idea. I think we've had \napproximately every possible bug that could come from inconsistancy \nbetween the files that give the parents and the index file. (I think Linus \ndidn't do it that way initially just because he was thinking of it as a \ncache, and there's little point in caching something small, and by the \ntime we started looking at it as primary information on its own, we'd \nstopped thinking about what should go in it.)\n\n> Another interesting consequence is if the development is a\n> single repository and linear, we did not even need any file in\n> $GIT_DIR/refs/ (\"branchless git\").  The commit recorded as the\n> topmost in the index file itself would have served as the tip of\n> the development, and we would have been able to tangle the\n> history starting from the commit in the index file.\n\nWell, you wouldn't be able to check out an old version and then return to \nthe present without dredging the objects database for the dangling commit. \nMy memory has gotten fuzzy, but I think HEAD may have originally been just \na file, and we effectively had this (except that HEAD and the index were \nnot the same file as far as the filesystem was concerned).\n\n> While we are doing a thought experiment, let's say we allow to\n> record more than one commits the current index is based upon.\n> 'git merge' would record all the parent commits there, so that\n> writing out the merge result out of the index file as a tree and\n> then recording these commits as parents would have been the way\n> to create a merge commit.  We would not need the auxiliary file\n> $GIT_DIR/MERGE_HEAD if we did so.\n>\n> In other words, if the index file recorded the commits its\n> contents were based upon, instead of being a staging area for a\n> new tree, it would have been a staging area for a new commit.\n> \n> Now, the latest proposal, borrowing your idea, records the\n> subproject commits bound to subdirectories in the index itself.\n> This is halfway to make the index file a staging area for the\n> next commit.  If we were to do that, we also *could* record the\n> commits the current index is based upon, so that it can truly be\n> used as a staging area to create a new commit, not just a tree.\n> \n> On the other hand, this could be a reason *not* to do the\n> `update-index --bind` to record the subproject information in\n> the index file.  An auxiliary file such as $GIT_DIR/bind might\n> be sufficient, just like $GIT_DIR/MERGE_HEAD has been good\n> enough for us so far.  One difference between MERGE_HEAD and\n> bind is that the former is very transient -- only exists during\n> a merge while the latter is persistent while the top commit is\n> checked out and being worked on.\n\nWe've been able to make MERGE_HEAD work, but I remember there being \nproblems even there when people tried to abandon merges by changing \nbranches. Do you see an advantage to having the index only record the \ninformation used for making a tree, and keeping the information for making \na commit in other files?\n\n> [Footnote]\n> \n> *1* One good property of \"gitlink\" approach is that we *could*\n> extend this blob-like object to store arbitrary human readable\n> information to represent a point-in-time from an arbitrary\n> foreign SCM system.  IOW, we do not necessarily have to require\n> `commit` line that name a git commit to be there.  It could say\n> \"Please slurp http://www.kernel.org/pub/software/.../git.tar.gz\n> and extract it in git/ directory\".\n> \n> Of course, for such a toplevel project commit, the tool may not\n> be able to do a checkout automatically and require the user to\n> cat-file the link, download a tarball and extract the subtree\n> there manually.\n> \n> The bound commit approach requires you to have git commit object\n> names on the `bind` lines, and it is fundamentally much harder\n> to extend it to allow interfacing with foreign (non-)SCM\n> systems.\n\nI don't think this would really be useful. The reason to have the included \nrevision stored in a way that's explicitly marked for git to find is that \ngit can do useful things with the information (such as checking it out for \nyou, but more importantly, making sure that changes to what revision \nyou're working with propagate to changes in what revision you specify \nshould be there). If the bound project is foreign, this clearly isn't \ngoing to happen, so there's not much point. For your example above, you \ncould just have a regular file, \"git/README\", with the content \"Please \ndownload http://.../git.tar.gz, and extract it here\", and it would be at \nleast as good.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"15073","messageId":"7vk6cqyc79.fsf@assigned-by-dhcp.cox.net","threadId":"3131","inReplyTo":"20060123125013.GA4472@igloo.ds.co.ug","subject":"Re: Notes on Subproject Support","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-01-23T19:30:02Z","receivedAt":"2006-01-23T19:30:02Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Martin Atukunda <matlads@dsmagic.com> writes:\n\n> This proposal doesn't seem to cator for the event when a directory is\n> renamed or moved to a different location, or am I missing something?\n\nFirst of all, please do not top post.\n\nSecond of all, please do not quote the whole thing.\n\nThird of all, if you quote, please read the parts you quote.\n\n>> Merging\n>> -------\n>> ...\n>> Such renaming can be handled by first moving the bind points in\n>> our branch, and redoing the merge (this is a rare operation\n>> anyway).  It might go like this:\n>> ...\n\nThis step describes how bind-point might be relocated prior to a\nmerge.\n"},{"id":"15074","messageId":"7vek2yxukm.fsf@assigned-by-dhcp.cox.net","threadId":"3131","inReplyTo":"Pine.LNX.4.64.0601231116550.25300@iabervon.org","subject":"Re: Notes on Subproject Support","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-01-24T01:50:49Z","receivedAt":"2006-01-24T01:50:49Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Daniel Barkalow <barkalow@iabervon.org> writes:\n\n> I think it would be a lot more fragile if switching branches requires \n> multiple programs interacting with the index file. If things get \n> interrupted after the tree is read but before the bindings are changed, \n> the user will probably generate an inconsistant commit or have to deal \n> with figuring out what's going on. It is a nice property of the current \n> system that the index file never exists under the usual filename without \n> being consistant.\n\nThat is certainly an issue, which we have had already for quite\nsome time, I am afraid.  We can get interrupted during \"switch\nbranches\" flow after read-tree -u -m but before updating HEAD.\nWe can also get interrupted during \"commit\" flow after writing\nthe commit object out before updating the ref pointed at by\nHEAD.  No?\n\nIf we are truly serious about solving the issue of getting\ninterrupted in the middle, I suspect we have to take the \"index\nis a staging area for the next commit\" approach I digressed into\nlast night.  It would involve introducing a git-atomic-checkout\ncommand to replace the current \"git-rev-parse, git-read-tree,\nthen git-symbolic-ref\" sequence in the checkout flow.  In the\ncommit flow, we would need git-commit-index command to replace\nthe current \"git-write-tree, git-commit-tree, then\ngit-update-ref\" sequence.\n\nI am not particularly opposed to that, but I suspect it might be\na moderate amount of work for very little gain.  Continuing with\nthe digression, the updated index file may contain:\n\n  1. list of <blob path, object name>\n  2. list of parent commit object names for the next commit\n  3. the name of the local branch to create the next commit on\n  4. for each bound path:\n     list of parent commit object names for that path.\n\n1. is what we have in the current (version 2) index file.\n\n2. contains:\n - 0 commit in an index file before the initial commit\n - 1 commit in an index file after a fresh checkout and records\n   the commit object name we checked out (replaces HEAD+heads/$branch)\n - 2 commits in an index file after 3-way read-tree, or more\n   during an Octopus merge (replaces HEAD+MERGE_HEAD)\n\n3. may not be needed, but if we did so, it would replace HEAD.\n\n4. is similar to 2 but for bound subprojects.  Usually we have\n   one commit per bound path to record the \"bind\" line commit we\n   read from the commit object after a fresh checkout.  During a\n   subproject merge, we would:\n   - start out with 1 commit read from the \"bind\" line;\n   - merging in another subproject commit would add that commit;\n   - when making a new subproject commit, the recorded commits\n     are used as its parents.\n"},{"id":"15075","messageId":"7v8xt6xuke.fsf@assigned-by-dhcp.cox.net","threadId":"3131","inReplyTo":"Pine.LNX.4.64.0601231200380.25300@iabervon.org","subject":"Re: Notes on Subproject Support","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-01-24T01:50:57Z","receivedAt":"2006-01-24T01:50:57Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Daniel Barkalow <barkalow@iabervon.org> writes:\n\n> ... Do you see an advantage to having the index only record the \n> information used for making a tree, and keeping the information for making \n> a commit in other files?\n\nIf somebody else already did the work and presented me two git\nimplementations, one with the index file capable of generating a\ntree and uses separate files to keep track of other information\nfor commits, and the other with the index file with everything\nneeded for a commit, I'd certainly take the latter.  In that\nsense, I do not see such an advantage at all.  The practical\nadvantage of keeping them separate is to keep things simple,\nminimizing the changes.  I see the subproject support as a\nsecondary issue, and so far I haven't found a reason convincing\nenough to tell me that it is better to put HEAD+heads/$branch\ninformation in the index itself when used in a subproject-less\nsetup.  It perhaps would make us more robust when interrupted in\nthe middle of switching branches or making a commit, but that is\nabout it (I do not particularly see that a serious problem).\n\n>> *1* One good property of \"gitlink\" approach is that we *could*\n>> extend this blob-like object to store arbitrary human readable\n>> information to represent a point-in-time from an arbitrary\n>> foreign SCM system.  IOW, we do not necessarily have to require\n>> `commit` line that name a git commit to be there.  It could say\n>> \"Please slurp http://www.kernel.org/pub/software/.../git.tar.gz\n>> and extract it in git/ directory\".\n>> ...\n> I don't think this would really be useful. The reason to have the included \n> revision stored in a way that's explicitly marked for git to find is that \n> git can do useful things with the information ...\n> but more importantly, making sure that changes to what revision \n> you're working with propagate to changes in what revision you specify \n> should be there)...\n\nMy example was taking things to the extreme to be illustrative.\n\nTo be more practical, it could have pointed at \"git-1.0.tar.gz\"\nor an \"svn://\" URL with explicit revision name, which ought to\nbe enough to recreate the exact state.\n"},{"id":"15078","messageId":"Pine.LNX.4.64.0601232220570.25300@iabervon.org","threadId":"3131","inReplyTo":"7vek2yxukm.fsf@assigned-by-dhcp.cox.net","subject":"Re: Notes on Subproject Support","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2006-01-24T04:22:41Z","receivedAt":"2006-01-24T04:22:41Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Mon, 23 Jan 2006, Junio C Hamano wrote:\n\n> Daniel Barkalow <barkalow@iabervon.org> writes:\n> \n> > I think it would be a lot more fragile if switching branches requires \n> > multiple programs interacting with the index file. If things get \n> > interrupted after the tree is read but before the bindings are changed, \n> > the user will probably generate an inconsistant commit or have to deal \n> > with figuring out what's going on. It is a nice property of the current \n> > system that the index file never exists under the usual filename without \n> > being consistant.\n> \n> That is certainly an issue, which we have had already for quite\n> some time, I am afraid.  We can get interrupted during \"switch\n> branches\" flow after read-tree -u -m but before updating HEAD.\n> We can also get interrupted during \"commit\" flow after writing\n> the commit object out before updating the ref pointed at by\n> HEAD.  No?\n\nThe switch branches one is accurate, but I think that, if we get \ninterrupted before updating the ref, the index will still be the same, and \nwe'll just have a dangling object (which, if we commit the same thing \nagain, will be the same object we generate).\n\nI suppose the existing branch switching isn't much less bad than the new \none would be, though. I sort of worry that rewriting the index file is \nmore likely to be interrupted than updating a ref, but that's probably not \nreally a significant difference.\n\n> If we are truly serious about solving the issue of getting\n> interrupted in the middle, I suspect we have to take the \"index\n> is a staging area for the next commit\" approach I digressed into\n> last night.  It would involve introducing a git-atomic-checkout\n> command to replace the current \"git-rev-parse, git-read-tree,\n> then git-symbolic-ref\" sequence in the checkout flow. \n\nWell, if the value of HEAD were in the index file, that would be \nsufficient to prevent anything actually bad from happening in the checkout \npath; if it gets interrupted, the index file's \"current commit\" field \nwould then not match the ref and it would be clear that the system was in \nan intermediate state. (It would appear like if you'd fetched into the \ncurrent branch without it doing the fast-forward.)\n\n> In the commit flow, we would need git-commit-index command to replace\n> the current \"git-write-tree, git-commit-tree, then\n> git-update-ref\" sequence.\n\nI don't think there's an issue here, anyway.\n\n> I am not particularly opposed to that, but I suspect it might be\n> a moderate amount of work for very little gain.  Continuing with\n> the digression, the updated index file may contain:\n> \n>   1. list of <blob path, object name>\n>   2. list of parent commit object names for the next commit\n>   3. the name of the local branch to create the next commit on\n>   4. for each bound path:\n>      list of parent commit object names for that path.\n> \n> 1. is what we have in the current (version 2) index file.\n> \n> 2. contains:\n>  - 0 commit in an index file before the initial commit\n>  - 1 commit in an index file after a fresh checkout and records\n>    the commit object name we checked out (replaces HEAD+heads/$branch)\n>  - 2 commits in an index file after 3-way read-tree, or more\n>    during an Octopus merge (replaces HEAD+MERGE_HEAD)\n> \n> 3. may not be needed, but if we did so, it would replace HEAD.\n\nI don't think it would be needed; it could certainly be passed in. \nActually not having HEAD would complicate a lot of programs that use HEAD \nbut don't currently read the index (and don't actually care about whether \nyou have the branch that you consider current actually checked out).\n\n> 4. is similar to 2 but for bound subprojects.  Usually we have\n>    one commit per bound path to record the \"bind\" line commit we\n>    read from the commit object after a fresh checkout.  During a\n>    subproject merge, we would:\n>    - start out with 1 commit read from the \"bind\" line;\n>    - merging in another subproject commit would add that commit;\n>    - when making a new subproject commit, the recorded commits\n>      are used as its parents.\n\nRight.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"}]}