{"thread":{"id":"3187","subject":"[RFC] shallow clone","startedAt":"2006-01-30T07:18:50Z","lastAt":"2006-02-02T19:31:35Z","messageCount":30,"participants":["Junio C Hamano","Johannes Schindelin","Simon Richter","Franck"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"15247","messageId":"7voe1uchet.fsf@assigned-by-dhcp.cox.net","threadId":"3187","inReplyTo":null,"subject":"[RFC] shallow clone","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-01-30T07:18:50Z","receivedAt":"2006-01-30T07:18:50Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Shallow History Cloning\n=======================\n\nOne good thing about git repository is that each clone is a\nfreestanding and complete entity, and you can keep developing in\nit offline, without talking to the outside world, knowing that\nyou can sync with them later when online.\n\nIt is also a bad thing.  It gives people working on projects\nwith long development history stored in CVS a heart attack when\nwe tell them that their clones need to store the whole history.\n\nThere was a suggestion by Linus to allow a partial clone using a\nsyntax like this:\n\n\t$ git clone --since=v2.6.14 git://.../linux-2.6/ master\n\nHere is an outline of what changes are needed to the current\ncore to do this.\n\n\nStrategy\n--------\n\nWe have `info/grafts` mechanism to fake parent information for\ncommit objects.  Using this facility, we could roughly do:\n\n. Download the full tree for v2.6.14 commit and store its\n  objects locally.\n\n. Set up `info/grafts` to lie to the local git that Linux kernel\n  history began at v2.6.14 version.\n\n. Run `git fetch git://.../linux-2.6 master`, with a local ref\n  pointing at v2.6.14 commit, to pretend that we have everything\n  up to v2.6.14 to `upload-pack` running on the other end.\n\n. Update the `origin` branch with the master commit object name\n  we just fetched from Linus.\n\nThere are some issues.\n\n. In the fetch above to obtain everything after v2.6.14, and\n  future runs of `git fetch origin`, if a blob that is in the\n  commit being fetched happens to match what used to be in a\n  commit that is older than v2.6.14 (e.g. a patch was reverted),\n  `upload-pack` running on the other end is free to omit sending\n  it, because we are telling it that we are up to date with\n  respect to v2.6.14.  Although I think the current `rev-list\n  --objects` implementation does not always do such a revert\n  optimization if the revert is to a blob in a revision that is\n  sufficiently old, it is free to optimize more aggressively in\n  the future.\n\n. Later when the user decides to fetch older history, the\n  operation can become a bit cumbersome.\n\nI think the latter one is cumbersome but is doable -- we could\ndo the equivalent of:\n\n\t$ git clone --since=v2.6.13 origin v2.6.14\n\nplace all the objects obtained by such a clone/fetch operation\nand remember that now we have history beginning at v2.6.13.  So\nlet's worry about that later.\n\nFor the first issue, we need to have the other end cooperate\nwhile fetching from it.  If the other end also thinks the\ndevelopment started at v2.6.14, even if we tell that we have the\nhistory up to v2.6.14 (or a commit we obtained since then),\nthere is no way for `upload-pack` running there to optimize too\nagressively and assume we have a blob that appeared in v2.6.13.\nMore simply, we do not have to tell them we have anything -- if\nthe other end thinks the epoch is at v2.6.14, only commits that\ncomes later will be sent to us.\n\n\nDesign\n------\n\nFirst, to bootstrap the process, we would need to add a way to\nobtain all objects associated with a commit.  We could do a new\nprogram, or we could implement this as a protocol extension to\n`upload-pack`.  My current inclination is the latter.\n\nWhen talking with `upload-pack` that supports this extension,\nthe downloader can give one commit object name and get a pack\nthat contains all the objects in the tree associated with that\ncommit, plus the commit object itself.  This is a rough\nequivalent of running the commit walker with the `-t` flag.\n\nAnother functionality we would need is to tell `upload-pack` to\nuse `info/grafts` of downloader's choice.  With this, after\nfetching the objects for v2.6.14 commit, the downloader can set\nup its own grafts file to cauterize the development history at\nv2.6.14, and tell the `upload-pack` to pretend the kernel\nhistory starts at that commit, while sending the tip of Linus'\ndevelopment track to us.\n\nUsing the extended protocol (let's call it 'shallow' extension),\na clone to create a repository that has only recent kernel\nhistory since v2.6.14 goes like this:\n\nThe first client is to fetch the v2.6.14 itself.\n\n[NOTE]\nMost likely this is not directly run by the user but is run as\nthe first command invoked by the shallow clone script.\n\n1. The `fetch-pack` command acquires a new option, `--single`:\n\n\t$ git-fetch-pack --single git://.../linux-2.6/ v2.6.14\n\n   This talks with `upload-pack` on the kernel.org server via\n   `git-daemon`.\n\n2. `upload-pack` tells the fetcher what commits it has,\n   what their refs are, and what protocol extensions it\n   supports, as usual.\n\n3. If it does not see `shallow` extension supported, there is no\n   way to get a single tree, so things fail here.  Otherwise, it\n   sends `single X{40}\\0` request, instead of the usual `want`\n   line.  The object name sent here is the desired commit.\n\n4. `upload-pack` notices this is a single commit request, and\n   sends an ACK if it can satisfy the request (or a NAK if it\n   can't, e.g. it does not have the asked commit).  Instead of\n   doing the usual `get_common_commits` followed by\n   `create_pack_file`, it does:\n\n\t$ git rev-list -n1 --objects $commit | git pack-object\n\n   and sends the result out.\n\n5. The fetcher checks the ACK and receives the objects.\n\nAfter the above exchange, we have downloaded v2.6.14 commit and\nits objects but not its history.  `git-fetch-pack` would output\nthe tag object name for `v2.6.14` and we would stash it away in\n`$GIT_DIR/FETCH_HEAD` as usual.  Then we set up `info/grafts`\nwith this:\n\n\t$ git rev-parse FETCH_HEAD^{commit} >\"$GIT_DIR/info/grafts\"\n\nThis cauterizes the history on our end.\n\nThe second phase of the shallow clone is to fetch the history\nsince v2.6.14 to the tip.\n\n1. The `fetch-pack` command is run as usual.  Most likely the\n   command line run by the shallow clone script would be:\n\n\t$ git fetch-pack git://.../linux-2.6/ master\n\n   Notice there is nothing magical about it.  It is just the\n   business as usual.\n\n2. `upload-pack` does its usual greeting to the downloader.\n\n3. We notice `shallow` extension again, and first send out\n   `graft X{40}\\0` request.  The syntax of graft request would\n   be `graft ` followed by one or more commit object names on a\n   line separated with SP.  After sending out all the needed\n   graft requests (in this example there is only one, to\n   cauterize the history at v2.6.14), it does the usual `want\n   X{40}\\0multi_ack` and a flush.\n\n4. `upload-pack` notices graft requests, reinitializes its graft\n   information with what it receives from the other end, and\n   then records `want`.\n\n5. After the above steps, the usual `upload-pack` vs\n   `fetch-pack` exchange continues and objects needed to\n   complete the Linus' tip of development trail for somebody who\n   has v2.6.14 are sent in a pack.  The difference from the\n   usual operation is that `upload-pack` during this run thinks\n   v2.6.14 commit does not have any parent.\n\nThe exact sequence from the second part of the initial \"shallow\nclone\" can be used for further updates.\n\nThere is a small issue about the actual implementation.  In the\nabove description I pretended that `upload-pack` can be told to\nuse phony grafts information, but in the current implementation\nthe program that needs to use phony grafts information is\n`rev-list` spawned from it.  We _could_ point GIT_GRAFT_FILE\nenvironment variable point at a temporary file while we do so,\nbut I'd like to avoid using a temporary file if possible, given\nthat `upload-pack` is run from `git-daemon`.  Maybe we could\ngive --read-graft-from-stdin flag to `rev-list` for this\npurpose.\n\n\nAnybody want to try?\n"},{"id":"15249","messageId":"Pine.LNX.4.63.0601301220420.6424@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3187","inReplyTo":"7voe1uchet.fsf@assigned-by-dhcp.cox.net","subject":"Re: [RFC] shallow clone","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-01-30T11:39:34Z","receivedAt":"2006-01-30T11:39:34Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 29 Jan 2006, Junio C Hamano wrote:\n\n> Strategy\n> --------\n> \n> We have `info/grafts` mechanism to fake parent information for\n> commit objects.  Using this facility, we could roughly do:\n> \n> . Download the full tree for v2.6.14 commit and store its\n>   objects locally.\n\nOn first read, I mistook \"tree\" for \"commit\"...\n\n> . Set up `info/grafts` to lie to the local git that Linux kernel\n>   history began at v2.6.14 version.\n\nMaybe also record this in .git/config, so that you can\n\n- disallow fetching from this repo, and\n- easily extend the shallow copy to a larger shallow one, or a full one.\n\n> . Run `git fetch git://.../linux-2.6 master`, with a local ref\n>   pointing at v2.6.14 commit, to pretend that we have everything\n>   up to v2.6.14 to `upload-pack` running on the other end.\n\nHow about refs/tags/start_shallow?\n\n> . Update the `origin` branch with the master commit object name\n>   we just fetched from Linus.\n> \n> Design\n> ------\n>\n> [...]\n>\n> Another functionality we would need is to tell `upload-pack` to\n> use `info/grafts` of downloader's choice.  With this, after\n> fetching the objects for v2.6.14 commit, the downloader can set\n> up its own grafts file to cauterize the development history at\n> v2.6.14, and tell the `upload-pack` to pretend the kernel\n> history starts at that commit, while sending the tip of Linus'\n> development track to us.\n\nWhy not just start another fetch? Then, \"have <refs/tags/start_shallow>\" \nwould be sent, and upload-pack does the right thing?\n\nIf you absolutely want to get only one pack, which then is stored as-is, \nupload-pack could start two rev-list processes: one for the tree and one \nfor all the rest.\n\n> [...]\n> \n> [NOTE]\n> Most likely this is not directly run by the user but is run as\n> the first command invoked by the shallow clone script.\n\nBetter make it an option to git-clone\n\n> 4. `upload-pack` notices this is a single commit request, and\n>    sends an ACK if it can satisfy the request (or a NAK if it\n>    can't, e.g. it does not have the asked commit).  Instead of\n>    doing the usual `get_common_commits` followed by\n>    `create_pack_file`, it does:\n> \n> \t$ git rev-list -n1 --objects $commit | git pack-object\n\nHere it could say\n\n(git rev-list -n1 --objects $commit_since; git rev-list --objects \n\t^$commit_since $commit) | git pack-object\n\nIf the former is still needed (e.g. for git-tar-remote-tree), we could \ndistinguish \"single <ref>\" and \"shallow <ref>\" commands.\n\n> [...]\n> \n> The second phase of the shallow clone is to fetch the history\n> since v2.6.14 to the tip.\n\nAs I outlined above, I don't see the need for this.\n\nCiao,\nDscho\n"},{"id":"15251","messageId":"43DDFF5C.30803@hogyros.de","threadId":"3187","inReplyTo":"Pine.LNX.4.63.0601301220420.6424@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: [RFC] shallow clone","fromName":"Simon Richter","fromEmail":"simon.richter@hogyros.de","sentAt":"2006-01-30T11:58:20Z","receivedAt":"2006-01-30T11:58:20Z","isPatch":false,"sender":{"key":"simon.richter@hogyros.de","avatar":"https://gravatar.com/avatar/1192aa9fa5dd19ce258b12b044cc27111dd7cdb58d920dc2123a02a24f55b5c5?d=mp&s=160"},"body":"Hi,\n\nJohannes Schindelin wrote:\n\n>>. Set up `info/grafts` to lie to the local git that Linux kernel\n>>  history began at v2.6.14 version.\n\n> Maybe also record this in .git/config, so that you can\n\nI like that \"config\" thing less and less every day. It appears to become \na kind of registry, where having dedicated files for specific \nfunctionality would provide the robustness of tools not having to touch \nthings they do not care about; but that's just personal opinion.\n\n> - disallow fetching from this repo, and\n\nWhy? It's perfectly acceptable to pull from an incomplete repo, as long \nas you don't care about the old history.\n\n> - easily extend the shallow copy to a larger shallow one, or a full one.\n\nHrm, I think there should also be a way to shrink a repo and \"forget\" \nold history occasionally (obviously, use of that feature would be highly \ndiscouraged).\n\n>>. Run `git fetch git://.../linux-2.6 master`, with a local ref\n>>  pointing at v2.6.14 commit, to pretend that we have everything\n>>  up to v2.6.14 to `upload-pack` running on the other end.\n\n> How about refs/tags/start_shallow?\n\nNo, as that would imply that cloning from such a repo is disallowed.\n\nIMO, it may be a lot more robust to just have a list of \"cutoff\" object \nids in .git/shallow instead of messing with grafts here, as adding or \nremoving a line from that file is an easier thing to do for porcelain \n(or by hand) than rewriting the grafts file. Whether that list would be \ninclusive or exclusive would need to be decided still.\n\n    Simon\n"},{"id":"15252","messageId":"Pine.LNX.4.63.0601301305100.20228@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3187","inReplyTo":"43DDFF5C.30803@hogyros.de","subject":"Re: [RFC] shallow clone","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-01-30T12:13:56Z","receivedAt":"2006-01-30T12:13:56Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 30 Jan 2006, Simon Richter wrote:\n\n> Johannes Schindelin wrote:\n> \n> > > . Set up `info/grafts` to lie to the local git that Linux kernel\n> > >  history began at v2.6.14 version.\n> \n> > Maybe also record this in .git/config, so that you can\n> \n> I like that \"config\" thing less and less every day. It appears to become a\n> kind of registry, where having dedicated files for specific functionality\n> would provide the robustness of tools not having to touch things they do not\n> care about; but that's just personal opinion.\n\nIt is becoming sort of a registry: it contains metadata about the current \nrepository, easily available to scripts and programs.\n\nI beg to differ on your personal opinion on the grounds that the \nrobustness comes from testing, not from diversity. I much prefer to have a \nwell tested config mechanism to having dozens of differently formatted \nfiles with less-than-well tested parsers.\n\nThank you for the insights in your personal opinion anyway.\n\n> > - disallow fetching from this repo, and\n> \n> Why? It's perfectly acceptable to pull from an incomplete repo, as long as you\n> don't care about the old history.\n\nRight. But should that be the default? I don't think so. Therefore: \ndisable it, and if the user is absolutely sure to do dumb things, she'll \nhave to enable it explicitely.\n\n> > - easily extend the shallow copy to a larger shallow one, or a full one.\n> \n> Hrm, I think there should also be a way to shrink a repo and \"forget\" old\n> history occasionally (obviously, use of that feature would be highly\n> discouraged).\n\nYes. And you need information about how shallow it used to be. My \nsuggestion was to store that information at a place specific to that \nrepository (see above).\n\n> > > . Run `git fetch git://.../linux-2.6 master`, with a local ref\n> > >  pointing at v2.6.14 commit, to pretend that we have everything\n> > >  up to v2.6.14 to `upload-pack` running on the other end.\n> \n> > How about refs/tags/start_shallow?\n> \n> No, as that would imply that cloning from such a repo is disallowed.\n\nSee above.\n\n> IMO, it may be a lot more robust to just have a list of \"cutoff\" object ids in\n> .git/shallow instead of messing with grafts here, as adding or removing a line\n> from that file is an easier thing to do for porcelain (or by hand) than\n> rewriting the grafts file. Whether that list would be inclusive or exclusive\n> would need to be decided still.\n\nThe functionality of cutoff objects is included in grafts functionality, \nso why should we spend time on reimplementing a subset of features?\n\nIMHO, adding and removing lines from scripts is fragile.\n\nI beg your pardon, you want to edit this information *by hand*? Wow.\n\nCiao,\nDscho\n"},{"id":"15254","messageId":"43DE13B4.8090403@hogyros.de","threadId":"3187","inReplyTo":"Pine.LNX.4.63.0601301305100.20228@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: [RFC] shallow clone","fromName":"Simon Richter","fromEmail":"simon.richter@hogyros.de","sentAt":"2006-01-30T13:25:08Z","receivedAt":"2006-01-30T13:25:08Z","isPatch":false,"sender":{"key":"simon.richter@hogyros.de","avatar":"https://gravatar.com/avatar/1192aa9fa5dd19ce258b12b044cc27111dd7cdb58d920dc2123a02a24f55b5c5?d=mp&s=160"},"body":"Hi,\n\nJohannes Schindelin wrote:\n\n[config as a registry]\n\n> It is becoming sort of a registry: it contains metadata about the current \n> repository, easily available to scripts and programs.\n\nProvided you have a parser that can handle it.\n\n> I beg to differ on your personal opinion on the grounds that the \n> robustness comes from testing, not from diversity. I much prefer to have a \n> well tested config mechanism to having dozens of differently formatted \n> files with less-than-well tested parsers.\n\nIndeed. But we already have a method for associating data values with \nkeys in a hierarchical namespace, and that one is pretty well tested. :-)\n\n>>Why? It's perfectly acceptable to pull from an incomplete repo, as long as you\n>>don't care about the old history.\n\n> Right. But should that be the default? I don't think so. Therefore: \n> disable it, and if the user is absolutely sure to do dumb things, she'll \n> have to enable it explicitely.\n\nWhat harm is done if I have an incomplete repository? It would probably \nmake more sense to emit a warning on clone and explain things if the \nuser tries to go to a version she doesn't have.\n\n>>Hrm, I think there should also be a way to shrink a repo and \"forget\" old\n>>history occasionally (obviously, use of that feature would be highly\n>>discouraged).\n\n> Yes. And you need information about how shallow it used to be. My \n> suggestion was to store that information at a place specific to that \n> repository (see above).\n\nIndeed, but you are keeping this information in two places, namely the \ngrafts file and the config file. This is asking for trouble if they ever \nget out of sync.\n\n>>>How about refs/tags/start_shallow?\n\n>>No, as that would imply that cloning from such a repo is disallowed.\n\n> See above.\n\nWell, I can however see the use case of a developer hosting an \nincomplete repo on a free web service and another developer wanting to \nmerge her changes into her (complete) repo. You would have to \nspecialcase this tag in the fetch operation to avoid copying it over.\n\nWhat's probably worse: You can only have a single cutoff point that way. \nYou probably want multiple in case you want to cut off at a place where \ndevelopment happened in multiple branches that got subsequently merged \ninside the window of objects you keep.\n\n> The functionality of cutoff objects is included in grafts functionality, \n> so why should we spend time on reimplementing a subset of features?\n\nI would ask for the grafts parser to add \"fake\" grafts when it \nencounters the \"shallow\" file. Otherwise, it would be hard to \ndistinguish between grafts the user made when doing interesting merges, \nand grafts that were created to build a shallow repo, because you would \nneed some heuristics to figure out the latter from the former if you \nwant to have a function in your porcelain to \"pull more/all objects\".\n\n> I beg your pardon, you want to edit this information *by hand*? Wow.\n\nYes. That is actually the reason I like git so much: I can repair it by \nhand if something breaks, and this can be done with simple commands. I \ncan remove an object id from a file with \"grep -v\" or perl. I would need \nto fire up an editor or hack a longer script if I wanted to fix \nsomething inside a complex file that does multiple things.\n\n    Simon\n"},{"id":"15268","messageId":"7v8xsxa70o.fsf@assigned-by-dhcp.cox.net","threadId":"3187","inReplyTo":"Pine.LNX.4.63.0601301220420.6424@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: [RFC] shallow clone","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-01-30T18:46:15Z","receivedAt":"2006-01-30T18:46:15Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n>> . Download the full tree for v2.6.14 commit and store its\n>>   objects locally.\n>\n> On first read, I mistook \"tree\" for \"commit\"...\n\nIt turns out that this 'single' request step is unneeded, as\nlong as we implement 'graft' requests.  We can then tell\n\"Cauterize at v2.6.14 and give me the master\" to `upload-pack`.\n`upload-pack` would run `rev-list --objects master`, tries to include\neverything that is reachable from \"master\", but notices that the\nv2.6.14 commit does not have any parent (thanks to the\ncustomized graft) and stops there -- the result is the history\nsince v2.6.14.\n\n>> . Set up `info/grafts` to lie to the local git that Linux kernel\n>>   history began at v2.6.14 version.\n>\n> Maybe also record this in .git/config, so that you can\n>\n> - disallow fetching from this repo, and\n> - easily extend the shallow copy to a larger shallow one, or a full one.\n\nI thought about that before I wrote the message, but it boils\ndown to grepping lines from grafts that have only one object\nname (i.e. cauterizing records), so it is redundant.\n\nAlso there is no strict reason to forbid cloning from such a\nshallow repository.  No harm is done as long as you make it\nclear to somebody who clones from you that what you have is a\nshallow copy, so that the cloned repository can cauterize\nhistory at appropriate places.\n\nA second generation clone, when cloning from a shallow\nrepository, needs to mark itself that it has the same or\nshallower history (otherwise a third generation clone from it\nwould not work), so the `upload-pack` protocol needs to be\nupdated to send grafts information the `upload-pack` side\nusually uses to the downloader even when 'graft' request is not\nused by the downloader.  But once it is done, you should be able\nto clone safely from a shallow repository and end up with a\nrepository with the same (or shallower -- if you asked to make a\nshallow clone from it) history.\n\n> Why not just start another fetch? Then, \"have <refs/tags/start_shallow>\" \n> would be sent, and upload-pack does the right thing?\n\nYes, almost.  We need to realize that `upload-pack` that hears\n\"have A, want B\" is allowed to omit objects that appear in\n`ls-tree B` output but not in `ls-tree A`.  \"have A\" means not\njust \"I have A\", but \"I have A and all of its ancestors\", so\njust sending \"have start_shallow\" (or start_shallow^ for that\nmatter) is not quite enough [*1*].\n\n> If you absolutely want to get only one pack, which then is stored as-is, \n> upload-pack could start two rev-list processes: one for the tree and one \n> for all the rest.\n\nThe message you are responding did two separate transfers (one\n'single', and another 'fetch'); I do not particularly mind doing\ntwo (it is just an initial clone anyway), but as I said it turns\nout that we do not need the initial 'single'.\n\n>> [NOTE]\n>> Most likely this is not directly run by the user but is run as\n>> the first command invoked by the shallow clone script.\n>\n> Better make it an option to git-clone\n\nProbably -- I was just outlining the lowest-level mechanism and\nhaven't thought much about the UI.\n\n[Footnote]\n\n*1* This is true even without more aggressive optimization by\nrev-list that does not exist there yet.  Here is a minimalistic\ndemonstration.  One file project with a handful straight-line\ncommits.  Each change to the file reverts the change made by the\nprevious commit.\n\n * The HEAD commit has \"white\", the HEAD~1 \"black\" and HEAD~2\n   \"white\".\n\n * We say we are interested in things since HEAD~2 (i.e. we\n   pretend that the history starts at HEAD~1 and it does not\n   have a parent) and ask for HEAD.\n\n * Notice that only one copy of the file appears in the output.\n   It is \"black\" blob.  We do not get \"white\" blob because we\n   are telling it that we _have_ HEAD~2.  The resulting set of\n   objects is not enough to check-out the HEAD commit.\n\nThis roughly corresponds to your \"have shallow_start\", but not\nquite -- in that sequence you have objects for HEAD~2 commit.\nBut the point is that I want to leave the door open for\noptimizing upload-pack, so that it can choose to omit objects\nthat do not appear in A when you say \"have A\", if the object\nappears in one of A's ancestors.\n\n-- >8 --\n#!/bin/sh\n\nrm -fr .git\n\ngit init-db\nzebra=white\necho $zebra >file\ngit add file\ngit commit -m initial\n\nfor i in 0 1 2 3 4 5\ndo\n\tcase $zebra in\n\twhite) zebra=black ;;\n\tblack) zebra=white ;;\n\tesac\n\techo $zebra >file\n\tgit commit -a -m \"$i $zebra\"\ndone\ngit rev-list --objects HEAD~2..HEAD |\ngit name-rev --stdin\n"},{"id":"15273","messageId":"7v64o18qn4.fsf@assigned-by-dhcp.cox.net","threadId":"3187","inReplyTo":"43DDFF5C.30803@hogyros.de","subject":"Re: [RFC] shallow clone","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-01-30T19:25:19Z","receivedAt":"2006-01-30T19:25:19Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Simon Richter <Simon.Richter@hogyros.de> writes:\n\n>> - disallow fetching from this repo, and\n>\n> Why? It's perfectly acceptable to pull from an incomplete repo, as\n> long as you don't care about the old history.\n\nI agree.  As long as the cloned one can record itself as a\nshallow one (and with what epochs), I do not see a reason to\nforbid second generation clone from a shallow repository.\n\n> Hrm, I think there should also be a way to shrink a repo and \"forget\"\n> old history occasionally (obviously, use of that feature would be\n> highly discouraged).\n\nI do not think of a reason to discourage it, and I think you can\ndo the \"forgetting\" part with the current set of tools.  Choose\nappropriate cauterizing points, set up info/grafts and running\n\"repack -a -d\" would be sufficient.\n\n> IMO, it may be a lot more robust to just have a list of \"cutoff\"\n> object ids in .git/shallow instead of messing with grafts here, as\n> adding or removing a line from that file is an easier thing to do for\n> porcelain (or by hand) than rewriting the grafts file. Whether that\n> list would be inclusive or exclusive would need to be decided still.\n\nI would rather not to have .git/shallow nor .git/shallow_start.\n\nCauterizing is not any more special than other grafts entries.\nIf you have grafted historical kernel repository behind the\nofficial kernel repository with 2.6.12-rc2 epoch, I do not think\nof any reason to forbid people from cloning such with the\ngrafts.  \n"},{"id":"15274","messageId":"7vzmld7c2g.fsf@assigned-by-dhcp.cox.net","threadId":"3187","inReplyTo":"Pine.LNX.4.63.0601301305100.20228@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: [RFC] shallow clone","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-01-30T19:25:27Z","receivedAt":"2006-01-30T19:25:27Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n>> > - disallow fetching from this repo, and\n>> \n>> Why? It's perfectly acceptable to pull from an incomplete\n>> repo, as long as you don't care about the old history.\n>\n> Right. But should that be the default? I don't think so. Therefore: \n> disable it, and if the user is absolutely sure to do dumb things, she'll \n> have to enable it explicitely.\n\nIf the downstream person wants to have a shallow history of post\nX.org X server core to further hack on it, I do not think of a\nreason why we would want to refuse her from cloning a repository\nof a fellow developer who has already done such a shallow copy.\n\nIf such a clone is done without telling the downstream that the\nresult is a shallow one, it is \"dumb\".  I would agree it should\nnot be done.  We need to propagate the grafts to the downstream\nwhen a clone is done because of this.\n\nBy the way, please refrain from discussing .git/config vs\n.git/eparate-config-files issue in this thread.  My personal\nfeeling so far is that the information current graft represents\nis good enough to support shallow clones, and if not we can\nextend its semantics to support such.  It can be discussed\nindependently if it is a good idea to move the final result\n(grafts with updated semantics) to config file.  Even if we end\nup not doing any of the shallow cloning support we have been\ndiscussing, moving the information in .git/info/grafts to config\nmight make sense.  The issue is tangential.\n"},{"id":"15285","messageId":"cda58cb80601310037s58989b26s@mail.gmail.com","threadId":"3187","inReplyTo":"7v64o18qn4.fsf@assigned-by-dhcp.cox.net","subject":"Re: [RFC] shallow clone","fromName":"Franck","fromEmail":"vagabon.xyz@gmail.com","sentAt":"2006-01-31T08:37:38Z","receivedAt":"2006-01-31T08:37:38Z","isPatch":false,"sender":{"key":"vagabon.xyz@gmail.com","avatar":null},"body":"2006/1/30, Junio C Hamano <junkio@cox.net>:\n> Simon Richter <Simon.Richter@hogyros.de> writes:\n>\n> >> - disallow fetching from this repo, and\n> >\n> > Why? It's perfectly acceptable to pull from an incomplete repo, as\n> > long as you don't care about the old history.\n>\n> I agree.  As long as the cloned one can record itself as a\n> shallow one (and with what epochs), I do not see a reason to\n> forbid second generation clone from a shallow repository.\n>\n\nI agree too\n\n> Cauterizing is not any more special than other grafts entries.\n> If you have grafted historical kernel repository behind the\n> official kernel repository with 2.6.12-rc2 epoch, I do not think\n> of any reason to forbid people from cloning such with the\n> grafts.\n>\n\nI built my public repository from a cautorized one and everybody who\nis pulling from mine is aware of the lack of the full history but they\nactually don't care. If someone is pulling from my repo, he actually\nwants to work on my project which do not need any old thing...\n\nThanks\n--\n               Franck\n"},{"id":"15286","messageId":"7vvew03hls.fsf@assigned-by-dhcp.cox.net","threadId":"3187","inReplyTo":"cda58cb80601310037s58989b26s@mail.gmail.com","subject":"Re: [RFC] shallow clone","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-01-31T08:51:43Z","receivedAt":"2006-01-31T08:51:43Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Franck <vagabon.xyz@gmail.com> writes:\n\n> I built my public repository from a cautorized one and everybody who\n> is pulling from mine is aware of the lack of the full history but they\n> actually don't care. If someone is pulling from my repo, he actually\n> wants to work on my project which do not need any old thing...\n\nMind writing up a howto on the topic?\n\n - How things are set up using the current tool.\n - How others initially clone from you.\n - How others update (pull) from you.\n - What are the pitfalls you and others need to avoid\n   (i.e. operations that involve old history)\n\nI brought this up, because lack of official support of shallow\ncloning was cited as one of the showstopper for a project that\nonce considered switching to git but didn't, from a mailing list\nresearch.\n"},{"id":"15287","messageId":"cda58cb80601310100o6ca1f0a3g@mail.gmail.com","threadId":"3187","inReplyTo":"43DF1F1D.1060704@innova-card.com","subject":"Re: [RFC] shallow clone","fromName":"Franck","fromEmail":"vagabon.xyz@gmail.com","sentAt":"2006-01-31T09:00:16Z","receivedAt":"2006-01-31T09:00:16Z","isPatch":false,"sender":{"key":"vagabon.xyz@gmail.com","avatar":null},"body":"2006/1/31, Franck Bui-Huu <fbh.work@gmail.com>:\n> Junio C Hamano wrote:\n> > Shallow History Cloning\n> > =======================\n> >\n> > One good thing about git repository is that each clone is a\n> > freestanding and complete entity, and you can keep developing in\n> > it offline, without talking to the outside world, knowing that\n> > you can sync with them later when online.\n> >\n\ncould we be able to make a public repository from such repo ?\n\n> > It is also a bad thing.  It gives people working on projects\n> > with long development history stored in CVS a heart attack when\n> > we tell them that their clones need to store the whole history.\n> >\n\nyeah and I haven't survive :)\nI didn't notice that other people were asking for this feature, that's great !\n\n> > There was a suggestion by Linus to allow a partial clone using a\n> > syntax like this:\n\n[snip]\n\n> >\n> > There are some issues.\n> >\n> > . In the fetch above to obtain everything after v2.6.14, and\n> >   future runs of `git fetch origin`, if a blob that is in the\n> >   commit being fetched happens to match what used to be in a\n> >   commit that is older than v2.6.14 (e.g. a patch was reverted),\n> >   `upload-pack` running on the other end is free to omit sending\n> >   it, because we are telling it that we are up to date with\n> >   respect to v2.6.14.  Although I think the current `rev-list\n> >   --objects` implementation does not always do such a revert\n> >   optimization if the revert is to a blob in a revision that is\n> >   sufficiently old, it is free to optimize more aggressively in\n> >   the future.\n> >\n\noops, I wasn't aware of that. I still can resolve this issue by hand, no ?\n\n> > . Later when the user decides to fetch older history, the\n> >   operation can become a bit cumbersome.\n> >\n\n[snip]\n\n> >\n> > Design\n> > ------\n> >\n> > First, to bootstrap the process, we would need to add a way to\n> > obtain all objects associated with a commit.  We could do a new\n> > program, or we could implement this as a protocol extension to\n> > `upload-pack`.  My current inclination is the latter.\n\nis the document in \"Documentation/technical/pack-protocol.txt\"\nuptodate ? I can't find anything on multi_ack for example.\n\n> >\n> > When talking with `upload-pack` that supports this extension,\n> > the downloader can give one commit object name and get a pack\n> > that contains all the objects in the tree associated with that\n> > commit, plus the commit object itself.  This is a rough\n> > equivalent of running the commit walker with the `-t` flag.\n\n[snip]\n\n> >\n> >\n> > Anybody want to try?\n> >\n\nwell, you made almost the job with your analysis, but I've never took\na look to git deep internals and with my lack of time, it would take\ntoo much time...\n\nThanks\n--\n               Franck\n"},{"id":"15290","messageId":"7vmzhc1wz6.fsf_-_@assigned-by-dhcp.cox.net","threadId":"3187","inReplyTo":"7v8xsxa70o.fsf@assigned-by-dhcp.cox.net","subject":"[PATCH] Shallow clone: low level machinery.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-01-31T11:02:37Z","receivedAt":"2006-01-31T11:02:37Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"This adds --shallow=refname option to git-clone-pack, and\nextends upload-pack protocol with \"shallow\" extension.\n\nAn example:\n\n\t$ mkdir junk && cd junk && git init-db\n\t$ git clone-pack --shallow=refs/heads/master ../git.git master\n\nThis creates a very shallow clone of my repository.  It says\n\"pretend refs/heads/master commit is the beginning of time, and\nclone your master branch\".  As before, clone-pack with explicit\nhead name outputs the commit object name and refname to the\nstandard output instead of creating the branch.  The command\ncreates a .git/info/grafts file to cauterize the history at that\ncommit as well.\n\nI think upload-pack side is more or less ready to be debugged,\nbut the client side is highly experimental.  It has quite\nserious limitations and is more of a proof of correctness at the\nprotocol extension level than for practical use:\n\n - Currently it can take only one ---shallow option.\n\n - It has to be spelled in full (refs/heads/master, not\n   \"master\").\n\n - It has to be included as part of explicit refname list.\n\n - There is no matching --shallow in git-fetch-pack.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n\n---\n\n cache.h       |    9 +++\n clone-pack.c  |   69 ++++++++++++++++++++++-\n commit-tree.c |    5 --\n commit.c      |  174 +++++++++++++++++++++++++++++++++++++++------------------\n commit.h      |   14 +++++\n connect.c     |   24 ++++++++\n object.c      |    7 ++\n object.h      |    2 +\n upload-pack.c |   94 +++++++++++++++++++++++++++++--\n 9 files changed, 331 insertions(+), 67 deletions(-)\n\n75f1f4871277f403991c771eb642bdbd6fe82021\ndiff --git a/cache.h b/cache.h\nindex bdbe2d6..18d4cdb 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -111,11 +111,18 @@ static inline unsigned int create_ce_mod\n extern struct cache_entry **active_cache;\n extern unsigned int active_nr, active_alloc, active_cache_changed;\n \n+/*\n+ * Having more than two parents is not strange at all, and this is\n+ * how multi-way merges are represented.\n+ */\n+#define MAXPARENT (16)\n+\n #define GIT_DIR_ENVIRONMENT \"GIT_DIR\"\n #define DEFAULT_GIT_DIR_ENVIRONMENT \".git\"\n #define DB_ENVIRONMENT \"GIT_OBJECT_DIRECTORY\"\n #define INDEX_ENVIRONMENT \"GIT_INDEX_FILE\"\n #define GRAFT_ENVIRONMENT \"GIT_GRAFT_FILE\"\n+#define GRAFT_INFO_ENVIRONMENT \"GIT_GRAFT_INFO\"\n \n extern char *get_git_dir(void);\n extern char *get_object_directory(void);\n@@ -296,6 +303,8 @@ struct ref {\n \tchar name[FLEX_ARRAY]; /* more */\n };\n \n+extern void send_graft_info(int);\n+\n extern int git_connect(int fd[2], char *url, const char *prog);\n extern int finish_connect(pid_t pid);\n extern int path_match(const char *path, int nr, char **match);\ndiff --git a/clone-pack.c b/clone-pack.c\nindex f634431..c1708d5 100644\n--- a/clone-pack.c\n+++ b/clone-pack.c\n@@ -1,15 +1,76 @@\n #include \"cache.h\"\n #include \"refs.h\"\n #include \"pkt-line.h\"\n+#include \"commit.h\"\n \n static const char clone_pack_usage[] =\n-\"git-clone-pack [--exec=<git-upload-pack>] [<host>:]<directory> [<heads>]*\";\n+\"git-clone-pack [--shallow=name] [--exec=<git-upload-pack>] [<host>:]<directory> [<heads>]*\";\n static const char *exec = \"git-upload-pack\";\n+static char *shallow = NULL;\n+\n+static void shallow_exchange(int fd[2], struct ref *ref)\n+{\n+\tchar line[1024];\n+\tchar *graft_file;\n+\tFILE *fp;\n+\tint i, j;\n+\n+\twhile (ref) {\n+\t\tif (!strcmp(ref->name, shallow))\n+\t\t\tbreak;\n+\t\tref = ref->next;\n+\t}\n+\tif (!ref)\n+\t\tdie(\"No matching ref specified for shallow clone %s\",\n+\t\t    shallow);\n+\tif (!server_supports(\"shallow\"))\n+\t\tdie(\"The other end does not support shallow clone\");\n+\tpacket_write(fd[1], \"shallow\\n\");\n+\tpacket_flush(fd[1]);\n+\n+\t/* Read their graft */\n+\tprepare_commit_graft();\n+\tfor (;;) {\n+\t\tint len;\n+\t\tlen = packet_read_line(fd[0], line, sizeof(line));\n+\t\tif (!len)\n+\t\t\tbreak;\n+\t\tadd_graft_info(line);\n+\t}\n+\t/* And cauterize at --shallow=<sha1> */\n+\tsprintf(line, \"%s\\n\", sha1_to_hex(ref->old_sha1));\n+\tadd_graft_info(line);\n+\n+\t/* tell ours */\n+\tpacket_write(fd[1], \"custom\\n\");\n+\tsend_graft_info(fd[1]);\n+\tpacket_flush(fd[1]);\n+\n+\t/* write out ours */\n+\tgraft_file = get_graft_file();\n+\tfp = fopen(graft_file, \"w\");\n+\tif (!fp)\n+\t\tdie(\"cannot update grafts!\");\n+\n+\tfor (i = 0; i < commit_graft_nr; i++) {\n+\t\tstruct commit_graft *g = commit_graft[i];\n+\t\tfputs(sha1_to_hex(g->sha1), fp);\n+\t\tfor (j = 0; j < g->nr_parent; j++) {\n+\t\t\tfputc(' ', fp);\n+\t\t\tfputs(sha1_to_hex(g->parent[j]), fp);\n+\t\t}\n+\t\tfputc('\\n', fp);\n+\t}\n+\tfclose(fp);\n+}\n \n static void clone_handshake(int fd[2], struct ref *ref)\n {\n \tunsigned char sha1[20];\n \n+\tif (shallow)\n+\t\tshallow_exchange(fd, ref);\n+\n \twhile (ref) {\n \t\tpacket_write(fd[1], \"want %s\\n\", sha1_to_hex(ref->old_sha1));\n \t\tref = ref->next;\n@@ -160,6 +221,10 @@ int main(int argc, char **argv)\n \t\t\t\texec = arg + 7;\n \t\t\t\tcontinue;\n \t\t\t}\n+\t\t\tif (!strncmp(\"--shallow=\", arg, 10)) {\n+\t\t\t\tshallow = arg + 10;\n+\t\t\t\tcontinue;\n+\t\t\t}\n \t\t\tusage(clone_pack_usage);\n \t\t}\n \t\tdest = arg;\n@@ -167,6 +232,8 @@ int main(int argc, char **argv)\n \t\tnr_heads = argc - i - 1;\n \t\tbreak;\n \t}\n+\tif (shallow && !nr_heads)\n+\t\tdie(\"shallow clone needs an explicit head name\");\n \tif (!dest)\n \t\tusage(clone_pack_usage);\n \tpid = git_connect(fd, dest, exec);\ndiff --git a/commit-tree.c b/commit-tree.c\nindex 4634b50..cbf2979 100644\n--- a/commit-tree.c\n+++ b/commit-tree.c\n@@ -53,11 +53,6 @@ static void check_valid(unsigned char *s\n \tfree(buf);\n }\n \n-/*\n- * Having more than two parents is not strange at all, and this is\n- * how multi-way merges are represented.\n- */\n-#define MAXPARENT (16)\n static unsigned char parent_sha1[MAXPARENT][20];\n \n static const char commit_tree_usage[] = \"git-commit-tree <sha1> [-p <sha1>]* < changelog\";\ndiff --git a/commit.c b/commit.c\nindex 97205bf..a862287 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -102,12 +102,8 @@ static unsigned long parse_commit_date(c\n \treturn date;\n }\n \n-static struct commit_graft {\n-\tunsigned char sha1[20];\n-\tint nr_parent;\n-\tunsigned char parent[0][20]; /* more */\n-} **commit_graft;\n-static int commit_graft_alloc, commit_graft_nr;\n+struct commit_graft **commit_graft;\n+int commit_graft_alloc, commit_graft_nr;\n \n static int commit_graft_pos(const unsigned char *sha1)\n {\n@@ -128,62 +124,104 @@ static int commit_graft_pos(const unsign\n \treturn -lo - 1;\n }\n \n-static void prepare_commit_graft(void)\n+int add_graft_info(char *buf)\n {\n-\tchar *graft_file = get_graft_file();\n-\tFILE *fp = fopen(graft_file, \"r\");\n+\t/* The format is just \"Commit Parent1 Parent2 ...\\n\" */\n+\tint len = strlen(buf);\n+\tint i;\n+\tstruct commit_graft *graft = NULL;\n+\n+\tif (buf[len-1] == '\\n')\n+\t\tbuf[--len] = 0;\n+\tif (buf[0] == '#')\n+\t\treturn 0;\n+\tif ((len + 1) % 41) {\n+\tbad_graft_data:\n+\t\terror(\"bad graft data: %s\", buf);\n+\t\tfree(graft);\n+\t\treturn -1;\n+\t}\n+\ti = (len + 1) / 41 - 1;\n+\tgraft = xmalloc(sizeof(*graft) + 20 * i);\n+\tgraft->nr_parent = i;\n+\tif (get_sha1_hex(buf, graft->sha1))\n+\t\tgoto bad_graft_data;\n+\tfor (i = 40; i < len; i += 41) {\n+\t\tif (buf[i] != ' ')\n+\t\t\tgoto bad_graft_data;\n+\t\tif (get_sha1_hex(buf + i + 1, graft->parent[i/41]))\n+\t\t\tgoto bad_graft_data;\n+\t}\n+\ti = commit_graft_pos(graft->sha1);\n+\tif (0 <= i) {\n+\t\tfree(commit_graft[i]);\n+\t\tcommit_graft[i] = graft;\n+\t\treturn 0;\n+\t}\n+\ti = -i - 1;\n+\tif (commit_graft_alloc <= ++commit_graft_nr) {\n+\t\tcommit_graft_alloc = alloc_nr(commit_graft_alloc);\n+\t\tcommit_graft = xrealloc(commit_graft,\n+\t\t\t\t\tsizeof(*commit_graft) *\n+\t\t\t\t\tcommit_graft_alloc);\n+\t}\n+\tif (i < commit_graft_nr)\n+\t\tmemmove(commit_graft + i + 1,\n+\t\t\tcommit_graft + i,\n+\t\t\t(commit_graft_nr - i - 1) *\n+\t\t\tsizeof(*commit_graft));\n+\tcommit_graft[i] = graft;\n+\treturn 0;\n+}\n+\n+void clear_commit_graft(void)\n+{\n+\tint i;\n+\tfor (i = 0; i < commit_graft_nr; i++)\n+\t\tfree(commit_graft[i]);\n+\tfree(commit_graft);\n+\tcommit_graft_nr = commit_graft_alloc = 0;\n+\tcommit_graft = NULL;\n+}\n+\n+void prepare_commit_graft(void)\n+{\n+\tchar *graft_file;\n+\tFILE *fp;\n \tchar buf[1024];\n+\n+\tif (getenv(GRAFT_INFO_ENVIRONMENT)) {\n+\t\tchar *cp, *ep;\n+\t\tfor (cp = getenv(GRAFT_INFO_ENVIRONMENT);\n+\t\t     *cp;\n+\t\t     cp = ep) {\n+\t\t\tint more = 0;\n+\t\t\tep = strchr(cp, '\\n');\n+\t\t\tif (ep) {\n+\t\t\t\tmore = 1;\n+\t\t\t\t*ep = '\\0';\n+\t\t\t}\n+\t\t\telse {\n+\t\t\t\tep = cp + strlen(cp);\n+\t\t\t}\n+\t\t\tif (ep != cp)\n+\t\t\t\tadd_graft_info(cp);\n+\t\t\tif (!more)\n+\t\t\t\tbreak;\n+\t\t\t*ep = '\\n';\n+\t\t\tep++;\n+\t\t}\n+\t\treturn;\n+\t}\n+\tgraft_file = get_graft_file();\n+\tfp = fopen(graft_file, \"r\");\n \tif (!fp) {\n-\t\tcommit_graft = (struct commit_graft **) \"hack\";\n+\t\tcommit_graft = (struct commit_graft **) xmalloc(1);\n \t\treturn;\n \t}\n-\twhile (fgets(buf, sizeof(buf), fp)) {\n-\t\t/* The format is just \"Commit Parent1 Parent2 ...\\n\" */\n-\t\tint len = strlen(buf);\n-\t\tint i;\n-\t\tstruct commit_graft *graft = NULL;\n+\twhile (fgets(buf, sizeof(buf), fp))\n+\t\tadd_graft_info(buf);\n \n-\t\tif (buf[len-1] == '\\n')\n-\t\t\tbuf[--len] = 0;\n-\t\tif (buf[0] == '#')\n-\t\t\tcontinue;\n-\t\tif ((len + 1) % 41) {\n-\t\tbad_graft_data:\n-\t\t\terror(\"bad graft data: %s\", buf);\n-\t\t\tfree(graft);\n-\t\t\tcontinue;\n-\t\t}\n-\t\ti = (len + 1) / 41 - 1;\n-\t\tgraft = xmalloc(sizeof(*graft) + 20 * i);\n-\t\tgraft->nr_parent = i;\n-\t\tif (get_sha1_hex(buf, graft->sha1))\n-\t\t\tgoto bad_graft_data;\n-\t\tfor (i = 40; i < len; i += 41) {\n-\t\t\tif (buf[i] != ' ')\n-\t\t\t\tgoto bad_graft_data;\n-\t\t\tif (get_sha1_hex(buf + i + 1, graft->parent[i/41]))\n-\t\t\t\tgoto bad_graft_data;\n-\t\t}\n-\t\ti = commit_graft_pos(graft->sha1);\n-\t\tif (0 <= i) {\n-\t\t\terror(\"duplicate graft data: %s\", buf);\n-\t\t\tfree(graft);\n-\t\t\tcontinue;\n-\t\t}\n-\t\ti = -i - 1;\n-\t\tif (commit_graft_alloc <= ++commit_graft_nr) {\n-\t\t\tcommit_graft_alloc = alloc_nr(commit_graft_alloc);\n-\t\t\tcommit_graft = xrealloc(commit_graft,\n-\t\t\t\t\t\tsizeof(*commit_graft) *\n-\t\t\t\t\t\tcommit_graft_alloc);\n-\t\t}\n-\t\tif (i < commit_graft_nr)\n-\t\t\tmemmove(commit_graft + i + 1,\n-\t\t\t\tcommit_graft + i,\n-\t\t\t\t(commit_graft_nr - i - 1) *\n-\t\t\t\tsizeof(*commit_graft));\n-\t\tcommit_graft[i] = graft;\n-\t}\n \tfclose(fp);\n }\n \n@@ -288,6 +326,30 @@ int parse_commit(struct commit *item)\n \treturn ret;\n }\n \n+static void reparse_commit_parents(struct object *o)\n+{\n+\tstruct commit *c;\n+\tstruct commit_list *parents;\n+\tif ((o->type != commit_type) || !o->parsed)\n+\t\treturn;\n+\tc = (struct commit *)o;\n+\tparents = c->parents;\n+\to->parsed = 0;\n+\twhile (parents) {\n+\t\tstruct commit_list *next = parents->next;\n+\t\tfree(parents);\n+\t\tparents = next;\n+\t}\n+\tc->parents = NULL;\n+\tfree(c->buffer);\n+\tc->buffer = NULL;\n+}\n+\n+void reparse_all_parsed_commits(void)\n+{\n+\tfor_each_object(reparse_commit_parents);\n+}\n+\n struct commit_list *commit_list_insert(struct commit *item, struct commit_list **list_p)\n {\n \tstruct commit_list *new_list = xmalloc(sizeof(struct commit_list));\ndiff --git a/commit.h b/commit.h\nindex 986b22d..abc5b9e 100644\n--- a/commit.h\n+++ b/commit.h\n@@ -17,6 +17,20 @@ struct commit {\n \tchar *buffer;\n };\n \n+struct commit_graft {\n+\tunsigned char sha1[20];\n+\tint nr_parent;\n+\tunsigned char parent[0][20]; /* more */\n+};\n+\n+extern struct commit_graft **commit_graft;\n+extern int commit_graft_alloc, commit_graft_nr;\n+\n+extern void prepare_commit_graft(void);\n+extern void clear_commit_graft(void);\n+extern int add_graft_info(char *);\n+extern void reparse_all_parsed_commits(void);\n+\n extern int save_commit_buffer;\n extern const char *commit_type;\n \ndiff --git a/connect.c b/connect.c\nindex 3f2d65c..046d1da 100644\n--- a/connect.c\n+++ b/connect.c\n@@ -3,6 +3,7 @@\n #include \"pkt-line.h\"\n #include \"quote.h\"\n #include \"refs.h\"\n+#include \"commit.h\"\n #include <sys/wait.h>\n #include <sys/socket.h>\n #include <netinet/in.h>\n@@ -298,6 +299,29 @@ int match_refs(struct ref *src, struct r\n \treturn 0;\n }\n \n+void send_graft_info(int outfd)\n+{\n+\tint i, j;\n+\tchar packet_buf[41*MAXPARENT], *buf;\n+\n+\tfor (i = 0; i < commit_graft_nr; i++) {\n+\t\tstruct commit_graft *g = commit_graft[i];\n+\t\tbuf = packet_buf;\n+\t\tmemcpy(buf, sha1_to_hex(g->sha1), 40);\n+\t\tbuf += 40;\n+\t\tif (MAXPARENT <= g->nr_parent)\n+\t\t\tdie(\"insanely big octopus graft with %d parents: %s\",\n+\t\t\t    g->nr_parent, sha1_to_hex(g->sha1));\n+\t\tfor (j = 0; j < g->nr_parent; j++) {\n+\t\t\t*buf++ = ' ';\n+\t\t\tmemcpy(buf, sha1_to_hex(g->parent[j]), 40);\n+\t\t\tbuf += 40;\n+\t\t}\n+\t\t*buf = 0;\n+\t\tpacket_write(outfd, \"%s\\n\", packet_buf);\n+\t}\n+}\n+\n enum protocol {\n \tPROTO_LOCAL = 1,\n \tPROTO_SSH,\ndiff --git a/object.c b/object.c\nindex 1577f74..bbcfcd8 100644\n--- a/object.c\n+++ b/object.c\n@@ -252,3 +252,10 @@ int object_list_contains(struct object_l\n \t}\n \treturn 0;\n }\n+\n+void for_each_object(void (*fn)(struct object *))\n+{\n+\tint i;\n+\tfor (i = 0; i < nr_objs; i++)\n+\t\tfn(objs[i]);\n+}\ndiff --git a/object.h b/object.h\nindex 0e76182..b4c9729 100644\n--- a/object.h\n+++ b/object.h\n@@ -55,4 +55,6 @@ unsigned object_list_length(struct objec\n \n int object_list_contains(struct object_list *list, struct object *obj);\n \n+void for_each_object(void (*)(struct object *));\n+\n #endif /* OBJECT_H */\ndiff --git a/upload-pack.c b/upload-pack.c\nindex d198055..90ea549 100644\n--- a/upload-pack.c\n+++ b/upload-pack.c\n@@ -13,11 +13,16 @@ static const char upload_pack_usage[] = \n #define WANTED (1U << 2)\n #define MAX_HAS 256\n #define MAX_NEEDS 256\n-static int nr_has = 0, nr_needs = 0, multi_ack = 0, nr_our_refs = 0;\n+#define MAX_PARENTS 20\n+static int nr_has = 0, nr_needs = 0, nr_our_refs = 0;\n static unsigned char has_sha1[MAX_HAS][20];\n static unsigned char needs_sha1[MAX_NEEDS][20];\n static unsigned int timeout = 0;\n \n+/* protocol extensions */\n+static int multi_ack = 0;\n+static int using_custom_graft = 0;\n+\n static void reset_timeout(void)\n {\n \talarm(timeout);\n@@ -163,6 +168,77 @@ static int get_common_commits(void)\n \t}\n }\n \n+static void exchange_grafts(void)\n+{\n+\tint len;\n+\tchar line[41*MAX_PARENTS];\n+\n+\t/* We heard \"shallow\"; drop up to the next flush */\n+\tfor (;;) {\n+\t\tlen = packet_read_line(0, line, sizeof(line));\n+\t\treset_timeout();\n+\t\tif (!len)\n+\t\t\tbreak;\n+\t}\n+\n+\t/* Send our graft */\n+\tprepare_commit_graft();\n+\tsend_graft_info(1);\n+\tpacket_flush(1);\n+\n+\t/* For precise common commits discovery, we need to use\n+\t * the graft information we received from them.\n+\t * But this is expensive, so the downloader first says\n+\t * if it wants to use our graft as is.\n+\t */\n+\tlen = packet_read_line(0, line, sizeof(line));\n+\treset_timeout();\n+\tif (!len)\n+\t\t; /* use ours as is */\n+\telse if (!strcmp(line, \"custom\\n\")) {\n+\t\tusing_custom_graft = 1;\n+\t\tclear_commit_graft();\n+\t\tfor (;;) {\n+\t\t\tlen = packet_read_line(0, line, sizeof(line));\n+\t\t\treset_timeout();\n+\t\t\tif (!len)\n+\t\t\t\tbreak;\n+\t\t\tif (add_graft_info(line))\n+\t\t\t\tdie(\"Bad graft line %s\", line);\n+\t\t}\n+\t\t/* And using that, we prepare our end. */\n+\t\treparse_all_parsed_commits();\n+\t}\n+\telse\n+\t\tdie(\"expected 'custom', got '%s'\", line);\n+}\n+\n+static void setup_custom_graft(void)\n+{\n+\tchar *graft_env = strdup(GRAFT_INFO_ENVIRONMENT \"=\");\n+\tint envlen = strlen(graft_env);\n+\tint i, j;\n+\n+\tfor (i = 0; i < commit_graft_nr; i++) {\n+\t\tstruct commit_graft *g = commit_graft[i];\n+\t\tchar buf[41*MAX_PARENTS], *ptr;\n+\t\tptr = buf;\n+\t\tmemcpy(ptr, sha1_to_hex(g->sha1), 40);\n+\t\tptr += 40;\n+\t\tfor (j = 0; j < g->nr_parent; j++) {\n+\t\t\t*ptr++ = ' ';\n+\t\t\tmemcpy(ptr, sha1_to_hex(g->parent[j]), 40);\n+\t\t\tptr += 40;\n+\t\t}\n+\t\t*ptr++ = '\\n';\n+\t\t*ptr = 0;\n+\t\tgraft_env = xrealloc(graft_env, envlen + (ptr - buf));\n+\t\tmemcpy(graft_env + envlen, buf, ptr - buf + 1);\n+\t\tenvlen += ptr - buf;\n+\t}\n+\tputenv(graft_env);\n+}\n+\n static int receive_needs(void)\n {\n \tstatic char line[1000];\n@@ -180,16 +256,22 @@ static int receive_needs(void)\n \t\tsha1_buf = dummy;\n \t\tif (needs == MAX_NEEDS) {\n \t\t\tfprintf(stderr,\n-\t\t\t\t\"warning: supporting only a max of %d requests. \"\n+\t\t\t\t\"warning: supporting only a max of \"\n+\t\t\t\t\"%d requests. \"\n \t\t\t\t\"sending everything instead.\\n\",\n \t\t\t\tMAX_NEEDS);\n \t\t}\n \t\telse if (needs < MAX_NEEDS)\n \t\t\tsha1_buf = needs_sha1[needs];\n \n-\t\tif (strncmp(\"want \", line, 5) || get_sha1_hex(line+5, sha1_buf))\n+\t\tif (!strcmp(\"shallow\\n\", line)) {\n+\t\t\texchange_grafts();\n+\t\t\tcontinue;\n+\t\t}\n+\t\tif (strncmp(\"want \", line, 5) ||\n+\t\t    get_sha1_hex(line+5, sha1_buf))\n \t\t\tdie(\"git-upload-pack: protocol error, \"\n-\t\t\t    \"expected to get sha, not '%s'\", line);\n+\t\t\t    \"expected to get want-sha1, not '%s'\", line);\n \t\tif (strstr(line+45, \"multi_ack\"))\n \t\t\tmulti_ack = 1;\n \n@@ -213,7 +295,7 @@ static int receive_needs(void)\n \n static int send_ref(const char *refname, const unsigned char *sha1)\n {\n-\tstatic char *capabilities = \"multi_ack\";\n+\tstatic char *capabilities = \"multi_ack shallow\";\n \tstruct object *o = parse_object(sha1);\n \n \tif (capabilities)\n@@ -243,6 +325,8 @@ static int upload_pack(void)\n \tif (!nr_needs)\n \t\treturn 0;\n \tget_common_commits();\n+\tif (using_custom_graft)\n+\t\tsetup_custom_graft();\n \tcreate_pack_file();\n \treturn 0;\n }\n-- \n1.1.6.gefef\n"},{"id":"15291","messageId":"cda58cb80601310311v45531fb5u@mail.gmail.com","threadId":"3187","inReplyTo":"7vvew03hls.fsf@assigned-by-dhcp.cox.net","subject":"Re: [RFC] shallow clone","fromName":"Franck","fromEmail":"vagabon.xyz@gmail.com","sentAt":"2006-01-31T11:11:53Z","receivedAt":"2006-01-31T11:11:53Z","isPatch":false,"sender":{"key":"vagabon.xyz@gmail.com","avatar":null},"body":"2006/1/31, Junio C Hamano <junkio@cox.net>:\n> Franck <vagabon.xyz@gmail.com> writes:\n>\n> > I built my public repository from a cautorized one and everybody who\n> > is pulling from mine is aware of the lack of the full history but they\n> > actually don't care. If someone is pulling from my repo, he actually\n> > wants to work on my project which do not need any old thing...\n>\n> Mind writing up a howto on the topic?\n\nok I'll try to sum-up something this week, hope my bad english will be\nunderstandable...\n\n>\n>  - How things are set up using the current tool.\n>  - How others initially clone from you.\n>  - How others update (pull) from you.\n>  - What are the pitfalls you and others need to avoid\n>    (i.e. operations that involve old history)\n\nactually I just discovered one thanks to your first email for this\nthread about reverted commit...So I'm not very the one for this\nsection...\n\n>\n> I brought this up, because lack of official support of shallow\n> cloning was cited as one of the showstopper for a project that\n> once considered switching to git but didn't, from a mailing list\n> research.\n\nagain I wasn't aware that this feature is really needed...\n\nthanks\n--\n               Franck\n"},{"id":"15292","messageId":"Pine.LNX.4.63.0601311127490.25248@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3187","inReplyTo":"7vzmld7c2g.fsf@assigned-by-dhcp.cox.net","subject":"Re: [RFC] shallow clone","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-01-31T11:28:39Z","receivedAt":"2006-01-31T11:28:39Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 30 Jan 2006, Junio C Hamano wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> >> > - disallow fetching from this repo, and\n> >> \n> >> Why? It's perfectly acceptable to pull from an incomplete\n> >> repo, as long as you don't care about the old history.\n> >\n> > Right. But should that be the default? I don't think so. Therefore: \n> > disable it, and if the user is absolutely sure to do dumb things, she'll \n> > have to enable it explicitely.\n> \n> If the downstream person wants to have a shallow history of post\n> X.org X server core to further hack on it, I do not think of a\n> reason why we would want to refuse her from cloning a repository\n> of a fellow developer who has already done such a shallow copy.\n\nOkay. But in their case, they'll probably do what was done with Linux: \nstart afresh. If you want to have the old history, you can import it and \nmerge it via a graft.\n\n> If such a clone is done without telling the downstream that the\n> result is a shallow one, it is \"dumb\".  I would agree it should\n> not be done.\n\nThat was my point. As long as you don't make sure the client handles the \nshallow upstream gracefully, it is dangerous. At the moment, there are too \nmany code parts relying on the completeness of the repository (local and \nremote).\n\nSince I wrote this, I realized that the problem I saw is not limited to \nshallow upstream, but there is a subtle issue with shallow downstreams, \ntoo:\n\nJust imagine this: Alice starts a project, Bob makes a shallow copy from \nit when Alice just reverted an experimental feature. Then, Alice decides \nthe experimental feature was not bad at all and reverts the revert. Bob \npulls from Alice: Alice's upload-pack assumes Bob already has the original \nfiles (now re-reverted), and Bob ends up with a broken repository.\n\nWhile writing the last paragraph, it became clear to me that the shallow \nthing is very fragile: IMHO it is impossible to be fully backwards \ncompatible (remember: you should not force anybody to upgrade).\n\n> By the way, please refrain from discussing .git/config vs \n> .git/eparate-config-files issue in this thread.\n\nOkay. I will shut up on that issue.\n\n> My personal feeling so far is that the information current graft \n> represents is good enough to support shallow clones, and if not we can \n> extend its semantics to support such.\n\nNo. The grafts are more powerful. I have quite a few repos here in which I \nheavily work with grafts, and they are no cutoffs for shallow repos. They \nare hard links between different lines of development. For example, I use \nthem to map merges in cvsimported projects, thus fixing a shortcoming of \nCVS. Also, you can \"add\" history.\n\nIf you now rely on the grafts file to determine what was a cutoff, you may \nwell end up with bogus cutoffs.\n\nCiao,\nDscho\n"},{"id":"15294","messageId":"43DF608C.1060201@hogyros.de","threadId":"3187","inReplyTo":"Pine.LNX.4.63.0601311127490.25248@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: [RFC] shallow clone","fromName":"Simon Richter","fromEmail":"simon.richter@hogyros.de","sentAt":"2006-01-31T13:05:16Z","receivedAt":"2006-01-31T13:05:16Z","isPatch":false,"sender":{"key":"simon.richter@hogyros.de","avatar":"https://gravatar.com/avatar/1192aa9fa5dd19ce258b12b044cc27111dd7cdb58d920dc2123a02a24f55b5c5?d=mp&s=160"},"body":"Hi,\n\nJohannes Schindelin wrote:\n\n>>If the downstream person wants to have a shallow history of post\n>>X.org X server core to further hack on it, I do not think of a\n>>reason why we would want to refuse her from cloning a repository\n>>of a fellow developer who has already done such a shallow copy.\n\n> Okay. But in their case, they'll probably do what was done with Linux: \n> start afresh. If you want to have the old history, you can import it and \n> merge it via a graft.\n\nWell, in the Linux case the problem was not knowing what the SHA1 sum of \nthe entire Linux history was. In the shallow repo case we know it, so \nthere is no point in throwing away that information.\n\n>>If such a clone is done without telling the downstream that the\n>>result is a shallow one, it is \"dumb\".  I would agree it should\n>>not be done.\n\n> That was my point. As long as you don't make sure the client handles the \n> shallow upstream gracefully, it is dangerous. At the moment, there are too \n> many code parts relying on the completeness of the repository (local and \n> remote).\n\nWell, the important thing would be that commands that can work (a merge \nonly needs to find the most recent common ancestor, etc) do work, and \ncommands that cannot (\"log\") emit sensible diagnostics.\n\n> Just imagine this: Alice starts a project, Bob makes a shallow copy from \n> it when Alice just reverted an experimental feature. Then, Alice decides \n> the experimental feature was not bad at all and reverts the revert. Bob \n> pulls from Alice: Alice's upload-pack assumes Bob already has the original \n> files (now re-reverted), and Bob ends up with a broken repository.\n\nI know far too little about the internal workings for that, but I'd \nassume that in this case Bob's copy starts at the commit that was never \nin question (and he never saw the reverted commit), and Alice's contains \na commit on top of that. That one should work. But the other way 'round \nis problematic, when Bob starts with a commit that has been reverted in \nAlice's repository. The solution is for Bob to ask Alice's repo for the \ncommon ancestor of his shallow base and Alice's HEAD. Alice's repo can, \nhowever, fail to deliver these if there has been a purge since, in that \ncase, stuff needs to be merged by hand (but you already have a problem \nif someone clones your repo before you revert changes, so no regression \nhere).\n\n> If you now rely on the grafts file to determine what was a cutoff, you may \n> well end up with bogus cutoffs.\n\nExactly that was my concern earlier; my database design gut feeling \ntells me that information duplication is not good either, hence my \nsuggestion to split off these grafts into a separate file in order to \nmark them as cutoff points.\n\n    Simon\n"},{"id":"15296","messageId":"Pine.LNX.4.63.0601311422400.7918@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3187","inReplyTo":"43DF608C.1060201@hogyros.de","subject":"Re: [RFC] shallow clone","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-01-31T13:31:03Z","receivedAt":"2006-01-31T13:31:03Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 31 Jan 2006, Simon Richter wrote:\n\n> Well, the important thing would be that commands that can work (a merge only\n> needs to find the most recent common ancestor, etc) do work, and commands that\n> cannot (\"log\") emit sensible diagnostics.\n\nNo it would not.\n\nA commit is a very small object which points (among others) to a tree \nobject.\n\nA tree object corresponds to a directory (that is, it can point to a \nnumber of tree and blob objects).\n\nA blob object corresponds to a file (that is, git never parses its \ncontents).\n\nIf two separate revisions contain the same file (i.e. same contents), this \nis not duplicated, but the corresponding tree objects point to the same \nobject.\n\nIf you pull, upload-pack will think you have *every* object depending on \nevery ref you have stored.\n\nSay you have three revisions, A -> B -> C, and A and C contain the \nsame file bla.txt, and the client says it has B, the upstream upload-pack \nassumes you have bla.txt.\n\n> I know far too little about the internal workings for that, [...]\n\nI hope I clarified the important aspect.\n\n> > If you now rely on the grafts file to determine what was a cutoff, you may\n> > well end up with bogus cutoffs.\n> \n> Exactly that was my concern earlier; my database design gut feeling tells me\n> that information duplication is not good either, [...]\n\nYou only have two choices: you proposed code duplication, and yours truly \nproposed data duplication.\n\nAs is known from good database design: a few redundancies here and there \nare typically needed for good performance.\n\nCiao,\nDscho\n"},{"id":"15297","messageId":"Pine.LNX.4.63.0601311449040.8033@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3187","inReplyTo":"7vmzhc1wz6.fsf_-_@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] Shallow clone: low level machinery.","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-01-31T13:58:21Z","receivedAt":"2006-01-31T13:58:21Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\napart from my thinking this is not backward-compatible (you are supposed \nto be able to pull from a complete repo, even if it has a \nnon-shallow-capable upload-pack), here are my comments:\n\n- it is good that MAXPARENT and struct commit_graft are in more public \n\tplaces now.\n\n- reparse_* is misleading. Nothing is reparsed, but rather \"unparsed\".\n\n- I'd hesitate to let git-daemon write temporary files. That is a whole \n\tnew can of security worms.\n\n- It looks wrong to me to define MAX_PARENTS as 20 in upload-pack.c, when \n\tMAXPARENT is defined as 16 in cache.h.\n\n- The custom_graft issue could be handled in a more elegant manner if \n\tgit was lib'ified (no temporary file). Since that is already the \n\tplan, why not do that first, and come back later?\n\nCiao,\nDscho\n"},{"id":"15298","messageId":"Pine.LNX.4.63.0601311518120.9824@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3187","inReplyTo":"7v8xsxa70o.fsf@assigned-by-dhcp.cox.net","subject":"Re: [RFC] shallow clone","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-01-31T14:20:36Z","receivedAt":"2006-01-31T14:20:36Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 30 Jan 2006, Junio C Hamano wrote:\n\n> We need to realize that `upload-pack` that hears\n> \"have A, want B\" is allowed to omit objects that appear in\n> `ls-tree B` output but not in `ls-tree A`.  \"have A\" means not\n> just \"I have A\", but \"I have A and all of its ancestors\", so\n> just sending \"have start_shallow\" (or start_shallow^ for that\n> matter) is not quite enough.\n\nSo how about adding a \"have-single A\" which would be translated to \n\"git-rev-list ~A\", which in turn would only mark the tree and its \nchildren, but not the parents?\n\nCiao,\nDscho\n"},{"id":"15299","messageId":"43DF72E6.4050802@hogyros.de","threadId":"3187","inReplyTo":"Pine.LNX.4.63.0601311422400.7918@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: [RFC] shallow clone","fromName":"Simon Richter","fromEmail":"simon.richter@hogyros.de","sentAt":"2006-01-31T14:23:34Z","receivedAt":"2006-01-31T14:23:34Z","isPatch":false,"sender":{"key":"simon.richter@hogyros.de","avatar":"https://gravatar.com/avatar/1192aa9fa5dd19ce258b12b044cc27111dd7cdb58d920dc2123a02a24f55b5c5?d=mp&s=160"},"body":"Hi,\n\nJohannes Schindelin wrote:\n\n> If you pull, upload-pack will think you have *every* object depending on \n> every ref you have stored.\n\nAh, okay. That was the missing information, thanks.\n\n> You only have two choices: you proposed code duplication, and yours truly \n> proposed data duplication.\n\nErm, if there are multiple places for parsing a grafts file, that needs \nto be addressed as well.\n\n> As is known from good database design: a few redundancies here and there \n> are typically needed for good performance.\n\nSure, but only if you can \"rebuild\" all the redundant information reliably.\n\n    Simon\n"},{"id":"15309","messageId":"7vd5i81e4e.fsf@assigned-by-dhcp.cox.net","threadId":"3187","inReplyTo":"Pine.LNX.4.63.0601311449040.8033@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: [PATCH] Shallow clone: low level machinery.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-01-31T17:49:53Z","receivedAt":"2006-01-31T17:49:53Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> apart from my thinking this is not backward-compatible (you are supposed \n> to be able to pull from a complete repo, even if it has a \n> non-shallow-capable upload-pack), here are my comments:\n\nIt cannot do a shallow clone against older servers, no.  I think\nit should be able to do a full clone from older servers, but I\nneed to double check -- at least that is how I meant to write\nthat thing but it was late night ;-).\n\n> - it is good that MAXPARENT and struct commit_graft are in more public \n> \tplaces now.\n>\n> - reparse_* is misleading. Nothing is reparsed, but rather \"unparsed\".\n\nI meant to reparse them thear but forgot.  Will remember to fix.\n\n> - I'd hesitate to let git-daemon write temporary files. That is a whole \n> \tnew can of security worms.\n>\n> - The custom_graft issue could be handled in a more elegant manner if \n> \tgit was lib'ified (no temporary file). Since that is already the \n> \tplan, why not do that first, and come back later?\n\nThat is why it does not write any temporary files.  It\nintroduces a way to read graft information from an environment\nvariable.\n\n> - It looks wrong to me to define MAX_PARENTS as 20 in upload-pack.c, when \n> \tMAXPARENT is defined as 16 in cache.h.\n\nThis is remnant from my earlier one that did not move MAXPARENT\nout from commit-tree I forgot to clean up before calling it a\nday.  Will remember to clean up.\n"},{"id":"15311","messageId":"Pine.LNX.4.63.0601311904410.10944@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3187","inReplyTo":"7vd5i81e4e.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] Shallow clone: low level machinery.","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-01-31T18:06:28Z","receivedAt":"2006-01-31T18:06:28Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 31 Jan 2006, Junio C Hamano wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> > apart from my thinking this is not backward-compatible (you are supposed \n> > to be able to pull from a complete repo, even if it has a \n> > non-shallow-capable upload-pack), here are my comments:\n> \n> It cannot do a shallow clone against older servers, no.\n\nWorse, you cannot pull from older servers into shallow repos.\n\n> > - The custom_graft issue could be handled in a more elegant manner if \n> > \tgit was lib'ified (no temporary file). Since that is already the \n> > \tplan, why not do that first, and come back later?\n> \n> That is why it does not write any temporary files.  It\n> introduces a way to read graft information from an environment\n> variable.\n\nOoops. I only saw that you setup_custom_grafts and assumed wrongly.\n\nCiao,\nDscho\n"},{"id":"15313","messageId":"7vzmlcz28x.fsf@assigned-by-dhcp.cox.net","threadId":"3187","inReplyTo":"Pine.LNX.4.63.0601311904410.10944@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: [PATCH] Shallow clone: low level machinery.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-01-31T18:22:22Z","receivedAt":"2006-01-31T18:22:22Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> Worse, you cannot pull from older servers into shallow repos.\n\n\"have X\" means different thing if you do not have matching\ngrafts information, so I suspect that is fundamentally\nunsolvable.\n\nI am not sure you can convince \"git-rev-list ^A\" to mean \"not at\nA but things before that is still interesting\", especially when\nyou give many other heads to start traversing from, but if you\ncan, then you can do things at rev-list command line parameter\nlevel without doing the \"exchange and use the same grafts\"\ntrickery.  That _might_ be easier to implement but I do not see\nan obvious correctness guarantee in the approach.\n\nImplementation bugs aside, it is obvious the things _would_ work \ncorrectly with \"exchange and use the same grafts\" approach.\n"},{"id":"15327","messageId":"7vr76oun9o.fsf@assigned-by-dhcp.cox.net","threadId":"3187","inReplyTo":"7v8xsxa70o.fsf@assigned-by-dhcp.cox.net","subject":"Re: [RFC] shallow clone","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-01-31T20:59:31Z","receivedAt":"2006-01-31T20:59:31Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"This is whacky, but another completely different strategy is to\nintroduce remote alternates.\n\nIf you can allow objects/info/alternates to name a repository\nthat is not on the local disk, we can set the original remote\nrepository we \"clone\" from as one of the alternates, and teach\nread_sha1_file() to locally cache objects we read from remote\nalternates.\n\nAfter such a \"shallow clone\", the user may want to prime the\ncache by something like:\n\n\t$ git-rev-list --objects v2.6.14..master |\n          git-pack-objects --stdout >/dev/null\n\nbefore going offline.  Obviously you can keep the resulting pack\ninstead of leaving things loose.\n\nI am not seriously advocating this yet -- adding calls to http\nand git transfer machinery in read_sha1_file(), which is as low\nlevel as you can go, is not something I have guts to do at the\nmoment.\n"},{"id":"15395","messageId":"Pine.LNX.4.63.0602011528030.28923@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3187","inReplyTo":"7vzmlcz28x.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] Shallow clone: low level machinery.","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-02-01T14:33:51Z","receivedAt":"2006-02-01T14:33:51Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 31 Jan 2006, Junio C Hamano wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> > Worse, you cannot pull from older servers into shallow repos.\n> \n> \"have X\" means different thing if you do not have matching\n> grafts information, so I suspect that is fundamentally\n> unsolvable.\n\nIf the shallow-capable client could realize that the server is not \nshallow-capable *and* the local repo is shallow, and refuse to operate \n(unless called with \"-f\", in which case the result may or may not be a \nbroken repo, which has to be fixed up manually by copying \nover ORIG_HEAD to HEAD).\n\nOf course, the client has to know that the local repo is shallow, which it \nmust not determine by looking at the grafts file.\n \n> I am not sure you can convince \"git-rev-list ^A\" to mean \"not at\n> A but things before that is still interesting\", especially when\n> you give many other heads to start traversing from, but if you\n> can, then you can do things at rev-list command line parameter\n> level without doing the \"exchange and use the same grafts\"\n> trickery.  That _might_ be easier to implement but I do not see\n> an obvious correctness guarantee in the approach.\n\nIf you introduce a different \"have X\" -- like \"have-no-parent X\" -- and \nteach git-rev-list that \"~A\" means \"traverse the tree of A, but not A's \nparents\", you'd basically have everything you need, right?\n\n> Implementation bugs aside, it is obvious the things _would_ work \n> correctly with \"exchange and use the same grafts\" approach.\n\nYes, I agree. But again, the local repo has to know which grafts were \nintroduced by making the repo shallow.\n\nCiao,\nDscho\n"},{"id":"15397","messageId":"Pine.LNX.4.63.0602011545390.28923@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3187","inReplyTo":"7vr76oun9o.fsf@assigned-by-dhcp.cox.net","subject":"Re: [RFC] shallow clone","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-02-01T14:47:24Z","receivedAt":"2006-02-01T14:47:24Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 31 Jan 2006, Junio C Hamano wrote:\n\n> This is whacky, but another completely different strategy is to\n> introduce remote alternates.\n\nI'd rather go with the original plan. After all, you do not really need \nthe cut-off commit objects. All needed objects are available on the server \nside: it just has to have a way to know which ones to send.\n\nCiao,\nDscho\n"},{"id":"15423","messageId":"7vbqxqbz9q.fsf@assigned-by-dhcp.cox.net","threadId":"3187","inReplyTo":"Pine.LNX.4.63.0602011528030.28923@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: [PATCH] Shallow clone: low level machinery.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-02-01T20:27:29Z","receivedAt":"2006-02-01T20:27:29Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n>> > Worse, you cannot pull from older servers into shallow repos.\n>> \n>> \"have X\" means different thing if you do not have matching\n>> grafts information, so I suspect that is fundamentally\n>> unsolvable.\n>\n> If the shallow-capable client could realize that the server is not \n> shallow-capable *and* the local repo is shallow, and refuse to operate \n> (unless called with \"-f\", in which case the result may or may not be a \n> broken repo, which has to be fixed up manually by copying \n> over ORIG_HEAD to HEAD).\n\n\"If ... refuse to operate\" then?  If \"Then that is OK\" is what\nyou meant to say I agree (I meant to code the client code that\nway but I started only with the initial clone).  I said\n\"fundamentally unsolvable\" because I thought you wanted it to do\nsomething sensible without refusing even in such a case.\n\n> Of course, the client has to know that the local repo is shallow, which it \n> must not determine by looking at the grafts file.\n\nSorry, I fail to understand this requirement.  Why is it \"it must not\"?\n\n> If you introduce a different \"have X\" -- like \"have-no-parent X\" -- and \n> teach git-rev-list that \"~A\" means \"traverse the tree of A, but not A's \n> parents\", you'd basically have everything you need, right?\n\nIf you have such a modified rev-list, yes.  I was having doubts\nabout keeping an obvious correctness guarantee when doing such\n\"rev-list ~A\".\n\n> Yes, I agree. But again, the local repo has to know which grafts were \n> introduced by making the repo shallow.\n\nI am not sure I understand.  grafts are grafts are grafts.  If\nthe other side has grafts to connect otherwise unrelated commit\nobjects, I suspect the cloner needs to know about them, all of\nthem, in order to use the resulting clone.  Also the upstream\nside would need to know the altered world view the cloner has to\nadjust the commit ancestry graph, at least during the cloning\nand fetching, and I do not think it should be limited only to\ncauterizign entries created by earlier shallow clone operations.\nManually created cauterizing entries should also count (for that\nmatter, grafts to stitch unrelated lines together), No?\n"},{"id":"15455","messageId":"Pine.LNX.4.63.0602020113200.30910@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3187","inReplyTo":"7vbqxqbz9q.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] Shallow clone: low level machinery.","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-02-02T00:48:31Z","receivedAt":"2006-02-02T00:48:31Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 1 Feb 2006, Junio C Hamano wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> >> > Worse, you cannot pull from older servers into shallow repos.\n> >> \n> >> \"have X\" means different thing if you do not have matching\n> >> grafts information, so I suspect that is fundamentally\n> >> unsolvable.\n> >\n> > If the shallow-capable client could realize that the server is not \n> > shallow-capable *and* the local repo is shallow, and refuse to operate \n> > (unless called with \"-f\", in which case the result may or may not be a \n> > broken repo, which has to be fixed up manually by copying \n> > over ORIG_HEAD to HEAD).\n> \n> \"If ... refuse to operate\" then?\n\nJust skip the \"If\". I'll start to enclose all emails I write in \n<tired>..</tired> blocks.\n\n> > Of course, the client has to know that the local repo is shallow, which it \n> > must not determine by looking at the grafts file.\n> \n> Sorry, I fail to understand this requirement.  Why is it \"it must not\"?\n\nSee below.\n\n> > If you introduce a different \"have X\" -- like \"have-no-parent X\" -- and \n> > teach git-rev-list that \"~A\" means \"traverse the tree of A, but not A's \n> > parents\", you'd basically have everything you need, right?\n> \n> If you have such a modified rev-list, yes.  I was having doubts\n> about keeping an obvious correctness guarantee when doing such\n> \"rev-list ~A\".\n\nI think it would be trivial: just resolve ~A to the tree A points to:\n\n-- snip --\n[PATCH] rev-list: Support \"~treeish\"\n\nNow, \"git rev-list --objects ~some_rev\" traverses just the tree of\nsome_rev.\n\nSigned-off-by: Johannes Schindelin <Johannes.Schindelin@gmx.de>\n\n---\n\n rev-list.c |   23 +++++++++++++++++++++++\n 1 files changed, 23 insertions(+), 0 deletions(-)\n\n43267e65c9ad933ad1a49005c4b61c23adaec372\ndiff --git a/rev-list.c b/rev-list.c\nindex 8012762..a196110 100644\n--- a/rev-list.c\n+++ b/rev-list.c\n@@ -720,6 +720,24 @@ static void handle_one_commit(struct com\n \tcommit_list_insert(com, lst);\n }\n \n+static void handle_tree(const unsigned char *sha1)\n+{\n+\tstruct object *object;\n+\n+\tobject = parse_object(sha1);\n+\tif (!object)\n+\t\tdie(\"bad object %s\", sha1_to_hex(sha1));\n+\n+\tif (object->type == tree_type)\n+\t\tadd_pending_object(object, \"\");\n+\telse if (object->type == commit_type) {\n+\t\tstruct commit *commit = (struct commit *)object;\n+\t\tif (parse_commit(commit) < 0)\n+\t\t\tdie(\"unable to parse commit %s\", sha1_to_hex(sha1));\n+\t\tadd_pending_object(&(commit->tree->object), \"\");\n+\t}\n+}\n+\n /* for_each_ref() callback does not allow user data -- Yuck. */\n static struct commit_list **global_lst;\n \n@@ -865,6 +883,11 @@ int main(int argc, const char **argv)\n \t\t\tflags = UNINTERESTING;\n \t\t\targ++;\n \t\t\tlimited = 1;\n+\t\t} else if (*arg == '~') {\n+\t\t\tif (get_sha1(arg + 1, sha1) < 0)\n+\t\t\t\tdie(\"cannot get '%s'\", arg);\n+\t\t\thandle_tree(sha1);\n+\t\t\tcontinue;\n \t\t}\n \t\tif (get_sha1(arg, sha1) < 0) {\n \t\t\tstruct stat st;\n-- \n1.1.4.g9bd9d-dirty\n-- snap --\n\n> > Yes, I agree. But again, the local repo has to know which grafts were \n> > introduced by making the repo shallow.\n> \n> I am not sure I understand.  grafts are grafts are grafts.\n\nExactly. And grafts are grafts are not necessarily cutoffs.\n\nNow, is it possible that a fetch does something unintended, when there are \ngrafts which are not cutoffs? I don't know yet, but I think so.\n\nCiao,\nDscho\n"},{"id":"15457","messageId":"7vr76m36ge.fsf@assigned-by-dhcp.cox.net","threadId":"3187","inReplyTo":"Pine.LNX.4.63.0602020113200.30910@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: [PATCH] Shallow clone: low level machinery.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-02-02T01:17:05Z","receivedAt":"2006-02-02T01:17:05Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n>> If you have such a modified rev-list, yes.  I was having doubts\n>> about keeping an obvious correctness guarantee when doing such\n>> \"rev-list ~A\".\n>\n> I think it would be trivial: just resolve ~A to the tree A points to:\n\n<tired> Hmph.  I thought you meant \"have-only A\" to mean similar\nto \"have A\" but additionally \"do not assume I have things behind\nA\", and are going to extend rev-list to support ~A syntax to do\nthat.  I am a bit surprised to see your \"rev-list ~A\" is to\ninclude A, not exclude A and not what are behind A.  Where is\nthe connection between this and \"have-only A\"?  </tired> ;-)\n\n>> > Yes, I agree. But again, the local repo has to know which grafts were \n>> > introduced by making the repo shallow.\n>> \n>> I am not sure I understand.  grafts are grafts are grafts.\n>\n> Exactly. And grafts are grafts are not necessarily cutoffs.\n>\n> Now, is it possible that a fetch does something unintended, when there are \n> grafts which are not cutoffs? I don't know yet, but I think so.\n\nI think we are disagreeing, so \"not Exactly\".  I meant \"grafts\nare grafts, there is no cutoffs, they are also just grafts\".  So\nthe answer to your question is \"it does not matter\".\n"},{"id":"15497","messageId":"Pine.LNX.4.63.0602021932450.16426@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3187","inReplyTo":"7vr76m36ge.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] Shallow clone: low level machinery.","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-02-02T18:44:00Z","receivedAt":"2006-02-02T18:44:00Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 1 Feb 2006, Junio C Hamano wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> > I think it would be trivial: just resolve ~A to the tree A points to:\n> \n> <tired> Hmph.  I thought you meant \"have-only A\" to mean similar\n> to \"have A\" but additionally \"do not assume I have things behind\n> A\", and are going to extend rev-list to support ~A syntax to do\n> that.  I am a bit surprised to see your \"rev-list ~A\" is to\n> include A, not exclude A and not what are behind A.  Where is\n> the connection between this and \"have-only A\"?  </tired> ;-)\n\n<tired> My patch was wrong. You'd have to introduce a new flag saying: \nTraverse this commit, but mark its parents as uninteresting. </tired>\n\n> > Now, is it possible that a fetch does something unintended, when there are \n> > grafts which are not cutoffs? I don't know yet, but I think so.\n> \n> I think we are disagreeing, so \"not Exactly\".  I meant \"grafts\n> are grafts, there is no cutoffs, they are also just grafts\".  So\n> the answer to your question is \"it does not matter\".\n\nScenario: I have cvsimported a project. Using a graft, I told git that a \ncertain commit is indeed a merge between two branches. That is, in \naddition to the parent the commit objects tells us about, it has another \nparent which was tip of another branch.\n\nHow would this graft be interpreted by the server we want to pull from? As \nif we had cut off the history. Which we did not. In effect, we could be \nsent many, many objects we already have.\n\nCiao,\nDscho\n"},{"id":"15499","messageId":"7v3bj1r208.fsf@assigned-by-dhcp.cox.net","threadId":"3187","inReplyTo":"Pine.LNX.4.63.0602021932450.16426@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: [PATCH] Shallow clone: low level machinery.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-02-02T19:31:35Z","receivedAt":"2006-02-02T19:31:35Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> Scenario: I have cvsimported a project. Using a graft, I told git that a \n> certain commit is indeed a merge between two branches. That is, in \n> addition to the parent the commit objects tells us about, it has another \n> parent which was tip of another branch.\n>\n> How would this graft be interpreted by the server we want to pull from? As \n> if we had cut off the history. Which we did not. In effect, we could be \n> sent many, many objects we already have.\n\nI thought the protocol is sending the full graft file both ways.\nThe uploader says \"here are the grafts I have and use\", and the\ndownloader modifies it and sends back what grafts it wants to\nbe used during the common revision discovery (aka building\nrev-list parameters).  The most important modification during\nthis exchange is to cauterize the history at --since=v2.6.14\ncommit (or tag).\n\nThe uploader may not have the fake parent you grafted onto a\ncommit.  You may have a graft entry that says commit W has X, Y\nand Z as its parents, when its real parent is only X.  Y may be\nsome other commit in the project (i.e. the other end knows about\nit but it is not a real parent of W), and Z may be from a\ndevelopment track that the uploader has not even heard of.  You\nmay say a commit V does not have parent but that commit itself\nis from a separate development track the uploader does not know\nabout.\n\nThe uploader, however, should be able to at least honour, modulo\nimplementation bugs ;-), \"X and Y are both parents of W\" part.\nJust ignoring V and Z and keeping usable part of information\nwould be a reasonable fallback position [*1*].  And that should\nnot result in a \"many objects\" situation when the downloader\nsays \"Now I happen to have W, do not send things reachable from\nit\".  The uploader side should be able to omit what are\nreachable from X or Y even though it cannot exclude things\nreachable from Z.  Because the uploader does not even have Z,\nthere is no reason to worry about things reachable from Z being\nsent unnecessarily to the downloader.\n\nAt least that was the intention.  \"graft\" messages are not about\nsending \"here are the cut-off points\"; it is to agree on the\ngraft information both ends use during the common revision\ncomputation.  The experimental code does not treat cut-offs any\ndifferently other grafts.\n\n\n[Footnote]\n\n*1* we might want to enhance the \"shallow\" protocol further to\ndo this exchange slightly differently.  The downloader first\nsends its grafts (which may contain parents or graft/cutoff\npoints that uploader does not have), and the uploader adjusts\nthe received grafts for commits like V and parents like Z and\nthen add its own grafts.  The result is sent back to the\ndownloader and that becomes the common set of grafts in effect\nduring the common revision discovery.  This would contain\ncommits and parents that the downloader does not yet have but\nthat is not a problem for common revision discovery.  After the\ntransfer is done, the downloader would adjust its \"graft\" file\nif it made a new shallow clone, but otherwise it should not use\nthe information it received from the uploader, because things\nlike V and Z are not in this list.  I _think_ it would suffice\nto look at each graft entry and to add that entry locally if it\ntalks about a commit the downloader does not have in its graft\nfile.\n"}]}