{"thread":{"id":"6752","subject":"restriction of pulls","startedAt":"2007-02-09T10:49:12Z","lastAt":"2007-02-12T14:29:39Z","messageCount":13,"participants":["Christoph Duelli","Jakub Narebski","Johannes Schindelin","Rogan Dawes","Andy Parkins"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"34055","messageId":"200702091149.12462.duelli@melosgmbh.de","threadId":"6752","inReplyTo":null,"subject":"restriction of pulls","fromName":"Christoph Duelli","fromEmail":"duelli@melosgmbh.de","sentAt":"2007-02-09T10:49:12Z","receivedAt":"2007-02-09T10:49:12Z","isPatch":false,"sender":{"key":"duelli@melosgmbh.de","avatar":null},"body":"Is it possible to restrict a chechout, clone or a later pull to some \nsubdirectory of a repository?\n(Background: using subversion (or cvs), it is possible to do a file or \ndirectory-restricted update.)\n\nSay, I have a repository containing 2 (mostly) independent projects A and B \n(in separate) directories:\n- R\n  -  A\n  -  B\nIs it possibly to pull all the changes made to B, but not those made to A. \n(Yes, I know that this causes trouble if there are dependencies into A.)\n\n\nRegards\n-- \nChristoph Duelli\nMELOS GmbH\n"},{"id":"34058","messageId":"eqhl87$ut3$1@sea.gmane.org","threadId":"6752","inReplyTo":"200702091149.12462.duelli@melosgmbh.de","subject":"Re: restriction of pulls","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-02-09T11:19:05Z","receivedAt":"2007-02-09T11:19:05Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Christoph Duelli wrote:\n\n> Is it possible to restrict a chechout, clone or a later pull to some \n> subdirectory of a repository?\n> (Background: using subversion (or cvs), it is possible to do a file or \n> directory-restricted update.)\n> \n> Say, I have a repository containing 2 (mostly) independent projects\n> A and B (in separate) directories:\n> - R\n>   -  A\n>   -  B\n> Is it possibly to pull all the changes made to B, but not those made to A. \n> (Yes, I know that this causes trouble if there are dependencies into A.)\n\nNo, it is not possible. Moreover, it is not sensible, as it breaks atomicity\nof a commit. Well, you can hack, but...\n\nThat said, there is experimental submodule (subproject) support\n  http://git.or.cz/gitwiki/SubprojectSupport\n  http://git.admingilde.org/tali/git.git/module2\n(there was also proposal of more lightweight submodule support, but I don't\nhave a link to it), and you should have set A and B as submodules\n(subprojects).\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"34061","messageId":"Pine.LNX.4.63.0702091554160.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6752","inReplyTo":"200702091149.12462.duelli@melosgmbh.de","subject":"Re: restriction of pulls","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-09T14:54:40Z","receivedAt":"2007-02-09T14:54:40Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 9 Feb 2007, Christoph Duelli wrote:\n\n> Is it possible to restrict a chechout, clone or a later pull to some \n> subdirectory of a repository?\n\nNo. In git, a revision really is a revision, and not a group of file \nrevisions.\n\nCiao,\nDscho\n"},{"id":"34063","messageId":"45CC941E.9030808@dawes.za.net","threadId":"6752","inReplyTo":"Pine.LNX.4.63.0702091554160.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: restriction of pulls","fromName":"Rogan Dawes","fromEmail":"lists@dawes.za.net","sentAt":"2007-02-09T15:32:46Z","receivedAt":"2007-02-09T15:32:46Z","isPatch":false,"sender":{"key":"lists@dawes.za.net","avatar":null},"body":"Johannes Schindelin wrote:\n> Hi,\n> \n> On Fri, 9 Feb 2007, Christoph Duelli wrote:\n> \n>> Is it possible to restrict a chechout, clone or a later pull to some \n>> subdirectory of a repository?\n> \n> No. In git, a revision really is a revision, and not a group of file \n> revisions.\n> \n> Ciao,\n> Dscho\n> \n\nI thought about how this might be implemented, although I'm not entirely \nsure how efficient this will be.\n\nOne obstacle to implementing partial checkouts is that one does not know \nwhich objects have changed or been deleted. One way of addressing this \nis to keep a record of the hashes of all the objects that were NOT \nchecked out. (If one does not check out part of a directory, simply \nstore the hash of the top level, and you do not need to store the child \nhashes.) This record would be a kind of \"negative index\".\n\nWhen deciding what to check in, or which files are modified, one would \ncheck the \"negative index\" first to see if an entry exists. If not, only \nthen would you check the filesystem to see if modification times have \nchanged. With the \"negative index\", and the files in the file system, \none would be able to construct new commits, without any problem.\n\nIt would also require an updated transfer protocol, which would allow \nthe client to specify a tag/commit, then walk the tree that it points to \nto find the portion that the client is looking for, then pull only those \nobjects (and possibly their history). This is likely to be VERY \ninefficient in terms of round trips, at least initially.\n\nThis might be able to benefit from the shallow checkout support that was \nrecently implemented.\n\nComments?\n\nRogan\n"},{"id":"34065","messageId":"200702091619.23058.andyparkins@gmail.com","threadId":"6752","inReplyTo":"45CC941E.9030808@dawes.za.net","subject":"Re: restriction of pulls","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-02-09T16:19:20Z","receivedAt":"2007-02-09T16:19:20Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Friday 2007 February 09 15:32, Rogan Dawes wrote:\n\n> One obstacle to implementing partial checkouts is that one does not know\n> which objects have changed or been deleted. One way of addressing this\n\nWhy would you want to do a partial checkout.  I used subversion for a long \ntime before git, which does to partial checkouts and it's a nightmare.\n\nThings like this\n\n cd dir1/\n edit files\n cd ../dir2\n edit files\n svn commit\n * committed revision 100\n\nKABLAM!  Disaster.  Revision 100 no longer compiles/runs.  The changes in dir1 \nand dir2 were complimentary changes (say like renaming a function and then \nthe places that call that function).\n\nI didn't even notice how awful it was until I started using git and had a VCS \nthat did the right thing.\n\nIn every way that matters you can do a partial checkout - I can pull any \nversion of any file out of the repository.  However, it should certainly not \nbe the case that git records that fact.\n\nI think what you're actually after (from your description) is a shallow clone.  \nI believe that went in a while ago from Dscho.\n\n $ git clone --depth=5 <someurl>\n\nWill fetch only the last 5 revisions from the remote.  The other half to that \nis a shallow-by-tree clone; that is anathema to git as there is no such thing \nas a partial tree.  Submodule support is what you want, but that's still in \ndevelopment.\n\nThe only piece that (I think) is missing to get the functionality you want is \na kind of lazy transfer mode.  For something like, say, the kde repository \nyou can do\n\n svn checkout svn://svn.kde.org/kde/some/deep/path/in/the/project\n\nAnd just get that directory - i.e. you don't have to pay the cost of \ndownloading the whole of KDE.  Git can't do that; however, I think one day it \nwill be able to by choosing not to download every object from the remote.\n\n\nAndy\n-- \nDr Andy Parkins, M Eng (hons), MIEE\nandyparkins@gmail.com\n"},{"id":"34066","messageId":"45CCA301.4060504@dawes.za.net","threadId":"6752","inReplyTo":"200702091619.23058.andyparkins@gmail.com","subject":"Re: restriction of pulls","fromName":"Rogan Dawes","fromEmail":"lists@dawes.za.net","sentAt":"2007-02-09T16:36:17Z","receivedAt":"2007-02-09T16:36:17Z","isPatch":false,"sender":{"key":"lists@dawes.za.net","avatar":null},"body":"Andy Parkins wrote:\n> On Friday 2007 February 09 15:32, Rogan Dawes wrote:\n> \n>> One obstacle to implementing partial checkouts is that one does not know\n>> which objects have changed or been deleted. One way of addressing this\n> \n> Why would you want to do a partial checkout.  I used subversion for a long \n> time before git, which does to partial checkouts and it's a nightmare.\n> \n> Things like this\n> \n>  cd dir1/\n>  edit files\n>  cd ../dir2\n>  edit files\n>  svn commit\n>  * committed revision 100\n> \n> KABLAM!  Disaster.  Revision 100 no longer compiles/runs.  The changes in dir1 \n> and dir2 were complimentary changes (say like renaming a function and then \n> the places that call that function).\n\nPlease note that my suggestion does NOT imply allowing partial checkins \n(or if it does, it was not my intention)\n\nWhat I am trying to support is Jon Smirl's description of how some \nMozilla contributors work, specifically the documentation folks.\n\nThey do not have any need to look at the actual code, but simply limit \nthemselves to the files in the doc/ directory.\n\nSupporting a partial checkout of this doc/ directory would allow them to \nget a \"check in\"-able subdirectory, without having to download the rest \nof the source.\n\nWhat I intended to convey was that when determining which files have \nchanged, and presenting them to the user to decide whether to commit \nthem or not, the filesystem-walker would first check the \"negative \nindex\" to see if that directory/file had been explicitly excluded from \nthe checkout. This implies that they did not (and do not intend to) \nmodify that portion of the tree. Which implies that the committer can \nthen construct a complete view of the entire tree (now including the \nchanges that were made in the partial checkout) by resolving the \nmodified files with the knowledge of the hashes of the unmodified \nfiles/trees.\n\n> \n> In every way that matters you can do a partial checkout - I can pull any \n> version of any file out of the repository.  However, it should certainly not \n> be the case that git records that fact.\n\nWhy not? If you only want to modify that file, does it not make sense \nthat you can just check out that file, modify it, and check it back in?\n\nOr at least if not check it in, construct a diff for mailing to the \nmaintainer?\n\nOr even, allowing the maintainer to pull/merge the changes from the \ncontributor, even though the contributor doesn't necessarily have all \nthe blobs required to make up the tree he is committing? They should all \nbe available from the \"alternate\" if required.\n\nRogan\n"},{"id":"34067","messageId":"200702091645.33384.andyparkins@gmail.com","threadId":"6752","inReplyTo":"45CCA301.4060504@dawes.za.net","subject":"Re: restriction of pulls","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-02-09T16:45:31Z","receivedAt":"2007-02-09T16:45:31Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Friday 2007 February 09 16:36, Rogan Dawes wrote:\n\n> Please note that my suggestion does NOT imply allowing partial checkins\n> (or if it does, it was not my intention)\n\nMy apologies then; I did misunderstand.\n\n\n> > In every way that matters you can do a partial checkout - I can pull any\n> > version of any file out of the repository.  However, it should certainly\n> > not be the case that git records that fact.\n>\n> Why not? If you only want to modify that file, does it not make sense\n> that you can just check out that file, modify it, and check it back in?\n\nSorry - what I meant was that it shouldn't record that you checked out \nrevision 74 of that file and retain a link from the current version to that \nold version.\n\n\nAndy\n-- \nDr Andy Parkins, M Eng (hons), MIEE\nandyparkins@gmail.com\n"},{"id":"34069","messageId":"45CCB041.1000500@dawes.za.net","threadId":"6752","inReplyTo":"200702091645.33384.andyparkins@gmail.com","subject":"Re: restriction of pulls","fromName":"Rogan Dawes","fromEmail":"lists@dawes.za.net","sentAt":"2007-02-09T17:32:49Z","receivedAt":"2007-02-09T17:32:49Z","isPatch":false,"sender":{"key":"lists@dawes.za.net","avatar":null},"body":"Andy Parkins wrote:\n> On Friday 2007 February 09 16:36, Rogan Dawes wrote:\n> \n>> Please note that my suggestion does NOT imply allowing partial checkins\n>> (or if it does, it was not my intention)\n> \n> My apologies then; I did misunderstand.\n> \nThat'll teach me to be more clear ;-)\n\n> \n>>> In every way that matters you can do a partial checkout - I can pull any\n>>> version of any file out of the repository.  However, it should certainly\n>>> not be the case that git records that fact.\n>> Why not? If you only want to modify that file, does it not make sense\n>> that you can just check out that file, modify it, and check it back in?\n> \n> Sorry - what I meant was that it shouldn't record that you checked out \n> revision 74 of that file and retain a link from the current version to that \n> old version.\n\nWell, the new commit would have the previous commit as its direct \nparent, even though it may not have all the blobs to support it.\n\nWhich implies that all the git merge semantics should still work, \nassuming that the person actually doing the merge has all the necessary \nobjects to resolve any conflicts. (Which does not necessarily imply that \nhe has ALL of the objects in the tree, just those that are implicated in \nany conflicts).\n\nSo, for example, the doc team may have a documentation maintainer who \nhas the entire doc/ directory, who resolves any submissions from the doc \nteam, and feeds that up into the master tree. And all of this could be \ndone by means of pulls by the upstream maintainers.\n\nRogan\n"},{"id":"34095","messageId":"200702100959.40401.andyparkins@gmail.com","threadId":"6752","inReplyTo":"45CCB041.1000500@dawes.za.net","subject":"Re: restriction of pulls","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-02-10T09:59:38Z","receivedAt":"2007-02-10T09:59:38Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Friday 2007, February 09, Rogan Dawes wrote:\n\n> Well, the new commit would have the previous commit as its direct\n> parent, even though it may not have all the blobs to support it.\n\nThis I agree with; this seems like the way that a partial checkout would \nbe supported.\n\nAs you say - there would be no need to have the blobs available for \nobjects you aren't altering.  Unfortunately, it seems like it would be \na huge amount of work to actually do.\n\n\nAndy\n-- \nDr Andrew Parkins, M Eng (Hons), AMIEE\nandyparkins@gmail.com\n"},{"id":"34107","messageId":"Pine.LNX.4.63.0702101533060.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6752","inReplyTo":"45CC941E.9030808@dawes.za.net","subject":"Re: restriction of pulls","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-10T14:50:11Z","receivedAt":"2007-02-10T14:50:11Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 9 Feb 2007, Rogan Dawes wrote:\n\n> Johannes Schindelin wrote:\n> > \n> > On Fri, 9 Feb 2007, Christoph Duelli wrote:\n> > \n> > > Is it possible to restrict a chechout, clone or a later pull to some \n> > > subdirectory of a repository?\n> > \n> > No. In git, a revision really is a revision, and not a group of file \n> > revisions.\n> \n> I thought about how this might be implemented, although I'm not entirely \n> sure how efficient this will be.\n\nThere are basically three ways I can think of:\n\n- rewrite the commit objects on the fly. You might want to avoid the use \nof the pack protocol here (i.e. use HTTP or FTP transport).\n\n- try to teach git a way to ignore certain missing objects and \ndirectories. This might be involved, but you could extend upload-pack \neasily with a new extension for that.\n\n(my favourite:)\n- use git-split to create a new branch, which only contains doc/. Do work \nonly on that branch, and merge into mainline from time to time.\n\nIf you don't need the history, you don't need to git-split the branch.\n\nYou only need to make sure that the newly created branch is _not_ branched \noff of mainline, since the next merge would _delete_ all files outside of \ndoc/ (merge would see that the files exist in mainline, and existed in the \ncommon ancestor, too, so would think that the files were deleted in the \ndoc branch).\n\nCiao,\nDscho\n"},{"id":"34282","messageId":"45D07296.7070804@dawes.za.net","threadId":"6752","inReplyTo":"Pine.LNX.4.63.0702101533060.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: restriction of pulls","fromName":"Rogan Dawes","fromEmail":"discard@dawes.za.net","sentAt":"2007-02-12T13:58:46Z","receivedAt":"2007-02-12T13:58:46Z","isPatch":false,"sender":{"key":"discard@dawes.za.net","avatar":null},"body":"Johannes Schindelin wrote:\n> Hi,\n> \n> On Fri, 9 Feb 2007, Rogan Dawes wrote:\n> \n>> Johannes Schindelin wrote:\n>>> On Fri, 9 Feb 2007, Christoph Duelli wrote:\n>>>\n>>>> Is it possible to restrict a chechout, clone or a later pull to some \n>>>> subdirectory of a repository?\n>>> No. In git, a revision really is a revision, and not a group of file \n>>> revisions.\n>> I thought about how this might be implemented, although I'm not entirely \n>> sure how efficient this will be.\n> \n> There are basically three ways I can think of:\n> \n> - rewrite the commit objects on the fly. You might want to avoid the use \n> of the pack protocol here (i.e. use HTTP or FTP transport).\n> \n> - try to teach git a way to ignore certain missing objects and \n> directories. This might be involved, but you could extend upload-pack \n> easily with a new extension for that.\n> \n> (my favourite:)\n> - use git-split to create a new branch, which only contains doc/. Do work \n> only on that branch, and merge into mainline from time to time.\n> \n> If you don't need the history, you don't need to git-split the branch.\n> \n> You only need to make sure that the newly created branch is _not_ branched \n> off of mainline, since the next merge would _delete_ all files outside of \n> doc/ (merge would see that the files exist in mainline, and existed in the \n> common ancestor, too, so would think that the files were deleted in the \n> doc branch).\n> \n> Ciao,\n> Dscho\n> \n\nYour third option sounds quite clever, apart from the problem of \nattributing a commit and a commit message to someone, when the actual \ncommit doesn't match what they actually did :-(\n\nAs well as wondering what happens when they check out a few more files. \nDo we rewrite those commits as well? What happens if the user has made \nsome commits already? What happens if they have already sent those \nupstream? etc.\n\nI think the best solution is ultimately to make git able to cope with \ncertain missing objects.\n\nI started writing this in response to another message, but it will do \nfine here, too:\n\nThe description I give here will likely horrify people in terms of \ncommunications inefficiency, but I'm sure that can be improved.\n\nScenario:\n\nA user sees a documentation bug in a git-managed project, and decides \nthat she wants to do something about it. Since she is not on the fastest \nof connections, she'd like to reduce the checkout to a reasonable \nminimum, while still working with the git tools.\n\nViewing the repo layout using gitweb, she sees that all the \ndocumentation is stored in the docs/ directory from the root.\n\nSo, she creates a local repo to work in:\n\n$ git init-db\n\nShe configures her local repo to reference the source one:\n\n(Hypothetical syntax)\n$ git clone --reference http://example.com/project.git \\\n     http://example.com/project.git\n\nSince the reference and repo are the same (and non-local), git doesn't \nactually download anything, other than the current heads (and maybe tags).\n\nShe then does a partial checkout of the master branch, but only the \ndocs/ directory:\n\n$ git checkout -p master docs/\n\nThe -p flag indicates that this is a partial checkout of master. Git \nrecords that the current HEAD is \"master\", checks out the docs/ \ndirectory, and removes any other files in the working directory (that it \nknew about from the existing index, if any - I'm not suggesting that it \nshould arbitrarily delete files!)\n\nThe checkout process goes as follows: Resolve the <treeish> that HEAD \npoints to, and retrieve it from the upstream repo if it does not exist \nlocally. Continue requesting only the necessary tree and blob objects to \nsatisfy the requested checkout. i.e. From the first tree, identify the \ndocs/ directory. Then request only that tree object. Continue to \ndownload tree and blob objects until the entire docs/ directory can be \ncreated in the working directory.\n\nThis will likely require a new index file format, that simply stores the \nhashes of objects (blobs or trees) that have not been checked out, as \nwell as the current file's stat information.\n\nNow create a \"negative index\" (pindex?) that has details about the other \nfiles and directories that were not checked out. Obviously, this does \nnot need to recurse into directories that were not checked out. Simply \nhaving the hash of the parent directory in the pindex is sufficient \ninformation to reconstruct a new index. (This might require a new index \nformat that does not include all known files, but simply stores the hash \nof the unchecked-out tree or blob.)\n\nThen creating a new commit would require creating the necessary blobs \nfor changed files, new tree objects for trees that change, and a commit \nobject.\n\nAs far as I can tell, that could then be pushed/pulled/merged using the \nexisting tools, without any problems.\n\nRogan\n"},{"id":"34283","messageId":"Pine.LNX.4.63.0702121508360.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6752","inReplyTo":"45D07296.7070804@dawes.za.net","subject":"Re: restriction of pulls","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-12T14:13:44Z","receivedAt":"2007-02-12T14:13:44Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 12 Feb 2007, Rogan Dawes wrote:\n\n> Johannes Schindelin wrote:\n> > \n> > (my favourite:)\n> > - use git-split to create a new branch, which only contains doc/. Do work\n> > only on that branch, and merge into mainline from time to time.\n> \n> Your third option sounds quite clever, apart from the problem of attributing a\n> commit and a commit message to someone, when the actual commit doesn't match\n> what they actually did :-(\n\nThis problem is not related to subprojects at all. If the commit message \ndoes not match the patch, you are always fscked.\n\n> As well as wondering what happens when they check out a few more files. Do we\n> rewrite those commits as well? What happens if the user has made some commits\n> already? What happens if they have already sent those upstream? etc.\n\nI think you misunderstood. My favourite option would make docs a \n_separate_ project, with its own history. It just happens to be pulled \nfrom time to time, just like git-gui, gitk and git-fast-import in git.git.\n\n> I think the best solution is ultimately to make git able to cope with \n> certain missing objects.\n\nHmm. I am not convinced. On nice thing about git is its level of \nintegrity. Which means that no random objects are missing.\n\n> I started writing this in response to another message, but it will do fine\n> here, too:\n> \n> The description I give here will likely horrify people in terms of\n> communications inefficiency, but I'm sure that can be improved.\n>\n> [goes on... and describes the lazy clone!]\n\nAFAICT this really is the lazy clone. And it was already determined that \nit is all to easy to pull in all commit objects by accident. Which boils \ndown to a substantial chunk of the repository.\n\nBut if you want to play with it: by all means, go ahead. It might just be \nthat you overcome the fundamental difficulties, and we get something nice \nout of it.\n\nCiao,\nDscho\n"},{"id":"34286","messageId":"45D079D3.2020500@dawes.za.net","threadId":"6752","inReplyTo":"Pine.LNX.4.63.0702121508360.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: restriction of pulls","fromName":"Rogan Dawes","fromEmail":"lists@dawes.za.net","sentAt":"2007-02-12T14:29:39Z","receivedAt":"2007-02-12T14:29:39Z","isPatch":false,"sender":{"key":"lists@dawes.za.net","avatar":null},"body":"Johannes Schindelin wrote:\n> Hi,\n> \n> On Mon, 12 Feb 2007, Rogan Dawes wrote:\n> \n>> Johannes Schindelin wrote:\n>>> (my favourite:)\n>>> - use git-split to create a new branch, which only contains doc/. Do work\n>>> only on that branch, and merge into mainline from time to time.\n>> Your third option sounds quite clever, apart from the problem of attributing a\n>> commit and a commit message to someone, when the actual commit doesn't match\n>> what they actually did :-(\n> \n> This problem is not related to subprojects at all. If the commit message \n> does not match the patch, you are always fscked.\n\nWell, I was thinking about the fact that the files originally checked in \nwill not match the files \"checked in\" in the rewritten commit.\n\n>> As well as wondering what happens when they check out a few more files. Do we\n>> rewrite those commits as well? What happens if the user has made some commits\n>> already? What happens if they have already sent those upstream? etc.\n> \n> I think you misunderstood. My favourite option would make docs a \n> _separate_ project, with its own history. It just happens to be pulled \n> from time to time, just like git-gui, gitk and git-fast-import in git.git.\n\nI see. However, that does not allow for the random single-file checkout \nscenario I sketched out. Which may or may not be common/desirable, but \nit is an extreme case of the partial checkout, without fixed delineation.\n\n>> I think the best solution is ultimately to make git able to cope with \n>> certain missing objects.\n> \n> Hmm. I am not convinced. On nice thing about git is its level of \n> integrity. Which means that no random objects are missing.\n\nGood point. :-(\n\n>> I started writing this in response to another message, but it will do fine\n>> here, too:\n>>\n>> The description I give here will likely horrify people in terms of\n>> communications inefficiency, but I'm sure that can be improved.\n>>\n>> [goes on... and describes the lazy clone!]\n >\n> AFAICT this really is the lazy clone. And it was already determined that \n> it is all to easy to pull in all commit objects by accident. Which boils \n> down to a substantial chunk of the repository.\n>\n\nNot so much a lazy clone as a partial clone. It is only in the \"clone\", \n\"fetch\" or \"checkout\" code paths that new objects will be retrieved from \nthe source repo. Things like \"git log\"/\"git show\" would not do so, and \nwould be required to handle missing objects gracefully.\n\n> But if you want to play with it: by all means, go ahead. It might just be \n> that you overcome the fundamental difficulties, and we get something nice \n> out of it.\n> \n> Ciao,\n> Dscho\n> \n\nMaybe ;-) We'll see if I get any time for it.\n\nRogan\n"}]}