{"thread":{"id":"26222","subject":"clone breaks replace","startedAt":"2011-01-06T21:00:24Z","lastAt":"2011-01-15T05:27:41Z","messageCount":34,"participants":["Phillip Susi","Jonathan Nieder","Junio C Hamano","Stephen Bash","Jeff King","Christian Couder","Johannes Sixt"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"159068","messageId":"4D262D68.2050804@cfl.rr.com","threadId":"26222","inReplyTo":null,"subject":"clone breaks replace","fromName":"Phillip Susi","fromEmail":"psusi@cfl.rr.com","sentAt":"2011-01-06T21:00:24Z","receivedAt":"2011-01-06T21:00:24Z","isPatch":false,"sender":{"key":"psusi@cfl.rr.com","avatar":null},"body":"I've been experimenting with git replace to remove ancient history, and\nI have found that cloning a repository breaks replace.  I read about\nthis process at http://progit.org/2010/03/17/replace.html.  I managed to\ncorrectly add a replace commit that truncates the history and contains\ninstructions where you can find it, and running git log only goes back\nto the replacement commit, unless you add --no-replace-objects, which\ncauses it to show the original full history.\n\nThe problem is that when I clone the repository, I expect the clone to\ncontain only history up to the replacement record, and not the old\nhistory before that.  Instead, the clone contains only the full original\nhistory, and the replacement ref is not imported at all.  A git replace\nin the new clone shows nothing.\n\nShouldn't clone copy .git/refs/replace?\n"},{"id":"159071","messageId":"20110106213338.GA15325@burratino","threadId":"26222","inReplyTo":"4D262D68.2050804@cfl.rr.com","subject":"Re: clone breaks replace","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2011-01-06T21:33:38Z","receivedAt":"2011-01-06T21:33:38Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Phillip Susi wrote:\n\n> I managed to\n> correctly add a replace commit that truncates the history and contains\n> instructions where you can find it, and running git log only goes back\n> to the replacement commit, unless you add --no-replace-objects, which\n> causes it to show the original full history.\n\nBefore I get to your real question: this seems a bit backwards.  Let\nme say a few words about why.\n\nIn the days before replacement refs (and today, too), each commit\nname described not only the state of a tree at a moment but the\nhistory that led up to it.  In fact you can see this somewhat directly:\ngiven two distinct commits A and B if you try\n\n\t$ git cat-file commit A >a.commit\n\t$ git cat-file commit B >b.commit\n\t$ diff -u a.commit b.commit\n\nthen you will see precisely what can make them different:\n\n - the author's name and email and the date of authorship\n - the committer's name and email and the date committed\n - the names of the parent commits, describing the history\n - the name of a tree, describing the content\n - the log message, including its encoding\n\nThe commit name is a hash of that information (see git-hash-object(1))\nand an invariant maintained is \"if a repository has access to commit A,\nit has access to its parents, their parents, and so on\".  This invariant\nis maintained during object transfer and garbage collection and relied\non by object transfer and revision traversal.\n\nThe beauty of replacement refs is that they can be easily added or\nremoved without breaking this invariant.  And a replacement ref is an\nactual reference into history, so garbage collection does not remove\nthose commits and the repository keeps enough information to traverse\nboth the modified and unmodified history.\n\nTherefore if you want clients to be able to choose between a minimal\nhistory and a larger one to save bandwidth, it has to work like this\n\n - to get the minimal history, fetch _without_ any replacement refs\n - to get the full history, fetch the replacement refs on top of that.\n\nbecause an additional reference can only increase the number of\nobjects to be downloaded.\n\n> The problem is that when I clone the repository, I expect the clone to\n> contain only history up to the replacement record, and not the old\n> history before that.  Instead, the clone contains only the full original\n> history, and the replacement ref is not imported at all.  A git replace\n> in the new clone shows nothing.\n>\n> Shouldn't clone copy .git/refs/replace?\n\nWith that in mind, I suspect the best way to achieve what you are\nlooking for is the following:\n\n 1. Make a big, ugly history (branch \"big\").  Presumably this part's\n    already done.\n\n 2. Find the part you want to get rid of and make appropriate\n    replacement refs so \"gitk big\" shows what you want it to.\n\n 3. Use \"git filter-branch\" to make that history a reality (branch\n    \"simpler\").  Remove the replacement refs.\n\n 4. Use \"git replace\" to graft back on the pieces you cauterized.\n    Publish the result.\n\n 5. Perhaps also run and publish \"git replace big simpler\", so\n    contributors of branches based against the old 'big' can merge\n    your latest changes from 'simpler'.  Encourage contributors to\n    use 'git rebase' or 'git filter-branch' to rebase their\n    contributions against the new, simpler history.\n\nDoes that make sense?\n\nJonathan\n"},{"id":"159078","messageId":"7v7hehae2s.fsf@alter.siamese.dyndns.org","threadId":"26222","inReplyTo":"20110106213338.GA15325@burratino","subject":"Re: clone breaks replace","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-01-06T21:59:23Z","receivedAt":"2011-01-06T21:59:23Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n> Therefore if you want clients to be able to choose between a minimal\n> history and a larger one to save bandwidth, it has to work like this\n>\n>  - to get the minimal history, fetch _without_ any replacement refs\n>  - to get the full history, fetch the replacement refs on top of that.\n>\n> because an additional reference can only increase the number of\n> objects to be downloaded.\n\nVery nicely and clearly put.  Can we have this somewhere in the docs?\n"},{"id":"159142","messageId":"4D276CD2.60607@cfl.rr.com","threadId":"26222","inReplyTo":"20110106213338.GA15325@burratino","subject":"Re: clone breaks replace","fromName":"Phillip Susi","fromEmail":"psusi@cfl.rr.com","sentAt":"2011-01-07T19:43:14Z","receivedAt":"2011-01-07T19:43:14Z","isPatch":false,"sender":{"key":"psusi@cfl.rr.com","avatar":null},"body":"On 1/6/2011 4:33 PM, Jonathan Nieder wrote:\n> Therefore if you want clients to be able to choose between a minimal\n> history and a larger one to save bandwidth, it has to work like this\n> \n>  - to get the minimal history, fetch _without_ any replacement refs\n>  - to get the full history, fetch the replacement refs on top of that.\n> \n> because an additional reference can only increase the number of\n> objects to be downloaded.\n\nThis seems backwards.  The original commit links to its parent and\ntherefore, the full history trail going back.  The reason you add the\nreplacement record is to get rid of that parent link, thus truncating\nthe history.  Therefore, if you fetch the original record that still has\nthe reference to its parent, and not the replacement record, you end up\nwith the full history.  Ergo, to get only the truncated history, you\nmust fetch the replacement record, and pay attention to it to stop\nfetching commits older than the truncation point.\n\n>  3. Use \"git filter-branch\" to make that history a reality (branch\n>     \"simpler\").  Remove the replacement refs.\n\nIsn't the whole purpose of using replace to avoid having to use\nfilter-branch, which throws out all of the existing commit records, and\ncreates an entirely new commit chain that is slightly modified?\n\n>  4. Use \"git replace\" to graft back on the pieces you cauterized.\n>     Publish the result.\n\nIf you are going to use filter-branch, then what do you need to replace?\n And publishing the result of a replace seems to have no effect, since\nother people do not get the replace ref when they clone.\n\n>  5. Perhaps also run and publish \"git replace big simpler\", so\n>     contributors of branches based against the old 'big' can merge\n>     your latest changes from 'simpler'.  Encourage contributors to\n>     use 'git rebase' or 'git filter-branch' to rebase their\n>     contributions against the new, simpler history.\n\nAgain, the entire point of replace seems to be to AVOID having to go\nthrough the hassle of having to rebase or filter-branch.  Isn't that\nexactly how you would accomplish this before replace was added?\n"},{"id":"159157","messageId":"20110107205103.GC4629@burratino","threadId":"26222","inReplyTo":"4D276CD2.60607@cfl.rr.com","subject":"Re: clone breaks replace","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2011-01-07T20:51:03Z","receivedAt":"2011-01-07T20:51:03Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Phillip Susi wrote:\n\n> Isn't the whole purpose of using replace to avoid having to use\n> filter-branch, which throws out all of the existing commit records, and\n> creates an entirely new commit chain that is slightly modified?\n\nNo.  What documentation suggested that?  Maybe it can be fixed.\n\nThe original purpose of grafts (the ideological ancestor of\nreplacement refs) was to serve a very particular use case.  Sit down\nby the fire, if you will, and...\n\nGit had just came into existence and pack files did not exist yet.  A\nfull import of the Linux kernel history was possible but the result\nwas enormous and not something ready to be imposed on all Linux\ncontributors.  So what can one do?\n\n $ git show -s v2.6.12-rc2^0\n commit 1da177e4c3f41524e886b7f1b8a0c1fc7321cac2\n Author: Linus Torvalds <torvalds@ppc970.osdl.org>\n Date:   Sat Apr 16 15:20:36 2005 -0700\n\n     Linux-2.6.12-rc2\n\n     Initial git repository build. I'm not bothering with the full history,\n     even though we have it. We can create a separate \"historical\" git\n     archive of that later if we want to, and in the meantime it's about\n     3.2GB when imported into git - space that would just make the early\n     git days unnecessarily complicated, when we don't have a lot of good\n     infrastructure for it.\n\n     Let it rip!\n\nFast forward three months, and there is discussion[1] about what to do\nwith the historical git archive.  A clever idea: teach git to _pretend_\nthat the historical archive is the parent to v2.6.12-rc2, so\n\"git log --grep\", \"gitk\", and so on work as they ought to.\n\nSo grafts were born.  One of the nicest advantages of grafts is that\nthey make it easy to do complex history surgery: make some grafts ---\ncut here, paste there --- and then run \"git filter-branch\" to make it\npermanent.\n\nBut grafts have a serious problem.\n\nTransport machinery needs to ignore grafts --- otherwise, the two ends\nof a connection could have different ideas of the history preceding a\ncommit, resulting in confusion and breakage.  A fix to that was\nfinally grafted on a few years later (see also [2]).\n\n $ GIT_NOTES_REF=refs/remotes/charon/notes/full \\\n   git log --grep=graft --grep=repack --all-match --no-merges\n [...]\n     git repack: keep commits hidden by a graft\n [...]\n     Archived-At: <http://thread.gmane.org/gmane.comp.version-control.git/123874>\n\nThere is also the problem that grafts are too \"raw\": it is very easy\nto make a graft pointing to a nonexistent object, say.  And meanwhile\ngit has no native support for transfering grafts over the wire.\n\nIn that context there emerged the nicer (imho) refs/replace mechanism:\n\n - reachability checking and transport machinery can treat them like\n   all other references --- no need for low-level tools to pay\n   attention to the artificial history;\n - easy to script around with \"git replace\" and \"git for-each-ref\"\n - can choose to fetch or not fetch with the usual\n   \"git fetch repo refs/replace/*:refs/replace/*\" syntax\n\nCommon applications:\n\n - locally staging history changes that will later be made permanent\n   with \"git filter-branch\";\n - grafting on additional (historical) history;\n - replacing ancient broken commits with fixed ones, for use by \"git\n   bisect\".\n\nHope that helps,\nJonathan\n\n[1] http://thread.gmane.org/gmane.comp.version-control.git/6470/focus=6484\nfound with \"git log --grep=graft --reverse\"\n[2] http://thread.gmane.org/gmane.comp.version-control.git/37744/focus=37908\n"},{"id":"159159","messageId":"1351312.105443.1294434905621.JavaMail.root@mail.hq.genarts.com","threadId":"26222","inReplyTo":"20110107205103.GC4629@burratino","subject":"Re: clone breaks replace","fromName":"Stephen Bash","fromEmail":"bash@genarts.com","sentAt":"2011-01-07T21:15:05Z","receivedAt":"2011-01-07T21:15:05Z","isPatch":false,"sender":{"key":"bash@genarts.com","avatar":null},"body":"----- Original Message -----\n> From: \"Jonathan Nieder\" <jrnieder@gmail.com>\n> To: \"Phillip Susi\" <psusi@cfl.rr.com>\n> Cc: git@vger.kernel.org, \"Christian Couder\" <chriscool@tuxfamily.org>\n> Sent: Friday, January 7, 2011 3:51:03 PM\n> Subject: Re: clone breaks replace\n> Phillip Susi wrote:\n> \n> > Isn't the whole purpose of using replace to avoid having to use\n> > filter-branch, which throws out all of the existing commit records,\n> > and creates an entirely new commit chain that is slightly modified?\n> \n> No. What documentation suggested that? Maybe it can be fixed.\n\nI'll chime in here as another person who read the ProGit blog entry on git-replace [1] and came to the same conclusion Phillip (and I'm guessing others) did.  OTOH when I attempted to read the actual git-replace manpage, I got completely lost, so I retained my (apparently incorrect) understanding from ProGit.\n\nThanks,\nStephen\n\n[1] http://progit.org/2010/03/17/replace.html\n"},{"id":"159160","messageId":"20110107213426.GB8693@burratino","threadId":"26222","inReplyTo":"20110107205103.GC4629@burratino","subject":"Re: clone breaks replace","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2011-01-07T21:34:26Z","receivedAt":"2011-01-07T21:34:26Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Jonathan Nieder wrote:\n\n> Transport machinery needs to ignore grafts --- otherwise, the two ends\n> of a connection could have different ideas of the history preceding a\n> commit, resulting in confusion and breakage.  A fix to that was\n> finally grafted on a few years later (see also [2]).\n\nSorry, I walked away mid-paragraph and left out a crucial piece when I\nreturned.  Because transport machinery ignores grafts, garbage\ncollection must make sure not to remove pieces of the non-artificial\nhistory.  It is the garbage collection that Dscho fixed with\nv1.6.4-rc3~7^2.\n\nSorry for the nonsense.\n"},{"id":"159161","messageId":"4D278930.7010100@cfl.rr.com","threadId":"26222","inReplyTo":"20110107205103.GC4629@burratino","subject":"Re: clone breaks replace","fromName":"Phillip Susi","fromEmail":"psusi@cfl.rr.com","sentAt":"2011-01-07T21:44:16Z","receivedAt":"2011-01-07T21:44:16Z","isPatch":false,"sender":{"key":"psusi@cfl.rr.com","avatar":null},"body":"On 1/7/2011 3:51 PM, Jonathan Nieder wrote:\n> Phillip Susi wrote:\n> \n>> Isn't the whole purpose of using replace to avoid having to use\n>> filter-branch, which throws out all of the existing commit records, and\n>> creates an entirely new commit chain that is slightly modified?\n> \n> No.  What documentation suggested that?  Maybe it can be fixed.\n\nIt's just what made sense to me.  If you can modify the history with\nfilter-branch, then you don't need replace refs.  The downside to\nfilter-branch is that it breaks people tracking your repository, since\nthe history they had been tracking is thrown out and replaced with a\ncompletely new commit chain that looks similar, but as far as git is\nconcerned, is unrelated to the original.  Replace refs seem to have been\ncreated to allow you to accomplish the goal of modifying an old commit\nrecord, but without having to rewrite that and all subsequent commits,\ncausing breakage.\n\n>  - can choose to fetch or not fetch with the usual\n>    \"git fetch repo refs/replace/*:refs/replace/*\" syntax\n\nIt seems like this should be the default behavior.  Or perhaps\nrefs/replace should be forked into one meant to be private, and one\nmeant to be public, and fetched by default.  Or maybe it should be\nfetched by default, but not pushed, so you have to explicitly push\nreplacements to the public mirror that you intend for public\nconsumption.  Having the replace only apply locally and still needing to\nfilter-branch to make the change visible to the public seems to render\nthe replace somewhat pointless.\n\nTake the kernel history as an example, only imagine that Linus did not\noriginally make that first commit leaving out the prior history, but\nwants to go back and fix it now.  He can do it with a replace, but then\nif he runs filter-branch as you suggest to make the change 'real', then\neveryone tracking his tree will fail the next time they try to pull.\nYou could get the same result without replace, so why bother?\n\nIf the replace was fetched by default, the people already tracking would\nget it the next time they pull and would not have a problem.  If they\nwanted to see the old history, then they would already have it in the\nrepository and just need to add --no-replace-objects to see it, or run\ngit log on the original commit id that the replace record should refer\nyou to ( in the comments ).  Those cloning the repository for the first\ntime would get it, and avoid fetching all of the old history since they\nwould be using the replace record in place of the original commit.\n"},{"id":"159163","messageId":"20110107214907.GA9194@burratino","threadId":"26222","inReplyTo":"4D278930.7010100@cfl.rr.com","subject":"Re: clone breaks replace","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2011-01-07T21:49:07Z","receivedAt":"2011-01-07T21:49:07Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Phillip Susi wrote:\n\n> Take the kernel history as an example, only imagine that Linus did not\n> originally make that first commit leaving out the prior history, but\n> wants to go back and fix it now.  He can do it with a replace, but then\n> if he runs filter-branch as you suggest to make the change 'real', then\n> everyone tracking his tree will fail the next time they try to pull.\n> You could get the same result without replace, so why bother?\n>\n> If the replace was fetched by default, the people already tracking would\n> get it the next time they pull and would not have a problem.\n\nInteresting.  I hadn't thought about this detail before.\n\n> Those cloning the repository for the first\n> time would get it, and avoid fetching all of the old history since they\n> would be using the replace record in place of the original commit.\n\nNo, it doesn't work that way.  Imagine for a moment that each commit\nobject actually contains all of its ancestors.  That isn't precisely\nright but in a way it is close.\n\nTo change the ancestry of a commit, you really do need to change its\nname.  If you disagree, feel free to try it and I'd be glad to help\nwhere I can with the coding if the design is sane.  Deal?\n\nMaybe it would be nice if git replace worked that way, but that would\nbe fundamentally a _different_ feature.\n"},{"id":"159166","messageId":"4D278F1D.4020706@cfl.rr.com","threadId":"26222","inReplyTo":"20110107214907.GA9194@burratino","subject":"Re: clone breaks replace","fromName":"Phillip Susi","fromEmail":"psusi@cfl.rr.com","sentAt":"2011-01-07T22:09:33Z","receivedAt":"2011-01-07T22:09:33Z","isPatch":false,"sender":{"key":"psusi@cfl.rr.com","avatar":null},"body":"On 1/7/2011 4:49 PM, Jonathan Nieder wrote:\n> No, it doesn't work that way.  Imagine for a moment that each commit\n> object actually contains all of its ancestors.  That isn't precisely\n> right but in a way it is close.\n> \n> To change the ancestry of a commit, you really do need to change its\n> name.  If you disagree, feel free to try it and I'd be glad to help\n> where I can with the coding if the design is sane.  Deal?\n\nThat's why a replace record seems to be the perfect solution.  The\noriginal record still references the old history, but you ignore it in\nfavor of the replacement, which does not.  Thus you have a choice; you\nignore the replacement and use the original with the full history\nattached, or you respect the replacement and the history is truncated.\n\nAs long as git-upload-pack respects the replacement, then new checkouts\nwill ignore the old history.  You could then create a new historical\nbranch that points to the parent commit of the replaced one, and tell\npeople to fetch that branch to get the old history, or pass\n--no-replace-objects over the wire to git-upload-pack.\n"},{"id":"159168","messageId":"20110107220942.GB10343@sigill.intra.peff.net","threadId":"26222","inReplyTo":"20110107214907.GA9194@burratino","subject":"Re: clone breaks replace","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-01-07T22:09:42Z","receivedAt":"2011-01-07T22:09:42Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jan 07, 2011 at 03:49:07PM -0600, Jonathan Nieder wrote:\n\n> Phillip Susi wrote:\n> \n> > Take the kernel history as an example, only imagine that Linus did not\n> > originally make that first commit leaving out the prior history, but\n> > wants to go back and fix it now.  He can do it with a replace, but then\n> > if he runs filter-branch as you suggest to make the change 'real', then\n> > everyone tracking his tree will fail the next time they try to pull.\n> > You could get the same result without replace, so why bother?\n> >\n> > If the replace was fetched by default, the people already tracking would\n> > get it the next time they pull and would not have a problem.\n> \n> Interesting.  I hadn't thought about this detail before.\n\nI think there are two separate issues here:\n\n  1. Should transport protocols respect replacements (i.e., if you\n     truncate history with a replacement object and I fetch from you,\n     should you get the full history or the truncated one)?\n\n  2. Should clone fetch refs from refs/replace (either by default, or\n     with an option)?\n\nBased on previous discussions, I think the answer to the first is no.\nThe resulting repo violates a fundamental assumption of git. Yes,\nbecause of the replacement object, many things will still work. But many\nparts of git intentionally do not respect replacement, and they will be\nbroken.\n\nInstead, I think of replacements as a specific view into history, not a\nfundamental history-changing operation itself. Which means you can never\nsave bandwidth or space by truncating history with replacements. You can\nonly give somebody the full history, and share with them your view. If\nyou want to truncate, you must rewrite history[1].\n\nWhich leads to the second question. It is basically a matter of saying\n\"do you want to fetch the view that upstream has\"? I can definitely see\nthat being useful, and meriting an option. However, it may or may not be\nworth turning on by default, as upstream's view may be confusing.\n\n-Peff\n\n[1] Actually, what we are talking about it basically shallow clone.\n    Which does do exactly this truncation, but does not use the replace\n    mechanism. So it _is_ possible, but lots of things need to be\n    tweaked to understand the shallow-ness. Perhaps in the long run\n    making git understand replacement-truncated repos with missing\n    objects would be a good thing, and shallow clones can be implemented\n    simply as a special case of that. It would probably make the code a\n    bit cleaner.\n"},{"id":"159176","messageId":"7vmxnc48yt.fsf@alter.siamese.dyndns.org","threadId":"26222","inReplyTo":"20110107220942.GB10343@sigill.intra.peff.net","subject":"Re: clone breaks replace","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-01-07T22:58:34Z","receivedAt":"2011-01-07T22:58:34Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n>   2. Should clone fetch refs from refs/replace (either by default, or\n>      with an option)?\n> ...\n> Which leads to the second question. It is basically a matter of saying\n> \"do you want to fetch the view that upstream has\"? I can definitely see\n> that being useful, and meriting an option. However, it may or may not be\n> worth turning on by default, as upstream's view may be confusing.\n\nI think that should be stated a bit differently.  \"Do you want to fetch\nthe view that the upstream offers as an option, and if you want, which\nones (meaning: there could be more than one replacement grafts to give\ndifferent views)?\"\n\nAnd as an optional view, I would say it is perfectly Ok to fetch whichever\nview you want as a separate step after the initial clone.\n"},{"id":"159190","messageId":"4D27B33C.2020907@cfl.rr.com","threadId":"26222","inReplyTo":"20110107220942.GB10343@sigill.intra.peff.net","subject":"Re: clone breaks replace","fromName":"Phillip Susi","fromEmail":"psusi@cfl.rr.com","sentAt":"2011-01-08T00:43:40Z","receivedAt":"2011-01-08T00:43:40Z","isPatch":false,"sender":{"key":"psusi@cfl.rr.com","avatar":null},"body":"On 01/07/2011 05:09 PM, Jeff King wrote:\n> I think there are two separate issues here:\n>\n>    1. Should transport protocols respect replacements (i.e., if you\n>       truncate history with a replacement object and I fetch from you,\n>       should you get the full history or the truncated one)?\n>\n>    2. Should clone fetch refs from refs/replace (either by default, or\n>       with an option)?\n>\n> Based on previous discussions, I think the answer to the first is no.\n> The resulting repo violates a fundamental assumption of git. Yes,\n> because of the replacement object, many things will still work. But many\n> parts of git intentionally do not respect replacement, and they will be\n> broken.\n\nWhat parts do not respect replacement?  More importantly, what parts \nwill be broken?  The man page seems to indicate that about the only \nthing that does not by default is reachability testing, which to me \nmeans fsck and prune.  It seems to be the purpose of replace to \n/prevent/ breakage and be respected by default, unless doing so would \ncause harm, which is why fsck and prune do not.\n\n> Instead, I think of replacements as a specific view into history, not a\n> fundamental history-changing operation itself. Which means you can never\n> save bandwidth or space by truncating history with replacements. You can\n> only give somebody the full history, and share with them your view. If\n> you want to truncate, you must rewrite history[1].\n\nRight, but if you only care about that view, then there is no need to \nwaste bandwidth fetching the original one.  It goes without saying that \npeople pulling from the repository mainly care about the view upstream \nchooses to publish.  Upstream can choose to rewrite, which will cause \nbreakage and is a sort of sneaky way to hide the original history, or \nthey can use replace, which avoids the breakage and gives the client the \nchoice of which view to use.\n"},{"id":"159300","messageId":"20110111053653.GB10094@sigill.intra.peff.net","threadId":"26222","inReplyTo":"7vmxnc48yt.fsf@alter.siamese.dyndns.org","subject":"Re: clone breaks replace","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-01-11T05:36:54Z","receivedAt":"2011-01-11T05:36:54Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jan 07, 2011 at 02:58:34PM -0800, Junio C Hamano wrote:\n\n> Jeff King <peff@peff.net> writes:\n> \n> >   2. Should clone fetch refs from refs/replace (either by default, or\n> >      with an option)?\n> > ...\n> > Which leads to the second question. It is basically a matter of saying\n> > \"do you want to fetch the view that upstream has\"? I can definitely see\n> > that being useful, and meriting an option. However, it may or may not be\n> > worth turning on by default, as upstream's view may be confusing.\n> \n> I think that should be stated a bit differently.  \"Do you want to fetch\n> the view that the upstream offers as an option, and if you want, which\n> ones (meaning: there could be more than one replacement grafts to give\n> different views)?\"\n\nSure, I think that is a sane way for the user to think about it, but do\nwe actually support multiple views? I thought replacement objects were\nall or nothing.\n\n-Peff\n"},{"id":"159301","messageId":"20110111054735.GC10094@sigill.intra.peff.net","threadId":"26222","inReplyTo":"4D27B33C.2020907@cfl.rr.com","subject":"Re: clone breaks replace","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-01-11T05:47:35Z","receivedAt":"2011-01-11T05:47:35Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jan 07, 2011 at 07:43:40PM -0500, Phillip Susi wrote:\n\n> >Based on previous discussions, I think the answer to the first is no.\n> >The resulting repo violates a fundamental assumption of git. Yes,\n> >because of the replacement object, many things will still work. But many\n> >parts of git intentionally do not respect replacement, and they will be\n> >broken.\n> \n> What parts do not respect replacement?  More importantly, what parts\n> will be broken?  The man page seems to indicate that about the only\n> thing that does not by default is reachability testing, which to me\n> means fsck and prune.  It seems to be the purpose of replace to\n> /prevent/ breakage and be respected by default, unless doing so would\n> cause harm, which is why fsck and prune do not.\n\nOff the top of my head, I don't know. I suspect it would take somebody\nwriting a patch to create such an incomplete repository (or making one\nmanually) and seeing how badly things broke. Maybe nothing would, and I\nam being overly conservative. It just makes me nervous to start\nviolating what has always been a fundamental assumption about the object\ndatabase (though as I pointed out, we did start violating it with\nshallow clones, so maybe it is not so bad).\n\n> >Instead, I think of replacements as a specific view into history, not a\n> >fundamental history-changing operation itself. Which means you can never\n> >save bandwidth or space by truncating history with replacements. You can\n> >only give somebody the full history, and share with them your view. If\n> >you want to truncate, you must rewrite history[1].\n> \n> Right, but if you only care about that view, then there is no need to\n> waste bandwidth fetching the original one.  It goes without saying\n> that people pulling from the repository mainly care about the view\n> upstream chooses to publish.  Upstream can choose to rewrite, which\n> will cause breakage and is a sort of sneaky way to hide the original\n> history, or they can use replace, which avoids the breakage and gives\n> the client the choice of which view to use.\n\nOnce you have fetched with that view, how locked into that view are you?\nCertainly you can never push to or be the fetch remote for another\nrepository that does not want to respect that view, because you simply\ndon't have the objects to complete the history for them.\n\nBut what about deepening your own repo? In your proposal, I contact the\nserver and ask for the replacement refs along with the branch refs. For\nthe history of the branches, it gives me the truncated version with the\nreplacement objects, right? Now how do I go back later and say \"I'm\ninterested in getting the rest of history, give me the real one\"?\n\nI guess you can get the parent pointer from the real, \"non-replaced\"\nobject and ask for it. But you can't ask for a specific commit, so for\nevery such truncation, the parent needs to publish an extra ref (but\n_not_ make it one of the ones fetched by default, or it would nullify\nyour original shallow fetch), and we need to contact them and find that\nref.\n\nSo I guess it's do-able, but there are a few interesting corners. I\nthink somebody would need to whip up a proof of concept patch to explore\nthose corners.\n\n-Peff\n"},{"id":"159305","messageId":"20110111065244.GB8631@burratino","threadId":"26222","inReplyTo":"20110111054735.GC10094@sigill.intra.peff.net","subject":"Re: clone breaks replace","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2011-01-11T06:52:44Z","receivedAt":"2011-01-11T06:52:44Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Jeff King wrote:\n> On Fri, Jan 07, 2011 at 07:43:40PM -0500, Phillip Susi wrote:\n\n>> What parts do not respect replacement?  More importantly, what parts\n>> will be broken?\n[...]\n> Off the top of my head, I don't know. I suspect it would take somebody\n> writing a patch to create such an incomplete repository (or making one\n> manually) and seeing how badly things broke.\n\nI have two worries:\n\n - first, how easily can the replacement be undone? (as you mention\n   below)\n - second, what happens if the two ends of transport have different\n   replacements?\n\nThat second worry is the more major in my opinion.  Shallow clones are\na different story --- they do not fundamentally change the history and\nthey have special support in git protocol.  It is possible to punt on\nboth by saying that (1) replacements _cannot_ be undone --- a second\nreplacement is needed --- and (2) the receiving end of a connection is\nnot allowed to have any replacements for objects in common that the\nsending end does not have, but then does that buy you anything\nsignificant over a filter-branch?\n"},{"id":"159344","messageId":"4D2C7611.6060204@cfl.rr.com","threadId":"26222","inReplyTo":"20110111054735.GC10094@sigill.intra.peff.net","subject":"Re: clone breaks replace","fromName":"Phillip Susi","fromEmail":"psusi@cfl.rr.com","sentAt":"2011-01-11T15:24:01Z","receivedAt":"2011-01-11T15:24:01Z","isPatch":false,"sender":{"key":"psusi@cfl.rr.com","avatar":null},"body":"On 1/11/2011 12:47 AM, Jeff King wrote:\n> Once you have fetched with that view, how locked into that view are you?\n> Certainly you can never push to or be the fetch remote for another\n> repository that does not want to respect that view, because you simply\n> don't have the objects to complete the history for them.\n\nIf you want to fetch the original history, then it is as simple as git\n--no-replace-objects fetch.  Unless of course, the upstream repository\nactually removed the original history ( or you are pulling from someone\nelse who only pulled the truncated history ), possibly transplanting it\nto a historical repository that they should refer you to in the message\nof the replace commit.  Then you just fetch from there instead, and\nviola!  You have the complete original history.\n\n> I guess you can get the parent pointer from the real, \"non-replaced\"\n> object and ask for it. But you can't ask for a specific commit, so for\n> every such truncation, the parent needs to publish an extra ref (but\n> _not_ make it one of the ones fetched by default, or it would nullify\n> your original shallow fetch), and we need to contact them and find that\n> ref.\n\nYes, either a new branch or separate historical repository could be\npublished to pull the original history from, or git would need to pass\nthe --no-replace-objects flag to git-upload-pack on the server, causing\nit to ignore the replace and send the original history.\n"},{"id":"159345","messageId":"4D2C7948.6080304@cfl.rr.com","threadId":"26222","inReplyTo":"20110111065244.GB8631@burratino","subject":"Re: clone breaks replace","fromName":"Phillip Susi","fromEmail":"psusi@cfl.rr.com","sentAt":"2011-01-11T15:37:44Z","receivedAt":"2011-01-11T15:37:44Z","isPatch":false,"sender":{"key":"psusi@cfl.rr.com","avatar":null},"body":"On 1/11/2011 1:52 AM, Jonathan Nieder wrote:\n> I have two worries:\n> \n>  - first, how easily can the replacement be undone? (as you mention\n>    below)\n\ngit replace -d id, or git --no-replace-objects.  It also might be nice\nto add a new switch to git replace to disable a replace without deleting\nit, so that it can later be enabled again.\n\n>  - second, what happens if the two ends of transport have different\n>    replacements?\n\nThen you have a conflict, just like if the two ends have different tags\nwith the same name.\n\n> That second worry is the more major in my opinion.  Shallow clones are\n> a different story --- they do not fundamentally change the history and\n> they have special support in git protocol.  It is possible to punt on\n> both by saying that (1) replacements _cannot_ be undone --- a second\n> replacement is needed --- and (2) the receiving end of a connection is\n> not allowed to have any replacements for objects in common that the\n> sending end does not have, but then does that buy you anything\n> significant over a filter-branch?\n\nOne of the major advantages of replacements is that they can easily be\nundone, so defeating that would be silly.  Just like with conflicting\ntags, if the receiving end has conflicting replacements, they will be\nkept instead of the remote version and a warning issued.  If you want\nthe remote version, delete your local one and fetch again.\n\nWhat it buys you over filter-branch is:\n\n1)  Those tracking your repo don't have breakage when they next fetch\nbecause the chain of commits they were tracking has been destroyed and\nreplaced by a completely different one.\n\n2)  It is obvious when a replace has been done, and the original is\nstill available.  This is good for auditing and traceability.  Paper\ntrails are good.\n\n3)  Inserting a replace record takes a lot less cpu and IO than\nfilter-branch rewriting the entire chain.\n"},{"id":"159349","messageId":"20110111173922.GB1833@sigill.intra.peff.net","threadId":"26222","inReplyTo":"4D2C7611.6060204@cfl.rr.com","subject":"Re: clone breaks replace","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-01-11T17:39:22Z","receivedAt":"2011-01-11T17:39:22Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jan 11, 2011 at 10:24:01AM -0500, Phillip Susi wrote:\n\n> Yes, either a new branch or separate historical repository could be\n> published to pull the original history from, or git would need to pass\n> the --no-replace-objects flag to git-upload-pack on the server, causing\n> it to ignore the replace and send the original history.\n\nAFAIK, git can't pass --no-replace-objects to the server over git:// (or\nsmart http). You would need a protocol extension.\n\nAnd here's another corner case I thought of:\n\nSuppose you have some server S1 with this history:\n\n  A--B--C--D\n\nand a replace object truncating history to look like:\n\n  B'--C--D\n\nYou clone from S1 and have only commits B', C, and D (or maybe even B,\ndepending on the implementation). But definitely not A, nor its\nassociated tree and blobs.\n\nNow you want to fetch from another server S2, which built some commits\non the original history:\n\n  A--B--C--D--E--F\n\nYou and S2 negotiate that you both have D, which implies that you have\nall of the ancestors of D. S2 therefore sends you a thin pack containing\nE and F, which may contain deltas against objects found in D or its\nancestors. Some of which may be only in A, which means you do not have\nthem.\n\nAside from fetching the entire real history, the only solution is that\nyou somehow have to communicate to S2 exactly which objects you have,\npresumably by telling them which replacements you have used to arrive at\nthe object set you have. Which in the general case would mean actually\nshipping them your replacement refs and objects (simply handling the\nspecial case of commit truncation isn't sufficient; you could have\nreplaced any object with any other one).\n\n-Peff\n"},{"id":"159350","messageId":"7vr5cj49vi.fsf@alter.siamese.dyndns.org","threadId":"26222","inReplyTo":"20110111053653.GB10094@sigill.intra.peff.net","subject":"Re: clone breaks replace","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-01-11T17:40:17Z","receivedAt":"2011-01-11T17:40:17Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> Sure, I think that is a sane way for the user to think about it, but do\n> we actually support multiple views? I thought replacement objects were\n> all or nothing.\n\nIt is not implausible for a long running large project to restart their\nhistory from a physical root commit every year, stiching the year-long\nsegments together at their ends with replacements, to make a default clone\nto get a year's worth of the most recent history while allowing people to\nget more by asking, no?\n\nOf course, if you trust shallow-clones, you do not have to do that kind of\nhistory surgery ;-).\n"},{"id":"159352","messageId":"20110111175031.GA2085@sigill.intra.peff.net","threadId":"26222","inReplyTo":"7vr5cj49vi.fsf@alter.siamese.dyndns.org","subject":"Re: clone breaks replace","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-01-11T17:50:31Z","receivedAt":"2011-01-11T17:50:31Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jan 11, 2011 at 09:40:17AM -0800, Junio C Hamano wrote:\n\n> Jeff King <peff@peff.net> writes:\n> \n> > Sure, I think that is a sane way for the user to think about it, but do\n> > we actually support multiple views? I thought replacement objects were\n> > all or nothing.\n> \n> It is not implausible for a long running large project to restart their\n> history from a physical root commit every year, stiching the year-long\n> segments together at their ends with replacements, to make a default clone\n> to get a year's worth of the most recent history while allowing people to\n> get more by asking, no?\n\nOh, absolutely I think it is reasonable. I just meant that we do not\nhave a convenient way of saying \"fetch these replace objects, but only\nuse this particular subset\". I think you are stuck with something manual\nlike:\n\n  # grab \"view\" from upstream and name it; let's imagine it links 2010\n  # history into 2009\n  git fetch origin refs/replace/$sha1 refs/views/2009/$sha1\n\n  # now we feel like using them\n  git for-each-ref --shell --format='%(refname)' refs/views/2009 |\n    while read ref; do\n      git update-ref \"refs/replace/${ref#refs/views/2009}\" \"$ref\"\n    done\n\nWhich is a little overkill for the simple example you gave, but would\nalso handle something as complex as a view like \"pretend the foo/\nsubtree never existed\" or even \"pretend the foo/ subtree existed all\nalong\".\n\nNot that I'm sure such things are actually sane to do, performance-wise.\nThe replace system is fast, but it was designed for a handful of\nobjects, not hundreds or thousands.\n\nAnyway. My point is that we don't have the porcelain to do something\nlike managing views or enabling/disabling them in a sane manner.\n\n-Peff\n"},{"id":"159354","messageId":"20110111175621.GC15133@burratino","threadId":"26222","inReplyTo":"20110111175031.GA2085@sigill.intra.peff.net","subject":"Re: clone breaks replace","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2011-01-11T17:56:21Z","receivedAt":"2011-01-11T17:56:21Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Jeff King wrote:\n\n> I think you are stuck with something manual\n> like:\n> \n>   # grab \"view\" from upstream and name it; let's imagine it links 2010\n>   # history into 2009\n>   git fetch origin refs/replace/$sha1 refs/views/2009/$sha1\n> \n>   # now we feel like using them\n>   git for-each-ref --shell --format='%(refname)' refs/views/2009 |\n>     while read ref; do\n>       git update-ref \"refs/replace/${ref#refs/views/2009}\" \"$ref\"\n>     done\n> \n> Which is a little overkill for the simple example you gave, but would\n> also handle something as complex as a view like \"pretend the foo/\n> subtree never existed\" or even \"pretend the foo/ subtree existed all\n> along\".\n> \n> Not that I'm sure such things are actually sane to do, performance-wise.\n> The replace system is fast, but it was designed for a handful of\n> objects, not hundreds or thousands.\n> \n> Anyway. My point is that we don't have the porcelain to do something\n> like managing views or enabling/disabling them in a sane manner.\n\nMaybe something like\n\n\tgit fetch origin refs/views/2009/*:refs/replace/*\n\nexcept that that does not provide a nice way to remove to replace\nrefs when done.\n\nA potential usability enhancement might be to allow additional\nreplacement hierarchies to be requested on a per command basis, like\n\n\tGIT_REPLACE_REFS=refs/remotes/origin/views/2009 gitk --all\n\nalong the lines of GIT_NOTES_REF.\n"},{"id":"159356","messageId":"20110111180332.GD1833@sigill.intra.peff.net","threadId":"26222","inReplyTo":"20110111175621.GC15133@burratino","subject":"Re: clone breaks replace","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-01-11T18:03:32Z","receivedAt":"2011-01-11T18:03:32Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jan 11, 2011 at 11:56:21AM -0600, Jonathan Nieder wrote:\n\n> Maybe something like\n> \n> \tgit fetch origin refs/views/2009/*:refs/replace/*\n\nHeh, yeah, that is much simpler than what I did. :)\n\n> A potential usability enhancement might be to allow additional\n> replacement hierarchies to be requested on a per command basis, like\n> \n> \tGIT_REPLACE_REFS=refs/remotes/origin/views/2009 gitk --all\n> \n> along the lines of GIT_NOTES_REF.\n\nYes, that is a much better solution, IMHO.\n\n-Peff\n"},{"id":"159357","messageId":"20110111182225.GE15133@burratino","threadId":"26222","inReplyTo":"4D2C7948.6080304@cfl.rr.com","subject":"Re: clone breaks replace","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2011-01-11T18:22:25Z","receivedAt":"2011-01-11T18:22:25Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nThoughts on use cases.  Jeff already explained the main protocol\nproblem to be solved very well (thanks!).\n\nPhillip Susi wrote:\n\n> 1)  Those tracking your repo don't have breakage when they next fetch\n> because the chain of commits they were tracking has been destroyed and\n> replaced by a completely different one.\n\nThis does not require transport respecting replacements.  Just start\na new line of history and teach \"git pull\" to pull replacement refs\nfirst when requested in the refspec.\n\nIt could work like this:\n\n\talice$ git branch historical\n\talice$ git checkout --orphan newline\n\talice$ git branch newroot\n\talice$ ... hack hack hack ...\n\talice$ git replace newroot historical\n\talice$ git push world refs/replace/* +HEAD:master\n\n\tbob$ git remote show origin\n\t  URL: git://git.alice.example.com/project.git\n\t  Ref specifier: refs/replace/*:refs/replace/* refs/heads/*:refs/remotes/origin/*\n\t  HEAD branch: master\n\t  Remote branch:\n\t    master tracked\n\t  Local branch configured for 'git pull':\n\t    master merges with remote master\n\tbob$ git pull\n\tremote: Counting objects: 18, done.\n\tremote: Compressing objects: 100% (11/11), done.\n\tremote: Total 11 (delta 8), reused 0 (delta 0)\n\tUnpacking objects: 100% (11/11), done.\n\tFrom git://git.alice.example.com/project.git\n\t * [new replacement]      87a8c7yc65c87c98c87c6a87c8a     -> replace/87a8c7yc65c87c98c87c6a87c8a\n\t   a78c9df..8c98df9  master     -> origin/master\n\n> 2)  It is obvious when a replace has been done, and the original is\n> still available.  This is good for auditing and traceability.  Paper\n> trails are good.\n\nWith the method you are suggesting, others do _not_ always have the\noriginal still available.  After I fetch from you with\n--respect-hard-replacements, then while I am on an airplane I will\nhave this hard replacement ref staring at me that I cannot remove.\n\nIf the original goes missing or gets corrupted on the few machines\nthat had it, the hard replacement ref is permanent.\n\n> 3)  Inserting a replace record takes a lot less cpu and IO than\n> filter-branch rewriting the entire chain.\n\nIf the modified history is much shorter than the original (as in the\nuse case you described), would building it really take so much CPU and\nI/O?  Moreover, is the extra CPU time to keep checking all the\nreplacements on the client side worth saving that one-time CPU time\nexpenditure on the server?\n\nIf (and only if) so then I see how that could be an advantage.\n\nSorry for the longwinded message.  Hope that helps,\nJonathan\n"},{"id":"159362","messageId":"4D2CA4AC.8060005@cfl.rr.com","threadId":"26222","inReplyTo":"20110111182225.GE15133@burratino","subject":"Re: clone breaks replace","fromName":"Phillip Susi","fromEmail":"psusi@cfl.rr.com","sentAt":"2011-01-11T18:42:52Z","receivedAt":"2011-01-11T18:42:52Z","isPatch":false,"sender":{"key":"psusi@cfl.rr.com","avatar":null},"body":"On 1/11/2011 1:22 PM, Jonathan Nieder wrote:\n>> 1)  Those tracking your repo don't have breakage when they next fetch\n>> because the chain of commits they were tracking has been destroyed and\n>> replaced by a completely different one.\n> \n> This does not require transport respecting replacements.  Just start\n> a new line of history and teach \"git pull\" to pull replacement refs\n> first when requested in the refspec.\n\nThat's what I've been saying.  My statement that you quote above is\nstating why git replace is better than git filter-branch.\n\n>> 2)  It is obvious when a replace has been done, and the original is\n>> still available.  This is good for auditing and traceability.  Paper\n>> trails are good.\n> \n> With the method you are suggesting, others do _not_ always have the\n> original still available.  After I fetch from you with\n> --respect-hard-replacements, then while I am on an airplane I will\n> have this hard replacement ref staring at me that I cannot remove.\n\nThey may not have it in their local repository, but it is clear that\nthere IS an original history, and the replace record comment should tell\nthem from where they can fetch it, and those tracking the repository\nbefore the replace was added already have it.\n\nUsing filter-branch on the other hand, is a sort of dirty hack that\nviolates the integrity constrains normally in place, and can leave you\nwith a history that has no indication that there ever was more.\n\n> If the original goes missing or gets corrupted on the few machines\n> that had it, the hard replacement ref is permanent.\n\nI think it goes without saying that if you loose part of the repository,\nand there are no other copies, then you have lost part of the repository.\n\n> If the modified history is much shorter than the original (as in the\n> use case you described), would building it really take so much CPU and\n> I/O?  Moreover, is the extra CPU time to keep checking all the\n> replacements on the client side worth saving that one-time CPU time\n> expenditure on the server?\n\nIt would take more than just inserting the replace record.  I'm not sure\nwhat you mean by \"keep checking all the replacements on the client side\".\n"},{"id":"159371","messageId":"201101112032.42880.chriscool@tuxfamily.org","threadId":"26222","inReplyTo":"20110111175621.GC15133@burratino","subject":"Re: clone breaks replace","fromName":"Christian Couder","fromEmail":"chriscool@tuxfamily.org","sentAt":"2011-01-11T19:32:42Z","receivedAt":"2011-01-11T19:32:42Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"Hi,\n\nOn Tuesday 11 January 2011 18:56:21 Jonathan Nieder wrote:\n> \n> A potential usability enhancement might be to allow additional\n> replacement hierarchies to be requested on a per command basis, like\n> \n> \tGIT_REPLACE_REFS=refs/remotes/origin/views/2009 gitk --all\n> \n> along the lines of GIT_NOTES_REF.\n\nYes, it should not be much work to implement GIT_REPLACE_REFS like the above, \nbut I think it should accept a list of ref directories, for example:\n\nGIT_REPLACE _REFS=\".:bisect:refs/remotes/origin/views/2009\"\n\nBest regards,\nChristian.\n"},{"id":"159373","messageId":"201101112048.57326.j6t@kdbg.org","threadId":"26222","inReplyTo":"20110111173922.GB1833@sigill.intra.peff.net","subject":"Re: clone breaks replace","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2011-01-11T19:48:57Z","receivedAt":"2011-01-11T19:48:57Z","isPatch":false,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"On Dienstag, 11. Januar 2011, Jeff King wrote:\n> On Tue, Jan 11, 2011 at 10:24:01AM -0500, Phillip Susi wrote:\n> > Yes, either a new branch or separate historical repository could be\n> > published to pull the original history from, or git would need to pass\n> > the --no-replace-objects flag to git-upload-pack on the server, causing\n> > it to ignore the replace and send the original history.\n>\n> AFAIK, git can't pass --no-replace-objects to the server over git:// (or\n> smart http). You would need a protocol extension.\n\nWhy would you have to? git-upload-pack never looks at replacement objects.\n\n> And here's another corner case I thought of:\n>\n> Suppose you have some server S1 with this history:\n>\n>   A--B--C--D\n>\n> and a replace object truncating history to look like:\n>\n>   B'--C--D\n>\n> You clone from S1 and have only commits B', C, and D (or maybe even B,\n> depending on the implementation). But definitely not A, nor its\n> associated tree and blobs.\n\nWhy so? Cloning transfers the database using git-upload-pack, \ngit-pack-objects, git-index-pack, and git-unpack-objects. All of them have \nobject replacements disabled. (And AFAICS, there is no possibility to \n*enable* it.)\n\nTherefore, after cloning you get\n\n A--B--C--D\n\nand perhaps also the replacement object B'.\n\nHint: git grep read_replace_refs\n\n-- Hannes\n"},{"id":"159375","messageId":"20110111195107.GA18714@sigill.intra.peff.net","threadId":"26222","inReplyTo":"201101112048.57326.j6t@kdbg.org","subject":"Re: clone breaks replace","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-01-11T19:51:07Z","receivedAt":"2011-01-11T19:51:07Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jan 11, 2011 at 08:48:57PM +0100, Johannes Sixt wrote:\n\n> On Dienstag, 11. Januar 2011, Jeff King wrote:\n> > On Tue, Jan 11, 2011 at 10:24:01AM -0500, Phillip Susi wrote:\n> > > Yes, either a new branch or separate historical repository could be\n> > > published to pull the original history from, or git would need to pass\n> > > the --no-replace-objects flag to git-upload-pack on the server, causing\n> > > it to ignore the replace and send the original history.\n> >\n> > AFAIK, git can't pass --no-replace-objects to the server over git:// (or\n> > smart http). You would need a protocol extension.\n> \n> Why would you have to? git-upload-pack never looks at replacement objects.\n\nI think you missed the first part of this discussion. Phillip is\nproposing that it should, and I am arguing against it.\n\n-Peff\n"},{"id":"159376","messageId":"201101112100.32083.j6t@kdbg.org","threadId":"26222","inReplyTo":"20110111195107.GA18714@sigill.intra.peff.net","subject":"Re: clone breaks replace","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2011-01-11T20:00:31Z","receivedAt":"2011-01-11T20:00:31Z","isPatch":false,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"On Dienstag, 11. Januar 2011, Jeff King wrote:\n> I think you missed the first part of this discussion. Phillip is\n> proposing that it should, and I am arguing against it.\n\nYou're right, sorry for the noise. Now I understand this three-word-subject.\n\n-- Hannes\n"},{"id":"159380","messageId":"4D2CBC1A.9000302@cfl.rr.com","threadId":"26222","inReplyTo":"201101112100.32083.j6t@kdbg.org","subject":"Re: clone breaks replace","fromName":"Phillip Susi","fromEmail":"psusi@cfl.rr.com","sentAt":"2011-01-11T20:22:50Z","receivedAt":"2011-01-11T20:22:50Z","isPatch":false,"sender":{"key":"psusi@cfl.rr.com","avatar":null},"body":"On 1/11/2011 3:00 PM, Johannes Sixt wrote:\n> On Dienstag, 11. Januar 2011, Jeff King wrote:\n>> I think you missed the first part of this discussion. Phillip is\n>> proposing that it should, and I am arguing against it.\n> \n> You're right, sorry for the noise. Now I understand this three-word-subject.\n\nWhat it really comes down to is that you can use replace locally to\nmodify your history and it works great.  As soon as someone clones from\nyou though, they don't get the replace and so they end up with a\ndifferent history than you see.\n\nI suggested that git-upload-pack should respect replace records by\ndefault, so that people cloning your repository will get the same\nreplaced history instead of the original.\n\nIt seems that the recommended use of replace is to locally append\nhistory back on, after it has been removed upstream with git\nfilter-branch.  Using filter-branch is bad, so it makes more sense to me\nto do the remove with git replace, and then if you want to add it back,\nyou just have to disable the replace ( and maybe fetch additional objects ).\n\nThe one problem that has come up is that when you fetch and tell the\nserver you have a commit after the replace, it assumes that you also\nhave the commits prior to the replace and may delta against objects you\ndo not have.  Fixing that would require informing the server of any\nreplacements you have, and it being able to use that information to\navoid deltas against objects hidden by the replace.\n\nDoes that sound like a pretty good summary to everyone?\n"},{"id":"159381","messageId":"20110111205043.GA19928@burratino","threadId":"26222","inReplyTo":"4D2CBC1A.9000302@cfl.rr.com","subject":"Re: clone breaks replace","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2011-01-11T20:50:43Z","receivedAt":"2011-01-11T20:50:43Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Phillip Susi wrote:\n\n> It seems that the recommended use of replace is to locally append\n> history back on, after it has been removed upstream with git\n> filter-branch.  Using filter-branch is bad, so it makes more sense to me\n> to do the remove with git replace, and then if you want to add it back,\n> you just have to disable the replace ( and maybe fetch additional objects ).\n>\n> The one problem that has come up is that when you fetch and tell the\n> server you have a commit after the replace, it assumes that you also\n> have the commits prior to the replace and may delta against objects you\n> do not have.  Fixing that would require informing the server of any\n> replacements you have, and it being able to use that information to\n> avoid deltas against objects hidden by the replace.\n>\n> Does that sound like a pretty good summary to everyone?\n\nYes, except for \"Using filter-branch is bad\".  Using filter-branch is\nnot bad.  Also there are many recommended uses of replace: for example,\nto swap out a commit that builds for one that doesn't when using \"git\nbisect\", or to stage history changes before making them permanent with\nfilter-branch.\n"},{"id":"293733","messageId":"4D2CFD0A.1060901@cfl.rr.com","threadId":"26222","inReplyTo":"20110111205043.GA19928@burratino","subject":"Re: clone breaks replace","fromName":"Phillip Susi","fromEmail":"psusi@cfl.rr.com","sentAt":"2011-01-12T00:59:54Z","receivedAt":"2011-01-12T00:59:54Z","isPatch":false,"sender":{"key":"psusi@cfl.rr.com","avatar":null},"body":"On 01/11/2011 03:50 PM, Jonathan Nieder wrote:\n> Yes, except for \"Using filter-branch is bad\".  Using filter-branch is\n> not bad.\n\nIt is bad because it breaks people tracking your branch, and violates \nthe immutability of history.\n"},{"id":"159506","messageId":"20110114205308.GA15286@burratino","threadId":"26222","inReplyTo":"4D2CFD0A.1060901@cfl.rr.com","subject":"small downloads and immutable history (Re: clone breaks replace)","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2011-01-14T20:53:08Z","receivedAt":"2011-01-14T20:53:08Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Phillip Susi wrote:\n> On 01/11/2011 03:50 PM, Jonathan Nieder wrote:\n\n>> Yes, except for \"Using filter-branch is bad\".  Using filter-branch is\n>> not bad.\n>\n> It is bad because it breaks people tracking your branch, and\n> violates the immutability of history.\n\nAh, I forgot the use case.  If you are using this to at long last get\npast the limitations (e.g., inability to push) of \"fetch --depth\",\nthen yes, rewriting existing history is bad.\n\nSo what's left is some way to make the \"have\" part of transport\nnegotiation make sense in this context.  I'll be happy if it happens.\n\nThanks for clarifying.\nJonathan\n\n[note: if you occasionally use\n\n git commit; # new commit\n git tag tmp\n git checkout --orphan newroot\n git replace newroot tmp\n git tag -d tmp\n\nso the history without replacement refs is short, no rewriting of\nhistory has to take place.  Some testing and tweaking might be\nrequired to make \"git pull\" continue to fast-forward.]\n"},{"id":"159518","messageId":"4D31304D.2000102@cfl.rr.com","threadId":"26222","inReplyTo":"20110114205308.GA15286@burratino","subject":"Re: small downloads and immutable history (Re: clone breaks replace)","fromName":"Phillip Susi","fromEmail":"psusi@cfl.rr.com","sentAt":"2011-01-15T05:27:41Z","receivedAt":"2011-01-15T05:27:41Z","isPatch":false,"sender":{"key":"psusi@cfl.rr.com","avatar":null},"body":"On 01/14/2011 03:53 PM, Jonathan Nieder wrote:\n> Ah, I forgot the use case.  If you are using this to at long last get\n> past the limitations (e.g., inability to push) of \"fetch --depth\",\n> then yes, rewriting existing history is bad.\n\nI'm not really talking about using --depth, but more of the project \ndeciding to truncate the history in the central repository.\n\n> So what's left is some way to make the \"have\" part of transport\n> negotiation make sense in this context.  I'll be happy if it happens.\n\nGood point.  Whether local history is short because of --depth or \nreplace records, the same problem arises; the negotiation needs to be \nable to exclude older objects that are not present locally, rather than \nassuming that the client has the entire history if it has any at all. \nIt seems like this should just require sending the server and end point \nin addition to a start point.  In other words, not just send ID of the \nmost recent commit, but also the oldest that it has on hand, so that the \nserver can be sure that it does not deltafy against objects prior to \nthat commit.\n"}]}