{"thread":{"id":"119","subject":"Re: [darcs-devel] Darcs and git: plan of action","startedAt":"2005-04-18T23:33:46Z","lastAt":"2005-04-20T17:11:55Z","messageCount":18,"participants":["linux@horizon.com","Ray Lee","Kevin Smith","David Roundy","Patrick McFarland","Tupshin Harper","Juliusz Chroboczek","Ralph Corderoy"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"730","messageId":"20050418210436.23935.qmail@science.horizon.com","threadId":"119","inReplyTo":null,"subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"","fromEmail":"linux@horizon.com","sentAt":null,"receivedAt":"2005-04-18T23:33:46Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"> Hell no.\n> \n> The commit _does_ specify the patch uniquely and exactly, so I really \n> don't see the point. You can always get the patch by just doing a\n>\n> \tgit diff $parent_tree $thistree\n>\n> so putting the patch in the comment is not an option.\n\nEr... no.\n\nOne of darcs' big points is that it has at least two fundamentally\ndifferent *kinds* of patches.  One is the classic diff(1) style.\n\nThe other is \"replace very instace of identifier `foo` with identifier`bar`\".\n\nNote that merging such a patch with another that adds a new instance\nof \"foo\" has a quite different effect from a similar diff-style patch.\nEven though both have identical effects on the tree to which they were\ninitially merged.\n\nAnd darcs is specifically intended to support additional kinds of patches.\nAgain, all in order that the patch can work better when applied to\ntrees *other* that the one it was originally developed against.\n\n\nAnyway, the point is that, in the darcs world, it is NOT possible to\nreconstruct a patch from the before and after trees.\n"},{"id":"733","messageId":"1113869248.23938.94.camel@orca.madrabbit.org","threadId":"119","inReplyTo":"20050418210436.23935.qmail@science.horizon.com","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"Ray Lee","fromEmail":"ray-lk@madrabbit.org","sentAt":"2005-04-19T00:07:28Z","receivedAt":"2005-04-19T00:07:28Z","isPatch":false,"sender":{"key":"ray-lk@madrabbit.org","avatar":null},"body":"On Mon, 2005-04-18 at 21:04 +0000, linux@horizon.com wrote:\n> The other is \"replace very instace of identifier `foo` with identifier`bar`\".\n\nThat could be derived, however, by a particularly smart parser [1].\nAlternately, that itself could be embedded in the comment for patches\nsourced from darcs. Of course, that means patches from others are less\ncommutable than from other darcs users, but that's the price you'd pay\nfor relying on the user to explicitly note a token rename.\n\n  [1] An example: http://minnie.tuhs.org/Programs/Ctcompare/index.html\n\nAs for \"darcs mv\", that can be derived from the before/after pictures of\nthe trees.\n\n> And darcs is specifically intended to support additional kinds of patches.\n\nAnything missing out of what I listed above? (darcs has adddir and\naddfile, IIRC, but those are trivially discovered via inspection of the\ntrees as well, I think.)\n\n> Anyway, the point is that, in the darcs world, it is NOT possible to\n> reconstruct a patch from the before and after trees.\n\nNot yet, and maybe not ever, but I think we can certainly get closer to\ndiscovering what the coder was thinking during a changeset.\n\nRay\n\n"},{"id":"748","messageId":"42645969.2090609@qualitycode.com","threadId":"119","inReplyTo":"1113869248.23938.94.camel@orca.madrabbit.org","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"Kevin Smith","fromEmail":"yarcs@qualitycode.com","sentAt":"2005-04-19T01:05:45Z","receivedAt":"2005-04-19T01:05:45Z","isPatch":false,"sender":{"key":"yarcs@qualitycode.com","avatar":null},"body":"Ray Lee wrote:\n> On Mon, 2005-04-18 at 21:04 +0000, linux@horizon.com wrote:\n> \n>>The other is \"replace very instace of identifier `foo` with identifier`bar`\".\n> \n> \n> That could be derived, however, by a particularly smart parser [1].\n\nNo, it can't. Seriously. A darcs replace patch is encoded as rules, not\neffects, and it is impossible to derive the rules just by looking at the\nresults. Not difficult. Impossible. You could guess, but that's not good\nenough for darcs to be able to reliably commute the patches later.\n\nI am curious whether Linus's suggestion about including the\ncorresponding darcs patch id in the git commit comments would be good\nenough.\n\n> As for \"darcs mv\", that can be derived from the before/after pictures of\n> the trees.\n\nPerhaps. If a file is moved and edited within the same commit, I'm not\nsure that you can be certain whether it was done with d 'darcs mv' or\nnot. Requiring separate checkins for the rename and the subsequent\nmodify would make things easier on SCM's, but is impractical in real\nlife. Automated refactoring tools, for example, perform the\nrename+modify as an atomic operation.\n\nNow, git might not need to deal with any of this, because it only needs\nto work with the kernel project. But darcs does have to deal with this\nwide range of uses, as does just about any other SCM.\n\nI'm *not* advocating cluttering up git with features that are not\ndirectly needed for kernel development. I'm just trying to clarify the\nfacts so everyone can understand what darcs is trying to do.\n\nKevin\n"},{"id":"756","messageId":"1113874931.23938.111.camel@orca.madrabbit.org","threadId":"119","inReplyTo":"42645969.2090609@qualitycode.com","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"Ray Lee","fromEmail":"ray-lk@madrabbit.org","sentAt":"2005-04-19T01:42:11Z","receivedAt":"2005-04-19T01:42:11Z","isPatch":false,"sender":{"key":"ray-lk@madrabbit.org","avatar":null},"body":"On Mon, 2005-04-18 at 21:05 -0400, Kevin Smith wrote:\n> >>The other is \"replace very instace of identifier `foo` with identifier`bar`\".\n> > That could be derived, however, by a particularly smart parser [1].\n> \n> No, it can't. Seriously. A darcs replace patch is encoded as rules, not\n> effects, and it is impossible to derive the rules just by looking at the\n> results. Not difficult. Impossible.\n\nOkay, either I'm a sight stupider than I thought, or I'm not\ncommunicating well. Same net effect either way, I 'spose.\n\nIf I do a token replace in an editor (say one of those fancy new-fangled\nrefactoring thangs, or good ol' vi), a token-level comparator can\ndiscover what I did. That link I sent is an example of one such beast.\n\n> You could guess, but that's not good\n> enough for darcs to be able to reliably commute the patches later.\n\nWho said anything about guessing? If a user replaces all instances of\nfoo with bar, that's as close to proof as you can ever get, without\nrecording intent of the user at the time it's done. Now, I realize that\ndarcs *does* record intent, but I claim that's immaterial.\n\nPerhaps I'm clueless; it's happened before, I'm resigned to it happening\nagain. So, tell it to me with full jargon, if you will. When it comes\ndown to brass tacks, why does my suggestion place weaker guarantees\nabout the quality of the resulting patch operator?\n\n> > As for \"darcs mv\", that can be derived from the before/after pictures of\n> > the trees.\n> \n> Perhaps. If a file is moved and edited within the same commit, I'm not\n> sure that you can be certain whether it was done with d 'darcs mv' or\n> not.\n\nAgreed. But then you go lart the committer of that patch.\n\n> Requiring separate checkins for the rename and the subsequent\n> modify would make things easier on SCM's, but is impractical in real\n> life.\n\nEh? Why? \"darcs mv\" *is* a commit. Just because it doesn't seem to look\nlike one doesn't change the fact that you just invoked the SCM.\n\n> Automated refactoring tools, for example, perform the\n> rename+modify as an atomic operation.\n\nAnd that's harder, I agree. But unless I'm missing some nifty\nrefactoring editor out there that integrates with darcs during the edit\nsession, the user *still* has to tell the SCM about the rename manually.\n\n> Now, git might not need to deal with any of this, because it only needs\n> to work with the kernel project.\n\nIt'd be unfortunate if git were limited to such a small developer base.\n\n> I'm *not* advocating cluttering up git with features that are not\n> directly needed for kernel development.\n\nI'm not claiming you are. We want the same thing -- a nuanced SCM that\ncan take some of the drudge-work away from this stuff.\n\nRay\n\n"},{"id":"767","messageId":"4264677A.9090003@qualitycode.com","threadId":"119","inReplyTo":"1113874931.23938.111.camel@orca.madrabbit.org","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"Kevin Smith","fromEmail":"yarcs@qualitycode.com","sentAt":"2005-04-19T02:05:46Z","receivedAt":"2005-04-19T02:05:46Z","isPatch":false,"sender":{"key":"yarcs@qualitycode.com","avatar":null},"body":"Ray Lee wrote:\n> On Mon, 2005-04-18 at 21:05 -0400, Kevin Smith wrote:\n> \n>>>>The other is \"replace very instace of identifier `foo` with identifier`bar`\".\n>>>\n>>>That could be derived, however, by a particularly smart parser [1].\n>>\n>>No, it can't. Seriously. A darcs replace patch is encoded as rules, not\n>>effects, and it is impossible to derive the rules just by looking at the\n>>results. Not difficult. Impossible.\n> \n> \n> Okay, either I'm a sight stupider than I thought, or I'm not\n> communicating well. Same net effect either way, I 'spose.\n> \n> If I do a token replace in an editor (say one of those fancy new-fangled\n> refactoring thangs, or good ol' vi), a token-level comparator can\n> discover what I did. That link I sent is an example of one such beast.\n\nThe big feature of a darcs replace patch is that it works forward and\nbackward in time. Let me try to come up with an example that can help\nexplain it. Hopefully I'll get it right. Let's start with a file like\nthis that exists in a project for which both you and I have darcs repos:\n\ncat\ndog\nfish\n\nNow, you change it to:\n\ncat dog\ndog\nfish\n\nwhile I simultaneously do a replace of \"dog\" with \"plant\", resulting in:\n\ncat\nplant\nfish\n\nWe merge. The final result in both of our trees is:\n\ncat plant\nplant\nfish\n\nNotice that just by looking at my diffs, you can't tell that I used a\nreplace operation. I didn't just replace the instances of \"dog\" that\nwere in my file at that moment. I conceptually replaced all instances,\nincluding ones that aren't there yet.\n\nNow, I should mention here that I personally dislike the replace\noperation, and I think it is more dangerous than helpful. However, other\ndarcs users are quite happy with it, and it certainly is a creative and\npowerful feature.\n\nOther creative patch types have also been dreamed of. For example, a\npowerful language-specific refactoring operation has been discussed as a\nfar-future possibility. That would be safe, and cool.\n\n>>Automated refactoring tools, for example, perform the\n>>rename+modify as an atomic operation.\n> \n> And that's harder, I agree. But unless I'm missing some nifty\n> refactoring editor out there that integrates with darcs during the edit\n> session, the user *still* has to tell the SCM about the rename manually.\n\nAlthough there are no such nifty refactoring tools available today, they\nwill exist at some point. If they existed today, the world would be a\nbetter place.\n\nEven without tools, many shops have policies against checking in code\nthat won't compile. If you rename a java class, you must simultaneously\nperform the rename and modify the class name inside. If you commit\nbetween those steps, it's broken. [I do realize that the kernel doesn't\nhave java code, by the way.]\n\nI should also mention that I currently believe that Linus is correct\nthat explicit rename tracking is not required for git. I have every hope\nthat his plan for handling the more general case of \"moved text\" will\ntake care of renames as a side effect. I don't know if that will be\nsufficient to allow a two-way lossless gateway between git and darcs or\nother systems that do track renames explicitly.\n\nKevin\n"},{"id":"812","messageId":"20050419110521.GC28269@abridgegame.org","threadId":"119","inReplyTo":"1113874931.23938.111.camel@orca.madrabbit.org","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"David Roundy","fromEmail":"droundy@abridgegame.org","sentAt":"2005-04-19T11:05:26Z","receivedAt":"2005-04-19T11:05:26Z","isPatch":false,"sender":{"key":"droundy@abridgegame.org","avatar":"https://gravatar.com/avatar/e8bcfd76f63303732bdfcdba6fc8ac6ccdff8f5a224a25ffaca24bd8a4c4571f?d=mp&s=160"},"body":"On Mon, Apr 18, 2005 at 06:42:11PM -0700, Ray Lee wrote:\n> On Mon, 2005-04-18 at 21:05 -0400, Kevin Smith wrote:\n> > You could guess, but that's not good enough for darcs to be able to\n> > reliably commute the patches later.\n>\n> Who said anything about guessing? If a user replaces all instances of\n> foo with bar, that's as close to proof as you can ever get, without\n> recording intent of the user at the time it's done. Now, I realize that\n> darcs *does* record intent, but I claim that's immaterial.\n\nThe problem is, how do you know how to define a token? That's also included\nin a darcs patch.  And a darcs user may choose not to use a replace patch,\nif (for example) he's renaming a local variable, since he might not want to\nmess with other functions in the same file.\n\nGuessing the author's intent cannot reliably reproduce the author's stated\nintent.  Either we need to include that information in one form or another\n(and in one location or another), or we've got to simply disallow replaces\n(and moves?) when interacting with git.\n-- \nDavid Roundy\nhttp://www.darcs.net\n"},{"id":"872","messageId":"200504191808.26559.pmcfarland@downeast.net","threadId":"119","inReplyTo":"4264677A.9090003@qualitycode.com","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"Patrick McFarland","fromEmail":"pmcfarland@downeast.net","sentAt":"2005-04-19T22:08:20Z","receivedAt":"2005-04-19T22:08:20Z","isPatch":false,"sender":{"key":"pmcfarland@downeast.net","avatar":null},"body":"On Monday 18 April 2005 10:05 pm, Kevin Smith wrote:\n> The big feature of a darcs replace patch is that it works forward and\n> backward in time. Let me try to come up with an example that can help\n> explain it. Hopefully I'll get it right. Let's start with a file like\n> this that exists in a project for which both you and I have darcs repos:\n>\n> cat\n> dog\n> fish\n>\n> Now, you change it to:\n>\n> cat dog\n> dog\n> fish\n>\n> while I simultaneously do a replace of \"dog\" with \"plant\", resulting in:\n>\n> cat\n> plant\n> fish\n>\n> We merge. The final result in both of our trees is:\n>\n> cat plant\n> plant\n> fish\n>\n> Notice that just by looking at my diffs, you can't tell that I used a\n> replace operation. I didn't just replace the instances of \"dog\" that\n> were in my file at that moment. I conceptually replaced all instances,\n> including ones that aren't there yet.\n\nI think that's the best explanation of how it works. And that is partially why \ndarcs is so powerful.\n\n-- \nPatrick \"Diablo-D3\" McFarland || pmcfarland@downeast.net\n\"Computer games don't affect kids; I mean if Pac-Man affected us as kids, we'd \nall be running around in darkened rooms, munching magic pills and listening to\nrepetitive electronic music.\" -- Kristian Wilson, Nintendo, Inc, 1989\n"},{"id":"884","messageId":"1113950442.29444.31.camel@orca.madrabbit.org","threadId":"119","inReplyTo":"4264677A.9090003@qualitycode.com","subject":"Re: Darcs and git: plan of action","fromName":"Ray Lee","fromEmail":"ray-lk@madrabbit.org","sentAt":"2005-04-19T22:40:42Z","receivedAt":"2005-04-19T22:40:42Z","isPatch":false,"sender":{"key":"ray-lk@madrabbit.org","avatar":null},"body":"(Sorry for the delayed reply -- I'm living on tape delay for a bit.)\n\nOn Mon, 2005-04-18 at 22:05 -0400, Kevin Smith wrote:\n> >>>>The other is \"replace very instace of identifier `foo` with identifier`bar`\".\n> >>>\n> >>>That could be derived, however, by a particularly smart parser [1].\n> >>\n> >>No, it can't. Seriously. A darcs replace patch is encoded as rules, not\n> >>effects, and it is impossible to derive the rules just by looking at the\n> >>results. Not difficult. Impossible.\n> >  \n> > If I do a token replace in an editor (say one of those fancy new-fangled\n> > refactoring thangs, or good ol' vi), a token-level comparator can\n> > discover what I did. That link I sent is an example of one such beast.\n> \n> The big feature of a darcs replace patch is that it works forward and\n> backward in time.\n\nThat's *not* a feature of the token replace patch, however. That's a\nfeature of the darcs commutation machinery, correct? (With the obvious\ncaveat that darcs can only *do* the commutation if it has correctly\nnuanced darcs-style token replace patches, rather than mere ASCII\ntextual diffs.)\n\n> Let me try to come up with an example that can help\n> explain it. Hopefully I'll get it right. Let's start with a file like\n> this that exists in a project for which both you and I have darcs repos:\n> \n> cat\n> dog\n> fish\n> \n> Now, you change it to:\n> \n> cat dog\n> dog\n> fish\n> \n> while I simultaneously do a replace of \"dog\" with \"plant\", resulting in:\n> \n> cat\n> plant\n> fish\n> \n> We merge. The final result in both of our trees is:\n> \n> cat plant\n> plant\n> fish\n\nOkay, that all makes sense.\n\n> Notice that just by looking at my diffs, you can't tell that I used a\n> replace operation.\n\nHere's where we disagree. If you checkpoint your tree before the\nreplace, and immediately after, the only differences in the\nsource-controlled files would be due to the replace. And since the\nlanguage of the file is known (and thereby the tokenization -- it *is*\nwell-defined), then a tokenizer that compares the before and after trees\n(for just the files that changed, obviously), can discover what you did,\nand promote the mere ASCII diff into a token-replace diff. (The same\nsort of idea could be done for reindention, I'd hope.)\n\n> I didn't just replace the instances of \"dog\" that\n> were in my file at that moment. I conceptually replaced all instances,\n> including ones that aren't there yet.\n\nWell yes, that's exactly what we want. And the key point of all of this\nis that there's no magic here. The darcs machinery does all the\ncommutations such that the patches can wiggle together without\nconflicts. To do it's job, of course, it needs nuanced patches, rather\nthan the quite literal ones generated by diff.\n\nWe agree on everything except that it's provable that one can discover a\nreplace operation, given a before and after tree.\n\n> Now, I should mention here that I personally dislike the replace\n> operation, and I think it is more dangerous than helpful. However, other\n> darcs users are quite happy with it, and it certainly is a creative and\n> powerful feature.\n\nIt's creative alright, though I had the same misgivings. In my common\ncode workflow, I almost never have global tokens -- all my replaces\nwould be per function, so I never saw an opportunity to use it when I\nwas screwing around with darcs.\n\n> Other creative patch types have also been dreamed of. For example, a\n> powerful language-specific refactoring operation has been discussed as a\n> far-future possibility. That would be safe, and cool.\n\n<subliminal> indention patch type, indention patch type... </subliminal>\n\n> > > Automated refactoring tools, for example, perform the\n> > > rename+modify as an atomic operation.\n> > [...]\n> Although there are no such nifty refactoring tools available today, they\n> will exist at some point.\n\nYeah, I spent some time drooling over the refactoring editors before\nslapping myself and deciding I'd wait for others to live on that\nbleeding edge for a while. I've had to clean up too much code from other\npeople.\n\n> Even without tools, many shops have policies against checking in code\n> that won't compile. If you rename a java class, you must simultaneously\n> perform the rename and modify the class name inside. If you commit\n> between those steps, it's broken.\n\nI'm trying hard to find a nice way to say that's silly. I'm failing. My\nsuggestion in that case would be that the local coder commit many\npatches to a local repository, one of which is the rename. Then upon\ncompletion of the refactoring, the set of patches is committed to the\ngroup repository. Tags before and after preserve the repository's\nprecondition that it always compiles.\n\n> [I do realize that the kernel doesn't have java code, by the way.]\n\nDon't worry, I didn't think that you did :-).\n\nRay\n"},{"id":"895","messageId":"42658D95.7020404@tupshin.com","threadId":"119","inReplyTo":"1113950442.29444.31.camel@orca.madrabbit.org","subject":"Re: Darcs and git: plan of action","fromName":"Tupshin Harper","fromEmail":"tupshin@tupshin.com","sentAt":"2005-04-19T23:00:37Z","receivedAt":"2005-04-19T23:00:37Z","isPatch":false,"sender":{"key":"tupshin@tupshin.com","avatar":null},"body":"Ray Lee wrote:\n\n>Here's where we disagree. If you checkpoint your tree before the\n>replace, and immediately after, the only differences in the\n>source-controlled files would be due to the replace.\n>\nThis is assuming that you only have one replace and no other operations \nrecorded in the patch. If you have multiple replaces or a replace and a \ntraditional diff  recorded in the same patch, then this is not true.\n\n> And since the\n>language of the file is known (and thereby the tokenization -- it *is*\n>well-defined), then a tokenizer that compares the before and after trees\n>(for just the files that changed, obviously), can discover what you did,\n>and promote the mere ASCII diff into a token-replace diff. (The same\n>sort of idea could be done for reindention, I'd hope.)\n>  \n>\nSee above for one set of limitations on this. A more fundamental problem \ncomes back to intent. If I have a file \"foo\" before:\na1\na2\nand after:\nb1\nb2\nis that a \"replace [_a-zA-Z0-9] a b foo\" patch, or is that a\n-a1\n-a2\n+b1\n+b2\npatch? Note that this comes down to heuristics, and no matter what you \nuse, you will be wrong sometimes,  *and* the choice that is made can \nsubstantively affect the contents of the repository after additional \npatches are applied.\n\n>We agree on everything except that it's provable that one can discover a\n>replace operation, given a before and after tree.\n>  \n>\nIt's provable that you can not.\n\n-Tupshin\n"},{"id":"897","messageId":"42658E38.1020406@qualitycode.com","threadId":"119","inReplyTo":"1113950442.29444.31.camel@orca.madrabbit.org","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"Kevin Smith","fromEmail":"yarcs@qualitycode.com","sentAt":"2005-04-19T23:03:20Z","receivedAt":"2005-04-19T23:03:20Z","isPatch":false,"sender":{"key":"yarcs@qualitycode.com","avatar":null},"body":"Ray Lee wrote:\n> On Mon, 2005-04-18 at 22:05 -0400, Kevin Smith wrote:\n> \n>>Notice that just by looking at my diffs, you can't tell that I used a\n>>replace operation.\n> \n> \n> Here's where we disagree. If you checkpoint your tree before the\n> replace, and immediately after, the only differences in the\n> source-controlled files would be due to the replace. \n\nBut I might have manually changed those tokens, or I might have done it\nwith a replace operation. Just looking at the diffs, those two cases\nwould look identical and be indistinguishable. The only way to know\nwhether or not a darcs replace was done was to look at the patch metadata.\n\nPop quiz:\n\nHere is revision 1 of my file:\n\n    abcde\n\nHere is revision 2:\n\n    wow\n\nNow, did I do that with a darcs replace, or just by typing?\n\nKevin\n"},{"id":"899","messageId":"1113951972.29444.42.camel@orca.madrabbit.org","threadId":"119","inReplyTo":"42658E38.1020406@qualitycode.com","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"Ray Lee","fromEmail":"ray-lk@madrabbit.org","sentAt":"2005-04-19T23:06:11Z","receivedAt":"2005-04-19T23:06:11Z","isPatch":false,"sender":{"key":"ray-lk@madrabbit.org","avatar":null},"body":"On Tue, 2005-04-19 at 19:03 -0400, Kevin Smith wrote:\n> Pop quiz:\n> Here is revision 1 of my file:\n>     abcde\n> \n> Here is revision 2:\n>     wow\n\n> Now, did I do that with a darcs replace, or just by typing?\n\nI'm still not communicating well.\n\nGive me a case where assuming it's a replace will do the wrong thing,\nfor C code, where it's a variable or function name.\n\nRay\n\n"},{"id":"905","messageId":"1113952916.29444.60.camel@orca.madrabbit.org","threadId":"119","inReplyTo":"42658D95.7020404@tupshin.com","subject":"Re: Darcs and git: plan of action","fromName":"Ray Lee","fromEmail":"ray-lk@madrabbit.org","sentAt":"2005-04-19T23:21:56Z","receivedAt":"2005-04-19T23:21:56Z","isPatch":false,"sender":{"key":"ray-lk@madrabbit.org","avatar":null},"body":"On Tue, 2005-04-19 at 16:00 -0700, Tupshin Harper wrote:\n> Ray Lee wrote:\n> \n> >Here's where we disagree. If you checkpoint your tree before the\n> >replace, and immediately after, the only differences in the\n> >source-controlled files would be due to the replace.\n> >\n> This is assuming that you only have one replace and no other operations \n> recorded in the patch. If you have multiple replaces or a replace and a \n> traditional diff  recorded in the same patch, then this is not true.\n\nI had a precondition on my argument (not quoted), that the code was\ncheckpointed before and after. Obviously, a large set of changes in one\npatch is a problem. However, a darcs replace is (effectively) a commit\non its own, so I was limiting myself to the same situation under a\ndifferent system.\n\n> A more fundamental problem comes back to intent. If I have a file\n> \"foo\" before:\n> a1\n> a2\n> and after:\n> b1\n> b2\n> is that a \"replace [_a-zA-Z0-9] a b foo\" patch, or is that a\n> -a1\n> -a2\n> +b1\n> +b2\n> patch?\n\nOkay, so in reading the online darcs manual (yet) again, I now see that\nit allows regular expressions for the match and replace, which means\nmultiple unique tokens could change atomically. (Does anyone actually\n*use* regexes? Sounds like a cannon that'd be hard to aim.)\n\nRegardless, I only care about code, not free text. If it's in a language\nthat doesn't do some use-'em-as-you-need-'em duck typing spiel\n(<cough>python</cough), then the context of your patch (namely, the\nfile) already has those tokens somewhere in them. And I bet that if\n*you* looked at that file, you could tell if it was a replace or a mere\ntextual diff. Am I wrong?\n\n> Note that this comes down to heuristics, and no matter what you \n> use, you will be wrong sometimes,  *and* the choice that is made can \n> substantively affect the contents of the repository after additional \n> patches are applied.\n\nUnless I'm missing something, the darcs replace patch can already do the\nwrong thing. If I do a replace patch on a variable introduced in a local\ntree, then do a darcs replace on it before committing it to a shared\nrepository, and coder B introduces a variable of the same original name\nin my copy, then there's a chance that the replace patch will\nincorrectly apply upon his newly introduced variable. No?\n\n> It's provable that you can not.\n\nI'm still not seeing the problem, at least when it comes to ANSI C.\n\nRay\n"},{"id":"908","messageId":"426594F9.4090002@tupshin.com","threadId":"119","inReplyTo":"1113951972.29444.42.camel@orca.madrabbit.org","subject":"Re: Darcs and git: plan of action","fromName":"Tupshin Harper","fromEmail":"tupshin@tupshin.com","sentAt":"2005-04-19T23:32:09Z","receivedAt":"2005-04-19T23:32:09Z","isPatch":false,"sender":{"key":"tupshin@tupshin.com","avatar":null},"body":"Ray Lee wrote:\n\n> I'm still not communicating well.\n>\n>Give me a case where assuming it's a replace will do the wrong thing,\n>for C code, where it's a variable or function name.\n>\n>Ray\n>\n>-\n>\nI think you are communicating fine, but not fully understanding darcs.\n\ntry this:\ninitial patch creates hello.c\n#include <stdio.h>\n\nint main(int argc, char *argv[])\n{\n  printf(\"Hello world!\\n\");\n  return 0;\n}\n\nsecond patch:\nreplace ./hello.c [A-Za-z_0-9] world universe\n\nthird patch, for conceptual clarity, created in another repository that \nhad seen the first patch, but not the second (adds function wide_world):\nhunk ./hello.c 3\n+void wide_world()\n+{\n+  printf(\"Hello wide world\\n\");\n+}\n+\nhunk ./hello.c 11\n+  wide_world();\n}\n\nIf patch2 was a replace patch, then the result of running the combined 3 \npatch version would be:\nHello universe!\nHello wide universe\n\nbut if patch2 was a non-replace patch, then the result would be:\nHello universe!\nHello wide world\n\n-Tupshin\n"},{"id":"910","messageId":"42659678.9090605@tupshin.com","threadId":"119","inReplyTo":"1113952916.29444.60.camel@orca.madrabbit.org","subject":"Re: Darcs and git: plan of action","fromName":"Tupshin Harper","fromEmail":"tupshin@tupshin.com","sentAt":"2005-04-19T23:38:32Z","receivedAt":"2005-04-19T23:38:32Z","isPatch":false,"sender":{"key":"tupshin@tupshin.com","avatar":null},"body":"Ray Lee wrote:\n\n>it allows regular expressions for the match and replace, which means\n>multiple unique tokens could change atomically. (Does anyone actually\n>*use* regexes? Sounds like a cannon that'd be hard to aim.)\n>  \n>\nYes, and replace patches need to be used very carefully.\n\n>Regardless, I only care about code, not free text. If it's in a language\n>that doesn't do some use-'em-as-you-need-'em duck typing spiel\n>(<cough>python</cough), then the context of your patch (namely, the\n>file) already has those tokens somewhere in them. And I bet that if\n>*you* looked at that file, you could tell if it was a replace or a mere\n>textual diff. Am I wrong?\n>  \n>\nYes. See my hello world example from my last email.\n\n>\n>Unless I'm missing something, the darcs replace patch can already do the\n>wrong thing. \n>\nYes, depending on how you define wrong. Darcs replace is fully \npredictable, and poorly chosen replaces can lead to incorrect results \nafter future patches are applied.\n\n>If I do a replace patch on a variable introduced in a local\n>tree, then do a darcs replace on it before committing it to a shared\n>repository, and coder B introduces a variable of the same original name\n>in my copy, then there's a chance that the replace patch will\n>incorrectly apply upon his newly introduced variable. No?\n>  \n>\nAbsolutely correct, and the exact reason why replace patches need to be \nused *very* selectively.\n\n>  \n>\n>>It's provable that you can not.\n>>    \n>>\n>\n>I'm still not seeing the problem, at least when it comes to ANSI C.\n>\n>Ray\n>  \n>\nSee hello world example in my other email. You can argue that it is an \nexisting problem in darcs, but really, it just points out the fact that \na computer is *incapable* of knowing whether it is safe to use a replace \npatch based on a diff because replace patches are dangerous if not used \nintelligently.\n\n-Tupshin\n"},{"id":"922","messageId":"1113959503.29444.91.camel@orca.madrabbit.org","threadId":"119","inReplyTo":"426594F9.4090002@tupshin.com","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"Ray Lee","fromEmail":"ray-lk@madrabbit.org","sentAt":"2005-04-20T01:11:43Z","receivedAt":"2005-04-20T01:11:43Z","isPatch":false,"sender":{"key":"ray-lk@madrabbit.org","avatar":null},"body":"Thanks for your patience.\n\nOn Tue, 2005-04-19 at 16:32 -0700, Tupshin Harper wrote:\n> >Give me a case where assuming it's a replace will do the wrong thing,\n> >for C code, where it's a variable or function name.\n\n> try this:\n> initial patch creates hello.c\n> #include <stdio.h>\n> \n> int main(int argc, char *argv[])\n> {\n>   printf(\"Hello world!\\n\");\n>   return 0;\n> }\n> \n> second patch:\n> replace ./hello.c [A-Za-z_0-9] world universe\n\nAha! Okay, I now see at least part of issue: we're using different\ndefinitions of 'token.' Yours is quite sensible, in that it matches the\ndarcs syntax. However, I'm claiming a token is defined by the file's\nlanguage, and that a replace patch on anything but a token as per those\nlanguage standards is a silly thing.\n\nIn your example, I'd claim you did an inter-token edit, as the natural\ntoken there was \"Hello world!\\n\".\n\nWith that, let me restate what I think is possible.\n\nOne should be able to discover renames (replaces) of user identifiers in\nC code programmatically. Is that everything darcs replace does?\nObviously not. Is that what users would usually *want*? If I were using\nit, that's what I'd want (especially including the limited scope of\nreplacement -- user identifiers such as variable or function names,\netc.). But then I'm not a lurker on the darcs user list, so I don't know\nhow usage of darcs replace plays out in actual practice.\n\nSo, it's a subset. Is it a useful subset? Yes, as it addresses what\nhappens during refactoring, which is when I'd usually see this getting\nused. (Syntactically ignorant search and replace is so, y'know,\n*1970s*.)\n\nAny clearer?\n\nRay\n\n"},{"id":"953","messageId":"7i3btlubpm.fsf@lanthane.pps.jussieu.fr","threadId":"119","inReplyTo":"1113959503.29444.91.camel@orca.madrabbit.org","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"Juliusz Chroboczek","fromEmail":"juliusz.chroboczek@pps.jussieu.fr","sentAt":"2005-04-20T07:52:21Z","receivedAt":"2005-04-20T07:52:21Z","isPatch":false,"sender":{"key":"juliusz.chroboczek@pps.jussieu.fr","avatar":null},"body":"> However, I'm claiming a token is defined by the file's language, and\n> that a replace patch on anything but a token as per those language\n> standards is a silly thing.\n\nPlease recall the context of this discussion: getting Darcs to grok\ngit repositories.\n\nYou are arguing that it should be possible to design a set of\nheuristics that Do The Right Thing often enough.  And you are probably\nright.\n\nBut the point is immaterial as nobody has stepped up to implement in\nDarcs the sort of heuristics you have in mind.  Partly because nobody\nhas time, but mostly because we don't like heuristics, we prefer Darcs\nto remain deterministic.\n\nSo while yes, it might be possible to get about using heuristics, it\nseems rather unlikely that that's what we'll do.\n\n                                        Juliusz\n\n"},{"id":"972","messageId":"20050420115547.GI29945@abridgegame.org","threadId":"119","inReplyTo":"1113959503.29444.91.camel@orca.madrabbit.org","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"David Roundy","fromEmail":"droundy@abridgegame.org","sentAt":"2005-04-20T11:55:52Z","receivedAt":"2005-04-20T11:55:52Z","isPatch":false,"sender":{"key":"droundy@abridgegame.org","avatar":"https://gravatar.com/avatar/e8bcfd76f63303732bdfcdba6fc8ac6ccdff8f5a224a25ffaca24bd8a4c4571f?d=mp&s=160"},"body":"On Tue, Apr 19, 2005 at 06:11:43PM -0700, Ray Lee wrote:\n> > second patch:\n> > replace ./hello.c [A-Za-z_0-9] world universe\n> \n> Aha! Okay, I now see at least part of issue: we're using different\n> definitions of 'token.' Yours is quite sensible, in that it matches the\n> darcs syntax. However, I'm claiming a token is defined by the file's\n> language, and that a replace patch on anything but a token as per those\n> language standards is a silly thing.\n\nThe trouble is that a token based on language standards is also wrong,\nunless your file at all times is syntactically correct.  It also means (for\nC in particular) that the result of the token replace isn't uniquely\ndetermined by the combination of the token replace patch and the file it\napplies to, since you need parse any header files in order to tokenize the\nC file.  In the case of header files, it may not be possible to tokenize\nthem uniquely, since they may tokenize differently depending on what other\nheader files are included before them.  And of course, none of this may be\npossible if you haven't run autoconf and configure, since you may not\nactually *have* the header files in the first place...\n\nIn a (reasonably) general-purpose tool like darcs, I think it's better to\nstick with a simpler definition of token that doesn't require a complete\nintegrated development environment.\n\nIt's also true that often you want to modify headers and string contents\nsimultaneously with the change of the code itself.  When I replace\nget_pseudowavefunction with get_atomic_orbital, I also want to modify\n\n// We call get_pseudowavefunction to get the atomic orbital...\n\nand\n\nprintf(\"Error in get_pseudowavefunction!\\n\");\n-- \nDavid Roundy\nhttp://www.darcs.net\n"},{"id":"1004","messageId":"200504201711.j3KHBtn22691@blake.inputplus.co.uk","threadId":"119","inReplyTo":"1113951972.29444.42.camel@orca.madrabbit.org","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"Ralph Corderoy","fromEmail":"ralph@inputplus.co.uk","sentAt":"2005-04-20T17:11:55Z","receivedAt":"2005-04-20T17:11:55Z","isPatch":false,"sender":{"key":"ralph@inputplus.co.uk","avatar":null},"body":"\nHi Ray,\n\n> Give me a case where assuming it's a replace will do the wrong thing,\n> for C code, where it's a variable or function name.\n\nHow about two patches.\n\n    1.  s/foo/bar/ throughout file because foo() has been decided upon\n    as the name of a new globally visible forthcoming function but was\n    already in use as a static function.\n\n    2.  Add definition of new foo().\n\nPatch 1 mustn't be a `darcs replace' despite it changing every occurence\nof the C token foo into bar.\n\nCheers,\n\n\nRalph.\n\n"}]}