{"thread":{"id":"102","subject":"Re: Darcs and git: plan of action","startedAt":"2005-04-18T12:20:21Z","lastAt":"2005-04-20T11:29:43Z","messageCount":17,"participants":["David Roundy","Linus Torvalds","Ray Lee","Juliusz Chroboczek","Petr Baudis","Tupshin Harper"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"652","messageId":"20050418122011.GA13769@abridgegame.org","threadId":"102","inReplyTo":"7ivf6lm594.fsf@lanthane.pps.jussieu.fr","subject":"Re: Darcs and git: plan of action","fromName":"David Roundy","fromEmail":"droundy@abridgegame.org","sentAt":"2005-04-18T12:20:21Z","receivedAt":"2005-04-18T12:20:21Z","isPatch":false,"sender":{"key":"droundy@abridgegame.org","avatar":"https://gravatar.com/avatar/e8bcfd76f63303732bdfcdba6fc8ac6ccdff8f5a224a25ffaca24bd8a4c4571f?d=mp&s=160"},"body":"Linus and gittish people,\n\nI'm cc'ing you on this email, since Juliusz had some interesting ideas as\nto how darcs could interact with git, which then gave me an idea concerning\nwhich I'd like feedback from you.  In particular, it would make life (that\nis, life interacting back and forth with git) easier if we were to embed\ndarcs patches in their entirety in the git comment block.  It's a bit of an\nugly idea, but would greatly simplify the two-way interaction between git\nand darcs, since no information would be lost when a darcs patch was merged\ninto git.  See below for the discussion.\n\nAs I say, it's a bit ugly, and before we explore the idea further, it would\nbe nice to know if this would cause Linus to vomit in disgust and/or refuse\npatches from darcs users.  Another slightly less noxious possibility would\nbe to store the darcs patch as a \"hidden\" file, if git were given the\nconcept of commit-specific files.  So then we could include in the commit\nlog something like \"Darcs-patch:\n780c057447d4feef015a905aaf6c87db894ff58c\".  We could do this silently,\nexcept that I wonder if fsck would delete these files, since they aren't\npointed to by any trees.\n\nOn Mon, Apr 18, 2005 at 12:02:15AM +0200, Juliusz Chroboczek wrote:\n> David,\n> \n> I've read git over the week-end.  I think I can see where it's coming\n> from.\n> \n> Git is basically a (userspace) filesystem with support for efficiently\n> finding identical objects.  It's both simple and generic enough to be\n> usable by us.\n\nRight.\n\n> You mentioned that you'd like to use git as a cache for Darcs; and I\n> don't think I agree.  Caches are tricky -- they need to be kept in\n> synch -- and they might result in unexpected performance (you need to\n> update both the native and the cached data structures on every\n> modification).\n\nIt's true that we'd need to keep the cache in sync, which would mean making\nsure it gets updated with every repository-modifying darcs command, but\nwe've already got a cache that has those properties, and it seems like\nmodifying the interface to deal with a more complex cache would be\nrelatively straightforward, and would likely have other advantages, such as\nif we wanted to implement a per-file cache to speed up annotate (since the\nspeed of annotate seems to be a relatively common concern).\n\nBasically, I'm imagining that we'd have to replace writePristine and\nwrite_dirty_Pristine with the applyPristine that Ian implemented for\nefficiency reasons.  So we'd write to pristine by throwing patches at it,\nand letting it do what it pleases with them.  Then we'd read from Pristine\nas usual--but we might want to add interfaces for reading slurpies of older\nversions from the pristine cache.  This would again be a helpful interface\nanyways, since it might allow us, for example, to use checkpoints when\nreading older versions.\n\n> I'd rather remodularise Darcs so that the on-disk patch representation\n> is decoupled from the in-memory representation, so that we can use\n> various backends in the same way as we use the native repository\n> format.\n\nThe problem I have with this is that \"other\" repository formats (e.g. git)\nstore \"tree versions\", not \"changes\", and I think it would be fragile to\ntry to store \"changes\" (in the darcs sense) in them.\n\n> As you seem motivated by git (my motivation is slightly different -- I\n> want to be able to pull from Arch and other widespread systems with\n> dysfunctional user interfaces), I suggest that we start with that.\n\nI see.  You're thinking of using darcs as a client for other SCMs.  That's\nsort of how I'm thinking of darcs interacting with git, so we aren't so far\noff in terms of goals.  My hope would tend to be that people would coalesce\naround git--since Linus will be using git.  If everyone can interoperate\nwith git, we'd be able to interoperate with everyone, in a sense, anyways.\n\n> I suggest we do the following:\n> \n>  1. remove the assumption that patch IDs have a fixed format.  Patch\n>     IDs should be opaque blobs of binary data that Darcs only compares\n>     for equality.\n\nI'm not really comfortable with this, although I can see that there is an\nappeal to it, and that something like it may turn out to be necesary for\ninteracting with systems for which we can't create a simple mapping of\npatch IDs.\n\n>  2. get Darcs to pull from git.  By restricting ourselves to a fairly\n>     simple command, this should be doable in finite time.\n\nOkay, this is definitely a good goal.  See below for thoughts on how this\nshould be accomplished.\n\n>  3. allow a patch to have multiple IDs; if the IDs associated to two\n>     patches are not disjoint, then the patches are the same patch.\n\nThis I find a bit confusing.  So a patch can have two IDs, presumably\nsomething like a \"darcs ID\" and a \"git ID\"? I can see that this might\nsimplify some things, but am not sure how it would work.  The IDs would\nhave to have a hierarchy, so that you wouldn't ever end up with the \"same\"\npatch having disjoint IDs in two cases.\n\n>  4. allow applying to git repos of non-merger patches.\n\nHere's where I think I'd differ.  I think when dealing with git (and\nprobably also with *any* other SCM (arch being a possible exception), we\nneed to consider the exchange medium to be not a patch, but a tag.  Git\nonly knows about \"versions\" of the tree, which in darcs terminology is a\ntag.  It *does* know about the (possibly multiple) parents of a given\nversion, so we have a \"context\" for the patch--provided those two (or\none...) parents are treated as tags.\n\nSo in pulling from git, I'd treat each git change as a patch followed by a\ntag.  When pulling from git, unfortunately, the contents of that patch will\nbe determined by our diff algorithm, so if we want long-term stability we\nmight need to mummify a variant of the diff algorithm that we agree not to\nchange, and to always use when computing patches from a git archive.  This\ntagging (and I imagine the tags will look something like\n\"git:0c16636264037e8b5ccd38b28ecd191aebc67389\") will mean that we can\ncreate a single-patch darcs \"patch bundle\" for any given git commit.  Which\nis to say, that we'll be able to \"see\" a git repository as an odd-looking\ndarcs repository.\n\nThis means that getting a fresh darcs repository from git would potentially\ninvolve a whole lot of merging...\n\nPutting darcs patches *into* git is more complicated, since we'll want to\nget them back again without modification.  Normal \"hunk\" patches would be\nno problem, provided we never change our diff algorithm (which has been\ndiscussed recently, in the context of making hunks better align with blocks\nof code).  We could perhaps tell users not to use \"replace\" patches.  But\navoiding \"mv\" patches would be downright silly.  So we're somehow going to\nhave to either sneak this sort of metadata into the git repository, or\nwe're going to have to store just the darcs \"patch ID\" in the git\nrepository, and require that darcs users get the actual patch from\nsomewhere else.  I had been imagining the latter, but now I'm wondering if\nthe former is a reasonable possibility.\n\nLinus has said that he figures an SCM needs to be built on top of git, and\nthat SCM--rather than git itself--would be the one that would know about\nthings like file renames, probably by storing some sort of rename metadata.\nI wonder if we could perhaps store the entire darcs patch in the git\ncommit? It seems a bit abusive, but would certainly be the easiest way to\ninterface losslessly with git.  So when we pull from git, we'd look in the\ncommit log for the magic words indicating that this is really a darcs\npatch.  If so, we could handle it natively.  If not, we'd know it was\nactually a \"gittish\" entity, and requires that we diff a couple of trees to\nfind the actual patch to be used with darcs' patch theory.\n\nThe ugliness of this idea is that it involves storing redundant\ninformation.  And I think we'll have a bit of an excercise in commutation\nand merging when we get the patches from a git repository, since in git\nthey'll be stored in tree form, but that's something we'll have to do even\nto create a read-only darcs mirror of a git repository.\n\nWe could perhaps alleviate the pain by perhaps not including the actual\ncontents of new or deleted files in the patch, but instead retrieve those\nfrom git directly.  But that might be more trouble than it's worth, at\nleast for the first sketch.\n\n>  5. think about mergers.\n\nSince git stores a branched history rather than a linear history, I'm not\nsure that we'll ever need to store mergers in git.  Instead, we could just\ncommute until the mergers disappear (which might be a bit scary), and then\nstore *that* in git.  On the other hand, if this is to inefficient, and if\nwe store actual darcs patches in git, then we wouldn't perhaps need to\nworry about mergers as a special case.\n\n> Whether we end up with a useful implementation of Darcs/git or not,\n> this will result in a more modular Darcs, and hence one that will be\n> easier to optimise.\n> \n> What do you think?\n\nErr, I think I've answered that one... :)\n-- \nDavid Roundy\nhttp://www.darcs.net\n"},{"id":"664","messageId":"Pine.LNX.4.58.0504180832330.7211@ppc970.osdl.org","threadId":"102","inReplyTo":"20050418122011.GA13769@abridgegame.org","subject":"Re: Darcs and git: plan of action","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-18T15:38:25Z","receivedAt":"2005-04-18T15:38:25Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 18 Apr 2005, David Roundy wrote:\n> \n> I'm cc'ing you on this email, since Juliusz had some interesting ideas as\n> to how darcs could interact with git, which then gave me an idea concerning\n> which I'd like feedback from you.  In particular, it would make life (that\n> is, life interacting back and forth with git) easier if we were to embed\n> darcs patches in their entirety in the git comment block.\n\nHell no.\n\nThe commit _does_ specify the patch uniquely and exactly, so I really \ndon't see the point. You can always get the patch by just doing a\n\n\tgit diff $parent_tree $thistree\n\nso putting the patch in the comment is not an option.\n\nThen you can use the patch to index to whatever extra \"darcs index\" \ninformation you want to.\n\n> As I say, it's a bit ugly, and before we explore the idea further, it would\n> be nice to know if this would cause Linus to vomit in disgust and/or refuse\n> patches from darcs users.\n\nThat's definitely the case. I will _not_ be taking random files etc just \nto keep other peoples stuff straightened up.\n\nIf you want to add a \"log ID\", you can certainly do that, but the data the \nID refers to is _you_ data, and will not go into the git archive. So:\n\n> Another slightly less noxious possibility would\n> be to store the darcs patch as a \"hidden\" file, if git were given the\n> concept of commit-specific files.\n\nNo, git will not track commit-specific files. There's the comment section,\nand that _is_ the commit-specific file. But I will refuse to take any\ncomments that aren't just human-readable explanations, together with maybe \none extra line of\n\n\t# Darcs ID: 780c057447d4feef015a905aaf6c87db894ff58c\n\n(others will want to track _their_ PR numbers etc) and that's it. The \nactual darcs data that that ID refers to can obviously be maintained in \n_another_ git archive, but it's not one I'm going to carry about.\n\n\t\t\tLinus\n"},{"id":"677","messageId":"1113849317.23938.48.camel@orca.madrabbit.org","threadId":"102","inReplyTo":"20050418122011.GA13769@abridgegame.org","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"Ray Lee","fromEmail":"ray-lk@madrabbit.org","sentAt":"2005-04-18T18:35:16Z","receivedAt":"2005-04-18T18:35:16Z","isPatch":false,"sender":{"key":"ray-lk@madrabbit.org","avatar":null},"body":"On Mon, 2005-04-18 at 08:20 -0400, David Roundy wrote:\n> Putting darcs patches *into* git is more complicated, since we'll want to\n> get them back again without modification.  Normal \"hunk\" patches would be\n> no problem, provided we never change our diff algorithm (which has been\n> discussed recently, in the context of making hunks better align with blocks\n> of code).  We could perhaps tell users not to use \"replace\" patches.  But\n> avoiding \"mv\" patches would be downright silly.\n\nOkay, I still haven't used git yet (and have only toyed around with\ndarcs for a bit), so take what I'm saying with a grain of salt.\nRegardless, I think you may be asking the wrong question. The tracking\nof renames was bandied about pretty thoroughly on-list from Wednesday\nthrough Friday (for far better commentary and insight, see Linus'\nmessages with subject: Merge with git-pasky II.)\n\ngit does track changesets that describe the parent tree(s) and the\nresult. The trees track filenames and hashes. So, doing a fairly\nstraightforward compare on two trees will let you immediately discover\nrenames that have occurred, as the filename in the tree changed while\nthe hash didn't.\n\nSo, the question then becomes, can an outside tool cheaply derive all\nthe information that darcs would need to perform it's work? The renames\nshould be easy, as long as no content changed during the rename. As for\ntoken replacement (and whitespace changes, etc.), that could be\ndiscovered via domain-specific parsers (something specific per language,\nfor example). Linus tossed a link to one such tool (hmm, where was it.\nSheesh. You sure right a lot, dude :-).)\n\n\thttp://minnie.tuhs.org/Programs   (see Ctcompare)\n\n...which should be viewed more as a proof-of-concept than a mergeable\ncode-set. It does show that diff's vocabulary is sadly lacking in\nexpressiveness, and improving that, I think, would be a useful area to\nexpend effort. \n\nAgain, I may be off here, especially considering I've a backlog of a\ncouple hundred messages to read since the weekend. (You guys need to go\noutside more often.)\n\nRay\n\n"},{"id":"744","messageId":"7iy8bf7fh2.fsf@lanthane.pps.jussieu.fr","threadId":"102","inReplyTo":"20050418122011.GA13769@abridgegame.org","subject":"Re: Darcs and git: plan of action","fromName":"Juliusz Chroboczek","fromEmail":"juliusz.chroboczek@pps.jussieu.fr","sentAt":"2005-04-19T00:55:05Z","receivedAt":"2005-04-19T00:55:05Z","isPatch":false,"sender":{"key":"juliusz.chroboczek@pps.jussieu.fr","avatar":null},"body":"[Using git as a backend for Darcs.]\n\n> The problem I have with this is that \"other\" repository formats (e.g. git)\n> store \"tree versions\", not \"changes\", and I think it would be fragile to\n> try to store \"changes\" (in the darcs sense) in them.\n\nNot really; a Darcs patch is just a pair of two git versions (from and\nto).  Which is why Darcs needs to support arbitrarily formatted patch\nids -- a patch originating from git will be identified by a pair of\ngit hashes.\n\nObviously, we'll need to think harder when pushing from darcs into git\n(we'll need to preserve the Darcs patch id somehow), but it's premature\nto worry about that right now.\n\n>>  1. remove the assumption that patch IDs have a fixed format.  Patch\n>>     IDs should be opaque blobs of binary data that Darcs only compares\n>>     for equality.\n\n> I'm not really comfortable with this,\n\nWhy?\n\n>>  3. allow a patch to have multiple IDs; if the IDs associated to two\n>>     patches are not disjoint, then the patches are the same patch.\n>\n> This I find a bit confusing.  So a patch can have two IDs, presumably\n> something like a \"darcs ID\" and a \"git ID\"? I can see that this might\n> simplify some things, but am not sure how it would work.  The IDs would\n> have to have a hierarchy, so that you wouldn't ever end up with the \"same\"\n> patch having disjoint IDs in two cases.\n\nIt's a case of ``don't do that''.\n\nSuppose I record a patch in Darcs; it gets a Darcs id.  I push it into\ngit, at which point it gets a git id, whether we want it to or not.\nWhat do we do when we pull that patch back into darcs?\n\nEither we arbitrarily discard one of the ids (which one?), or we keep\nboth.  If there's more pulling/pushing going on on the git side, we\ndefinitely need to keep both.\n\n> Here's where I think I'd differ.\n\nSame to you ;-)\n\n> I think when dealing with git (and probably also with *any* other\n> SCM (arch being a possible exception), we need to consider the\n> exchange medium to be not a patch, but a tag.\n\nWe're thinking in opposite directions -- you're thinking of the alien\nversions as integrals of Darcs patches, I'm thinking of Darcs patches\nas derivatives of alien versions.\n\n  You:  alien version = Darcs tag\n\n  Me:   Darcs patch = pair of successive alien versions\n\nMy gut instinct is that the second model can be made to work almost\nseamlessly, unlike the first one.  But that's just a guess.\n\n> if we want long-term stability we might need to mummify a variant of\n> the diff algorithm that we agree not to change,\n\nGood point, noted.\n\n> But avoiding \"mv\" patches would be downright silly.\n\nAye, that will require some metadata on the git side (the hack,\nsuggested by Linus, of using git hashes to notice moves won't work).\nHappily, it's premature to worry about that, too.\n\n                                        Juliusz\n"},{"id":"757","messageId":"1113874991.23938.113.camel@orca.madrabbit.org","threadId":"102","inReplyTo":"7iy8bf7fh2.fsf@lanthane.pps.jussieu.fr","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"Ray Lee","fromEmail":"ray@madrabbit.org","sentAt":"2005-04-19T01:43:11Z","receivedAt":"2005-04-19T01:43:11Z","isPatch":false,"sender":{"key":"ray@madrabbit.org","avatar":null},"body":"On Tue, 2005-04-19 at 02:55 +0200, Juliusz Chroboczek wrote:\n> > But avoiding \"mv\" patches would be downright silly.\n> \n> Aye, that will require some metadata on the git side (the hack,\n> suggested by Linus, of using git hashes to notice moves won't work).\n\nOkay, I'm coming to believe I missed something. So, why won't it work?\n\nRay\n\n"},{"id":"792","messageId":"7i7jizyy4i.fsf@lanthane.pps.jussieu.fr","threadId":"102","inReplyTo":"1113874991.23938.113.camel@orca.madrabbit.org","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"Juliusz Chroboczek","fromEmail":"juliusz.chroboczek@pps.jussieu.fr","sentAt":"2005-04-19T08:22:21Z","receivedAt":"2005-04-19T08:22:21Z","isPatch":false,"sender":{"key":"juliusz.chroboczek@pps.jussieu.fr","avatar":null},"body":"> > Aye, that will require some metadata on the git side (the hack,\n> > suggested by Linus, of using git hashes to notice moves won't work).\n\n> So, why won't it work?\n\nBecause two files can legitimately have identical contents without\nbeing ``the same'' file from the VC system's point of view.\n\nIn other words, two files may happen to have the same contents but\nhave distinct histories.\n\n                                        Juliusz\n\n\n"},{"id":"809","messageId":"20050419104252.GA28269@abridgegame.org","threadId":"102","inReplyTo":"Pine.LNX.4.58.0504180832330.7211@ppc970.osdl.org","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"David Roundy","fromEmail":"droundy@abridgegame.org","sentAt":"2005-04-19T10:42:53Z","receivedAt":"2005-04-19T10:42:53Z","isPatch":false,"sender":{"key":"droundy@abridgegame.org","avatar":"https://gravatar.com/avatar/e8bcfd76f63303732bdfcdba6fc8ac6ccdff8f5a224a25ffaca24bd8a4c4571f?d=mp&s=160"},"body":"On Mon, Apr 18, 2005 at 08:38:25AM -0700, Linus Torvalds wrote:\n> On Mon, 18 Apr 2005, David Roundy wrote:\n> > .... In particular, it would make life (that is, life interacting back\n> > and forth with git) easier if we were to embed darcs patches in their\n> > entirety in the git comment block.\n> \n> Hell no.\n\nI was afraid that would be the response...\n\n> The commit _does_ specify the patch uniquely and exactly, so I really \n> don't see the point. You can always get the patch by just doing a\n> \n> \tgit diff $parent_tree $thistree\n> \n> so putting the patch in the comment is not an option.\n\nThe issue is that in darcs the parent and child trees *don't* uniquely or\nexactly specify the patch.  In fact, even the output of git diff will\ndepend on what version of diff you're using (e.g. if someone were to use\nBSD diff rather than GNU diff).\n\n> > As I say, it's a bit ugly, and before we explore the idea further, it would\n> > be nice to know if this would cause Linus to vomit in disgust and/or refuse\n> > patches from darcs users.\n> \n> That's definitely the case. I will _not_ be taking random files etc just \n> to keep other peoples stuff straightened up.\n\nOkay.\n\n> > Another slightly less noxious possibility would be to store the darcs\n> > patch as a \"hidden\" file, if git were given the concept of\n> > commit-specific files.\n> \n> No, git will not track commit-specific files. There's the comment\n> section, and that _is_ the commit-specific file. But I will refuse to\n> take any comments that aren't just human-readable explanations, together\n> with maybe one extra line of\n> \n> \t# Darcs ID: 780c057447d4feef015a905aaf6c87db894ff58c\n> \n> (others will want to track _their_ PR numbers etc) and that's it. The \n> actual darcs data that that ID refers to can obviously be maintained in \n> _another_ git archive, but it's not one I'm going to carry about.\n\nThe trouble is that the philosophy of darcs and git are about as orthogonal\nas one can come.  Git treats the content as fundamental, where in darcs the\nchanges are fundamental.  Since in darcs there can be different changes\nthat lead from the same parent to the same child--and these differences are\nmeaningful when merges happen---when interacting with git, we either need\nto restrict darcs to only describe changes in a way that can be uniquely\ndetermined by a parent and child, or we need to have extra metadata\nsomewhere.\n\nFor bidirectional functionality, we either need to avoid the use of\nadvanced darcs features, or we need to include that information in git\nsomehow, or we need to keep a parallel darcs archive holding that\ninformation.\n\nWould a small amount of human-readable change information be acceptable in\nthe free-form comment area? In the rename thread I got the impression this\nwould be okay for renames.  For example,\n\nrename foo bar\n\nor (this is less important, but you might consider it to be a useful\nhuman-readable comment)\n\nreplace [_a-zA-Z0-9] old_variable new_variable file/path\n\nCurrently these two patch types account for almost the sum total of the\ncases where different patches lead to the same resulting trees.\n-- \nDavid Roundy\n"},{"id":"811","messageId":"20050419110407.GB28269@abridgegame.org","threadId":"102","inReplyTo":"7iy8bf7fh2.fsf@lanthane.pps.jussieu.fr","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"David Roundy","fromEmail":"droundy@abridgegame.org","sentAt":"2005-04-19T11:04:12Z","receivedAt":"2005-04-19T11:04:12Z","isPatch":false,"sender":{"key":"droundy@abridgegame.org","avatar":"https://gravatar.com/avatar/e8bcfd76f63303732bdfcdba6fc8ac6ccdff8f5a224a25ffaca24bd8a4c4571f?d=mp&s=160"},"body":"On Tue, Apr 19, 2005 at 02:55:05AM +0200, Juliusz Chroboczek wrote:\n> [Using git as a backend for Darcs.]\n...\n> >>  1. remove the assumption that patch IDs have a fixed format.  Patch\n> >>  IDs should be opaque blobs of binary data that Darcs only compares\n> >>  for equality.\n> \n> > I'm not really comfortable with this,\n> \n> Why?\n\nI'm not clear why it would be necesary, and it takes the only immutable\npiece of information regarding a patch, and makes it variable.  Just seems\ndangerous and complicated, and I'm not sure why we'd need to do it.\n\n> Suppose I record a patch in Darcs; it gets a Darcs id.  I push it into\n> git, at which point it gets a git id, whether we want it to or not.\n> What do we do when we pull that patch back into darcs?\n> \n> Either we arbitrarily discard one of the ids (which one?), or we keep\n> both.  If there's more pulling/pushing going on on the git side, we\n> definitely need to keep both.\n\nOr alternatively, we could have a one-to-one mapping between git IDs and\ndarcs IDs, which is what I'd do.\n\n> > I think when dealing with git (and probably also with *any* other SCM\n> > (arch being a possible exception), we need to consider the exchange\n> > medium to be not a patch, but a tag.\n> \n> We're thinking in opposite directions -- you're thinking of the alien\n> versions as integrals of Darcs patches, I'm thinking of Darcs patches\n> as derivatives of alien versions.\n> \n>   You:  alien version = Darcs tag\n> \n>   Me:   Darcs patch = pair of successive alien versions\n> \n> My gut instinct is that the second model can be made to work almost\n> seamlessly, unlike the first one.  But that's just a guess.\n\nThe problem is that there is no sequence of alien versions that one can\ndifferentiate.  Git has a branched history, with each version that follows\na merge having multiple parents.  How do you define that change?  It's easy\nenough to do if we tag each git version in darcs, since we know what the\ntwo parents are, and we know what the final state is, but there *is* no\ntranslation from a single git ID either to a single patch(1) patch, or to a\nsingle darcs patch--unless you treat its parents as tags.\n\nThe key is that we can't make git work like darcs, so we'll have to make\ndarcs work like git.  If we do it right (automatically tagging like crazy\npeople), darcs users between themselves can cherry-pick all they like,\nwithout introducing inconsistencies or losing interoperability with git.\n\nTo summarize how I'd see the mapping between git information and darcs, a\ngit commit would be composed of one darcs patch and one darcs tag.  With\nthis mapping, I don't believe we lose any information, and I believe we'll\nbe able to (except that patches would have to be uniquely determined by a\npair of trees) simply translate the darcs system right back again, since\nit's a one-to-one correspondence of information.\n\nMy proposed mapping:\n\ntree 6ff0e9f3d131bd110d32829f0b14f07da8313c45\n# This is a darcs tag ID\nparent abd62b9caee377595a9bf75f363328c82a38f86e\n# This is the context of both a patch and tag.\nauthor James Bottomley <James.Bottomley@SteelEye.com> 1113879319 -0700\n# This is the author and date of the patch\ncommitter Linus Torvalds <torvalds@ppc970.osdl.org.(none)> 1113879319 -0700\n# This is the author and date of the tag\n# Everything below would be the name and long comment of the patch\n\n[PATCH] SCSI trees, merges and git status\n\nDoing the latest SCSI merge exposed two bugs in your merge script:\n\n1) It doesn't like a completely new directory (the misc tree contains a\n   new drivers/scsi/lpfc)\n2) the merge testing logic is wrong.  You only want to exit 1 if the\n   merge fails. \n\n\n-- \nDavid Roundy\nhttp://www.darcs.net\n"},{"id":"813","messageId":"7i4qe3x8ig.fsf@lanthane.pps.jussieu.fr","threadId":"102","inReplyTo":"20050419110407.GB28269@abridgegame.org","subject":"Re: Darcs and git: plan of action","fromName":"Juliusz Chroboczek","fromEmail":"juliusz.chroboczek@pps.jussieu.fr","sentAt":"2005-04-19T12:20:55Z","receivedAt":"2005-04-19T12:20:55Z","isPatch":false,"sender":{"key":"juliusz.chroboczek@pps.jussieu.fr","avatar":null},"body":"[Removing Linus from CC, keeping the Git list -- or should we remove it?]\n\n> I'm not clear why it would be necesary, and it takes the only immutable\n> piece of information regarding a patch, and makes it variable.\n\nEr... I'm not suggesting to make it variable, just to make it an\nopaque blob of bytes (still immutable).  I see from the examples you\ngive below that you agree that the format needs extending, so I\nsuspect we're actually agreeing here, just failing to communicate.\n\nabout having multiple ids per patch:\n\n> Or alternatively, we could have a one-to-one mapping between git IDs and\n> darcs IDs, which is what I'd do.\n\nOkay, you've convinced me.  It's much simpler that way, we'll see how\nwell it works.\n\n> The problem is that there is no sequence of alien versions that one can\n> differentiate.  Git has a branched history, with each version that follows\n> a merge having multiple parents.\n\nYep.  I've just realised that this morning.  Is there some notion of\n``primary parent'' as in Arch?  Can a changeset have 0 parents?\n\n> If we do it right (automatically tagging like crazy people), darcs\n> users between themselves can cherry-pick all they like, without\n> introducing inconsistencies or losing interoperability with git.\n\nYou've lost me here.  How can you cherry-pick if every tag depends on\nthe preceding patches?  Or are you thinking of pulling just the patch\nand not the tag -- in that case, what happens when you push to git a\nDarcs patch that depends on a patch that originated with git?\n\nI've started interfacing Haskell with git this week-end, that's\nsomething we'll need whichever model we choose.  We should be able to\nstart playing with actually modifying Darcs after next week-end.\n\n                                        Juliusz\n"},{"id":"814","messageId":"20050419122518.GD12757@pasky.ji.cz","threadId":"102","inReplyTo":"7i4qe3x8ig.fsf@lanthane.pps.jussieu.fr","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-19T12:25:18Z","receivedAt":"2005-04-19T12:25:18Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Tue, Apr 19, 2005 at 02:20:55PM CEST, I got a letter\nwhere Juliusz Chroboczek <Juliusz.Chroboczek@pps.jussieu.fr> told me that...\n> > The problem is that there is no sequence of alien versions that one can\n> > differentiate.  Git has a branched history, with each version that follows\n> > a merge having multiple parents.\n> \n> Yep.  I've just realised that this morning.  Is there some notion of\n> ``primary parent'' as in Arch?  Can a changeset have 0 parents?\n\nYes, the root commit. Usually, there is only one, but there may be\nmultiple of them theoretically.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"828","messageId":"Pine.LNX.4.58.0504190749030.19286@ppc970.osdl.org","threadId":"102","inReplyTo":"20050419104252.GA28269@abridgegame.org","subject":"Re: Darcs and git: plan of action","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-19T14:55:44Z","receivedAt":"2005-04-19T14:55:44Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 19 Apr 2005, David Roundy wrote:\n> \n> Would a small amount of human-readable change information be acceptable in\n> the free-form comment area? In the rename thread I got the impression this\n> would be okay for renames.  For example,\n> \n> rename foo bar\n\nSure. That's human-readable and meaningful, as in \"it actually makes sense \nas a commit comment regardless of any darcs issues\". As does:\n\n> replace [_a-zA-Z0-9] old_variable new_variable file/path\n\nwhich is almost so (a human would have written \"rename old to new\", but\nthe above isn't _that_ different).\n\nHOWEVER, then the requirement would be that we'd never have complex\ncombinations of the above. Ie having 2-5 lines of something like that is\n\"human-readable\". Having 10+ lines of the above is not. See?\n\nI have this suspicion that the \"replace\" thing often ends up being done on\ndozens of files, and I don't want to have dozens of lines of stuff that\nends up really being machine-readable. But if it's ok to depend on the\ncontent changes (you _do_ see which files changed) together with a single\nline of \"replace [token-def] xxx yyy\", then hell yes - I consider that to\nbe useful information even outside of git.\n\n(In other words: if it looks like something a careful human _could_ have\nwritten, it's certainly ok. But if it looks like something a careful human\nwould have used a script to generate 40 entries of, it's bad).\n\n\t\tLinus\n"},{"id":"831","messageId":"426532D5.3040306@tupshin.com","threadId":"102","inReplyTo":"Pine.LNX.4.58.0504190749030.19286@ppc970.osdl.org","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"Tupshin Harper","fromEmail":"tupshin@tupshin.com","sentAt":"2005-04-19T16:33:25Z","receivedAt":"2005-04-19T16:33:25Z","isPatch":false,"sender":{"key":"tupshin@tupshin.com","avatar":null},"body":"Linus Torvalds wrote:\n\n>(In other words: if it looks like something a careful human _could_ have\n>written, it's certainly ok. But if it looks like something a careful human\n>would have used a script to generate 40 entries of, it's bad).\n>\n>\t\tLinus\n>  \n>\nThis is the way that darcs would currently represent a \"darcs replace \nfoo bar\" on 15 files, which is obviously exactly what you are objecting to:\n[global foo to bar\ntupshin@tupshin.com**20050419155539] {\nreplace ./dir1/file1 [A-Za-z_0-9] foo bar\nreplace ./dir1/file2 [A-Za-z_0-9] foo bar\nreplace ./dir1/file3 [A-Za-z_0-9] foo bar\nreplace ./dir1/file4 [A-Za-z_0-9] foo bar\nreplace ./dir1/file5 [A-Za-z_0-9] foo bar\nreplace ./dir2/file1 [A-Za-z_0-9] foo bar\nreplace ./dir2/file2 [A-Za-z_0-9] foo bar\nreplace ./dir2/file3 [A-Za-z_0-9] foo bar\nreplace ./dir2/file4 [A-Za-z_0-9] foo bar\nreplace ./dir2/file5 [A-Za-z_0-9] foo bar\nreplace ./dir3/file1 [A-Za-z_0-9] foo bar\nreplace ./dir3/file2 [A-Za-z_0-9] foo bar\nreplace ./dir3/file3 [A-Za-z_0-9] foo bar\nreplace ./dir3/file4 [A-Za-z_0-9] foo bar\nreplace ./dir3/file5 [A-Za-z_0-9] foo bar\n}\n\nI see two possible complementary ways to address this:\n1) allow something akin to the above form in git free-form comments as a \n*technical* solution, while leaving it up to the individual repository \nowner whether to accept such patches on aesthetic grounds.\n2) explore adding a different format to darcs that would allow a files \naffected to be represented more compactly.\n\n\nI suspect that any use of wildcards in a new format would be impossible \nfor darcs since it wouldn't allow darcs to construct dependencies, \nthough I'll leave it to david to respond to that.\n\nAt a minimum, something like:\nreplace ./dir1/[file1|file2|file3|file4|file5] [A-Za-z_0-9] foo bar\nreplace ./dir2/[file1|file2|file3|file4|file5] [A-Za-z_0-9] foo bar\nreplace ./dir3/[file1|file2|file3|file4|file5] [A-Za-z_0-9] foo bar\nshould be pretty feasible.\n\nI don't believe, however, that it would ever be 100% reliable to try to \nlook at a one line replace description and combine it with the actual \nchanges and end up with a correct darcs replace patch.\n\n-Tupshin\n"},{"id":"832","messageId":"Pine.LNX.4.58.0504190943060.19286@ppc970.osdl.org","threadId":"102","inReplyTo":"426532D5.3040306@tupshin.com","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-19T16:49:12Z","receivedAt":"2005-04-19T16:49:12Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 19 Apr 2005, Tupshin Harper wrote:\n> \n> I suspect that any use of wildcards in a new format would be impossible \n> for darcs since it wouldn't allow darcs to construct dependencies, \n> though I'll leave it to david to respond to that.\n\nNote that git _does_ very efficiently (and I mean _very_) expose the \nchanged files.\n\nSo if this kind of darcs patch is always the same pattern just repeated\nover <n> files, then you really don't need to even list the files at all.  \nGit gives you a very efficient file listing by just doing a \"diff-tree\"  \n(which does not diff the _contents_ - it really just gives you a pretty\nmuch zero-cost \"which files changed\" listing).\n\nSo that combination would be 100% reliable _if_ you always split up darcs \npatches to \"common elements\". \n\nAnd note that there does not have to be a 1:1 relationship between a git\ncommit and a darcs patch. For example, say that you have a darcs patch\nthat does a combination of \"change token x to token y in 100 files\" and\n\"rename file a into b\". I don't know if you do those kind of \"combination \npatches\" at all, but if you do, why not just split them up into two? That \nway the list of files changed _does_ 100% determine the list of files for \nthe token exchange.\n\n\t\tLinus\n"},{"id":"923","messageId":"1113960166.29444.102.camel@orca.madrabbit.org","threadId":"102","inReplyTo":"7i7jizyy4i.fsf@lanthane.pps.jussieu.fr","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"Ray Lee","fromEmail":"ray@madrabbit.org","sentAt":"2005-04-20T01:22:45Z","receivedAt":"2005-04-20T01:22:45Z","isPatch":false,"sender":{"key":"ray@madrabbit.org","avatar":null},"body":"On Tue, 2005-04-19 at 10:22 +0200, Juliusz Chroboczek wrote:\n> > > Aye, that will require some metadata on the git side (the hack,\n> > > suggested by Linus, of using git hashes to notice moves won't work).\n> \n> > So, why won't it work?\n> \n> Because two files can legitimately have identical contents without\n> being ``the same'' file from the VC system's point of view.\n> \n> In other words, two files may happen to have the same contents but\n> have distinct histories.\n\nEh, let's not talk using integral/summation view across all the patches\nthat ever could have come in against the file. We're hamstringing\nourselves if we do that, and it's not what darcs does. darcs looks at a\ndifferential view of the changes, and for a mv, it looks at it when it\nhappens.\n\ndarcs does a \"darcs mv\" to commit a \"file move patch\" to whatever\nlogging or patch repository it keeps below the surface.\n\nThe equivalent in git would be to have a given tree, move a file via\nbash's mv, and then checkpoint a new tree. (I'm sure there's details in\nthere, but that's plumbing, and what we have Petr for.)\n\nA differential comparison of the two trees shows no content changed, but\na file label was modified. Ergo, a rename occurred.\n\nQED.\n\n~r.\n\n"},{"id":"969","messageId":"20050420111446.GE29945@abridgegame.org","threadId":"102","inReplyTo":"Pine.LNX.4.58.0504190943060.19286@ppc970.osdl.org","subject":"Re: Darcs and git: plan of action","fromName":"David Roundy","fromEmail":"droundy@abridgegame.org","sentAt":"2005-04-20T11:14:52Z","receivedAt":"2005-04-20T11:14:52Z","isPatch":false,"sender":{"key":"droundy@abridgegame.org","avatar":"https://gravatar.com/avatar/e8bcfd76f63303732bdfcdba6fc8ac6ccdff8f5a224a25ffaca24bd8a4c4571f?d=mp&s=160"},"body":"On Tue, Apr 19, 2005 at 09:49:12AM -0700, Linus Torvalds wrote:\n> On Tue, 19 Apr 2005, Tupshin Harper wrote:\n> > I suspect that any use of wildcards in a new format would be impossible\n> > for darcs since it wouldn't allow darcs to construct dependencies,\n> > though I'll leave it to david to respond to that.\n> \n> Note that git _does_ very efficiently (and I mean _very_) expose the \n> changed files.\n> \n> So if this kind of darcs patch is always the same pattern just repeated\n> over <n> files, then you really don't need to even list the files at all.\n> Git gives you a very efficient file listing by just doing a \"diff-tree\"\n> (which does not diff the _contents_ - it really just gives you a pretty\n> much zero-cost \"which files changed\" listing).\n\nThe catch is that it's possible to have a darcs patch that doesn't change\nany files, or that affects files without changing them.  If I rename\nfunction foo to bar, I might want to do\n\ndarcs replace foo bar *.c\n\nwhich would issue a replace on all files, which means that when this patch\nis merged with any patches that add occurrences of foo in a file, that will\nget modified to a bar, regardless of whether there was previously an\noccurrence of foo in that file.\n\nI think we might (when working with git--it would be problematic within\ndarcs straight) be able to work out some sort of a wildcard replace\nscheme, so it could be something like\n\nreplace foo bar in: mm/*.c\n\nThe regexp bit could be left out, if we restrict the definition of \"tokens\"\nin token replaces--which probably isn't a troublesome limitation.  By\ndefault darcs uses two tokenizing schemes, one which allows \".\" in tokens\n(usually relevant in Makefiles), and one which doesn't, and basically\nmatches C identifiers.  We could allow for both of these if we had a second\noption:\n\nreplace filename foo.h bar.h in: mm/*.c\n\nWe'd just need to expand the wildcards when translating from the git\nrepository into darcs patches.\n\n> So that combination would be 100% reliable _if_ you always split up darcs \n> patches to \"common elements\". \n> \n> And note that there does not have to be a 1:1 relationship between a git\n> commit and a darcs patch. For example, say that you have a darcs patch\n> that does a combination of \"change token x to token y in 100 files\" and\n> \"rename file a into b\". I don't know if you do those kind of \"combination \n> patches\" at all, but if you do, why not just split them up into two? That \n> way the list of files changed _does_ 100% determine the list of files for \n> the token exchange.\n\nWe do allow multiple sorts of changes (in darcs terminology, multiple\n\"primitive patches\") in a single patch.\n\nOne *could* have multiple git commits for a single darcs patch, but that\nseems ugly and I'd rather avoid it.  In my view, revision control system is\nmore about communication than history (which is why by default, darcs\ndoesn't \"do\" history), and grouping changes together is how we express\nwhich changes \"go together\".  Of course, we could still have a grouping at\na higher level, so that a single \"changeset\" could consist of multiple git\ncommits (for example by recognizing that identical commit logs mean that\nit's a single change), but that adds a layer of complexity that I'd like to\navoid if possible.\n-- \nDavid Roundy\nhttp://www.darcs.net\n"},{"id":"970","messageId":"20050420111847.GF29945@abridgegame.org","threadId":"102","inReplyTo":"20050419122518.GD12757@pasky.ji.cz","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"David Roundy","fromEmail":"droundy@abridgegame.org","sentAt":"2005-04-20T11:18:51Z","receivedAt":"2005-04-20T11:18:51Z","isPatch":false,"sender":{"key":"droundy@abridgegame.org","avatar":"https://gravatar.com/avatar/e8bcfd76f63303732bdfcdba6fc8ac6ccdff8f5a224a25ffaca24bd8a4c4571f?d=mp&s=160"},"body":"On Tue, Apr 19, 2005 at 02:25:18PM +0200, Petr Baudis wrote:\n> Dear diary, on Tue, Apr 19, 2005 at 02:20:55PM CEST, I got a letter\n> where Juliusz Chroboczek <Juliusz.Chroboczek@pps.jussieu.fr> told me that...\n> > > The problem is that there is no sequence of alien versions that one\n> > > can differentiate.  Git has a branched history, with each version\n> > > that follows a merge having multiple parents.\n> > \n> > Yep.  I've just realised that this morning.  Is there some notion of\n> > ``primary parent'' as in Arch?  Can a changeset have 0 parents?\n> \n> Yes, the root commit. Usually, there is only one, but there may be\n> multiple of them theoretically.\n\nIncidentally (and completely off-topic for this thread), wouldn't there be\na sha1 tree hash corresponding to a completely empty directory, and\ncouldn't one use that as the parent for the root? Would there be any reason\nto do so? Just a silly thought...\n-- \nDavid Roundy\nhttp://www.darcs.net\n"},{"id":"971","messageId":"20050420112942.GG29945@abridgegame.org","threadId":"102","inReplyTo":"7i4qe3x8ig.fsf@lanthane.pps.jussieu.fr","subject":"Re: [darcs-devel] Darcs and git: plan of action","fromName":"David Roundy","fromEmail":"droundy@abridgegame.org","sentAt":"2005-04-20T11:29:43Z","receivedAt":"2005-04-20T11:29:43Z","isPatch":false,"sender":{"key":"droundy@abridgegame.org","avatar":"https://gravatar.com/avatar/e8bcfd76f63303732bdfcdba6fc8ac6ccdff8f5a224a25ffaca24bd8a4c4571f?d=mp&s=160"},"body":"On Tue, Apr 19, 2005 at 02:20:55PM +0200, Juliusz Chroboczek wrote:\n> [Removing Linus from CC, keeping the Git list -- or should we remove it?]\n\nI think leaving much of this on git would be appropriate, since there are\nissues of how to relate to git that should be relevant.\n\n> > If we do it right (automatically tagging like crazy people), darcs\n> > users between themselves can cherry-pick all they like, without\n> > introducing inconsistencies or losing interoperability with git.\n> \n> You've lost me here.  How can you cherry-pick if every tag depends on\n> the preceding patches?  Or are you thinking of pulling just the patch\n> and not the tag -- in that case, what happens when you push to git a\n> Darcs patch that depends on a patch that originated with git?\n\nYes, I'm thinking of pulling patches from one darcs repo to another.  If we\ncherry-pick in this way, we need to create a \"git-tag\" for each patch that\nwe pull without its associated tag.  To git, this would look like two\nseparate changes that have the same commit log, except that they have\ndifferent parents and different commiters and commit dates.\n\nI don't think this will be a problem for git, and since darcs will\nrecognize the two patches as the identical darcs patch (we'll need to put\nsomewhere in the git commit log a magic word indicating that this patch\noriginated in darcs), there won't be a problem for darcs either.\n\nIn case I haven't been clear (which seems likely), the scenario is that\ndarcs user 1 makes the following changes to his darcs version of a\ngit-based repository:\n\nchanges in 1: A -> B\ntags in 1:    A1   B1\n\nDarcs user 2 wants B, but not A, and didn't do any development:\n\nchanges in 2: B\ntags in 2:    B2\n\nUser 2 pushes to git, and now git has (where P is the parent of both of the\nabove):\n\ngit:\nP -> B/B2  (where B/B2 is the commit log with B2 as \"committer info\" and B\n            as the \"author info and long comment)\n\nUser 1 pushes (everything) to git and merges the two (patch M, which has\ntwo parents, B1 and B2:\n\ngit:\n\n   ->B/B2---------\n  /               \\\nP--> A/A1 -> B/B1---> M\n\nIt's a little lame, and if user 2 doesn't do any real work, the git-using\nperson might be annoyed, but I think it's doable.\n-- \nDavid Roundy\nhttp://www.darcs.net\n"}]}