{"thread":{"id":"29821","subject":"[RFC] \"Remote helper for Subversion\" project","startedAt":"2012-03-03T12:27:26Z","lastAt":"2012-03-27T03:58:52Z","messageCount":22,"participants":["David Barr","Jonathan Nieder","Andrew Sayers","Stephen Bash","Nathan Gray","Sam Vilain","Phil Hord","Ramkumar Ramachandra"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"185984","messageId":"1330777646-28381-1-git-send-email-davidbarr@google.com","threadId":"29821","inReplyTo":null,"subject":"[RFC] \"Remote helper for Subversion\" project","fromName":"David Barr","fromEmail":"davidbarr@google.com","sentAt":"2012-03-03T12:27:26Z","receivedAt":"2012-03-03T12:27:26Z","isPatch":false,"sender":{"key":"davidbarr@google.com","avatar":"https://avatars.githubusercontent.com/u/220594?v=4"},"body":"---\n SoC-2012-Ideas.md |   26 ++++++++++++++++++++++++++\n 1 files changed, 26 insertions(+), 0 deletions(-)\n\n This is simply the direct translation of last year's project idea.\n This project make significant incremental progess each year.\n I'm seeking feedback from all involved on setting the direction.\n\n --\n David Barr\n\ndiff --git a/SoC-2012-Ideas.md b/SoC-2012-Ideas.md\nindex 5e83342..4c2ab05 100644\n--- a/SoC-2012-Ideas.md\n+++ b/SoC-2012-Ideas.md\n@@ -182,3 +182,29 @@ this project.\n \n Proposed by: Thomas Rast  \n Possible mentor(s): Thomas Rast\n+\n+Remote helper for Subversion\n+------------------------------------\n+\n+Write a remote helper for Subversion. While a lot of the underlying\n+infrastructure work was completed last year, the remote helper itself\n+is essentially incomplete. Major work includes:\n+\n+* Understanding revision mapping and building a revision-commit mapper.\n+\n+* Working through transport and fast-import related plumbing, changing\n+  whatever is necessary.\n+\n+* Getting an Git-to-SVN converter merged.\n+\n+* Building the remote helper itself.\n+\n+Goal: Build a full-featured bi-directional `git-remote-svn` and get it\n+      merged into upstream Git.  \n+Language: C  \n+See: [A note on SVN history][SVN history], [svnrdump][].  \n+Proposed by: David Barr  \n+Possible mentors: Jonathan Nieder, Sverre Rabbelier, David Barr\n+\n+[SVN history]: http://article.gmane.org/gmane.comp.version-control.git/150007\n+[svnrdump]: http://svn.apache.org/repos/asf/subversion/trunk/subversion/svnrdump\n-- \n1.7.9\n"},{"id":"185985","messageId":"CAFfmPPMPDCKjAmZ85Cj1cdT2yAUykm9sV6a66zXeFRmYfrmtjg@mail.gmail.com","threadId":"29821","inReplyTo":"1330777646-28381-1-git-send-email-davidbarr@google.com","subject":"Re: [RFC] \"Remote helper for Subversion\" project","fromName":"David Barr","fromEmail":"davidbarr@google.com","sentAt":"2012-03-03T12:41:31Z","receivedAt":"2012-03-03T12:41:31Z","isPatch":false,"sender":{"key":"davidbarr@google.com","avatar":"https://avatars.githubusercontent.com/u/220594?v=4"},"body":"On Sat, Mar 3, 2012 at 11:27 PM, David Barr <davidbarr@google.com> wrote:\n> ---\n>  SoC-2012-Ideas.md |   26 ++++++++++++++++++++++++++\n>  1 files changed, 26 insertions(+), 0 deletions(-)\n>\n>  This is simply the direct translation of last year's project idea.\n>  This project make significant incremental progess each year.\n>  I'm seeking feedback from all involved on setting the direction.\n>\n>  --\n>  David Barr\n>\n> diff --git a/SoC-2012-Ideas.md b/SoC-2012-Ideas.md\n> index 5e83342..4c2ab05 100644\n> --- a/SoC-2012-Ideas.md\n> +++ b/SoC-2012-Ideas.md\n> @@ -182,3 +182,29 @@ this project.\n>\n>  Proposed by: Thomas Rast\n>  Possible mentor(s): Thomas Rast\n> +\n> +Remote helper for Subversion\n> +------------------------------------\n> +\n> +Write a remote helper for Subversion. While a lot of the underlying\n> +infrastructure work was completed last year, the remote helper itself\n> +is essentially incomplete. Major work includes:\n> +\n> +* Understanding revision mapping and building a revision-commit mapper.\n> +\n> +* Working through transport and fast-import related plumbing, changing\n> +  whatever is necessary.\n> +\n> +* Getting an Git-to-SVN converter merged.\n> +\n> +* Building the remote helper itself.\n> +\n> +Goal: Build a full-featured bi-directional `git-remote-svn` and get it\n> +      merged into upstream Git.\n> +Language: C\n> +See: [A note on SVN history][SVN history], [svnrdump][].\n> +Proposed by: David Barr\n> +Possible mentors: Jonathan Nieder, Sverre Rabbelier, David Barr\n> +\n> +[SVN history]: http://article.gmane.org/gmane.comp.version-control.git/150007\n> +[svnrdump]: http://svn.apache.org/repos/asf/subversion/trunk/subversion/svnrdump\n> --\n> 1.7.9\n>\n\n+cc: Ramkumar Ramachandra, Sam Vilain, Stephen Bash\nI wasn't even close to \"all involved.\n\n--\nDavid Barr\n"},{"id":"186024","messageId":"20120304075424.GI14725@burratino","threadId":"29821","inReplyTo":"CAFfmPPMPDCKjAmZ85Cj1cdT2yAUykm9sV6a66zXeFRmYfrmtjg@mail.gmail.com","subject":"Re: [RFC] \"Remote helper for Subversion\" project","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2012-03-04T07:54:25Z","receivedAt":"2012-03-04T07:54:25Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"David Barr wrote:\n> On Sat, Mar 3, 2012 at 11:27 PM, David Barr <davidbarr@google.com> wrote:\n\n>> --- a/SoC-2012-Ideas.md\n>> +++ b/SoC-2012-Ideas.md\n>> @@ -182,3 +182,29 @@ this project.\n>>\n>>  Proposed by: Thomas Rast\n>>  Possible mentor(s): Thomas Rast\n>> +\n>> +Remote helper for Subversion\n>> +------------------------------------\n>> +\n>> +Write a remote helper for Subversion. While a lot of the underlying\n>> +infrastructure work was completed last year, the remote helper itself\n>> +is essentially incomplete. Major work includes:\n\nBy the way, didn't we have a remote-svn prototype?  I'm happy to merge\nany old hacky thing for staging in contrib/svn-fe, as long as it is\nnot documented in a misleading way.\n\n(More generally, if anyone wants to resend useful svn-fe patches, that\nwill help a lot.)\n\n>> +\n>> +* Understanding revision mapping and building a revision-commit mapper.\n\nDoes this mean creating commit notes to record which subversion rev\ncorresponds to each commit, and marks or lightweight tags going the\nother way?\n\n>> +\n>> +* Working through transport and fast-import related plumbing, changing\n>> +  whatever is necessary.\n\nI think Dmitry and Sverre took care of most of this.\n\n>> +\n>> +* Getting an Git-to-SVN converter merged.\n\nProbably could fill a summer in itself.  In previous starts I think\nthere was some complexity creep. :/\n\n http://thread.gmane.org/gmane.comp.version-control.git/170290\n http://thread.gmane.org/gmane.comp.version-control.git/170551\n\n>> +\n>> +* Building the remote helper itself.\n>> +\n>> +Goal: Build a full-featured bi-directional `git-remote-svn` and get it\n>> +      merged into upstream Git.\n\nSure would be neat. ;-)  Another nice piece to build would be branch\ntracking / follow_parent heuristics.\n\nThanks,\nJonathan\n"},{"id":"186026","messageId":"CAFfmPPPs0FRbT-i+ZwBLNSca330Eo7thjNxDt3hJf0yUATthtQ@mail.gmail.com","threadId":"29821","inReplyTo":"20120304075424.GI14725@burratino","subject":"Re: [RFC] \"Remote helper for Subversion\" project","fromName":"David Barr","fromEmail":"davidbarr@google.com","sentAt":"2012-03-04T10:37:28Z","receivedAt":"2012-03-04T10:37:28Z","isPatch":false,"sender":{"key":"davidbarr@google.com","avatar":"https://avatars.githubusercontent.com/u/220594?v=4"},"body":"On Sun, Mar 4, 2012 at 6:54 PM, Jonathan Nieder <jrnieder@gmail.com> wrote:\n> David Barr wrote:\n>> On Sat, Mar 3, 2012 at 11:27 PM, David Barr <davidbarr@google.com> wrote:\n>\n>>> --- a/SoC-2012-Ideas.md\n>>> +++ b/SoC-2012-Ideas.md\n>>> @@ -182,3 +182,29 @@ this project.\n>>>\n>>>  Proposed by: Thomas Rast\n>>>  Possible mentor(s): Thomas Rast\n>>> +\n>>> +Remote helper for Subversion\n>>> +------------------------------------\n>>> +\n>>> +Write a remote helper for Subversion. While a lot of the underlying\n>>> +infrastructure work was completed last year, the remote helper itself\n>>> +is essentially incomplete. Major work includes:\n>\n> By the way, didn't we have a remote-svn prototype?  I'm happy to merge\n> any old hacky thing for staging in contrib/svn-fe, as long as it is\n> not documented in a misleading way.\n>\n> (More generally, if anyone wants to resend useful svn-fe patches, that\n> will help a lot.)\n\nFound at former SoC2011Projects wiki page:\n(http://git.wiki.kernel.org/articles/s/o/c/SoC2011Projects_b1f9.html#Remote_helper_for_Subversion_and_git-svn)\n[vcs-svn, svn-fe: add a couple of\noptions](http://thread.gmane.org/gmane.comp.version-control.git/176578)\n[remote-svn-alpha\nupdates](http://thread.gmane.org/gmane.comp.version-control.git/176617)\n\nThe introduction should be rephrased to include Dmitry's progression.\n\n>>> +* Understanding revision mapping and building a revision-commit mapper.\n>\n> Does this mean creating commit notes to record which subversion rev\n> corresponds to each commit, and marks or lightweight tags going the\n> other way?\n\nYes. I think once again, Dmitry produced a good prototype for this component.\nHowever, I think it also potentially incorporates git-svn style\nslicing of history.\nThat's a significant chunk of work.\n\n>>> +* Working through transport and fast-import related plumbing, changing\n>>> +  whatever is necessary.\n>\n> I think Dmitry and Sverre took care of most of this.\n\nDitto.\n\n>>> +* Getting an Git-to-SVN converter merged.\n>\n> Probably could fill a summer in itself.  In previous starts I think\n> there was some complexity creep. :/\n>\n>  http://thread.gmane.org/gmane.comp.version-control.git/170290\n>  http://thread.gmane.org/gmane.comp.version-control.git/170551\n\nThis is my preferred focus, and is a sufficient project in its own right.\n\n>>> +* Building the remote helper itself.\n>>> +\n>>> +Goal: Build a full-featured bi-directional `git-remote-svn` and get it\n>>> +      merged into upstream Git.\n>\n> Sure would be neat. ;-)  Another nice piece to build would be branch\n> tracking / follow_parent heuristics.\n\nAs noted earlier, the remote helper itself is now half-complete.\nI do think the immediate goal should be bi-direction.\nThe remainder is porting git-svn logic to the new helper.\nHowever, it would be interesting to see what's missing with respect to porting\n\n--\nDavid Barr\n"},{"id":"186040","messageId":"4F536FE9.1050000@pileofstuff.org","threadId":"29821","inReplyTo":"CAFfmPPPs0FRbT-i+ZwBLNSca330Eo7thjNxDt3hJf0yUATthtQ@mail.gmail.com","subject":"Re: [RFC] \"Remote helper for Subversion\" project","fromName":"Andrew Sayers","fromEmail":"andrew-git@pileofstuff.org","sentAt":"2012-03-04T13:36:41Z","receivedAt":"2012-03-04T13:36:41Z","isPatch":false,"sender":{"key":"andrew-git@pileofstuff.org","avatar":null},"body":"Hi guys,\n\nI made a few little git contributions a couple of years back, before\nbeing swallowed up by my old job.  I've become quite interested in\nsvn->git translation since starting a new job, but I'm delurking earlier\nthan planned so please bear with me as all I have right now is an\ninteresting proof of concept and the possibility of some free time some\nday.  If things go well, I hope to help out with the\nbranching-and-merging part of the problem, so I'll describe what I've\ndone and how it might affect SoC projects.  Apologies in advance for\nhijacking the thread :)\n\nI mentioned proof-of-concept code for the work I've done - I'd like to\nget confirmation from my new employer before putting it all online, so\nhopefully I can make it available in a few days.\n\nWhile researching the problem, I found Stephen Bash's original\nproposal[1] and snerp-vortex[2] quite inspiring, but wasn't able to find\nany details on SoC-related work in the branching-and-merging department\n- hopefully the following isn't just a retread of ideas developed since\nthen.  I've concentrated on importing from SVN so far, but have kept an\neye on update and half an eye on bi-direction in the hopes of being\nuseful there some day.\n\nIt seems to me the \"svn export\" and \"git import\" steps make most sense\nas two unrelated projects.  Snerp-vortex and Stephen's scripts both cut\nthe history import problem at that point, as do svn-fe/git-fast-import\nwith code import.  Exporting SVN history is a messy and sometimes\nproject-specific job, so allowing a project to concentrate on that part\nmakes it possible for SVN experts to use all their skills without having\nto learn git plumbing before they make their first commit (much respect\nto Stephen for managing that feat BTW).\n\n\nI've written a proof-of-concept history converter that can be split into\nthree parts: a format for describing SVN history; a large, often messy\nPerl program that writes files in that format; and a small Perl program\nthat reads the format and translates it into git.  With hindsight, Perl\nis the right language for the SVN exporter, but the git importer would\nhave been better written in C.\n\n\nPersonally, I think SVN export will always need a strong manual\ncomponent to get the best results, so I've put quite a bit of work into\ndesigning a good SVN history format.  Like git-fast-import, it's an\nASCII format designed both for human and machine consumption:\n\nIn r1, create branch \"trunk\"\nIn r10, create branch \"branches/foo\" as \"foo\" from \"trunk\" r9\nIn r12, create tag \"tags/version1\" as \"version1\" from \"trunk\" r11\nIn r12, deactivate \"tags/version1\"\n# blank lines and lines beginning with a '#' are ignored\nIn r15, merge \"branches/foo\" r14 into \"trunk\"\nIn r15, delete \"branches/foo\"\n\nThe above has been designed as an abstract representation of SVN\nhistory, with as little git-specific content as possible.  This turns\nout to make the problem a bit easier to think about, a bit easier to\nimplement, and potentially useful to other DVCSs some day.  I wanted to\nenable svn-merge2git[3]-style extraction of merge info from logs, but\nfound that even a message as clear as \"merge r123 from trunk\" could mean\n\"cherry-pick revision 123\", \"merge everything up to revision 123\",\n\"merge everything up to the last revision that touched trunk before\n123\", \"merge everything up to 132\", or any number of other things.  As\nsuch, the format is designed to allow copious comments and ease of\nbulk-editing with regexps and a powerful editor.\n\nThe format feels relatively complete to me, and in any case not a good\ncandidate for an SoC project because it's all about experience and\nnothing to do with raw talent.\n\n\nOnce the format is defined, git import is fairly straightforward.\nProof-of-concept code to follow, but it's really just a wrapper around\ngit-commit-tree, git-mktag etc.  I wrote this in Perl thinking it would\nrelate somehow to git-svn, but eventually realised it didn't and that a\nfew hundred calls to (plumbing) processes per second isn't so good for\nperformance.  The only interesting part of the problem is how to tackle\nSVN tags.  I went for an ambitious approach, making normal tags where\npossible and downgrading them to lightweight tags when necessary.  This\ndoes involve managing something that is effectively a branch in\nrefs/tags/, but what else is an SVN tag but a branch in the wrong namespace?\n\nI'm afraid I don't know enough about SoC to say whether rewriting git\nimport in C would make a good project - on the one hand, it's a smallish\nbit of work that would make a student learn a lot of good git code, on\nthe other hand it would mostly just involve fixing up a bit of ugly Perl\nwith little chance to show any creativity.\n\n\nSVN export is much more complicated, and has taken most of my time.\nAlthough there's no way to tell automatically which directories are\nbranches, I realised you can detect trunks automatically about 90% of\nthe time by looking for where files/directories are first created, and\ncan detect non-trunk branches at least 99% of the time once you know the\ntrunks.  Trunk detection is normally just a convenience, but can be a\nlifesaver when importing from a sufficiently messy repository.  There's\na lot you can do with merge detection, but between svn:mergeinfo,\nsvnmerge.py[4], svk[5] and log messages I realise I've only scratched\nthe surface.  In many ways, this is a classic Perl problem - take a\nbunch of messy textual input, string it all together and make some neat\ntextual output.  The code I've got right now seems fine for anything up\nto about 20,000 commits, but eats way too much memory for huge repos.\nDepending on how much time I get, I might have tackled that problem by\nthe time I put this code online.\n\nAssuming I have enough time to work on SVN export, I would be loathed to\nput it on anyone else in the near future.  I could see endless\noptimisations being spun off once the code is mature, but right now it's\nan idiosyncratic collection of untested mostly-working bits that need\nserious attention before I'd be happy warping a young mind on it.\n\n\nI hope this information dump explains where I am and how I can help.  As\nI say, I can't yet promise how much (if any) time I'll get to work on\nthis in future, but I hope to help with the history translation process\nif circumstances allow.  More importantly to this thread, I hope this\ngives some ideas about SoC projects and places (not) to put attention.\n\n\t- Andrew Sayers\n\n[1] http://comments.gmane.org/gmane.comp.version-control.git/158940\n[2] https://github.com/rcaputo/snerp-vortex\n[3] http://repo.or.cz/w/svn-merge2git.git\n[4] http://www.orcaware.com/svn/wiki/Svnmerge.py\n[5] http://search.cpan.org/dist/SVK/\n"},{"id":"186041","messageId":"20120304162322.GB17923@burratino","threadId":"29821","inReplyTo":"CAFfmPPPs0FRbT-i+ZwBLNSca330Eo7thjNxDt3hJf0yUATthtQ@mail.gmail.com","subject":"Re: [RFC] \"Remote helper for Subversion\" project","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2012-03-04T16:23:22Z","receivedAt":"2012-03-04T16:23:22Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"David Barr wrote:\n> On Sun, Mar 4, 2012 at 6:54 PM, Jonathan Nieder <jrnieder@gmail.com> wrote:\n\n>> (More generally, if anyone wants to resend useful svn-fe patches, that\n>> will help a lot.)\n>\n> Found at former SoC2011Projects wiki page:\n> (http://git.wiki.kernel.org/articles/s/o/c/SoC2011Projects_b1f9.html#Remote_helper_for_Subversion_and_git-svn)\n> [vcs-svn, svn-fe: add a couple of\n> options](http://thread.gmane.org/gmane.comp.version-control.git/176578)\n> [remote-svn-alpha\n> updates](http://thread.gmane.org/gmane.comp.version-control.git/176617)\n\nDo you mean these are patches that should be applied?  New emails\ncontaining a git url or, even better, the actual patch are best, since\nit means I can be sure I am looking at the latest or at least the\nintended version of the change.\n\n[...]\n> However, I think it also potentially incorporates git-svn style\n> slicing of history.\n\nDo I understand correctly that you mean paying attention to copy-from\ninformation, like \"svn log\" does?  (For example, making cloning\n\n\tsvn::http://svn.example.com/project/branches/feature\n\nwhen branches/feature was originally copied from trunk involve\ngrabbing \"http://svn.example.com/project/trunk\" in early revs?)\n\n[...]\n> The remainder is porting git-svn logic to the new helper.\n> However, it would be interesting to see what's missing with respect to porting\n\nWhile git-svn can be useful for inspiration when wondering \"how could\nI possibly solve such-and-such problem\", I'm not sure feature-parity\nwith git-svn is too important.  After all, people needing git-svn\nfeatures can still use git-svn.\n\nI say this since git-svn has lots of features we are missing:\nnot discarding unhandled properties (important), shared history with\nmultiple branches, author mapping, fetching and pushing svn:mergeinfo\ninformation, partial clone via a path-ignore regex, choice of\ntimezone, filename reencoding, manual svn:ignore-to-gitignore\nconversion, svn-compatible \"log\" and \"blame\" output, custom git<->svn\nbranchname mappings, and so on.  The ability to track one branch,\nincluding push support, with a linear history would be exciting\nalready and doesn't require all that.\n\nCheers,\nJonathan\n"},{"id":"186120","messageId":"3c2ab05e-b2af-4df4-bca6-ff5512b0c73e@mail","threadId":"29821","inReplyTo":"4F536FE9.1050000@pileofstuff.org","subject":"Approaches to SVN to Git conversion (was: Re: [RFC] \"Remote helper for Subversion\" project)","fromName":"Stephen Bash","fromEmail":"bash@genarts.com","sentAt":"2012-03-05T15:27:27Z","receivedAt":"2012-03-05T15:27:27Z","isPatch":false,"sender":{"key":"bash@genarts.com","avatar":null},"body":"All-\n\nThis turned out to be longer than I intended, but actually summarizes some of my modern thoughts on SVN to Git conversion (as always, I'm more interested in one time migration than bidirectional operation, so read with a grain of salt).\n\nMore inline...\n\n----- Original Message -----\n> From: \"Andrew Sayers\" <andrew-git@pileofstuff.org>\n> Sent: Sunday, March 4, 2012 8:36:41 AM\n> Subject: Re: [RFC] \"Remote helper for Subversion\" project\n> \n> ... snip ...\n> \n> While researching the problem, I found Stephen Bash's original\n> proposal[1] and snerp-vortex[2] quite inspiring, but wasn't able to\n> find any details on SoC-related work in the branching-and-merging\n> department - hopefully the following isn't just a retread of ideas\n> developed since then.  I've concentrated on importing from SVN so far,\n> but have kept an eye on update and half an eye on bi-direction in the\n> hopes of being useful there some day.\n> \n> It seems to me the \"svn export\" and \"git import\" steps make most sense\n> as two unrelated projects.  Snerp-vortex and Stephen's scripts both\n> cut the history import problem at that point, as do\n> svn-fe/git-fast-import with code import.  Exporting SVN history is a\n> messy and sometimes project-specific job, so allowing a project to\n> concentrate on that part makes it possible for SVN experts to use all\n> their skills without having to learn git plumbing before they make\n> their first commit (much respect to Stephen for managing that feat\n> BTW).\n\nAfter many a long conversation with Ram and Jonathan (and others), I'm actually going the other direction.  My current thinking (and this is very much open for discussion) is that as long as the SVN properties are available (especially the copyfrom information) Git has just as much information (if not more) to reconstruct the SVN history as SVN does.  (And going through our messy history I haven't found any counterpoint to this yet)\n\n> I've written a proof-of-concept history converter that can be split\n> into three parts: a format for describing SVN history; a large, often\n> messy Perl program that writes files in that format; and a small Perl\n> program that reads the format and translates it into git.  With\n> hindsight, Perl is the right language for the SVN exporter, but the\n> git importer would have been better written in C.\n> \n> Personally, I think SVN export will always need a strong manual\n> component to get the best results, so I've put quite a bit of work\n> into designing a good SVN history format.  Like git-fast-import, it's\n> an ASCII format designed both for human and machine consumption...\n\nFirst, I'm very impressed that you managed to get a language like this up and working.  It could prove very useful going forward.  On the flip side, from my experiments over the last year I've actually been leaning toward a solution that is more implicit than explicit.  Taking git-svn as a model, I've been trying to define a mapping system (in Perl):\n\n  my %branch_spec = { '/trunk/projname' => 'master',\n                      '/branches/*/projname' => '/refs/heads/*' };\n  my %tag_spec = { '/tags/*/projname' => '/refs/tags/*' };\n\n(See [1] for notes on our SVN structure.  In a std-layout-style repo it would be:\n\n  my %branch_spec = { '/projname/trunk' => 'master',\n                      '/projname/branches/*' => '/refs/heads/*' };\n  my %tag_spec = { '/projname/tags/*' => '/refs/tags/*' };)\n\nNow I know this simple mapping will fail as I get further in our history -- in particular we have one branch that came from:\n\n  svn cp $SVN_REPO/trunk/ $SVN_REPO/foo  # OOPS! not in branches!\n  svn mv $SVN_REPO/foo $SVN_REPO/branches/foo\n\n>From an automation perspective, I expect the first svn operation to produce an error saying \"Possible branch created at /foo from known branch trunk (master), but doesn't match any known branch spec\", while the second (if \"continue-on-error\" is turned on) would error \"Branch /branches/foo created from unknown branch located at /foo\".  It's then up to the user to modify the branch map to something that accounts for this behavior:\n\n  my %branch_spec = { '/trunk/projname' => 'master',\n                      '/branches/*/projname' => '/refs/heads/*',\n                      '/foo' => '/refs/heads/foo' };\n  my %tag_spec = { '/tags/*/projname' => '/refs/tags/*' };\n\nSo in this case I'm making an explicit branch mapping, but the use of glob-style syntax allows the user to catch larger classes of branches if desired (I'll also note that depending on the implementation /foo may need to map to /refs/heads/bad-foo so that /branches/foo can map to /refs/heads/foo, but my intention has been to squash empty commits, so it's possible the name conflict is a non-issue since content didn't change).  I also know that we have some copy operations that are just weird, so it might be helpful to have an ignore mechanism that tells the system to ignore copies into/out-of certain SVN paths.\n\n> Once the format is defined, git import is fairly straightforward.\n> Proof-of-concept code to follow, but it's really just a wrapper around\n> git-commit-tree, git-mktag etc.  I wrote this in Perl thinking it\n> would relate somehow to git-svn, but eventually realised it didn't and\n> that a few hundred calls to (plumbing) processes per second isn't so\n> good for performance.  The only interesting part of the problem is how\n> to tackle SVN tags.  I went for an ambitious approach, making normal\n> tags where possible and downgrading them to lightweight tags when\n> necessary.  This does involve managing something that is effectively a\n> branch in refs/tags/, but what else is an SVN tag but a branch in the\n> wrong namespace?\n\nI don't understand how \"normal\" and \"lightweight\" apply in this situation?  As I mentioned before I'd like to squash empty commits (in the case of a one-time migration, in the bidirectional case it's probably easier not to), so many SVN tagging operations wouldn't produce new commits, and the (technically) correct commit is tagged.  In the case of actual content changes in a tag's life, I think it's up to the user to decide between three options:\n\n  1) only retain the last SVN tag\n  2) tag using the git-svn-style 'tagname@rev' for all but the last\n  3) Do (2), but move older tags to some hidden namespace (refs/hidden/tags or the like)\n\nOption (3) is predicated on gc searching accepting all subdirectories of refs/ as valid (it did this when I wrote my original scripts, and I don't believe this behavior has changed).  For a one-time migration I think all three of these options can be implemented using annotated tags.  In the bidirectional case things get murky (maybe always tag with tagname@rev and hope for tab completion?).\n\n> ... snip ...\n>\n> SVN export is much more complicated, and has taken most of my time.\n> Although there's no way to tell automatically which directories are\n> branches, I realised you can detect trunks automatically about 90% of\n> the time by looking for where files/directories are first created, and\n> can detect non-trunk branches at least 99% of the time once you know\n> the trunks.  Trunk detection is normally just a convenience, but can\n> be a lifesaver when importing from a sufficiently messy repository.\n> There's a lot you can do with merge detection, but between\n> svn:mergeinfo, svnmerge.py[4], svk[5] and log messages I realise I've\n> only scratched the surface.  In many ways, this is a classic Perl\n> problem - take a bunch of messy textual input, string it all together\n> and make some neat textual output.  The code I've got right now seems\n> fine for anything up to about 20,000 commits, but eats way too much\n> memory for huge repos.  Depending on how much time I get, I might have\n> tackled that problem by the time I put this code online.\n\nBranch detection falls into my branch mapping mentioned above.  I realized that for as much as I slaved over finding every last branch (regardless of location) in our repo, it really came down to \"list the known branches, find copies of those, if any copies leave the list of known branches warn the user/revise branch spec, repeat\".  Now there are probably SVN repositories out there that can't write a well-formed branch spec (I had one at a previous job where the definition of a branch changed halfway through history), but I'm attempting to convince myself something is better than nothing and we'll catch the exceptions as they come up.\n\nMerge detection and translation is (IMO) not a well formed problem.  As you point out, there's no real way to know if it's a real merge or a cherry-pick.  It occurred to me last night after reading your e-mail a tool could attempt to look at the diffs in the branch history and compare that with the diff created by the merge, but that gets into all kinds of diff machinery/text processing hijinks that I don't want to contemplate (others might be more willing).\n \n> [1] http://comments.gmane.org/gmane.comp.version-control.git/158940\n> [2] https://github.com/rcaputo/snerp-vortex\n> [3] http://repo.or.cz/w/svn-merge2git.git\n> [4] http://www.orcaware.com/svn/wiki/Svnmerge.py\n> [5] http://search.cpan.org/dist/SVK/\n\nThanks for the brain dump.  I hope this brain dump in response is helpful in someway.  I seem to revisit this topic about once every three months or so, and this was a good chance to share some of my recent revelations.\n\nThanks,\nStephen\n"},{"id":"186155","messageId":"4F554BE4.5010401@pileofstuff.org","threadId":"29821","inReplyTo":"3c2ab05e-b2af-4df4-bca6-ff5512b0c73e@mail","subject":"Re: Approaches to SVN to Git conversion","fromName":"Andrew Sayers","fromEmail":"andrew-git@pileofstuff.org","sentAt":"2012-03-05T23:27:32Z","receivedAt":"2012-03-05T23:27:32Z","isPatch":false,"sender":{"key":"andrew-git@pileofstuff.org","avatar":null},"body":"On 05/03/12 15:27, Stephen Bash wrote:\n> All-\n> \n> This turned out to be longer than I intended, but actually summarizes some of my modern thoughts on SVN to Git conversion (as always, I'm more interested in one time migration than bidirectional operation, so read with a grain of salt).\n> \n> More inline...\n> \n> ----- Original Message -----\n>> From: \"Andrew Sayers\" <andrew-git@pileofstuff.org>\n>> Sent: Sunday, March 4, 2012 8:36:41 AM\n>> Subject: Re: [RFC] \"Remote helper for Subversion\" project\n>>\n>> ... snip ...\n>>\n>> While researching the problem, I found Stephen Bash's original\n>> proposal[1] and snerp-vortex[2] quite inspiring, but wasn't able to\n>> find any details on SoC-related work in the branching-and-merging\n>> department - hopefully the following isn't just a retread of ideas\n>> developed since then.  I've concentrated on importing from SVN so far,\n>> but have kept an eye on update and half an eye on bi-direction in the\n>> hopes of being useful there some day.\n>>\n>> It seems to me the \"svn export\" and \"git import\" steps make most sense\n>> as two unrelated projects.  Snerp-vortex and Stephen's scripts both\n>> cut the history import problem at that point, as do\n>> svn-fe/git-fast-import with code import.  Exporting SVN history is a\n>> messy and sometimes project-specific job, so allowing a project to\n>> concentrate on that part makes it possible for SVN experts to use all\n>> their skills without having to learn git plumbing before they make\n>> their first commit (much respect to Stephen for managing that feat\n>> BTW).\n> \n> After many a long conversation with Ram and Jonathan (and others), I'm actually going the other direction.  My current thinking (and this is very much open for discussion) is that as long as the SVN properties are available (especially the copyfrom information) Git has just as much information (if not more) to reconstruct the SVN history as SVN does.  (And going through our messy history I haven't found any counterpoint to this yet)\n\nI agree that git can be taught a superset of the information in SVN, but\nyou'll need absolutely all SVN properties available - somewhere out\nthere, someone has a created a merge script that sets \"my-merge-info:\nrevisions 1 to 10 from trunk (inclusive)\".  Building on your point\nbelow, a git-based converter could take an ambiguous message like \"merge\nrevision 123 from trunk\" and compare the diff for this commit against\nthe diff for r123, for r1:r123, for r1:r132 and so on until it found a\ngood match.\n\nI'm personally more interested in extracting SVN history from an SVN\ndump than from git (which I'll discuss in a moment), but these\napproaches sound fundamentally compatible - it sounds like we agree that\none part of the problem is a is a simple process that needs some\noptimisation work and the other is a good candidate for full employment\ntheorem[1].  Putting an interface between them means we can write one\nimplementation for the bit with an obvious right answer, and continue\nexperimenting with different solutions for the other bit.\n\nI described the work I've done as a solution with three parts (SVN\nexport, file format, Git import), but it's equally valid to call them\nthree different projects that were initially developed in parallel.  I'd\nbe quite happy for the format and import parts to move towards becoming\npart of the wider \"remote helper\" project, and the export part to become\na CPAN module that's just one of many programs that writes the\nhistory-import format (a bit like how svn-fe is one of many programs\nthat writes the git-fast-import format).\n\nI wrote my SVN exporter based on SVN dumps for three reasons - I figured\npeople switching from SVN would be more comfortable customising a\nsolution that only used technologies they understood, I figured it might\nbe useful to Mercurial or Bazaar some day if it was DVCS-neutral, and I\nhave to use SVN for my day job so I'm more interested in getting a good\nmigration story today than a great one tomorrow.\n\nMy instinct is that under ideal conditions, the SVN exporter I've got\nwill initially surge forward, build up a great collection of test cases,\nthen gradually sink under the weight of hacks needed to pass the tests.\n This strikes me as a good reason to isolate the exporter from git\nitself, but also a necessary step in getting SVN import done - there's\nno way to really know how to architect a good solution until we can\nbuild a spec from real world test cases.\n\n> \n>> I've written a proof-of-concept history converter that can be split\n>> into three parts: a format for describing SVN history; a large, often\n>> messy Perl program that writes files in that format; and a small Perl\n>> program that reads the format and translates it into git.  With\n>> hindsight, Perl is the right language for the SVN exporter, but the\n>> git importer would have been better written in C.\n>>\n>> Personally, I think SVN export will always need a strong manual\n>> component to get the best results, so I've put quite a bit of work\n>> into designing a good SVN history format.  Like git-fast-import, it's\n>> an ASCII format designed both for human and machine consumption...\n> \n> First, I'm very impressed that you managed to get a language like this up and working.  It could prove very useful going forward.  On the flip side, from my experiments over the last year I've actually been leaning toward a solution that is more implicit than explicit.  Taking git-svn as a model, I've been trying to define a mapping system (in Perl):\n\nJust to be clear, the language described above goes after the clever\nbranch-detection work described below - all branches need to be\nspecified explicitly in the format so that it can be a simple mechanical\njob.\n\n> \n>   my %branch_spec = { '/trunk/projname' => 'master',\n>                       '/branches/*/projname' => '/refs/heads/*' };\n>   my %tag_spec = { '/tags/*/projname' => '/refs/tags/*' };\n> \n> (See [1] for notes on our SVN structure.  In a std-layout-style repo it would be:\n> \n>   my %branch_spec = { '/projname/trunk' => 'master',\n>                       '/projname/branches/*' => '/refs/heads/*' };\n>   my %tag_spec = { '/projname/tags/*' => '/refs/tags/*' };)\n> \n> Now I know this simple mapping will fail as I get further in our history -- in particular we have one branch that came from:\n> \n>   svn cp $SVN_REPO/trunk/ $SVN_REPO/foo  # OOPS! not in branches!\n>   svn mv $SVN_REPO/foo $SVN_REPO/branches/foo\n> \n>>From an automation perspective, I expect the first svn operation to produce an error saying \"Possible branch created at /foo from known branch trunk (master), but doesn't match any known branch spec\", while the second (if \"continue-on-error\" is turned on) would error \"Branch /branches/foo created from unknown branch located at /foo\".  It's then up to the user to modify the branch map to something that accounts for this behavior:\n> \n>   my %branch_spec = { '/trunk/projname' => 'master',\n>                       '/branches/*/projname' => '/refs/heads/*',\n>                       '/foo' => '/refs/heads/foo' };\n>   my %tag_spec = { '/tags/*/projname' => '/refs/tags/*' };\n> \n> So in this case I'm making an explicit branch mapping, but the use of glob-style syntax allows the user to catch larger classes of branches if desired (I'll also note that depending on the implementation /foo may need to map to /refs/heads/bad-foo so that /branches/foo can map to /refs/heads/foo, but my intention has been to squash empty commits, so it's possible the name conflict is a non-issue since content didn't change).  I also know that we have some copy operations that are just weird, so it might be helpful to have an ignore mechanism that tells the system to ignore copies into/out-of certain SVN paths.\n> \n\nI started with an approach like you describe, but as you say it winds up\nin a mess of special cases.  A friend pointed me to Perl's catalyst\nrepository[2], which is a wonderful haven of every mad SVN thing ever\ndreamt up.  That got me playing with more general heuristics, and while\nwriting this e-mail I think I've finally nailed it.  What do you say to\ndefining SVN branches like this:\n\nA directory is a branch if...\n1. it is not a subdirectory of an existing branch; and\n2. either:\n2a. it is in a list of branches specified by the user, or\n2b. it is copied from a (subdirectory of a) branch\n\nI'll have to go and play with the implementation details, but I don't\nsee any misbehaviour jumping out of this approach.  Rule 1 discounts the\n\"svn cp /branches/foo /trunk/libfoo\" pattern, but I'm fine with that\nbecause I don't think anyone really has a good answer there yet.  Rule\n2a allows the user as much input as they want but only requires most\npeople to say that /trunk is a branch.  Finally, rule 2b counts the \"svn\ncp /trunk/libfoo /branches/foo\" pattern - I'm personally not bothered by\nthe asymmetry there compared to rule 1.\n\nI'm afraid I've spent all my free time tonight writing this e-mail, but\nthe proof-of-concept code is only based on a sloppy half-realisation of\nthe above so probably wouldn't be that enlightening anyway.\n\n>> Once the format is defined, git import is fairly straightforward.\n>> Proof-of-concept code to follow, but it's really just a wrapper around\n>> git-commit-tree, git-mktag etc.  I wrote this in Perl thinking it\n>> would relate somehow to git-svn, but eventually realised it didn't and\n>> that a few hundred calls to (plumbing) processes per second isn't so\n>> good for performance.  The only interesting part of the problem is how\n>> to tackle SVN tags.  I went for an ambitious approach, making normal\n>> tags where possible and downgrading them to lightweight tags when\n>> necessary.  This does involve managing something that is effectively a\n>> branch in refs/tags/, but what else is an SVN tag but a branch in the\n>> wrong namespace?\n> \n> I don't understand how \"normal\" and \"lightweight\" apply in this situation?  As I mentioned before I'd like to squash empty commits (in the case of a one-time migration, in the bidirectional case it's probably easier not to), so many SVN tagging operations wouldn't produce new commits, and the (technically) correct commit is tagged.  In the case of actual content changes in a tag's life, I think it's up to the user to decide between three options:\n> \n>   1) only retain the last SVN tag\n>   2) tag using the git-svn-style 'tagname@rev' for all but the last\n>   3) Do (2), but move older tags to some hidden namespace (refs/hidden/tags or the like)\n> \n> Option (3) is predicated on gc searching accepting all subdirectories of refs/ as valid (it did this when I wrote my original scripts, and I don't believe this behavior has changed).  For a one-time migration I think all three of these options can be implemented using annotated tags.  In the bidirectional case things get murky (maybe always tag with tagname@rev and hope for tab completion?).\n> \n\nI didn't explain this particularly well, as it's based largely on the\nvague desire to make update work some day.  Imagine the user does this:\n\n* git svn-pull # get tags/foo, a candidate for an annotated tag\n... time passes ...\n* git svn-pull # tags/foo has now been updated in another revision\n\nIf we create an annotated tag in step 1, what do we do in step 2?  You\ncan't make the tag object the parent of a new revision, so you need to\ndo something unpleasant.  The solution I proposed was to convert the tag\nmessage to a commit message (i.e. pretend a lightweight tag had been\ncreated all along), then add another commit on top of it and make a\nlightweight tag from the new commit (i.e. treat it like a branch).  In\nretrospect that's far too much magic without user involvement - a better\nsolution would be to give the user this option along with the ones you\noutlined, and let git-config remember their preference if they want.\n\n\t- Andrew\n\n[1] http://en.wikipedia.org/wiki/Full_employment_theorem\n[2] http://dev.catalyst.perl.org/repos/bast/\n"},{"id":"186217","messageId":"9130e486-21bd-4c8c-9647-b627dbc1e5c6@mail","threadId":"29821","inReplyTo":"4F554BE4.5010401@pileofstuff.org","subject":"Re: Approaches to SVN to Git conversion","fromName":"Stephen Bash","fromEmail":"bash@genarts.com","sentAt":"2012-03-06T14:36:31Z","receivedAt":"2012-03-06T14:36:31Z","isPatch":false,"sender":{"key":"bash@genarts.com","avatar":null},"body":"----- Original Message -----\n> From: \"Andrew Sayers\" <andrew-git@pileofstuff.org>\n> Sent: Monday, March 5, 2012 6:27:32 PM\n> Subject: Re: Approaches to SVN to Git conversion\n> \n> > My current thinking (and this is very much open for discussion) is\n> > that as long as the SVN properties are available (especially the\n> > copyfrom information) Git has just as much information (if not more)\n> > to reconstruct the SVN history as SVN does.  (And going through our\n> > messy history I haven't found any counterpoint to this yet)\n> \n> I agree that git can be taught a superset of the information in SVN,\n> but you'll need absolutely all SVN properties available...\n\nI'm pretty sure Jonathan won't be happy with anything less ;)\n\n> I wrote my SVN exporter based on SVN dumps for three reasons - I\n> figured people switching from SVN would be more comfortable\n> customising a solution that only used technologies they understood, I\n> figured it might be useful to Mercurial or Bazaar some day if it was\n> DVCS-neutral, and I have to use SVN for my day job so I'm more\n> interested in getting a good migration story today than a great one\n> tomorrow.\n\nThe multiple systems argument is a good one.\n\n> >   my %branch_spec = { '/trunk/projname' => 'master',\n> >                       '/branches/*/projname' => '/refs/heads/*' };\n> >   my %tag_spec = { '/tags/*/projname' => '/refs/tags/*' };\n> > \n> > Now I know this simple mapping will fail as I get further in our\n> > history -- in particular we have one branch that came from:\n> > \n> >   svn cp $SVN_REPO/trunk/ $SVN_REPO/foo  # OOPS! not in branches!\n> >   svn mv $SVN_REPO/foo $SVN_REPO/branches/foo\n> > \n> > It's then up to the user to modify the branch\n> > map to something that accounts for this behavior:\n> > \n> >   my %branch_spec = { '/trunk/projname' => 'master',\n> >                       '/branches/*/projname' => '/refs/heads/*',\n> >                       '/foo' => '/refs/heads/foo' };\n> >   my %tag_spec = { '/tags/*/projname' => '/refs/tags/*' };\n> \n> I started with an approach like you describe, but as you say it winds\n> up in a mess of special cases.  A friend pointed me to Perl's catalyst\n> repository[2], which is a wonderful haven of every mad SVN thing ever\n> dreamt up.  That got me playing with more general heuristics, and\n> while writing this e-mail I think I've finally nailed it.  What do you\n> say to defining SVN branches like this:\n> \n> A directory is a branch if...\n> 1. it is not a subdirectory of an existing branch; and\n> 2. either:\n> 2a. it is in a list of branches specified by the user, or\n> 2b. it is copied from a (subdirectory of a) branch\n\nI think I started with a very similar set of rules...  Looking at my code now I'm having a hard time summarizing them (probably because they evolved with the code, so what started simple morphed into something pretty complicated).  I guess as long as the user has the option to say \"no, don't treat this copy as a branch\" (or equivalently the Git side of things has a way to say \"ignore this branch\") these rules would be okay.  But at that point we're back to a list of exceptions -- really we're arguing white-list vs black-list... I eventually chose to go the white-list route for our conversion after starting with black-list (a white-list that still required a few manual edits before manipulating the Git history).  So take that single data point for what it's worth.\n\n> > > Once the format is defined, git import is fairly straightforward.\n> > > Proof-of-concept code to follow, but it's really just a wrapper\n> > > around git-commit-tree, git-mktag etc.  I wrote this in Perl\n> > > thinking it would relate somehow to git-svn, but eventually\n> > > realised it didn't and that a few hundred calls to (plumbing)\n> > > processes per second isn't so good for performance.  The only\n> > > interesting part of the problem is how to tackle SVN tags.  I went\n> > > for an ambitious approach, making normal tags where possible and\n> > > downgrading them to lightweight tags when necessary.  This does\n> > > involve managing something that is effectively a branch in\n> > > refs/tags/, but what else is an SVN tag but a branch in the wrong\n> > > namespace?\n> > \n> > I don't understand how \"normal\" and \"lightweight\" apply in this\n> > situation? ... In the case of actual content changes in a tag's\n> > life, I think it's up to the user to decide between three options:\n> > \n> >   1) only retain the last SVN tag\n> >   2) tag using the git-svn-style 'tagname@rev' for all but the last\n> >   3) Do (2), but move older tags to some hidden namespace\n> >      (refs/hidden/tags or the like)\n> > \n> > ... In the bidirectional case things get murky (maybe always tag\n> > with tagname@rev and hope for tab completion?).\n> \n> I didn't explain this particularly well, as it's based largely on the\n> vague desire to make update work some day.  Imagine the user does\n> this:\n> \n> * git svn-pull # get tags/foo, a candidate for an annotated tag\n> ... time passes ...\n> * git svn-pull # tags/foo has now been updated in another revision\n> \n> If we create an annotated tag in step 1, what do we do in step 2?  You\n> can't make the tag object the parent of a new revision, so you need to\n> do something unpleasant.  The solution I proposed was to convert the\n> tag message to a commit message (i.e. pretend a lightweight tag had\n> been created all along), then add another commit on top of it and make\n> a lightweight tag from the new commit (i.e. treat it like a branch).\n> In retrospect that's far too much magic without user involvement - a\n> better solution would be to give the user this option along with the\n> ones you outlined, and let git-config remember their preference if\n> they want.\n\nOkay, that's what I thought you meant (and what I classified as a bidirectional problem, but I guess it's not strictly a bidirectional problem, but a one-time migration does not have the problem).  If you want to continue to update Git from SVN there are two cases to consider:\n\n  1) Each Git repository *only* talks to SVN\n  2) The Git repository is cloned for further use \n     (So the chain is something like SVN->Git->Git)\n\nIn (1) your lightweight tag solution is probably okay (but I'm pretty sure creating/deleting annotated tags would behave the same way because no one else sees the Git tag object).  In (2) I think there would still be a tag conflict when the upstream Git repo replaces a lightweight tag and the downstream repo attempts to fetch it.  I don't know what the fetch/pull machinery does when there's a lightweight tag conflict (I'm guessing either bails out or keeps the local one?).  Case (2) motivates me to say always generate (annotated?) tags named tagname@rev so there can be no conflicts.  In that case the only difference I see is if we create an empty Git commit with the tag message plus a lightweight tag or tag the original commit with an annotated tag (I think it's fairly obvious I'm a fan of\n  the latter).\n\n> [1] http://en.wikipedia.org/wiki/Full_employment_theorem\n> [2] http://dev.catalyst.perl.org/repos/bast/\n\nThanks,\nStephen\n"},{"id":"186240","messageId":"CA+7g9Jwb=7wH7R3=ShhOGMdHXWmq4ZahocpaEuJdf+yBfCpA8A@mail.gmail.com","threadId":"29821","inReplyTo":"3c2ab05e-b2af-4df4-bca6-ff5512b0c73e@mail","subject":"Re: Approaches to SVN to Git conversion (was: Re: [RFC] \"Remote helper for Subversion\" project)","fromName":"Nathan Gray","fromEmail":"n8gray@n8gray.org","sentAt":"2012-03-06T19:29:59Z","receivedAt":"2012-03-06T19:29:59Z","isPatch":false,"sender":{"key":"n8gray@n8gray.org","avatar":"https://avatars.githubusercontent.com/u/82794?v=4"},"body":"Hi everyone,\n\nSoon I'm going to be undertaking a migration of a subproject from a\nvery messy multiproject SVN repo to git, so this is a topic that's\nquite near to my heart at the moment.  More inline...\n\nOn Mon, Mar 5, 2012 at 7:27 AM, Stephen Bash <bash@genarts.com> wrote:\n>\n> ----- Original Message -----\n>> From: \"Andrew Sayers\" <andrew-git@pileofstuff.org>\n>> Sent: Sunday, March 4, 2012 8:36:41 AM\n>> Subject: Re: [RFC] \"Remote helper for Subversion\" project\n>>\n\n[snip]\n\n>> Personally, I think SVN export will always need a strong manual\n>> component to get the best results, so I've put quite a bit of work\n>> into designing a good SVN history format.  Like git-fast-import, it's\n>> an ASCII format designed both for human and machine consumption...\n>\n> First, I'm very impressed that you managed to get a language like this up and working.  It could prove very useful going forward.  On the flip side, from my experiments over the last year I've actually been leaning toward a solution that is more implicit than explicit.  Taking git-svn as a model, I've been trying to define a mapping system (in Perl):\n>\n>  my %branch_spec = { '/trunk/projname' => 'master',\n>                      '/branches/*/projname' => '/refs/heads/*' };\n>  my %tag_spec = { '/tags/*/projname' => '/refs/tags/*' };\n\nThe problem of specifying and detecting branches is a major problem in\nmy upcoming conversion.  We've got toplevel trunk/branches/tags\ndirectories but underneath \"branches\" it's a free-for-all:\n\n/branches/codenameA/{projectA,projectB,projectC}\n/branches/codenameB   (actually a branch of projectA)\n/branches/developers/joe/frobnicator-experiment (also a branch of projectA)\n\nClearly there's no simple regex that's going to capture this, so I'm\nreduced to listing every branch of projectA, which is tedious and\nerror-prone.  However, what *would* work fabulously well for me is\n\"marker file\" detection.  Every copy of projectA has a certain file at\nit's root.  Let's call it \"markerFile.txt\".  What I'd really love is a\nway to say:\n\nmy %branch_markers = {'/branches/**/markerFile.txt' => '/refs/heads/**'}\n\nI'm using ** to signify that this may match multiple path components\n(sorry, I don't know perl glob syntax).  A branch point is any\nrevision that creates a new file that matches the marker pattern.\n\nIdeally one could use logical connectives like AND and OR to specify a\nset of patterns that could account for marker files changing over the\nhistory of the project, but for my purposes that wouldn't be necessary\n-- we've got a well-defined marker that's always present.\n\nFor bonus points I'd like to be able to speed things up by excluding\nknown-bad markers.  Say projectB has a file \"badMarker.txt\" at its\nroot and I don't want to import projectB into my new repo.  Maybe I\ncould specify:\n\nmy %branch_spec = {\n        '/branches/**/markerFile.txt' => '/refs/heads/**',\n        '/branches/**/badMarker.txt' => '!'}\n\nI'm assuming that it would be helpful for the script to have this\ninformation (e.g. it could stop recursive searches when badMarker.txt\nis found), but maybe that's not the case.\n\nI'd welcome any comments or (especially!) code to try out.  ;^)\n\nCheers,\n-Nathan\n\n-- \nhttp://n8gray.org\n"},{"id":"186249","messageId":"ab5eb5a7-a446-4dc3-b8e8-e3f7ec306452@mail","threadId":"29821","inReplyTo":"CA+7g9Jwb=7wH7R3=ShhOGMdHXWmq4ZahocpaEuJdf+yBfCpA8A@mail.gmail.com","subject":"Re: Approaches to SVN to Git conversion (was: Re: [RFC] \"Remote helper for Subversion\" project)","fromName":"Stephen Bash","fromEmail":"bash@genarts.com","sentAt":"2012-03-06T20:35:45Z","receivedAt":"2012-03-06T20:35:45Z","isPatch":false,"sender":{"key":"bash@genarts.com","avatar":null},"body":"\n\n----- Original Message -----\n> From: \"Nathan Gray\" <n8gray@n8gray.org>\n> Sent: Tuesday, March 6, 2012 2:29:59 PM\n> Subject: Re: Approaches to SVN to Git conversion (was: Re: [RFC] \"Remote helper for Subversion\" project)\n> \n> > > Personally, I think SVN export will always need a strong manual\n> > > component to get the best results, so I've put quite a bit of work\n> > > into designing a good SVN history format.  Like git-fast-import,\n> > > it's an ASCII format designed both for human and machine\n> > > consumption...\n> >\n> > First, I'm very impressed that you managed to get a language like\n> > this up and working.  It could prove very useful going forward.   On\n> > the flip side, from my experiments over the last year I've actually\n> > been leaning toward a solution that is more implicit than explicit.\n> > Taking git-svn as a model, I've been trying to define a mapping\n> > system (in Perl):\n> >\n> >  my %branch_spec = { '/trunk/projname' => 'master',\n> >                      '/branches/*/projname' => '/refs/heads/*' };\n> >  my %tag_spec = { '/tags/*/projname' => '/refs/tags/*' };\n> \n> The problem of specifying and detecting branches is a major problem in\n> my upcoming conversion.  We've got toplevel trunk/branches/tags\n> directories but underneath \"branches\" it's a free-for-all:\n> \n> /branches/codenameA/{projectA,projectB,projectC}\n> /branches/codenameB   (actually a branch of projectA)\n> /branches/developers/joe/frobnicator-experiment (also a branch of\n> projectA)\n> \n> Clearly there's no simple regex that's going to capture this, so I'm\n> reduced to listing every branch of projectA, which is tedious and\n> error-prone.  However, what *would* work fabulously well for me is\n> \"marker file\" detection.  Every copy of projectA has a certain file at\n> it's root.  Let's call it \"markerFile.txt\".  What I'd really love is a\n> way to say:\n> \n> my %branch_markers = {'/branches/**/markerFile.txt' =>\n>                       '/refs/heads/**'}\n\nOoo...  I like it.  I hadn't hit on this idea yet, but it certainly is a very helpful heuristic.  I doubt I'd have any sort of demo code for you in the near future, but it's definitely an idea to roll into the mix.\n\nThanks,\nStephen\n"},{"id":"186268","messageId":"4F5690FB.9060800@pileofstuff.org","threadId":"29821","inReplyTo":"CA+7g9Jwb=7wH7R3=ShhOGMdHXWmq4ZahocpaEuJdf+yBfCpA8A@mail.gmail.com","subject":"Re: Approaches to SVN to Git conversion","fromName":"Andrew Sayers","fromEmail":"andrew-git@pileofstuff.org","sentAt":"2012-03-06T22:34:35Z","receivedAt":"2012-03-06T22:34:35Z","isPatch":false,"sender":{"key":"andrew-git@pileofstuff.org","avatar":null},"body":"I've now added a bit of documentation and uploaded my code to github:\nhttps://github.com/andrew-sayers/Proof-of-concept-History-Converter\n\nI haven't attached it here because the code isn't at a stage where it\nwould be useful to review line-by-line.  Comments are welcome if you\nreally want to though :)\n\nsvn-branch-export.pl makes heavy use of SVN::Dump.  You may want to get\nthe latest version from github if speed is important to you:\nhttps://github.com/book/SVN-Dump/ - many thanks to Philippe Bruhat for\naccepting my performance patch so quickly.\n\nHere are some particular gripes I have with the code I've uploaded:\n\ngit-branch-import.pl gets the revision number by parsing out the\n\"git-svn-id\" in commit messages - as I mentioned earlier, I started off\nthinking this script would be closely related to git-svn somehow.  In\nhindsight it would be better to read revision numbers from the marks\nfile exported by git-fast-import.\n\nBranch History Format has some git-specific stuff in the setup section.\n I didn't think about this in too much detail while writing it, but\nDVCS-neutrality would be better served by turning these into\ncommand-line options.\n\nAs mentioned before, branch detection in svn-branch-export.pl is rather\nmuddled, as my understanding of the problem evolved significantly while\nwriting it.\n\nsvn-branch-export.pl half-heartedly uses a configure/make/make install\nanalogy to describe its behaviour - I'm increasingly sure this is\ngimmicky and awful, rather than a neat explanatory trick.\n\nsvn-branch-export.pl exposes a lot of config values (e.g. \"log_style\")\nthat just bulk up the implementation and create space for bugs to creep\nin without adding much actual value.  They should be removed.\n\nOn 06/03/12 19:29, Nathan Gray wrote:\n<snip>\n> \n> The problem of specifying and detecting branches is a major problem in\n> my upcoming conversion.  We've got toplevel trunk/branches/tags\n> directories but underneath \"branches\" it's a free-for-all:\n> \n> /branches/codenameA/{projectA,projectB,projectC}\n> /branches/codenameB   (actually a branch of projectA)\n> /branches/developers/joe/frobnicator-experiment (also a branch of projectA)\n> \n> Clearly there's no simple regex that's going to capture this, so I'm\n> reduced to listing every branch of projectA, which is tedious and\n> error-prone.  However, what *would* work fabulously well for me is\n> \"marker file\" detection.  Every copy of projectA has a certain file at\n> it's root.  Let's call it \"markerFile.txt\".  What I'd really love is a\n> way to say:\n\nThis is quite close to the implementation I've got.  The SVN exporter\nruns in two stages:\n\nIn the first stage, the script treats any non-blacklisted file as a\nmarker file, but only looks for trunk branches.  It looks all through\nthe history, traces back through the copyfroms, and tries to find the\noriginal directory associated with the file.  Usually it decides that\nthe only branch without a copyfrom is /trunk.  Searching just for trunks\nwith this weak heuristic makes it much easier to hand-verify the result.\n\nIn the second stage, the script looks through the history again, tracing\nthe copies of known branches in a slightly less clever way than\ndescribed in my previous e-mail.  There's no need for marker files this\ntime round, as we just assume any `svn cp /trunk\n/directory/not/within/a/branch` is a new branch.  In my experiments this\nhas been a pretty solid way of detecting branches without too much human\ninput - I might be missing something (or have mis-explained something),\nbut I'd be interested to hear examples of where this would go wrong.\nHaving said that, here's a dodgy example I'd like to pre-emptively defend:\n\n\tsvn add tronk\n\tsvn ci -m \"Created trunk\" # r1\n\tsvn cp tronk trunk\n\tsvn ci -m \"D'oh\" # r2\n\tsvn rm tronk\n\tsvn add trunk/markerFile.txt\n\tsvn ci -m \"Double d'oh!\" # r3\n\nYou could argue that the correct branch history description for the\nabove would be:\n\n\tIn r3, create branch \"trunk\"\n\nIn other words, ignore everything that happened before the marker file\nwas created.  However, I would argue the following representation is\nmore correct:\n\n\tIn r1, create branch \"tronk\"\n\tIn r2, create branch \"trunk\" from \"tronk\" r1\n\tIn r3, delete branch \"tronk\"\n\nThe branch history format supports the \"delete branch\" command (remove\nthe branch entirely) as well as the more common \"deactivate branch\"\n(keep the branch but don't accept any new commits) specifically to deal\nwith this sort of weirdness.  Creating a branch then deleting it keeps\nthe r1 revision log intact as part of the \"trunk\" branch, without\nleaving any useless branches lying around.\n\n\t- Andrew\n"},{"id":"186280","messageId":"4F56A4DF.8060807@vilain.net","threadId":"29821","inReplyTo":"ab5eb5a7-a446-4dc3-b8e8-e3f7ec306452@mail","subject":"Re: [spf:guess] Re: Approaches to SVN to Git conversion (was: Re: [RFC] \"Remote helper for Subversion\" project)","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2012-03-06T23:59:27Z","receivedAt":"2012-03-06T23:59:27Z","isPatch":false,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"On 3/6/12 12:35 PM, Stephen Bash wrote:\n>> The problem of specifying and detecting branches is a major problem in\n>> my upcoming conversion.  We've got toplevel trunk/branches/tags\n>> directories but underneath \"branches\" it's a free-for-all:\n>>\n>> /branches/codenameA/{projectA,projectB,projectC}\n>> /branches/codenameB   (actually a branch of projectA)\n>> /branches/developers/joe/frobnicator-experiment (also a branch of\n>> projectA)\n>>\n>> Clearly there's no simple regex that's going to capture this, so I'm\n>> reduced to listing every branch of projectA, which is tedious and\n>> error-prone.  However, what *would* work fabulously well for me is\n>> \"marker file\" detection.  Every copy of projectA has a certain file at\n>> it's root.  Let's call it \"markerFile.txt\".  What I'd really love is a\n>> way to say:\n>>\n>> my %branch_markers = {'/branches/**/markerFile.txt' =>\n>>                        '/refs/heads/**'}\n>\n> Ooo...  I like it.  I hadn't hit on this idea yet, but it certainly is a very helpful heuristic.  I doubt I'd have any sort of demo code for you in the near future, but it's definitely an idea to roll into the mix.\n\nWhat I did for the Perl Perforce conversion is make this a multi–step \nprocess; first, the heuristic goes through and detects branches and \nmerge parents.  Then you do the actual export.  If, however, the \nheuristic gets it wrong, then you can manually override the branch \ndetection for a particular revision, which invalidates all of the \n_automatic_ decisions made for later revisions the next time you run it.\n\nEven with all of the information in Postgres, and much of the hard work \npushed into the Postgres engine, and Postgres tuned for OLAP, this was \nthe slowest part of the operation.  For a 30,000–odd revision Perforce \nrepository.\n\nThe manual input is extremely useful for bespoke conversions; there will \nalways be warts in the history and no heuristic is perfect (even if you \ncan supply your own set of expressions, a way to override it for just \none revision is handy).\n\nJust to revise, the steps in git-p4raw, are:\n\n* load metadata (git-p4raw load ; git-p4raw check)\n* load blobs (git-p4raw export-blobs)\n* find project roots (git-p4raw find-branches)\n\n   Project root decisions can be overridden, in git-p4raw this was \nthrough a DB insert, but all this consisted of was inserting (revision, \nbranch) tuples into the appropriate table so a front–end would be \ntrivial.  As you suggest, a custom heuristic is also an option but the \nmost flexible solution is just being able to override the decisions made \nfor a particular revision.\n\n* detect project merges (also done by git-p4raw find-branches)\n\nDetecting merge parents used a heuristic based on the per–file \nintegration records and a computation based on an internal diff-tree \nwhich produced a list of files that would have needed resolving.  This \none I actually used enough to bother implementing a front–end for:\n\n   git-p4raw graft REV PARENT PARENT\n\nWhere 'PARENT' could be another project root (revision/branch location), \nor it could be a git commit ID (for the inevitable occasion where you \nneed to manually graft on some history).  This interface allows you to \ndo several things:\n\n   1. mark a merge which was not recorded correctly in history\n   2. un–mark a merge which was detected/recorded incorrectly\n   3. skip bad sections of history, for instance squash merging merges \nwhich happened over several commits (SVN and Perforce, of course, \nsupport insane piecemeal merging prohibited by git)\n\n* the actual fast-import exporter.\n\n   git-p4raw export-commits 1..5000\n\nThere was also an important reverse operation:\n\n   git-p4raw unexport-commits 2500\n\nWhich moved all of the exported refs backwards, deleted ones which \ndidn't exist at revision 2500.\n\nOnce the data has been mined, the actual exporting can proceed very \nfast.  Eg, on my laptop I could easily be topping 300 commits per second \nwhich makes for a nice export/examine/rewind/adjust cycle.\n\nFor more information,\n\n   git clone git://github.com/samv/git-p4raw\n   cd git-p4raw\n   perldoc git-p4raw\n\nThe \"Game plan.\" section of the POD is particularly relevant.  Remember \nthat SVN is very similar to Perforce in virtually all of its design \ndetails so this tool, its database schema, and implementation are all \nvery relevant to the design of the new svn-fe importer.\n\nSam\n"},{"id":"186318","messageId":"4F5780F0.5080901@vilain.net","threadId":"29821","inReplyTo":"4F5690FB.9060800@pileofstuff.org","subject":"Re: Approaches to SVN to Git conversion","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2012-03-07T15:38:24Z","receivedAt":"2012-03-07T15:38:24Z","isPatch":false,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"On 3/6/12 2:34 PM, Andrew Sayers wrote:\n> I've now added a bit of documentation and uploaded my code to github:\n> https://github.com/andrew-sayers/Proof-of-concept-History-Converter\n>\n> I haven't attached it here because the code isn't at a stage where it\n> would be useful to review line-by-line.  Comments are welcome if you\n> really want to though :)\n\nI just took a look at your readme—did you consider writing the tool to \nwork against an svn-fe import, rather than using SVN::Dump? Do you think \nit could be adjusted to be like that?\n\nSam\n"},{"id":"186334","messageId":"4F57C50A.9010701@pileofstuff.org","threadId":"29821","inReplyTo":"4F5780F0.5080901@vilain.net","subject":"Re: Approaches to SVN to Git conversion","fromName":"Andrew Sayers","fromEmail":"andrew-git@pileofstuff.org","sentAt":"2012-03-07T20:28:58Z","receivedAt":"2012-03-07T20:28:58Z","isPatch":false,"sender":{"key":"andrew-git@pileofstuff.org","avatar":null},"body":"On 07/03/12 15:38, Sam Vilain wrote:\n> On 3/6/12 2:34 PM, Andrew Sayers wrote:\n>> I've now added a bit of documentation and uploaded my code to github:\n>> https://github.com/andrew-sayers/Proof-of-concept-History-Converter\n>>\n>> I haven't attached it here because the code isn't at a stage where it\n>> would be useful to review line-by-line.  Comments are welcome if you\n>> really want to though :)\n> \n> I just took a look at your readme—did you consider writing the tool to\n> work against an svn-fe import, rather than using SVN::Dump? Do you think\n> it could be adjusted to be like that?\n\nI did consider writing svn-branch-export.pl against a branch created by\nsvn-fe, but right now it doesn't provide enough information to do a good\njob (e.g. copyfrom properties).  I understand that support is in the\nworks, but this project is more about getting a scrappy end-to-end\nsolution so we can see what the issues are (is there any demand for\nDVCS-neutral SVN history export?  What are the hard cases and how do you\nrepresent them?).  I'm keen to make sure that documentation and tests\nare done in such a way that a future git-based exporter could use them\nwithout relying on any of the actual code.\n\nI also considered writing git-branch-import.pl against the raw svn-fe\noutput.  As well as the technical issues with this approach, I felt like\nthese were better tackled as orthogonal problems.  Producing an accurate\nrepresentation of the SVN history is a very different problem to\nproducing a user-friendly representation, and separating those concerns\nseems like it will make life easier down the line.  For example, a\nuser-friendly representation might convert svn:ignore properties to\n.gitignore files, but that would make bidirection hard to implement\nwithout an accurate representation in the middle.\n\n\t- Andrew\n"},{"id":"186366","messageId":"4F57DBF0.4060101@pileofstuff.org","threadId":"29821","inReplyTo":"4F56A4DF.8060807@vilain.net","subject":"Re: [spf:guess] Re: Approaches to SVN to Git conversion (was: Re: [RFC] \"Remote helper for Subversion\" project)","fromName":"Andrew Sayers","fromEmail":"andrew-git@pileofstuff.org","sentAt":"2012-03-07T22:06:40Z","receivedAt":"2012-03-07T22:06:40Z","isPatch":false,"sender":{"key":"andrew-git@pileofstuff.org","avatar":null},"body":"It sounds like we've approached two similar problems in similar ways, so\nI'm curious about the differences where they exist.  I've been reading\nthis message of yours from 18 months ago alongside this thread:\nhttp://article.gmane.org/gmane.comp.version-control.git/150007\nUnfortunately these comprise everything I know about Perforce.\n\nI notice that git-p4raw stores all of its data in Postgres and provides\na programmatic interface for querying it, whereas I've focussed on\nproviding ASCII interfaces at relevant points.  I can see how a DB store\nwould help manage the amount of data you'd need to process in a big\nrepository, but were there any other issues that drove you down this\nroute?  Did you consider a text-based interface?\n\nOn 06/03/12 23:59, Sam Vilain wrote:\n<snip>\n> What I did for the Perl Perforce conversion is make this a multi–step\n> process; first, the heuristic goes through and detects branches and\n> merge parents.  Then you do the actual export.  If, however, the\n> heuristic gets it wrong, then you can manually override the branch\n> detection for a particular revision, which invalidates all of the\n> _automatic_ decisions made for later revisions the next time you run it.\n\nCould you give an example of overriding branch/merge detection?  It\nsounds like you're saying that if there's some problem detecting merge\nparents in an early revision, then all future merges are ignored by the\nscript.\n\n<snip>\n> The manual input is extremely useful for bespoke conversions; there will\n> always be warts in the history and no heuristic is perfect (even if you\n> can supply your own set of expressions, a way to override it for just\n> one revision is handy).\n\nAgain, would you mind providing a few examples?  It sounds like you have\nsome edge cases that could be handled by extending the branch history\nformat, but I'd like to pin it down a bit more before discussing solutions.\n\n<snip>\n>   3. skip bad sections of history, for instance squash merging merges\n> which happened over several commits (SVN and Perforce, of course,\n> support insane piecemeal merging prohibited by git)\n\nThis is an excellent point I've stumbled past in my experiments without\nrealising what I was seeing.  A simple SVN example might look like this:\n\n\tsvn add trunk branches\n\tsvn add trunk/foo trunk/bar\n\tsvn ci -m \"Initial revision\" # r1\n\n\tsvn cp trunk branches/my_branch\n\tsvn ci -m \"Created my_branch\" # r2\n\n\t# edit files in my_branch\n\n\tsvn merge branches/my_branch/foo trunk/foo\n\tsvn ci -m \"Merge my_branch -> trunk (1/3)\" # r11\n\n\tsvn merge branches/my_branch/bar trunk/bar\n\tsvn ci -m \"Merge my_branch -> trunk (2/3)\" # r12\n\n\tsvn cp branches/my_branch/new_file trunk/new_file\n\tsvn ci -m \"Merge my_branch -> trunk (3/3)\" # r13\n\nThis strikes me as a sensibly cautious workflow in SVN, where merge\nconflicts are common and changes are hard to revert.  The best\nrepresentation for this in the current branch history format would be\nsomething like this:\n\n\tIn r1, create branch \"trunk\"\n\tIn r2, create branch \"branches/my_branch\" from \"trunk\"\n\tIn r13, merge \"branches/my_branch\" r13 into \"trunk\"\n\nIn other words, pretend r11 and r12 are just normal commits, and that\nr13 is a full merge.  A more useful (and arguably more accurate)\nrepresentation would be possible if we extended the format a bit:\n\n\tIn r1, create branch \"trunk\"\n\tIn r2, create branch \"branches/my_branch\" from \"trunk\"\n\tIn r12, squash changes in \"branches/my_branch\"\n\tIn r13, squash changes in \"branches/my_branch\"\n\tIn r13, merge \"branches/my_branch\" r13 into \"trunk\"\n\nAdding \"squash\" and \"fixup\" commands would let us represent the whole\nmessy business as a single commit, which is closer to what the user was\ntrying to say even if it's further from what they actually had to say.\n\n\t- Andrew\n"},{"id":"186352","messageId":"CABURp0pLnMdFVgQ2+S_rW-KEjUBjKbEbMepEdPUkRJD+JQ_Ehg@mail.gmail.com","threadId":"29821","inReplyTo":"4F5690FB.9060800@pileofstuff.org","subject":"Re: Approaches to SVN to Git conversion","fromName":"Phil Hord","fromEmail":"phil.hord@gmail.com","sentAt":"2012-03-07T22:33:39Z","receivedAt":"2012-03-07T22:33:39Z","isPatch":false,"sender":{"key":"phil.hord@gmail.com","avatar":"https://avatars.githubusercontent.com/u/123908?v=4"},"body":"On Tue, Mar 6, 2012 at 5:34 PM, Andrew Sayers\n<andrew-git@pileofstuff.org> wrote:\n> This is quite close to the implementation I've got.  The SVN exporter\n> runs in two stages:\n>\n> In the first stage, the script treats any non-blacklisted file as a\n> marker file, but only looks for trunk branches.  It looks all through\n> the history, traces back through the copyfroms, and tries to find the\n> original directory associated with the file.  Usually it decides that\n> the only branch without a copyfrom is /trunk.  Searching just for trunks\n> with this weak heuristic makes it much easier to hand-verify the result.\n>\n> In the second stage, the script looks through the history again, tracing\n> the copies of known branches in a slightly less clever way than\n> described in my previous e-mail.  There's no need for marker files this\n> time round, as we just assume any `svn cp /trunk\n> /directory/not/within/a/branch` is a new branch.  In my experiments this\n> has been a pretty solid way of detecting branches without too much human\n> input - I might be missing something (or have mis-explained something),\n> but I'd be interested to hear examples of where this would go wrong.\n\nI think what you're describing would work perfectly for my weird svn\nrepo.  I have branches named like this:\n\nbranches/developer/hordp/foo\nbranches/developer/hordp/bar\netc.\n\nSince these were created with 'svn cp' originally, they would be\nproperly considered branches by your algorithm, right?    If so,\nsweet!\n\n> Having said that, here's a dodgy example I'd like to pre-emptively defend:\n>\n>        svn add tronk\n>        svn ci -m \"Created trunk\" # r1\n>        svn cp tronk trunk\n>        svn ci -m \"D'oh\" # r2\n>        svn rm tronk\n>        svn add trunk/markerFile.txt\n>        svn ci -m \"Double d'oh!\" # r3\n>\n> You could argue that the correct branch history description for the\n> above would be:\n>\n>        In r3, create branch \"trunk\"\n>\n> In other words, ignore everything that happened before the marker file\n> was created.  However, I would argue the following representation is\n> more correct:\n>\n>        In r1, create branch \"tronk\"\n>        In r2, create branch \"trunk\" from \"tronk\" r1\n>        In r3, delete branch \"tronk\"\n>\n\nI prefer your interpretation. It doesn't look dodgy at all.\n\nPhil\n"},{"id":"186353","messageId":"CA+7g9JzETuynGMCRo1MLuNErFiFc3AmhGS6Hr+jO-hoV2j4JDg@mail.gmail.com","threadId":"29821","inReplyTo":"4F5690FB.9060800@pileofstuff.org","subject":"Re: Approaches to SVN to Git conversion","fromName":"Nathan Gray","fromEmail":"n8gray@n8gray.org","sentAt":"2012-03-07T23:08:20Z","receivedAt":"2012-03-07T23:08:20Z","isPatch":false,"sender":{"key":"n8gray@n8gray.org","avatar":"https://avatars.githubusercontent.com/u/82794?v=4"},"body":"On Tue, Mar 6, 2012 at 2:34 PM, Andrew Sayers\n<andrew-git@pileofstuff.org> wrote:\n[snip]\n> On 06/03/12 19:29, Nathan Gray wrote:\n> <snip>\n>>\n>> The problem of specifying and detecting branches is a major problem in\n>> my upcoming conversion.  We've got toplevel trunk/branches/tags\n>> directories but underneath \"branches\" it's a free-for-all:\n>>\n>> /branches/codenameA/{projectA,projectB,projectC}\n>> /branches/codenameB   (actually a branch of projectA)\n>> /branches/developers/joe/frobnicator-experiment (also a branch of projectA)\n>>\n>> Clearly there's no simple regex that's going to capture this, so I'm\n>> reduced to listing every branch of projectA, which is tedious and\n>> error-prone.  However, what *would* work fabulously well for me is\n>> \"marker file\" detection.  Every copy of projectA has a certain file at\n>> it's root.  Let's call it \"markerFile.txt\".  What I'd really love is a\n>> way to say:\n>\n> This is quite close to the implementation I've got.  The SVN exporter\n> runs in two stages:\n>\n> In the first stage, the script treats any non-blacklisted file as a\n> marker file, but only looks for trunk branches.  It looks all through\n> the history, traces back through the copyfroms, and tries to find the\n> original directory associated with the file.  Usually it decides that\n> the only branch without a copyfrom is /trunk.  Searching just for trunks\n> with this weak heuristic makes it much easier to hand-verify the result.\n\nI'm not sure I understand.  So if I have /trunk/projectA and\n/trunk/projectB then do I have to blacklist /trunk/projectB to extract\nonly projectA's history?  Assuming it's always lived there will your\ncode detect /trunk/projectA as the \"trunk?\"  Would it be possible to\nspecify /trunk/projectA directly instead of blacklisting everything\nelse?\n\n> In the second stage, the script looks through the history again, tracing\n> the copies of known branches in a slightly less clever way than\n> described in my previous e-mail.  There's no need for marker files this\n> time round, as we just assume any `svn cp /trunk\n> /directory/not/within/a/branch` is a new branch.  In my experiments this\n> has been a pretty solid way of detecting branches without too much human\n> input - I might be missing something (or have mis-explained something),\n> but I'd be interested to hear examples of where this would go wrong.\n\nThat sounds pretty good, but it should probably also be transitive,\ni.e. `svn cp /any/known/branch/root /some/new/path` is also a new\nbranch.  Sometimes we'll spin off hotfix branches from release\nbranches, for example.\n\nI'll have to give your code a try and see how it works.\n\nCheers,\n-n8\n\n-- \nhttp://n8gray.org\n"},{"id":"186355","messageId":"4F57EC04.8060705@vilain.net","threadId":"29821","inReplyTo":"4F57DBF0.4060101@pileofstuff.org","subject":"Re: [spf:guess,iffy] Re: [spf:guess] Re: Approaches to SVN to Git conversion (was: Re: [RFC] \"Remote helper for Subversion\" project)","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2012-03-07T23:15:16Z","receivedAt":"2012-03-07T23:15:16Z","isPatch":false,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"On 3/7/12 2:06 PM, Andrew Sayers wrote:\n> It sounds like we've approached two similar problems in similar ways, so\n> I'm curious about the differences where they exist.  I've been reading\n> this message of yours from 18 months ago alongside this thread:\n> http://article.gmane.org/gmane.comp.version-control.git/150007\n> Unfortunately these comprise everything I know about Perforce.\n\nRight, I went into more detail back then than I did with my more recent \nmessage.\n\n> I notice that git-p4raw stores all of its data in Postgres and provides\n> a programmatic interface for querying it, whereas I've focussed on\n> providing ASCII interfaces at relevant points.  I can see how a DB store\n> would help manage the amount of data you'd need to process in a big\n> repository, but were there any other issues that drove you down this\n> route?  Did you consider a text-based interface?\n\nI wrote it like this mostly because the source metadata was already in a \ntabular form.  It allowed me to load the data, and then convert \ndeductions I could make of the data into unique and foreign key \nconstraints.  It provided me with ACID semantics to make it so that if \nmy program ran and failed the changes would not be applied.  Despite the \npopular opinion of \"web–scale\" technologists, databases do have large \nadvantages over unstructured hierarchical data :-).\n\nI didn't really intend to provide a programmatic interface, that was a \nset of user tools.  The SQL store is the programmatic interface :)\n\n>> What I did for the Perl Perforce conversion is make this a multi–step\n>> process; first, the heuristic goes through and detects branches and\n>> merge parents.  Then you do the actual export.  If, however, the\n>> heuristic gets it wrong, then you can manually override the branch\n>> detection for a particular revision, which invalidates all of the\n>> _automatic_ decisions made for later revisions the next time you run it.\n>\n> Could you give an example of overriding branch/merge detection?  It\n> sounds like you're saying that if there's some problem detecting merge\n> parents in an early revision, then all future merges are ignored by the\n> script.\n\nThe wrong decision can make things much worse down the line.  With the \nPerl history, the repository was about 350MB of pack, until I got the \nmerge history correct.  Afterwards, it packed down to about 70MB.  This \nis because there was a lot of criss–cross merging, and by marking them \ncorrectly git's repack algorithm was more able to locate similar blobs \nand compress correctly.  The pack size was not the goal, but a good \nverification that I had brought the correct commits together in history.\n\nThe bigger problems with it range from thinking changes are merged in \nyour branch which weren't really, or depending on how branch detection \netc works, getting thrown off completely and emitting garbage branch \nhistories.  So, it does help to be able to \"rewind\" the heuristics, poke \ninformation in and then resume again and see if things are improved. \nThe information could be inserted into a single file which has \nconfigured the entire import, and also serves as a set of notes as to \nthe amendments carried out.  I was happy with a database dump :-).\n\n> <snip>\n>> The manual input is extremely useful for bespoke conversions; there will\n>> always be warts in the history and no heuristic is perfect (even if you\n>> can supply your own set of expressions, a way to override it for just\n>> one revision is handy).\n>\n> Again, would you mind providing a few examples?  It sounds like you have\n> some edge cases that could be handled by extending the branch history\n> format, but I'd like to pin it down a bit more before discussing solutions.\n\nThere's a few,\n\n* a branch contains a subproject and is merged into a subtree\n* someone puts a \"README\" or similar file in a funny place, which isn't \ninside a project root\n* someone starts a project with no files in its root directory\n* someone records a merge incorrectly (or using a young or middle–aged \nSVN which didn't record merges).  You don't want your annotate to hit a \nmerge commit which isn't recorded as a merge, and then have to go \nhunting around in history for the real origin of a line of code\n* the piecemeal merge case you have seen yourself.\n\nIt's just very useful to be able to reparent during the data mining stage.\n\n> <snip>\n>>    3. skip bad sections of history, for instance squash merging merges\n>> which happened over several commits (SVN and Perforce, of course,\n>> support insane piecemeal merging prohibited by git)\n>\n> This is an excellent point I've stumbled past in my experiments without\n> realising what I was seeing.  A simple SVN example might look like this:\n>\n> \tsvn add trunk branches\n> \tsvn add trunk/foo trunk/bar\n> \tsvn ci -m \"Initial revision\" # r1\n>\n> \tsvn cp trunk branches/my_branch\n> \tsvn ci -m \"Created my_branch\" # r2\n>\n> \t# edit files in my_branch\n>\n> \tsvn merge branches/my_branch/foo trunk/foo\n> \tsvn ci -m \"Merge my_branch ->  trunk (1/3)\" # r11\n>\n> \tsvn merge branches/my_branch/bar trunk/bar\n> \tsvn ci -m \"Merge my_branch ->  trunk (2/3)\" # r12\n>\n> \tsvn cp branches/my_branch/new_file trunk/new_file\n> \tsvn ci -m \"Merge my_branch ->  trunk (3/3)\" # r13\n>\n> This strikes me as a sensibly cautious workflow in SVN, where merge\n> conflicts are common and changes are hard to revert.  The best\n> representation for this in the current branch history format would be\n> something like this:\n>\n> \tIn r1, create branch \"trunk\"\n> \tIn r2, create branch \"branches/my_branch\" from \"trunk\"\n> \tIn r13, merge \"branches/my_branch\" r13 into \"trunk\"\n>\n> In other words, pretend r11 and r12 are just normal commits, and that\n> r13 is a full merge.  A more useful (and arguably more accurate)\n> representation would be possible if we extended the format a bit:\n>\n> \tIn r1, create branch \"trunk\"\n> \tIn r2, create branch \"branches/my_branch\" from \"trunk\"\n> \tIn r12, squash changes in \"branches/my_branch\"\n> \tIn r13, squash changes in \"branches/my_branch\"\n> \tIn r13, merge \"branches/my_branch\" r13 into \"trunk\"\n>\n> Adding \"squash\" and \"fixup\" commands would let us represent the whole\n> messy business as a single commit, which is closer to what the user was\n> trying to say even if it's further from what they actually had to say.\n\nRight, you see the problem.\n\nI think your text syntax is fine so long as it is precise enough, and \nsimilar to what I mention earlier in this e–mail with having a single \nfile to drive a conversion run.  That really is the kind of input data \nthat I had, it's just that I set it up as a useful set of commands.\n\nSam.\n"},{"id":"186356","messageId":"4F57F01C.8080400@pileofstuff.org","threadId":"29821","inReplyTo":"CA+7g9JzETuynGMCRo1MLuNErFiFc3AmhGS6Hr+jO-hoV2j4JDg@mail.gmail.com","subject":"Re: Approaches to SVN to Git conversion","fromName":"Andrew Sayers","fromEmail":"andrew-git@pileofstuff.org","sentAt":"2012-03-07T23:32:44Z","receivedAt":"2012-03-07T23:32:44Z","isPatch":false,"sender":{"key":"andrew-git@pileofstuff.org","avatar":null},"body":"On 07/03/12 23:08, Nathan Gray wrote:\n<snip>\n> \n> I'm not sure I understand.  So if I have /trunk/projectA and\n> /trunk/projectB then do I have to blacklist /trunk/projectB to extract\n> only projectA's history?  Assuming it's always lived there will your\n> code detect /trunk/projectA as the \"trunk?\"  Would it be possible to\n> specify /trunk/projectA directly instead of blacklisting everything\n> else?\n\nPlease do try it, but the process should go something like this for you:\n\n1. run the SVN export \"configure\" stage - this reads through your repo\n   and suggests two trunks - \"/trunk/projectA\" and \"/trunk/projectB\".\n   You can explicitly ignore \"/trunk/projectB\", but at present there's\n   no way to ignore trunks by default.  No particular reason, I just\n   hadn't thought to add it :)\n\n2. run the SVN export \"make\" stage - this looks through whichever\n   trunks you've specified, and tracks the branches coming from it.  I\n   didn't explain this correctly in my previous e-mail, but yes this is\n   transitive - branches from branches from branches from trunk are\n   tracked in the appropriate way.\n\n3. edit the file created in stage 2.  If you wanted to ignore a\n   specific branch from (a branch from...) trunk/projectA, your best\n   bet is to exercise your text-fu on this file\n\n4. Import the history into git\n\nI'll be interested to hear how you and Phil get on, as it sounds like\nyes this approach should work for both of your repos.\n\n\t- Andrew\n"},{"id":"186493","messageId":"4F591BE0.4070709@pileofstuff.org","threadId":"29821","inReplyTo":"4F57EC04.8060705@vilain.net","subject":"Re: [spf:guess,iffy] Re: [spf:guess] Re: Approaches to SVN to Git conversion (was: Re: [RFC] \"Remote helper for Subversion\" project)","fromName":"Andrew Sayers","fromEmail":"andrew-git@pileofstuff.org","sentAt":"2012-03-08T20:51:44Z","receivedAt":"2012-03-08T20:51:44Z","isPatch":false,"sender":{"key":"andrew-git@pileofstuff.org","avatar":null},"body":"Thanks - this has really helped my thoughts to crystalise.\nHere's my plan at this point:\n\n1. Create an \"SVN History description\" project\n\nThis will build on the ASCII format I've been proposing so far.  The\ngoal will be to produce a human- and machine-readable format that\ndescribes SVN history in terms of an idealised version control system;\nand to produce a set of tests that any SVN history exporter can use as a\ntesting framework.\n\n2. Create an SVN history exporter\n\nThis will build on the svn-branch-export.pl script I previously made\navailable.  That script ran in exactly two passes (\"configure\" and\n\"make\") which wrote pointlessly different file formats.  The new script\nwill accept an SVN history file as input and create another SVN history\nfile as output, allowing users to iteratively improve the file as Sam\ndescribed.\n\n3. Create an SVN history importer for git\n\nThis will resemble the git-branch-import.pl script I previously made\navailable, but written in C based on the final SVN history format.  This\nthread has convinced me this would be a nice little SoC project, and\nI'll propose it in another thread if I've got project 1 to a reasonable\nstate before it's too late.  Failing that, and if nobody else wants to\ntake this project, I'll have a go myself some day when project 2 is\napproaching completion.\n\nMy next step will be to write up the SVN history work thrown up by this\nthread.  I'll come back to the list for advice when I've got something\npresentable.\n\n\t- Andrew\n"},{"id":"187811","messageId":"CALkWK0mMHjYpmxckanoPbqD=TrKR-i+7SjBEf2-8cV0kvPx0qA@mail.gmail.com","threadId":"29821","inReplyTo":"20120304075424.GI14725@burratino","subject":"Re: [RFC] \"Remote helper for Subversion\" project","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2012-03-27T03:58:52Z","receivedAt":"2012-03-27T03:58:52Z","isPatch":false,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Hi,\n\nJonathan Nieder wrote:\n> David Barr wrote:\n>> On Sat, Mar 3, 2012 at 11:27 PM, David Barr <davidbarr@google.com> wrote:\n>>> +\n>>> +* Getting an Git-to-SVN converter merged.\n>\n> Probably could fill a summer in itself.  In previous starts I think\n> there was some complexity creep. :/\n>\n>  http://thread.gmane.org/gmane.comp.version-control.git/170290\n>  http://thread.gmane.org/gmane.comp.version-control.git/170551\n\nI've been meaning to finish this off for sometime now- probably as an\nSoC project this summer?\n\n>>> +\n>>> +* Building the remote helper itself.\n>>> +\n>>> +Goal: Build a full-featured bi-directional `git-remote-svn` and get it\n>>> +      merged into upstream Git.\n>\n> Sure would be neat. ;-)  Another nice piece to build would be branch\n> tracking / follow_parent heuristics.\n\nThis doesn't sound awfully complicated; if the Git -> SVN converter\ngets merged early on, I might get a chance to work on this as well.\n\n    Ram\n"}]}