{"thread":{"id":"26264","subject":"working with a large repository and git svn","startedAt":"2011-01-12T01:27:10Z","lastAt":"2011-01-16T03:32:57Z","messageCount":12,"participants":["Joe Corneli","Wesley J. Landaker","Jonathan Nieder","Ramkumar Ramachandra","Michael Haggerty"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"159392","messageId":"AANLkTimKbS3ECzOaGtNgvx7DThJGH_DkPmg4ehKXGtwc@mail.gmail.com","threadId":"26264","inReplyTo":null,"subject":"working with a large repository and git svn","fromName":"Joe Corneli","fromEmail":"holtzermann17@gmail.com","sentAt":"2011-01-12T01:27:10Z","receivedAt":"2011-01-12T01:27:10Z","isPatch":false,"sender":{"key":"holtzermann17@gmail.com","avatar":"https://gravatar.com/avatar/64bf524b41adf9303ce90e13a7eba2c3927aea526b9dc016d96ede8e03bc65f5?d=mp&s=160"},"body":"Greetings -\n\nI am experiencing trouble with git svn, trying to import a\nlarge repository (7.9 gigs, ~54000 commits) from Git into\nSVN.\n\nThis has failed in a couple of different ways, depending\non the operating environment.  With Git version 1.7.3.5\nrunning on Ubuntu 9.10, in the final step\n\n  git svn dcommit --no-rebase\n\nof the formula described below, I get:\n\n failing with \"Can't fork at /usr/share/perl/5.10.0/Git.pm line 1261.\"\n\nafter committing just over 2000 revisions.\n\nPreviously, on Mac OS X 10.6.4 with git version 1.7.3.4,\nit made it through about 18000 commits before failing with\nsome other error.  (I don't have that one recorded at the\nmoment.)\n\nSeparately from the latest attempt, I tried repacking the\nrepository before doing the \"git svn\" stuff, with\n\n  git repack -a -d --depth=250 --window=250 -f\n\nbut that also failed (\"pack-objects died of signal 11\").\n\nAny tips for dealing with new, large, repositories would\nbe appreciated.  The sequence of commands I used are below\nthe 8<.\n\nThanks,\nJoe\n\n8<-8<-8<-8<-8<-8<-8<-8<-8<-8<-8<-8<-8<-8<-8<-8<-8<-8<-8<-8<-8<\n\n## Creating an svn repo\n\n$ mkdir repo;\n$ svnadmin create repo;\n$ mkdir init;\n$ touch init/README;\n$ svn import init file://`pwd`/repo/init -m \"Initial import\";\n$ svn checkout file://`pwd`/repo/init working;\n\n## THIS PART FOLLOWS THE MODEL SUGGESTED BY THE FOLKS AT code.google.com\n## (/Users/jac2349/planetary/destination/ IS THE LOCATION OF MY GIT REPO.)\n\n$ mkdir cloning\n$ cd cloning\n$ git svn clone file:///Users/jac2349/planetary/repo/init\n$ cd init\n$ git fetch git:///Users/jac2349/planetary/destination/.git\n\n$ git branch tmp $(cut -b-40 .git/FETCH_HEAD)\n$ git tag -a -m \"Last fetch\" last tmp\n\n$ INIT_COMMIT=$(git log tmp --pretty=format:%H | tail -1)\n$ git checkout $INIT_COMMIT .\n$ git commit -C $INIT_COMMIT\n\n$ git rebase master tmp\n$ git branch -M tmp master\n\n$ git svn dcommit --no-rebase\n\n$ mv .git/refs/tags/newlast .git/refs/tags/last\n\n## BTW, THE --no-rebase FLAG KEEPS IT FROM BEING IMPOSSIBLY SLOW!\n"},{"id":"159404","messageId":"201101120830.47016.wjl@icecavern.net","threadId":"26264","inReplyTo":"AANLkTimKbS3ECzOaGtNgvx7DThJGH_DkPmg4ehKXGtwc@mail.gmail.com","subject":"Re: working with a large repository and git svn","fromName":"Wesley J. Landaker","fromEmail":"wjl@icecavern.net","sentAt":"2011-01-12T15:30:45Z","receivedAt":"2011-01-12T15:30:45Z","isPatch":false,"sender":{"key":"wjl@icecavern.net","avatar":"https://avatars.githubusercontent.com/u/67229?v=4"},"body":"On Tuesday, January 11, 2011 18:27:10 Joe Corneli wrote:\n> I am experiencing trouble with git svn, trying to import a\n> large repository (7.9 gigs, ~54000 commits) from Git into\n> SVN.\n> \n> This has failed in a couple of different ways, depending\n> on the operating environment.  With Git version 1.7.3.5\n> running on Ubuntu 9.10, in the final step\n> \n>   git svn dcommit --no-rebase\n> \n> of the formula described below, I get:\n> \n>  failing with \"Can't fork at /usr/share/perl/5.10.0/Git.pm line 1261.\"\n> \n> after committing just over 2000 revisions.\n\nI haven't tried importing 8 GB from Git to Subversion, but I have used Git \nagainst existing huge Subversion repositories that are >= 10 GB with little \ntrouble, other than that it takes forever because Subversion is slow.\n\nHere are some thoughts on how I'd approach what you are doing. Realize that \nno matter what, it's still probably going to take \"forever\" (e.g. run it \nover the weekend).\n\n  1) Sounds like git-svn is running out of resources on your machine -- \nthat's probably a bug, but work around it: Don't dcommit all 20000 revisions \nat once. Maybe write a shell script that goes through and dcommits a 100 \ncommits at a time.\n\n  2) Do you need the full history to be in SVN? Can you rebase/squash large \nparts together and thus need to commit less revisions in the first place?\n\n  3) I love git-svn for working with Subversion repositories, but you could \nconsider a different tool, like tailor, if you can't make git-svn do what \nyou want. I have also heard talk (but I don't know the state of things) of \npeople working on a fast-import tool for SVN, so you could git-fast-export \nand svn-fast-import in a big batch.\n\n  4) Does 8 GB of data really belong in the same repository? Maybe it should \nreally be split up and used with git submodules or SVN externals? That may \nmake things easier to work with in the long term.\n\n  5) Do you really want to be going from Git, to Subversion? That seems like \na big step backwards. =)\n\nIn any case, good luck!\n"},{"id":"159418","messageId":"AANLkTi=uuBuunYmwmLYD_vUnPkDBk9YDLtATw9GtX33z@mail.gmail.com","threadId":"26264","inReplyTo":"201101120830.47016.wjl@icecavern.net","subject":"Re: working with a large repository and git svn","fromName":"Joe Corneli","fromEmail":"holtzermann17@gmail.com","sentAt":"2011-01-13T00:54:27Z","receivedAt":"2011-01-13T00:54:27Z","isPatch":false,"sender":{"key":"holtzermann17@gmail.com","avatar":"https://gravatar.com/avatar/64bf524b41adf9303ce90e13a7eba2c3927aea526b9dc016d96ede8e03bc65f5?d=mp&s=160"},"body":">  1) Sounds like git-svn is running out of resources on your machine --\n> that's probably a bug, but work around it: Don't dcommit all 20000 revisions\n> at once. Maybe write a shell script that goes through and dcommits a 100\n> commits at a time.\n\nHm, I found a related blog post here, but designed for interactive use:\nhttp://fredericiana.com/2009/12/31/partial-svn-dcommit-with-git/\nCould you give me a more detailed hint about how to do what you suggested?\n\n>  2) Do you need the full history to be in SVN? Can you rebase/squash large\n> parts together and thus need to commit less revisions in the first place?\n\nMaybe.  We want a tool for managing the entire history, and Git seems\nlike a good tool for that.  At the same time, checking out the entire\nhistory can take a long time - if we could just check out just the\nlatest files and check them back in in a sensible way, that would be\ngood - SVN does seem suitable for that purpose.  If there's a git-only\nway to do this I'd be happy to know about that as well!\n\n>  3) I love git-svn for working with Subversion repositories, but you could\n> consider a different tool, like tailor, if you can't make git-svn do what\n> you want.\n\nTried it, but it didn't even get through the initiation phase.  I\nasked for help in the relevant mailing list.\n\n> people working on a fast-import tool for SVN, so you could git-fast-export\n> and svn-fast-import in a big batch.\n\nNot finding these.\n\n>  4) Does 8 GB of data really belong in the same repository? Maybe it should\n> really be split up and used with git submodules or SVN externals? That may\n> make things easier to work with in the long term.\n\nProbably true.  if there was a nice way to give each *file* its own\nassociated \"repository\", then stitch these together into packets (even\n\"on demand\"), that would be cool.  I was assuming we could do fancy\nstuff like this as \"future work\" however - and it would seem that if\nwe use a completely git-based solution we'll be there.\n\n>  5) Do you really want to be going from Git, to Subversion? That seems like\n> a big step backwards. =)\n\nIf there's a good way to just pull down the latest revision into a\nworking copy and be able to push that back to the repo that would be\nnice.  This doesn't seem to be the Git way, but for an 8 gig repo it's\nprobably pretty important feature.  Thoughts?\n\nThanks,\nJoe\n"},{"id":"159423","messageId":"20110113032300.GB9184@burratino","threadId":"26264","inReplyTo":"201101120830.47016.wjl@icecavern.net","subject":"Re: working with a large repository and git svn","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2011-01-13T03:23:00Z","receivedAt":"2011-01-13T03:23:00Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Wesley J. Landaker wrote:\n\n>   3) I love git-svn for working with Subversion repositories, but you could \n> consider a different tool, like tailor, if you can't make git-svn do what \n> you want. I have also heard talk (but I don't know the state of things) of \n> people working on a fast-import tool for SVN, so you could git-fast-export \n> and svn-fast-import in a big batch.\n\nI think the state of the art is currently git2svn[1] + \"svnrdump load\".\nThis requires permission to change properties on the svn repo, just\nlike svnsync would.\n\nHope that helps,\nJonathan\n\n[1] http://repo.or.cz/w/git2svn.git\n"},{"id":"159461","messageId":"AANLkTikCvjDqUpL-=srVKcMQx+NM6bV7FabmJ+4sPqD7@mail.gmail.com","threadId":"26264","inReplyTo":"20110113032300.GB9184@burratino","subject":"Re: working with a large repository and git svn","fromName":"Joe Corneli","fromEmail":"holtzermann17@gmail.com","sentAt":"2011-01-14T07:43:19Z","receivedAt":"2011-01-14T07:43:19Z","isPatch":false,"sender":{"key":"holtzermann17@gmail.com","avatar":"https://gravatar.com/avatar/64bf524b41adf9303ce90e13a7eba2c3927aea526b9dc016d96ede8e03bc65f5?d=mp&s=160"},"body":"> I think the state of the art is currently git2svn\n\nThanks, that did indeed work, though, for the record it uses committer\nname and email in the log that it generates, not author name and\nemail, but no worries!\n\nJoe\n"},{"id":"159463","messageId":"20110114080554.GA1735@kytes","threadId":"26264","inReplyTo":"AANLkTikCvjDqUpL-=srVKcMQx+NM6bV7FabmJ+4sPqD7@mail.gmail.com","subject":"Re: working with a large repository and git svn","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2011-01-14T08:05:57Z","receivedAt":"2011-01-14T08:05:57Z","isPatch":false,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Hi Joe,\n\nJoe Corneli writes:\n> > I think the state of the art is currently git2svn\n> \n> Thanks, that did indeed work, though, for the record it uses committer\n> name and email in the log that it generates, not author name and\n> email, but no worries!\n\nThat should be easy enough to fix with something like this (warning:\nuntested). A more elegant solution would actually use some sort of\nuser-configurable mapping from Git authors/ committers to SVN authors\nthough.\n\nSigned-off-by: Ramkumar Ramachandra <artagnon@gmail.com>\n--8<--\ndiff --git a/git2svn b/git2svn\nindex 2380775..3856696 100755\n--- a/git2svn\n+++ b/git2svn\n@@ -261,12 +261,8 @@ COMMAND: while (!eof(IN)) {\n \t    $commit{Mark} = $1;\n \t    $next = next_line($IN);\n \t}\n-\tif ($next =~ m/author +(.*)/) {\n-\t    $commit{Author} = $1;\n-\t    $next = next_line($IN);\n-\t}\n-\tunless ($next =~ m/committer +(.+) +<([^>]+)> +(\\d+) +[+-](\\d+)$/) {\n-\t    die \"missing comitter: $_\";\n+\tunless ($next =~ m/author +(.+) +<([^>]+)> +(\\d+) +[+-](\\d+)$/) {\n+\t    die \"missing author: $_\";\n \t}\n \n \t$commit{CommitterName} = $1;\n@@ -275,6 +271,9 @@ COMMAND: while (!eof(IN)) {\n \t$commit{CommitterTZ} = $4;\n \n \t$next = next_line($IN);\n+\tif ($next =~ m/committer +(.*)/) {\n+\t    $next = next_line($IN);\n+\t}\n \tmy $log = read_data($IN, $next);\n \n \t$next = next_line($IN);\n"},{"id":"159464","messageId":"20110114082931.GC11343@burratino","threadId":"26264","inReplyTo":"20110114080554.GA1735@kytes","subject":"Re: working with a large repository and git svn","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2011-01-14T08:29:31Z","receivedAt":"2011-01-14T08:29:31Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Ramkumar Ramachandra wrote:\n> Joe Corneli writes:\n\n>>> I think the state of the art is currently git2svn\n>>\n>> Thanks, that did indeed work, though, for the record it uses committer\n>> name and email in the log that it generates, not author name and\n>> email, but no worries!\n>\n> That should be easy enough to fix with something like this (warning:\n> untested). A more elegant solution would actually use some sort of\n> user-configurable mapping from Git authors/ committers to SVN authors\n> though.\n\nThanks for the cc.  (cc-ing lha, as I should have before.)\n\nI suppose if svn will show only one of the two (committer and author)\nthen it is better to show the author.  Possible complications:\n\n. The author lines in fast-import streams are optional.\n\n. Existing users of the incremental import facility might not want the\n  meaning of svn:author to change between imports.  _If_ that is a\n  problem then a command-line option to switch behaviors might help.\n\n. Is svn okay with non-monotonic dates?  (If not, then the committer\n  date would need to be used.)\n\nModulo those complications I like the idea.  (Though I haven't read\nthe implementation, which follows for reference.)\n\n> \n> Signed-off-by: Ramkumar Ramachandra <artagnon@gmail.com>\n> --8<--\n> diff --git a/git2svn b/git2svn\n> index 2380775..3856696 100755\n> --- a/git2svn\n> +++ b/git2svn\n> @@ -261,12 +261,8 @@ COMMAND: while (!eof(IN)) {\n>  \t    $commit{Mark} = $1;\n>  \t    $next = next_line($IN);\n>  \t}\n> -\tif ($next =~ m/author +(.*)/) {\n> -\t    $commit{Author} = $1;\n> -\t    $next = next_line($IN);\n> -\t}\n> -\tunless ($next =~ m/committer +(.+) +<([^>]+)> +(\\d+) +[+-](\\d+)$/) {\n> -\t    die \"missing comitter: $_\";\n> +\tunless ($next =~ m/author +(.+) +<([^>]+)> +(\\d+) +[+-](\\d+)$/) {\n> +\t    die \"missing author: $_\";\n>  \t}\n>  \n>  \t$commit{CommitterName} = $1;\n> @@ -275,6 +271,9 @@ COMMAND: while (!eof(IN)) {\n>  \t$commit{CommitterTZ} = $4;\n>  \n>  \t$next = next_line($IN);\n> +\tif ($next =~ m/committer +(.*)/) {\n> +\t    $next = next_line($IN);\n> +\t}\n>  \tmy $log = read_data($IN, $next);\n>  \n>  \t$next = next_line($IN);\n"},{"id":"159471","messageId":"4D30162F.5060408@alum.mit.edu","threadId":"26264","inReplyTo":"20110114082931.GC11343@burratino","subject":"Re: working with a large repository and git svn","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2011-01-14T09:23:59Z","receivedAt":"2011-01-14T09:23:59Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 01/14/2011 09:29 AM, Jonathan Nieder wrote:\n> . Is svn okay with non-monotonic dates?  (If not, then the committer\n>   date would need to be used.)\n\nSubversion can tolerate non-monotonic dates with one caveat: it breaks\nthe find-revision-by-date feature (e.g., \"svn update -r '{2010-12-25}'\")\nfor the time intervals with non-monotonic dates.  This is a seldom-used\nfeature and therefore its sacrifice is often accepted, for example when\nthe history of the Subversion project itself was migrated into the\nApache project's Subversion repository.\n\nMichael\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"159475","messageId":"20110114101636.GA22970@kytes","threadId":"26264","inReplyTo":"F0299861-B36C-459C-972E-856212A92615@kth.se","subject":"[PATCH] Optionally parse author information","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2011-01-14T10:16:38Z","receivedAt":"2011-01-14T10:16:38Z","isPatch":true,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"When creating a new commit, instead of picking up the SVN author from\nthe committer's email, pick it up from the author's email, when\npossible. Also add a new command-line switch '--ignore-author' to\nforce older behavior for backward compatibilty.\n\nNoticed-by: Joe Corneli <holtzermann17@gmail.com>\nSigned-off-by: Ramkumar Ramachandra <artagnon@gmail.com>\n---\n git2svn |   25 +++++++++++++++++++------\n 1 files changed, 19 insertions(+), 6 deletions(-)\n\ndiff --git a/git2svn b/git2svn\nindex 2380775..8ef55f1 100755\n--- a/git2svn\n+++ b/git2svn\n@@ -36,7 +36,7 @@ use Pod::Usage;\n my $IN;\n my $OUT;\n \n-my ($help, $verbose, $keeplogs, $no_load);\n+my ($help, $verbose, $keeplogs, $no_load, $ignore_author);\n \n # svn\n my $svntree = \"repro\";\n@@ -200,6 +200,7 @@ $result = GetOptions (\"git-branch=s\" => \\$branch,\n \t\t      \"svn-prefix=s\" => \\$basedir,\n \t\t      \"keep-logs\" => \\$keeplogs,\n \t\t      \"no-load\" => \\$no_load,\n+\t\t      \"ignore-author\" => \\$ignore_author,\n \t\t      \"verbose+\" => \\$verbose,\n \t\t      \"help\" => \\$help) or pod2usage(2);\n \n@@ -261,12 +262,15 @@ COMMAND: while (!eof(IN)) {\n \t    $commit{Mark} = $1;\n \t    $next = next_line($IN);\n \t}\n-\tif ($next =~ m/author +(.*)/) {\n-\t    $commit{Author} = $1;\n+\tif ($next =~ m/author +(.+) +<([^>]+)> +(\\d+) +[+-](\\d+)$/) {\n+\t    $commit{AuthorName} = $1;\n+\t    $commit{AuthorEmail} = $2;\n+\t    $commit{AuthorWhen} = $3;\n+\t    $commit{AuthorTZ} = $4;\n \t    $next = next_line($IN);\n \t}\n \tunless ($next =~ m/committer +(.+) +<([^>]+)> +(\\d+) +[+-](\\d+)$/) {\n-\t    die \"missing comitter: $_\";\n+\t    die \"missing committer: $_\";\n \t}\n \n \t$commit{CommitterName} = $1;\n@@ -291,11 +295,15 @@ COMMAND: while (!eof(IN)) {\n \t    strftime(\"%Y-%m-%dT%H:%M:%S.000000Z\", \n \t\t     gmtime($commit{CommitterWhen}));\n \n-\tmy $author = \"(no author)\";\n+\tmy $author = \"git2svn-dump\";\n \tif ($commit{CommitterEmail} =~ m/([^@]+)/) {\n \t    $author = $1;\n \t}\n-\t$author = \"git2svn-dump\" if ($author eq \"(no author)\");\n+\tunless ($ignore_author) {\n+\t    if ($commit{AuthorEmail} =~ m/([^@]+)/) {\n+\t        $author = $1;\n+\t    }\n+\t}\n \n \tmy $props = \"\";\n \t$props .= prop(\"svn:author\", $author);\n@@ -486,6 +494,11 @@ match the default GIT branch (master).\n \n Don't load the svn repository or update the syncpoint tagname.\n \n+=item B<--ignore-author>\n+\n+Ignore \"author\" lines in the fast-import stream. Use \"committer\"\n+information instead.\n+\n =item B<--keep-logs>\n \n Don't delete the logs in $CWD/.data on success.\n-- \n1.7.4.rc1.7.g2cf08.dirty\n"},{"id":"159538","messageId":"AANLkTi=ddJYT8YiUDYy80xobkxJnvuREN-09=464P_vB@mail.gmail.com","threadId":"26264","inReplyTo":"20110114101636.GA22970@kytes","subject":"Re: [PATCH] Optionally parse author information","fromName":"Joe Corneli","fromEmail":"holtzermann17@gmail.com","sentAt":"2011-01-16T02:17:22Z","receivedAt":"2011-01-16T02:17:22Z","isPatch":true,"sender":{"key":"holtzermann17@gmail.com","avatar":"https://gravatar.com/avatar/64bf524b41adf9303ce90e13a7eba2c3927aea526b9dc016d96ede8e03bc65f5?d=mp&s=160"},"body":"I tested it, and it seems to use email handle instead of author name\n(perhaps that's intentional, though in my case it's not so desirable)\nbut, quite critically, it gets the dates wrong:\n\n~/pmhistory.svn$ svn log -l 5\n------------------------------------------------------------------------\nr53127 | majordomo | 2011-01-10 18:31:58 -0500 (Mon, 10 Jan 2011) | 1 line\n\n\n------------------------------------------------------------------------\nr53126 | majordomo | 2011-01-10 18:31:58 -0500 (Mon, 10 Jan 2011) | 1 line\n\n\n------------------------------------------------------------------------\nr53125 | majordomo | 2011-01-10 18:31:58 -0500 (Mon, 10 Jan 2011) | 1 line\n\n\n------------------------------------------------------------------------\nr53124 | majordomo | 2011-01-10 18:31:57 -0500 (Mon, 10 Jan 2011) | 1 line\n\n\n------------------------------------------------------------------------\nr53123 | majordomo | 2011-01-10 18:31:57 -0500 (Mon, 10 Jan 2011) | 1 line\n\n~/pmhistory.git$ git log -5\ncommit 411b8698e494ee12799300611fed0c8029e76ad3\nAuthor: milogardner <majordomo@planetmath.org>\nDate:   Thu Dec 16 14:11:57 2010 +0000\n\ncommit d12f8472cc06feec1a0e3a652e4ac14d7869fb3f\nAuthor: milogardner <majordomo@planetmath.org>\nDate:   Thu Dec 16 14:00:13 2010 +0000\n\ncommit 5febb4767563255280d95091ff9b2b0207042071\nAuthor: Mathprof <majordomo@planetmath.org>\nDate:   Wed Dec 15 23:02:47 2010 +0000\n\ncommit 61b11b97c4e503c353af5c1cd68e17b053d12b8e\nAuthor: pahio <majordomo@planetmath.org>\nDate:   Mon Dec 13 17:59:12 2010 +0000\n\ncommit 14d66e2dcd6151eb7214a9afbab159459912da6d\nAuthor: pahio <majordomo@planetmath.org>\nDate:   Mon Dec 13 17:54:21 2010 +0000\n"},{"id":"159540","messageId":"20110116025707.GB28452@burratino","threadId":"26264","inReplyTo":"AANLkTi=ddJYT8YiUDYy80xobkxJnvuREN-09=464P_vB@mail.gmail.com","subject":"Re: [PATCH] Optionally parse author information","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2011-01-16T02:57:07Z","receivedAt":"2011-01-16T02:57:07Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Joe Corneli wrote:\n\n> I tested it, and it seems to use email handle instead of author name\n> (perhaps that's intentional, though in my case it's not so desirable)\n\nGood point.  Presumably git2svn is using the local part of the email\naddress to mimic svn's default behavior of using one's username.\n\nOther possibilities:\n\n - email address (e.g., majordomo@planetmath.org, as in most google\n   repositories)\n - display name (e.g., \"Joe Corneli\").  I don't know if svn-related\n   tools or scripts assume that svn:author doesn't contain spaces.\n - full ident string (e.g., \"Joe Corneli <majordomo@planetmath.org>\")\n - whatever the operator wants (mapping specified in authors file).\n\nMy guess: an \"authors file\" facility would be needed to cover all\ncases, but whichever rule you want to implement short of that could\nalso be useful.\n\n> but, quite critically, it gets the dates wrong:\n> \n> ~/pmhistory.svn$ svn log -l 5\n> ------------------------------------------------------------------------\n> r53127 | majordomo | 2011-01-10 18:31:58 -0500 (Mon, 10 Jan 2011) | 1 line\n[...]\n> ~/pmhistory.git$ git log -5\n> commit 411b8698e494ee12799300611fed0c8029e76ad3\n> Author: milogardner <majordomo@planetmath.org>\n> Date:   Thu Dec 16 14:11:57 2010 +0000\n\nMaybe it is using the committer date (as shown by \"git log --format=fuller\")\nand someone rebased recently.  If you don't care about svn's '{date}'\nconstruct working (meaning out-of-order dates are ok) then author date\nmight be more suitable.  Presumably the important thing is for it to\nbe consistent.\n\nHope that helps,\nJonathan\n"},{"id":"159541","messageId":"20110116033253.GA22707@kytes","threadId":"26264","inReplyTo":"AANLkTi=ddJYT8YiUDYy80xobkxJnvuREN-09=464P_vB@mail.gmail.com","subject":"Re: [PATCH] Optionally parse author information","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2011-01-16T03:32:57Z","receivedAt":"2011-01-16T03:32:57Z","isPatch":true,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Hi Joe,\n\nJoe Corneli writes:\n> I tested it, and it seems to use email handle instead of author name\n> (perhaps that's intentional, though in my case it's not so desirable)\n> but, quite critically, it gets the dates wrong:\n\nYes. I didn't change it's core behavior- it used to extract the\ninformation from the committer's email address previously; I just\nchanged it to use the author's email address. For dates, it uses\ncommitter dates again, and this is probably desirable: author dates\naren't necessarily monotonic, and this can break some functionality in\nSVN.\n\nOfcourse, a lot more is possible with an patch that allows users to\nconfigure all these things. Until then, I recommend that you just edit\nthe source to achieve the desired results.\n\n-- Ram\n"}]}