{"thread":{"id":"15187","subject":"[RFC] mtn to git conversion script","startedAt":"2008-08-24T09:18:50Z","lastAt":"2008-11-11T16:40:46Z","messageCount":21,"participants":["Felipe Contreras","Miklos Vajna","Johannes Schindelin","Shawn O. Pearce","Brian Downing","Anand Kumria","Jakub Narebski","Thomas Moschny","Juan Jose Comellas"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"88355","messageId":"94a0d4530808240218j4bedbe3di99303da9addc93a4@mail.gmail.com","threadId":"15187","inReplyTo":null,"subject":"[RFC] mtn to git conversion script","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2008-08-24T09:18:50Z","receivedAt":"2008-08-24T09:18:50Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"Hi,\n\nI developed a script that converts a monotone repository into a git\none (exact clone), I want to contribute it so everybody can use it.\n\nHowever, I might have not done it correctly.\n\nThis is the gist of the script:\n\nmtn update --revision #{@id} --reallyquiet\ngit ls-files --modified --others --exclude-standard -z | git\nupdate-index --add --remove -z --stdin\ngit write-tree\ngit write-raw < /tmp/commit.txt\ngit update-ref refs/mtn/#{@id} #{@git_id}\n\nbranches.each do |e|\n    git update-ref refs/heads/#{e} #{@git_id}\nend\n\nI wrote \"git write-raw\" which takes the commit text as is, and puts it\ninto the repository.\n\nI've read about 'fast-import' but I'm not sure if it would be more\nefficient, because you would have to parse the output of different mtn\ntools.\n\nWhat do you think? Does it makes sense to have a 'write-raw' command?\nOr should I somehow use 'fast-import'?\n\nBest regards.\n\n-- \nFelipe Contreras\n"},{"id":"88361","messageId":"20080824131405.GJ23800@genesis.frugalware.org","threadId":"15187","inReplyTo":"94a0d4530808240218j4bedbe3di99303da9addc93a4@mail.gmail.com","subject":"Re: [RFC] mtn to git conversion script","fromName":"Miklos Vajna","fromEmail":"vmiklos@frugalware.org","sentAt":"2008-08-24T13:14:05Z","receivedAt":"2008-08-24T13:14:05Z","isPatch":false,"sender":{"key":"vmiklos@frugalware.org","avatar":"https://gravatar.com/avatar/401c1cbbb3a5d13e650c691a2c71d6fd0b80df1a01bc74d9f1972675dd58f2bd?d=mp&s=160"},"body":"On Sun, Aug 24, 2008 at 12:18:50PM +0300, Felipe Contreras <felipe.contreras@gmail.com> wrote:\n> What do you think? Does it makes sense to have a 'write-raw' command?\n> Or should I somehow use 'fast-import'?\n\nYes, you should. ;-)\n\nThe syntax of it is not so hard, see for example 'git fast-export\nHEAD~2..' on a git repo and you'll see.\n\nThis should help a lot if you are like me, who likes to learn from\nexamples.\n"},{"id":"88375","messageId":"alpine.DEB.1.00.0808242017520.24820@pacific.mpi-cbg.de.mpi-cbg.de","threadId":"15187","inReplyTo":"20080824131405.GJ23800@genesis.frugalware.org","subject":"Re: [RFC] mtn to git conversion script","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-08-24T18:19:23Z","receivedAt":"2008-08-24T18:19:23Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 24 Aug 2008, Miklos Vajna wrote:\n\n> On Sun, Aug 24, 2008 at 12:18:50PM +0300, Felipe Contreras <felipe.contreras@gmail.com> wrote:\n> > What do you think? Does it makes sense to have a 'write-raw' command? \n> > Or should I somehow use 'fast-import'?\n> \n> Yes, you should. ;-)\n> \n> The syntax of it is not so hard, see for example 'git fast-export\n> HEAD~2..' on a git repo and you'll see.\n> \n> This should help a lot if you are like me, who likes to learn from\n> examples.\n\nHeh.  I am glad you like fast-export.  To be honest, I never intended to \nuse fast-export for anything else than as an example how to drive \nfast-import... :-)\n\nCiao,\nDscho\n"},{"id":"88377","messageId":"94a0d4530808241133n5cc9f17arc79a1a5013187869@mail.gmail.com","threadId":"15187","inReplyTo":"20080824131405.GJ23800@genesis.frugalware.org","subject":"Re: [RFC] mtn to git conversion script","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2008-08-24T18:33:54Z","receivedAt":"2008-08-24T18:33:54Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Sun, Aug 24, 2008 at 4:14 PM, Miklos Vajna <vmiklos@frugalware.org> wrote:\n> On Sun, Aug 24, 2008 at 12:18:50PM +0300, Felipe Contreras <felipe.contreras@gmail.com> wrote:\n>> What do you think? Does it makes sense to have a 'write-raw' command?\n>> Or should I somehow use 'fast-import'?\n>\n> Yes, you should. ;-)\n>\n> The syntax of it is not so hard, see for example 'git fast-export\n> HEAD~2..' on a git repo and you'll see.\n>\n> This should help a lot if you are like me, who likes to learn from\n> examples.\n\nIs it possible to create a fast-import from the index? I realize this\nis not the best thing to do, but for now I would like to do that.\n\nBest regards.\n\n-- \nFelipe Contreras\n"},{"id":"88387","messageId":"20080824193732.GO23800@genesis.frugalware.org","threadId":"15187","inReplyTo":"alpine.DEB.1.00.0808242017520.24820@pacific.mpi-cbg.de.mpi-cbg.de","subject":"Re: [RFC] mtn to git conversion script","fromName":"Miklos Vajna","fromEmail":"vmiklos@frugalware.org","sentAt":"2008-08-24T19:37:32Z","receivedAt":"2008-08-24T19:37:32Z","isPatch":false,"sender":{"key":"vmiklos@frugalware.org","avatar":"https://gravatar.com/avatar/401c1cbbb3a5d13e650c691a2c71d6fd0b80df1a01bc74d9f1972675dd58f2bd?d=mp&s=160"},"body":"On Sun, Aug 24, 2008 at 08:19:23PM +0200, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> Heh.  I am glad you like fast-export.  To be honest, I never intended to \n> use fast-export for anything else than as an example how to drive \n> fast-import... :-)\n\nIt's much more, git-bzr's bi-directional operation would be impossible\nwithout it. ;-)\n"},{"id":"88404","messageId":"20080824224658.GA16590@spearce.org","threadId":"15187","inReplyTo":"94a0d4530808241133n5cc9f17arc79a1a5013187869@mail.gmail.com","subject":"Re: [RFC] mtn to git conversion script","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-24T22:46:59Z","receivedAt":"2008-08-24T22:46:59Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Felipe Contreras <felipe.contreras@gmail.com> wrote:\n> On Sun, Aug 24, 2008 at 4:14 PM, Miklos Vajna <vmiklos@frugalware.org> wrote:\n> > On Sun, Aug 24, 2008 at 12:18:50PM +0300, Felipe Contreras <felipe.contreras@gmail.com> wrote:\n> >> What do you think? Does it makes sense to have a 'write-raw' command?\n> >> Or should I somehow use 'fast-import'?\n> >\n> > Yes, you should. ;-)\n> >\n> > The syntax of it is not so hard, see for example 'git fast-export\n> > HEAD~2..' on a git repo and you'll see.\n> >\n> > This should help a lot if you are like me, who likes to learn from\n> > examples.\n> \n> Is it possible to create a fast-import from the index? I realize this\n> is not the best thing to do, but for now I would like to do that.\n\nNo, fast-import uses its own internal structure and avoids the\nindex file.\n\nAlso, look at `git-hash-objects -w` as a replacement for your\ngit-write-raw tool if you aren't going to use git-fast-import.\n\n-- \nShawn.\n"},{"id":"88412","messageId":"94a0d4530808241745r3f2bdb56q9cfa8bc61f79223e@mail.gmail.com","threadId":"15187","inReplyTo":"20080824224658.GA16590@spearce.org","subject":"Re: [RFC] mtn to git conversion script","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2008-08-25T00:45:11Z","receivedAt":"2008-08-25T00:45:11Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Mon, Aug 25, 2008 at 1:46 AM, Shawn O. Pearce <spearce@spearce.org> wrote:\n> Felipe Contreras <felipe.contreras@gmail.com> wrote:\n>> On Sun, Aug 24, 2008 at 4:14 PM, Miklos Vajna <vmiklos@frugalware.org> wrote:\n>> > On Sun, Aug 24, 2008 at 12:18:50PM +0300, Felipe Contreras <felipe.contreras@gmail.com> wrote:\n>> >> What do you think? Does it makes sense to have a 'write-raw' command?\n>> >> Or should I somehow use 'fast-import'?\n>> >\n>> > Yes, you should. ;-)\n>> >\n>> > The syntax of it is not so hard, see for example 'git fast-export\n>> > HEAD~2..' on a git repo and you'll see.\n>> >\n>> > This should help a lot if you are like me, who likes to learn from\n>> > examples.\n>>\n>> Is it possible to create a fast-import from the index? I realize this\n>> is not the best thing to do, but for now I would like to do that.\n>\n> No, fast-import uses its own internal structure and avoids the\n> index file.\n\nYeah, I knew that, but wanted to just replace the 'write-raw' command.\nTo avoid doing unnecessary changes.\n\n> Also, look at `git-hash-objects -w` as a replacement for your\n> git-write-raw tool if you aren't going to use git-fast-import.\n\nAwesome, but I just did it properly :)\n\nA few comments regarding fast-import:\n\nWhy the distinction between 'from' and 'merge'? Doesn't it make more\nsense to use 'parent' for both?\n\nI'm doing: commit refs/mtn/d137c7046bae7e4a0144fee82bfce8061f61e3b3\n\nSo I was expecing this to work:\nfrom mtn/d137c7046bae7e4a0144fee82bfce8061f61e3b3\n\nBut it didn't, probably because the commit hasn't actually been\ncommitted. Wouldn't it make sense to store it as a temporal commit so\nmy script doesn't have to deal with that?\n\nAnyway, very nice tool. It's going much faster (1h) compared to before (1 day).\n\nBest regards.\n\n-- \nFelipe Contreras\n"},{"id":"88472","messageId":"20080825163530.GJ31114@lavos.net","threadId":"15187","inReplyTo":"94a0d4530808240218j4bedbe3di99303da9addc93a4@mail.gmail.com","subject":"Re: [RFC] mtn to git conversion script","fromName":"Brian Downing","fromEmail":"bdowning@lavos.net","sentAt":"2008-08-25T16:35:31Z","receivedAt":"2008-08-25T16:35:31Z","isPatch":false,"sender":{"key":"bdowning@lavos.net","avatar":"https://avatars.githubusercontent.com/u/366426?v=4"},"body":"On Sun, Aug 24, 2008 at 12:18:50PM +0300, Felipe Contreras wrote:\n> I developed a script that converts a monotone repository into a git\n> one (exact clone), I want to contribute it so everybody can use it.\n> \n> This is the gist of the script:\n> \n> mtn update --revision #{@id} --reallyquiet\n> git ls-files --modified --others --exclude-standard -z | git\n> update-index --add --remove -z --stdin\n> git write-tree\n> git write-raw < /tmp/commit.txt\n> git update-ref refs/mtn/#{@id} #{@git_id}\n> \n> branches.each do |e|\n>     git update-ref refs/heads/#{e} #{@git_id}\n> end\n\nYou definitely want to use fast-import, but you probably want to do\nsomething a lot closer to fast-export for monotone (read: use its\nautomate stdio interface and avoid expensive calls).\n\nHere's a simple monotone to git converter I wrote.  You'll need the\nMonotone::AutomateStdio perl module to use it (which I think I got it\nfrom monotone's net.venge.monotone.contrib.lib.automate-stdio branch).\nIt is very fast; it can convert the OpenEmbedded repo in something like\n5-10 minutes on my machine.\n\nNote that for monotone export to go fast you absolutely /must/ avoid the\nget_manifest operation.  In my converter I use the revision information\ndirectly.  Getting the renames right with this is a little tricky; IIRC,\nthe ordering that works is:\n\n* Rename all renamed files, innermost files first, to temporary names.\n* Delete all deleted files, innermost first.\n* Rename all temporary names to permanent names, outermost first.\n* Add all new/modified files.\n\nConveniently, all of the above can be done by directly emitting\nfast-import commands, so you don't have to keep track of trees directly.\n(With one exception, which I'll elaborate on in a different email.)\n\n-bcd\n\n\n#!/usr/bin/perl\n# Copyright (C) 2007-2008  Brian Downing\n# This program is licensed under version 2 of the GNU GPL.\n\nuse strict;\nuse Monotone::AutomateStdio;\nuse Date::Parse;\n\nmy $D = 0;\n\nmy $m = Monotone::AutomateStdio->new($ARGV[0]);\n\nmy $revlist = [];\n$m->graph($revlist);\nmy $sorted = [];\n\nfor my $rev (@$revlist) {\n    push(@$sorted, $rev->{revision_id});\n}\n\nmy $leaves = [];\n$m->leaves($leaves);\n\n$m->toposort($sorted, @$sorted);\n\nmy $marks = {};\nmy $mark = 1;\nmy $c = 0;\n\nsub quote_file {\n    $_ = shift;\n    return $_;\n    s/\\\\/\\\\\\\\/g;\n    s/\\n/\\\\n/g;\n    s/\"/\\\\\"/g;\n    return qq(\"$_\");\n}\n\nsub lprint {\n    my $fh = shift;\n    print @_ if $D;\n    print $fh @_;\n}\n\nsub lprintf {\n    my $fh = shift;\n    printf @_ if $D;\n    printf $fh @_;\n}\n\nmy $tmptag = \"624d893e-ae1a-42d8-90a9-926a6ceffae8\";\n\nopen my $fi, '|-', 'git-fast-import --export-marks=file';\nfor my $rev (@$sorted) {\n    my ($time, $author, $msg) = (\"0\", \"__UNKNOWN__\", \"__UNKNOWN__\");\n    my @certs;\n    my @branches;\n    $m->certs(\\@certs, $rev);\n    for my $cert (@certs) {\n        my ($n, $v) = ($cert->{name}, $cert->{value});\n        $author = $v if ($n eq 'author');\n        $time = $v if ($n eq 'date');\n        $msg = $v if ($n eq 'changelog');\n        push(@branches, $v) if ($n eq 'branch');\n    }\n    my $email = $author;\n    $msg .= \"\\nmtn-revision: $rev\\n\";\n    for my $b (sort @branches) {\n        $msg .= \"mtn-branch: $b\\n\";\n    }\n    $time = str2time($time, 'UTC');\n    my $mfest = [];\n    $m->get_revision($mfest, $rev);\n    my $orcount = 0;\n    my $add_files = {};\n    my $add_dirs = {};\n    my $delete_files = {};\n    my $from_tmpnames = {};\n    my $to_tmpnames = {};\n    my $curtmp = 0;\n    my @parents;\n    for my $e (@$mfest) {\n        if ($e->{type} eq 'old_revision') {\n            push(@parents, $e->{revision_id});\n            ++$orcount;\n        } \n        next if $orcount > 1;\n        if ($e->{type} eq 'add_file' || $e->{type} eq 'patch') {\n            my $id = $e->{file_id} || $e->{to_file_id};\n            $add_files->{$e->{name}} = $id;\n            unless ($marks->{$id}) {\n                my $data;\n                $m->get_file(\\$data, $id);\n                print \"new file $id\\n\" if $D;\n                print $fi \"blob\\n\";\n                my $len = length($data);\n                print $fi \"mark :$mark\\n\";\n                $marks->{$id} = $mark++;\n                print $fi \"data $len\\n$data\\n\";\n            }\n        } elsif ($e->{type} eq 'add_dir') {\n            $add_dirs->{$e->{name}} = 1;\n        } elsif ($e->{type} eq 'delete') {\n            $delete_files->{$e->{name}} = 1;\n        } elsif ($e->{type} eq 'rename') {\n            $curtmp++;\n            $from_tmpnames->{$e->{from_name}} = \"__tmp_${tmptag}_$curtmp\";\n            $to_tmpnames->{$e->{to_name}} = \"__tmp_${tmptag}_$curtmp\";\n        }\n    }\n    printf(\"rev $rev (%d/%d, %.2f%)\\n\",\n           ++$c, scalar(@$sorted), 100*$c/scalar(@$sorted));\n    print $fi \"reset refs/import\\n\" unless @parents;\n    lprint $fi, \"commit refs/import\\n\";\n    print $fi \"mark :$mark\\n\";\n    $marks->{$rev} = $mark++;\n    if ($author =~ m(\\s*(.*?\\S)\\s*<(.*)>\\s*)) {\n        $author = $1;\n        $email = $2;\n    }\n    $author =~ s/[<>]/_/g;\n    $email =~ s/[<>]/_/g;\n    $author =~ s/@.*//;\n    print $fi \"committer $author <$email> $time +0000\\n\";\n    my $len = length($msg);\n    print $fi \"data $len\\n$msg\\n\";\n    my $from = \"from\";\n    for my $p (@parents) {\n        lprint $fi, \"$from :$marks->{$p}\\n\";\n        $from = \"merge\";\n    }\n    for my $f (sort { length($b) <=> length ($a) } keys %$from_tmpnames) {\n        lprintf($fi, \"R %s %s\\n\",\n                quote_file($f), quote_file($from_tmpnames->{$f}));\n    }\n    for my $f (sort { length($b) <=> length ($a) } keys %$delete_files) {\n        lprintf($fi, \"D %s\\n\", quote_file($f));\n    }\n    for my $f (sort { length($a) <=> length ($b) } keys %$to_tmpnames) {\n        lprintf($fi, \"R %s %s\\n\",\n                quote_file($to_tmpnames->{$f}), quote_file($f));\n    }\n    for my $f (keys %$add_files) {\n        lprintf($fi, \"M 0644 :%s %s\\n\",\n                $marks->{$add_files->{$f}}, quote_file($f));\n    }\n    for my $f (keys %$add_dirs) {\n        $f .= \"/\" if $f;\n        lprintf($fi, \"M 0644 inline %s\\n\", quote_file(\"$f.gitignore\"));\n        lprint($fi, \"data 0\\n\\n\");\n    }\n    print $fi \"\\n\";\n}\n\nmy $branches = {};\nfor my $rev (@$leaves) {\n    my $branch;\n    my @certs;\n    $m->certs(\\@certs, $rev);\n    for my $cert (@certs) {\n        my ($n, $v) = ($cert->{name}, $cert->{value});\n        $branch = $v if ($n eq 'branch');\n    }\n    my $r = $branches->{$branch};\n    $branches->{$branch}--;\n    if ($marks->{$rev}) {\n        print $fi \"reset refs/heads/$branch$r\\n\";\n        print $fi \"from :$marks->{$rev}\\n\\n\";\n    }\n}\n\nclose $fi;\n"},{"id":"88474","messageId":"20080825164153.GK31114@lavos.net","threadId":"15187","inReplyTo":"20080825163530.GJ31114@lavos.net","subject":"Re: [RFC] mtn to git conversion script","fromName":"Brian Downing","fromEmail":"bdowning@lavos.net","sentAt":"2008-08-25T16:41:53Z","receivedAt":"2008-08-25T16:41:53Z","isPatch":false,"sender":{"key":"bdowning@lavos.net","avatar":"https://avatars.githubusercontent.com/u/366426?v=4"},"body":"On Mon, Aug 25, 2008 at 11:35:31AM -0500, Brian Downing wrote:\n> Note that for monotone export to go fast you absolutely /must/ avoid the\n> get_manifest operation.  In my converter I use the revision information\n> directly.  Getting the renames right with this is a little tricky; IIRC,\n> the ordering that works is:\n> \n> * Rename all renamed files, innermost files first, to temporary names.\n> * Delete all deleted files, innermost first.\n> * Rename all temporary names to permanent names, outermost first.\n> * Add all new/modified files.\n> \n> Conveniently, all of the above can be done by directly emitting\n> fast-import commands, so you don't have to keep track of trees directly.\n> (With one exception, which I'll elaborate on in a different email.)\n\nThe exception is one commit in monotone's repository.  There was\nactually a commit that did:\n\n    rename '/' '/something'\n    add '/other'\n\nMonotone can apparently handle that, but git fast-import cannot, last I\nchecked.  One would have to \"know\" what all the files were and recreate\nthem by hand, which was what fast-import's move/copy commands were\nsupposed to avoid.\n\nObviously this use case is not too important to me, as patches have not\nbeen forthcoming to fix this, but I figured I'd mention it in case it's\nimportant to somebody else.\n\n-bcd\n"},{"id":"88507","messageId":"94a0d4530808251347g4d6246bv7ebd5cc86294dd05@mail.gmail.com","threadId":"15187","inReplyTo":"20080825163530.GJ31114@lavos.net","subject":"Re: [RFC] mtn to git conversion script","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2008-08-25T20:47:53Z","receivedAt":"2008-08-25T20:47:53Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Mon, Aug 25, 2008 at 7:35 PM, Brian Downing <bdowning@lavos.net> wrote:\n> On Sun, Aug 24, 2008 at 12:18:50PM +0300, Felipe Contreras wrote:\n>> I developed a script that converts a monotone repository into a git\n>> one (exact clone), I want to contribute it so everybody can use it.\n>>\n>> This is the gist of the script:\n>>\n>> mtn update --revision #{@id} --reallyquiet\n>> git ls-files --modified --others --exclude-standard -z | git\n>> update-index --add --remove -z --stdin\n>> git write-tree\n>> git write-raw < /tmp/commit.txt\n>> git update-ref refs/mtn/#{@id} #{@git_id}\n>>\n>> branches.each do |e|\n>>     git update-ref refs/heads/#{e} #{@git_id}\n>> end\n>\n> You definitely want to use fast-import, but you probably want to do\n> something a lot closer to fast-export for monotone (read: use its\n> automate stdio interface and avoid expensive calls).\n>\n> Here's a simple monotone to git converter I wrote.  You'll need the\n> Monotone::AutomateStdio perl module to use it (which I think I got it\n> from monotone's net.venge.monotone.contrib.lib.automate-stdio branch).\n> It is very fast; it can convert the OpenEmbedded repo in something like\n> 5-10 minutes on my machine.\n\nInteresting, how many commits?\n\n> Note that for monotone export to go fast you absolutely /must/ avoid the\n> get_manifest operation.  In my converter I use the revision information\n> directly.  Getting the renames right with this is a little tricky; IIRC,\n> the ordering that works is:\n>\n> * Rename all renamed files, innermost files first, to temporary names.\n> * Delete all deleted files, innermost first.\n> * Rename all temporary names to permanent names, outermost first.\n> * Add all new/modified files.\n>\n> Conveniently, all of the above can be done by directly emitting\n> fast-import commands, so you don't have to keep track of trees directly.\n> (With one exception, which I'll elaborate on in a different email.)\n\nI guess I haven't stumbled upon that problem yet =/\n\nBest regards.\n\n-- \nFelipe Contreras\n"},{"id":"88511","messageId":"20080825210932.GL31114@lavos.net","threadId":"15187","inReplyTo":"94a0d4530808251347g4d6246bv7ebd5cc86294dd05@mail.gmail.com","subject":"Re: [RFC] mtn to git conversion script","fromName":"Brian Downing","fromEmail":"bdowning@lavos.net","sentAt":"2008-08-25T21:09:32Z","receivedAt":"2008-08-25T21:09:32Z","isPatch":false,"sender":{"key":"bdowning@lavos.net","avatar":"https://avatars.githubusercontent.com/u/366426?v=4"},"body":"On Mon, Aug 25, 2008 at 11:47:53PM +0300, Felipe Contreras wrote:\n> On Mon, Aug 25, 2008 at 7:35 PM, Brian Downing <bdowning@lavos.net> wrote:\n> > Here's a simple monotone to git converter I wrote.  You'll need the\n> > Monotone::AutomateStdio perl module to use it (which I think I got it\n> > from monotone's net.venge.monotone.contrib.lib.automate-stdio branch).\n> > It is very fast; it can convert the OpenEmbedded repo in something like\n> > 5-10 minutes on my machine.\n> \n> Interesting, how many commits?\n\n:; git rev-list --all | wc -l\n23498 revisions\n:; git ls-tree -r org.openembedded.stable | wc -l\n17502 files in HEAD\n\n(Some of those files are .gitignore files, which I create in every\ndirectory to hold open empty ones.)\n\n-bcd\n"},{"id":"88973","messageId":"g95eoo$5ok$8@ger.gmane.org","threadId":"15187","inReplyTo":"94a0d4530808241745r3f2bdb56q9cfa8bc61f79223e@mail.gmail.com","subject":"Re: [RFC] mtn to git conversion script","fromName":"Anand Kumria","fromEmail":"wildfire@progsoc.org","sentAt":"2008-08-28T05:57:44Z","receivedAt":"2008-08-28T05:57:44Z","isPatch":false,"sender":{"key":"wildfire@progsoc.org","avatar":null},"body":"\nHi Felipe,\n\nOn Mon, 25 Aug 2008 03:45:11 +0300, Felipe Contreras wrote:\n\n> \n> Anyway, very nice tool. It's going much faster (1h) compared to before\n> (1 day).\n\nWill you be submitting this as something for/to contrib?\n\nThanks,\nAnand\n"},{"id":"88982","messageId":"g95j3v$5ok$9@ger.gmane.org","threadId":"15187","inReplyTo":"20080825163530.GJ31114@lavos.net","subject":"Re: [RFC] mtn to git conversion script","fromName":"Anand Kumria","fromEmail":"wildfire@progsoc.org","sentAt":"2008-08-28T07:11:59Z","receivedAt":"2008-08-28T07:11:59Z","isPatch":false,"sender":{"key":"wildfire@progsoc.org","avatar":null},"body":"Hi Brian,\n\nOn Mon, 25 Aug 2008 11:35:31 -0500, Brian Downing wrote:\n\n[snip - Convenient mtn -> git converter ]\n\nI think you need to add a Signed-off-by: line in order for Junio to be \nable to take this and put into the contrib section.\n\nThanks,\nAnand\n"},{"id":"88996","messageId":"94a0d4530808280203o6d97f69we4768115e12800c2@mail.gmail.com","threadId":"15187","inReplyTo":"g95eoo$5ok$8@ger.gmane.org","subject":"Re: [Monotone-devel] Re: [RFC] mtn to git conversion script","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2008-08-28T09:03:29Z","receivedAt":"2008-08-28T09:03:29Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Thu, Aug 28, 2008 at 8:57 AM, Anand Kumria <wildfire@progsoc.org> wrote:\n>\n> Hi Felipe,\n>\n> On Mon, 25 Aug 2008 03:45:11 +0300, Felipe Contreras wrote:\n>\n>>\n>> Anyway, very nice tool. It's going much faster (1h) compared to before\n>> (1 day).\n>\n> Will you be submitting this as something for/to contrib?\n\nYes, that's the plan.\n\nHowever, I still don't have something that creates an exact clone with\nfast-import.\n\nAlso, I'm trying different ways to see what would be most efficient.\nRight now it's a combination of Ruby + C, but once I get it working\nI'll post it to the OE mailing lists to see if it works fine for them\ntoo.\n\nOnce the design is good enough I might move everything to C.\n\nBest regards.\n\n-- \nFelipe Contreras\n"},{"id":"89748","messageId":"94a0d4530809040243k49635fd3kfef1ee22a6865e98@mail.gmail.com","threadId":"15187","inReplyTo":"94a0d4530808280203o6d97f69we4768115e12800c2@mail.gmail.com","subject":"Re: [Monotone-devel] Re: [RFC] mtn to git conversion script","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2008-09-04T09:43:09Z","receivedAt":"2008-09-04T09:43:09Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Thu, Aug 28, 2008 at 12:03 PM, Felipe Contreras\n<felipe.contreras@gmail.com> wrote:\n> On Thu, Aug 28, 2008 at 8:57 AM, Anand Kumria <wildfire@progsoc.org> wrote:\n>>\n>> Hi Felipe,\n>>\n>> On Mon, 25 Aug 2008 03:45:11 +0300, Felipe Contreras wrote:\n>>\n>>>\n>>> Anyway, very nice tool. It's going much faster (1h) compared to before\n>>> (1 day).\n>>\n>> Will you be submitting this as something for/to contrib?\n>\n> Yes, that's the plan.\n>\n> However, I still don't have something that creates an exact clone with\n> fast-import.\n>\n> Also, I'm trying different ways to see what would be most efficient.\n> Right now it's a combination of Ruby + C, but once I get it working\n> I'll post it to the OE mailing lists to see if it works fine for them\n> too.\n>\n> Once the design is good enough I might move everything to C.\n>\n> Best regards.\n\nOk, now the basics seem to be working. So I'm uploading some code if\nanyone wants to take a look.\n\nThe C code is generating a topologically sorted list of revisions, and\nstoring the relevant information (certs and parents) separately. This\ncode is very fast. It's using GLib and sqlite3, so probably the GLib\nstuff should be converted to use libgit.\nhttp://gist.github.com/8742\n\nThe Ruby code takes that input and generates an output suitable for\nfast-import. It would be tedious to port the parsing stuff to C, but\nstraightforward.\nhttp://gist.github.com/8740\n\nA patch for fast-import is required, I already submitted it.\n\nComments?\n\n-- \nFelipe Contreras\n"},{"id":"89753","messageId":"m3vdxcp5gv.fsf@localhost.localdomain","threadId":"15187","inReplyTo":"94a0d4530809040243k49635fd3kfef1ee22a6865e98@mail.gmail.com","subject":"Re: [Monotone-devel] Re: [RFC] mtn to git conversion script","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-09-04T10:31:52Z","receivedAt":"2008-09-04T10:31:52Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Felipe Contreras\" <felipe.contreras@gmail.com> writes:\n\n> Ok, now the basics seem to be working. So I'm uploading some code if\n> anyone wants to take a look.\n> \n> The C code is generating a topologically sorted list of revisions, and\n> storing the relevant information (certs and parents) separately. This\n> code is very fast. It's using GLib and sqlite3, so probably the GLib\n> stuff should be converted to use libgit.\n> http://gist.github.com/8742\n> \n> The Ruby code takes that input and generates an output suitable for\n> fast-import. It would be tedious to port the parsing stuff to C, but\n> straightforward.\n> http://gist.github.com/8740\n> \n> A patch for fast-import is required, I already submitted it.\n> \n> Comments?\n\nIf you feel like it is good enough, could you add information about\nthis tool to Git Wiki:\n  http://git.or.cz/gitwiki/InterfacesFrontendsAndTools\nin the \"Interaction with other Revision Control Systems\" section?\n\nTIA\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"89757","messageId":"200809041250.17715.thomas.moschny@gmx.de","threadId":"15187","inReplyTo":"94a0d4530809040243k49635fd3kfef1ee22a6865e98@mail.gmail.com","subject":"Re: [Monotone-devel] Re: [RFC] mtn to git conversion script","fromName":"Thomas Moschny","fromEmail":"thomas.moschny@gmx.de","sentAt":"2008-09-04T10:50:09Z","receivedAt":"2008-09-04T10:50:09Z","isPatch":false,"sender":{"key":"thomas.moschny@gmx.de","avatar":null},"body":"On Thu, Sep 4, 2008, Felipe Contreras wrote:\n> Ok, now the basics seem to be working. So I'm uploading some code if\n> anyone wants to take a look.\n>\n> The C code is generating a topologically sorted list of revisions, and\n> storing the relevant information (certs and parents) separately. This\n> code is very fast. It's using GLib and sqlite3, so probably the GLib\n> stuff should be converted to use libgit.\n> http://gist.github.com/8742\n\nYou shouldn't access Monotone's sqlite database directly, for various reasons. \nUse the Automation Interface instead, see \nhttp://www.monotone.ca/docs/Automation.html#Automation. Using 'mtn automate \nstdio', you can feed an arbitrary amount of commands to one single running mtn \nprocess.\n\n- Thomas\n\n"},{"id":"89762","messageId":"94a0d4530809040621h40111cwf412bc0f8811d9db@mail.gmail.com","threadId":"15187","inReplyTo":"m3vdxcp5gv.fsf@localhost.localdomain","subject":"Re: [Monotone-devel] Re: [RFC] mtn to git conversion script","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2008-09-04T13:21:48Z","receivedAt":"2008-09-04T13:21:48Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Thu, Sep 4, 2008 at 1:31 PM, Jakub Narebski <jnareb@gmail.com> wrote:\n> \"Felipe Contreras\" <felipe.contreras@gmail.com> writes:\n>\n>> Ok, now the basics seem to be working. So I'm uploading some code if\n>> anyone wants to take a look.\n>>\n>> The C code is generating a topologically sorted list of revisions, and\n>> storing the relevant information (certs and parents) separately. This\n>> code is very fast. It's using GLib and sqlite3, so probably the GLib\n>> stuff should be converted to use libgit.\n>> http://gist.github.com/8742\n>>\n>> The Ruby code takes that input and generates an output suitable for\n>> fast-import. It would be tedious to port the parsing stuff to C, but\n>> straightforward.\n>> http://gist.github.com/8740\n>>\n>> A patch for fast-import is required, I already submitted it.\n>>\n>> Comments?\n>\n> If you feel like it is good enough, could you add information about\n> this tool to Git Wiki:\n>  http://git.or.cz/gitwiki/InterfacesFrontendsAndTools\n> in the \"Interaction with other Revision Control Systems\" section?\n\nNot yet.\n\nIt still needs to add the branches, tags, and HEAD.\n\n-- \nFelipe Contreras\n"},{"id":"89764","messageId":"94a0d4530809040629y2f2d9c74v311b83afebba0051@mail.gmail.com","threadId":"15187","inReplyTo":"200809041250.17715.thomas.moschny@gmx.de","subject":"Re: [Monotone-devel] Re: [RFC] mtn to git conversion script","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2008-09-04T13:29:31Z","receivedAt":"2008-09-04T13:29:31Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Thu, Sep 4, 2008 at 1:50 PM, Thomas Moschny <thomas.moschny@gmx.de> wrote:\n> On Thu, Sep 4, 2008, Felipe Contreras wrote:\n>> Ok, now the basics seem to be working. So I'm uploading some code if\n>> anyone wants to take a look.\n>>\n>> The C code is generating a topologically sorted list of revisions, and\n>> storing the relevant information (certs and parents) separately. This\n>> code is very fast. It's using GLib and sqlite3, so probably the GLib\n>> stuff should be converted to use libgit.\n>> http://gist.github.com/8742\n>\n> You shouldn't access Monotone's sqlite database directly, for various reasons.\n> Use the Automation Interface instead, see\n> http://www.monotone.ca/docs/Automation.html#Automation. Using 'mtn automate\n> stdio', you can feed an arbitrary amount of commands to one single running mtn\n> process.\n\nI use mtn stdio when needed, that is, when doing it manually would be\ntoo complicated (get_file). Doing it directly with sqlite3 is *very*\nfast, I don't see any reason to not to do it.\n\nFeel free to modify the code for stdio.\n\n-- \nFelipe Contreras\n"},{"id":"95462","messageId":"1c3be50f0811110830y5045717csc5ddb1c0576cb046@mail.gmail.com","threadId":"15187","inReplyTo":"20080825163530.GJ31114@lavos.net","subject":"Re: [RFC] mtn to git conversion script","fromName":"Juan Jose Comellas","fromEmail":"juanjo@comellas.org","sentAt":"2008-11-11T16:30:53Z","receivedAt":"2008-11-11T16:30:53Z","isPatch":false,"sender":{"key":"juanjo@comellas.org","avatar":"https://gravatar.com/avatar/0933b4001eab527017ad3554b501a677366d69c489dfa89d09027f7390c02222?d=mp&s=160"},"body":"I made some modifications to the script that converts Monotone repositories\nto Git to make it work with what I had. I also added support for renaming\ncommit authors. To use the modified script just call it passing the name of\nthe repository file as the first argument. You can add a second optional\nargument with the name of the file that holds the authors' names and email\naddresses. In this file you should have one line per commit author with the\nfollowing format:\n\nFirstname Lastname <email@example.com>\n\nThis script still uses the AutomateStdio.pm Perl module that can be found in\nthe net.venge.monotone.contrib.lib.automate-stdio branch of Monotone's main\nrepository.\n\nPS. I'm no Perl guru so there might be some bugs lurking in the code I\nadded. It did work for my repositories, though.\n\n\nOn Mon, Aug 25, 2008 at 2:35 PM, Brian Downing <bdowning@lavos.net> wrote:\n\n> On Sun, Aug 24, 2008 at 12:18:50PM +0300, Felipe Contreras wrote:\n> > I developed a script that converts a monotone repository into a git\n> > one (exact clone), I want to contribute it so everybody can use it.\n> >\n> > This is the gist of the script:\n> >\n> > mtn update --revision #{@id} --reallyquiet\n> > git ls-files --modified --others --exclude-standard -z | git\n> > update-index --add --remove -z --stdin\n> > git write-tree\n> > git write-raw < /tmp/commit.txt\n> > git update-ref refs/mtn/#{@id} #{@git_id}\n> >\n> > branches.each do |e|\n> >     git update-ref refs/heads/#{e} #{@git_id}\n> > end\n>\n> You definitely want to use fast-import, but you probably want to do\n> something a lot closer to fast-export for monotone (read: use its\n> automate stdio interface and avoid expensive calls).\n>\n> Here's a simple monotone to git converter I wrote.  You'll need the\n> Monotone::AutomateStdio perl module to use it (which I think I got it\n> from monotone's net.venge.monotone.contrib.lib.automate-stdio branch).\n> It is very fast; it can convert the OpenEmbedded repo in something like\n> 5-10 minutes on my machine.\n>\n> Note that for monotone export to go fast you absolutely /must/ avoid the\n> get_manifest operation.  In my converter I use the revision information\n> directly.  Getting the renames right with this is a little tricky; IIRC,\n> the ordering that works is:\n>\n> * Rename all renamed files, innermost files first, to temporary names.\n> * Delete all deleted files, innermost first.\n> * Rename all temporary names to permanent names, outermost first.\n> * Add all new/modified files.\n>\n> Conveniently, all of the above can be done by directly emitting\n> fast-import commands, so you don't have to keep track of trees directly.\n> (With one exception, which I'll elaborate on in a different email.)\n>\n> -bcd\n>\n\n\n#!/usr/bin/perl\n# Copyright (C) 2007-2008  Brian Downing\n# This program is licensed under version 2 of the GNU GPL.\n\nuse strict;\nuse Monotone::AutomateStdio;\nuse Date::Parse;\n\nmy $D = 0;\n\nmy $m = Monotone::AutomateStdio->new($ARGV[0]);\n\nmy $revlist = [];\n$m->graph($revlist);\nmy $sorted = [];\n\nfor my $rev (@$revlist) {\n    push(@$sorted, $rev->{revision_id});\n}\n\nmy $leaves = [];\n$m->leaves($leaves);\n\n$m->toposort($sorted, @$sorted);\n\nmy $marks = {};\nmy $mark = 1;\nmy $c = 0;\nmy %authors = {};\n\nsub quote_file {\n    $_ = shift;\n    return $_;\n    s/\\\\/\\\\\\\\/g;\n    s/\\n/\\\\n/g;\n    s/\"/\\\\\"/g;\n    return qq(\"$_\");\n}\n\nsub lprint {\n    my $fh = shift;\n    print @_ if $D;\n    print $fh @_;\n}\n\nsub lprintf {\n    my $fh = shift;\n    printf @_ if $D;\n    printf $fh @_;\n}\n\nsub load_authors {\n    my %authors = ();\n    my $filename = shift(@_);\n    open my $fi, '<', $filename or die \"Could not open authors map file $filename\\n\"; \n    for my $line (<$fi>) {\n        if ($line =~ m/(.*) <(.*)>/) {\n            if ($2) {\n                $authors{$2} = $1;\n            }\n        }\n    }\n    return %authors;\n}\n\nsub author_name {\n    my $email = shift(@_);\n    my $name = $authors{$email};\n    if ($name) {\n        return $name;\n    } else {\n        if ($email =~ m/(.+)\\@.+/) {\n            return $1;\n        } else {\n            return $email;\n        }\n    }\n}\n\nmy $tmptag = \"624d893e-ae1a-42d8-90a9-926a6ceffae8\";\n\nif ($ARGV[1]) {\n    %authors = load_authors($ARGV[1]);\n} else {\n    %authors = {};\n}\nopen my $fi, '|-', 'git-fast-import --export-marks=file';\n# open my $fi, '>fast-import.dump';\nbinmode $fi;\nfor my $rev (@$sorted) {\n    my ($time, $author, $msg) = (\"0\", \"__UNKNOWN__\", \"__UNKNOWN__\");\n    my @certs;\n    my @branches;\n    $m->certs(\\@certs, $rev);\n    for my $cert (@certs) {\n        my ($n, $v) = ($cert->{name}, $cert->{value});\n        $author = $v if ($n eq 'author');\n        $time = $v if ($n eq 'date');\n        $msg = $v if ($n eq 'changelog');\n        push(@branches, $v) if ($n eq 'branch');\n    }\n    my $email = $author;\n    $msg .= \"\\nmtn-revision: $rev\\n\";\n    for my $b (sort @branches) {\n        $msg .= \"mtn-branch: $b\\n\";\n    }\n    $time = str2time($time, 'UTC');\n    my $mfest = [];\n    $m->get_revision($mfest, $rev);\n    my $orcount = 0;\n    my $add_files = {};\n    my $add_dirs = {};\n    my $delete_files = {};\n    my $from_tmpnames = {};\n    my $to_tmpnames = {};\n    my $curtmp = 0;\n    my @parents;\n    for my $e (@$mfest) {\n        if ($e->{type} eq 'old_revision') {\n            push(@parents, $e->{revision_id});\n            ++$orcount;\n        } \n        next if $orcount > 1;\n        if ($e->{type} eq 'add_file' || $e->{type} eq 'patch') {\n            my $id = $e->{file_id} || $e->{to_file_id};\n            $add_files->{$e->{name}} = $id;\n            unless ($marks->{$id}) {\n                my $data;\n                $m->get_file(\\$data, $id);\n                print \"new file $id\\n\" if $D;\n                print $fi \"blob\\n\";\n                my $len = length($data);\n                print $fi \"mark :$mark\\n\";\n                $marks->{$id} = $mark++;\n                print $fi \"data $len\\n$data\\n\";\n                #print $fi \"data $len\\n\";\n                #print $fi pack('C', $data);\n                #print $fi \"\\n\";\n            }\n        } elsif ($e->{type} eq 'add_dir') {\n            $add_dirs->{$e->{name}} = 1;\n        } elsif ($e->{type} eq 'delete') {\n            $delete_files->{$e->{name}} = 1;\n        } elsif ($e->{type} eq 'rename') {\n            $curtmp++;\n            $from_tmpnames->{$e->{from_name}} = \"__tmp_${tmptag}_$curtmp\";\n            $to_tmpnames->{$e->{to_name}} = \"__tmp_${tmptag}_$curtmp\";\n        }\n    }\n    printf(\"rev $rev (%d/%d, %.2f%)\\n\",\n           ++$c, scalar(@$sorted), 100*$c/scalar(@$sorted));\n    print $fi \"reset refs/import\\n\" unless @parents;\n    lprint $fi, \"commit refs/import\\n\";\n    print $fi \"mark :$mark\\n\";\n    $marks->{$rev} = $mark++;\n    if ($author =~ m(\\s*(.*?\\S)\\s*<(.*)>\\s*)) {\n        # $author = $1;\n        $email = $2;\n    }\n    # $author =~ s/[<>]/_/g;\n    $email =~ s/[<>]/_/g;\n    # $author =~ s/@.*//;\n    $author = author_name($email);\n    print $fi \"committer $author <$email> $time +0000\\n\";\n    my $len = length($msg);\n    print $fi \"data $len\\n$msg\\n\";\n    my $from = \"from\";\n    for my $p (@parents) {\n        # lprint $fi, \"$from :$marks->{$p}\\n\";\n        my $parent_mark = $marks->{$p};\n        if ($parent_mark) {\n            lprint $fi, \"$from :$parent_mark\\n\";\n            $from = \"merge\";\n        }\n    }\n    for my $f (sort { length($b) <=> length ($a) } keys %$from_tmpnames) {\n        lprintf($fi, \"R %s %s\\n\",\n                quote_file($f), quote_file($from_tmpnames->{$f}));\n    }\n    for my $f (sort { length($b) <=> length ($a) } keys %$delete_files) {\n        lprintf($fi, \"D %s\\n\", quote_file($f));\n    }\n    for my $f (sort { length($a) <=> length ($b) } keys %$to_tmpnames) {\n        lprintf($fi, \"R %s %s\\n\",\n                quote_file($to_tmpnames->{$f}), quote_file($f));\n    }\n    for my $f (keys %$add_files) {\n        lprintf($fi, \"M 0644 :%s %s\\n\",\n                $marks->{$add_files->{$f}}, quote_file($f));\n    }\n    for my $f (keys %$add_dirs) {\n        $f .= \"/\" if $f;\n        lprintf($fi, \"M 0644 inline %s\\n\", quote_file(\"$f.gitignore\"));\n        lprint($fi, \"data 0\\n\\n\");\n    }\n    print $fi \"\\n\";\n}\n\nmy $branches = {};\nfor my $rev (@$leaves) {\n    my $branch;\n    my @certs;\n    $m->certs(\\@certs, $rev);\n    for my $cert (@certs) {\n        my ($n, $v) = ($cert->{name}, $cert->{value});\n        $branch = $v if ($n eq 'branch');\n    }\n    my $r = $branches->{$branch};\n    $branches->{$branch}--;\n    if ($marks->{$rev}) {\n        print $fi \"reset refs/heads/$branch$r\\n\";\n        print $fi \"from :$marks->{$rev}\\n\\n\";\n    }\n}\n\nclose $fi;\n\n\n_______________________________________________\nMonotone-devel mailing list\nMonotone-devel@nongnu.org\nhttp://lists.nongnu.org/mailman/listinfo/monotone-devel\n"},{"id":"95463","messageId":"1c3be50f0811110840p7dff8972r2a5ef7193a0306c2@mail.gmail.com","threadId":"15187","inReplyTo":"20080825163530.GJ31114@lavos.net","subject":"Re: [RFC] mtn to git conversion script","fromName":"Juan Jose Comellas","fromEmail":"juanjo@comellas.org","sentAt":"2008-11-11T16:40:46Z","receivedAt":"2008-11-11T16:40:46Z","isPatch":false,"sender":{"key":"juanjo@comellas.org","avatar":"https://gravatar.com/avatar/0933b4001eab527017ad3554b501a677366d69c489dfa89d09027f7390c02222?d=mp&s=160"},"body":"I made some modifications to the script that converts Monotone\nrepositories to Git to make it work with what I had. I also added\nsupport for renaming commit authors.\n\nTo use the modified script just call it passing the name of the\nrepository file as the first argument. You can add a second optional\nargument with the name of the file that holds the authors' names and\nemail addresses. In this file you should have one line per commit\nauthor with the following format:\n\nFirstname Lastname <email@example.com>\n\nThis script still uses the AutomateStdio.pm Perl module that can be\nfound in the net.venge.monotone.contrib.lib.automate-stdio branch of\nMonotone's main repository.\n\nI'm no Perl guru so there might be some bugs lurking in the code I\nadded. It did work for my repositories, though.\n\nPS. Resending because I mistakenly sent the previous message as HTML mail.\n\n\nOn Mon, Aug 25, 2008 at 2:35 PM, Brian Downing <bdowning@lavos.net> wrote:\n>\n> On Sun, Aug 24, 2008 at 12:18:50PM +0300, Felipe Contreras wrote:\n> > I developed a script that converts a monotone repository into a git\n> > one (exact clone), I want to contribute it so everybody can use it.\n> >\n> > This is the gist of the script:\n> >\n> > mtn update --revision #{@id} --reallyquiet\n> > git ls-files --modified --others --exclude-standard -z | git\n> > update-index --add --remove -z --stdin\n> > git write-tree\n> > git write-raw < /tmp/commit.txt\n> > git update-ref refs/mtn/#{@id} #{@git_id}\n> >\n> > branches.each do |e|\n> >     git update-ref refs/heads/#{e} #{@git_id}\n> > end\n>\n> You definitely want to use fast-import, but you probably want to do\n> something a lot closer to fast-export for monotone (read: use its\n> automate stdio interface and avoid expensive calls).\n>\n> Here's a simple monotone to git converter I wrote.  You'll need the\n> Monotone::AutomateStdio perl module to use it (which I think I got it\n> from monotone's net.venge.monotone.contrib.lib.automate-stdio branch).\n> It is very fast; it can convert the OpenEmbedded repo in something like\n> 5-10 minutes on my machine.\n>\n> Note that for monotone export to go fast you absolutely /must/ avoid the\n> get_manifest operation.  In my converter I use the revision information\n> directly.  Getting the renames right with this is a little tricky; IIRC,\n> the ordering that works is:\n>\n> * Rename all renamed files, innermost files first, to temporary names.\n> * Delete all deleted files, innermost first.\n> * Rename all temporary names to permanent names, outermost first.\n> * Add all new/modified files.\n>\n> Conveniently, all of the above can be done by directly emitting\n> fast-import commands, so you don't have to keep track of trees directly.\n> (With one exception, which I'll elaborate on in a different email.)\n>\n> -bcd\n\n\n#!/usr/bin/perl\n# Copyright (C) 2007-2008  Brian Downing\n# This program is licensed under version 2 of the GNU GPL.\n\nuse strict;\nuse Monotone::AutomateStdio;\nuse Date::Parse;\n\nmy $D = 0;\n\nmy $m = Monotone::AutomateStdio->new($ARGV[0]);\n\nmy $revlist = [];\n$m->graph($revlist);\nmy $sorted = [];\n\nfor my $rev (@$revlist) {\n    push(@$sorted, $rev->{revision_id});\n}\n\nmy $leaves = [];\n$m->leaves($leaves);\n\n$m->toposort($sorted, @$sorted);\n\nmy $marks = {};\nmy $mark = 1;\nmy $c = 0;\nmy %authors = {};\n\nsub quote_file {\n    $_ = shift;\n    return $_;\n    s/\\\\/\\\\\\\\/g;\n    s/\\n/\\\\n/g;\n    s/\"/\\\\\"/g;\n    return qq(\"$_\");\n}\n\nsub lprint {\n    my $fh = shift;\n    print @_ if $D;\n    print $fh @_;\n}\n\nsub lprintf {\n    my $fh = shift;\n    printf @_ if $D;\n    printf $fh @_;\n}\n\nsub load_authors {\n    my %authors = ();\n    my $filename = shift(@_);\n    open my $fi, '<', $filename or die \"Could not open authors map file $filename\\n\"; \n    for my $line (<$fi>) {\n        if ($line =~ m/(.*) <(.*)>/) {\n            if ($2) {\n                $authors{$2} = $1;\n            }\n        }\n    }\n    return %authors;\n}\n\nsub author_name {\n    my $email = shift(@_);\n    my $name = $authors{$email};\n    if ($name) {\n        return $name;\n    } else {\n        if ($email =~ m/(.+)\\@.+/) {\n            return $1;\n        } else {\n            return $email;\n        }\n    }\n}\n\nmy $tmptag = \"624d893e-ae1a-42d8-90a9-926a6ceffae8\";\n\nif ($ARGV[1]) {\n    %authors = load_authors($ARGV[1]);\n} else {\n    %authors = {};\n}\nopen my $fi, '|-', 'git-fast-import --export-marks=file';\n# open my $fi, '>fast-import.dump';\nbinmode $fi;\nfor my $rev (@$sorted) {\n    my ($time, $author, $msg) = (\"0\", \"__UNKNOWN__\", \"__UNKNOWN__\");\n    my @certs;\n    my @branches;\n    $m->certs(\\@certs, $rev);\n    for my $cert (@certs) {\n        my ($n, $v) = ($cert->{name}, $cert->{value});\n        $author = $v if ($n eq 'author');\n        $time = $v if ($n eq 'date');\n        $msg = $v if ($n eq 'changelog');\n        push(@branches, $v) if ($n eq 'branch');\n    }\n    my $email = $author;\n    $msg .= \"\\nmtn-revision: $rev\\n\";\n    for my $b (sort @branches) {\n        $msg .= \"mtn-branch: $b\\n\";\n    }\n    $time = str2time($time, 'UTC');\n    my $mfest = [];\n    $m->get_revision($mfest, $rev);\n    my $orcount = 0;\n    my $add_files = {};\n    my $add_dirs = {};\n    my $delete_files = {};\n    my $from_tmpnames = {};\n    my $to_tmpnames = {};\n    my $curtmp = 0;\n    my @parents;\n    for my $e (@$mfest) {\n        if ($e->{type} eq 'old_revision') {\n            push(@parents, $e->{revision_id});\n            ++$orcount;\n        } \n        next if $orcount > 1;\n        if ($e->{type} eq 'add_file' || $e->{type} eq 'patch') {\n            my $id = $e->{file_id} || $e->{to_file_id};\n            $add_files->{$e->{name}} = $id;\n            unless ($marks->{$id}) {\n                my $data;\n                $m->get_file(\\$data, $id);\n                print \"new file $id\\n\" if $D;\n                print $fi \"blob\\n\";\n                my $len = length($data);\n                print $fi \"mark :$mark\\n\";\n                $marks->{$id} = $mark++;\n                print $fi \"data $len\\n$data\\n\";\n                #print $fi \"data $len\\n\";\n                #print $fi pack('C', $data);\n                #print $fi \"\\n\";\n            }\n        } elsif ($e->{type} eq 'add_dir') {\n            $add_dirs->{$e->{name}} = 1;\n        } elsif ($e->{type} eq 'delete') {\n            $delete_files->{$e->{name}} = 1;\n        } elsif ($e->{type} eq 'rename') {\n            $curtmp++;\n            $from_tmpnames->{$e->{from_name}} = \"__tmp_${tmptag}_$curtmp\";\n            $to_tmpnames->{$e->{to_name}} = \"__tmp_${tmptag}_$curtmp\";\n        }\n    }\n    printf(\"rev $rev (%d/%d, %.2f%)\\n\",\n           ++$c, scalar(@$sorted), 100*$c/scalar(@$sorted));\n    print $fi \"reset refs/import\\n\" unless @parents;\n    lprint $fi, \"commit refs/import\\n\";\n    print $fi \"mark :$mark\\n\";\n    $marks->{$rev} = $mark++;\n    if ($author =~ m(\\s*(.*?\\S)\\s*<(.*)>\\s*)) {\n        # $author = $1;\n        $email = $2;\n    }\n    # $author =~ s/[<>]/_/g;\n    $email =~ s/[<>]/_/g;\n    # $author =~ s/@.*//;\n    $author = author_name($email);\n    print $fi \"committer $author <$email> $time +0000\\n\";\n    my $len = length($msg);\n    print $fi \"data $len\\n$msg\\n\";\n    my $from = \"from\";\n    for my $p (@parents) {\n        # lprint $fi, \"$from :$marks->{$p}\\n\";\n        my $parent_mark = $marks->{$p};\n        if ($parent_mark) {\n            lprint $fi, \"$from :$parent_mark\\n\";\n            $from = \"merge\";\n        }\n    }\n    for my $f (sort { length($b) <=> length ($a) } keys %$from_tmpnames) {\n        lprintf($fi, \"R %s %s\\n\",\n                quote_file($f), quote_file($from_tmpnames->{$f}));\n    }\n    for my $f (sort { length($b) <=> length ($a) } keys %$delete_files) {\n        lprintf($fi, \"D %s\\n\", quote_file($f));\n    }\n    for my $f (sort { length($a) <=> length ($b) } keys %$to_tmpnames) {\n        lprintf($fi, \"R %s %s\\n\",\n                quote_file($to_tmpnames->{$f}), quote_file($f));\n    }\n    for my $f (keys %$add_files) {\n        lprintf($fi, \"M 0644 :%s %s\\n\",\n                $marks->{$add_files->{$f}}, quote_file($f));\n    }\n    for my $f (keys %$add_dirs) {\n        $f .= \"/\" if $f;\n        lprintf($fi, \"M 0644 inline %s\\n\", quote_file(\"$f.gitignore\"));\n        lprint($fi, \"data 0\\n\\n\");\n    }\n    print $fi \"\\n\";\n}\n\nmy $branches = {};\nfor my $rev (@$leaves) {\n    my $branch;\n    my @certs;\n    $m->certs(\\@certs, $rev);\n    for my $cert (@certs) {\n        my ($n, $v) = ($cert->{name}, $cert->{value});\n        $branch = $v if ($n eq 'branch');\n    }\n    my $r = $branches->{$branch};\n    $branches->{$branch}--;\n    if ($marks->{$rev}) {\n        print $fi \"reset refs/heads/$branch$r\\n\";\n        print $fi \"from :$marks->{$rev}\\n\\n\";\n    }\n}\n\nclose $fi;\n"}]}