{"thread":{"id":"7115","subject":"mercurial to git","startedAt":"2007-03-06T21:06:29Z","lastAt":"2007-03-17T11:37:26Z","messageCount":19,"participants":["Rocco Rutte","Theodore Tso","Josef Sipek","Shawn O. Pearce","Linus Torvalds","Len Brown","Simon 'corecode' Schubert"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"36471","messageId":"20070306210629.GA42331@peter.daprodeges.fqdn.th-h.de","threadId":"7115","inReplyTo":null,"subject":"mercurial to git","fromName":"Rocco Rutte","fromEmail":"pdmef@gmx.net","sentAt":"2007-03-06T21:06:29Z","receivedAt":"2007-03-06T21:06:29Z","isPatch":false,"sender":{"key":"pdmef@gmx.net","avatar":null},"body":"Hi,\n\nattached are two files of take #1 of writing a hg2git converter/tracker \nusing git-fast-import. It basically works so use at your own risk and \nsend patches... :)\n\n\"Basically\" means that it gets tags, branches and merges right (working \ntree md5 sums match after imports). It also means that it is horribly \nslow for the repos I tested it own (only mutt and hg-crew).\n\nThe performance bottleneck is hg exporting data, as discovered by people \non #mercurial, the problem is not really fixable and is due to hg's \nrevlog handling. As a result, I needed to let the script feed the full \ncontents of the repository at each revision we walk (i.e. all for the \ninitial import) into git-fast-import. This is horribly slow. For mutt \nwhich contains several tags, a handfull of branches and only 5k commits \nthis takes roughly two hours at 1 commit/sec. My earlier version not \nusing 'deleteall' and feeding only files that changed took 15 minutes \nalltogether, git-fast-import from a textfile 1 min 30 sec.\n\nAs I'll use this my for daily work (more or less), I'll think I'll \n\"maintain\" and keep improving it, so if anyone has comments, critics, \nhints, patches, ...\n\nSomewhat related: It would be really nice to teach git-fast-import to \ninit from a previously saved mark file. Right now I use hg revision \nnumbers as marks, let git-fast-import save them, and read them back next \ntime. These are needed to map hg revisions to git SHA1s in case I need \nto reference something in an incremental import from an earlier run. It \nwould be nice if git-fast-import could do this on its own so that all \nconsumers can benefit and can have persistent marks accross sessions.\n\nAbout the attached files: hg2git.py is the worker script using the \nmercurial python package so that no more slow shell or pipes including \nfork are needed for the raw export, hg2git.sh is a convenience shell \nwrapper taking core of the state files for incremental imports.\n\n   bye, Rocco\n-- \n:wq!\n\n\n#!/usr/bin/env python\n\n# Copyright (c) 2007 Rocco Rutte <pdmef@gmx.net>\n# License: GPLv2\n\n\"\"\"hg2git.py - A mercurial-to-git filter for git-fast-import(1)\nUsage: hg2git.py <hg repo url> <marks file> <heads file> <tip file>\n\"\"\"\n\nfrom mercurial import repo,hg,cmdutil,util,ui,revlog\nfrom tempfile import mkstemp\nimport re\nimport sys\nimport os\n\n# silly regex to see if user field has email address\nuser_re=re.compile('[^<]+ <[^>]+>$')\n# git branch for hg's default 'HEAD' branch\ncfg_master='master'\n# insert 'checkpoint' command after this many commits\ncfg_checkpoint_count=1000\n\ndef usage(ret):\n  sys.stderr.write(__doc__)\n  return ret\n\ndef setup_repo(url):\n  myui=ui.ui()\n  return myui,hg.repository(myui,url)\n\ndef get_changeset(ui,repo,revision):\n  def get_branch(name):\n    if name=='HEAD':\n      name=cfg_master\n    return name\n  def fixup_user(user):\n    if user_re.match(user)==None:\n      if '@' not in user:\n        return user+' <none@none>'\n      return user+' <'+user+'>'\n    return user\n  node=repo.lookup(revision)\n  (manifest,user,(time,timezone),files,desc,extra)=repo.changelog.read(node)\n  tz=\"%+03d%02d\" % (-timezone / 3600, ((-timezone % 3600) / 60))\n  branch=get_branch(extra.get('branch','master'))\n  return (manifest,fixup_user(user),(time,tz),files,desc,branch,extra)\n\ndef gitmode(x):\n  return x and '100755' or '100644'\n\ndef wr(msg=''):\n  print msg\n  #map(lambda x: sys.stderr.write('\\t[%s]\\n' % x),msg.split('\\n'))\n\ndef checkpoint(count):\n  count=count+1\n  if count%cfg_checkpoint_count==0:\n    sys.stderr.write(\"Checkpoint after %d commits\\n\" % count)\n    wr('checkpoint')\n    wr()\n  return count\n\ndef get_parent_mark(parent,marks):\n  p=marks.get(str(parent),None)\n  if p==None:\n    # if we didn't see parent previously, assume we saw it in this run\n    p=':%d' % (parent+1)\n  return p\n\ndef export_commit(ui,repo,revision,marks,heads,last,max,count):\n  sys.stderr.write('Exporting revision %d (tip %d) as [:%d]\\n' % (revision,max,revision+1))\n\n  (_,user,(time,timezone),files,desc,branch,_)=get_changeset(ui,repo,revision)\n  parents=repo.changelog.parentrevs(revision)\n\n  # we need this later to write out tags\n  marks[str(revision)]=':%d'%(revision+1)\n\n  wr('commit refs/heads/%s' % branch)\n  wr('mark :%d' % (revision+1))\n  wr('committer %s %d %s' % (user,time,timezone))\n  wr('data %d' % (len(desc)+1)) # wtf?\n  wr(desc)\n  wr()\n\n  src=heads.get(branch,'')\n  link=''\n  if src!='':\n    # if we have a cached head, this is an incremental import: initialize it\n    # and kill reference so we won't init it again\n    wr('from %s' % src)\n    heads[branch]=''\n  elif not heads.has_key(branch) and revision>0:\n    # newly created branch and not the first one: connect to parent\n    tmp=get_parent_mark(parents[0],marks)\n    wr('from %s' % tmp)\n    sys.stderr.write('Link new branch [%s] to parent [%s]\\n' %\n        (branch,tmp))\n    link=tmp # avoid making a merge commit for branch fork\n\n  if parents:\n    l=last.get(branch,revision)\n    for p in parents:\n      # 1) as this commit implicitely is the child of the most recent\n      #    commit of this branch, ignore this parent\n      # 2) ignore nonexistent parents\n      # 3) merge otherwise\n      if p==l or p==revision or p<0:\n        continue\n      tmp=get_parent_mark(p,marks)\n      # if we fork off a branch, don't merge via 'merge' as we have\n      # 'from' already above\n      if tmp==link:\n        continue\n      sys.stderr.write('Merging branch [%s] with parent [%s] from [r%d]\\n' %\n          (branch,tmp,p))\n      wr('merge %s' % tmp)\n\n  last[branch]=revision\n  heads[branch]=''\n\n  # just wipe the branch clean, all full manifest contents\n  wr('deleteall')\n\n  ctx=repo.changectx(str(revision))\n  man=ctx.manifest()\n\n  #for f in man.keys():\n  #  fctx=ctx.filectx(f)\n  #  d=fctx.data()\n  #  wr('M %s inline %s' % (gitmode(man.execf(f)),f))\n  #  wr('data %d' % len(d)) # had some trouble with size()\n  #  wr(d)\n\n  for fctx in ctx.filectxs():\n    f=fctx.path()\n    d=fctx.data()\n    wr('M %s inline %s' % (gitmode(man.execf(f)),f))\n    wr('data %d' % len(d)) # had some trouble with size()\n    wr(d)\n\n  wr()\n  return checkpoint(count)\n\ndef export_tags(ui,repo,cache,count):\n  l=repo.tagslist()\n  for tag,node in l:\n    if tag=='tip':\n      continue\n    rev=repo.changelog.rev(node)\n    ref=cache.get(str(rev),None)\n    if ref==None:\n      sys.stderr.write('Failed to find reference for creating tag'\n          ' %s at r%d\\n' % (tag,rev))\n      continue\n    (_,user,(time,timezone),_,desc,branch,_)=get_changeset(ui,repo,rev)\n    sys.stderr.write('Exporting tag [%s] at [hg r%d] [git %s]\\n' % (tag,rev,ref))\n    wr('tag %s' % tag)\n    wr('from %s' % ref)\n    wr('tagger %s %d %s' % (user,time,timezone))\n    msg='hg2git created tag %s for hg revision %d on branch %s on (summary):\\n\\t%s' % (tag,\n        rev,branch,desc.split('\\n')[0])\n    wr('data %d' % (len(msg)+1))\n    wr(msg)\n    wr()\n    count=checkpoint(count)\n  return count\n\ndef load_cache(filename):\n  cache={}\n  if not os.path.exists(filename):\n    return cache\n  f=open(filename,'r')\n  l=0\n  for line in f.readlines():\n    l+=1\n    fields=line.split(' ')\n    if fields==None or not len(fields)==2 or fields[0][0]!=':':\n      sys.stderr.write('Invalid file format in [%s], line %d\\n' % (filename,l))\n      continue\n    # put key:value in cache, key without ^:\n    cache[fields[0][1:]]=fields[1].split('\\n')[0]\n  f.close()\n  return cache\n\ndef save_cache(filename,cache):\n  f=open(filename,'w+')\n  map(lambda x: f.write(':%s %s\\n' % (str(x),str(cache.get(x)))),cache.keys())\n  f.close()\n\ndef verify_heads(ui,repo,cache):\n  def getsha1(branch):\n    f=open(os.getenv('GIT_DIR','/dev/null')+'/refs/heads/'+branch)\n    sha1=f.readlines()[0].split('\\n')[0]\n    f.close()\n    return sha1\n\n  for b in cache.keys():\n    sys.stderr.write('Verifying branch [%s]\\n' % b)\n    sha1=getsha1(b)\n    c=cache.get(b)\n    if sha1!=c:\n      sys.stderr.write('Warning: Branch [%s] modified outside hg2git:'\n        '\\n%s (repo) != %s (cache)\\n' % (b,sha1,c))\n  return True\n\nif __name__=='__main__':\n  if len(sys.argv)!=6: sys.exit(usage(1))\n  repourl,m,marksfile,headsfile,tipfile=sys.argv[1:]\n  _max=int(m)\n\n  marks_cache=load_cache(marksfile)\n  heads_cache=load_cache(headsfile)\n  state_cache=load_cache(tipfile)\n\n  ui,repo=setup_repo(repourl)\n\n  if not verify_heads(ui,repo,heads_cache):\n    sys.exit(1)\n\n  tip=repo.changelog.count()\n\n  min=int(state_cache.get('tip',0))\n  max=_max\n  if _max<0:\n    max=tip\n\n  c=int(state_cache.get('count',0))\n  last={}\n  for rev in range(min,max):\n    c=export_commit(ui,repo,rev,marks_cache,heads_cache,last,tip,c)\n\n  c=export_tags(ui,repo,marks_cache,c)\n\n  state_cache['tip']=max\n  state_cache['count']=c\n  state_cache['repo']=repourl\n  save_cache(tipfile,state_cache)\n"},{"id":"36475","messageId":"20070306215459.GI18370@thunk.org","threadId":"7115","inReplyTo":"20070306210629.GA42331@peter.daprodeges.fqdn.th-h.de","subject":"Re: mercurial to git","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2007-03-06T21:54:59Z","receivedAt":"2007-03-06T21:54:59Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Tue, Mar 06, 2007 at 09:06:29PM +0000, Rocco Rutte wrote:\n> \n> attached are two files of take #1 of writing a hg2git converter/tracker \n> using git-fast-import. It basically works so use at your own risk and \n> send patches... :)\n\nI was actually thinking about doing this too, but apparently you beat\nme too it.  :-)\n\n> The performance bottleneck is hg exporting data, as discovered by people \n> on #mercurial, the problem is not really fixable and is due to hg's \n> revlog handling. As a result, I needed to let the script feed the full \n> contents of the repository at each revision we walk (i.e. all for the \n> initial import) into git-fast-import. This is horribly slow. For mutt \n> which contains several tags, a handfull of branches and only 5k commits \n> this takes roughly two hours at 1 commit/sec. My earlier version not \n> using 'deleteall' and feeding only files that changed took 15 minutes \n> alltogether, git-fast-import from a textfile 1 min 30 sec.\n\nHmm.... the way I was planning on handling the performance bottleneck\nwas to use \"hg manifest --debug <rev>\" and diffing the hashes against\nits parents.  Using \"hg manifest\" only hits .hg/00manifest.[di] and\n.hg/00changelog.[di] files, so it's highly efficient.  With the\n--debug option to hg manifest (not needed on some earlier versions of\nhg, but it seems to be needed on the latest development version of\nhg), it outputs the mode and SHA1 hash of the files, so it becomes\neasy to see which files were changed relative to the revision's\nparent(s).\n\nOnce we know which files we need to feed to git-fast-import, it's just\na matter of using \"hg cat -r <rev> <pathname>\" to feed the individual\nchanged file to git-fast-import.  For each file, you only have to\ntouch .hg/data/pathane.[di] files.  So this should allow us to feed\ninput into git-fast-important without needing to feed the full\ncontents of the repository for each revision.\n\nThe other thing that I've been working in my design is how to make the\nconverter to be bidrectional.  That is, if a changelog is made on the\nhg repository, it should be possible to push it over to the git\nrepository, and vice versa, if there are changes made in the git\nrepository, it should be possible to push it back to git.  \n\nIn order to do this it becomes necessary to special case the .hgrc\nfile, and in fact we need to make sure that the .hgrc file does *not*\nshow up in the git repository, but the contents of the .hgrc file\nneeds to be stored in the state file that lives alongside the git and\nhg repositories.\n\nRegards,\n\t\t\t\t\t\t- Ted\n"},{"id":"36484","messageId":"20070306224747.GB42331@peter.daprodeges.fqdn.th-h.de","threadId":"7115","inReplyTo":"20070306215459.GI18370@thunk.org","subject":"Re: mercurial to git","fromName":"Rocco Rutte","fromEmail":"pdmef@gmx.net","sentAt":"2007-03-06T22:47:47Z","receivedAt":"2007-03-06T22:47:47Z","isPatch":false,"sender":{"key":"pdmef@gmx.net","avatar":null},"body":"Hi,\n\n* Theodore Tso [07-03-06 16:54:59 -0500] wrote:\n\n>Hmm.... the way I was planning on handling the performance bottleneck\n>was to use \"hg manifest --debug <rev>\" and diffing the hashes against\n>its parents.  Using \"hg manifest\" only hits .hg/00manifest.[di] and\n>.hg/00changelog.[di] files, so it's highly efficient.  With the\n>--debug option to hg manifest (not needed on some earlier versions of\n>hg, but it seems to be needed on the latest development version of\n>hg), it outputs the mode and SHA1 hash of the files, so it becomes\n>easy to see which files were changed relative to the revision's\n>parent(s).\n\nI started getting/looking at hg a few days ago, mainly at the source \nonly so that I likely miss some things...\n\nHmm. I'll need to further read the hg source to see how they do it. I \nnow switched to defaulting to use the hg changes for normal changesets \nand the full manifest for merges. That's a huge boost already. Your \napproach sounds even better... so that I'll use it. :)\n\n   bye, Rocco\n-- \n:wq!\n"},{"id":"36490","messageId":"20070306230802.GA17226@filer.fsl.cs.sunysb.edu","threadId":"7115","inReplyTo":"20070306215459.GI18370@thunk.org","subject":"Re: mercurial to git","fromName":"Josef Sipek","fromEmail":"jsipek@fsl.cs.sunysb.edu","sentAt":"2007-03-06T23:08:02Z","receivedAt":"2007-03-06T23:08:02Z","isPatch":false,"sender":{"key":"jsipek@fsl.cs.sunysb.edu","avatar":null},"body":"On Tue, Mar 06, 2007 at 04:54:59PM -0500, Theodore Tso wrote:\n...\n> The other thing that I've been working in my design is how to make the\n> converter to be bidrectional.\n \nA while back, I tried to write an extension to mercurial that would export a\nhg repo using the git protocol. One side-effect was that it converted the\nentire repository to a git repo with many loose objects.\n\nIt \"worked\" (I never finished it enough) on a small repo in a bidirectional\nway.\n\nI'll try to dig up the code, and put it up somewhere...\n\nJosef \"Jeff\" Sipek.\n\n-- \nI'm somewhere between geek and normal.\n\t\t- Linus Torvalds\n"},{"id":"36504","messageId":"20070307001105.GJ18370@thunk.org","threadId":"7115","inReplyTo":"20070306230802.GA17226@filer.fsl.cs.sunysb.edu","subject":"Re: mercurial to git","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2007-03-07T00:11:05Z","receivedAt":"2007-03-07T00:11:05Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Tue, Mar 06, 2007 at 06:08:02PM -0500, Josef Sipek wrote:\n> I'll try to dig up the code, and put it up somewhere...\n\nHere's a hacked up version of Stelian Pop's converter code that I used\nfor an initial test conversion of e2fsprogs from hg to git.  The main\nimprovements from Stelian's is that it's a bit faster by caching the\nresults of \"hg log\", and that it handles parses the Signed-off-by:\nheaders to feed in into the ChangeSet's Author identity (as distinct\nfrom the committer identity, which it gets from the hg information).\n\nThe other change which I added was add a pretty kludgy committer name\ncannonicalizer, since there the commiter information dates is pretty\ngrotty.  That's because the e2fsprogs source repository has over the\nyears been converted from CVS, to BitKeeper, to Mercurial, and now at\nsome point soon when I'm happy with a decent hg-to-git tool, to git.\n\nMy plan was to rewrite the converter to call Mercurial's python\nclasses directly (using the equivalent python code to 'hg manifest'\nand 'hg cat' to speed things up enormously, compared to checking out\neach revision one at a time and then using git to figure out which\nfiles had been added/changed/deleted), and to interface it into\ngit-fast-import, and make the necessary changes (including more\nintelligent handling of .hgtags) so that the conversion could be\nbidrectional.\n\nBut if I can convince someone else to do the work, especially if their\nconverter handles the Signed-off-by: parsing, and making sure the\nauthor and commit dates are properly set, that would certainly be a\nbonus.  :-)\n\n\t\t\t\t\t\t- Ted\n\nP.S.  Oh yes, my plan was to use Python's ConfigParser class to store\nthe author cannonicalization information, instead of hard-coding the\ndata into the python script.  Code snippets to do this available on\nrequest; it was pretty trivial to do.\n\n#! /usr/bin/python\n\n\"\"\" hg-to-git.py - A Mercurial to GIT converter\n\n    Copyright (C)2007 Stelian Pop <stelian@xxxxxxxxxx>\n\n    This program is free software; you can redistribute it and/or modify\n    it under the terms of the GNU General Public License as published by\n    the Free Software Foundation; either version 2, or (at your option)\n    any later version.\n\n    This program is distributed in the hope that it will be useful,\n    but WITHOUT ANY WARRANTY; without even the implied warranty of\n    MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the\n    GNU General Public License for more details.\n\n    You should have received a copy of the GNU General Public License\n    along with this program; if not, write to the Free Software\n    Foundation, Inc., 675 Mass Ave, Cambridge, MA 02139, USA.\n\"\"\"\n\nimport os, os.path, sys\nimport tempfile, popen2, pickle, getopt\nimport re\n\n# Maps hg version -> git version\nhgvers = {}\n# List of children for each hg revision\nhgchildren = {}\n# Current branch for each hg revision\nhgbranch = {}\n\n#------------------------------------------------------------------------------\n\ndef usage():\n\n        print \"\"\"\\\n%s: [OPTIONS] <hgprj>\n\noptions:\n    -s, --gitstate=FILE: name of the state to be saved/read\n                         for incrementals\n\nrequired:\n    hgprj:  name of the HG project to import (directory)\n\"\"\" % sys.argv[0]\n\n#------------------------------------------------------------------------------\n\ndef getgitenv(user, author, date):\n    env = ''\n    if author == '':\n\tauthor = user\n\n    elems = re.compile('(.*?)\\s+<(.*)>').match(user)\n    if elems:\n        env += 'export GIT_COMMITER_NAME=\"%s\" ;' % elems.group(1)\n        env += 'export GIT_COMMITER_EMAIL=\"%s\" ;' % elems.group(2)\n    else:\n        env += 'export GIT_COMMITER_NAME=\"%s\" ;' % user\n        env += 'export GIT_COMMITER_EMAIL= ;'\n\n    elems = re.compile('(.*?)\\s+<(.*)>').match(author)\n    if elems:\n        env += 'export GIT_AUTHOR_NAME=\"%s\" ;' % elems.group(1)\n        env += 'export GIT_AUTHOR_EMAIL=\"%s\" ;' % elems.group(2)\n    else:\n        env += 'export GIT_AUTHOR_NAME=\"%s\" ;' % author\n        env += 'export GIT_AUTHOR_EMAIL= ;'\n\n    env += 'export GIT_AUTHOR_DATE=\"%s\" ;' % date\n    env += 'export GIT_COMMITTER_DATE=\"%s\" ;' % date\n    return env\n\n#------------------------------------------------------------------------------\n\nstate = ''\n\ntry:\n    opts, args = getopt.getopt(sys.argv[1:], 's:t:', ['gitstate=', 'tempdir='])\n    for o, a in opts:\n        if o in ('-s', '--gitstate'):\n            state = a\n            state = os.path.abspath(state)\n\n    if len(args) != 1:\n        raise('params')\nexcept:\n    usage()\n    sys.exit(1)\n\nhgprj = args[0]\nos.chdir(hgprj)\n\nif state:\n    if os.path.exists(state):\n        print 'State does exist, reading'\n        f = open(state, 'r')\n        hgvers = pickle.load(f)\n    else:\n        print 'State does not exist, first run'\n\ntip = os.popen('hg tip | head -1 | cut -f 2 -d :').read().strip()\nprint 'tip is', tip\n\n# Calculate the branches\nprint 'analysing the branches...'\nhgchildren[\"0\"] = ()\nhgbranch[\"0\"] = \"master\"\nfor cset in range(1, int(tip) + 1):\n    hgchildren[str(cset)] = ()\n    prnts = os.popen('hg log -r %d | grep ^parent: | cut -f 2 -d :' % cset).readlines()\n    if len(prnts) > 0:\n        parent = prnts[0].strip()\n    else:\n        parent = str(cset - 1)\n    hgchildren[parent] += ( str(cset), )\n    if len(prnts) > 1:\n        mparent = prnts[1].strip()\n        hgchildren[mparent] += ( str(cset), )\n    else:\n        mparent = None\n\n    if mparent:\n        # For merge changesets, take either one, preferably the 'master' branch\n        if hgbranch[mparent] == 'master':\n            hgbranch[str(cset)] = 'master'\n        else:\n            hgbranch[str(cset)] = hgbranch[parent]\n    else:\n        # Normal changesets\n        # For first children, take the parent branch, for the others create a new branch\n        if hgchildren[parent][0] == str(cset):\n            hgbranch[str(cset)] = hgbranch[parent]\n        else:\n            hgbranch[str(cset)] = \"branch-\" + str(cset)\n\nif not hgvers.has_key(\"0\"):\n    print 'creating repository'\n    os.system('git-init-db')\n\n# loop through every hg changeset\nfor cset in range(int(tip) + 1):\n\n    # incremental, already seen\n    if hgvers.has_key(str(cset)):\n        continue\n\n    # get info\n    prnts = os.popen('hg log -r %d | grep ^parent: | cut -f 2 -d :' % cset).readlines()\n    if len(prnts) > 0:\n        parent = prnts[0].strip()\n    else:\n        parent = str(cset - 1)\n    if len(prnts) > 1:\n        mparent = prnts[1].strip()\n    else:\n        mparent = None\n\n    (fdlog, filelog) = tempfile.mkstemp()\n    logtxt = os.popen('hg log -r %d -v' % cset).read().strip()\n    os.write(fdlog, logtxt)\n    os.close(fdlog)\n\n    (fdcomment, filecomment) = tempfile.mkstemp()\n    csetcomment = os.popen('grep -v ^changeset: < %s | grep -v ^parent: | grep -v ^user: | grep -v ^date | grep -v ^files: | grep -v ^description: | grep -v ^tag:' % filelog).read().strip()\n    os.write(fdcomment, csetcomment)\n    os.close(fdcomment)\n\n    date = os.popen('grep -m 1 ^date: < %s | cut -f 2- -d :' % filelog).read().strip()\n\n    tag = os.popen('grep -m 1 ^tag: < %s | cut -f 2- -d :' % filelog).read().strip()\n\n    user = os.popen('grep -m 1 ^user: < %s | cut -f 2- -d :' % filelog).read().strip()\n    if user == 'tytso@mit.edu':\n\tuser = \"Theodore Ts'o <tytso@mit.edu>\"\n    if user == 'tytso@think.thunk.org':\n\tuser = \"Theodore Ts'o <tytso@mit.edu>\"\n    if user == 'tytso@snap.thunk.org':\n\tuser = \"Theodore Ts'o <tytso@mit.edu>\"\n    if user == 'tytso@fs.thunk.org':\n\tuser = \"Theodore Ts'o <tytso@mit.edu>\"\n    if user == 'tytso@voltaire.debian.org':\n\tuser = \"Theodore Ts'o <tytso@mit.edu>\"\n    if user == 'tytso@who-could-of.thunk.org':\n\tuser = \"Theodore Ts'o <tytso@mit.edu>\"\n    if user == 'tytso@universal.(none)':\n\tuser = \"Theodore Ts'o <tytso@mit.edu>\"\n    if user == 'tytso@theodore-tsos-computer.local':\n\tuser = \"Theodore Ts'o <tytso@mit.edu>\"\n\n    if user == 'adilger@clusterfs.com':\n\tuser = \"Andreas Dilger <adilger@clusterfs.com>\"\n    if user == 'adilger@lynx.adilger.int':\n\tuser = \"Andreas Dilger <adilger@clusterfs.com>\"\n    if user == 'root@lynx.adilger.int':\n\tuser = \"Andreas Dilger <adilger@clusterfs.com>\"\n    if user == 'matthias.andree@gmx.de':\n\tuser = \"Matthias Andree <matthias.andree@gmx.de>\"\n\n    if user == 'laptop@duncow.home.oldelvet.org.uk':\n\tuser = \"Richard Mortimer <richm@oldelvet.org.uk>\"\n\n    if user == 'sct@redhat.com':\n\tuser = 'Stephen Tweedie <sct@redhat.com>'\n    if user == 'sct@sisko.scot.redhat.com':\n\tuser = 'Stephen Tweedie <sct@redhat.com>'\n\n    if user == 'paubert@gra-vd1.iram.es':\n\tuser = 'Gabriel Paubert <paubert@iram.es>'\n\n    author = os.popen('grep -m 1 ^Signed-off-by: < %s | cut -f 2- -d :' % filelog).read().strip()\n    if author == '\"Theodore Ts\\'o\" <tytso@mit.edu>':\n\tauthor = \"Theodore Ts'o <tytso@mit.edu>\"\n\n    os.unlink(filelog)\n\n    print '-----------------------------------------'\n    print 'cset:', cset\n    print 'branch:', hgbranch[str(cset)]\n    print 'user:', user\n    print 'author:', author\n    print 'date:', date\n    print 'comment:', csetcomment\n    print 'parent:', parent\n    if mparent:\n        print 'mparent:', mparent\n    if tag:\n        print 'tag:', tag\n    print '-----------------------------------------'\n\n    # checkout the parent if necessary\n    if cset != 0:\n        if hgbranch[str(cset)] == \"branch-\" + str(cset):\n            print 'creating new branch', hgbranch[str(cset)]\n            os.system('git-checkout -b %s %s' % (hgbranch[str(cset)], hgvers[parent]))\n        else:\n            print 'checking out branch', hgbranch[str(cset)]\n            os.system('git-checkout %s' % hgbranch[str(cset)])\n\n    # merge\n    if mparent:\n        if hgbranch[parent] == hgbranch[str(cset)]:\n            otherbranch = hgbranch[mparent]\n        else:\n            otherbranch = hgbranch[parent]\n        print 'merging', otherbranch, 'into', hgbranch[str(cset)]\n        os.system(getgitenv(user, author, date) + 'git-merge --no-commit -s ours \"\" %s %s' % (hgbranch[str(cset)], otherbranch))\n\n    # remove everything except .git and .hg directories\n    os.system('find . \\( -path \"./.hg\" -o -path \"./.git\" \\) -prune -o ! -name \".\" -print | xargs rm -rf')\n\n    # repopulate with checkouted files\n    os.system('hg update -C %d' % cset)\n\n    # add new files\n    os.system('git-ls-files -x .hg --others | git-update-index --add --stdin')\n    # delete removed files\n    os.system('git-ls-files -x .hg --deleted | git-update-index --remove --stdin')\n\n    # commit\n    os.system(getgitenv(user, author, date) + 'git-commit -a -F %s' % filecomment)\n    os.unlink(filecomment)\n\n    # tag\n    if tag and tag != 'tip':\n        os.system(getgitenv(user, author, date) + 'git-tag %s' % tag)\n\n    # delete branch if not used anymore...\n    if mparent and len(hgchildren[str(cset)]):\n        print \"Deleting unused branch:\", otherbranch\n        os.system('git-branch -d %s' % otherbranch)\n\n    # retrieve and record the version\n    vvv = os.popen('git-show | head -1').read()\n    vvv = vvv[vvv.index(' ') + 1 : ].strip()\n    print 'record', cset, '->', vvv\n    hgvers[str(cset)] = vvv\n\nos.system('git-repack -a -d')\n\n# write the state for incrementals\nif state:\n    print 'Writing state'\n    f = open(state, 'w')\n    pickle.dump(hgvers, f)\n\n# vim: et ts=8 sw=4 sts=4\n"},{"id":"36549","messageId":"20070307155929.GD27596@spearce.org","threadId":"7115","inReplyTo":"20070306210629.GA42331@peter.daprodeges.fqdn.th-h.de","subject":"Re: mercurial to git","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-03-07T15:59:29Z","receivedAt":"2007-03-07T15:59:29Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Rocco Rutte <pdmef@gmx.net> wrote:\n> The performance bottleneck is hg exporting data, as discovered by people \n> on #mercurial, the problem is not really fixable and is due to hg's \n> revlog handling. As a result, I needed to let the script feed the full \n> contents of the repository at each revision we walk (i.e. all for the \n> initial import) into git-fast-import.\n\nI thought that hg stored file revisions such that each source file\n(e.g. foo.c) had its own revision file (e.g. foo.revdata) and that\nevery revision of foo.c was stored in that one file, ordered from\noldest to newest?  If that is the case why not strip all of those\ninto fast-import up front, doing one source file at a time as a\nhuge series of blobs and mark them, then do the commit/trees later\non using only the marks?\n\nOr am I just missing something about hg?\n\n> This is horribly slow. For mutt \n> which contains several tags, a handfull of branches and only 5k commits \n> this takes roughly two hours at 1 commit/sec.\n\nNot fast-import's fault.  ;-)\n\n> Somewhat related: It would be really nice to teach git-fast-import to \n> init from a previously saved mark file. Right now I use hg revision \n> numbers as marks, let git-fast-import save them, and read them back next \n> time. These are needed to map hg revisions to git SHA1s in case I need \n> to reference something in an incremental import from an earlier run. It \n> would be nice if git-fast-import could do this on its own so that all \n> consumers can benefit and can have persistent marks accross sessions.\n\nSure, that sounds pretty easy.  I'll try to work that up later\ntoday or tomorrow.\n \n-- \nShawn.\n"},{"id":"36579","messageId":"20070307231421.GF27922@spearce.org","threadId":"7115","inReplyTo":"20070306210629.GA42331@peter.daprodeges.fqdn.th-h.de","subject":"Re: mercurial to git","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-03-07T23:14:21Z","receivedAt":"2007-03-07T23:14:21Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Rocco Rutte <pdmef@gmx.net> wrote:\n> Somewhat related: It would be really nice to teach git-fast-import to \n> init from a previously saved mark file. Right now I use hg revision \n> numbers as marks, let git-fast-import save them, and read them back next \n> time. These are needed to map hg revisions to git SHA1s in case I need \n> to reference something in an incremental import from an earlier run. It \n> would be nice if git-fast-import could do this on its own so that all \n> consumers can benefit and can have persistent marks accross sessions.\n\nDone.  See the new --import-marks option.\n\nThe following changes since commit c390ae97beb9e8cdab159b593ea9659e8096c4db:\n  Li Yang (1):\n        gitweb: Change to use explicitly function call cgi->escapHTML()\n\nare found in the git repository at:\n\n  git://repo.or.cz:/git/fastimport.git\n\nShawn O. Pearce (3):\n      Preallocate memory earlier in fast-import\n      Use atomic updates to the fast-import mark file\n      Allow fast-import frontends to reload the marks table\n\n Documentation/git-fast-import.txt |   13 +++++-\n fast-import.c                     |   85 +++++++++++++++++++++++++++++-------\n t/t9300-fast-import.sh            |    8 ++++\n 3 files changed, 88 insertions(+), 18 deletions(-)\n\n-- \nShawn.\n"},{"id":"36598","messageId":"20070308085613.GA2881@peter.daprodeges.fqdn.th-h.de","threadId":"7115","inReplyTo":"20070307155929.GD27596@spearce.org","subject":"Re: mercurial to git","fromName":"Rocco Rutte","fromEmail":"pdmef@gmx.net","sentAt":"2007-03-08T08:56:14Z","receivedAt":"2007-03-08T08:56:14Z","isPatch":false,"sender":{"key":"pdmef@gmx.net","avatar":null},"body":"Hi,\n\n* Shawn O. Pearce [07-03-07 10:59:29 -0500] wrote:\n\n>I thought that hg stored file revisions such that each source file\n>(e.g. foo.c) had its own revision file (e.g. foo.revdata) and that\n>every revision of foo.c was stored in that one file, ordered from\n>oldest to newest?  If that is the case why not strip all of those\n>into fast-import up front, doing one source file at a time as a\n>huge series of blobs and mark them, then do the commit/trees later\n>on using only the marks?\n\n>Or am I just missing something about hg?\n\nI don't want to use anything except the hg mecurial API so that in \ntheory the importer could work even for remote hg repositories.\n\nBut the \"blob feed\" approach doesn't seem perfectly right to me \nespecially for incremental imports. There would have to be state files \nand internal tables telling what revisions of what files there are with \nwhat content. With thousands of files I think this gets quite messy to \nfind even the minimum set to start of with for an incremental import. \nAlso, you can already specify up to which revision to import so it would \nget even more complicated.\n\n   bye, Rocco\n-- \n:wq!\n"},{"id":"36599","messageId":"20070308090107.GB2881@peter.daprodeges.fqdn.th-h.de","threadId":"7115","inReplyTo":"20070306215459.GI18370@thunk.org","subject":"Re: mercurial to git","fromName":"Rocco Rutte","fromEmail":"pdmef@gmx.net","sentAt":"2007-03-08T09:01:07Z","receivedAt":"2007-03-08T09:01:07Z","isPatch":false,"sender":{"key":"pdmef@gmx.net","avatar":null},"body":"Hi,\n\n* Theodore Tso [07-03-06 16:54:59 -0500] wrote:\n\n>Hmm.... the way I was planning on handling the performance bottleneck\n>was to use \"hg manifest --debug <rev>\" and diffing the hashes against\n>its parents.  Using \"hg manifest\" only hits .hg/00manifest.[di] and\n>.hg/00changelog.[di] files, so it's highly efficient.  With the\n>--debug option to hg manifest (not needed on some earlier versions of\n>hg, but it seems to be needed on the latest development version of\n>hg), it outputs the mode and SHA1 hash of the files, so it becomes\n>easy to see which files were changed relative to the revision's\n>parent(s).\n\n>Once we know which files we need to feed to git-fast-import, it's just\n>a matter of using \"hg cat -r <rev> <pathname>\" to feed the individual\n>changed file to git-fast-import.\n\nI've done that now and the repositories come out as before in about 10 \nminutes. Also I sanitized the tags handling and will push out the \nchanged version somewhere soon.\n\n   bye, Rocco\n-- \n:wq!\n"},{"id":"36617","messageId":"20070308104947.GC2881@peter.daprodeges.fqdn.th-h.de","threadId":"7115","inReplyTo":"20070306210629.GA42331@peter.daprodeges.fqdn.th-h.de","subject":"Re: mercurial to git","fromName":"Rocco Rutte","fromEmail":"pdmef@gmx.net","sentAt":"2007-03-08T10:49:47Z","receivedAt":"2007-03-08T10:49:47Z","isPatch":false,"sender":{"key":"pdmef@gmx.net","avatar":null},"body":"Hi,\n\n* Rocco Rutte [07-03-06 21:06:29 +0000] wrote:\n\n[...]\n\nI've now pushed the changes out to:\n\n   http://repo.or.cz/w/hg2git.git\n\nI don't know that the status is and/or future plans are for:\n\n   http://repo.or.cz/w/fast-export.git\n\n...but these two seem worth combining, IMHO.\n\nI haven't followed git development consequently lately, so are there any \nplans of including these or replacing current importers by these?\n\n   bye, Rocco\n-- \n:wq!\n"},{"id":"37133","messageId":"20070315002505.GA31770@thunk.org","threadId":"7115","inReplyTo":"20070314111257.GA4526@peter.daprodeges.fqdn.th-h.de","subject":"Re: mercurial to git","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2007-03-15T00:25:07Z","receivedAt":"2007-03-15T00:25:07Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Wed, Mar 14, 2007 at 11:12:57AM +0000, Rocco Rutte wrote:\n> \n> I tried the import on the e2fsprogs repo and the files come out \n> identical, authors/comitters look okay to me, too.\n\nVery cool!  It looks like some of the git author dates are only\ngetting set if the -s flag is set.  Was that intentional?\n\n\t\t\t\t\t\t- Ted\n"},{"id":"37155","messageId":"20070315101913.GA9831@peter.daprodeges.fqdn.th-h.de","threadId":"7115","inReplyTo":"20070315002505.GA31770@thunk.org","subject":"Re: mercurial to git","fromName":"Rocco Rutte","fromEmail":"pdmef@gmx.net","sentAt":"2007-03-15T10:19:13Z","receivedAt":"2007-03-15T10:19:13Z","isPatch":false,"sender":{"key":"pdmef@gmx.net","avatar":null},"body":"Hi,\n\n* Theodore Tso [07-03-14 20:25:07 -0400] wrote:\n>On Wed, Mar 14, 2007 at 11:12:57AM +0000, Rocco Rutte wrote:\n\nI failed to send a response to the list and it went Theodore privately \nonly, sorry. I merged hg2git into fast-export.git at repo.or.cz and \nnamed it 'hg-fast-export' to match with the other importers there. It \nnow can parse Signed-off-by lines and supports author maps (as \ngit-cvsimport and git-svnimport do, same syntax).\n\n>> I tried the import on the e2fsprogs repo and the files come out \n>> identical, authors/comitters look okay to me, too.\n\n>Very cool!  It looks like some of the git author dates are only\n>getting set if the -s flag is set.  Was that intentional?\n\nFor which changesets exactly? The script only attempts to write out the \n'author' command if -s (for parsing signed-off-by) is given. But for \nboth commands the time information written out are identical and are \nexactly what hg gives us. So the bug must be elsewhere.\n\n   bye, Rocco\n-- \n:wq!\n"},{"id":"37161","messageId":"20070315141227.GA18416@thunk.org","threadId":"7115","inReplyTo":"20070315101913.GA9831@peter.daprodeges.fqdn.th-h.de","subject":"Re: mercurial to git","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2007-03-15T14:12:27Z","receivedAt":"2007-03-15T14:12:27Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Thu, Mar 15, 2007 at 10:19:13AM +0000, Rocco Rutte wrote:\n> I failed to send a response to the list and it went Theodore privately \n> only, sorry. I merged hg2git into fast-export.git at repo.or.cz and \n> named it 'hg-fast-export' to match with the other importers there. It \n> now can parse Signed-off-by lines and supports author maps (as \n> git-cvsimport and git-svnimport do, same syntax).\n\nBTW, there are a number of places where the old name (hg2git) is still\nbeing used for filenames, et.al, because $PFX is still being set to\nhg2git.  \n\n> For which changesets exactly? The script only attempts to write out the \n> 'author' command if -s (for parsing signed-off-by) is given. But for \n> both commands the time information written out are identical and are \n> exactly what hg gives us. So the bug must be elsewhere.\n\nAll of them.  :-)\n\nUpon doing more investigation, the failure case seems to be if -A is\nspecified but NOT -s.  Comare:\n\n(Generated using: hg-fast-export.sh -A ../e2fsprogs.authors -r ../e2fsprogs)\ncommit b584b9c57ecbbeef91970ca2924d66662029ab29\nAuthor: Theodore Ts'o <tytso@mit.edu>\nDate:   Thu Jan 1 00:00:00 1970 +0000  <=============\n\nwith\n\n(Generated using: hg-fast-export.sh -s -A ../e2fsprogs.authors -r ../e2fsprogs)\ncommit 9e9a5867e4d4985bde6d6be072efb96e901e08cc\nAuthor: Theodore Ts'o <tytso@mit.edu>\nDate:   Wed Mar 7 08:09:10 2007 -0500  <=============\n\nThe date seems to be correctly generated using\n\n\thg-fast-export.sh -s -A ../e2fsprogs.authors -r ../e2fsprogs\n\thg-fast-export.sh -s -r ../e2fsprogs\n\thg-fast-export.sh -r ../e2fsprogs\n\nIt seems to be this combination of options:\n\n\thg-fast-export.sh -A ../e2fsprogs.authors -r ../e2fsprogs\n\nWhere all of the dates end up being Jan 1, 1970.\n\nRegards,\n\n\t\t\t\t\t\t- Ted\n"},{"id":"37165","messageId":"20070315151936.GA31087@peter.daprodeges.fqdn.th-h.de","threadId":"7115","inReplyTo":"20070315141227.GA18416@thunk.org","subject":"Re: mercurial to git","fromName":"Rocco Rutte","fromEmail":"pdmef@gmx.net","sentAt":"2007-03-15T15:19:37Z","receivedAt":"2007-03-15T15:19:37Z","isPatch":false,"sender":{"key":"pdmef@gmx.net","avatar":null},"body":"Hi,\n\n* Theodore Tso [07-03-15 10:12:27 -0400] wrote:\n>On Thu, Mar 15, 2007 at 10:19:13AM +0000, Rocco Rutte wrote:\n>> I failed to send a response to the list and it went Theodore privately \n>> only, sorry. I merged hg2git into fast-export.git at repo.or.cz and \n>> named it 'hg-fast-export' to match with the other importers there. It \n>> now can parse Signed-off-by lines and supports author maps (as \n>> git-cvsimport and git-svnimport do, same syntax).\n\n>BTW, there are a number of places where the old name (hg2git) is still\n>being used for filenames, et.al, because $PFX is still being set to\n>hg2git.\n\nI know and intend to leave it that way as the filenames are shorter.\n\n>The date seems to be correctly generated using\n\n>\thg-fast-export.sh -s -A ../e2fsprogs.authors -r ../e2fsprogs\n>\thg-fast-export.sh -s -r ../e2fsprogs\n>\thg-fast-export.sh -r ../e2fsprogs\n\n>It seems to be this combination of options:\n\n>\thg-fast-export.sh -A ../e2fsprogs.authors -r ../e2fsprogs\n\n>Where all of the dates end up being Jan 1, 1970.\n\nHmm. Strange, I cannot reproduce this. Can you mail me your authors file \nalong with version information privately please?\n\n   bye, Rocco\n-- \n:wq!\n"},{"id":"37168","messageId":"Pine.LNX.4.64.0703150854520.3816@woody.linux-foundation.org","threadId":"7115","inReplyTo":"20070315141227.GA18416@thunk.org","subject":"Re: mercurial to git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-03-15T15:56:59Z","receivedAt":"2007-03-15T15:56:59Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 15 Mar 2007, Theodore Tso wrote:\n>\n> (Generated using: hg-fast-export.sh -A ../e2fsprogs.authors -r ../e2fsprogs)\n> commit b584b9c57ecbbeef91970ca2924d66662029ab29\n> Author: Theodore Ts'o <tytso@mit.edu>\n> Date:   Thu Jan 1 00:00:00 1970 +0000  <=============\n\nIs the committer date ok (does it even exist)? Use \"--pretty=fuller\" or \nperhaps even \"--pretty=raw\" to see both author and committer date.\n\nThe normal log only shows author date, since that's usually the one people \ncare about (git itself doesn't at all, it uses the committer date-stamp to \ndo it's \"choose most recent path to go down\").\n\n\t\tLinus\n"},{"id":"37177","messageId":"20070315210406.GA8568@thunk.org","threadId":"7115","inReplyTo":"20070315094434.GA4425@peter.daprodeges.fqdn.th-h.de","subject":"Re: mercurial to git","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2007-03-15T21:04:07Z","receivedAt":"2007-03-15T21:04:07Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"Hopefully you won't mind that I'm adding the git list back to the cc\nline, since it would be useful for others to provide some feedback.\n\nOn Thu, Mar 15, 2007 at 09:44:35AM +0000, Rocco Rutte wrote:\n> >So I'll go try it out in the near future.  Are you planning on being\n> >able to make it be bi-directional?  (i.e., so that changes in the git\n> >tree can get propagated back to the hg tree?)\n> \n\n> But as there's no hg-fast-import, I think git to hg not so trivial to \n> implement and convert-repo already exists, so I'd rather prefer \n> extending it to do the job.\n\nActually, there *is* an hg-fast-import.  It exists in the hg sources\nin contrib/convert-repo, and it is being used in production to do\nincremental conversion from the Linux kernel git tree to an hg tree.\nSo it does handle octopus merges already (it has to, the ACPI folks\nare very ocotpus merge happy :-).\n\n> However, I never even used hg and have only some knowledge about the API \n> so that I see some difficulties and need more time to think about it \n> (e.g. how to detect whether a change in hg originates at git and vice \n> versa, what to do with octopus merges, cherry-picks, etc).\n\nSo actually I have thought about this a fair amount, so if you don't\nmind my pontificating a bit.   :-)\n\nAt the highest architectural viewpoint, there are three levels of\ndifficulty of SCM conversions:\n\nA) One-way conversion utilities.  Examples of this would be the\n\thg2git, hg-fast-import scripts that convert from hg to git,\n\tand the convert-repo script which will convert from git to hg.\n\nB) Single point bidrectional conversion.  At this level, the hg/git\n\tgateway will run on a single machine, and with a state file,\n\tcan recognize new git changesets, and create a functionally\n\tequivalent hg changeset and commit it to the hg repository,\n\tand can also recognize new hg changeset, and create a\n\tfunctionaly equivalent git changeset, and commit it to the git\n\trepository.  \n\nC) Multisite bidirectional conversion.  At this level, multiple users\n\tbe gatewaying between the two DSCM systems, and as long as\n\tthey are using the same configuration parameters (more on this\n\tin a moment), if user A converts changeset from hg to git, and\n\tthat changeset is passed along via git to user B, who then\n\trunning the birectional gateway program, converts it back from\n\tgit to hg, the hg changeset is identical so that hg recognizes\n\tis the same changeset when it gets propgated back to user A.\n\n(C) would be the ideal, given the distributed nature of hg and git.\nIt is also the most difficult, because it means that we need to be\nable to do a lossless, one-to-one conversion of a Changeset.  It is\nalso somewhat at odds with doing author mapping and signed-off-by\nparsing, since that could prevent a reversible transformation.\nHowever, what may very well be common for projects is for them to\nstart with (B), and to convert over some of the historical changesets,\nand then later on allow multiple users to clone from the two git/hg\nrepositories and then do the multisite conversion.\n\nSo what that also means is that even if we only do (B) at first, it\nmight be useful if we have some of the characteristics needed to\neventually get to (C), even if we can't get there right away.\n\nSo more practially, here are some of the things that we would need to\ndo, looking at hg-fast-export:\n\n*) Change the index/marks file to map between hg SHA hash ID's instead\nof the small integer ordinals.  This is useful for enabling multisite\nconversion, but it is also useful for tracking tag changes in .hgtags.\n\n*) Have a mode so that instead of only checking changes greater than\nlast run, to simply iterate over all changesets in mercurial and check\nto see if hg SHA1 commit ID is already in the marks file; if so, skip\nit.  \n\n*) Have a mode where the COMMITER id is \"hg2git\" and the COMMITER_DATE\nis the same as the AUTHOR_DATE (so that the changelog converesion is\nthe same no matter where or who does the converation).  This is mainly\nto enable multisite converstaion.\n\n\t\t\t\t\t\t- Ted\n"},{"id":"37180","messageId":"20070315220705.GD31087@peter.daprodeges.fqdn.th-h.de","threadId":"7115","inReplyTo":"20070315210406.GA8568@thunk.org","subject":"Re: mercurial to git","fromName":"Rocco Rutte","fromEmail":"pdmef@gmx.net","sentAt":"2007-03-15T22:07:05Z","receivedAt":"2007-03-15T22:07:05Z","isPatch":false,"sender":{"key":"pdmef@gmx.net","avatar":null},"body":"Hi,\n\n* Theodore Tso [07-03-15 17:04:07 -0400] wrote:\n>Hopefully you won't mind that I'm adding the git list back to the cc\n>line, since it would be useful for others to provide some feedback.\n\nNot at all. Just wondering when others would get too bored... :)\n\n>Actually, there *is* an hg-fast-import.  It exists in the hg sources\n>in contrib/convert-repo, and it is being used in production to do\n>incremental conversion from the Linux kernel git tree to an hg tree.\n>So it does handle octopus merges already (it has to, the ACPI folks\n>are very ocotpus merge happy :-).\n\nI know convert-repo and like it as a starting point. But it has some \nproblems like not properly creating hg branches, can import only one \nbranch at a time which must also be checkout out on the git side, etc.\n\nWith 'hg-fast-import' I meant something like git-fast-import where \nclients can feed in more raw data instead of preparing each commit on \nits own and comitting it.\n\n[...]\n\n>So more practially, here are some of the things that we would need to\n>do, looking at hg-fast-export:\n\n>*) Change the index/marks file to map between hg SHA hash ID's instead\n>of the small integer ordinals.  This is useful for enabling multisite\n>conversion, but it is also useful for tracking tag changes in .hgtags.\n\nThe small numbers are the hg revision numbers which we'll need for \ngit-fast-import. Ideally git-fast-import would allow us to use anything \nfor a mark we want as long as it's unique. But I'm sure there's a cheap \nway of mapping revision to SHA1 in the hg API.\n\nSo, if anybody wants to join in writing up such a hybrid system, I'm for \nit. :)\n\n   bye, Rocco\n-- \n"},{"id":"37201","messageId":"200703160053.54699.lenb@kernel.org","threadId":"7115","inReplyTo":"20070315210406.GA8568@thunk.org","subject":"Re: mercurial to git","fromName":"Len Brown","fromEmail":"lenb@kernel.org","sentAt":"2007-03-16T04:53:54Z","receivedAt":"2007-03-16T04:53:54Z","isPatch":false,"sender":{"key":"lenb@kernel.org","avatar":null},"body":"On Thursday 15 March 2007 17:04, Theodore Tso wrote:\n> So it does handle octopus merges already (it has to, the ACPI folks\n> are very ocotpus merge happy :-).\n\nWell, just to set the record straight...\n\nSo yes, I did a 12-way merge in the kernel a long while back on a lark.\nI don't generally do them any more in the official kernel tree\nbecause I think they make bisect more complicated than it needs to be.\n\ncheers,\n-Len\n"},{"id":"37310","messageId":"45FBD2F6.30701@fs.ei.tum.de","threadId":"7115","inReplyTo":"20070315220705.GD31087@peter.daprodeges.fqdn.th-h.de","subject":"Re: mercurial to git","fromName":"Simon 'corecode' Schubert","fromEmail":"corecode@fs.ei.tum.de","sentAt":"2007-03-17T11:37:26Z","receivedAt":"2007-03-17T11:37:26Z","isPatch":false,"sender":{"key":"corecode@fs.ei.tum.de","avatar":"https://gravatar.com/avatar/eff9dbf0cdac0d1e6a6cd7ed0e50763edcb376b493b5253a35ff167918ad79e1?d=mp&s=160"},"body":"Rocco Rutte wrote:\n> With 'hg-fast-import' I meant something like git-fast-import where \n> clients can feed in more raw data instead of preparing each commit on \n> its own and comitting it.\n\nI wrote something like that for fromcvs: <http://ww2.fs.ei.tum.de/~corecode/hg/fromcvs?f=7992019d6861;file=tohg.py>\n\nit still communicates blobs via files, because that's what hg wants to see.  performance is quite okay.\n\ncheers\n  simon\n\n-- \nServe - BSD     +++  RENT this banner advert  +++    ASCII Ribbon   /\"\\\nWork - Mac      +++  space for low €€€ NOW!1  +++      Campaign     \\ /\nParty Enjoy Relax   |   http://dragonflybsd.org      Against  HTML   \\\nDude 2c 2 the max   !   http://golden-apple.biz       Mail + News   / \\\n\n"}]}