{"thread":{"id":"11115","subject":"git-svn: .git/svn disk usage","startedAt":"2007-12-03T06:17:28Z","lastAt":"2007-12-06T06:47:17Z","messageCount":10,"participants":["Ollie Wild","Pascal Obry","David Brown","Kelvie Wong","David Voit","Karl Hasselström","Eric Wong","Steven Grimm"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"61735","messageId":"65dd6fd50712022217l5f807f31pf3f00d82c3dccf5c@mail.gmail.com","threadId":"11115","inReplyTo":null,"subject":"git-svn: .git/svn disk usage","fromName":"Ollie Wild","fromEmail":"aaw@google.com","sentAt":"2007-12-03T06:17:28Z","receivedAt":"2007-12-03T06:17:28Z","isPatch":false,"sender":{"key":"aaw@google.com","avatar":null},"body":"Hi,\n\nI've been using git-svn to mirror the gcc repository at\nsvn://gcc.gnu.org/svn/gcc.  Recently, I noticed that my .git directory\nis consuming 11GB of disk space.  Digging further, I discovered that\n9.8GB of this is attributable to the .git/svn directory (which\nincludes 200 branches and 2,588 tags).  Given that my .git/objects\ndirectory is 652MB, it seems that it ought to be possible to store\nthis information in a more compact form.\n\nI'm curious if other developers have run into this issue.  If so, are\nthere any proposals / plans for improving the storage of git-svn\nmetadata?\n\nThanks,\nOllie\n"},{"id":"61738","messageId":"4753A43F.9060303@obry.net","threadId":"11115","inReplyTo":"65dd6fd50712022217l5f807f31pf3f00d82c3dccf5c@mail.gmail.com","subject":"Re: git-svn: .git/svn disk usage","fromName":"Pascal Obry","fromEmail":"pascal@obry.net","sentAt":"2007-12-03T06:37:51Z","receivedAt":"2007-12-03T06:37:51Z","isPatch":false,"sender":{"key":"pascal@obry.net","avatar":"https://avatars.githubusercontent.com/u/467069?v=4"},"body":"Ollie,\n\n> I'm curious if other developers have run into this issue.  If so, are\n> there any proposals / plans for improving the storage of git-svn\n> metadata?\n\nDid you run \"git gc\" after importing code form the subversion\nrepository? On my side I found that it has reduced drastically the size\nof the local Git repository.\n\nPascal.\n\n-- \n\n--|------------------------------------------------------\n--| Pascal Obry                           Team-Ada Member\n--| 45, rue Gabriel Peri - 78114 Magny Les Hameaux FRANCE\n--|------------------------------------------------------\n--|              http://www.obry.net\n--| \"The best way to travel is by means of imagination\"\n--|\n--| gpg --keyserver wwwkeys.pgp.net --recv-key C1082595\n"},{"id":"61739","messageId":"20071203064603.GA18583@old.davidb.org","threadId":"11115","inReplyTo":"4753A43F.9060303@obry.net","subject":"Re: git-svn: .git/svn disk usage","fromName":"David Brown","fromEmail":"git@davidb.org","sentAt":"2007-12-03T06:46:03Z","receivedAt":"2007-12-03T06:46:03Z","isPatch":false,"sender":{"key":"git@davidb.org","avatar":"https://gravatar.com/avatar/94c86a2938470a74c2eac5e2b69afc0871f79a660295c02219597aba8cb101c1?d=mp&s=160"},"body":"On Mon, Dec 03, 2007 at 07:37:51AM +0100, Pascal Obry wrote:\n>Ollie,\n>\n>> I'm curious if other developers have run into this issue.  If so, are\n>> there any proposals / plans for improving the storage of git-svn\n>> metadata?\n>\n>Did you run \"git gc\" after importing code form the subversion\n>repository? On my side I found that it has reduced drastically the size\n>of the local Git repository.\n\nI think the original poster is probably finding the space in the .git/svn\ndirectory.  'git-svn' keeps an index file for every branch in SVN.\n\nI suspect it does this for speed, at least on a large import, since the SVN\ncommits will come across numerically, affecting the branches out of order.\n\nHowever, the index could fairly easily be extracted from git (since that is\nwhat it normally does).  In this case, where all of the indexes take\nsignificant space if this is worth it.\n\nOllie, if you look in these svn branch directories, is most of the space\ntaken up with files called 'index'?\n\nBrowsing through the few svn clones that I have, the space seems to be\nroughly split between 'index' files and 'unhandled.log' files.\n\nDave\n"},{"id":"61740","messageId":"94ccbe710712022253t3b1834bcs9facbf34717c7faa@mail.gmail.com","threadId":"11115","inReplyTo":"20071203064603.GA18583@old.davidb.org","subject":"Re: git-svn: .git/svn disk usage","fromName":"Kelvie Wong","fromEmail":"kelvie@ieee.org","sentAt":"2007-12-03T06:53:08Z","receivedAt":"2007-12-03T06:53:08Z","isPatch":false,"sender":{"key":"kelvie@ieee.org","avatar":null},"body":"I'm going to have to say this is due to the unhandled.log as well.\n\nJust gzip -9 it (AFAIK it's not used for anything, but keep it just in case).\n\nOn Dec 2, 2007 10:46 PM, David Brown <git@davidb.org> wrote:\n> On Mon, Dec 03, 2007 at 07:37:51AM +0100, Pascal Obry wrote:\n> >Ollie,\n> >\n> >> I'm curious if other developers have run into this issue.  If so, are\n> >> there any proposals / plans for improving the storage of git-svn\n> >> metadata?\n> >\n> >Did you run \"git gc\" after importing code form the subversion\n> >repository? On my side I found that it has reduced drastically the size\n> >of the local Git repository.\n>\n> I think the original poster is probably finding the space in the .git/svn\n> directory.  'git-svn' keeps an index file for every branch in SVN.\n>\n> I suspect it does this for speed, at least on a large import, since the SVN\n> commits will come across numerically, affecting the branches out of order.\n>\n> However, the index could fairly easily be extracted from git (since that is\n> what it normally does).  In this case, where all of the indexes take\n> significant space if this is worth it.\n>\n> Ollie, if you look in these svn branch directories, is most of the space\n> taken up with files called 'index'?\n>\n> Browsing through the few svn clones that I have, the space seems to be\n> roughly split between 'index' files and 'unhandled.log' files.\n>\n> Dave\n>\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n\n\n\n-- \nKelvie Wong\n"},{"id":"61793","messageId":"65dd6fd50712030935p242895fvc2e4576448868403@mail.gmail.com","threadId":"11115","inReplyTo":"20071203064603.GA18583@old.davidb.org","subject":"Re: git-svn: .git/svn disk usage","fromName":"Ollie Wild","fromEmail":"aaw@google.com","sentAt":"2007-12-03T17:35:22Z","receivedAt":"2007-12-03T17:35:22Z","isPatch":false,"sender":{"key":"aaw@google.com","avatar":null},"body":"On Dec 2, 2007 10:46 PM, David Brown <git@davidb.org> wrote:\n>\n> Ollie, if you look in these svn branch directories, is most of the space\n> taken up with files called 'index'?\n\nI'm seeing the following breakdown:\n\n4.3G index\n77M  unhandled.log\n5.5G .rev_db.138bc75d-0d04-0410-961f-82ee72b054a4\n\nWhat exactly are the index and .rev_db files used for?\n\nOllie\n"},{"id":"61811","messageId":"loom.20071203T182924-435@post.gmane.org","threadId":"11115","inReplyTo":"65dd6fd50712022217l5f807f31pf3f00d82c3dccf5c@mail.gmail.com","subject":"Re: git-svn: .git/svn disk usage","fromName":"David Voit","fromEmail":"david.voit@gmail.com","sentAt":"2007-12-03T18:51:44Z","receivedAt":"2007-12-03T18:51:44Z","isPatch":false,"sender":{"key":"david.voit@gmail.com","avatar":null},"body":"Ollie Wild <aaw <at> google.com> writes:\n\n> \n> Hi,\n> \n> I've been using git-svn to mirror the gcc repository at\n> svn://gcc.gnu.org/svn/gcc.  Recently, I noticed that my .git directory\n> is consuming 11GB of disk space.  Digging further, I discovered that\n> 9.8GB of this is attributable to the .git/svn directory (which\n> includes 200 branches and 2,588 tags).  Given that my .git/objects\n> directory is 652MB, it seems that it ought to be possible to store\n> this information in a more compact form.\n> \n> I'm curious if other developers have run into this issue.  If so, are\n> there any proposals / plans for improving the storage of git-svn\n> metadata?\n> \n> Thanks,\n> Ollie\n> \n\nHi all,\n\nI've seen the same effect, so i tried to reduce the size of the revdb and made a\nnew format:\nFirst, in the bin files the sha1 are stored as hexvalues not as ascii, this\nreduces the a single sha1 from 41 bytes to 20.\nSecond, only save the non-zero commits, thats what the idx are used for.\nA idx file has three 32bit integers per entry.\nThe first integer represents the first zero svn revision, the second the last\nzero revision and the last integer is the position of the next non-zero revision\nin the bin.\n\nExample:\nRevision 0-373006 are zero revision and 373007 is the first actualy used revision\nand 373008-373623 are again zero revisions\nthe idx has the following content:\n0 373006 0\n373007 373007 1\n\nand the bin only saves\n59037b8043268c9ca0d87ba86519ed0b5358c8a1\neef3f7e25993a46e3c4242aa502d93e909b08c57\n\nThe format currently used produce a 373624*41bytes large file.\n\nUsed on a git-svn clone here, i get:\nThe results are:\nold:\n1,1G    hadoop (1004M   svn/)\nnew:\n47M     hadoop (5,9M    svn/)\n\nin detail:\n\n.git/svn/trunk:\nold:\n-rw-r--r-- 1 david david 24M 2007-11-29 10:26\n.rev_db.13f79535-47bb-0310-9956-ffa450edef68\n-rw-r--r-- 1 david david 75K 2007-11-29 10:26\n.rev_db.7fecf15c-03ad-4724-994c-e980afa7160c\nnew:\n-rw-r--r-- 1 david david 32K 2007-12-03 18:40\nrevdb-13f79535-47bb-0310-9956-ffa450edef68.bin\n-rw-r--r-- 1 david david 18K 2007-12-03 18:40\nrevdb-13f79535-47bb-0310-9956-ffa450edef68.idx\n-rw-r--r-- 1 david david  32K 2007-12-03 18:44\nrevdb-7fecf15c-03ad-4724-994c-e980afa7160c.bin\n-rw-r--r-- 1 david david 2,0K 2007-12-03 18:44\nrevdb-7fecf15c-03ad-4724-994c-e980afa7160c.idx\n\nHere a example sourcecode to test this idea:\n\nI try to integrate this in git-svn this week.\n\nNOTE: I'm not a perl hacker, so use at your own risk.\n\nBye David\nps.: I'm not a member of this list please reply directly to me.\n\nmigrate.pl:\n$uuid = \"7fecf15c-03ad-4724-994c-e980afa7160c\";\n\nopen (NONZERO, '.rev_db.'.$uuid);\nopen (IDX, '>revdb-'.$uuid.'.idx');\nopen (BIN, '>revdb-'.$uuid.'.bin');\n\n$first_zero = 0;\n$pos = 0;\n$rev = 0;\n\nwhile ($sha1 = <NONZERO>) {\n\nchomp($sha1);\n\nif ($sha1 ne (\"0\" x 40))\n{\n        print BIN pack(\"H40\", $sha1);\n\n        if ($first_zero != $rev)\n        {\n                print IDX pack(\"N N N\", $first_zero, ($rev-1), $pos);\n        }\n\n        $first_zero=$rev+1;\n        $pos++;\n}\n\n$rev++;\n\n}\n\nclose(BIN);\nclose(IDX);\nclose(NONZERO);\n\nparse.pl:\nuse strict;\nuse Fcntl;\n\nmy(@index, $buf, $i);\nmy($uuid) = \"13f79535-47bb-0310-9956-ffa450edef68\";\nmy($db_path) = \"revdb-$uuid.bin\";\n\nsysopen(IDX, \"revdb-$uuid.idx\", O_RDONLY);\n\nwhile (sysread(IDX, $buf, 12)) {\n   my($minrev, $maxrev, $pos) = unpack(\"N N N\", $buf);\n\n   push @index, [$minrev, $maxrev, $pos];\n}\n\nclose(IDX);\n\nmy($lastindex) = scalar(@index)-1;\nmy($lastindexpos) = $index[$lastindex][2];\nmy($lastindexrev) = $index[$lastindex][1];\n\nmy @stat = stat $db_path;\n($stat[7] % 20) == 0 or die \"$db_path inconsistent size: $stat[7]\\n\";\nmy ($maxrev) = ($stat[7]/20)-($lastindexpos)+($lastindexrev);\nmy ($minrev) = $index[0][1]+1;\n\nmy($cachestep) = int((scalar(@index))/9);\nmy(@cache);\nfor (my($i)=0; $i < scalar(@index); $i += $cachestep) {\n   push @cache, [$index[$i][0], $i];\n}\n\nmy($lastsearch) = 0;\n\nsub pos2sha1 {\n   my($pos) = @_;\n   sysopen(BINDB, $db_path, O_RDONLY);\n\n   sysseek(BINDB, $pos, 0);\n\n   my($buf);\n   sysread(BINDB, $buf, 20);\n\n   return unpack (\"H40\", $buf);\n\n   close(BINDB);\n}\n\nsub get_sha1 {\n   my($rev) = @_;\n   my($i) = 0;\n\n   if (($rev <= 0) || ($rev > $maxrev) || $rev <= $index[0][1]) {\n      return (\"0\" x 40);\n   }\n\n   if ($rev > $lastindexrev) {\n      my($pos) = (((($rev-1) - $lastindexrev)+$lastindexpos))*20;\n\n      return pos2sha1($pos);\n   }\n\n   if(($rev >= $index[$lastsearch][0] && $rev <= $index[$lastsearch][1]) ||\n($rev >= $index[$lastsearch+1][0] && $rev <= $index[$lastsearch+1][1])) {\n      return (\"0\" x 40);\n   }\n   elsif ($rev > $index[$lastsearch][1] && $rev < $index[$lastsearch+1][0]) {\n      my($pos) = (($rev-1) - $index[$lastsearch][1] + $index[$lastsearch][2]) * 20;\n\n      return pos2sha1($pos);\n   }\n   elsif($rev > $index[$lastsearch+1][1] && $rev < $index[$lastsearch+2][0]) {\n      $lastsearch++;\n\n      my($pos) = (($rev-1) - $index[$lastsearch][1] + $index[$lastsearch][2]) * 20;\n\n      return pos2sha1($pos);\n   }\n   elsif($lastsearch != 0 && $rev > $index[$lastsearch-1][1] && $rev <\n$index[$lastsearch][0]) {\n      $lastsearch--;\n\n      my($pos) = (($rev-1) - $index[$lastsearch][1] + $index[$lastsearch][2]) * 20;\n\n      return pos2sha1($pos);\n   }\n\n   my($l, $r);\n   $l = 0;\n   $r = scalar(@index)-1;\n\n   my($lastcache) = scalar(@cache)-1;\n\n   for (my($i)=0; $i <= $lastcache; $i++) {\n      if ($rev >= $cache[$i][0]) {\n         $l = $cache[$i][1];\n      }\n\n      if ($rev < $cache[$lastcache-$i][0]) {\n         $r = $cache[$lastcache-$i][1];\n      }\n   }\n\n   if ($rev <= $index[$l][1]) {\n      return (\"0\" x 40);\n   }\n\n   while ($l <= $r) {\n      $i = int(($l + $r)/2);\n\n      if ($rev >= $index[$i][0]  && $rev <= $index[$i][1]) {\n         $lastsearch = $i;\n\n         return (\"0\" x 40);\n      }\n      elsif ($rev <= $index[$i][0]) {\n         $r = $i-1;\n      }\n      elsif ($rev >= $index[$i+1][0]) {\n         $l = $i+1;\n      } \n      else {\n         $lastsearch = $i;\n         my($pos) = (($rev-1) - $index[$i][1] + $index[$i][2]) * 20;\n\n         return pos2sha1($pos);\n      }\n   }\n\n   return (\"0\" x 40);\n}\n\nfor (my($i)=$maxrev; $i >= $minrev; $i--) {\n   my($sha1) = get_sha1($i);\n   if ($sha1 ne (\"0\" x 40)) {\n      print \"$i = $sha1\\n\";\n   }\n}\n"},{"id":"61882","messageId":"20071204082940.GA24630@diana.vm.bytemark.co.uk","threadId":"11115","inReplyTo":"65dd6fd50712030935p242895fvc2e4576448868403@mail.gmail.com","subject":"Re: git-svn: .git/svn disk usage","fromName":"Karl Hasselström","fromEmail":"kha@treskal.com","sentAt":"2007-12-04T08:29:40Z","receivedAt":"2007-12-04T08:29:40Z","isPatch":false,"sender":{"key":"kha@treskal.com","avatar":"https://gravatar.com/avatar/f0120c734b5279b345075a28521e1ac66acb20c9913ffe9bf6ae97e53f7f3f13?d=mp&s=160"},"body":"On 2007-12-03 09:35:22 -0800, Ollie Wild wrote:\n\n> I'm seeing the following breakdown:\n>\n> 4.3G index\n> 77M  unhandled.log\n> 5.5G .rev_db.138bc75d-0d04-0410-961f-82ee72b054a4\n>\n> What exactly are the index and .rev_db files used for?\n\nThe indexes are just normal git index files, one for each branch and\ntag. They're used to speed up importing new commits to the branch or\ntag.\n\nMy guess is that the performance impact of deleting them between\ngit-svn runs would be very small, since recreating an index is cheap,\nand we'd still get the speed benefit when importing several revisions\nto a branch in the same run. And it'd be a very small code change too,\nI think.\n\nIf nothing else, it's insane to keep the index for the tags. :-)\n\n-- \nKarl Hasselström, kha@treskal.com\n      www.treskal.com/kalle\n"},{"id":"61987","messageId":"20071205085451.GA347@soma","threadId":"11115","inReplyTo":"loom.20071203T182924-435@post.gmane.org","subject":"Re: git-svn: .git/svn disk usage","fromName":"Eric Wong","fromEmail":"normalperson@yhbt.net","sentAt":"2007-12-05T08:54:52Z","receivedAt":"2007-12-05T08:54:52Z","isPatch":false,"sender":{"key":"e@80x24.org","avatar":null},"body":"David Voit <david.voit@gmail.com> wrote:\n> Ollie Wild <aaw <at> google.com> writes:\n> \n> > \n> > Hi,\n> > \n> > I've been using git-svn to mirror the gcc repository at\n> > svn://gcc.gnu.org/svn/gcc.  Recently, I noticed that my .git directory\n> > is consuming 11GB of disk space.  Digging further, I discovered that\n> > 9.8GB of this is attributable to the .git/svn directory (which\n> > includes 200 branches and 2,588 tags).  Given that my .git/objects\n> > directory is 652MB, it seems that it ought to be possible to store\n> > this information in a more compact form.\n> > \n> > I'm curious if other developers have run into this issue.  If so, are\n> > there any proposals / plans for improving the storage of git-svn\n> > metadata?\n> > \n> > Thanks,\n> > Ollie\n> > \n> \n> Hi all,\n> \n> I've seen the same effect, so i tried to reduce the size of the revdb and made a\n> new format:\n> First, in the bin files the sha1 are stored as hexvalues not as ascii, this\n> reduces the a single sha1 from 41 bytes to 20.\n> Second, only save the non-zero commits, thats what the idx are used for.\n> A idx file has three 32bit integers per entry.\n> The first integer represents the first zero svn revision, the second the last\n> zero revision and the last integer is the position of the next non-zero revision\n> in the bin.\n> \n> Example:\n> Revision 0-373006 are zero revision and 373007 is the first actualy used revision\n> and 373008-373623 are again zero revisions\n> the idx has the following content:\n> 0 373006 0\n> 373007 373007 1\n> \n> and the bin only saves\n> 59037b8043268c9ca0d87ba86519ed0b5358c8a1\n> eef3f7e25993a46e3c4242aa502d93e909b08c57\n\nI'd very much like rev_db to be smaller, but I find the idea of the data\nrelying on a separate index too fragile and difficult to recover\nfrom if corruption occurs (mainly for --no-metadata users).\n\nThe rev_db is simply a lookup for mapping SVN revision numbers to\ngit commit SHA1s.\n\nI have an idea for a more compact .rev_db format:\n\n  All records are 24 bytes:\n    4 bytes for a 32-bit integer representing the SVN revision\n    20 bytes for the git commit SHA1\n\n  rev_db is an append-only format, so the 32-bit integer will be\n  monotonically increasing over time, which allows:\n\n  Lookups by revision number done via binary search:\n\n  Which means empty revisions never need to be entered anymore.\n\nOf course there needs to be a migration strategy for existing\nrepositories (mainly the ones using --no-metadata), too.\n\nUsers not using --no-metadata (nor the option for svk metadata) can just\nremove their .rev_db* files and git-svn will automatically recreate them\nas needed.\n\n> The format currently used produce a 373624*41bytes large file.\n> \n> Used on a git-svn clone here, i get:\n> The results are:\n> old:\n> 1,1G    hadoop (1004M   svn/)\n> new:\n> 47M     hadoop (5,9M    svn/)\n\nVery nice reduction!\n\n> Here a example sourcecode to test this idea:\n> \n> I try to integrate this in git-svn this week.\n> \n> NOTE: I'm not a perl hacker, so use at your own risk.\n> \n> Bye David\n> ps.: I'm not a member of this list please reply directly to me.\n\nIf you don't have time, I'll try to implement my ideas sometime this\nweek or weekend (assuming I have time, too).\n\n-- \nEric Wong\n"},{"id":"62058","messageId":"6D1288C9-8FD7-40CB-BA0B-0032F8D2DA6A@midwinter.com","threadId":"11115","inReplyTo":"20071205085451.GA347@soma","subject":"Re: git-svn: .git/svn disk usage","fromName":"Steven Grimm","fromEmail":"koreth@midwinter.com","sentAt":"2007-12-05T21:30:10Z","receivedAt":"2007-12-05T21:30:10Z","isPatch":false,"sender":{"key":"koreth@midwinter.com","avatar":"https://gravatar.com/avatar/71b4d2e8b62f168bdc9e9205341159e3567003b4f9e2127c617c5fa0a1f5bad2?d=mp&s=160"},"body":"How about using git itself to keep some of this information? I'll just  \nthrow this idea out there; might or might not make any actual sense.\n\nCreate a new \"git-svn metadata\" branch. This branch contains a fake  \ndirectory (never intended for checkout, though you could do it) that  \nhas a \"file\" for each svn revision. The filename is just the svn  \nrevision number, maybe divided into subdirectories in case you want to  \ncheck the branch out for debugging purposes or whatever. The contents  \nare the git commit SHA1 and whatever other metadata you want to keep  \nin the future.\n\nThe advantage of doing it this way? You can pass around svn metadata  \nusing the normal git fetch/push tools, query the metadata using \"git  \nshow\", etc. In terms of data integrity, it's as secure as anything  \nelse in a git repository, much more so than a separately maintained db  \nfile under .git.\n\nAlong similar lines, a separate branch where the filenames are commit  \nSHA1s and the file contents are the stuff that currently gets written  \ninto the git-svn-id: lines would mean no more need to rewrite history  \nwhen doing dcommit, and thus easier mixing of native git workflows and  \ninteractions with an svn repository.\n\nIt would be great if you could clone a git-svn repository and then do  \n\"git svn dcommit\" from the clone, secure in the knowledge that things  \nwill stay consistent even if the origin gets your changes via \"git svn  \nfetch\" rather than from you.\n\n-Steve\n"},{"id":"62121","messageId":"20071206064717.GA7744@muzzle","threadId":"11115","inReplyTo":"6D1288C9-8FD7-40CB-BA0B-0032F8D2DA6A@midwinter.com","subject":"Re: git-svn: .git/svn disk usage","fromName":"Eric Wong","fromEmail":"normalperson@yhbt.net","sentAt":"2007-12-06T06:47:17Z","receivedAt":"2007-12-06T06:47:17Z","isPatch":false,"sender":{"key":"e@80x24.org","avatar":null},"body":"Steven Grimm <koreth@midwinter.com> wrote:\n> How about using git itself to keep some of this information? I'll just  \n> throw this idea out there; might or might not make any actual sense.\n> \n> Create a new \"git-svn metadata\" branch. This branch contains a fake  \n> directory (never intended for checkout, though you could do it) that  \n> has a \"file\" for each svn revision. The filename is just the svn  \n> revision number, maybe divided into subdirectories in case you want to  \n> check the branch out for debugging purposes or whatever. The contents  \n> are the git commit SHA1 and whatever other metadata you want to keep  \n> in the future.\n> \n> The advantage of doing it this way? You can pass around svn metadata  \n> using the normal git fetch/push tools, query the metadata using \"git  \n> show\", etc. In terms of data integrity, it's as secure as anything  \n> else in a git repository, much more so than a separately maintained db  \n> file under .git.\n\nI've thought of doing the way you describe in the past, too.\n\nHowever, a missing ref to the tree you proposed would mean that the\nmetadata becomes inaccessible unless git-svn-id: lines are retained.\n\nRight now there's a single ref for all data and metadata.  Going to\ntwo refs would mean those two refs would always need to be in sync\nwith each other.\n\nThe basic idea of the git-svn-id: lines is that with the default\nsettings, the .rev_db files are deletable and can be regenerated from\nthat metadata.  git-svn will automatically re-create .rev_db files it\ncannot find.\n\nThis is why the rev_db code in git-svn uses slow, synchronous writes iff\nsvk props or no-metadata is enabled; and fast, assynchronous writes\nwhen the user sticks with the git-svn defaults.\n\n> Along similar lines, a separate branch where the filenames are commit  \n> SHA1s and the file contents are the stuff that currently gets written  \n> into the git-svn-id: lines would mean no more need to rewrite history  \n> when doing dcommit, and thus easier mixing of native git workflows and  \n> interactions with an svn repository.\n\nThe current dcommit still has the advantage that commit times match\nthose in the SVN repository.\n\n> It would be great if you could clone a git-svn repository and then do  \n> \"git svn dcommit\" from the clone, secure in the knowledge that things  \n> will stay consistent even if the origin gets your changes via \"git svn  \n> fetch\" rather than from you.\n\nIt's actually doable after the [svn-remote \"...\"] section .git/config is\ncopied and the refs/remotes/* structure is cloned via git.\n\nThe [svn-remote \"...\"] information can be regenerated based on\ngit-svn-id: lines (there's no automated way to do that, currently).\n\n-- \nEric Wong\n"}]}