{"thread":{"id":"4442","subject":"Figured out how to get Mozilla into git","startedAt":"2006-06-09T02:17:00Z","lastAt":"2006-06-19T05:02:55Z","messageCount":67,"participants":["Jon Smirl","Nicolas Pitre","Martin Langhoff","Pavel Roskin","Jakub Narebski","Linus Torvalds","Greg KH","Carl Worth","Junio C Hamano","Rogan Dawes","Timo Hirvonen","Petr Baudis","Lars Johannsen","Paul Mackerras"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"21462","messageId":"9e4733910606081917l11354e49q25f0c4aea40618ea@mail.gmail.com","threadId":"4442","inReplyTo":null,"subject":"Figured out how to get Mozilla into git","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-06-09T02:17:00Z","receivedAt":"2006-06-09T02:17:00Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"I was able to import Mozilla into SVN without problem, it just occured\nto me to then import the SVN repository in git. The import has been\nrunning a few hours now and it is up to the year 2000 (starts in\n1998). Since I haven't hit any errors yet it will probably finish ok.\nI should have the results in the morning. I wonder how long it will\ntake to start gitk on a 10GB repository.\n\nOnce I get this monster into git, are there tools that will let me\nkeep it in sync with Mozilla CVS?\nSVN renamed numeric branches to this form, unlabeled-3.7.24, so that\nmay be a problem.\n\nAny advice on how to pack this to make it run faster?\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"21463","messageId":"Pine.LNX.4.64.0606082253430.19403@localhost.localdomain","threadId":"4442","inReplyTo":"9e4733910606081917l11354e49q25f0c4aea40618ea@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2006-06-09T02:56:53Z","receivedAt":"2006-06-09T02:56:53Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 8 Jun 2006, Jon Smirl wrote:\n\n> I was able to import Mozilla into SVN without problem, it just occured\n> to me to then import the SVN repository in git. The import has been\n> running a few hours now and it is up to the year 2000 (starts in\n> 1998). Since I haven't hit any errors yet it will probably finish ok.\n> I should have the results in the morning. I wonder how long it will\n> take to start gitk on a 10GB repository.\n\nBefore you do so consider repacking the repository with \n\n  git-repack -a -f -d && git-prune-packed\n\n\nNicolas\n"},{"id":"21464","messageId":"46a038f90606082006t5c6a5623q4b9cf7b036dad1e5@mail.gmail.com","threadId":"4442","inReplyTo":"9e4733910606081917l11354e49q25f0c4aea40618ea@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-06-09T03:06:11Z","receivedAt":"2006-06-09T03:06:11Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"Jon,\n\noh, I went back to a cvsimport that I started a couple days ago.\nCompleted with no problems...\n\nLast commit:\ncommit 5ecb56b9c4566618fad602a8da656477e4c6447a\nAuthor: wtchang%redhat.com <wtchang%redhat.com>\nDate:   Fri Jun 2 17:20:37 2006 +0000\n\n    Import NSPR 4.6.2 and NSS 3.11.1\n\nmozilla.git$ du -sh .git/\n2.0G    .git/\n\nIt took\n43492.19user 53504.77system 40:23:49elapsed 66%CPU (0avgtext+0avgdata\n0maxresident)k\n0inputs+0outputs (77334major+3122469478minor)pagefaults 0swaps\n\n> I should have the results in the morning. I wonder how long it will\n> take to start gitk on a 10GB repository.\n\nHopefully not that big :) -- anyway, just do gitk --max-count=1000\n\n> Once I get this monster into git, are there tools that will let me\n> keep it in sync with Mozilla CVS?\n\nIf you use git-cvsimport, you can safely re-run it on a cronjob to\nkeep it in sync. Not too sure about the cvs2svn => git-svnimport,\nthough git-svnimport does support incremental imports.\n\n> SVN renamed numeric branches to this form, unlabeled-3.7.24, so that\n> may be a problem.\n\nOuch,\n\n> Any advice on how to pack this to make it run faster?\n\ngit-repack -a -d but it OOMs on my 2GB+2GBswap machine :(\n\n\nmartin\n"},{"id":"21465","messageId":"20060608231200.4bkoc8sggk88k0ow@webmail.spamcop.net","threadId":"4442","inReplyTo":"9e4733910606081917l11354e49q25f0c4aea40618ea@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Pavel Roskin","fromEmail":"proski@gnu.org","sentAt":"2006-06-09T03:12:00Z","receivedAt":"2006-06-09T03:12:00Z","isPatch":false,"sender":{"key":"proski@gnu.org","avatar":null},"body":"Hi Jon,\n\nQuoting Jon Smirl <jonsmirl@gmail.com>:\n\n> I was able to import Mozilla into SVN without problem, it just occured\n> to me to then import the SVN repository in git.\n\nI feel bad that I didn't suggest it before.  That's quite expected.  Subversion\nwas created by  CVS developers with the intention of replacing CVS.  cvs2svn\nwas written by the same CVS developers, who paid attention to all CVS quirks. \ncvs2svn is quite mature and it has a testsuite, if I remember correctly.\n\nMy concern is how well a Subversion repository can be mapped to git considering\nthat Subversion is branch agnostic.  But if it works for Mozilla, this approach\ncould be recommended for anything big and serious.\n\n> The import has been\n> running a few hours now and it is up to the year 2000 (starts in\n> 1998). Since I haven't hit any errors yet it will probably finish ok.\n> I should have the results in the morning. I wonder how long it will\n> take to start gitk on a 10GB repository.\n\nThat's the \"raison d'etre\" of qgit.  I don't know if gitk has anything that qgit\ndoesn't, except bisecting.\n\n> Once I get this monster into git, are there tools that will let me\n> keep it in sync with Mozilla CVS?\n\nIdeally, make Mozilla developers use git :-)\n\n> SVN renamed numeric branches to this form, unlabeled-3.7.24, so that\n> may be a problem.\n\nI think git-svn is supposed to do the svn->git part, but I'm afraid it will need\nsome work to do it effectively.  Google search for \"cvs2svn incremental\" brings\nsome patches.  cvsup can be used to synchronize the CVS repository.\n\n--\nRegards,\nPavel Roskin\n"},{"id":"21466","messageId":"9e4733910606082028k37f6d915m26009e0d5011808b@mail.gmail.com","threadId":"4442","inReplyTo":"46a038f90606082006t5c6a5623q4b9cf7b036dad1e5@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-06-09T03:28:30Z","receivedAt":"2006-06-09T03:28:30Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 6/8/06, Martin Langhoff <martin.langhoff@gmail.com> wrote:\n> Jon,\n>\n> oh, I went back to a cvsimport that I started a couple days ago.\n> Completed with no problems...\n\nI am using cvsps-2.1-3.fc5, the last time I tried it died in the\nmiddle of the import. I don't remember why it died. Which cvsps are\nyou using? You're saying that it can handle the whole Mozilla CVS now,\nright? I will build a new cvsps from CVS and start it running tonight.\n\n> If you use git-cvsimport, you can safely re-run it on a cronjob to\n> keep it in sync. Not too sure about the cvs2svn => git-svnimport,\n> though git-svnimport does support incremental imports.\n\nI would much rather get a direct CVS import working so that I can do\nincremental updates. I went the SVN route because it was the only\nthing I could get working.\n\n> > Any advice on how to pack this to make it run faster?\n>\n> git-repack -a -d but it OOMs on my 2GB+2GBswap machine :(\n\nWe are all having problems getting this to run on 32 bit machines with\nthe 3-4GB process size limitations.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"21469","messageId":"e6b798$td3$1@sea.gmane.org","threadId":"4442","inReplyTo":"9e4733910606082028k37f6d915m26009e0d5011808b@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-06-09T07:17:02Z","receivedAt":"2006-06-09T07:17:02Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Jon Smirl wrote:\n\n\n>> git-repack -a -d but it OOMs on my 2GB+2GBswap machine :(\n> \n> We are all having problems getting this to run on 32 bit machines with\n> the 3-4GB process size limitations.\n\nIs that expected (for 10GB repository if I remember correctly), or is there\nsome way to avoid this OOM?\n\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"21471","messageId":"Pine.LNX.4.64.0606090745390.5498@g5.osdl.org","threadId":"4442","inReplyTo":"e6b798$td3$1@sea.gmane.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-09T15:01:56Z","receivedAt":"2006-06-09T15:01:56Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 9 Jun 2006, Jakub Narebski wrote:\n> Jon Smirl wrote:\n> \n> >> git-repack -a -d but it OOMs on my 2GB+2GBswap machine :(\n> > \n> > We are all having problems getting this to run on 32 bit machines with\n> > the 3-4GB process size limitations.\n> \n> Is that expected (for 10GB repository if I remember correctly), or is there\n> some way to avoid this OOM?\n\nWell, to some degree, the VM limitations are inevitable with huge packs.\n\nThe original idea for packs was to avoid making one huge pack, partly \nbecause it was expected to be really really slow to generate (so \nincremental repacking was a much better strategy), but partly simply \nbecause trying to map one huge pack is really hard to do.\n\nFor various reasons, we ended up mostly using a single pack most of the \ntime: it's the most efficient model when the project is reasonably sized, \nand it turns out that with the delta re-use, repacking even moderately \nlarge projects like the kernel doesn't actually take all that long.\n\nBut the fact that we ended up mostly using a single pack for the kernel, \nfor example, doesn't mean that the fundamental reasons that git supports \nmultiple packs would somehow have gone away. At some point, the project \ngets large enough that one single pack simply isn't reasonable.\n\nSo a single 2GB pack is already very much pushing it. It's really really \nhard to map in a 2GB file on a 32-bit platform: your VM is usually \nfragmented enough that it simply isn't practical. In fact, I think the \nlimit for _practical_ usage of single packs is probably somewhere in the \nhalf-gig region, unless you just have 64-bit machines.\n\nAnd yes, I realize that the \"single pack\" thing actually ends up having \nbecome a fact for cloning, for example. Originally, cloning would unpack \non the receiving end, and leave the repacking to happen there, but that \nobviously sucked. So now when we clone, we always get a single pack. That \ncan absolutely be a problem.\n\nI don't know what the right solution is. Single packs _are_ very useful, \nespecially after a clone. So it's possible that we should just make the \npack-reading code be able to map partial packs. But the point is that \nthere are certainly ways we can fix this - it's not _really_ fundamental.\n\nIt's going to complicate it a bit (damn, how I hate 32-bit VM \nlimitations), but the good news is that the whole git model of \"everything \nis an individual object\" means that it's a very _local_ decision: it will \nprobably be painful to re-do some of the pack reading code and have a LRU \nof pack _fragments_ instead of a LRU of packs, but it's only going to \naffect a small part of git, and everything else will never even see it.\n\nSo large packs are not really a fundamental problem, but right now we have \nsome practical issues with them.\n\n(It's not _just_ packs: running out of memory is also because of \ngit-rev-list --objects being pretty memory hungry. I've improved the \nmemory usage several times by over 50%, but people keep trying larger \nprojects. It used to be that I considered the kernel a large history, now \nwe're talking about things that have ten times the number of objects).\n\nMartin - do you have some place to make that big mozilla repo available? \nIt would be a good test-case.. \n\n\t\t\tLinus\n"},{"id":"21472","messageId":"Pine.LNX.4.64.0606091127540.19403@localhost.localdomain","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606090745390.5498@g5.osdl.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2006-06-09T16:11:42Z","receivedAt":"2006-06-09T16:11:42Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 9 Jun 2006, Linus Torvalds wrote:\n\n> \n> \n> On Fri, 9 Jun 2006, Jakub Narebski wrote:\n> > Jon Smirl wrote:\n> > \n> > >> git-repack -a -d but it OOMs on my 2GB+2GBswap machine :(\n> > > \n> > > We are all having problems getting this to run on 32 bit machines with\n> > > the 3-4GB process size limitations.\n> > \n> > Is that expected (for 10GB repository if I remember correctly), or is there\n> > some way to avoid this OOM?\n\nWhat was that 10GB related to, exactly?  The original CVS repo, or the \nunpacked GIT repo?\n\n> So a single 2GB pack is already very much pushing it. It's really really \n> hard to map in a 2GB file on a 32-bit platform: your VM is usually \n> fragmented enough that it simply isn't practical. In fact, I think the \n> limit for _practical_ usage of single packs is probably somewhere in the \n> half-gig region, unless you just have 64-bit machines.\n\nSure, but have we already reached that size?\n\nThe historic Linux repo currently repacks itself into a ~175MB pack for \n63428 commits.\n\nThe current Linux repo is ~103MB with a much shorter history (27153 \ncommits).\n\nGiven the above we can estimate the size of the kernel repository after \nx commits as follows:\n\n\tslope = (175 - 103) / (63428 - 27153) = approx 2KB per commit\n\n\tinitial size = 175 - .001985 * 63428 = 49MB\n\nSo the initial kernel commit is about 49MB in size which is coherent \nwith the corresponding compressed tarball.  Subsequent commits are 2KB \nin size on average.  Given that it will take about 233250 commits before \nthe kernel reaches the half gigabyte pack file, and given the current \ncommit rate (approx 23700 commits per year), that means we still have \nnearly 9 years to go.  And at that point 64-bit machines are likely to \nbe the norm.\n\nSo given those numbers I don't think this is really an issue.  The Linux \nkernel is a rather huge and pretty active project to base comparisons \nagainst.  The Mozilla repository might be difficult to import and \nrepack, but once repacked it should still be pretty usable now even on a \n32-bit machine even with a single pack.\n\nOtherwise that should be quite easy to add a batch size argument to \ngit-repack so git-rev-list and git-pack-objects are called multiple \ntimes with sequential commit \nranges to create a repo with multiple packs.\n\n\nNicolas\n"},{"id":"21473","messageId":"Pine.LNX.4.64.0606090926550.5498@g5.osdl.org","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606091127540.19403@localhost.localdomain","subject":"Re: Figured out how to get Mozilla into git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-09T16:30:35Z","receivedAt":"2006-06-09T16:30:35Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 9 Jun 2006, Nicolas Pitre wrote:\n> \n> > So a single 2GB pack is already very much pushing it. It's really really \n> > hard to map in a 2GB file on a 32-bit platform: your VM is usually \n> > fragmented enough that it simply isn't practical. In fact, I think the \n> > limit for _practical_ usage of single packs is probably somewhere in the \n> > half-gig region, unless you just have 64-bit machines.\n> \n> Sure, but have we already reached that size?\n\nNot for the Linux repos.\n\nBut apparently the mozilla repo ends up being 2GB in git. From Martin:\n\n  >> oh, I went back to a cvsimport that I started a couple days ago.\n  >> Completed with no problems...\n  >> \n  >> Last commit:\n  >> commit 5ecb56b9c4566618fad602a8da656477e4c6447a\n  >> Author: wtchang%redhat.com <wtchang%redhat.com>\n  >> Date:   Fri Jun 2 17:20:37 2006 +0000\n  >> \n  >>    Import NSPR 4.6.2 and NSS 3.11.1\n  >> \n  >> mozilla.git$ du -sh .git/\n  >> 2.0G    .git/\n\nnow that was done with _incremental_ repacking (ie his .git directory\nwon't be just one large pack), but I bet that if you were to clone it\n(without using the \"-l\" flag or rsync/http), you'd end up with serious\ntrouble because of the single-pack limit.\n\nSo we're starting to see archives where single packs are problematic for\na 32-bit architecture. \n\n\t\t\tLinus\n"},{"id":"21474","messageId":"e6ca1d$2u6$1@sea.gmane.org","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606091127540.19403@localhost.localdomain","subject":"Re: Figured out how to get Mozilla into git","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-06-09T17:10:14Z","receivedAt":"2006-06-09T17:10:14Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Nicolas Pitre wrote:\n\n> What was that 10GB related to, exactly?  The original CVS repo, or the \n> unpacked GIT repo?\n\nErm, Subversion repository, result of cvs2svn conversion:\n\nJon Smirl> I wonder how long it will take to start gitk on a 10GB \nJon Smirl> repository.\n\n(in first post in this thread).\n\n> Otherwise that should be quite easy to add a batch size argument to \n> git-repack so git-rev-list and git-pack-objects are called multiple \n> times with sequential commit ranges to create a repo with multiple\n> packs. \n\nGood idea. In addition to best size pack limted by 32bit and/or RAM size +\nswap size limit, there are (rare) limits of maximum filesize on filesystem,\ne.g. FAT28^W FAT32.\n\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"21475","messageId":"Pine.LNX.4.64.0606091326550.2703@localhost.localdomain","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606090926550.5498@g5.osdl.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2006-06-09T17:38:13Z","receivedAt":"2006-06-09T17:38:13Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 9 Jun 2006, Linus Torvalds wrote:\n\n> \n> \n> On Fri, 9 Jun 2006, Nicolas Pitre wrote:\n> > \n> > > So a single 2GB pack is already very much pushing it. It's really really \n> > > hard to map in a 2GB file on a 32-bit platform: your VM is usually \n> > > fragmented enough that it simply isn't practical. In fact, I think the \n> > > limit for _practical_ usage of single packs is probably somewhere in the \n> > > half-gig region, unless you just have 64-bit machines.\n> > \n> > Sure, but have we already reached that size?\n> \n> Not for the Linux repos.\n> \n> But apparently the mozilla repo ends up being 2GB in git. From Martin:\n> \n>   >> oh, I went back to a cvsimport that I started a couple days ago.\n>   >> Completed with no problems...\n>   >> \n>   >> Last commit:\n>   >> commit 5ecb56b9c4566618fad602a8da656477e4c6447a\n>   >> Author: wtchang%redhat.com <wtchang%redhat.com>\n>   >> Date:   Fri Jun 2 17:20:37 2006 +0000\n>   >> \n>   >>    Import NSPR 4.6.2 and NSS 3.11.1\n>   >> \n>   >> mozilla.git$ du -sh .git/\n>   >> 2.0G    .git/\n\nHe also sais:\n\n| git-repack -a -d but it OOMs on my 2GB+2GBswap machine :(\n\n> now that was done with _incremental_ repacking (ie his .git directory\n> won't be just one large pack),\n\nSo given the nature of packs, incrementally packing an imported \nrepository _might_ cause worse problems since each pack must be self \nreferenced by definition.  That means you may end up with multiple \nrevisions of the same file distributed amongst as many packs hence none \nof those revisions are ever deltified, and to repack that you currently \nhave to mmap all those packs at once.\n\n> but I bet that if you were to clone it\n> (without using the \"-l\" flag or rsync/http), you'd end up with serious\n> trouble because of the single-pack limit.\n\nMaybe that single pack would instead be under the 512MB limit?  I'd be \ncurious to know.\n\n> So we're starting to see archives where single packs are problematic for\n> a 32-bit architecture. \n\nDepending on the operation, the single pack might actually be better, \nespecially for a full clone where everything gets mapped.  Multiple \npacks will always take more space, which is fine if you don't need \naccess to all objects at once since individual packs are small, but the \nwhole of them (when repacking or cloning) isn't.\n\n\nNicolas\n"},{"id":"21476","messageId":"Pine.LNX.4.64.0606091047080.5498@g5.osdl.org","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606091326550.2703@localhost.localdomain","subject":"Re: Figured out how to get Mozilla into git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-09T17:49:28Z","receivedAt":"2006-06-09T17:49:28Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 9 Jun 2006, Nicolas Pitre wrote:\n>\n> Maybe that single pack would instead be under the 512MB limit?  I'd be \n> curious to know.\n\nPossible, but not likely, and with \"git repack -a -d\" running out of \nmemory, we clearly already have a problem in checking that.\n\nThat is most likely git-rev-list, though. Which is why I'd like to just \nrsync the repo, and run git-rev-list on it, and see what else I can shave \noff ;)\n\n> > So we're starting to see archives where single packs are problematic for\n> > a 32-bit architecture. \n> \n> Depending on the operation, the single pack might actually be better, \n\nAbsolutely. Which is why I said we probably need to do a LRU on pack \nfragments rather than full packs when we do the pack memory mapping.\n\n\t\tLinus\n"},{"id":"21477","messageId":"9e4733910606091113vdc6ab06l2d3582cb82b8fd09@mail.gmail.com","threadId":"4442","inReplyTo":"46a038f90606082006t5c6a5623q4b9cf7b036dad1e5@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-06-09T18:13:36Z","receivedAt":"2006-06-09T18:13:36Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 6/8/06, Martin Langhoff <martin.langhoff@gmail.com> wrote:\n> mozilla.git$ du -sh .git/\n> 2.0G    .git/\n\nThat looks too small. My svn git import is 2.7GB and the source CVS is\n3.0GB. The svn import wasn't finished when I stopped it.\n\nMy cvsps process is still running from last night. The error file is\n341MB. How big is it when the conversion is finished? My machine is\nswapping to death.\n\nI'm still attracted to the cvs2svn tool. It handled everything right\nthe first time and it only needs 100MB to run. It is also a lot\nfaster. cvsps and parsecvs both need gigabytes of RAM to run. I'll\nlook at cvs2svn some more but I still need to figure out more about\nlow level git and learn Python.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"21481","messageId":"Pine.LNX.4.64.0606091158460.5498@g5.osdl.org","threadId":"4442","inReplyTo":"9e4733910606091113vdc6ab06l2d3582cb82b8fd09@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-09T19:00:42Z","receivedAt":"2006-06-09T19:00:42Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 9 Jun 2006, Jon Smirl wrote:\n> \n> That looks too small. My svn git import is 2.7GB and the source CVS is\n> 3.0GB. The svn import wasn't finished when I stopped it.\n\nGit is much better at packing than either CVS or SVN. Get used to it ;)\n\n> My cvsps process is still running from last night. The error file is\n> 341MB. How big is it when the conversion is finished? My machine is\n> swapping to death.\n\nDo you have all the cvsps patches? There's a few important ones floating \naround, and David Mansfield never did a 2.2 release..\n\nI'm pretty sure Martin doesn't run plain 2.1.\n\n\t\tLinus\n"},{"id":"21488","messageId":"9e4733910606091317p26d66579mdf93db293f93fb50@mail.gmail.com","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606091158460.5498@g5.osdl.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-06-09T20:17:26Z","receivedAt":"2006-06-09T20:17:26Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 6/9/06, Linus Torvalds <torvalds@osdl.org> wrote:\n>\n>\n> On Fri, 9 Jun 2006, Jon Smirl wrote:\n> >\n> > That looks too small. My svn git import is 2.7GB and the source CVS is\n> > 3.0GB. The svn import wasn't finished when I stopped it.\n>\n> Git is much better at packing than either CVS or SVN. Get used to it ;)\n\nThe git tree that Martin got from cvsps is much smaller that the git\ntree I got from going to svn then to git.  I don't why the trees are\n700KB different, it may be different amounts of packing, or one of the\nconversion tools is losing something.\n\nEarlier he said:\n>git-repack -a -d but it OOMs on my 2GB+2GBswap machine :(\n\n> > My cvsps process is still running from last night. The error file is\n> > 341MB. How big is it when the conversion is finished? My machine is\n> > swapping to death.\n>\n> Do you have all the cvsps patches? There's a few important ones floating\n> around, and David Mansfield never did a 2.2 release..\n\nI am running cvsps-2.1-3.fc5 so I may be wasting my time. Error out is\n535MB now.\nHe sent me some git patches, but none for cvsps.\n\n> I'm pretty sure Martin doesn't run plain 2.1.\n\nI haven't come up with anything that is likely to result in Mozilla\nswitching over to git. Right now it takes three days to convert the\ntree. The tree will have to be run in parallel for a while to convince\neveryone to switch. I don't have a solution to keeping it in sync in\nnear real time (commits would still go to CVS). Most Mozilla\ndevelopers are interested but the infrastructure needs some help.\n\nMartin has also brought up the problem with needing a partial clone so\nthat everyone doesn't have to bring down the entire repository. A\ntrunk checkout is 340MB and Martin's git tree is 2GB (mine 2.7GB).  A\nkernel tree is only 680M.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"21493","messageId":"Pine.LNX.4.64.0606091331170.5498@g5.osdl.org","threadId":"4442","inReplyTo":"9e4733910606091317p26d66579mdf93db293f93fb50@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-09T20:40:49Z","receivedAt":"2006-06-09T20:40:49Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 9 Jun 2006, Jon Smirl wrote:\n>\n> > Git is much better at packing than either CVS or SVN. Get used to it ;)\n> \n> The git tree that Martin got from cvsps is much smaller that the git\n> tree I got from going to svn then to git.  I don't why the trees are\n> 700KB different, it may be different amounts of packing, or one of the\n> conversion tools is losing something.\n\n.. or one of them is adding something.\n\nFor example, it may well be that cvs2svn does a lot more commits or \nsomething like that.\n\nThat said, I don't even see where git-svn packs anythign at all, and \nyou're absolutely right that when/how you repack can make a huge \ndifference to disk usage, much more so than any importer details.\n\n> > Do you have all the cvsps patches? There's a few important ones floating\n> > around, and David Mansfield never did a 2.2 release..\n> \n> I am running cvsps-2.1-3.fc5 so I may be wasting my time. Error out is\n> 535MB now.\n> He sent me some git patches, but none for cvsps.\n\nI've got a couple, but I was hoping David would do a cvsps-2.2. I have \nthis dim memory of him saying he had done some other improvements too.\n\n> I haven't come up with anything that is likely to result in Mozilla\n> switching over to git. Right now it takes three days to convert the\n> tree. The tree will have to be run in parallel for a while to convince\n> everyone to switch. I don't have a solution to keeping it in sync in\n> near real time (commits would still go to CVS). Most Mozilla\n> developers are interested but the infrastructure needs some help.\n\nSure. That said, I pretty much guarantee that the size issues will be much \nmuch worse for any other distributed SCM. \n\nIf Mozilla doesn't need the distributed thing, then SVN is probably the \nbest choice. It's still a total piece of crap, but hey, if crap (== \ncentralized) is what people are used to, a few billion flies can't be \nwrong ;)\n\nIf you got your import done, is there some place I can rsync it from, and \nat least I can make sure that everything works fine for a repo that size.. \nOne day the Mozilla people will notice that they really _really_ want the \ndistribution, and they'll figure out quickly enough that SVK doesn't cut \nit, I suspect.\n\n\t\tLinus\n"},{"id":"21495","messageId":"e6cmjn$no5$1@sea.gmane.org","threadId":"4442","inReplyTo":"9e4733910606091317p26d66579mdf93db293f93fb50@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-06-09T20:44:48Z","receivedAt":"2006-06-09T20:44:48Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Jon Smirl wrote:\n\n> Martin has also brought up the problem with needing a partial clone so\n> that everyone doesn't have to bring down the entire repository. A\n> trunk checkout is 340MB and Martin's git tree is 2GB (mine 2.7GB).  A\n> kernel tree is only 680M.\n\nPartial/shallow nor lazy clone we don't have (although there might be some\nshallow clone partial solutions in topic branches and/or patches flying\naround in git mailing list). Yet.\n\nBut you can do what was done for Linux kernel: split repository into current\nand historical, and you can join them (join the history) if needed using\ngrafts. And even if one need historical repository, it is neede to\nclone/copy only _once_. With alternatives (using historical repository as\none of alternatives for current repository) someone who has both\nrepositories does need only a little more space, I think, than if one used\nsingle repository.\n\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"21498","messageId":"9e4733910606091356w391b4fdao23db5b2ce3c3e282@mail.gmail.com","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606091331170.5498@g5.osdl.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-06-09T20:56:17Z","receivedAt":"2006-06-09T20:56:17Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 6/9/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> > I haven't come up with anything that is likely to result in Mozilla\n> > switching over to git. Right now it takes three days to convert the\n> > tree. The tree will have to be run in parallel for a while to convince\n> > everyone to switch. I don't have a solution to keeping it in sync in\n> > near real time (commits would still go to CVS). Most Mozilla\n> > developers are interested but the infrastructure needs some help.\n>\n> Sure. That said, I pretty much guarantee that the size issues will be much\n> much worse for any other distributed SCM.\n>\n> If Mozilla doesn't need the distributed thing, then SVN is probably the\n> best choice. It's still a total piece of crap, but hey, if crap (==\n> centralized) is what people are used to, a few billion flies can't be\n> wrong ;)\n\nThey need the distributed thing whether they realize it or not. Some\nof the external projects like songbird and nvu are vulnerable to drift\nsince they are running their own repositories.  Once a  few\nmove/renames happen they can't easily stay in sync anymore. It has\nbeen over a year since NVU was merged back into the trunk.\n\nThat is the same reason I want it, so that I can work on stuff locally\nand have a repository. The core staff doesn't have this problem\nbecause they can make all the branches they want in the main\nrepository.\n\n> If you got your import done, is there some place I can rsync it from, and\n> at least I can make sure that everything works fine for a repo that size..\n> One day the Mozilla people will notice that they really _really_ want the\n> distribution, and they'll figure out quickly enough that SVK doesn't cut\n> it, I suspect.\n\nIt would be better to rsync Martins copy, he has a lot more bandwidth.\nIt will take over a day to copy it off my cable modem. I'm signed up\nto get FIOS as soon as they turn it on in my neighborhood, it's\nalready wired on the poles.\n\n\n>\n>                 Linus\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"21501","messageId":"Pine.LNX.4.64.0606091655420.2703@localhost.localdomain","threadId":"4442","inReplyTo":"9e4733910606091317p26d66579mdf93db293f93fb50@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2006-06-09T21:05:56Z","receivedAt":"2006-06-09T21:05:56Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 9 Jun 2006, Jon Smirl wrote:\n\n> I haven't come up with anything that is likely to result in Mozilla\n> switching over to git. Right now it takes three days to convert the\n> tree. The tree will have to be run in parallel for a while to convince\n> everyone to switch. I don't have a solution to keeping it in sync in\n> near real time (commits would still go to CVS). Most Mozilla\n> developers are interested but the infrastructure needs some help.\n\nThis is true.  GIT is still evolving and certainly needs work to cope \nwith environments and datasets that were never tested before.  The \nMozilla repo is one of those and we're certainly interested into making \nit work well.  GIT might not be right for it just yet, but if you could \nlet us rsync your converted repo to play with that might help us work on \nproper fixes for that kind of repo.\n\n> Martin has also brought up the problem with needing a partial clone so\n> that everyone doesn't have to bring down the entire repository.\n\nIf it can be repacked into a single pack that size might get much \nsmaller too.\n\n\nNicolas\n"},{"id":"21504","messageId":"9e4733910606091446u5660d7b0q1f2e118abcd057c9@mail.gmail.com","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606091655420.2703@localhost.localdomain","subject":"Re: Figured out how to get Mozilla into git","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-06-09T21:46:47Z","receivedAt":"2006-06-09T21:46:47Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 6/9/06, Nicolas Pitre <nico@cam.org> wrote:\n> On Fri, 9 Jun 2006, Jon Smirl wrote:\n>\n> > I haven't come up with anything that is likely to result in Mozilla\n> > switching over to git. Right now it takes three days to convert the\n> > tree. The tree will have to be run in parallel for a while to convince\n> > everyone to switch. I don't have a solution to keeping it in sync in\n> > near real time (commits would still go to CVS). Most Mozilla\n> > developers are interested but the infrastructure needs some help.\n>\n> This is true.  GIT is still evolving and certainly needs work to cope\n> with environments and datasets that were never tested before.  The\n> Mozilla repo is one of those and we're certainly interested into making\n> it work well.  GIT might not be right for it just yet, but if you could\n> let us rsync your converted repo to play with that might help us work on\n> proper fixes for that kind of repo.\n\nI'm rebuilding it on my shared hosting account at dreamhost.com. I'll\nsee if I can get it built before they notice and kill my process. My\naccount there is on a 4GB quad xeon box so hopefully it can convert\nthe tree faster. My account has 1TB download per month so rsync will\nbe ok. Not bad for $12 the first year.\n\nIt would take over a day to rsync it off from my home machine.\n\n> > Martin has also brought up the problem with needing a partial clone so\n> > that everyone doesn't have to bring down the entire repository.\n>\n> If it can be repacked into a single pack that size might get much\n> smaller too.\n>\n>\n> Nicolas\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"21505","messageId":"Pine.LNX.4.64.0606091450180.5498@g5.osdl.org","threadId":"4442","inReplyTo":"9e4733910606091356w391b4fdao23db5b2ce3c3e282@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-09T21:57:58Z","receivedAt":"2006-06-09T21:57:58Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 9 Jun 2006, Jon Smirl wrote:\n> \n> They need the distributed thing whether they realize it or not. Some\n> of the external projects like songbird and nvu are vulnerable to drift\n> since they are running their own repositories.  Once a  few\n> move/renames happen they can't easily stay in sync anymore. It has\n> been over a year since NVU was merged back into the trunk.\n> \n> That is the same reason I want it, so that I can work on stuff locally\n> and have a repository. The core staff doesn't have this problem\n> because they can make all the branches they want in the main\n> repository.\n\nYes. Anyway, I think we'll get git working well for repositories that \nsize, and eventually the core developers will notice how much better it \nis.\n\nIn the meantime, the fact that git-cvsimport can be done incrementally \nmeans that once we have the silly pack-file-mapping details worked out, it \nshould be perfectly fine to run the 3-day import just once, and then work \non it incrementally afterwards without any real problems.\n\nSo people like you who want to work on it off-line using a distributed \nsystem _can_ do so, realistically. Maybe not practically _today_, but I \ndon't think the git issues are serious enough that we'd be talking about \n\"months from now\", but more of a \"in a week or so we migh have something \nthat works fine for your case\".\n\n[ They had this long discussion about languages on #monotone the other \n  day, and the reason I'll take C over anything else any day is the fact \n  that a well-written C program is literally only limited by hardware, \n  never by the language. The poor python/perl guys may write things more \n  quickly, but when they hit a language wall, they hit it. \n\n  I think we've got an excellent data model, and handling even something \n  huge like the _whole_ history of mozilla doesn't look very daunting at \n  all. I just want to have a real test-case to motivate me to look at the \n  problems. ]\n\n> It would be better to rsync Martins copy, he has a lot more bandwidth.\n> It will take over a day to copy it off my cable modem. I'm signed up\n> to get FIOS as soon as they turn it on in my neighborhood, it's\n> already wired on the poles.\n\nSure. I actually just have regular 128kbps DSL myself. I guess I should \nupgrade to 256 (the downside of having deer munching on the roses in our \nback yard is that I don't think I even have the option for anything \nfaster), but I'm so damn well distributed that the slow 128kbps is \nactually more than enough - everything serious I do is local anyway.\n\nSo it will take me quite some time to download 2GB+, regardless of how fat \na pipe the other end has ;)\n\n\t\tLinus\n"},{"id":"21506","messageId":"Pine.LNX.4.64.0606091511190.5498@g5.osdl.org","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606091450180.5498@g5.osdl.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-09T22:17:56Z","receivedAt":"2006-06-09T22:17:56Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 9 Jun 2006, Linus Torvalds wrote:\n> \n> Sure. I actually just have regular 128kbps DSL myself.\n\nNot bits, bytes. 128_KB_/s, of course. Actually, it's slightly more. \nSomething like 146KB/s, I guess that comes to 1.5Mbps.\n\nJust in case somebody thought I was living in a cave in the middle ages.\n\nAnyway, no nice 5Mbps cable for me.\n\n\t\tLinus\n"},{"id":"21508","messageId":"20060609231614.GN17807@kroah.com","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606091450180.5498@g5.osdl.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Greg KH","fromEmail":"greg@kroah.com","sentAt":"2006-06-09T23:16:14Z","receivedAt":"2006-06-09T23:16:14Z","isPatch":false,"sender":{"key":"greg@kroah.com","avatar":"https://gravatar.com/avatar/5bb5aa0cc2e01c00ec899d11130c07796bc186e465bae57bc34873b13b72c7c8?d=mp&s=160"},"body":"On Fri, Jun 09, 2006 at 02:57:58PM -0700, Linus Torvalds wrote:\n> > It would be better to rsync Martins copy, he has a lot more bandwidth.\n> > It will take over a day to copy it off my cable modem. I'm signed up\n> > to get FIOS as soon as they turn it on in my neighborhood, it's\n> > already wired on the poles.\n> \n> So it will take me quite some time to download 2GB+, regardless of how fat \n> a pipe the other end has ;)\n\nFed-Ex a DVD or two would probably be fastest :)\n"},{"id":"21509","messageId":"46a038f90606091637o6a0194d5yb413237253a372fc@mail.gmail.com","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606091450180.5498@g5.osdl.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-06-09T23:37:35Z","receivedAt":"2006-06-09T23:37:35Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"Apologies, I dropped out of the conversation -- Friday night drinks\n(NZ timezone) took over ;-)\n\nNow, back on track...\n\nOn 6/10/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> In the meantime, the fact that git-cvsimport can be done incrementally\n> means that once we have the silly pack-file-mapping details worked out, it\n> should be perfectly fine to run the 3-day import just once, and then work\n> on it incrementally afterwards without any real problems.\n\nExactly. The dog at this time is cvsps -- I also remember vague\npromises from a list regular of publishing a git repo with cvsps2.1 +\nsome patches from the list.\n\nIn any case, and for the record, my cvsps is 2.1 pristine. It handles\nthe mozilla repo alright, as long as I give it a lot of RAM. I _think_\nit slurped 3GB with the mozilla cvs.\n\nI want to review that cvs2svn importer, probably to steal the test\ncases and perhaps some logic to revamp/replace cvsps. The thing is --\nwe can't just drop/replace cvsimport because it does incrementals, so\ncontinuity and consistency are key. All the CVS imports have to take\nsome hard decisions when the data is bad -- however it is we fudge it,\nwe kind of want to fudge it consistently ;-)\n\n> So people like you who want to work on it off-line using a distributed\n> system _can_ do so, realistically. Maybe not practically _today_\n\nOther than \"don't run repack -a\", it's feasible. In fact, that's how I\nuse git 99% of the time -- to do DSCM stuff on projects that are using\nCVS, like Moodle.\n\n>   The poor python/perl guys may write things more\n>   quickly, but when they hit a language wall, they hit it.\n\nFlamebait anyone? ;-) It is a different kind of fun -- let's say that\non top of knowing the performance tricks (or, to be more hip: \"design\npatterns\") for the hardware and OS, you also end up learning the\nperformance tricks of the interpreter/vm/whatever.\n\n> > It would be better to rsync Martins copy, he has a lot more bandwidth.\n\nI'm coming down to the office now to pick up my laptop, and I'll rsync\nit out to our git machine (also NZ kernel mirror, bandwidth should be\ngood). That's one of the things I've discovered with these large\ntrees: for the initial publish action, I just use rsync or scp.\nPerhaps I'm doing it wrong, but git-push doesn't optimise the\n'initialise repo', and it take ages (and it this case, it'd probably\nOOM).\n\n> So it will take me quite some time to download 2GB+, regardless of how fat\n> a pipe the other end has ;)\n\nRight-o. Linus, Jon, can you guys then ping me when you have cloned it\nsafely so I can take it down again?\n\ncheers,\n\n\nmartin\n"},{"id":"21510","messageId":"Pine.LNX.4.64.0606091640200.5498@g5.osdl.org","threadId":"4442","inReplyTo":"46a038f90606091637o6a0194d5yb413237253a372fc@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-09T23:43:17Z","receivedAt":"2006-06-09T23:43:17Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 10 Jun 2006, Martin Langhoff wrote:\n> \n> Exactly. The dog at this time is cvsps -- I also remember vague\n> promises from a list regular of publishing a git repo with cvsps2.1 +\n> some patches from the list.\n\nAhh. cvsps doesn't do anything incrementally, does it?\n\nAlthough it _does_ build up a cache of sorts, I think. That's not the \nparts I actually ever ended up looking at.\n\nBut yeah, a cvsps that blows up to a gig of VM and takes half an hour to \nparse things just for an incremental update would be a problem.\n\n> In any case, and for the record, my cvsps is 2.1 pristine. It handles\n> the mozilla repo alright, as long as I give it a lot of RAM. I _think_\n> it slurped 3GB with the mozilla cvs.\n\nOh, wow. Every single repo I've seen ends up having tons of complaints \nfrom pristine cvsps, but maybe that's because I only end up looking at the \nones with problems ;)\n\n> I'm coming down to the office now to pick up my laptop, and I'll rsync\n> it out to our git machine (also NZ kernel mirror, bandwidth should be\n> good). That's one of the things I've discovered with these large\n> trees: for the initial publish action, I just use rsync or scp.\n> Perhaps I'm doing it wrong, but git-push doesn't optimise the\n> 'initialise repo', and it take ages (and it this case, it'd probably\n> OOM).\n> \n> > So it will take me quite some time to download 2GB+, regardless of how fat\n> > a pipe the other end has ;)\n> \n> Right-o. Linus, Jon, can you guys then ping me when you have cloned it\n> safely so I can take it down again?\n\nTell me where/when it is, and I'll start slurping. Will let you know when \nI'm done.\n\n\t\tLinus\n"},{"id":"21511","messageId":"9e4733910606091700s49018cd5p3b66f8ef51b22d2e@mail.gmail.com","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606091640200.5498@g5.osdl.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-06-10T00:00:36Z","receivedAt":"2006-06-10T00:00:36Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 6/9/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> On Sat, 10 Jun 2006, Martin Langhoff wrote:\n> > In any case, and for the record, my cvsps is 2.1 pristine. It handles\n> > the mozilla repo alright, as long as I give it a lot of RAM. I _think_\n> > it slurped 3GB with the mozilla cvs.\n>\n> Oh, wow. Every single repo I've seen ends up having tons of complaints\n> from pristine cvsps, but maybe that's because I only end up looking at the\n> ones with problems ;)\n\nAre we sure cvsps is ok? It is generating 500MB of warnings when I run it.\n\nI have cvsps running at dreamhost currently. I had to modify cvs,\ncvps, git, etc to not repsond to signals to keep them from killing\neverything.\n\nI can clone 2GB git tree there. Let me know when it is up.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"21512","messageId":"Pine.LNX.4.64.0606091710560.5498@g5.osdl.org","threadId":"4442","inReplyTo":"9e4733910606091700s49018cd5p3b66f8ef51b22d2e@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-10T00:11:27Z","receivedAt":"2006-06-10T00:11:27Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 9 Jun 2006, Jon Smirl wrote:\n> \n> Are we sure cvsps is ok? It is generating 500MB of warnings when I run it.\n\nDo they go away with these patches?\n\n\t\tLinus\n---\ncommit 3d1ebcef6b4f9f6c9064efd64da4dd30d93c3c96\nAuthor: Linus Torvalds <torvalds@g5.osdl.org>\nDate:   Wed Mar 22 17:20:20 2006 -0800\n\n    Fix branch ancestor calculation\n    \n    Not having any ancestor at all means that any valid ancestor (even of\n    \"depth 0\") is fine.\n    \n    Signed-off-by: Linus Torvalds <torvalds@osdl.org>\n\ndiff --git a/cvsps.c b/cvsps.c\nindex c22147e..2695a0f 100644\n--- a/cvsps.c\n+++ b/cvsps.c\n@@ -2599,7 +2599,7 @@ static void determine_branch_ancestor(Pa\n \t * note: rev is the pre-commit revision, not the post-commit\n \t */\n \tif (!head_ps->ancestor_branch)\n-\t    d1 = 0;\n+\t    d1 = -1;\n \telse if (strcmp(ps->branch, rev->branch) == 0)\n \t    continue;\n \telse if (strcmp(head_ps->ancestor_branch, \"HEAD\") == 0)\n\ncommit 82fcf7e31bbeae3b01a8656549e9b8fd89d598eb\nAuthor: Linus Torvalds <torvalds@g5.osdl.org>\nDate:   Wed Mar 22 11:23:37 2006 -0800\n\n    Improve handling of file collisions in the same patchset\n    \n    Take the file revision into account.\n\ndiff --git a/cvsps.c b/cvsps.c\nindex 1e64e3c..c22147e 100644\n--- a/cvsps.c\n+++ b/cvsps.c\n@@ -2384,8 +2384,31 @@ void patch_set_add_member(PatchSet * ps,\n     for (next = ps->members.next; next != &ps->members; next = next->next) \n     {\n \tPatchSetMember * m = list_entry(next, PatchSetMember, link);\n-\tif (m->file == psm->file && ps->collision_link.next == NULL) \n-\t\tlist_add(&ps->collision_link, &collisions);\n+\tif (m->file == psm->file) {\n+\t\tint order = compare_rev_strings(psm->post_rev->rev, m->post_rev->rev);\n+\n+\t\t/*\n+\t\t * Same revision too? Add it to the collision list\n+\t\t * if it isn't already.\n+\t\t */\n+\t\tif (!order) {\n+\t\t\tif (ps->collision_link.next == NULL)\n+\t\t\t\tlist_add(&ps->collision_link, &collisions);\n+\t\t\treturn;\n+\t\t}\n+\n+\t\t/*\n+\t\t * If this is an older revision than the one we already have\n+\t\t * in this patchset, just ignore it\n+\t\t */\n+\t\tif (order < 0)\n+\t\t\treturn;\n+\n+\t\t/*\n+\t\t * This is a newer one, remove the old one\n+\t\t */\n+\t\tlist_del(&m->link);\n+\t}\n     }\n \n     psm->ps = ps;\n\ncommit 534120d9a47062eecd7b53fd7ac0b70d97feb4fd\nAuthor: Linus Torvalds <torvalds@g5.osdl.org>\nDate:   Wed Mar 22 11:20:59 2006 -0800\n\n    Increase log-length limit to 64kB\n    \n    Yeah, it should be dynamic. I'm lazy.\n\ndiff --git a/cvsps_types.h b/cvsps_types.h\nindex b41e2a9..dba145d 100644\n--- a/cvsps_types.h\n+++ b/cvsps_types.h\n@@ -8,7 +8,7 @@ #define CVSPS_TYPES_H\n \n #include <time.h>\n \n-#define LOG_STR_MAX 32768\n+#define LOG_STR_MAX 65536\n #define AUTH_STR_MAX 64\n #define REV_STR_MAX 64\n #define MIN(a, b) ((a) < (b) ? (a) : (b))\n"},{"id":"21513","messageId":"9e4733910606091716q67d4c5f9ra807b712d871e562@mail.gmail.com","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606091710560.5498@g5.osdl.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-06-10T00:16:35Z","receivedAt":"2006-06-10T00:16:35Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"I'll apply and give it a test.\n\nThey look like this for most of them.\n\nWARNING: Invalid PatchSet 151492, Tag JSS_4_0_RTM:\n    security/coreconf/HP-UX.mk:1.8=after,\nsecurity/jss/org/mozilla/jss/crypto/KeyPairAlgorithm.java:1.5=before.\nTreated as 'before'\nWARNING: Invalid PatchSet 151492, Tag JSS_4_0_RTM:\n    security/coreconf/HP-UX.mk:1.8=after,\nsecurity/jss/org/mozilla/jss/crypto/KeyPairGenerator.java:1.5=before.\nTreated as 'before'\nWARNING: Invalid PatchSet 151492, Tag JSS_4_0_RTM:\n    security/coreconf/HP-UX.mk:1.8=after,\nsecurity/jss/org/mozilla/jss/crypto/KeyPairGeneratorSpi.java:1.3=before.\nTreated as 'before'\nWARNING: Invalid PatchSet 151492, Tag JSS_4_0_RTM:\n    security/coreconf/HP-UX.mk:1.8=after,\nsecurity/jss/org/mozilla/jss/crypto/KeyWrapAlgorithm.java:1.8=before.\nTreated as 'before'\nWARNING: Invalid PatchSet 151492, Tag JSS_4_0_RTM:\n    security/coreconf/HP-UX.mk:1.8=after,\nsecurity/jss/org/mozilla/jss/crypto/KeyWrapper.java:1.8=before.\nTreated as 'before'\nWARNING: Invalid PatchSet 151492, Tag JSS_4_0_RTM:\n    security/coreconf/HP-UX.mk:1.8=after,\nsecurity/jss/org/mozilla/jss/crypto/Makefile:1.2=before. Treated as\n'before'\nWARNING: Invalid PatchSet 151492, Tag JSS_4_0_RTM:\n    security/coreconf/HP-UX.mk:1.8=after,\nsecurity/jss/org/mozilla/jss/crypto/NoSuchItemOnTokenException.java:1.3=before.\nTreated as 'before'\n\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"21516","messageId":"9e4733910606091745m103a8f69ieff197b60d3b7597@mail.gmail.com","threadId":"4442","inReplyTo":"9e4733910606091716q67d4c5f9ra807b712d871e562@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-06-10T00:45:25Z","receivedAt":"2006-06-10T00:45:25Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"They must be running some kind of process accounting at my host. As\nsoon as I hit 500MB RAM I get killed immediately. It is not from a\nsignal, I'm catching all of those. Maybe some kind of process\naccounting.\n\nI get this on the console:\n[1]+  Killed\nCVSROOT=~/jonsmirl.dreamhosters.com/mozilla/ cvsps -x --norc -A\nmozilla >mozilla.cvsps 2>mozilla.cvspserr\n\nand nothing on stdout or stderr.\n\nkernel string:\n 2.4.29-grsec+w+fhs6b+gr0501+nfs+a32+++p4+sata+c4+gr2b-v6.189\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"21518","messageId":"46a038f90606091814n1922bf25l94d913238b260296@mail.gmail.com","threadId":"4442","inReplyTo":"46a038f90606082006t5c6a5623q4b9cf7b036dad1e5@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-06-10T01:14:13Z","receivedAt":"2006-06-10T01:14:13Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 6/9/06, Martin Langhoff <martin.langhoff@gmail.com> wrote:\n> mozilla.git$ du -sh .git/\n> 2.0G    .git/\n\nOk -- pushed the repository out to our mirror box. Try:\n\n   git-clone http://mirrors.catalyst.net.nz/pub/mozilla.git/\n\nNow, good news. No, _very_ good news. As I was rsync'ing this out, and\nlooking at the repo, suddently something was odd. Apparently after a\ngit-repack -a -d OOMd on me, and I had posted this message, I re-ran\nit.\n\n[As it happens I have been running several imports of gentoo and moz\nlately on thebox. It is entirely possible that cvsps or a stray\ngit-cvsimport was sitting on a whole lot of ram at the time]\n\nNow I don't know how much memory or time this took, but it clearly\ncompleted ok. And, it's now a single pack, weighting a grand total of\n617MB\n\nSo my comments about OOM'ing were wrong apparently. Hey, if the whole\nhistory is actually only 617MB, then initial checkouts are back to\nsomething reasonable, I'd say.\n\ncheers,\n\n\n\nmartin\n"},{"id":"21519","messageId":"46a038f90606091823u45dd3dffsc584c0d4e0128b4c@mail.gmail.com","threadId":"4442","inReplyTo":"9e4733910606091317p26d66579mdf93db293f93fb50@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-06-10T01:23:35Z","receivedAt":"2006-06-10T01:23:35Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 6/10/06, Jon Smirl <jonsmirl@gmail.com> wrote:\n> The git tree that Martin got from cvsps is much smaller that the git\n> tree I got from going to svn then to git.  I don't why the trees are\n> 700KB different, it may be different amounts of packing, or one of the\n> conversion tools is losing something.\n\nDon't read too much into that. Packing/repacking points make a _huge_\ndifference, and even if one of our trees is a bit corrupt, the\npacksizes should be about the same.\n\n(With the patches I sent you we _are_ choosing to ignore a few\nbranches that don't seem to make sense in cvsps output. These will\nshow up in the error output -- what I saw were very old, possibly\ncorrupt branches there, stuff I wouldn't shed a tear over, but it is\nworth reviewing).\n\n> I haven't come up with anything that is likely to result in Mozilla\n> switching over to git. Right now it takes three days to convert the\n> tree. The tree will have to be run in parallel for a while to convince\n> everyone to switch. I don't have a solution to keeping it in sync in\n> near real time (commits would still go to CVS). Most Mozilla\n> developers are interested but the infrastructure needs some help.\n\nDon't worry about the initial import time. Once you've done it, you\ncan run the incremental import (which will take a few minutes) even\nhourly to keep 'in sync'.\n\n> Martin has also brought up the problem with needing a partial clone so\n> that everyone doesn't have to bring down the entire repository. A\n> trunk checkout is 340MB and Martin's git tree is 2GB (mine 2.7GB).  A\n> kernel tree is only 680M.\n\nNow that I have managed to repack the repo, it is indeed back in the\n600M range. Actually, I just re-repacked, it took under a minute, and\nit shrank down to 607MB.\n\nYay.\n\nI'm sure that if you git-repack -a -d on a machine with plenty of\nmemory once or twice, we'll have matching packs.\n\ncheers,\n\n\n\nmartin\n"},{"id":"21520","messageId":"Pine.LNX.4.64.0606091825080.5498@g5.osdl.org","threadId":"4442","inReplyTo":"46a038f90606091814n1922bf25l94d913238b260296@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-10T01:33:00Z","receivedAt":"2006-06-10T01:33:00Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 10 Jun 2006, Martin Langhoff wrote:\n> \n> Now I don't know how much memory or time this took, but it clearly\n> completed ok. And, it's now a single pack, weighting a grand total of\n> 617MB\n\nOk, that's more than reasonable. That should be fairly easily mapped on a \n32-bit architecture without any huge problems, even with some VM \nfragmentation going on. It might be borderline (and you definitely want a \n3:1 VM user:kernel split), but considering that the original CVS archive \nwas apparently 3GB, having a single 617M pack-file is still pretty damn \ngood.  That's like 20% of the original, with all the obvious distribution \nadvantages.\n\nClearly this whole thing _does_ show that we could improve the process of \nimporting things from CVS a whole lot, and I assume your 617MB pack \ndoesn't have the nice name/email translations so it needs to be fixed up, \nbut it sounds like on the whole the core git design came through with \nshining colors, even if we may want to polish things up a bit ;)\n\nI'm downloading the thing right now.\n\n\t\t\tLinus\n"},{"id":"21522","messageId":"Pine.LNX.4.64.0606091837040.5498@g5.osdl.org","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606091825080.5498@g5.osdl.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-10T01:43:11Z","receivedAt":"2006-06-10T01:43:11Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 9 Jun 2006, Linus Torvalds wrote:\n> \n> That's like 20% of the original, with all the obvious distribution \n> advantages.\n\nBtw, does anybody know roughly how much data a initial \"cvs co\" takes on \nthe mozilla repo? Git will obviously get the whole history, and that will \ninevitably be bigger than getting a single check-out, but it's not \nnecessarily orders of magnitude bigger.\n\nIt could be that getting a whole git archive is not _that_ much more \nexpnsive than getting a single version, considering how well history \ncompresses (eg the kernel git arhive isn't orders of magnitude bigger than \na single compressed tar-ball of the sources).\n\nAt that point, it's probably a pretty usable alternative.\n\n(Although, to be fair, we almost certainly have to improve \"git-rev-list \n--objects --all\" performance on that thing, since that's going to \notherwise make it totally impossible to do initial clones using the native \ngit protocol, and make git look bad).\n\n\t\t\tLinus\n"},{"id":"21523","messageId":"9e4733910606091848r5fb4d565taabfc5198140daf2@mail.gmail.com","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606091837040.5498@g5.osdl.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-06-10T01:48:04Z","receivedAt":"2006-06-10T01:48:04Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 6/9/06, Linus Torvalds <torvalds@osdl.org> wrote:\n>\n>\n> On Fri, 9 Jun 2006, Linus Torvalds wrote:\n> >\n> > That's like 20% of the original, with all the obvious distribution\n> > advantages.\n>\n> Btw, does anybody know roughly how much data a initial \"cvs co\" takes on\n> the mozilla repo? Git will obviously get the whole history, and that will\n> inevitably be bigger than getting a single check-out, but it's not\n> necessarily orders of magnitude bigger.\n\n339MB for initial checkout\n\n> It could be that getting a whole git archive is not _that_ much more\n> expnsive than getting a single version, considering how well history\n> compresses (eg the kernel git arhive isn't orders of magnitude bigger than\n> a single compressed tar-ball of the sources).\n>\n> At that point, it's probably a pretty usable alternative.\n>\n> (Although, to be fair, we almost certainly have to improve \"git-rev-list\n> --objects --all\" performance on that thing, since that's going to\n> otherwise make it totally impossible to do initial clones using the native\n> git protocol, and make git look bad).\n>\n>                         Linus\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"21524","messageId":"Pine.LNX.4.64.0606091853180.5498@g5.osdl.org","threadId":"4442","inReplyTo":"9e4733910606091848r5fb4d565taabfc5198140daf2@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-10T01:59:37Z","receivedAt":"2006-06-10T01:59:37Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 9 Jun 2006, Jon Smirl wrote:\n> > \n> > Btw, does anybody know roughly how much data a initial \"cvs co\" takes on\n> > the mozilla repo? Git will obviously get the whole history, and that will\n> > inevitably be bigger than getting a single check-out, but it's not\n> > necessarily orders of magnitude bigger.\n> \n> 339MB for initial checkout\n\nAnd I think people run :pserver: with compression by default, so we're \nlikely talking about half that in actual download overhead, no?\n\nSo a git clone would be about (wild handwaving, don't look at all the \nassumptions) four times as expensive - assuming we only look at a poor DSL \nline as the expense - as an initial CVS co, but you'd get the _whole_ \nhistory. Which may or may not make up for it. For some people it will, for \nothers it won't.\n\nOf course, to make up for some of the initial costs, I suspect that some \npeople who are used to \"cvs update\" taking 15 minutes to update two files, \nit would be a serious relief to see the git kind of \"300 objects in five \nseconds\" kinds of pulls.\n\nAlthough I guess that's one of the CVS things that SVN improved on. At \nleast I'd hope so ;/\n\n\t\t\tLinus\n"},{"id":"21526","messageId":"9e4733910606091921o1d07826w8292dc22b1872345@mail.gmail.com","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606091853180.5498@g5.osdl.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-06-10T02:21:17Z","receivedAt":"2006-06-10T02:21:17Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 6/9/06, Linus Torvalds <torvalds@osdl.org> wrote:\n>\n>\n> On Fri, 9 Jun 2006, Jon Smirl wrote:\n> > >\n> > > Btw, does anybody know roughly how much data a initial \"cvs co\" takes on\n> > > the mozilla repo? Git will obviously get the whole history, and that will\n> > > inevitably be bigger than getting a single check-out, but it's not\n> > > necessarily orders of magnitude bigger.\n> >\n> > 339MB for initial checkout\n>\n> And I think people run :pserver: with compression by default, so we're\n> likely talking about half that in actual download overhead, no?\n>\n> So a git clone would be about (wild handwaving, don't look at all the\n> assumptions) four times as expensive - assuming we only look at a poor DSL\n> line as the expense - as an initial CVS co, but you'd get the _whole_\n> history. Which may or may not make up for it. For some people it will, for\n> others it won't.\n\nCould you clone the repo and delete changesets earlier than 2004? Then\nI would clone the small repo and work with it. Later I decide I want\nfull history, can I pull from a full repository at that point and get\nupdated? That would need a flag to trigger it since I don't want full\nhistory to come over if I am just getting updates from someone else's\ntree that has a full history.\n\n>\n> Of course, to make up for some of the initial costs, I suspect that some\n> people who are used to \"cvs update\" taking 15 minutes to update two files,\n> it would be a serious relief to see the git kind of \"300 objects in five\n> seconds\" kinds of pulls.\n\nNo more cvs diff taking four minutes to finish. I have to do that\nevery time I want to generate a 10 line patch. Diffs can run locally.\nNo more cvs update to replace files I deleted because I messed up\nedits in them. And I can have local branches, yeah!\n\nWhat are we going to do about the BEOS developers on Mozilla? There\nare a couple more obscure OSes.\n\n> Although I guess that's one of the CVS things that SVN improved on. At\n> least I'd hope so ;/\n>\n>                         Linus\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"21527","messageId":"9e4733910606091930j70da8ca5g1ec91c98ed9a2445@mail.gmail.com","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606091853180.5498@g5.osdl.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-06-10T02:30:07Z","receivedAt":"2006-06-10T02:30:07Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 6/9/06, Linus Torvalds <torvalds@osdl.org> wrote:\n>\n>\n> On Fri, 9 Jun 2006, Jon Smirl wrote:\n> > >\n> > > Btw, does anybody know roughly how much data a initial \"cvs co\" takes on\n> > > the mozilla repo? Git will obviously get the whole history, and that will\n> > > inevitably be bigger than getting a single check-out, but it's not\n> > > necessarily orders of magnitude bigger.\n> >\n> > 339MB for initial checkout\n\nI ran the checkout through bzip and it is 36.4MB, 46.4MB with zip.\nSo the ratio may be 15 to 1 for the cvs co vs git\n\n> And I think people run :pserver: with compression by default, so we're\n> likely talking about half that in actual download overhead, no?\n>\n> So a git clone would be about (wild handwaving, don't look at all the\n> assumptions) four times as expensive - assuming we only look at a poor DSL\n> line as the expense - as an initial CVS co, but you'd get the _whole_\n> history. Which may or may not make up for it. For some people it will, for\n> others it won't.\n>\n> Of course, to make up for some of the initial costs, I suspect that some\n> people who are used to \"cvs update\" taking 15 minutes to update two files,\n> it would be a serious relief to see the git kind of \"300 objects in five\n> seconds\" kinds of pulls.\n>\n> Although I guess that's one of the CVS things that SVN improved on. At\n> least I'd hope so ;/\n>\n>                         Linus\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"21528","messageId":"87y7w5lowc.wl%cworth@cworth.org","threadId":"4442","inReplyTo":"9e4733910606091921o1d07826w8292dc22b1872345@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Carl Worth","fromEmail":"cworth@cworth.org","sentAt":"2006-06-10T02:34:27Z","receivedAt":"2006-06-10T02:34:27Z","isPatch":false,"sender":{"key":"cworth@cworth.org","avatar":"https://gravatar.com/avatar/3746dc28cde609bdbd7f939058356e7e2bbd16d21e32274df0725eb3d998bc5b?d=mp&s=160"},"body":"On Fri, 9 Jun 2006 22:21:17 -0400, \"Jon Smirl\" wrote:\n> \n> Could you clone the repo and delete changesets earlier than 2004? Then\n> I would clone the small repo and work with it. Later I decide I want\n> full history, can I pull from a full repository at that point and get\n> updated? That would need a flag to trigger it since I don't want full\n> history to come over if I am just getting updates from someone else's\n> tree that has a full history.\n\nThis is clearly a desirable feature, and has been requested by several\npeople (including myself) looking to switch some large-ish histories\nfrom an existing system to git.\n\nIf you'd like to look through git archives for some discussion of the\nissues that would be involved here, look for \"shallow clone\".\n\nThere's a related proposal termed \"lazy clone\" for one that would pull\ndown missing objects as needed over the network.\n\nMy impression is that both things will eventually be implemented.\nThere's certainly nothing fundamental in git that will prevent them,\n(though there will be some interesting things to resolve as a real\npatch for this stuff is explored).\n\n-Carl\n"},{"id":"21529","messageId":"Pine.LNX.4.64.0606092000110.5498@g5.osdl.org","threadId":"4442","inReplyTo":"9e4733910606091921o1d07826w8292dc22b1872345@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-10T03:01:08Z","receivedAt":"2006-06-10T03:01:08Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 9 Jun 2006, Jon Smirl wrote:\n>\n> No more cvs diff taking four minutes to finish. I have to do that\n> every time I want to generate a 10 line patch. Diffs can run locally.\n> No more cvs update to replace files I deleted because I messed up\n> edits in them. And I can have local branches, yeah!\n\nMore importantly, when the CVS server is down (can you say \n\"sourceforge\"?), who cares?\n\n> What are we going to do about the BEOS developers on Mozilla? There\n> are a couple more obscure OSes.\n\nWell, the git cvsserver exporter apparently works well enough...\n\n\t\t\tLinus\n"},{"id":"21530","messageId":"Pine.LNX.4.64.0606092001590.5498@g5.osdl.org","threadId":"4442","inReplyTo":"87y7w5lowc.wl%cworth@cworth.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-10T03:08:54Z","receivedAt":"2006-06-10T03:08:54Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 9 Jun 2006, Carl Worth wrote:\n\n> On Fri, 9 Jun 2006 22:21:17 -0400, \"Jon Smirl\" wrote:\n> > \n> > Could you clone the repo and delete changesets earlier than 2004? Then\n> > I would clone the small repo and work with it. Later I decide I want\n> > full history, can I pull from a full repository at that point and get\n> > updated? That would need a flag to trigger it since I don't want full\n> > history to come over if I am just getting updates from someone else's\n> > tree that has a full history.\n> \n> This is clearly a desirable feature, and has been requested by several\n> people (including myself) looking to switch some large-ish histories\n> from an existing system to git.\n\nThe thing is, to some degree it's really fundamentally hard.\n\nIt's easy for a linear history. What you do for a linear history is to \njust get the top commit, and the tree associated with it, and then you \ncauterize the parent by just grafting it to go away. Boom. You're done.\n\nThe problems are that if the preceding history _wasn't_ linear (or, in \nfact, _subsequent_ development refers to it by having branched off at an \nearlier point), and you try to pull your updates, the other end (that \nknows about all the history) will assume you have all the history that you \ndon't have, and will send you a pack assuming that.\n\nWhich won't even necessarily have all the tree/blob objects (it assumed \nyou already had them), but more annoyingly, the history won't be \ncauterized, and you'll have dangling commits. Which you can cauterize by \nhand, of course, but you literally _will_ have to get the objects and \ncauterize the thing by hand.\n\nYou're right that it's not \"fundamentally impossible\" to do: the git \nformat certainly _allows_ it. But the git protocol handshake really does \nend up optimizing away all the unnecessary work by knowing that the other \nside will have all the shared history, so lacking the shared history will \nmean that you're a bit screwed.\n\nUsing the http protocol actually works. It doesn't do any handshake: it \nwill just fetch objects from the other end as it needs them. The downside, \nof course, is that it also doesn't understand packs, so if the source is \npacked (and it pretty much _will_ be, for any big source), you're going to \nend up getting it all _anyway_.\n\n\t\tLinus\n"},{"id":"21531","messageId":"46a038f90606092041neadcc54n2acb6272d1f71de7@mail.gmail.com","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606091853180.5498@g5.osdl.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-06-10T03:41:39Z","receivedAt":"2006-06-10T03:41:39Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 6/10/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> On Fri, 9 Jun 2006, Jon Smirl wrote:\n> > >\n> > > Btw, does anybody know roughly how much data a initial \"cvs co\" takes on\n> > > the mozilla repo? Git will obviously get the whole history, and that will\n> > > inevitably be bigger than getting a single check-out, but it's not\n> > > necessarily orders of magnitude bigger.\n> >\n> > 339MB for initial checkout\n>\n> And I think people run :pserver: with compression by default, so we're\n> likely talking about half that in actual download overhead, no?\n\nYes, most people have -z3, and I agree with you, on paper it sounds\nlike the cost is 1/4 of a git clone.\n\nHowever.\n\nThe CVS protocol is very chatty because the client _acts_ extremely\nstupid. It says, ok, I got here an empty directory, and the server\nwalks the client through every little step. And all that chatter is\nuncompressed cleartext under pserver.\n\nSo the per-file and per-directory overhead are significant. I can do a\ncvs checkout via pserver:localhost but I don't know off-the-cuff how\nto measure the traffic. Hints?\n\ncheers,\n\n\nmartin\n"},{"id":"21533","messageId":"7v3bedll62.fsf@assigned-by-dhcp.cox.net","threadId":"4442","inReplyTo":"46a038f90606092041neadcc54n2acb6272d1f71de7@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-06-10T03:55:01Z","receivedAt":"2006-06-10T03:55:01Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Martin Langhoff\" <martin.langhoff@gmail.com> writes:\n\n> Yes, most people have -z3, and I agree with you, on paper it sounds\n> like the cost is 1/4 of a git clone.\n>\n> However.\n>\n> The CVS protocol is very chatty because the client _acts_ extremely\n> stupid. It says, ok, I got here an empty directory, and the server\n> walks the client through every little step. And all that chatter is\n> uncompressed cleartext under pserver.\n>\n> So the per-file and per-directory overhead are significant. I can do a\n> cvs checkout via pserver:localhost but I don't know off-the-cuff how\n> to measure the traffic. Hints?\n\nIf you have an otherwise unused interface, you can look at\nifconfig output and see RX/TX bytes?  But that sounds very\ncrude.\n\nRunning it through a proxy perhaps?\n"},{"id":"21534","messageId":"Pine.LNX.4.64.0606092043460.5498@g5.osdl.org","threadId":"4442","inReplyTo":"46a038f90606092041neadcc54n2acb6272d1f71de7@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-10T04:02:34Z","receivedAt":"2006-06-10T04:02:34Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 10 Jun 2006, Martin Langhoff wrote:\n> \n> So the per-file and per-directory overhead are significant. I can do a\n> cvs checkout via pserver:localhost but I don't know off-the-cuff how\n> to measure the traffic. Hints?\n\nOver localhost, you won't see the biggest issue, which is just latency.\n\nThe git protocol should be absolutely <i>wonderful</i> with bad latency, \nbecause once the early bakc-and-forth on what each side has is done, \nthere's no synchronization any more - it's all just streaming, with \nfull-frame TCP.\n\nIf :pserver: does per-file \"hey, what are you up to\" kind of \nsyncronization, the big killer would be the latency from one end to the \nother, regardless of any throughput.\n\nYou can try to approximate the latency by just looking at the number of \npackets, and using a large MTU (and on localhost, the MTU will be pretty \nlarge - roughly 16kB. Don't count packet size at all, just count how many \npackets each protocol sends (both ways), ignoring packets that are just \nempty ACK's.\n\nI don't know how to build a tcpdump expression for \"TCP packet with an \nempty payload\", but I bet it's possible.\n\n[ And I won't guarantee that it's a wonderful approximation for \"network \n  cost\", but I think it's potentially a reasonably good one. It's totally \n  realistic to equate 32kB of _streaming_ data (two packets flowing in \n  one direction with no synchronization) with just a single byte of data \n  going back-and-forth synchronously ]\n\n\t\tLinus\n"},{"id":"21536","messageId":"Pine.LNX.4.64.0606092109380.5498@g5.osdl.org","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606092043460.5498@g5.osdl.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-10T04:11:58Z","receivedAt":"2006-06-10T04:11:58Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 9 Jun 2006, Linus Torvalds wrote:\n> \n> You can try to approximate the latency by just looking at the number of \n> packets, and using a large MTU (and on localhost, the MTU will be pretty \n> large - roughly 16kB. Don't count packet size at all, just count how many \n> packets each protocol sends (both ways), ignoring packets that are just \n> empty ACK's.\n\nBtw, the reason you should ignore empty acks is that they happen when you \nhave a nice streaming one-way thing, because the TCP rules say that you \nshould send an ACK every two full packets minimum, even if you have \nnothing to say.\n\nSo empty acks really approximate to \"streaming data\", while packets with \npayload _could_ obviously mean \"nice streaming data going both ways\", but \nalmost always end up being synchronization discussion of some sort.\n\n\t\tLinus\n"},{"id":"21539","messageId":"9e4733910606092302h646ff554p107564417183e350@mail.gmail.com","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606092109380.5498@g5.osdl.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-06-10T06:02:58Z","receivedAt":"2006-06-10T06:02:58Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"Here's a new transport problem. When using git-clone to fetch Martin's\ntree it kept failing for me at dreamhost. I had a parallel fetch\nrunning on my local machine which has a much slower net connection. It\nfinally finished and I am watching the end phase where it prints all\nof the 'walk' messages. The git-http-fetch process has jumped up to\n800MB in size after being 2MB during the download. dreamhost has a\n500MB process size limit so that is why my fetches kept failing there.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"21540","messageId":"7vr71xk047.fsf@assigned-by-dhcp.cox.net","threadId":"4442","inReplyTo":"9e4733910606092302h646ff554p107564417183e350@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-06-10T06:15:04Z","receivedAt":"2006-06-10T06:15:04Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Jon Smirl\" <jonsmirl@gmail.com> writes:\n\n> Here's a new transport problem. When using git-clone to fetch Martin's\n> tree it kept failing for me at dreamhost. I had a parallel fetch\n> running on my local machine which has a much slower net connection. It\n> finally finished and I am watching the end phase where it prints all\n> of the 'walk' messages. The git-http-fetch process has jumped up to\n> 800MB in size after being 2MB during the download. dreamhost has a\n> 500MB process size limit so that is why my fetches kept failing there.\n\nThe http-fetch process uses by mmaping the downloaded pack, and\nif I recall correctly we are talking about 600MB pack, so 500MB\nlimit sounds impossible, perhaps?\n"},{"id":"21541","messageId":"e6dvds$oes$1@sea.gmane.org","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606092001590.5498@g5.osdl.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-06-10T08:21:27Z","receivedAt":"2006-06-10T08:21:27Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Linus Torvalds wrote:\n\n\n> On Fri, 9 Jun 2006, Carl Worth wrote:\n> \n>> On Fri, 9 Jun 2006 22:21:17 -0400, \"Jon Smirl\" wrote:\n>> > \n>> > Could you clone the repo and delete changesets earlier than 2004? Then\n>> > I would clone the small repo and work with it. Later I decide I want\n>> > full history, can I pull from a full repository at that point and get\n>> > updated? That would need a flag to trigger it since I don't want full\n>> > history to come over if I am just getting updates from someone else's\n>> > tree that has a full history.\n>> \n>> This is clearly a desirable feature, and has been requested by several\n>> people (including myself) looking to switch some large-ish histories\n>> from an existing system to git.\n> \n> The thing is, to some degree it's really fundamentally hard.\n> \n> It's easy for a linear history. What you do for a linear history is to \n> just get the top commit, and the tree associated with it, and then you \n> cauterize the parent by just grafting it to go away. Boom. You're done.\n> \n> The problems are that if the preceding history _wasn't_ linear (or, in \n> fact, _subsequent_ development refers to it by having branched off at an \n> earlier point), and you try to pull your updates, the other end (that \n> knows about all the history) will assume you have all the history that you \n> don't have, and will send you a pack assuming that.\n\nCouldn't it be solved by enhancing initial handshake to send from puller\n(object receivier) to pullee (object sender) the contents of graft file, or\nbetter the contents of cauterizing graft file - without splitting graft\nfile we better have an option to send graft file or not, when graft file is\nused to join historical repository line of development not to cauterize\nhistory.\n\nThen the sender would use sent cauterizing history graft file for\ncalculating which objects to sedn _only_, \"in memory\" cauterizing it's own\nhistory.\n\nMain disadvantage is if one cauterized history too eagerly, and shallow\nclone history can lack merge bases, and have no way to get them _simply_\nusing this approach...\n\n\nNow I guess you would tell me why this very simple idea is stupid...\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"21542","messageId":"448A847C.20105@dawes.za.net","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606092001590.5498@g5.osdl.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Rogan Dawes","fromEmail":"lists@dawes.za.net","sentAt":"2006-06-10T08:36:12Z","receivedAt":"2006-06-10T08:36:12Z","isPatch":false,"sender":{"key":"lists@dawes.za.net","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Fri, 9 Jun 2006, Carl Worth wrote:\n> \n>> On Fri, 9 Jun 2006 22:21:17 -0400, \"Jon Smirl\" wrote:\n>>> Could you clone the repo and delete changesets earlier than 2004? Then\n>>> I would clone the small repo and work with it. Later I decide I want\n>>> full history, can I pull from a full repository at that point and get\n>>> updated? That would need a flag to trigger it since I don't want full\n>>> history to come over if I am just getting updates from someone else's\n>>> tree that has a full history.\n>> This is clearly a desirable feature, and has been requested by several\n>> people (including myself) looking to switch some large-ish histories\n>> from an existing system to git.\n> \n> The thing is, to some degree it's really fundamentally hard.\n> \n> It's easy for a linear history. What you do for a linear history is to \n> just get the top commit, and the tree associated with it, and then you \n> cauterize the parent by just grafting it to go away. Boom. You're done.\n> \n> The problems are that if the preceding history _wasn't_ linear (or, in \n> fact, _subsequent_ development refers to it by having branched off at an \n> earlier point), and you try to pull your updates, the other end (that \n> knows about all the history) will assume you have all the history that you \n> don't have, and will send you a pack assuming that.\n> \n> Which won't even necessarily have all the tree/blob objects (it assumed \n> you already had them), but more annoyingly, the history won't be \n> cauterized, and you'll have dangling commits. Which you can cauterize by \n> hand, of course, but you literally _will_ have to get the objects and \n> cauterize the thing by hand.\n> \n> You're right that it's not \"fundamentally impossible\" to do: the git \n> format certainly _allows_ it. But the git protocol handshake really does \n> end up optimizing away all the unnecessary work by knowing that the other \n> side will have all the shared history, so lacking the shared history will \n> mean that you're a bit screwed.\n\nHere's an idea. How about separating trees and commits from the actual \nblobs (e.g. in separate packs)? My reasoning is that the commits and \ntrees should only be a small portion of the overall repository size, and \nshould not be that expensive to transfer. (Of course, this is only a \nguess, and needs some numbers to back it up.)\n\nSo, a shallow clone would receive all of the tree objects, and all of \nthe commit objects, and could then request a pack containing the blobs \nrepresented by the current HEAD.\n\nIn this way, the user has a history that will show all of the commit \nmessages, and would be able to see _which_ files have changed over time \ne.g. gitk would still work - except for the actual file level diff, \"git \nlog\" should also still work, etc\n\nThis would also enable other optimisations.\n\nFor example, documentation people would only need to get the objects \nunder the doc/ tree, and would not need to actually check out the \nsource. Git could detect any actual changes by checking whether it has \nthe previous blob in its local repository, and whether the file exists \nlocally. Creating a patch would obviously require that the person checks \nout the previous version, but one could theoretically commit a new blob \nto a repo without having the previous one (not saying that this would be \na good idea, of course)\n\nThis would probably require Eric Biederman's \"direct access to blob\" \npatches, I guess, in order to be feasible.\n\nRegards,\n\nRogan\n"},{"id":"21544","messageId":"7vac8lidwi.fsf@assigned-by-dhcp.cox.net","threadId":"4442","inReplyTo":"e6dvds$oes$1@sea.gmane.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-06-10T09:00:13Z","receivedAt":"2006-06-10T09:00:13Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> writes:\n\n> Couldn't it be solved by enhancing initial handshake to send from puller\n> (object receivier) to pullee (object sender) the contents of graft file, or\n> better the contents of cauterizing graft file - without splitting graft\n> file we better have an option to send graft file or not, when graft file is\n> used to join historical repository line of development not to cauterize\n> history.\n>\n> Then the sender would use sent cauterizing history graft file for\n> calculating which objects to sedn _only_, \"in memory\" cauterizing it's own\n> history.\n>\n> Now I guess you would tell me why this very simple idea is stupid...\n\nIt is not stupid at all; what you said is actually on a correct\ntrack.  You indeed just reinvented a half of what I've outlined\nearlier for implementing shallow clone (the other half you\nmissed is that the graft exchange needs to happen both ways,\nlimiting the commit ancestry graph the both ends walk to the\nintersection of the fake view of the ancestry graph both ends\nhave, but that is a minor detail).\n\nThe problem is that what Linus described as \"fundamentally hard\"\nis not the initial \"shallow clone\" stage, but lies elsewhere.\nNamely, what to do after you create such a shallow clone and\nwhen you want to unplug an earlier cauterization points.\n\nIn order to unplug a cauterization point (a commit we faked to\nbe parentless earlier, whose parents and associated objects we\nought to have but we do not because we made a shallow clone),\nthe downloader needs to re-fetch that commit while temporarily\npretending that it does not have any objects that are newer,\nperhaps defining another earlier point as a new cauterization\npoint at the same time.  Git format allows for that, and the\nprotocol exchange certainly can be extensible to support\nsomething like that, but the design work would be quite\ninvolved.\n"},{"id":"21545","messageId":"7vzmglgyz0.fsf@assigned-by-dhcp.cox.net","threadId":"4442","inReplyTo":"448A847C.20105@dawes.za.net","subject":"Re: Figured out how to get Mozilla into git","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-06-10T09:08:03Z","receivedAt":"2006-06-10T09:08:03Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Rogan Dawes <lists@dawes.za.net> writes:\n\n> Here's an idea. How about separating trees and commits from the actual\n> blobs (e.g. in separate packs)?\n\nIf I remember my numbers correctly, trees for any project with a\nsize that matters contribute nonnegligible amount of the total\npack weight.  Perhaps 10-25%.\n\n> In this way, the user has a history that will show all of the commit\n> messages, and would be able to see _which_ files have changed over\n> time e.g. gitk would still work - except for the actual file level\n> diff, \"git log\" should also still work, etc\n\nI suspect it would make a very unpleasant system to use.\nSometimes \"git diff -p\" would show diffs, and other times it\nmysteriously complain saying that it lacks necessary blobs to do\nits job.  You cannot even run fsck and tell from its output\nwhich missing objects are OK (because you chose to create such a\nsparse repository) and which are real corruption.\n\nA shallow clone with explicit cauterization in grafts file at\nleast would not have that problem. Although the user will still\nnot see the exact same result as what would happen in a full\nrepository, at least we can say \"your git log ends at that\ncommit because your copy of the history does not go back beyond\nthat\" and the user would understand.\n"},{"id":"21548","messageId":"448ADB8A.3070506@dawes.za.net","threadId":"4442","inReplyTo":"7vzmglgyz0.fsf@assigned-by-dhcp.cox.net","subject":"Re: Figured out how to get Mozilla into git","fromName":"Rogan Dawes","fromEmail":"discard@dawes.za.net","sentAt":"2006-06-10T14:47:38Z","receivedAt":"2006-06-10T14:47:38Z","isPatch":false,"sender":{"key":"discard@dawes.za.net","avatar":null},"body":"Junio C Hamano wrote:\n> Rogan Dawes <lists@dawes.za.net> writes:\n> \n>> Here's an idea. How about separating trees and commits from the actual\n>> blobs (e.g. in separate packs)?\n> \n> If I remember my numbers correctly, trees for any project with a\n> size that matters contribute nonnegligible amount of the total\n> pack weight.  Perhaps 10-25%.\n\nOut of curiosity, do you think that it may be possible for tree objects \nto compress more/better if they are packed together? Or does the \nexisting pack compression logic already do the diff against similar tree \nobjects?\n\n>> In this way, the user has a history that will show all of the commit\n>> messages, and would be able to see _which_ files have changed over\n>> time e.g. gitk would still work - except for the actual file level\n>> diff, \"git log\" should also still work, etc\n> \n> I suspect it would make a very unpleasant system to use.\n> Sometimes \"git diff -p\" would show diffs, and other times it\n> mysteriously complain saying that it lacks necessary blobs to do\n> its job.  You cannot even run fsck and tell from its output\n> which missing objects are OK (because you chose to create such a\n> sparse repository) and which are real corruption.\n\nThe fsck problem could be worked around by maintaining a list of objects \nthat are explicitly not expected to be present. As the list gets shorter \n(perhaps as diffs are performed, other parts of the blob history are \nretrieved, etc), the list will get shorter until we have a complete \nclone of the original tree.\n\nOf course diffs against a version further back in the history would \nfail. But if you start with a checkout of a complete tree, any changes \nmade since that point would at least have one version to compare against.\n\nIn effect, what we would have is a caching repository (or as Jakub said, \na lazy clone). An initial checkout would effectively be pre-seeding the \ncache. One does not necessarily even need to get the complete set of \ncommit and tree objects, either. The bare minimum would probably be to \nget the HEAD commit, and the tree objects that correspond to that commit.\n\nAt that point, one could populate the \"uncached objects\" list with the \nparent commits. One would not be in a position to get any history at \nall, of course.\n\nAs the user performs various operations, e.g. git log, git could either \ngo and fetch the necessary objects (updating the uncached list as it \ngoes), or fail with a message such as \"Cannot perform the requested \noperation - required objects are not available\". (We may require another \nutility that would list the objects required for an operation, and \ncompare it against the list of \"uncached objects\", printing out a list \nof which are not yet available locally. I realise that this may be \nexpensive. Maybe a repo configuration option \"cached\" to enable or \ndisable this.)\n\nAs Jakub suggested, it would be necessary to configure the location of \nthe source for any missing objects, but that is probably in the repo \nconfig anyway.\n\n> A shallow clone with explicit cauterization in grafts file at\n> least would not have that problem. Although the user will still\n> not see the exact same result as what would happen in a full\n> repository, at least we can say \"your git log ends at that\n> commit because your copy of the history does not go back beyond\n> that\" and the user would understand.\n\nOr, we could say, perform the operation while you are online, and can \naccess the necessary objects. If the user has explicitly chosen to make \na lazy clone, then they should expect that at some point, whatever they \ndo may require them to be online to access items that they have not yet \ncloned.\n\nRogan\n"},{"id":"21549","messageId":"e6emmm$idv$1@sea.gmane.org","threadId":"4442","inReplyTo":"448ADB8A.3070506@dawes.za.net","subject":"Re: Figured out how to get Mozilla into git","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-06-10T14:58:43Z","receivedAt":"2006-06-10T14:58:43Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Rogan Dawes wrote:\n\n> Junio C Hamano wrote:\n>> Rogan Dawes <lists@dawes.za.net> writes:\n>> \n>>> Here's an idea. How about separating trees and commits from the actual\n>>> blobs (e.g. in separate packs)?\n>> \n>> If I remember my numbers correctly, trees for any project with a\n>> size that matters contribute nonnegligible amount of the total\n>> pack weight.  Perhaps 10-25%.\n> \n> Out of curiosity, do you think that it may be possible for tree objects \n> to compress more/better if they are packed together? Or does the \n> existing pack compression logic already do the diff against similar tree \n> objects?\n\nThe problem with compressing and deltafying trees is with sha1 objects\nidentifiers, I guess.\n\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"21550","messageId":"Pine.LNX.4.64.0606101113130.2703@localhost.localdomain","threadId":"4442","inReplyTo":"448ADB8A.3070506@dawes.za.net","subject":"Re: Figured out how to get Mozilla into git","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2006-06-10T15:14:42Z","receivedAt":"2006-06-10T15:14:42Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sat, 10 Jun 2006, Rogan Dawes wrote:\n\n> Out of curiosity, do you think that it may be possible for tree objects to\n> compress more/better if they are packed together? Or does the existing pack\n> compression logic already do the diff against similar tree objects?\n\nTree objects for the same directories are already packed and deltified \nagainst each other in a pack.\n\n\nNicolas\n"},{"id":"21551","messageId":"9e4733910606100844v5f4765d8o85c9a6f239faed43@mail.gmail.com","threadId":"4442","inReplyTo":"7vr71xk047.fsf@assigned-by-dhcp.cox.net","subject":"Re: Figured out how to get Mozilla into git","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-06-10T15:44:58Z","receivedAt":"2006-06-10T15:44:58Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 6/10/06, Junio C Hamano <junkio@cox.net> wrote:\n> \"Jon Smirl\" <jonsmirl@gmail.com> writes:\n>\n> > Here's a new transport problem. When using git-clone to fetch Martin's\n> > tree it kept failing for me at dreamhost. I had a parallel fetch\n> > running on my local machine which has a much slower net connection. It\n> > finally finished and I am watching the end phase where it prints all\n> > of the 'walk' messages. The git-http-fetch process has jumped up to\n> > 800MB in size after being 2MB during the download. dreamhost has a\n> > 500MB process size limit so that is why my fetches kept failing there.\n>\n> The http-fetch process uses by mmaping the downloaded pack, and\n> if I recall correctly we are talking about 600MB pack, so 500MB\n> limit sounds impossible, perhaps?\n\nThe fetch on my local machine failed too. It left nothing behind, now\nI have to download the 680MB again.\n\nwalk 1f19465388a4ef7aff7527a13f16122a809487d4\nwalk c3ca840256e3767d08c649f8d2761a1a887351ab\nwalk 7a74e42699320c02b814b88beadb1ae65009e745\nerror: Couldn't get\nhttp://mirrors.catalyst.net.nz/pub/mozilla.git//refs/tags/JS%5F1%5F7%5FALPHA%5FBASE\nfor tags/JS_1_7_ALPHA_BASE\nCouldn't resolve host 'mirrors.catalyst.net.nz'\nerror: Could not interpret tags/JS_1_7_ALPHA_BASE as something to pull\n[jonsmirl@jonsmirl mozgit]$ cg update\nThere is no GIT repository here (.git not found)\n[jonsmirl@jonsmirl mozgit]$ ls -a\n.  ..\n[jonsmirl@jonsmirl mozgit]$\n\n\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"21552","messageId":"20060610191552.7d5a44d9.tihirvon@gmail.com","threadId":"4442","inReplyTo":"9e4733910606100844v5f4765d8o85c9a6f239faed43@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Timo Hirvonen","fromEmail":"tihirvon@gmail.com","sentAt":"2006-06-10T16:15:52Z","receivedAt":"2006-06-10T16:15:52Z","isPatch":false,"sender":{"key":"tihirvon@gmail.com","avatar":null},"body":"\"Jon Smirl\" <jonsmirl@gmail.com> wrote:\n\n> On 6/10/06, Junio C Hamano <junkio@cox.net> wrote:\n> > \"Jon Smirl\" <jonsmirl@gmail.com> writes:\n> >\n> > > Here's a new transport problem. When using git-clone to fetch Martin's\n> > > tree it kept failing for me at dreamhost. I had a parallel fetch\n> > > running on my local machine which has a much slower net connection. It\n> > > finally finished and I am watching the end phase where it prints all\n> > > of the 'walk' messages. The git-http-fetch process has jumped up to\n> > > 800MB in size after being 2MB during the download. dreamhost has a\n> > > 500MB process size limit so that is why my fetches kept failing there.\n> >\n> > The http-fetch process uses by mmaping the downloaded pack, and\n> > if I recall correctly we are talking about 600MB pack, so 500MB\n> > limit sounds impossible, perhaps?\n> \n> The fetch on my local machine failed too. It left nothing behind, now\n> I have to download the 680MB again.\n\nThat's sad.  Could git-clone be changed to not remove .git directory if\nfetching objects fails (after other files in the .git directory have\nbeen fetched)?  You could then hopefully continue with git-pull.\n\n-- \nhttp://onion.dynserv.net/~timo/\n"},{"id":"21553","messageId":"Pine.LNX.4.64.0606101041490.5498@g5.osdl.org","threadId":"4442","inReplyTo":"448A847C.20105@dawes.za.net","subject":"Re: Figured out how to get Mozilla into git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-10T17:53:09Z","receivedAt":"2006-06-10T17:53:09Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 10 Jun 2006, Rogan Dawes wrote:\n>\n> Here's an idea. How about separating trees and commits from the actual blobs\n> (e.g. in separate packs)? My reasoning is that the commits and trees should\n> only be a small portion of the overall repository size, and should not be that\n> expensive to transfer. (Of course, this is only a guess, and needs some\n> numbers to back it up.)\n\nThe trees in particular are actually a pretty big part of the history. \n\nMore importantly, the blobs compress horribly badly in the absense of \nhistory - a _lot_ of the compression in git packing comes very much from \nthe fact that we do a good job at delta-compression.\n\nSo if you get all of the commit/tree history, but none of the blob \nhistory, you're actually not going to win that much space. As already \ndiscussed, the _whole_ history packed with git is usually not insanely \nbigger than just the whole unpacked tree (with no history at all).\n\nSo you'd think that getting just the top version of the tree would be a \nmuch bigger space-saving that it actually is. If you _also_ get all the \ntree and commit objects, the space saving is even less.\n\nI actually suspect that the most realistic way to handle this is to use \nthe \"fetch.c\" logic (ie the incremental fetcher used by http), and add \nsome mode to the git daemon where you fetch literally one object at a time \n(ie this would be totally _separate_ from the pack-file thing: you'd not \nask for \"git-upload-pack\", you'd ask for something like \n\"git-serve-objects\" instead).\n\nThe fetch.c logic really does allow for on-demand object fetching, and is \nthus much more suitable for incomplete repositories.\n\nHOWEVER. The fetch.c logic - by necessity - works on a object-by-object \nlevel. That means that you'd get no delta compression AT ALL, and I \nsuspect that the downside of that would be a factor of ten expansion or \nmore, which means that it would really not work that well in practice.\n\nIt might be worth testing, though. It would work fine for the \"after I \nhave the initial cauterized tree, fetch small incremental updates\" case. \nThe operative word here being \"small\" and \"incremental\", because I'm \npretty sure it really would suck for the case of a big fetch.\n\nBut it would be _simple_, which is why it's worth trying out. It also has \nthe advantage that it would solve the \"I had data corruption on my disk, \nand lost 100 objects, but all the the rest is fine\" issue. Again, that's \nnot something that the efficient packing protocol handles, exactly because \nit assumes full history, and uses that to do all its optimizations.\n\n\t\tLinus\n"},{"id":"21554","messageId":"9e4733910606101102k2a860cf3jd767331e6b5dcf10@mail.gmail.com","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606101041490.5498@g5.osdl.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-06-10T18:02:13Z","receivedAt":"2006-06-10T18:02:13Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"Here's a random idea, how about a tool that turns a real pack into one\nthat is segmented and then faults in segments if you do an operation\nthat needs the old segments? The full pack would always look like it\nis there even if it isn't. Something like gitk would be modified not\nto fault in the missing segments.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"21555","messageId":"448B1130.8020005@dawes.za.net","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606101041490.5498@g5.osdl.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Rogan Dawes","fromEmail":"lists@dawes.za.net","sentAt":"2006-06-10T18:36:32Z","receivedAt":"2006-06-10T18:36:32Z","isPatch":false,"sender":{"key":"lists@dawes.za.net","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Sat, 10 Jun 2006, Rogan Dawes wrote:\n>> Here's an idea. How about separating trees and commits from the actual blobs\n>> (e.g. in separate packs)? My reasoning is that the commits and trees should\n>> only be a small portion of the overall repository size, and should not be that\n>> expensive to transfer. (Of course, this is only a guess, and needs some\n>> numbers to back it up.)\n> \n> The trees in particular are actually a pretty big part of the history. \n> \n> More importantly, the blobs compress horribly badly in the absense of \n> history - a _lot_ of the compression in git packing comes very much from \n> the fact that we do a good job at delta-compression.\n> \n> So if you get all of the commit/tree history, but none of the blob \n> history, you're actually not going to win that much space. As already \n> discussed, the _whole_ history packed with git is usually not insanely \n> bigger than just the whole unpacked tree (with no history at all).\n> \n> So you'd think that getting just the top version of the tree would be a \n> much bigger space-saving that it actually is. If you _also_ get all the \n> tree and commit objects, the space saving is even less.\n> \n\nOne possibility, given that the full commit and tree history is so\nlarge, is simply to get the HEAD commit and the trees that the commit\ndepends directly on, rather than fetching them all up front.\n\n> I actually suspect that the most realistic way to handle this is to use \n> the \"fetch.c\" logic (ie the incremental fetcher used by http), and add \n> some mode to the git daemon where you fetch literally one object at a time \n> (ie this would be totally _separate_ from the pack-file thing: you'd not \n> ask for \"git-upload-pack\", you'd ask for something like \n> \"git-serve-objects\" instead).\n> \n> The fetch.c logic really does allow for on-demand object fetching, and is \n> thus much more suitable for incomplete repositories.\n> \n> HOWEVER. The fetch.c logic - by necessity - works on a object-by-object \n> level. That means that you'd get no delta compression AT ALL, and I \n> suspect that the downside of that would be a factor of ten expansion or \n> more, which means that it would really not work that well in practice.\n\nWould it be possible to add a mode where fetch.c is given a list of \ndesired objects, and returns a list of pointers to those objects? Then \ncallers that already have such a list could be modified to pass the \nwhole list at once, allowing at least SOME compression, and optimisation \nof round trips, etc? There would be a tradeoff in memory use, though, I \nguess.\n\nRogan\n"},{"id":"21556","messageId":"20060610183724.GE2609@pasky.or.cz","threadId":"4442","inReplyTo":"9e4733910606100844v5f4765d8o85c9a6f239faed43@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2006-06-10T18:37:24Z","receivedAt":"2006-06-10T18:37:24Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Sat, Jun 10, 2006 at 05:44:58PM CEST, I got a letter\nwhere Jon Smirl <jonsmirl@gmail.com> said that...\n> The fetch on my local machine failed too. It left nothing behind, now\n> I have to download the 680MB again.\n> \n> walk 1f19465388a4ef7aff7527a13f16122a809487d4\n> walk c3ca840256e3767d08c649f8d2761a1a887351ab\n> walk 7a74e42699320c02b814b88beadb1ae65009e745\n> error: Couldn't get\n> http://mirrors.catalyst.net.nz/pub/mozilla.git//refs/tags/JS%5F1%5F7%5FALPHA%5FBASE\n> for tags/JS_1_7_ALPHA_BASE\n> Couldn't resolve host 'mirrors.catalyst.net.nz'\n> error: Could not interpret tags/JS_1_7_ALPHA_BASE as something to pull\n> [jonsmirl@jonsmirl mozgit]$ cg update\n> There is no GIT repository here (.git not found)\n> [jonsmirl@jonsmirl mozgit]$ ls -a\n> .  ..\n> [jonsmirl@jonsmirl mozgit]$\n\n  You could try with cg-clone, which won't delete the repository if\nthings fail. It will clone only the master branch, though.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nA person is just about as big as the things that make them angry.\n"},{"id":"21557","messageId":"20060610185535.GB19919@mail.Lars-Johannsen.dk","threadId":"4442","inReplyTo":"9e4733910606100844v5f4765d8o85c9a6f239faed43@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Lars Johannsen","fromEmail":"mail@lars-johannsen.dk","sentAt":"2006-06-10T18:55:35Z","receivedAt":"2006-06-10T18:55:35Z","isPatch":false,"sender":{"key":"mail@lars-johannsen.dk","avatar":null},"body":"On (10/06/06 11:44), Jon Smirl wrote:\n> Date:\tSat, 10 Jun 2006 11:44:58 -0400\n> From:\t\"Jon Smirl\" <jonsmirl@gmail.com>\n> To:\t\"Junio C Hamano\" <junkio@cox.net>\n> Subject: Re: Figured out how to get Mozilla into git\n> Cc:\tgit@vger.kernel.org\n> \n> On 6/10/06, Junio C Hamano <junkio@cox.net> wrote:\n> >\"Jon Smirl\" <jonsmirl@gmail.com> writes:\n> >\n> >> Here's a new transport problem. When using git-clone to fetch Martin's\n> >> tree it kept failing for me at dreamhost. I had a parallel fetch\n> >> running on my local machine which has a much slower net connection. It\n> >> finally finished and I am watching the end phase where it prints all\n> >> of the 'walk' messages. The git-http-fetch process has jumped up to\n> >> 800MB in size after being 2MB during the download. dreamhost has a\n> >> 500MB process size limit so that is why my fetches kept failing there.\n> >\n> >The http-fetch process uses by mmaping the downloaded pack, and\n> >if I recall correctly we are talking about 600MB pack, so 500MB\n> >limit sounds impossible, perhaps?\n> \n> The fetch on my local machine failed too. It left nothing behind, now\n> I have to download the 680MB again.\n> \n> walk 1f19465388a4ef7aff7527a13f16122a809487d4\n> walk c3ca840256e3767d08c649f8d2761a1a887351ab\n> walk 7a74e42699320c02b814b88beadb1ae65009e745\n> error: Couldn't get\n> http://mirrors.catalyst.net.nz/pub/mozilla.git//refs/tags/JS%5F1%5F7%5FALPHA%5FBASE\n> for tags/JS_1_7_ALPHA_BASE\n> Couldn't resolve host 'mirrors.catalyst.net.nz'\n> error: Could not interpret tags/JS_1_7_ALPHA_BASE as something to pull\n> [jonsmirl@jonsmirl mozgit]$ cg update\n> There is no GIT repository here (.git not found)\n> [jonsmirl@jonsmirl mozgit]$ ls -a\n> .  ..\n> [jonsmirl@jonsmirl mozgit]$\n\nTo prevent repeat (on this repo) your could grab it with a browser:\n-mkdir tmp; cd tmp; git init-db;\n-copy  mirror../pu/mozilla.git/objects/*  to .git/objects/\n-copy   --||---.git/info/refs to refsinfo in tmp-dir\ngawk '{if  ($2 !~ /\\^\\{\\}$/) print $1 > sprintf(\".git/%s\",$2);}' refsinfo\n to extract branches and tags into ./git/refs/{heads,tags}\nstart playing (after a backup) with git-fsck-objects, git-checkout etc.\n \n-- \nLars Johannsen \nmail@Lars-johannsen.dk\n"},{"id":"21611","messageId":"Pine.LNX.4.64.0606111747110.2703@localhost.localdomain","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606091825080.5498@g5.osdl.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2006-06-11T22:00:13Z","receivedAt":"2006-06-11T22:00:13Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 9 Jun 2006, Linus Torvalds wrote:\n\n> \n> \n> On Sat, 10 Jun 2006, Martin Langhoff wrote:\n> > \n> > Now I don't know how much memory or time this took, but it clearly\n> > completed ok. And, it's now a single pack, weighting a grand total of\n> > 617MB\n> \n> Ok, that's more than reasonable. That should be fairly easily mapped on a \n> 32-bit architecture without any huge problems, even with some VM \n> fragmentation going on. It might be borderline (and you definitely want a \n> 3:1 VM user:kernel split), but considering that the original CVS archive \n> was apparently 3GB, having a single 617M pack-file is still pretty damn \n> good.  That's like 20% of the original, with all the obvious distribution \n> advantages.\n\nI played a bit with git-repack on that repo.  the git-pack-objects \nmemory usage grew to around 760MB (git-rev-list was less than that).  So \nLRU of partial pack mappings might bring that down significantly.\n\nThen I used git-repack -a -f --window=20 --depth=20 which produced a \nnice 468MB pack file along with the invariant 45MB index file for a \ngrand total of 535MB for the whole repo (the .git/refs/ directory alone \nstill occupies 17MB on disk).\n\nSo it is probably worth having deeper delta chains for large historic \nrepositories as the deep revisions are unlikely to be referenced that \noften while the saving is quite significant.\n\n\nNicolas\n"},{"id":"22036","messageId":"Pine.LNX.4.64.0606181223580.5498@g5.osdl.org","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606111747110.2703@localhost.localdomain","subject":"Re: Figured out how to get Mozilla into git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-18T19:26:24Z","receivedAt":"2006-06-18T19:26:24Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 11 Jun 2006, Nicolas Pitre wrote:\n> \n> Then I used git-repack -a -f --window=20 --depth=20 which produced a \n> nice 468MB pack file along with the invariant 45MB index file for a \n> grand total of 535MB for the whole repo (the .git/refs/ directory alone \n> still occupies 17MB on disk).\n\nBtw, can others with that mozilla repo confirm that a mozilla repository \nthat has been repacked seems to be entirely fine, but git-fsck-objects \n(with \"--full\", of course) will report\n\n\terror: Packfile .git/objects/pack/pack-06389c21fc3c4312cbc9a4ddde087c907c1a840b.pack SHA1 mismatch with itself\n\nfor me (the fsck then completes with no other errors what-so-ever, so the \ncontents are actually fine).\n\nOr is it just me?\n\n\t\tLinus\n"},{"id":"22042","messageId":"46a038f90606181440q4fd03bebl9495ace131eb958@mail.gmail.com","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606181223580.5498@g5.osdl.org","subject":"Re: Figured out how to get Mozilla into git","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-06-18T21:40:44Z","receivedAt":"2006-06-18T21:40:44Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 6/19/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> Or is it just me?\n\nNo problems here with my latest import run. fsck-objects --full comes\nclean, takes 14m:\n\n/usr/bin/time git-fsck-objects --full\n737.22user 38.79system 14:09.40elapsed 91%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (20807major+19483471minor)pagefaults 0swaps\n\nBTW, that import (with the latest code Junio has) took 37hs even with\nthe aggressive repack -a -d. I want to bench it dropping the -a from\nthe recurrring repack, and doing a final repack -a -d.\n\ncheers,\n\n\nmartin\n"},{"id":"22047","messageId":"Pine.LNX.4.64.0606181532130.5498@g5.osdl.org","threadId":"4442","inReplyTo":"46a038f90606181440q4fd03bebl9495ace131eb958@mail.gmail.com","subject":"Re: Figured out how to get Mozilla into git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-18T22:36:52Z","receivedAt":"2006-06-18T22:36:52Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 19 Jun 2006, Martin Langhoff wrote:\n> \n> No problems here with my latest import run. fsck-objects --full comes\n> clean, takes 14m:\n>\n> /usr/bin/time git-fsck-objects --full\n> 737.22user 38.79system 14:09.40elapsed 91%CPU (0avgtext+0avgdata 0maxresident)k\n> 0inputs+0outputs (20807major+19483471minor)pagefaults 0swaps\n\nIt takes much less than that for me: \n\n\t408.40user 32.56system 7:22.07elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n\t0inputs+0outputs (145major+13455672minor)pagefaults 0swaps\n\nand in particular note the much lower minor pagefaults number (which is a \nvery good approximation of total RSS). Mine is with all the memory \noptimizations in place, but I didn't see _that_ big of a difference, so \nthere's something else in addition.\n\nHowever, the fact that I get \"SHA1 mismatch with itself\" is strange. The \nre-pack will always re-generate the SHA1, so I worry that this is perhaps \nsome PPC-specific bug in SHA1 handling (and it's entirely possible that \nit's triggered by doing a SHA1 over a 500+MB area).\n\nThe fact that you don't see it is indicative that it's somehow specific to \nmy setup.\n\n> BTW, that import (with the latest code Junio has) took 37hs even with\n> the aggressive repack -a -d. I want to bench it dropping the -a from\n> the recurrring repack, and doing a final repack -a -d.\n\nYeah, that's probably the right thing to do. The \"-a\" is ok with tons of \nmemory, and I'm trying to make it ok with _less_ memory, but it's probably \njust not worth it.\n\n\t\tLinus\n"},{"id":"22048","messageId":"Pine.LNX.4.64.0606181543270.5498@g5.osdl.org","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606181532130.5498@g5.osdl.org","subject":"Broken PPC sha1.. (Re: Figured out how to get Mozilla into git)","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-18T22:51:02Z","receivedAt":"2006-06-18T22:51:02Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\n\nOn Sun, 18 Jun 2006, Linus Torvalds wrote:\n> \n> On Mon, 19 Jun 2006, Martin Langhoff wrote:\n> > \n> > No problems here with my latest import run. fsck-objects --full comes\n> > clean, takes 14m:\n> >\n> > /usr/bin/time git-fsck-objects --full\n> > 737.22user 38.79system 14:09.40elapsed 91%CPU (0avgtext+0avgdata 0maxresident)k\n> > 0inputs+0outputs (20807major+19483471minor)pagefaults 0swaps\n> \n> It takes much less than that for me: \n> \n> \t408.40user 32.56system 7:22.07elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n> \t0inputs+0outputs (145major+13455672minor)pagefaults 0swaps\n\nOk, re-building the thing with MOZILLA_SHA1=1 rather than my default \nPPC_SHA1=1 fixes the problem. I no longer get that \"SHA1 mismatch with \nitself\" on the pack-file.\n\nSadly, it also takes a _lot_ longer to fsck.\n\nPaul - I think the ppc SHA1_Update() overflows in 32 bits, when the length \nof the memory area to be checksummed is huge.\n\nIn particular, the pack-file is 535MB in size, and the way we check the \nSHA1 checksum is by just mapping it all, doing a single SHA1_Update() over \nthe whole pack-file, and comparing the end result with the internal SHA1 \nat the end of the pack-file.\n\nThe PPC SHA1_Update() function starts off with:\n\n\tint SHA1_Update(SHA_CTX *c, const void *ptr, unsigned long n)\n\t{\n\t...\n\t\tc->len += n << 3;\n\nwhich will obviously overflow if \"n\" is bigger than 29 bits, ie 512MB.\n\nSo doing the length in bits (or whatever that \"<<3\" is there for) doesn't \nseem to be such a great idea.\n\nI guess we could make the caller just always chunk it up, but wouldn't it \nbe nice to fix the PPC SHA1 implementation instead?\n\nThat said, the _only_ thing this will ever trigger on in practice is \nexactly this one case: a large packfile whose checksum was _correctly_ \ngenerated - because pack-file generation does it in IO chunks using the \ncsum-file interfaces - but that will be incorrectly checked because we \ncheck it all at once.\n\nSo as bugs go, it's a fairly benign one.\n\n\t\t\tLinus\n"},{"id":"22050","messageId":"17557.57564.267469.561683@cargo.ozlabs.ibm.com","threadId":"4442","inReplyTo":"Pine.LNX.4.64.0606181543270.5498@g5.osdl.org","subject":"[PATCH] Fix PPC SHA1 routine for large input buffers","fromName":"Paul Mackerras","fromEmail":"paulus@samba.org","sentAt":"2006-06-18T23:25:16Z","receivedAt":"2006-06-18T23:25:16Z","isPatch":true,"sender":{"key":"paulus@samba.org","avatar":"https://avatars.githubusercontent.com/u/1606439?v=4"},"body":"The PPC SHA1 routine had an overflow which meant that it gave\nincorrect results for input buffers >= 512MB.  This fixes it by\nensuring that the update of the total length in bits is done using\n64-bit arithmetic.\n\nSigned-off-by: Paul Mackerras <paulus@samba.org>\n---\nLinus Torvalds writes:\n\n> Paul - I think the ppc SHA1_Update() overflows in 32 bits, when the length \n> of the memory area to be checksummed is huge.\n\nYep.  I checked the assembly output of this, and it looks right, but I\nhaven't actually tested it by running it...\n\nPaul.\n\ndiff --git a/ppc/sha1.c b/ppc/sha1.c\nindex 5ba4fc5..0820398 100644\n--- a/ppc/sha1.c\n+++ b/ppc/sha1.c\n@@ -30,7 +30,7 @@ int SHA1_Update(SHA_CTX *c, const void *\n \tunsigned long nb;\n \tconst unsigned char *p = ptr;\n \n-\tc->len += n << 3;\n+\tc->len += (uint64_t) n << 3;\n \twhile (n != 0) {\n \t\tif (c->cnt || n < 64) {\n \t\t\tnb = 64 - c->cnt;\n"},{"id":"22052","messageId":"Pine.LNX.4.64.0606182145370.5498@g5.osdl.org","threadId":"4442","inReplyTo":"17557.57564.267469.561683@cargo.ozlabs.ibm.com","subject":"Re: [PATCH] Fix PPC SHA1 routine for large input buffers","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-19T05:02:55Z","receivedAt":"2006-06-19T05:02:55Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 19 Jun 2006, Paul Mackerras wrote:\n>\n> The PPC SHA1 routine had an overflow which meant that it gave\n> incorrect results for input buffers >= 512MB.  This fixes it by\n> ensuring that the update of the total length in bits is done using\n> 64-bit arithmetic.\n> \n> Signed-off-by: Paul Mackerras <paulus@samba.org>\n\nAcked-by: Linus Torvalds <torvalds@osdl.org>\n\nThis fixes git-fsck-objects for me on the mozilla archive, no more \ncomplaints about bad SHA1's.\n\nAnd yeah, now it's taking me 14 minutes too, so the 7-minute fsck was just \nbecause it didn't actually check the SHA1 of the large pack fully.\n\n(Which is actually good news - half of the time is literally checking the \npack integrity. That implies that the individual object integrity isn't as \ndominating as I thought it would be, and that things like hw-accelerated \nSHA1 engines will help with fsck. I'd not be surprised to see things like \nthat in a couple of years).\n\n\t\tLinus\n"}]}