{"thread":{"id":"5114","subject":"Creating objects manually and repack","startedAt":"2006-08-04T03:43:42Z","lastAt":"2006-08-05T05:52:03Z","messageCount":36,"participants":["Jon Smirl","Jeff King","Linus Torvalds","A Large Angry SCM","Rogan Dawes","Junio C Hamano","Carl Worth","Jakub Narebski","Martin Langhoff","Shawn Pearce"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"24680","messageId":"9e4733910608032043u689f431rc5408c6d89398142@mail.gmail.com","threadId":"5114","inReplyTo":null,"subject":"Creating objects manually and repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-08-04T03:43:42Z","receivedAt":"2006-08-04T03:43:42Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"I've made 500K object files with my cvs2svn front end. This is 500K of\nrevision files and no tree files. Now I run get-repack. It says done\ncounting zero objects. What needs to be update so that repack will\nfind all of my objects?\n\ngit-fsck isn't happy either since I have no HEAD.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"24682","messageId":"20060804035803.GA362@coredump.intra.peff.net","threadId":"5114","inReplyTo":"9e4733910608032043u689f431rc5408c6d89398142@mail.gmail.com","subject":"Re: Creating objects manually and repack","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2006-08-04T03:58:03Z","receivedAt":"2006-08-04T03:58:03Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Aug 03, 2006 at 11:43:42PM -0400, Jon Smirl wrote:\n\n> I've made 500K object files with my cvs2svn front end. This is 500K of\n> revision files and no tree files. Now I run get-repack. It says done\n> counting zero objects. What needs to be update so that repack will\n> find all of my objects?\n\ngit-repack starts at your heads and works its way down. You can either:\n  - make a dummy commit for a tree with all of your blobs:\n    $ while read sha1; do\n        echo -e \"100644 blob $sha1\\t$sha1\"\n      done <list_of_sha1s | git-update-index --index-info\n      tree=$(git-write-tree)\n      commit=$(git-commit-tree $tree)\n      git-update-ref HEAD $commit\n\n  - call git-pack-objects directly with a list of objects\n      git-pack-objects .git/objects/pack/pack <list_of_sha1s\n\nObviously the latter is simpler, but the former will also make\ngit-fsck-objects happy. Note that they're both untested, so there might\nbe typos.\n\n-Peff\n"},{"id":"24683","messageId":"Pine.LNX.4.64.0608032052210.4168@g5.osdl.org","threadId":"5114","inReplyTo":"9e4733910608032043u689f431rc5408c6d89398142@mail.gmail.com","subject":"Re: Creating objects manually and repack","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-08-04T04:01:23Z","receivedAt":"2006-08-04T04:01:23Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 3 Aug 2006, Jon Smirl wrote:\n>\n> I've made 500K object files with my cvs2svn front end. This is 500K of\n> revision files and no tree files. Now I run get-repack. It says done\n> counting zero objects. What needs to be update so that repack will\n> find all of my objects?\n\nJust enumerate them by hand, and pass the list off to git-pack-objects.\n\nIOW, you can _literally_ do something like this\n\n\t(cd .git/objects ; find . -type f -name '[0-9a-f]*' | tr -d '\\./') |\n\t\tgit-pack-objects tmp-pack\n\nand it will generate a pack-file and index (called \"tmp-pack-*.pack\" and \n\"tmp-pack-*.idx\" respectively) that contains all your lose objects.\n\nNow, that said, pack-file will generally _suck_ if you actually do it like \nthe above. You actually want to pass in the object names _together_ with \nthe filenames they were generated from, so that git-pack-objects can use \nits heuristics for finding good delta candidates.\n\nSo what you actually want to do is pass in a set of object names with the \nname of the file they came with (space in between). See for example\n\n\tgit-rev-list --objects HEAD^..\n\noutput for how something like that might look (git-pack-objects is \ndesigned to take the \"git-rev-list --objects\" output as its input).\n\n> git-fsck isn't happy either since I have no HEAD.\n\nYeah, you cannot (and mustn't) run anything like git-fsck-objects or \"git \nprune\" until you've connected them all up somehow.\n\n\t\tLinus\n"},{"id":"24684","messageId":"9e4733910608032124o5b5b69b5hda2eb8cb1e0ac959@mail.gmail.com","threadId":"5114","inReplyTo":"Pine.LNX.4.64.0608032052210.4168@g5.osdl.org","subject":"Re: Creating objects manually and repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-08-04T04:24:26Z","receivedAt":"2006-08-04T04:24:26Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"I am converting all of the revisions from each CVS file into git\nobjects the first time the file is parsed. The plan was to run repack\nafter each file is finished. That way it should be easy to figure out\nthe deltas since everything will be a variation on the same file.\n\nSo what's the best way to pack these objects, append them to the\nexisting pack and then clean everything up for the next file? I am\nparsing 120K CVS files containing over 1M revs.\n\nAfter I get all of the objects written and packed later code is going\nto write out the trees.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"24687","messageId":"Pine.LNX.4.64.0608032138330.4168@g5.osdl.org","threadId":"5114","inReplyTo":"9e4733910608032124o5b5b69b5hda2eb8cb1e0ac959@mail.gmail.com","subject":"Re: Creating objects manually and repack","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-08-04T04:46:58Z","receivedAt":"2006-08-04T04:46:58Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 4 Aug 2006, Jon Smirl wrote:\n>\n> I am converting all of the revisions from each CVS file into git\n> objects the first time the file is parsed. The plan was to run repack\n> after each file is finished. That way it should be easy to figure out\n> the deltas since everything will be a variation on the same file.\n\nSure. In that case, just list the object ID's in the exact same order you \ncreated them.\n\nBasically,as you create them, just keep a list of all ID's you've created, \nand every (say) 50,000 objects, just do a\n\n\techo all objects you've created | git-pack-objects new-pack\n\nand then move the new pack into place, and remove all the loose objects \n(don't even bother using \"git prune\" - just basically do something like\n\"rm -rf .git/objects/??\" to get rid of them).\n\n> So what's the best way to pack these objects, append them to the\n> existing pack and then clean everything up for the next file? I am\n> parsing 120K CVS files containing over 1M revs.\n\nYou'll want to repack every once in a while just to not ever have _tons_ \nof those loose objects around, but if you do it every 50,000 objects, \nyou'll have just twenty nice pack-files once you're done, containing all \none million objects, and you'll never have had more than ~200 files in any \nof the loose object subdirectories.\n\nOf course, you might want to make that \"every 50,000 object\" thing \ntunable, so that if you don't have a lot of memory for caching, you might \nwant to do it a bit more often just to make each repack go faster and not \nhave tons of IO. \n\nYou can then do a _full_ repack to get one big object, by just listing \nevery object you ever created (in creation order) to git-pack-objects, and \nthen you can replace all the twenty (smaller) pack-files with the \nresulting single bigger one.\n\nIn fact, at that point you no longer even need to worry about \"creation \norder\", since you've basically created all the deltas in the first phase, \nand regardless of ordering, when you then repack everything at the end, it \nwill re-use all earlier delta information.\n\n\t\tLinus\n"},{"id":"24688","messageId":"Pine.LNX.4.64.0608032150510.4168@g5.osdl.org","threadId":"5114","inReplyTo":"Pine.LNX.4.64.0608032138330.4168@g5.osdl.org","subject":"Re: Creating objects manually and repack","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-08-04T05:01:13Z","receivedAt":"2006-08-04T05:01:13Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 3 Aug 2006, Linus Torvalds wrote:\n> \n> Sure. In that case, just list the object ID's in the exact same order you \n> created them.\n\nBtw, you still want to give a filename for each object you've created, so \nthat the delta sorter does the right thing for the packing. It doesn't \nhave to be a _real_ filename - just make sure that each revision that \ncomes from the same file has a filename that matches all the other \nrevisions from that file.\n\nWhat the filename actually _is_ doesn't much matter, and it doesn't have \nto be the \"real\" filename that was associated with that set of revisions, \nsince we'll just end up hashing it anyway. So it could be some \"SVN inode \nnumber\" for that set of revisions or something, for all git-pack-object \ncares.\n\nSo you could just go through each SVN file in whatever the SVN database is \n(I don't know how SVN organizes it), generate every revisions for that \nfile, and pass in the SVN _database_ filename, rather than necessarily the \nfilename that that revision is actually associated with when checked out.\n\nSo for example, if SVN were to use the same kind of \"Attic/filename,v\" \nformat that CVS uses, there's no reason to worry what the real filename \nwas in any particular checked out tree, you could just pass \ngit-pack-objects a series of lines in the form of\n\n\t..\n\t<sha1-object-name-of-rev1> Attic/filename,v\n\t<sha1-object-name-of-rev2> Attic/filename,v\n\t<sha1-object-name-of-rev3> Attic/filename,v\n\t..\n\nas input on its stdin, and it will create a pack-file of all the objects \nyou name, and use the \"Attic/filename,v\" info as the deltifier hint to \nknow to do all the deltas of those revs against each other rather than \nagainst random other objects.\n\nThe fact that the file was actually checked out as \"src/filename\" (and, \nsince SVN supports renaming, it might have been checked out under any \nnumber of _other_ names over the history of the project) doesn't matter, \nand you don't need to even try to figure that out. git-pack-objects \nwouldn't care anyway.\n\n\t\t\tLinus\n"},{"id":"24689","messageId":"9e4733910608032211r3571f6adje52968ffcb689457@mail.gmail.com","threadId":"5114","inReplyTo":"Pine.LNX.4.64.0608032150510.4168@g5.osdl.org","subject":"Re: Creating objects manually and repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-08-04T05:11:56Z","receivedAt":"2006-08-04T05:11:56Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 8/4/06, Linus Torvalds <torvalds@osdl.org> wrote:\n>\n>\n> On Thu, 3 Aug 2006, Linus Torvalds wrote:\n> >\n> > Sure. In that case, just list the object ID's in the exact same order you\n> > created them.\n>\n> Btw, you still want to give a filename for each object you've created, so\n\nI'll add a file name hint.\n\nI'm converting the cvs2svn tool to do cvs2git.\n\nMartin has a copy of it up under git. I haven't checked in any of my\nchanges yet.\nhttp://git.catalyst.net.nz/gitweb?p=cvs2svn.git;a=summary\n\nIf you read the log it is obvious that these guys have done major work\nto deal with all kinds of broken CVS repositories. I want to piggyback\non that work and reuse their code that builds change sets.  So far\nthis is the only tool I have found that can import the Mozilla CVS\nwithout errors. Only problem is that it imports it to SVN instead of\ngit. I'm fixing that and learning Python at the same time.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"24704","messageId":"9e4733910608040740x23a8b0cs3bc276ef9e6fb8f7@mail.gmail.com","threadId":"5114","inReplyTo":"Pine.LNX.4.64.0608032150510.4168@g5.osdl.org","subject":"Re: Creating objects manually and repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-08-04T14:40:50Z","receivedAt":"2006-08-04T14:40:50Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"One thing is obvious, I need to tune the repacks to happen before\nthings spill out of the cache.  git repack-objects has been chugging\naway for 2hrs now at 2% CPU and 3000 io/sec. It is in one of those\nmodes where it went back to get the early stuff and in the process of\ngetting that it knocked the later stuff out of the cache basically\nrendering the cache useless.\n\nI'm making good progress with this. I have hit two bugs in cvs2svn\nthat I will need to get fixed. cvs2svn is claiming two of the ,v files\nto be invalid but to my eyes they look ok.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"24705","messageId":"9e4733910608040750g3f72c07ct43f54347e47f25b4@mail.gmail.com","threadId":"5114","inReplyTo":"9e4733910608040740x23a8b0cs3bc276ef9e6fb8f7@mail.gmail.com","subject":"Re: Creating objects manually and repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-08-04T14:50:48Z","receivedAt":"2006-08-04T14:50:48Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"The whole problem with CVS import is avoiding getting IO bound. Since\nMozilla CVS expands into 20GB when the revisions are separated out\ndoing all that IO takes a lot of time. When these imports take four\ndays it is all IO time, not CPU.\n\nCould repack-objects be modified to take the objects on stdin as I\ngenerate them instead of me putting them into the file system and then\ndeleting them? That model would avoid many gigabytes of IO.\n\nIt might work to just stream the output from zlib into repack-objects\nand let it recompute the object name.  Or could I just stream in the\nuncompressed objects? I can still compute the object sha name in my\ncode so that I can find it later.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"24706","messageId":"Pine.LNX.4.64.0608040818270.5167@g5.osdl.org","threadId":"5114","inReplyTo":"9e4733910608040750g3f72c07ct43f54347e47f25b4@mail.gmail.com","subject":"Re: Creating objects manually and repack","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-08-04T15:22:23Z","receivedAt":"2006-08-04T15:22:23Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 4 Aug 2006, Jon Smirl wrote:\n> \n> Could repack-objects be modified to take the objects on stdin as I\n> generate them instead of me putting them into the file system and then\n> deleting them? That model would avoid many gigabytes of IO.\n\nI'd suggest against it, but you can (and should) just repack often enough \nthat you shouldn't ever have gigabytes of objects \"in flight\". I'd have \nexpected that with a repack every few ten thousand files, and most files \nbeing on the order of a few kB, you'd have been more than ok, but \nespecially if you have large files, you may want to make things \"every <n> \nbytes\" rather than \"every <n> files\".\n\nYou _could_ also decide to create packs very aggressively indeed, and if \nyou do them quickly enough, the raw objects never even get written back to \ndisk before you delete them. That will leave you with a lot of packs, but \nyou could then \"repack the packs\" every once in a while.\n\nThat said, it's obviously not _impossible_ to do what you suggest, it's \njust major surgery to pack-objects (which I'm not going to have time to \ndo, since I'll be going on a vacation this weekend).\n\n\t\t\tLinus\n"},{"id":"24707","messageId":"9e4733910608040841v7f4f27efra63e5ead2656e07@mail.gmail.com","threadId":"5114","inReplyTo":"Pine.LNX.4.64.0608040818270.5167@g5.osdl.org","subject":"Re: Creating objects manually and repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-08-04T15:41:58Z","receivedAt":"2006-08-04T15:41:58Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 8/4/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> I'd suggest against it, but you can (and should) just repack often enough\n> that you shouldn't ever have gigabytes of objects \"in flight\". I'd have\n> expected that with a repack every few ten thousand files, and most files\n> being on the order of a few kB, you'd have been more than ok, but\n> especially if you have large files, you may want to make things \"every <n>\n> bytes\" rather than \"every <n> files\".\n\nHow about forking off a pack-objects and handing it one file name at a\ntime over a pipe. When I hand it the next file name I delete the first\nfile. Does pack-objects make multiple passes over the files? This\nmodel would let me hand it all 1M files.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"24714","messageId":"44D36F64.5040404@gmail.com","threadId":"5114","inReplyTo":"9e4733910608040841v7f4f27efra63e5ead2656e07@mail.gmail.com","subject":"Re: Creating objects manually and repack","fromName":"A Large Angry SCM","fromEmail":"gitzilla@gmail.com","sentAt":"2006-08-04T16:01:40Z","receivedAt":"2006-08-04T16:01:40Z","isPatch":false,"sender":{"key":"gitzilla@gmail.com","avatar":"https://gravatar.com/avatar/354625c442439908ff3dd99757dee330e29e9df7847472384faf7a00add247fb?d=mp&s=160"},"body":"Jon Smirl wrote:\n> On 8/4/06, Linus Torvalds <torvalds@osdl.org> wrote:\n>> I'd suggest against it, but you can (and should) just repack often enough\n>> that you shouldn't ever have gigabytes of objects \"in flight\". I'd have\n>> expected that with a repack every few ten thousand files, and most files\n>> being on the order of a few kB, you'd have been more than ok, but\n>> especially if you have large files, you may want to make things \"every \n>> <n>\n>> bytes\" rather than \"every <n> files\".\n> \n> How about forking off a pack-objects and handing it one file name at a\n> time over a pipe. When I hand it the next file name I delete the first\n> file. Does pack-objects make multiple passes over the files? This\n> model would let me hand it all 1M files.\n> \n\nWhy don't you just write the pack file directly? Pack files without \ndeltas have a very simple structure, and git-index-pack will create a \npack index file for the pack file you give it.\n"},{"id":"24715","messageId":"9e4733910608040911p443a1360k6d9d1aab00039100@mail.gmail.com","threadId":"5114","inReplyTo":"44D36F64.5040404@gmail.com","subject":"Re: Creating objects manually and repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-08-04T16:11:35Z","receivedAt":"2006-08-04T16:11:35Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 8/4/06, A Large Angry SCM <gitzilla@gmail.com> wrote:\n> Jon Smirl wrote:\n> > On 8/4/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> >> I'd suggest against it, but you can (and should) just repack often enough\n> >> that you shouldn't ever have gigabytes of objects \"in flight\". I'd have\n> >> expected that with a repack every few ten thousand files, and most files\n> >> being on the order of a few kB, you'd have been more than ok, but\n> >> especially if you have large files, you may want to make things \"every\n> >> <n>\n> >> bytes\" rather than \"every <n> files\".\n> >\n> > How about forking off a pack-objects and handing it one file name at a\n> > time over a pipe. When I hand it the next file name I delete the first\n> > file. Does pack-objects make multiple passes over the files? This\n> > model would let me hand it all 1M files.\n> >\n>\n> Why don't you just write the pack file directly? Pack files without\n> deltas have a very simple structure, and git-index-pack will create a\n> pack index file for the pack file you give it.\n\nThat is under consideration but the undeltafied pack is about 12GB and\nit takes forever (about a day) to deltafy it. I'm not convinced yet\nthat an undeltafied pack is any faster than just having the objects in\nthe directories.\n\nThe same data in a deltafied pack is 700MB. That is a tremendous\ndifference in the amount of IO needed. The strategy has to be to avoid\nIO, nothing I am doing is ever CPU bound.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"24718","messageId":"Pine.LNX.4.64.0608040931000.5167@g5.osdl.org","threadId":"5114","inReplyTo":"9e4733910608040911p443a1360k6d9d1aab00039100@mail.gmail.com","subject":"Re: Creating objects manually and repack","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-08-04T16:32:00Z","receivedAt":"2006-08-04T16:32:00Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 4 Aug 2006, Jon Smirl wrote:\n> \n> That is under consideration but the undeltafied pack is about 12GB and\n> it takes forever (about a day) to deltafy it. I'm not convinced yet\n> that an undeltafied pack is any faster than just having the objects in\n> the directories.\n\nYeah, I think it's worth it deltifying things early, as you seem to get \nall the object info in the right order anyway (ie you do the revisions for \none file in one go).\n\n\t\tLinus\n"},{"id":"24719","messageId":"44D37845.5010009@dawes.za.net","threadId":"5114","inReplyTo":"9e4733910608040841v7f4f27efra63e5ead2656e07@mail.gmail.com","subject":"Re: Creating objects manually and repack","fromName":"Rogan Dawes","fromEmail":"discard@dawes.za.net","sentAt":"2006-08-04T16:39:33Z","receivedAt":"2006-08-04T16:39:33Z","isPatch":false,"sender":{"key":"discard@dawes.za.net","avatar":null},"body":"Jon Smirl wrote:\n> On 8/4/06, Linus Torvalds <torvalds@osdl.org> wrote:\n>> I'd suggest against it, but you can (and should) just repack often enough\n>> that you shouldn't ever have gigabytes of objects \"in flight\". I'd have\n>> expected that with a repack every few ten thousand files, and most files\n>> being on the order of a few kB, you'd have been more than ok, but\n>> especially if you have large files, you may want to make things \"every \n>> <n>\n>> bytes\" rather than \"every <n> files\".\n> \n> How about forking off a pack-objects and handing it one file name at a\n> time over a pipe. When I hand it the next file name I delete the first\n> file. Does pack-objects make multiple passes over the files? This\n> model would let me hand it all 1M files.\n> \n\nI'd imagine that this would not necessarily save you a lot, if you have \nto write it to disk, and then read it back again. Your only chance here \nis if you stay in the buffer, and avoid actually writing to disk at all.\n\nOf course, using a ramdisk/tmpfs for your object directories might be \nenough to save you. Just use a symlink to tmpfs for the objects \ndirectory, and leave the pack files on persistent storage.\n\nThat doesn't answer your question about how many passes pack-objects \ndoes. Nicholas Pitre should be able to answer that.\n\nRogan\n"},{"id":"24720","messageId":"Pine.LNX.4.64.0608040945070.5167@g5.osdl.org","threadId":"5114","inReplyTo":"9e4733910608040841v7f4f27efra63e5ead2656e07@mail.gmail.com","subject":"Re: Creating objects manually and repack","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-08-04T16:53:09Z","receivedAt":"2006-08-04T16:53:09Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 4 Aug 2006, Jon Smirl wrote:\n> \n> How about forking off a pack-objects and handing it one file name at a\n> time over a pipe. When I hand it the next file name I delete the first\n> file. Does pack-objects make multiple passes over the files? This\n> model would let me hand it all 1M files.\n\npack-objects does actually make several (well, two) passes over the \nobjects right now, because it first does all the sorting based on object \nsize/type, and then does the actual deltifying pass. \n\nBut doing things one file-name at a time would certainly be fine. You can \neven do it with git-pack-objects running in parallel, ie you can do a\n\n\tfor_each_filename() {\n\t\tcvs-generate-objects filename | git-pack-objects filename\n\t\trm -rf .git/objects/??/\n\t}\n\nand then \"cvs-generate-objects\" should just make sure that it writes the \ngit object _before_ it actually outputs the object name on stdout.\n\nAnd if you do it this way, you won't even have to pass any filenames, \nsince git-pack-objects will only get objects for the same file, and will \ndo the right thing just sorting them by size.\n\nSo in the above kind of setting, the _only_ thing that \ncvs-generate-objects needs to do is:\n\n\tfor_each_rev(file) {\n\t\tunsigned char sha1[20];\n\t\tunsigned long len;\n\t\tvoid *buf;\n\n\t\t/* unpack the revision into memory */\n\t\tbuf = cvs_unpack_revision(&len);\n\n\t\t/* Write it out as a git blob file */\n\t\twrite_sha1_file(buf, len, \"blob\", sha1);\n\n\t\t/* Free the memory image */\n\t\tfree(buf);\n\n\t\t/* Tell git-pack-objects the name of the git blob */\n\t\tprintf(\"%s\\n\", sha1_to_hex(sha1));\n\t}\n\nand you're basically all done. The above would turn each *,v file into a \n*-<sha>.pack/*-<sha>.idx file pair, so you'd have exactly as many \npack-files as you have *,v files.\n\n\t\tLinus\n"},{"id":"24721","messageId":"9e4733910608040953p171e4a62p9670f614b33f93b2@mail.gmail.com","threadId":"5114","inReplyTo":"44D37845.5010009@dawes.za.net","subject":"Re: Creating objects manually and repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-08-04T16:53:28Z","receivedAt":"2006-08-04T16:53:28Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 8/4/06, Rogan Dawes <discard@dawes.za.net> wrote:\n> Jon Smirl wrote:\n> > On 8/4/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> >> I'd suggest against it, but you can (and should) just repack often enough\n> >> that you shouldn't ever have gigabytes of objects \"in flight\". I'd have\n> >> expected that with a repack every few ten thousand files, and most files\n> >> being on the order of a few kB, you'd have been more than ok, but\n> >> especially if you have large files, you may want to make things \"every\n> >> <n>\n> >> bytes\" rather than \"every <n> files\".\n> >\n> > How about forking off a pack-objects and handing it one file name at a\n> > time over a pipe. When I hand it the next file name I delete the first\n> > file. Does pack-objects make multiple passes over the files? This\n> > model would let me hand it all 1M files.\n> >\n>\n> I'd imagine that this would not necessarily save you a lot, if you have\n> to write it to disk, and then read it back again. Your only chance here\n> is if you stay in the buffer, and avoid actually writing to disk at all.\n\nIf I keep creating files, reading them and then deleting them then it\nis likely that the same blocks are being used over and over. Since the\nblocks are reused it will stop the cache thrashing. Some disk writes\nwill still happen but that is way better than doing 12GB of unique\nwrites followed by 12GB of reads. The 24GB of IO is all reads on small\nfiles so it is seek time limited since repack does writes in the\nmiddle of the reads.\n\n> Of course, using a ramdisk/tmpfs for your object directories might be\n> enough to save you. Just use a symlink to tmpfs for the objects\n> directory, and leave the pack files on persistent storage.\n\nThe unpacked set of objects is way to big to fit into RAM. Any scheme\nusing the unpacked objects will spill to disk.\n\n> That doesn't answer your question about how many passes pack-objects\n> does. Nicholas Pitre should be able to answer that.\n>\n> Rogan\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"24722","messageId":"Pine.LNX.4.64.0608040953260.5167@g5.osdl.org","threadId":"5114","inReplyTo":"44D36F64.5040404@gmail.com","subject":"Re: Creating objects manually and repack","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-08-04T16:56:05Z","receivedAt":"2006-08-04T16:56:05Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 4 Aug 2006, A Large Angry SCM wrote:\n> \n> Why don't you just write the pack file directly? Pack files without deltas\n> have a very simple structure, and git-index-pack will create a pack index file\n> for the pack file you give it.\n\nPack-files without deltas are really huge. You really really don't want to \ndo this for some medium-large file that has several thousand revisions.\n\nThe reason you want to generate the deltas early is that then, once you've \ngenerated all the simple and obvious deltas (and within each *,v file from \nCVS, they are all simple and obvious), doing a \"git repack -a -d\" will be \nable to re-use the deltas you found, making it a much cheaper operation.\n\nNOTE! For that \"git repack -a -d\" to work, you'd obviously only do it at \nthe very end, when you've tied together all the blobs with trees and \ncommits (since \"git repack\" wants to follow the reachability chain).\n\n\t\tLinus\n"},{"id":"24724","messageId":"9e4733910608041017v235da03ocd3eeeb0ba0e259b@mail.gmail.com","threadId":"5114","inReplyTo":"Pine.LNX.4.64.0608040945070.5167@g5.osdl.org","subject":"Re: Creating objects manually and repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-08-04T17:17:41Z","receivedAt":"2006-08-04T17:17:41Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 8/4/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> and you're basically all done. The above would turn each *,v file into a\n> *-<sha>.pack/*-<sha>.idx file pair, so you'd have exactly as many\n> pack-files as you have *,v files.\n\nI'll end up with 110,000 pack files. I suspect when I run repack over\nthat it is going to take 24hrs or more, but maybe not since everything\nmay be small enough to run in RAM. We'll also get to see the\nperformance of repack with 110K open file handles. How is it going to\nfigure out which file handle contains which objects?\n\nA new tool might help. It would concatenate the pack files (while\nadjusting the headers) and then build a single index. No attempt at\nsearching for deltas.\n\nTo initially build a single pack file it looks like I need a version\nof repack that works in a single pass over the input files. To make\nthings simple it would just delete the file when it has finished\nreading it. Since I'm passing in the revisions in optimal order\nsorting them probably hurts the pack size. The number of files in\nflight will be a function of the pipe buffer size and file names.\n\nI'll work on the tree writing code over the week end.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"24725","messageId":"Pine.LNX.4.64.0608041027530.5167@g5.osdl.org","threadId":"5114","inReplyTo":"9e4733910608041017v235da03ocd3eeeb0ba0e259b@mail.gmail.com","subject":"Re: Creating objects manually and repack","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-08-04T17:29:57Z","receivedAt":"2006-08-04T17:29:57Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 4 Aug 2006, Jon Smirl wrote:\n> \n> I'll end up with 110,000 pack files. I suspect when I run repack over\n> that it is going to take 24hrs or more, but maybe not since everything\n> may be small enough to run in RAM.\n\nYou may definitely want to pack the pack-files together every once in a \nwhile. Doing so is not that hard: just list all the objects in all the \npack-files you want to merge, which in turn is trivial from reading the \nindex of the pack-files (and then you do want to do the filename, \nalthough you can just use the pack-file name if you want to). \n\nBut yeah, it's going to be expensive whatever you do. It's a big repo.\n\n\t\tLinus\n"},{"id":"24727","messageId":"Pine.LNX.4.64.0608041052030.5167@g5.osdl.org","threadId":"5114","inReplyTo":"Pine.LNX.4.64.0608041027530.5167@g5.osdl.org","subject":"Re: Creating objects manually and repack","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-08-04T18:06:32Z","receivedAt":"2006-08-04T18:06:32Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 4 Aug 2006, Linus Torvalds wrote:\n> \n> You may definitely want to pack the pack-files together every once in a \n> while. Doing so is not that hard: just list all the objects in all the \n> pack-files you want to merge, which in turn is trivial from reading the \n> index of the pack-files (and then you do want to do the filename, \n> although you can just use the pack-file name if you want to). \n\nBtw, that index format is actually documented (and it really is _very_ \nsimple) in Documentation/technical/pack-format.txt.\n\nTo get a list of all object names in a pack-file, you'd basically do just\nsomething like the appended. So with this (let's call it \n\"git-list-objects\"), you could just do\n\n\tfor i in $packlist\n\tdo\n\t\tgit-list-objects $i.idx\n\tdone | git-pack-objects combined-pack\n\nand it would combine all the packs in \"$packlist\" into one new \n\"combined-pack-<sha1>\" pack.\n\nAnd no, I didn't actually _test_ any of this, but it looks pretty damn \nsimple.\n\n\t\tLinus\n\n----\n#include <unistd.h>\n#include <fcntl.h>\n#include <stdio.h>\n\n#define CHUNK (100)\n\nint main(int argc, char **argv)\n{\n\tstatic unsigned char buffer[24*CHUNK];\n\tconst char *name = argv[1];\n\tunsigned int n;\n\tint fd;\n\tint i;\n\n\tif (!name)\n\t\tdie(\"no filename!\");\n\n\tfd = open(name, O_RDONLY);\n\tif (fd < 0)\n\t\tperror(name);\n\n\t/* throw away the first-level fan-out */\n\tif (read(fd, buffer, 4*256) != 4*256)\n\t\tperror(\"read fan-out\");\n\n\tn = (buffer[4*255 + 0] << 24) +\n\t    (buffer[4*255 + 1] << 16) +\n\t    (buffer[4*255 + 2] << 8) +\n\t    (buffer[4*255 + 3] << 0);\n\n\tfor (i = 0; i < n; i += CHUNK) {\n\t\tint j, left = n - i;\n\t\tif (left > CHUNK)\n\t\t\tleft = CHUNK;\n\t\tif (read(fd, buffer, left*24) != left*24)\n\t\t\tperror(\"read chunk\");\n\t\tfor (j = 0; j < left; j++) {\n\t\t\tconst unsigned char *sha1;\n\t\t\tsha1 = buffer + j*24 + 4;\n\t\t\tprintf(\"%s %s\\n\", sha1_to_hex(sha1), name);\n\t\t}\n\t}\n\treturn 0;\n}\n"},{"id":"24728","messageId":"7v64h8l5om.fsf@assigned-by-dhcp.cox.net","threadId":"5114","inReplyTo":"Pine.LNX.4.64.0608041052030.5167@g5.osdl.org","subject":"Re: Creating objects manually and repack","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-08-04T18:24:57Z","receivedAt":"2006-08-04T18:24:57Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> On Fri, 4 Aug 2006, Linus Torvalds wrote:\n>> \n>> You may definitely want to pack the pack-files together every once in a \n>> while. Doing so is not that hard: just list all the objects in all the \n>> pack-files you want to merge, which in turn is trivial from reading the \n>> index of the pack-files (and then you do want to do the filename, \n>> although you can just use the pack-file name if you want to). \n\nThat would only work *once*, because the resulting pack would\nnow have blobs from two or more different files and you cannot\ntell them apart.  So in order to collapse 110k packs into one,\nyou would pack packs into one every 330 packs, create trees and\ncommits for connectivity, and run the final repack -a -d over\nthe result, or something like that, I suppose...\n\n> Btw, that index format is actually documented (and it really is _very_ \n> simple) in Documentation/technical/pack-format.txt.\n>\n> To get a list of all object names in a pack-file, you'd basically do just\n> something like the appended.\n\ngit-show-index?\n"},{"id":"24738","messageId":"Pine.LNX.4.64.0608041218390.5167@g5.osdl.org","threadId":"5114","inReplyTo":"7v64h8l5om.fsf@assigned-by-dhcp.cox.net","subject":"Re: Creating objects manually and repack","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-08-04T19:20:36Z","receivedAt":"2006-08-04T19:20:36Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 4 Aug 2006, Junio C Hamano wrote:\n> \n> That would only work *once*, because the resulting pack would\n> now have blobs from two or more different files and you cannot\n> tell them apart.\n\nYou don't care. You need to keep track of the blob names separately \n_anyway_: the pack information is not enough to re-create all the revision \ninfo.\n\nSo clearly, to create the tree and commit objects, the cvsimport really \nneeds to keep track of the objects it has created, and what their \nrelationship is, and it needs to do that separately. The pack-file just \ncontains the contents, so that you only ever afterwards need to worry \nabout the 20-byte SHA1, not the actual file itself.\n\n> > To get a list of all object names in a pack-file, you'd basically do just\n> > something like the appended.\n> \n> git-show-index?\n\nYeah, that might be good.\n\n\t\tLinus\n"},{"id":"24740","messageId":"87d5bgmh5r.wl%cworth@cworth.org","threadId":"5114","inReplyTo":"Pine.LNX.4.64.0608041218390.5167@g5.osdl.org","subject":"Re: Creating objects manually and repack","fromName":"Carl Worth","fromEmail":"cworth@cworth.org","sentAt":"2006-08-04T19:31:44Z","receivedAt":"2006-08-04T19:31:44Z","isPatch":false,"sender":{"key":"cworth@cworth.org","avatar":"https://gravatar.com/avatar/3746dc28cde609bdbd7f939058356e7e2bbd16d21e32274df0725eb3d998bc5b?d=mp&s=160"},"body":"On Fri, 4 Aug 2006 12:20:36 -0700 (PDT), Linus Torvalds wrote:\n> > > To get a list of all object names in a pack-file, you'd basically do just\n> > > something like the appended.\n> >\n> > git-show-index?\n>\n> Yeah, that might be good.\n\nThat clashes pretty badly with update-index. git-show-pack-index\nperhaps?\n\n-Carl\n"},{"id":"24743","messageId":"7vodv0i89m.fsf@assigned-by-dhcp.cox.net","threadId":"5114","inReplyTo":"87d5bgmh5r.wl%cworth@cworth.org","subject":"Re: Creating objects manually and repack","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-08-04T19:57:25Z","receivedAt":"2006-08-04T19:57:25Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Carl Worth <cworth@cworth.org> writes:\n\n> On Fri, 4 Aug 2006 12:20:36 -0700 (PDT), Linus Torvalds wrote:\n>> > > To get a list of all object names in a pack-file, you'd basically do just\n>> > > something like the appended.\n>> >\n>> > git-show-index?\n>>\n>> Yeah, that might be good.\n>\n> That clashes pretty badly with update-index. git-show-pack-index\n> perhaps?\n\nThere _already_ is a command called git-show-index, since early\nJuly last year ;-).\n"},{"id":"24745","messageId":"87bqr0mfgr.wl%cworth@cworth.org","threadId":"5114","inReplyTo":"7vodv0i89m.fsf@assigned-by-dhcp.cox.net","subject":"Re: Creating objects manually and repack","fromName":"Carl Worth","fromEmail":"cworth@cworth.org","sentAt":"2006-08-04T20:08:20Z","receivedAt":"2006-08-04T20:08:20Z","isPatch":false,"sender":{"key":"cworth@cworth.org","avatar":"https://gravatar.com/avatar/3746dc28cde609bdbd7f939058356e7e2bbd16d21e32274df0725eb3d998bc5b?d=mp&s=160"},"body":"On Fri, 04 Aug 2006 12:57:25 -0700, Junio C Hamano wrote:\n> > That clashes pretty badly with update-index. git-show-pack-index\n> > perhaps?\n>\n> There _already_ is a command called git-show-index, since early\n> July last year ;-).\n\nAh, don't mind me then. Just another one of my typical public displays\nof ignorance. Time to go back and work harder on memorizing that list\nof 120+ commands...\n\n-Carl\n"},{"id":"24746","messageId":"87ac6kmfg7.wl%cworth@cworth.org","threadId":"5114","inReplyTo":"7vodv0i89m.fsf@assigned-by-dhcp.cox.net","subject":"Re: Creating objects manually and repack","fromName":"Carl Worth","fromEmail":"cworth@cworth.org","sentAt":"2006-08-04T20:08:40Z","receivedAt":"2006-08-04T20:08:40Z","isPatch":false,"sender":{"key":"cworth@cworth.org","avatar":"https://gravatar.com/avatar/3746dc28cde609bdbd7f939058356e7e2bbd16d21e32274df0725eb3d998bc5b?d=mp&s=160"},"body":"On Fri, 04 Aug 2006 12:57:25 -0700, Junio C Hamano wrote:\n> > That clashes pretty badly with update-index. git-show-pack-index\n> > perhaps?\n>\n> There _already_ is a command called git-show-index, since early\n> July last year ;-).\n\nAh, don't mind me then. Just another one of my public displays of\nincompetence. Time to go back and work harder on memorizing that list\nof 120+ commands...\n\n-Carl\n"},{"id":"24747","messageId":"eb09ln$h7k$1@sea.gmane.org","threadId":"5114","inReplyTo":"7vodv0i89m.fsf@assigned-by-dhcp.cox.net","subject":"Re: Creating objects manually and repack","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-08-04T20:12:16Z","receivedAt":"2006-08-04T20:12:16Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Junio C Hamano wrote:\n\n> Carl Worth <cworth@cworth.org> writes:\n> \n>> On Fri, 4 Aug 2006 12:20:36 -0700 (PDT), Linus Torvalds wrote:\n>>> > > To get a list of all object names in a pack-file, you'd basically do\njust\n>>> > > something like the appended.\n>>> >\n>>> > git-show-index?\n>>>\n>>> Yeah, that might be good.\n>>\n>> That clashes pretty badly with update-index. git-show-pack-index\n>> perhaps?\n> \n> There _already_ is a command called git-show-index, since early\n> July last year ;-).\n\nParhaps it should be renamed then, for consistency? \n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"24749","messageId":"7vd5bgi6qt.fsf@assigned-by-dhcp.cox.net","threadId":"5114","inReplyTo":"eb09ln$h7k$1@sea.gmane.org","subject":"Re: Creating objects manually and repack","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-08-04T20:30:18Z","receivedAt":"2006-08-04T20:30:18Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> writes:\n\n> Junio C Hamano wrote:\n>\n>> There _already_ is a command called git-show-index, since early\n>> July last year ;-).\n>\n> Parhaps it should be renamed then, for consistency? \n\nThere isn't anything to make consistent.  We use the word \n\"index\" to mean both dircache and the pack index files.\nUsually you can tell between the two usage from context.\n\nWe could rename it to git-show-pack-index, but the command has\nprimarily been the debugging aid and not for real use, and I was\nactually thinking about removing it, perhaps until now ;-).\n"},{"id":"24750","messageId":"eb0b5r$lav$1@sea.gmane.org","threadId":"5114","inReplyTo":"7vd5bgi6qt.fsf@assigned-by-dhcp.cox.net","subject":"Re: Creating objects manually and repack","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-08-04T20:37:56Z","receivedAt":"2006-08-04T20:37:56Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Junio C Hamano wrote:\n\n> Jakub Narebski <jnareb@gmail.com> writes:\n> \n>> Junio C Hamano wrote:\n>>\n>>> There _already_ is a command called git-show-index, since early\n>>> July last year ;-).\n>>\n>> Parhaps it should be renamed then, for consistency? \n> \n> There isn't anything to make consistent.  We use the word \n> \"index\" to mean both dircache and the pack index files.\n> Usually you can tell between the two usage from context.\n\nOops. I should say, for better readibility (to not depend on context), then.\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"24792","messageId":"46a038f90608042115m71adc8ffo77de7940efa847a8@mail.gmail.com","threadId":"5114","inReplyTo":"9e4733910608041017v235da03ocd3eeeb0ba0e259b@mail.gmail.com","subject":"Re: Creating objects manually and repack","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-08-05T04:15:00Z","receivedAt":"2006-08-05T04:15:00Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 8/5/06, Jon Smirl <jonsmirl@gmail.com> wrote:\n> On 8/4/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> > and you're basically all done. The above would turn each *,v file into a\n> > *-<sha>.pack/*-<sha>.idx file pair, so you'd have exactly as many\n> > pack-files as you have *,v files.\n>\n> I'll end up with 110,000 pack files.\n\nThen just do it every 100 files, and you'll only have 1,100 pack\nfiles, and it'll be fine.\n\n> I suspect when I run repack over\n> that it is going to take 24hrs or more,\n\nProbably, but only the initial import has to incur that huge cost.\n\ncheers,\n\n\n\nmartin\n"},{"id":"24793","messageId":"9e4733910608042212p6bf56224ye0ecf3f06b2840cf@mail.gmail.com","threadId":"5114","inReplyTo":"46a038f90608042115m71adc8ffo77de7940efa847a8@mail.gmail.com","subject":"Re: Creating objects manually and repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-08-05T05:12:53Z","receivedAt":"2006-08-05T05:12:53Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 8/5/06, Martin Langhoff <martin.langhoff@gmail.com> wrote:\n> On 8/5/06, Jon Smirl <jonsmirl@gmail.com> wrote:\n> > On 8/4/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> > > and you're basically all done. The above would turn each *,v file into a\n> > > *-<sha>.pack/*-<sha>.idx file pair, so you'd have exactly as many\n> > > pack-files as you have *,v files.\n> >\n> > I'll end up with 110,000 pack files.\n>\n> Then just do it every 100 files, and you'll only have 1,100 pack\n> files, and it'll be fine.\n\nThis is something that has to be tuned. If you wait too long\neverything spills out of RAM and you go totally IO bound for days. If\nyou do it too often you end up with too many packs and it takes a day\nto repack them.\n\nIf I had a way to pipe the all of the objects into repack one at a\ntime without repack doing multiple passes none of this tuning would be\nnecessary. In this model the standalone objects never get created in\nthe first place. The fastest IO is IO that has been eliminated.\n\n> > I suspect when I run repack over\n> > that it is going to take 24hrs or more,\n>\n> Probably, but only the initial import has to incur that huge cost.\n\nMozilla developers aren't all rushing to switch to git. A switch needs\nto be as painless as possible. If things are too complex they simply\nwon't switch.\n\nSwitching Mozilla to git is going to require a sales job and proof\nthat the tools are reliable and better than CVS. Right now I can't\neven reliably import Mozilla CVS. One of the conditions for even\nconsidering git is that they can easily do the CVS import internally\nand verify it for accuracy.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"24795","messageId":"20060805052135.GA18679@spearce.org","threadId":"5114","inReplyTo":"9e4733910608042212p6bf56224ye0ecf3f06b2840cf@mail.gmail.com","subject":"Re: Creating objects manually and repack","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-08-05T05:21:36Z","receivedAt":"2006-08-05T05:21:36Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Jon Smirl <jonsmirl@gmail.com> wrote:\n> On 8/5/06, Martin Langhoff <martin.langhoff@gmail.com> wrote:\n> >On 8/5/06, Jon Smirl <jonsmirl@gmail.com> wrote:\n> >> On 8/4/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> >> > and you're basically all done. The above would turn each *,v file into \n> >a\n> >> > *-<sha>.pack/*-<sha>.idx file pair, so you'd have exactly as many\n> >> > pack-files as you have *,v files.\n> >>\n> >> I'll end up with 110,000 pack files.\n> >\n> >Then just do it every 100 files, and you'll only have 1,100 pack\n> >files, and it'll be fine.\n> \n> This is something that has to be tuned. If you wait too long\n> everything spills out of RAM and you go totally IO bound for days. If\n> you do it too often you end up with too many packs and it takes a day\n> to repack them.\n> \n> If I had a way to pipe the all of the objects into repack one at a\n> time without repack doing multiple passes none of this tuning would be\n> necessary. In this model the standalone objects never get created in\n> the first place. The fastest IO is IO that has been eliminated.\n\nI'm almost done with what I'm calling `git-fast-import`.  It takes\na stream of blobs on STDIN and writes the pack to a file, printing\nSHA1s in hex format to STDOUT.  The basic format for STDIN is a 4\nbyte length (native format) followed by that many bytes of blob data.\nIt prints the SHA1 for that blob to STDOUT, then waits for another\nlength.\n\nIt naively deltas each object against the prior object, thus it\nwould be best to feed it one ,v file at a time working from the most\nrecent revision back to the oldest revision.  This works well for\nan RCS file as that's the natural order to process the file in.  :-)\n\nWhen done you close STDIN and it'll rip through and update the pack\nobject count and the trailing checksum.  This should let you pack\nthe entire repository in delta format using only two passes over the\ndata: one to write out the pack file and one to compute its checksum.\n\n\nI'll post the code in a couple of hours.\n\n-- \nShawn.\n"},{"id":"24798","messageId":"9e4733910608042240u581dd23q3859ebcfe4268ce2@mail.gmail.com","threadId":"5114","inReplyTo":"20060805052135.GA18679@spearce.org","subject":"Re: Creating objects manually and repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-08-05T05:40:23Z","receivedAt":"2006-08-05T05:40:23Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 8/5/06, Shawn Pearce <spearce@spearce.org> wrote:\n> I'm almost done with what I'm calling `git-fast-import`.  It takes\n> a stream of blobs on STDIN and writes the pack to a file, printing\n> SHA1s in hex format to STDOUT.  The basic format for STDIN is a 4\n> byte length (native format) followed by that many bytes of blob data.\n> It prints the SHA1 for that blob to STDOUT, then waits for another\n> length.\n>\n> It naively deltas each object against the prior object, thus it\n> would be best to feed it one ,v file at a time working from the most\n> recent revision back to the oldest revision.  This works well for\n> an RCS file as that's the natural order to process the file in.  :-)\n\nI am already doing this.\n\n> When done you close STDIN and it'll rip through and update the pack\n> object count and the trailing checksum.  This should let you pack\n> the entire repository in delta format using only two passes over the\n> data: one to write out the pack file and one to compute its checksum.\n\nThinking about this some more, the existing repack code could be made\nto work with minor changes. I would like to feed repack 1M revisions\nwhich are sorted by file and then newest to oldest. The problem is\nthat my expanded revs take up 12GB disk space.\n\nHow about adding a flag to repack that simply says delete the objects\nwhen done with them? I'd still create all of the objects on disk.\nRepack would assume that they have at least been sorted by filename.\nSo repack could read in object names until it sees a change in the\nfile name, sort them by size, deltafy, write out the pack and then\ndelete the objects from that batch. Then repeat this process for the\nnext file name on stdin.\n\nI'm making two assumptions, first that blocks from a deleted file\ndon't get written to disk. And that by deleting the file the file\nsystem will use the same blocks over and over. If those assumptions\nare close to being true then the cache shouldn't thrash. They don't\nhave to be totally true, close is good enough.\n\nOf course eliminating the files all together will be even faster.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"24799","messageId":"20060805054655.GB18679@spearce.org","threadId":"5114","inReplyTo":"20060805052135.GA18679@spearce.org","subject":"Re: Creating objects manually and repack","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-08-05T05:46:55Z","receivedAt":"2006-08-05T05:46:55Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Shawn Pearce <spearce@spearce.org> wrote:\n> Jon Smirl <jonsmirl@gmail.com> wrote:\n> > On 8/5/06, Martin Langhoff <martin.langhoff@gmail.com> wrote:\n> > >On 8/5/06, Jon Smirl <jonsmirl@gmail.com> wrote:\n> > >> On 8/4/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> > >> > and you're basically all done. The above would turn each *,v file into \n> > >a\n> > >> > *-<sha>.pack/*-<sha>.idx file pair, so you'd have exactly as many\n> > >> > pack-files as you have *,v files.\n> > >>\n> > >> I'll end up with 110,000 pack files.\n> > >\n> > >Then just do it every 100 files, and you'll only have 1,100 pack\n> > >files, and it'll be fine.\n> > \n> > This is something that has to be tuned. If you wait too long\n> > everything spills out of RAM and you go totally IO bound for days. If\n> > you do it too often you end up with too many packs and it takes a day\n> > to repack them.\n> > \n> > If I had a way to pipe the all of the objects into repack one at a\n> > time without repack doing multiple passes none of this tuning would be\n> > necessary. In this model the standalone objects never get created in\n> > the first place. The fastest IO is IO that has been eliminated.\n> \n> I'm almost done with what I'm calling `git-fast-import`.\n\nOK, now I'm done.  I'm attaching the code.  Toss it into the Makefile\nas git-fast-import and recompile.\n\nI tested it with the following Perl script, feeding the Perl script\na list of files that I wanted blobs for on STDIN:\n\n\twhile (<>) {\n\t\tchop;\n\t\tprint pack('L', -s $_);\n\t\topen(F, $_);\n\t\tmy $buf;\n\t\tprint $buf while read(F,$buf,128*1024) > 0;\n\t\tclose F;\n\t}\n\nThis gave me an execution order of:\n\n\tfind . -name '*.c' | perl test.pl | git-fast-import in.pack\n\tgit-index-pack in.pack\n\nat which point in.pack claims to be a completely valid pack with an\nindex of in.idx.  Move these into .git/objects/pack, generate trees\nand commits, and run git-repack -a -d.  If the order you feed the\nobjects to git-fast-import in is reasonable (do one RCS file at a\ntime, feed most recent to least recent revisions) you may not get\nany major benefit from using -f during your final repack.\n\nThe code for git-fast-import could probably be tweaked to accept\ntrees and commits too, which would permit you to stream the entire\nCVS repository into a single pack file.  :-)\n\nI can't help you decompress the RCS files faster, but hopefully\nthis will help you generate the GIT pack faster.  Hopefully you\ncan make use of it!\n\n-- \nShawn.\n\n\n#include \"builtin.h\"\n#include \"cache.h\"\n#include \"object.h\"\n#include \"blob.h\"\n#include \"delta.h\"\n#include \"pack.h\"\n#include \"csum-file.h\"\n\nstatic int max_depth = 10;\nstatic unsigned long object_count;\nstatic int packfd;\nstatic int current_depth;\nstatic void *lastdat;\nstatic unsigned long lastdatlen;\nstatic unsigned char lastsha1[20];\n\nstatic ssize_t yread(int fd, void *buffer, size_t length)\n{\n\tssize_t ret = 0;\n\twhile (ret < length) {\n\t\tssize_t size = xread(fd, (char *) buffer + ret, length - ret);\n\t\tif (size < 0) {\n\t\t\treturn size;\n\t\t}\n\t\tif (size == 0) {\n\t\t\treturn ret;\n\t\t}\n\t\tret += size;\n\t}\n\treturn ret;\n}\n\nstatic ssize_t ywrite(int fd, void *buffer, size_t length)\n{\n\tssize_t ret = 0;\n\twhile (ret < length) {\n\t\tssize_t size = xwrite(fd, (char *) buffer + ret, length - ret);\n\t\tif (size < 0) {\n\t\t\treturn size;\n\t\t}\n\t\tif (size == 0) {\n\t\t\treturn ret;\n\t\t}\n\t\tret += size;\n\t}\n\treturn ret;\n}\n\nstatic unsigned long encode_header(enum object_type type, unsigned long size, unsigned char *hdr)\n{\n\tint n = 1;\n\tunsigned char c;\n\n\tif (type < OBJ_COMMIT || type > OBJ_DELTA)\n\t\tdie(\"bad type %d\", type);\n\n\tc = (type << 4) | (size & 15);\n\tsize >>= 4;\n\twhile (size) {\n\t\t*hdr++ = c | 0x80;\n\t\tc = size & 0x7f;\n\t\tsize >>= 7;\n\t\tn++;\n\t}\n\t*hdr = c;\n\treturn n;\n}\n\nstatic void write_blob (void *dat, unsigned long datlen)\n{\n\tz_stream s;\n\tvoid *out, *delta;\n\tunsigned char hdr[64];\n\tunsigned long hdrlen, deltalen;\n\n\tif (lastdat && current_depth < max_depth) {\n\t\tdelta = diff_delta(lastdat, lastdatlen,\n\t\t\tdat, datlen,\n\t\t\t&deltalen, 0);\n\t} else\n\t\tdelta = 0;\n\n\tmemset(&s, 0, sizeof(s));\n\tdeflateInit(&s, zlib_compression_level);\n\n\tif (delta) {\n\t\tcurrent_depth++;\n\t\ts.next_in = delta;\n\t\ts.avail_in = deltalen;\n\t\thdrlen = encode_header(OBJ_DELTA, deltalen, hdr);\n\t\tif (ywrite(packfd, hdr, hdrlen) != hdrlen)\n\t\t\tdie(\"Can't write object header: %s\", strerror(errno));\n\t\tif (ywrite(packfd, lastsha1, sizeof(lastsha1)) != sizeof(lastsha1))\n\t\t\tdie(\"Can't write object base: %s\", strerror(errno));\n\t} else {\n\t\tcurrent_depth = 0;\n\t\ts.next_in = dat;\n\t\ts.avail_in = datlen;\n\t\thdrlen = encode_header(OBJ_BLOB, datlen, hdr);\n\t\tif (ywrite(packfd, hdr, hdrlen) != hdrlen)\n\t\t\tdie(\"Can't write object header: %s\", strerror(errno));\n\t}\n\n\ts.avail_out = deflateBound(&s, s.avail_in);\n\ts.next_out = out = xmalloc(s.avail_out);\n\twhile (deflate(&s, Z_FINISH) == Z_OK)\n\t\t/* nothing */;\n\tdeflateEnd(&s);\n\n\tif (ywrite(packfd, out, s.total_out) != s.total_out)\n\t\tdie(\"Failed writing compressed data %s\", strerror(errno));\n\n\tfree(out);\n\tif (delta)\n\t\tfree(delta);\n}\n\nstatic void init_pack_header ()\n{\n\tconst char* magic = \"PACK\";\n\tunsigned long version = 2;\n\tunsigned long zero = 0;\n\n\tversion = htonl(version);\n\n\tif (ywrite(packfd, (char*)magic, 4) != 4)\n\t\tdie(\"Can't write pack magic: %s\", strerror(errno));\n\tif (ywrite(packfd, &version, 4) != 4)\n\t\tdie(\"Can't write pack version: %s\", strerror(errno));\n\tif (ywrite(packfd, &zero, 4) != 4)\n\t\tdie(\"Can't write 0 object count: %s\", strerror(errno));\n}\n\nstatic void fixup_header_footer ()\n{\n\tSHA_CTX c;\n\tchar hdr[8];\n\tunsigned char sha1[20];\n\tunsigned long cnt;\n\tchar *buf;\n\tsize_t n;\n\n\tif (lseek(packfd, 0, SEEK_SET) != 0)\n\t\tdie(\"Failed seeking to start: %s\", strerror(errno));\n\n\tSHA1_Init(&c);\n\tif (yread(packfd, hdr, 8) != 8)\n\t\tdie(\"Failed reading header: %s\", strerror(errno));\n\tSHA1_Update(&c, hdr, 8);\n\nfprintf(stderr, \"%lu objects\\n\", object_count);\n\tcnt = htonl(object_count);\n\tSHA1_Update(&c, &cnt, 4);\n\tif (ywrite(packfd, &cnt, 4) != 4)\n\t\tdie(\"Failed writing object count: %s\", strerror(errno));\n\n\tbuf = xmalloc(128 * 1024);\n\tfor (;;) {\n\t\tn = xread(packfd, buf, 128 * 1024);\n\t\tif (n <= 0)\n\t\t\tbreak;\n\t\tSHA1_Update(&c, buf, n);\n\t}\n\tfree(buf);\n\n\tSHA1_Final(sha1, &c);\n\tif (ywrite(packfd, sha1, sizeof(sha1)) != sizeof(sha1))\n\t\tdie(\"Failed writing pack checksum: %s\", strerror(errno));\n}\n\nint main (int argc, const char **argv)\n{\n\tpackfd = open(argv[1], O_RDWR|O_CREAT|O_TRUNC, 0666);\n\tif (packfd < 0)\n\t\tdie(\"Can't create pack file %s: %s\", argv[1], strerror(errno));\n\n\tinit_pack_header();\n\tfor (;;) {\n\t\tunsigned long datlen;\n\t\tint hdrlen;\n\t\tvoid *dat;\n\t\tchar hdr[128];\n\t\tunsigned char sha1[20];\n\t\tSHA_CTX c;\n\n\t\tif (yread(0, &datlen, 4) != 4)\n\t\t\tbreak;\n\n\t\tdat = xmalloc(datlen);\n\t\tif (yread(0, dat, datlen) != datlen)\n\t\t\tbreak;\n\n\t\thdrlen = sprintf(hdr, \"blob %lu\", datlen) + 1;\n\t\tSHA1_Init(&c);\n\t\tSHA1_Update(&c, hdr, hdrlen);\n\t\tSHA1_Update(&c, dat, datlen);\n\t\tSHA1_Final(sha1, &c);\n\n\t\twrite_blob(dat, datlen);\n\t\tobject_count++;\n\t\tprintf(\"%s\\n\", sha1_to_hex(sha1));\n\t\tfflush(stdout);\n\n\t\tif (lastdat)\n\t\t\tfree(lastdat);\n\t\tlastdat = dat;\n\t\tlastdatlen = datlen;\n\t\tmemcpy(lastsha1, sha1, sizeof(sha1));\n\t}\n\tfixup_header_footer();\n\tclose(packfd);\n\n\treturn 0;\n}\n"},{"id":"24801","messageId":"20060805055203.GC18679@spearce.org","threadId":"5114","inReplyTo":"9e4733910608042240u581dd23q3859ebcfe4268ce2@mail.gmail.com","subject":"Re: Creating objects manually and repack","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-08-05T05:52:03Z","receivedAt":"2006-08-05T05:52:03Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Jon Smirl <jonsmirl@gmail.com> wrote:\n> How about adding a flag to repack that simply says delete the objects\n> when done with them? I'd still create all of the objects on disk.\n> Repack would assume that they have at least been sorted by filename.\n> So repack could read in object names until it sees a change in the\n> file name, sort them by size, deltafy, write out the pack and then\n> delete the objects from that batch. Then repeat this process for the\n> next file name on stdin.\n> \n> I'm making two assumptions, first that blocks from a deleted file\n> don't get written to disk. And that by deleting the file the file\n> system will use the same blocks over and over. If those assumptions\n> are close to being true then the cache shouldn't thrash. They don't\n> have to be totally true, close is good enough.\n> \n> Of course eliminating the files all together will be even faster.\n\nSee the email I just sent you.  The only file being written is the\npack file that's being generated.  No temporary files, no temporary\ninodes, no temporary blocks.  Only two passes over the data: one to\nwrite it out and a second to generate the SHA1.  I do two passes\nvs. keep it all in memory to prevent the program from blowing out\non extremely large inputs.\n\nIt may be possible to tweak git-pack-objects to get what you propose\nabove, but to be honest I think the git-fast-import I just sent\nwas easier, especially as it avoids the temporary loose object stage.\n\n-- \nShawn.\n"}]}