{"thread":{"id":"4505","subject":"Repacking many disconnected blobs","startedAt":"2006-06-14T07:17:58Z","lastAt":"2006-06-14T21:20:40Z","messageCount":15,"participants":["Keith Packard","Shawn Pearce","Johannes Schindelin","Sergey Vlasov","Junio C Hamano","Linus Torvalds","Nicolas Pitre"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"21769","messageId":"1150269478.20536.150.camel@neko.keithp.com","threadId":"4505","inReplyTo":null,"subject":"Repacking many disconnected blobs","fromName":"Keith Packard","fromEmail":"keithp@keithp.com","sentAt":"2006-06-14T07:17:58Z","receivedAt":"2006-06-14T07:17:58Z","isPatch":false,"sender":{"key":"keithp@keithp.com","avatar":"https://gravatar.com/avatar/fa1f479cdd51322fe86215c955a81d296bbf66a1fe625f8a12d87a8ec7faf648?d=mp&s=160"},"body":"parsecvs scans every ,v file and creates a blob for every revision of\nevery file right up front. Once these are created, it discards the\nactual file contents and deals solely with the hash values.\n\nThe problem is that while this is going on, the repository consists\nsolely of disconnected objects, and I can't make git-repack put those\ninto pack objects. This leaves the directories bloated, and operations\nwithin the tree quite sluggish. I'm importing a project with 30000 files\nand 30000 revisions (the CVS repository is about 700MB), and after\nscanning the files, and constructing (in memory) a complete revision\nhistory, the actual construction of the commits is happening at about 2\nper second, and about 70% of that time is in the kernel, presumably\nplaying around in the repository.\n\nI'm assuming that if I could get these disconnected blobs all neatly\ntucked into a pack object, things might go a bit faster.\n-- \nkeith.packard@intel.com\n"},{"id":"21770","messageId":"20060614072923.GB13886@spearce.org","threadId":"4505","inReplyTo":"1150269478.20536.150.camel@neko.keithp.com","subject":"Re: Repacking many disconnected blobs","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-06-14T07:29:24Z","receivedAt":"2006-06-14T07:29:24Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Keith Packard <keithp@keithp.com> wrote:\n> parsecvs scans every ,v file and creates a blob for every revision of\n> every file right up front. Once these are created, it discards the\n> actual file contents and deals solely with the hash values.\n> \n> The problem is that while this is going on, the repository consists\n> solely of disconnected objects, and I can't make git-repack put those\n> into pack objects. This leaves the directories bloated, and operations\n> within the tree quite sluggish. I'm importing a project with 30000 files\n> and 30000 revisions (the CVS repository is about 700MB), and after\n> scanning the files, and constructing (in memory) a complete revision\n> history, the actual construction of the commits is happening at about 2\n> per second, and about 70% of that time is in the kernel, presumably\n> playing around in the repository.\n> \n> I'm assuming that if I could get these disconnected blobs all neatly\n> tucked into a pack object, things might go a bit faster.\n\nWhat about running git-update-index using .git/objects as the\ncurrent working directory and adding all files in ??/* into the\nindex, then git-write-tree that index and git-commit-tree the tree.\n\nWhen you are done you have a bunch of orphan trees and a commit\nbut these shouldn't be very big and I'd guess would prune out with\na repack if you don't hold a ref to the orphan commit.\n\n-- \nShawn.\n"},{"id":"21773","messageId":"Pine.LNX.4.63.0606141104050.15578@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"4505","inReplyTo":"20060614072923.GB13886@spearce.org","subject":"Re: Repacking many disconnected blobs","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-06-14T09:07:52Z","receivedAt":"2006-06-14T09:07:52Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 14 Jun 2006, Shawn Pearce wrote:\n\n> Keith Packard <keithp@keithp.com> wrote:\n> > parsecvs scans every ,v file and creates a blob for every revision of\n> > every file right up front. Once these are created, it discards the\n> > actual file contents and deals solely with the hash values.\n> > \n> > The problem is that while this is going on, the repository consists\n> > solely of disconnected objects, and I can't make git-repack put those\n> > into pack objects. This leaves the directories bloated, and operations\n> > within the tree quite sluggish. I'm importing a project with 30000 files\n> > and 30000 revisions (the CVS repository is about 700MB), and after\n> > scanning the files, and constructing (in memory) a complete revision\n> > history, the actual construction of the commits is happening at about 2\n> > per second, and about 70% of that time is in the kernel, presumably\n> > playing around in the repository.\n> > \n> > I'm assuming that if I could get these disconnected blobs all neatly\n> > tucked into a pack object, things might go a bit faster.\n> \n> What about running git-update-index using .git/objects as the\n> current working directory and adding all files in ??/* into the\n> index, then git-write-tree that index and git-commit-tree the tree.\n> \n> When you are done you have a bunch of orphan trees and a commit\n> but these shouldn't be very big and I'd guess would prune out with\n> a repack if you don't hold a ref to the orphan commit.\n\nAlternatively, you could construct fake trees like this:\n\nREADME/1.1.1.1\nREADME/1.2\nREADME/1.3\n...\n\ni.e. every file becomes a directory -- containing all the versions of that \nfile -- in the (virtual) tree, which you can point to by a temporary ref.\n\nCiao,\nDscho\n"},{"id":"21775","messageId":"20060614133741.73cd80cb.vsu@altlinux.ru","threadId":"4505","inReplyTo":"1150269478.20536.150.camel@neko.keithp.com","subject":"Re: Repacking many disconnected blobs","fromName":"Sergey Vlasov","fromEmail":"vsu@altlinux.ru","sentAt":"2006-06-14T09:37:41Z","receivedAt":"2006-06-14T09:37:41Z","isPatch":false,"sender":{"key":"vsu@altlinux.ru","avatar":"https://avatars.githubusercontent.com/u/616082?v=4"},"body":"On Wed, 14 Jun 2006 00:17:58 -0700 Keith Packard wrote:\n\n> parsecvs scans every ,v file and creates a blob for every revision of\n> every file right up front. Once these are created, it discards the\n> actual file contents and deals solely with the hash values.\n> \n> The problem is that while this is going on, the repository consists\n> solely of disconnected objects, and I can't make git-repack put those\n> into pack objects. This leaves the directories bloated, and operations\n> within the tree quite sluggish. I'm importing a project with 30000 files\n> and 30000 revisions (the CVS repository is about 700MB), and after\n> scanning the files, and constructing (in memory) a complete revision\n> history, the actual construction of the commits is happening at about 2\n> per second, and about 70% of that time is in the kernel, presumably\n> playing around in the repository.\n> \n> I'm assuming that if I could get these disconnected blobs all neatly\n> tucked into a pack object, things might go a bit faster.\n\ngit-repack.sh basically does:\n\n  git-rev-list --objects --all | git-pack-objects .tmp-pack\n\nWhen you have only disconnected blobs, obviously the first part does\nnot work - git-rev-list cannot find these blobs.  However, you can do\nthat part manually - e.g., when you add a blob, do:\n\n  fprintf(list_file, \"%s %s\\n\", sha1, path);\n\n(path should be a relative path in the repo without \",v\" or \"Attic\" -\nit is used for delta packing optimization, so getting it wrong will\nnot cause any corruption, but the pack may become significantly\nlarger).  You may output some duplicate sha1 values, but\ngit-pack-objects should handle duplicates correctly.\n\nThen just invoke \"git-pack-objects --non-empty .tmp_pack <list_file\";\nit will output the resulting pack sha1 to stdout.  Then you need to\nmove the pack into place and call git-prune-packed (which does not\nuse object lists, so it should work even with unreachable objects).\n\nYou may even want to repack more than once during the import;\nprobably the simplest way to do it is to truncate list_file after\neach repack and use \"git-pack-objects --incremental\".\n"},{"id":"21780","messageId":"7vejxrgbn1.fsf@assigned-by-dhcp.cox.net","threadId":"4505","inReplyTo":"Pine.LNX.4.63.0606141104050.15578@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Repacking many disconnected blobs","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-06-14T12:33:22Z","receivedAt":"2006-06-14T12:33:22Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> Alternatively, you could construct fake trees like this:\n>\n> README/1.1.1.1\n> README/1.2\n> README/1.3\n> ...\n>\n> i.e. every file becomes a directory -- containing all the versions of that \n> file -- in the (virtual) tree, which you can point to by a temporary ref.\n\nThat would not play well with the packing heuristics, I suspect.\nIf you reverse it to use rev/file-id, then the same files from\ndifferent revs would sort closer, though.\n"},{"id":"21787","messageId":"Pine.LNX.4.64.0606140826200.5498@g5.osdl.org","threadId":"4505","inReplyTo":"1150269478.20536.150.camel@neko.keithp.com","subject":"Re: Repacking many disconnected blobs","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-14T15:53:22Z","receivedAt":"2006-06-14T15:53:22Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 14 Jun 2006, Keith Packard wrote:\n>\n> parsecvs scans every ,v file and creates a blob for every revision of\n> every file right up front. Once these are created, it discards the\n> actual file contents and deals solely with the hash values.\n> \n> The problem is that while this is going on, the repository consists\n> solely of disconnected objects, and I can't make git-repack put those\n> into pack objects.\n\nOk. That's actually _easily_ rectifiable, because it turns out that your \nbehaviour is something that re-packing is actually really good at \nhandling.\n\nThe thing is, \"git repack\" (the wrapper function) is all about finding all \nthe heads of a repository, and then tellign the _real_ packing logic which \nobjects to pack.\n\nIn other words, it literally boils down to basically\n\n\tgit-rev-list --all --objects $rev_list |\n\t\tgit-pack-objects --non-empty $pack_objects .tmp-pack\n\nwhere \"$rev_list\" and \"$pack_objects\" are just extra flags to the two \nphases that you don't really care about.\n\nBut the important point to recognize is that the pack generation itself \ndoesn't care about reachability or anything else AT ALL. The pack is just \na jumble of objects, nothing more. Which is exactly what you want.\n\n> I'm assuming that if I could get these disconnected blobs all neatly\n> tucked into a pack object, things might go a bit faster.\n\nAbsolutely. And it's even easy.\n\nWhat you should do is to just generate a list of objects every once in a \nwhile, and pass that list off to \"git-pack-objects\", which will create a \npack-file for you. Then you just move the generated pack-file (and index \nfile) into the .git/objects/pack directory, and then you can run the \nnormal \"git-prune-packed\", and you're done.\n\nThere's just two small subtle points to look out for:\n\n - You can list the objects with \"most important first\" order first, if \n   you can.  That will improve locality later (the packing will try to \n   generate the pack so that the order you gave the objects in will be a \n   rough order of the resul - the first objects will be together at the \n   beginning, the last objects will be at the end)\n\n   This is not a huge deal. If you don't have a good order, give them in \n   any order, and then after you're done (and you do have branches and \n   tag-heads), the final repack (with a regular \"git repack\") will fix it \n   all up.\n\n   You'll still get all of the size/access advantage of packfiles without \n   this, it just won't have the additional \"nice IO patterns within the \n   packfile\" behaviour (which mainly matters for the cold-cache case, so \n   you may well not care).\n\n - append the filename the object is associated with to the object name on \n   the list, if at all possible. This is what git-pack-objects will use as \n   part of the heuristic for finding the deltas, so this is actually a big \n   deal. If you forget (or mess up) the filename, packing will still \n   _work_ - it's just a heuristic, after all, and there are a few others \n   too - but the pack-file will have inferior delta chains.\n\n   (The name doesn't have to be the \"real name\", it really only needs to \n   be something unique per *,v file, but real name is probably best)\n\n   The corollary to this is that it's better to generate the pack-file \n   from a list of every version of a few files than it is to generate it \n   from a few versions of every file. Ie, if you process things one file \n   at a time, and create every object for that file, that is actually good \n   for packing, since there will be the optimal delta opportunity.\n\nIn other words, you should just feed git-pack-file a list of objects in \nthe form \"<sha1><space><filename>\\n\", and git-pack-file will do the rest.\n\nJust as a stupid example, if you were to want to pack just the _tree_ that \nis the current version of a git archive, you'd do\n\n\tgit-rev-list --objects HEAD^{tree} |\n\t\tgit-pack-objects --non-empty .tmp-pack\n\nwhich you can try on the current git tree just to see (the first line will \ngenerate a list of all objects reachable from the current _tree_: no \nhistory at all, the second line will create two files under the name of  \n\".tmp-pack-<sha1-of-object-list>.{pack|idx}\".\n\nThe reason I suggest doing this for the current tree of the git archive is \nsimply that you can look at the git-rev-list output with \"less\", and see \nfor yourself what it actually does (and there are just a few hundred \nobjects there: a few tree objects, and the blob objects for every file in \nthe current HEAD).\n\nSo the git pack-format is actually _optimal_ for your particular case, \nexactly because the pack-files don't actually care about any high-level \nsemantics: all they contain is a list of objects.\n\nSo in phase 1, when you generate all the objects, the simplest thing to do \nis to literally just remember the last five thousand objects or so as you \ngenerate them, and when that array of objects fills up, you just start the \n\"git-pack-objects\" thing, and feed it the list of objects, move the \npack-file into .git/objects/pack/pack-... and do a \"git prune-packed\". \n\nThen you just continue.\n\nSo this should all fit the parsecvs approach very well indeed.\n\n\t\tLinus\n"},{"id":"21788","messageId":"1150307715.20536.166.camel@neko.keithp.com","threadId":"4505","inReplyTo":"Pine.LNX.4.64.0606140826200.5498@g5.osdl.org","subject":"Re: Repacking many disconnected blobs","fromName":"Keith Packard","fromEmail":"keithp@keithp.com","sentAt":"2006-06-14T17:55:15Z","receivedAt":"2006-06-14T17:55:15Z","isPatch":false,"sender":{"key":"keithp@keithp.com","avatar":"https://gravatar.com/avatar/fa1f479cdd51322fe86215c955a81d296bbf66a1fe625f8a12d87a8ec7faf648?d=mp&s=160"},"body":"On Wed, 2006-06-14 at 08:53 -0700, Linus Torvalds wrote:\n\n>  - You can list the objects with \"most important first\" order first, if \n>    you can.  That will improve locality later (the packing will try to \n>    generate the pack so that the order you gave the objects in will be a \n>    rough order of the resul - the first objects will be together at the \n>    beginning, the last objects will be at the end)\n\nI take every ,v file and construct blobs for every revision. If I\nunderstand this correctly, I should be shuffling the revisions so I send\nthe latest revision of every file first, then the next-latest revision.\nIt would be somewhat easier to just send the whole list of revisions for\nthe first file and then move to the next file, but if shuffling is what\nI want, I'll do that.\n\n>    The corollary to this is that it's better to generate the pack-file \n>    from a list of every version of a few files than it is to generate it \n>    from a few versions of every file. Ie, if you process things one file \n>    at a time, and create every object for that file, that is actually good \n>    for packing, since there will be the optimal delta opportunity.\n\nI assumed that was the case. Fortunately, I process each file\nseparately, so this matches my needs exactly. I should be able to report\non this shortly.\n\n-- \nkeith.packard@intel.com\n"},{"id":"21789","messageId":"Pine.LNX.4.64.0606141113130.5498@g5.osdl.org","threadId":"4505","inReplyTo":"1150307715.20536.166.camel@neko.keithp.com","subject":"Re: Repacking many disconnected blobs","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-14T18:18:15Z","receivedAt":"2006-06-14T18:18:15Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 14 Jun 2006, Keith Packard wrote:\n\n> On Wed, 2006-06-14 at 08:53 -0700, Linus Torvalds wrote:\n> \n> >  - You can list the objects with \"most important first\" order first, if \n> >    you can.  That will improve locality later (the packing will try to \n> >    generate the pack so that the order you gave the objects in will be a \n> >    rough order of the resul - the first objects will be together at the \n> >    beginning, the last objects will be at the end)\n> \n> I take every ,v file and construct blobs for every revision. If I\n> understand this correctly, I should be shuffling the revisions so I send\n> the latest revision of every file first, then the next-latest revision.\n> It would be somewhat easier to just send the whole list of revisions for\n> the first file and then move to the next file, but if shuffling is what\n> I want, I'll do that.\n\nYou don't _need_ to shuffle. As mentioned, it will only affect the \nlocation of the data in the pack-file, which in turn will mostly matter \nas an IO pattern thing, not anything really fundamental.  If the pack-file \nends up caching well, the IO patterns obviously will never matter.\n\nEventually, after the whole import has finished, and you do the final \nrepack, that one will do things in \"recency order\" (or \"global \nreachability order\" if you prefer), which means that all the objects in \nthe final pack will be sorted by how \"close\" they are to the top-of-tree. \n\nAnd that will happen regardless of what the intermediate ordering has \nbeen.\n\nSo if shuffling is inconvenient, just don't do it.\n\nOn the other hand, if you know that you generated the blobs \"oldest to \nnewest\", just print them in the reverse order when you end up repacking, \nand you're all done (if you just save the info into some array before you \nrepack, just walk the array backwards).\n\n\t\t\tLinus\n"},{"id":"21790","messageId":"Pine.LNX.4.64.0606141146310.5498@g5.osdl.org","threadId":"4505","inReplyTo":"Pine.LNX.4.64.0606141113130.5498@g5.osdl.org","subject":"Re: Repacking many disconnected blobs","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-14T18:52:21Z","receivedAt":"2006-06-14T18:52:21Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 14 Jun 2006, Linus Torvalds wrote:\n> \n> You don't _need_ to shuffle. As mentioned, it will only affect the \n> location of the data in the pack-file, which in turn will mostly matter \n> as an IO pattern thing, not anything really fundamental.  If the pack-file \n> ends up caching well, the IO patterns obviously will never matter.\n\nActually, thinking about it more, the way you do things, shuffling \nprobably won't even help.\n\nWhy? Because you'll obviously have multiple files, and even if each file \nwere to be sorted \"correctly\", the access patterns from any global \nstandpoint won't really matter, becase you'd probably bounce back and \nforth in the pack-file anyway.\n\nSo if anything, I would say\n\n - just dump them into the packfile in whatever order is most convenient\n\n - if you know that later phases will go through the objects and actually \n   use them (as opposed to just building trees out of their SHA1 values) \n   in some particular order, _that_ might be the ordering to use.\n\n - in many ways, getting good delta chains is _much_ more important, since \n   \"git repack -a -d\" will re-use good deltas from a previous pack, but \n   will _not_ care about any ordering in the old pack. As well as \n   obviously improving the size of the temporary pack-files anyway.\n\nI'll pontificate more if I can think of any other cases that might matter.\n\n\t\tLinus\n"},{"id":"21791","messageId":"1150311567.30681.28.camel@neko.keithp.com","threadId":"4505","inReplyTo":"Pine.LNX.4.64.0606141113130.5498@g5.osdl.org","subject":"Re: Repacking many disconnected blobs","fromName":"Keith Packard","fromEmail":"keithp@keithp.com","sentAt":"2006-06-14T18:59:27Z","receivedAt":"2006-06-14T18:59:27Z","isPatch":false,"sender":{"key":"keithp@keithp.com","avatar":"https://gravatar.com/avatar/fa1f479cdd51322fe86215c955a81d296bbf66a1fe625f8a12d87a8ec7faf648?d=mp&s=160"},"body":"On Wed, 2006-06-14 at 11:18 -0700, Linus Torvalds wrote:\n\n> You don't _need_ to shuffle. As mentioned, it will only affect the \n> location of the data in the pack-file, which in turn will mostly matter \n> as an IO pattern thing, not anything really fundamental.  If the pack-file \n> ends up caching well, the IO patterns obviously will never matter.\n\nOk, sounds like shuffling isn't necessary; the only benefit packing\ngains me is to reduce the size of each directory in the object store;\nthe process I follow is to construct blobs for every revision, then just\nuse the sha1 values to construct an index for each commit. I never\nactually look at the blobs myself, so IO access patterns aren't\nrelevant.\n\nRepacking after the import is completed should undo whatever horror show\nI've created in any case.\n\n-- \nkeith.packard@intel.com\n"},{"id":"21792","messageId":"Pine.LNX.4.64.0606141212190.5498@g5.osdl.org","threadId":"4505","inReplyTo":"1150311567.30681.28.camel@neko.keithp.com","subject":"Re: Repacking many disconnected blobs","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-14T19:18:18Z","receivedAt":"2006-06-14T19:18:18Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 14 Jun 2006, Keith Packard wrote:\n> \n> Ok, sounds like shuffling isn't necessary; the only benefit packing\n> gains me is to reduce the size of each directory in the object store;\n\nThere's actually a secondary benefit to packing that turned out to be much \nbigger from a performance standpoint: the size benefit coupled with the \nfact that it's all in one file ends up meaning that accessing packed \nobjects is _much_ faster than accessing individual files.\n\nThe Linux system call overhead is one of the lowest ones out there, but \nit's still much bigger than just a function call, and doing a full \npathname walk and open/close is bigger yet. In contrast, if you access \nlots of objects and they are all in a pack, you only end up doing one mmap \nand a page fault for each 4kB entry, and that's it.\n\nSo packing has a large performance benefit outside of the actual disk use \none, and to some degree that performance benefit is then further magnified \nby good locality (ie you get more effective objects per page fault), but \nin your case that locality issue is secondary.\n\nI assume that you never actually end up looking at the _contents_ of the \nobjects any more ever afterwards, because in a very real sense you're \nreally interested in the SHA1 names, right? All the latter phases of \nparsecvs will just use the SHA1 names directly, and never actually even \nopen the data (packed or not).\n\nSo in that sense, you only care about the disksize and a much improved \ndirectory walk from fewer files (until the repository has actually been \nfully created, at which point a repack will do the right thing).\n\n\t\t\tLinus\n"},{"id":"21793","messageId":"Pine.LNX.4.64.0606141514000.2703@localhost.localdomain","threadId":"4505","inReplyTo":"1150311567.30681.28.camel@neko.keithp.com","subject":"Re: Repacking many disconnected blobs","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2006-06-14T19:25:14Z","receivedAt":"2006-06-14T19:25:14Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 14 Jun 2006, Keith Packard wrote:\n\n> On Wed, 2006-06-14 at 11:18 -0700, Linus Torvalds wrote:\n> \n> > You don't _need_ to shuffle. As mentioned, it will only affect the \n> > location of the data in the pack-file, which in turn will mostly matter \n> > as an IO pattern thing, not anything really fundamental.  If the pack-file \n> > ends up caching well, the IO patterns obviously will never matter.\n> \n> Ok, sounds like shuffling isn't necessary; the only benefit packing\n> gains me is to reduce the size of each directory in the object store;\n> the process I follow is to construct blobs for every revision, then just\n> use the sha1 values to construct an index for each commit. I never\n> actually look at the blobs myself, so IO access patterns aren't\n> relevant.\n> \n> Repacking after the import is completed should undo whatever horror show\n> I've created in any case.\n\nThe only advantage of feeding object names from latest to oldest has to \ndo with the delta direction.  In doing so the delta are backward such \nthat objects with deeper delta chain are further back in history and \nthis is what you want in the final pack for faster access to the latest \nrevision.\n\nOf course the final repack will do that automatically, but only if you \nuse -a -f with git-repack.  But when -f is not provided then already \ndeltified objects from other packs are copied as is without any delta \ncomputation making the repack process lots faster.  In that case it \nmight be preferable that the reuse of already deltified data is made of \nbackward delta which is the reason you might consider feeding object in \nthe prefered order up front.\n\n\nNicolas\n"},{"id":"21802","messageId":"1150319115.30681.54.camel@neko.keithp.com","threadId":"4505","inReplyTo":"Pine.LNX.4.64.0606141514000.2703@localhost.localdomain","subject":"Re: Repacking many disconnected blobs","fromName":"Keith Packard","fromEmail":"keithp@keithp.com","sentAt":"2006-06-14T21:05:15Z","receivedAt":"2006-06-14T21:05:15Z","isPatch":false,"sender":{"key":"keithp@keithp.com","avatar":"https://gravatar.com/avatar/fa1f479cdd51322fe86215c955a81d296bbf66a1fe625f8a12d87a8ec7faf648?d=mp&s=160"},"body":"On Wed, 2006-06-14 at 15:25 -0400, Nicolas Pitre wrote:\n\n> The only advantage of feeding object names from latest to oldest has to \n> do with the delta direction.  In doing so the delta are backward such \n> that objects with deeper delta chain are further back in history and \n> this is what you want in the final pack for faster access to the latest \n> revision.\n\nOk, so I'm feeding them from latest to oldest along each branch, which\noptimizes only the 'master' branch, leaving other branches much further\ndown in the data file. That should mean repacking will help a lot for\nrepositories with many active branches.\n\n> In that case it \n> might be preferable that the reuse of already deltified data is made of \n> backward delta which is the reason you might consider feeding object in \n> the prefered order up front.\n\nHmm. As I'm deltafying along branches, the delta data should actually be\nfairly good; the only 'bad' result will be the sub-optimal object\nordering in the pack files. I'll experiment with some larger trees to\nsee how much additional savings the various repack options yield.\n\n-- \nkeith.packard@intel.com\n"},{"id":"21803","messageId":"Pine.LNX.4.64.0606141415540.5498@g5.osdl.org","threadId":"4505","inReplyTo":"1150319115.30681.54.camel@neko.keithp.com","subject":"Re: Repacking many disconnected blobs","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-14T21:17:27Z","receivedAt":"2006-06-14T21:17:27Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 14 Jun 2006, Keith Packard wrote:\n> \n> > In that case it \n> > might be preferable that the reuse of already deltified data is made of \n> > backward delta which is the reason you might consider feeding object in \n> > the prefered order up front.\n> \n> Hmm. As I'm deltafying along branches, the delta data should actually be\n> fairly good; the only 'bad' result will be the sub-optimal object\n> ordering in the pack files. I'll experiment with some larger trees to\n> see how much additional savings the various repack options yield.\n\nThe fact that git repacking sorts by filesize after it sorts by filename \nshould make this a non-issue: we always try to delta against the larger \nversion (where \"larger\" is not only almost invariable also \"newer\", but \nthe delta is simpler, since deleting data doesn't take up any space in the \ndelta, while adding data needs to ay what the data added was, of course).\n\n\t\tLinus\n"},{"id":"21804","messageId":"Pine.LNX.4.64.0606141716460.2703@localhost.localdomain","threadId":"4505","inReplyTo":"1150319115.30681.54.camel@neko.keithp.com","subject":"Re: Repacking many disconnected blobs","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2006-06-14T21:20:40Z","receivedAt":"2006-06-14T21:20:40Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 14 Jun 2006, Keith Packard wrote:\n\n> Hmm. As I'm deltafying along branches, the delta data should actually be\n> fairly good; the only 'bad' result will be the sub-optimal object\n> ordering in the pack files. I'll experiment with some larger trees to\n> see how much additional savings the various repack options yield.\n\nNote that the object list order is unlikely to affect pack size.  It is \nreally about optimizing the pack layout for subsequent access to it.\n\n\nNicolas\n"}]}