{"thread":{"id":"26353","subject":"[RFC] Add --create-cache to repack","startedAt":"2011-01-28T08:06:24Z","lastAt":"2011-01-31T21:48:08Z","messageCount":26,"participants":["Shawn O. Pearce","Johannes Sixt","Shawn Pearce","Nicolas Pitre","Jay Soffian","Junio C Hamano","A Large Angry SCM"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"159968","messageId":"1296201984-24426-1-git-send-email-spearce@spearce.org","threadId":"26353","inReplyTo":null,"subject":"[RFC] Add --create-cache to repack","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-01-28T08:06:24Z","receivedAt":"2011-01-28T08:06:24Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"A cache pack is all objects reachable from a single commit that is\npart of the project's stable history and won't disappear, and is\naccessible to all readers of the repository.  By containing only that\ncommit and its contents, if the commit is reached from a reference we\nknow immediately that the entire pack is also reachable.  To help\nensure this is true, the --create-cache flag looks for a commit along\nrefs/heads and refs/tags that is at least 1 month old, working under\nthe assumption that a commit this old won't be rebased or pruned.\n\nDuring a clone request if a commit is discovered that matches the\ncache pack, all newer objects can be enumerated using normal rules and\nsent to the client, and then the cache pack can be simply appended\nonto the end of the stream.  There is no need to enumerate the objects\nas the object count is in the header of the cache pack.  There is no\nneed to allocate all of the objects in the pack-objects process, which\nreduces its working set size, and its impact on busy servers.\n\nBy keeping the pack with a standard .keep file, later repacks of the\nrepository won't include these objects, which permits disk usage to\nstay within a reasonable factor of the repository size.\n\nBecause newer packed objects are not delta compressed against the\nolder cached pack, clients may receive a larger data transfer when the\ncached pack is simply appended onto the stream.  pack-objects could\nwork around this by constructing a thin pack, and adding the cache\npack's tip commit as the uninteresting/common base for the thin pack.\nThe references for the newer objects will point to older data behind\nthem so they will automatically use the larger REF_DELTA format.\n\nThis commit only adds the logic to git-repack to construct the cached\npack.  For example on a Linux kernel repository:\n\n  # Construct the initial cache pack\n  $ git repack --create-cache --cache-include=v2.6.11-tree\n\n  # Remove duplicated objects\n  $ git repack -a -d\n\nIf this is actually a good idea, pack-objects can later learn how to\nuse $GIT_DIR/objects/info/cached during revision traversal to know\nwhen a cached pack is found, and switch to the thin pack + cached pack\ntransfer method described above.\n\nThe cached pack is only useful for initial clones of a repository, and\nonly if object enumeration takes more than a few seconds.  However\ninitial clones of big projects like linux-2.6.git are killing some\ncommon mirror sites, so this could be one way to help them out.\n\nLater fetch-pack/upload-pack protocol could learn how to more\nintelligently use the cached pack in the data stream, allowing a\nclient whose connection has been broken to resume with a byte range\nrequest within the cached pack, assuming the pack is still present on\nthe server.  This can be validated by giving the client both the SHA-1\npack name, and the SHA-1 trailer of the pack content, and requiring\nthese to match on a byte range request.\n\nRepository owners may also enjoy having the cached pack as frequent\n`git gc` invocations will now have lower IO and CPU requirements due\nto the large pack having a .keep file.  In the future `git gc --auto`\ncould learn to suggest removing the .keep file and regenerating the\ncached pack once there is sufficient new content to make creating a\nnew pack worthwhile.\n---\n git-repack.sh |   57 +++++++++++++++++++++++++++++++++++++++++++++++++++------\n 1 files changed, 51 insertions(+), 6 deletions(-)\n\ndiff --git a/git-repack.sh b/git-repack.sh\nindex 624feec..7a7984c 100755\n--- a/git-repack.sh\n+++ b/git-repack.sh\n@@ -15,6 +15,9 @@ F               pass --no-reuse-object to git-pack-objects\n n               do not run git-update-server-info\n q,quiet         be quiet\n l               pass --local to git-pack-objects\n+create-cache    create a cached pack for older history\n+cache-include=  other objects to include in the cache\n+cache-age=      how old to start caching from\n  Packing constraints\n window=         size of the window used for delta compression\n window-memory=  same as the above, but limit memory size instead of entries count\n@@ -26,6 +29,7 @@ SUBDIRECTORY_OK='Yes'\n \n no_update_info= all_into_one= remove_redundant= unpack_unreachable=\n local= no_reuse= extra=\n+create_cache= cache_include= cache_age=1.month.ago\n while test $# != 0\n do\n \tcase \"$1\" in\n@@ -38,6 +42,11 @@ do\n \t-f)\tno_reuse=--no-reuse-delta ;;\n \t-F)\tno_reuse=--no-reuse-object ;;\n \t-l)\tlocal=--local ;;\n+\t--create-cache) create_cache=t ;;\n+\t--cache-age) cache_age=$2; shift ;;\n+\t--cache-include)\n+\t\tname=$(git rev-parse --verify $2)\n+\t\tcache_include=\"$cache_include $name\"; shift ;;\n \t--max-pack-size|--window|--window-memory|--depth)\n \t\textra=\"$extra $1=$2\"; shift ;;\n \t--) shift; break;;\n@@ -52,16 +61,19 @@ true)\n esac\n \n PACKDIR=\"$GIT_OBJECT_DIRECTORY/pack\"\n+INFODIR=\"$GIT_OBJECT_DIRECTORY/info\"\n PACKTMP=\"$PACKDIR/.tmp-$$-pack\"\n rm -f \"$PACKTMP\"-*\n trap 'rm -f \"$PACKTMP\"-*' 0 1 2 3 15\n \n # There will be more repacking strategies to come...\n-case \",$all_into_one,\" in\n-,,)\n+case \",$create_cache,$all_into_one,\" in\n+,t,,)\n+\t;;\n+,,,)\n \targs='--unpacked --incremental'\n \t;;\n-,t,)\n+,,t,)\n \targs= existing=\n \tif [ -d \"$PACKDIR\" ]; then\n \t\tfor e in `cd \"$PACKDIR\" && find . -type f -name '*.pack' \\\n@@ -84,9 +96,22 @@ esac\n \n mkdir -p \"$PACKDIR\" || exit\n \n-args=\"$args $local ${GIT_QUIET:+-q} $no_reuse$extra\"\n-names=$(git pack-objects --keep-true-parents --honor-pack-keep --non-empty --all --reflog $args </dev/null \"$PACKTMP\") ||\n-\texit 1\n+if [ -n \"$create_cache\" ]; then\n+\troot=$(git rev-list -n 1 --until=$cache_age --branches --tags --)\n+\targs=\"$args ${GIT_QUIET:+-q} $no_reuse$extra\"\n+\tnames=$( ( echo \"$root\";\n+\t\t       for name in $cache_include\n+\t\t       do\n+\t\t         echo \"$name\"\n+\t\t       done ) |\n+\t\tgit pack-objects --keep-true-parents --non-empty $args --revs \\\n+\t\t\"$PACKTMP\") ||\n+\t\texit 1\n+else\n+\targs=\"$args $local ${GIT_QUIET:+-q} $no_reuse$extra\"\n+\tnames=$(git pack-objects --keep-true-parents --honor-pack-keep --non-empty --all --reflog $args </dev/null \"$PACKTMP\") ||\n+\t\texit 1\n+fi\n if [ -z \"$names\" ]; then\n \tsay Nothing new to pack.\n fi\n@@ -151,6 +176,10 @@ do\n \tmv -f \"$PACKTMP-$name.pack\" \"$PACKDIR/pack-$name.pack\" &&\n \tmv -f \"$PACKTMP-$name.idx\"  \"$PACKDIR/pack-$name.idx\" ||\n \texit\n+\n+\tif [ -n \"$create_cache\" ]; then\n+\t\techo \"cache $root$cache_include\" >\"$PACKDIR/pack-$name.keep\"\n+\tfi\n done\n \n # Remove the \"old-\" files\n@@ -162,6 +191,22 @@ done\n \n # End of pack replacement.\n \n+# Update the cache list\n+if [ -n \"$create_cache\" ]; then\n+\tmkdir -p \"$INFODIR\" || exit\n+\t( echo \"+ $root\" &&\n+\t  for name in $cache_include\n+\t  do\n+\t    echo \"+ $name\"\n+\t  done\n+\t  for name in $names\n+\t  do\n+\t    echo \"P $name\"\n+\t  done ) >\"$INFODIR/cached\"\n+\techo \"Cached from:\"\n+\tgit log --pretty=format:'  [%h] %cd%n  %s' -1 \"$root\" --\n+fi\n+\n if test \"$remove_redundant\" = t\n then\n \t# We know $existing are all redundant.\n-- \n1.7.4.rc1.253.gb7420\n"},{"id":"159970","messageId":"4D42878E.2020502@viscovery.net","threadId":"26353","inReplyTo":"1296201984-24426-1-git-send-email-spearce@spearce.org","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Johannes Sixt","fromEmail":"j.sixt@viscovery.net","sentAt":"2011-01-28T09:08:30Z","receivedAt":"2011-01-28T09:08:30Z","isPatch":false,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 1/28/2011 9:06, schrieb Shawn O. Pearce:\n> A cache pack is all objects reachable from a single commit that is\n> part of the project's stable history and won't disappear, and is\n> accessible to all readers of the repository.  By containing only that\n> commit and its contents, if the commit is reached from a reference we\n> know immediately that the entire pack is also reachable.  To help\n> ensure this is true, the --create-cache flag looks for a commit along\n> refs/heads and refs/tags that is at least 1 month old, working under\n> the assumption that a commit this old won't be rebased or pruned.\n\nIn one of my repositories, I have two stable branches and a good score of\ntopic branches of various ages (a few hours up to two years 8). The topic\nbranches will either be dropped eventually, or rebased.\n\nWhat are the odds that this choice of a tip commit picks one that is in a\ntopic branch? Or is there no point in using --create-cache in a repository\nlike this?\n\n-- Hannes\n"},{"id":"159983","messageId":"AANLkTim+AUY9SdeAFfkny2_a3qQ9SCDLUHR3s9Q3M98u@mail.gmail.com","threadId":"26353","inReplyTo":"4D42878E.2020502@viscovery.net","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-01-28T14:37:22Z","receivedAt":"2011-01-28T14:37:22Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Fri, Jan 28, 2011 at 01:08, Johannes Sixt <j.sixt@viscovery.net> wrote:\n> Am 1/28/2011 9:06, schrieb Shawn O. Pearce:\n>> A cache pack is all objects reachable from a single commit that is\n>> part of the project's stable history and won't disappear, and is\n>> accessible to all readers of the repository.  By containing only that\n>> commit and its contents, if the commit is reached from a reference we\n>> know immediately that the entire pack is also reachable.  To help\n>> ensure this is true, the --create-cache flag looks for a commit along\n>> refs/heads and refs/tags that is at least 1 month old, working under\n>> the assumption that a commit this old won't be rebased or pruned.\n>\n> In one of my repositories, I have two stable branches and a good score of\n> topic branches of various ages (a few hours up to two years 8). The topic\n> branches will either be dropped eventually, or rebased.\n>\n> What are the odds that this choice of a tip commit picks one that is in a\n> topic branch? Or is there no point in using --create-cache in a repository\n> like this?\n\nArgh, you are right.  Its quite likely this would pick a topic\nbranch... and that isn't really what is desired.\n\nMy original concept here was for distribution point repositories,\nwhich are less likely to have these topic branches that will rebase\nand disappear.  Though git.git has one called \"pu\".  *sigh*\n\nA simple fix is to use --heads --tags by default like I do here, but\nmake the actual parameters we feed to rev-list configurable.  A\nrepository owner could select only the master branch as input to\nrev-list, making it less likely the topic branches would be\nconsidered.  Unfortunately that requires direct access to the\nrepository.  It fails for a site like GitHub, where you don't manage\nthe repository at all.\n\ngit.git also is problematic because of the man, html and todo\nbranches.  Branches that are disconnected from the main history but\nare very small (e.g. todo) might be selected instead and create a\nnearly useless cache file.  Fortunately disconnected branches could\neach have their own cache file (with only the inode overhead of having\nan additional 3 files per disconnected branch), and pack-objects could\nconcat all of those packs together when sending.  Its just a challenge\nto identify these branches and keep them from being used for that main\nproject pack.\n\n\nThis started because I was looking for a way to speed up clones coming\nfrom a JGit server.  Cloning the linux-2.6 repository is painful, it\ntakes a long time to enumerate the 1.8 million objects.  So I tried\nadding a cached list of objects reachable from a given commit, which\nspeeds up the enumeration phase, but JGit still needs to allocate all\nof the working set to track those objects, then go find them in packs\nand slice out each compressed form and reformat the headers on the\nwire.  Its a lot of redundant work when your kernel repository has\n360MB of data that you know a client needs if they have asked for your\nmaster branch with no \"have\" set.\n\nLater I realized, we can get rid of that cached list of objects and\njust use the pack itself.  Its far cleaner, as there is no redundant\ncache.  But either way (object list or pack) its a bit of a challenge\nto automatically identify the right starting points to use.  Linus\nTorvalds' linux-2.6 repository is the perfect case for the RFC I\nposted, its one branch with all of the history, and it never rewinds.\nBut maybe Linus is just very unique in this world.  :-)\n\n-- \nShawn.\n"},{"id":"159988","messageId":"4D42E1E3.4060808@viscovery.net","threadId":"26353","inReplyTo":"AANLkTim+AUY9SdeAFfkny2_a3qQ9SCDLUHR3s9Q3M98u@mail.gmail.com","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Johannes Sixt","fromEmail":"j.sixt@viscovery.net","sentAt":"2011-01-28T15:33:55Z","receivedAt":"2011-01-28T15:33:55Z","isPatch":false,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 1/28/2011 15:37, schrieb Shawn Pearce:\n> A simple fix is to use --heads --tags by default like I do here, but\n> make the actual parameters we feed to rev-list configurable.  A\n> repository owner could select only the master branch as input to\n> rev-list, making it less likely the topic branches would be\n> considered.  Unfortunately that requires direct access to the\n> repository.  It fails for a site like GitHub, where you don't manage\n> the repository at all.\n\nLet's define a ref hierarchy, refs/cache-pack, that names the cache pack\ntips. A cache pack would be generated for each ref found in that\nhierarchy. Then these commits are under user control even on github,\nbecause you can just push the refs. Junio would perhaps choose a release\ntag, and corresponding commits in the man and html histories. The choice\nwould not be completely automatic, though.\n\n-- Hannes\n"},{"id":"159998","messageId":"AANLkTi=Mu+tP7V2wWq9Fr3RVmdWhG7gGhj=LWLSqerZ_@mail.gmail.com","threadId":"26353","inReplyTo":"4D42E1E3.4060808@viscovery.net","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-01-28T18:22:17Z","receivedAt":"2011-01-28T18:22:17Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Fri, Jan 28, 2011 at 07:33, Johannes Sixt <j.sixt@viscovery.net> wrote:\n> Am 1/28/2011 15:37, schrieb Shawn Pearce:\n>> A simple fix is to use --heads --tags by default like I do here, but\n>> make the actual parameters we feed to rev-list configurable.  A\n>> repository owner could select only the master branch as input to\n>> rev-list, making it less likely the topic branches would be\n>> considered.  Unfortunately that requires direct access to the\n>> repository.  It fails for a site like GitHub, where you don't manage\n>> the repository at all.\n>\n> Let's define a ref hierarchy, refs/cache-pack, that names the cache pack\n> tips. A cache pack would be generated for each ref found in that\n> hierarchy. Then these commits are under user control even on github,\n> because you can just push the refs. Junio would perhaps choose a release\n> tag, and corresponding commits in the man and html histories. The choice\n> would not be completely automatic, though.\n\nThis is a good idea.  Perhaps we go slightly further and say:\n\n  refs/cache-pack/name-without-slash\n\n    This packs into its own pack file, as a single tip.\n\n  refs/cache-pack/group/a\n  refs/cache-pack/group/b\n\n   These pack into a pack file together.\n\nIf you have direct repository access, you can also just make one of\nthese a symbolic reference to a branch, e.g. refs/heads/master, and\nthen periodic `git repack --create-cache` invocations would pick up\nthe latest point.\n\n-- \nShawn.\n"},{"id":"159999","messageId":"alpine.LFD.2.00.1101281304270.8580@xanadu.home","threadId":"26353","inReplyTo":"AANLkTim+AUY9SdeAFfkny2_a3qQ9SCDLUHR3s9Q3M98u@mail.gmail.com","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2011-01-28T18:46:33Z","receivedAt":"2011-01-28T18:46:33Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 28 Jan 2011, Shawn Pearce wrote:\n\n> This started because I was looking for a way to speed up clones coming\n> from a JGit server.  Cloning the linux-2.6 repository is painful, it\n> takes a long time to enumerate the 1.8 million objects.  So I tried\n> adding a cached list of objects reachable from a given commit, which\n> speeds up the enumeration phase, but JGit still needs to allocate all\n> of the working set to track those objects, then go find them in packs\n> and slice out each compressed form and reformat the headers on the\n> wire.  Its a lot of redundant work when your kernel repository has\n> 360MB of data that you know a client needs if they have asked for your\n> master branch with no \"have\" set.\n> \n> Later I realized, we can get rid of that cached list of objects and\n> just use the pack itself.  Its far cleaner, as there is no redundant\n> cache.  But either way (object list or pack) its a bit of a challenge\n> to automatically identify the right starting points to use.  Linus\n> Torvalds' linux-2.6 repository is the perfect case for the RFC I\n> posted, its one branch with all of the history, and it never rewinds.\n> But maybe Linus is just very unique in this world.  :-)\n\nPlaying my old record again... I know.  But pack v4 should solve a big \npart of this enumeration cost.\n\nI've changed the format slightly again in my WIP branch.  The idea is to:\n\n1) Have a non compressed yet still really dense representation for tree \n   objects;\n\n2) do the same thing for the first part of commit objects, and only \n   deflate the free form text part.\n\nThere is nothing new here.  However, it should be possible to:\n\n3) replace all SHA1 references by an offset into the pack file directly, \n   just like we do for OFS_DELTA objects.  If the SHA1 is actually \n   needed then we can obtain it with a reverse lookup with given object offset \n   in the pack index file, but in practice that is not actually required that \n   often.\n\nSo walking the history graph and enumerating objects would require \nnothing more than simply following straight pointers in the pack data in \n99% of the cases.  No object decompression, no memory buffer \nallocation/deallocation to perform that decompression, no string parsing \nin the tree object case, etc. Only cross pack references would require a \nfull SHA1 based lookup like we do now.\n\nI still have to sit down and figure out the implications of this, \nespecially with forward references, meaning that the offset might have \nto be an object index so to allow for variable length encoding, and also \nto make sure index-pack can reconstruct the pack index.  But that would \nonly be an indirect lookup which shouldn't be significantly costly.\n\nSo that's the idea.  Keep the exact same functionality as we have now, \nwithout any need for cache management, but making the data structure in \na form that should improve object enumeration by some magnitude.\n\n\nNicolas\n"},{"id":"160002","messageId":"AANLkTi=f34Q2VUrzA0dEG0KCFcHcd_Yq=UN6RSDPVS+p@mail.gmail.com","threadId":"26353","inReplyTo":"4D42E1E3.4060808@viscovery.net","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Jay Soffian","fromEmail":"jaysoffian@gmail.com","sentAt":"2011-01-28T19:15:30Z","receivedAt":"2011-01-28T19:15:30Z","isPatch":false,"sender":{"key":"jaysoffian@gmail.com","avatar":"https://avatars.githubusercontent.com/u/155970?v=4"},"body":"On Fri, Jan 28, 2011 at 10:33 AM, Johannes Sixt <j.sixt@viscovery.net> wrote:\n> Let's define a ref hierarchy, refs/cache-pack, that names the cache pack\n> tips. A cache pack would be generated for each ref found in that\n> hierarchy. Then these commits are under user control even on github,\n> because you can just push the refs. Junio would perhaps choose a release\n> tag, and corresponding commits in the man and html histories. The choice\n> would not be completely automatic, though.\n\nThis is just for bare repos, right? Why not just use HEAD?\n\nj.\n"},{"id":"160001","messageId":"AANLkTikPcp5CUTWfhy6FYbCEkNG6epGBAMNT5vTfSbvy@mail.gmail.com","threadId":"26353","inReplyTo":"alpine.LFD.2.00.1101281304270.8580@xanadu.home","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-01-28T19:15:34Z","receivedAt":"2011-01-28T19:15:34Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Fri, Jan 28, 2011 at 10:46, Nicolas Pitre <nico@fluxnic.net> wrote:\n> On Fri, 28 Jan 2011, Shawn Pearce wrote:\n>\n>> This started because I was looking for a way to speed up clones coming\n>> from a JGit server.  Cloning the linux-2.6 repository is painful,\n...\n>> Later I realized, we can get rid of that cached list of objects and\n>> just use the pack itself.\n...\n> Playing my old record again... I know.  But pack v4 should solve a big\n> part of this enumeration cost.\n\nI've said the same thing for years myself.  As much as it would be\nnice to fix some of the decompression costs with pack v2/v3, v2/v3 is\nvery common in the wild, and a new pack encoding is going to be a\nfairly complex thing to get added to C Git.  And pack v4 doesn't\neliminate the enumeration, it just makes it faster.\n\n> So that's the idea.  Keep the exact same functionality as we have now,\n> without any need for cache management, but making the data structure in\n> a form that should improve object enumeration by some magnitude.\n\nThat's what I also liked about my --create-cache flag.  Its keeping\nthe same data we already have, in the same format we already have it\nin.  We're just making a more explicit statement that everything in\nsome pack is about as tightly compressed as it ever would be for a\nclient, and it isn't going to change anytime soon.  Thus we might as\nwell tag it with .keep to prevent repack of mucking with it, and we\ncan take advantage of this to serve the pack to clients very fast.\n\nOver breakfast this morning I made the point to Junio that with the\ncached pack and a slight network protocol change (enabled by a\ncapability of course) we could stop using pkt-line framing when\nsending the cached pack part of the stream, and just send the pack\ndirectly down the socket.  That changes the clone of a 400 MB project\nlike linux-2.6 from being a lot of user space stuff, to just being a\nsendfile() call for the bulk of the content.  I think we can just hand\noff the major streaming to the kernel.  (Part of the protocol change\nis we would need to use multiple SHA-1 checksums in the stream, so we\ndon't have to re-checksum the existing cached pack.)\n\n\nI love the idea of some of the concepts in pack v4.  I really do.  But\nthis sounds a lot simpler to implement, and it lets us completely\neliminate a massive amount of server processing (even under pack v4\nyou still have object enumeration), in exchange for what might be a\nfew extra MBs on the wire to the client due to slightly less good\ndeltas and the use of REF_DELTA in the thin pack used for the most\nrecent objects.  I don't envision this being used on projects smaller\nthan git.git itself, if you can gc --aggressive the whole thing in a\nminute the cached pack is probably pointless.  But if you have 400+\nMB, you want that to be network bound, and have almost no CPU impact\non the server.\n\nPlus we can safely do byte range requests for resumable clone within\nthe cached pack part of the stream.  And when pack v4 comes along, we\ncan use this same strategy for an equally large pack v4 pack.\n\n-- \nShawn.\n"},{"id":"160003","messageId":"AANLkTi=ZAPbp-qAtk=bCoi4=uEbu-F6w-8TfStXfstEL@mail.gmail.com","threadId":"26353","inReplyTo":"AANLkTi=f34Q2VUrzA0dEG0KCFcHcd_Yq=UN6RSDPVS+p@mail.gmail.com","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-01-28T19:19:40Z","receivedAt":"2011-01-28T19:19:40Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Fri, Jan 28, 2011 at 11:15, Jay Soffian <jaysoffian@gmail.com> wrote:\n> On Fri, Jan 28, 2011 at 10:33 AM, Johannes Sixt <j.sixt@viscovery.net> wrote:\n>> Let's define a ref hierarchy, refs/cache-pack, that names the cache pack\n>> tips. A cache pack would be generated for each ref found in that\n>> hierarchy. Then these commits are under user control even on github,\n>> because you can just push the refs. Junio would perhaps choose a release\n>> tag, and corresponding commits in the man and html histories. The choice\n>> would not be completely automatic, though.\n>\n> This is just for bare repos, right? Why not just use HEAD?\n\nEven on a bare repository a user might rewind his/her HEAD frequently.\n Caching from today's HEAD might not be ideal if you are about to\nrewrite the last 10 commits and push those again to the repository.\nThat's actually where the \"1.month.ago\" guess came from in the patch.\nIf we go back a little in history, the odds of a rewrite are reduced,\nand we're more likely to be able to reuse this pack.\n\nHEAD - X commits/X days might be a good approximation if there are no\nrefs/cache-pack *and* gc --auto notices there is \"enough\" content to\nsuggest creating a cached pack.  But I do like Johannes Sixt's\nrefs/cache-pack ref hierarchy as a way to configure this explicitly.\n\n-- \nShawn.\n"},{"id":"160008","messageId":"alpine.LFD.2.00.1101281502170.8580@xanadu.home","threadId":"26353","inReplyTo":"AANLkTikPcp5CUTWfhy6FYbCEkNG6epGBAMNT5vTfSbvy@mail.gmail.com","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2011-01-28T21:09:30Z","receivedAt":"2011-01-28T21:09:30Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 28 Jan 2011, Shawn Pearce wrote:\n\n> On Fri, Jan 28, 2011 at 10:46, Nicolas Pitre <nico@fluxnic.net> wrote:\n> > On Fri, 28 Jan 2011, Shawn Pearce wrote:\n> >\n> >> This started because I was looking for a way to speed up clones coming\n> >> from a JGit server.  Cloning the linux-2.6 repository is painful,\n> ...\n> >> Later I realized, we can get rid of that cached list of objects and\n> >> just use the pack itself.\n> ...\n> > Playing my old record again... I know.  But pack v4 should solve a big\n> > part of this enumeration cost.\n> \n> I've said the same thing for years myself.  As much as it would be\n> nice to fix some of the decompression costs with pack v2/v3, v2/v3 is\n> very common in the wild, and a new pack encoding is going to be a\n> fairly complex thing to get added to C Git.  And pack v4 doesn't\n> eliminate the enumeration, it just makes it faster.\n\nWell, you don't necessarily need pack v4 to be widely deployed for \npeople to benefit from it.  If it is available on servers such as \ngit.kernel.org then everybody will see their clone requests go faster.  \nSame principle as for the cache packs.\n\nAnd yes it doesn't eliminate the enumeration, but you can't eliminate it \nentirely either as many other operations do require object enumeration \ntoo, and those would be sped up as well.\n\nBut this is in fact orthogonal to the cache pack concept indeed.\n\n> That's what I also liked about my --create-cache flag.  Its keeping\n> the same data we already have, in the same format we already have it\n> in.  We're just making a more explicit statement that everything in\n> some pack is about as tightly compressed as it ever would be for a\n> client, and it isn't going to change anytime soon.  Thus we might as\n> well tag it with .keep to prevent repack of mucking with it, and we\n> can take advantage of this to serve the pack to clients very fast.\n\nI do agree on that point.   And I like it too.  However I'd prefer if \nthe whole thing wasn't created \"automatically\".  It's probably best if \nthe repository administrator decides explicitly what should go in such \ncached packs according to actual purpose and usage for good commit \nthresholds and branches.  Only a human can make that decision.\n\nI'd also recommend _not_ using the ref namespace for that.  Let's not \nmix up branching/tagging with what is effectively a storage \nimplementation issue. Linking the ref namespace with the actual packs \nthey refer to would be highly inelegant if the SHA1 of the pack has to \nbe part of the ref name.  Instead, I'd suggest simply listing all the \ncommit tips a cache pack contains in the .keep file directly instead.  \nThat would make it much easier to use with the object alternates too as \nthe alternate mechanism points to the object store of a foreign repo and \nnot to its refs.\n\n> Over breakfast this morning I made the point to Junio that with the\n> cached pack and a slight network protocol change (enabled by a\n> capability of course) we could stop using pkt-line framing when\n> sending the cached pack part of the stream, and just send the pack\n> directly down the socket.  That changes the clone of a 400 MB project\n> like linux-2.6 from being a lot of user space stuff, to just being a\n> sendfile() call for the bulk of the content.  I think we can just hand\n> off the major streaming to the kernel. \n\nWhile this might look like a good idea in theory, did you actually \nprofile it to see if that would make a noticeable difference?  The \npkt-line framing allows for asynchronous messages to be sent over a \nsideband, which you wouldn't be able to do anymore until the full 400 MB \nis received by the remote side.  Without concrete performance numbers \nI'm not convinced it is worth the maintenance cost for creating a \ndeviation in the protocol like this.\n\n> (Part of the protocol change\n> is we would need to use multiple SHA-1 checksums in the stream, so we\n> don't have to re-checksum the existing cached pack.)\n\n?? I don't follow you here.\n\n> I love the idea of some of the concepts in pack v4.  I really do.  But\n> this sounds a lot simpler to implement, and it lets us completely\n> eliminate a massive amount of server processing (even under pack v4\n> you still have object enumeration), in exchange for what might be a\n> few extra MBs on the wire to the client due to slightly less good\n> deltas and the use of REF_DELTA in the thin pack used for the most\n> recent objects.\n\nI agree.  And what I personally like the most is the fact that this can \nbe made transparent to clients using the existing network protocol \nunchanged.\n\n> Plus we can safely do byte range requests for resumable clone within\n> the cached pack part of the stream.\n\nThat part I'm not sure of.  We are still facing the same old issues \nhere, as some mirrors might have the same commit edges for a cache pack \nbut not necessarily the same packing result, etc.  So I'd keep that out \nof the picture for now.  The idea of being able to resume the transfer \nof a cache pack is good, however I'd make it into a totally separate \nservice outside git-upload-pack where the issue of validating and \nupdating content on both sides can be done efficiently without impacting \nthe upload-pack protocol.  There would be more than just the cache pack \nin play during a typical clone.\n\n> And when pack v4 comes along, we\n> can use this same strategy for an equally large pack v4 pack.\n\nAbsolutely.\n\n\nNicolas\n"},{"id":"160014","messageId":"AANLkTi=U7qRRij=BQXC1Goqa9toDFfaVKT=+-8zYxCcc@mail.gmail.com","threadId":"26353","inReplyTo":"alpine.LFD.2.00.1101281502170.8580@xanadu.home","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-01-29T01:32:01Z","receivedAt":"2011-01-29T01:32:01Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Fri, Jan 28, 2011 at 13:09, Nicolas Pitre <nico@fluxnic.net> wrote:\n> On Fri, 28 Jan 2011, Shawn Pearce wrote:\n>\n>> On Fri, Jan 28, 2011 at 10:46, Nicolas Pitre <nico@fluxnic.net> wrote:\n>> > On Fri, 28 Jan 2011, Shawn Pearce wrote:\n>> >\n>> >> This started because I was looking for a way to speed up clones coming\n>> >> from a JGit server.  Cloning the linux-2.6 repository is painful,\n\nWell, scratch the idea in this thread.  I think.\n\nI retested JGit vs. CGit on an identical linux-2.6 repository.  The\nrepository was fully packed, but had two pack files.  362M and 57M,\nand was created by packing a 1 month old master, marking it .keep, and\nthen repacking -a -d to get most recent last month into another pack.\nThis results in some files that should be delta compressed together\nbeing stored whole in the two packs (obviously).\n\nThe two implementations take the same amount of time to generate the\nclone.  3m28s / 3m22s for JGit, 3m23s for C Git.  The JGit created\npack is actually smaller 376.30 MiB vs. C Git's 380.59 MiB.  I point\nout this data because improvements made to JGit may show similar\nimprovements to CGit given how close they are in running time.\n\nI fully implemented the reuse of a cached pack behind a thin pack idea\nI was trying to describe in this thread.  It saved 1m7s off the JGit\nrunning time, but increased the data transfer by 25 MiB.  I didn't\nexpect this much of an increase, I honestly expected the thin pack\nportion to be well, thinner.  The issue is the thin pack cannot delta\nagainst all of the history, its only delta compressing against the tip\nof the cached pack.  So long-lived side branches that forked off an\nolder part of the history aren't delta compressing well, or at all,\nand that is significantly bloating the thin pack.  (Its also why that\n\"newer\" pack is 57M, but should be 14M if correctly combined with the\ncached pack.)  If I were to consider all of the objects in the cached\npack as potential delta base candidates for the thin pack, the entire\nbenefit of the cached pack disappears.\n\n\nWhich leaves me with dropping this idea.  I started it because I was\nactually looking for a way to speed up JGit.  But we're already\nroughly on-par with CGit performance.  Dropping 1m7s on a clone is\ngreat, but not at the expense of 6.5% larger network transfer.  For\nmost clients, 25 MiB of additional data transfer may be much more\nsignificant time than 1m7s saved doing server-side computation.\n\n>> That's what I also liked about my --create-cache flag.\n>\n> I do agree on that point.   And I like it too.\n\nI'm not sure I like it so much anymore.  :-)\n\nThe idea was half-baked, and came at the end of a long day, and after\nputting my cranky infant son down to sleep way past his normal bed\ntime.  I claim I was a sleep deprived new parent who wasn't thinking\nthings through enough before writing an email to git@vger.\n\n>> sendfile() call for the bulk of the content.  I think we can just hand\n>> off the major streaming to the kernel.\n>\n> While this might look like a good idea in theory, did you actually\n> profile it to see if that would make a noticeable difference?  The\n> pkt-line framing allows for asynchronous messages to be sent over a\n> sideband,\n\nNo, of course not.  The pkt-line framing is pretty low overhead, but\ncopying kernel buffer to userspace back to kernel buffer sort of sucks\nfor 400 MiB of data.  sendfile() on 400 MiB to a network socket is\nmuch easier when its all kernel space.  I figured, if it all worked\nout already to just dump the pack to the wire as-is, then we probably\nshould also try to go for broke and reduce the userspace copying.  It\nmight not matter to your desktop, but ask John Hawley (CC'd) about\nkernel.org and the git traffic volume he is serving.  They are doing\nmore than 1 million git:// requests per day now.\n\n>> Plus we can safely do byte range requests for resumable clone within\n>> the cached pack part of the stream.\n>\n> That part I'm not sure of.  We are still facing the same old issues\n> here, as some mirrors might have the same commit edges for a cache pack\n> but not necessarily the same packing result, etc.  So I'd keep that out\n> of the picture for now.\n\nI don't think its that hard.  If we modify the transfer protocol to\nallow the server to denote boundaries between packs, the server can\nsend the pack name (as in pack-$name.pack) and the pack SHA-1 trailer\nto the client.  A client asking for resume of a cached pack presents\nits original want list, these two SHA-1s, and the byte offset he wants\nto restart from.  The server validates the want set is still\nreachable, that the cached pack exists, and that the cached pack tips\nare reachable from current refs.  If all of that is true, it validates\nthe trailing SHA-1 in the pack matches what the client gave it.  If\nthat matches, it should be OK to resume transfer from where the client\nasked for.\n\nThen its up to the server administrators of a round-robin serving\ncluster to ensure that the same cached pack is available on all nodes,\nso that a resuming client is likely to have his request succeed.  This\nisn't impossible.  If the server operator cares they can keep the\nprior cached pack for several weeks after creating a newer cached\npack, giving clients plenty of time to resume a broken clone.  Disk is\nfairly inexpensive these days.\n\nBut its perhaps pointless, see above.  :-)\n\n-- \nShawn.\n"},{"id":"160016","messageId":"AANLkTimuW-7D4YA2jeF+y4DPE=CdqtL713MQK+1Gtp-d@mail.gmail.com","threadId":"26353","inReplyTo":"AANLkTi=U7qRRij=BQXC1Goqa9toDFfaVKT=+-8zYxCcc@mail.gmail.com","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-01-29T02:34:13Z","receivedAt":"2011-01-29T02:34:13Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Fri, Jan 28, 2011 at 17:32, Shawn Pearce <spearce@spearce.org> wrote:\n>\n> Well, scratch the idea in this thread.  I think.\n>\n> I retested JGit vs. CGit on an identical linux-2.6 repository.  The\n> repository was fully packed, but had two pack files.  362M and 57M,\n> and was created by packing a 1 month old master, marking it .keep, and\n> then repacking -a -d to get most recent last month into another pack.\n> This results in some files that should be delta compressed together\n> being stored whole in the two packs (obviously).\n>\n> The two implementations take the same amount of time to generate the\n> clone.  3m28s / 3m22s for JGit, 3m23s for C Git.  The JGit created\n> pack is actually smaller 376.30 MiB vs. C Git's 380.59 MiB.\n\nI just tried caching only the object list of what is reachable from a\nparticular commit.  The file is a small 20 byte header:\n\n  4 byte magic\n  4 byte version\n  4 byte number of commits (C)\n  4 byte number of trees (T)\n  4 byte number of blobs (B)\n\nThen C commit SHA-1s, followed by T tree SHA-1 + 4 byte path_hash,\nfollowed by B blob SHA-1 + 4 byte path_hash.  For any project the size\nis basically on par with the .idx file for the pack v1 format, so ~41\nMB for linux-2.6.  The file is stored as\n$GIT_OBJECT_DIRECTORY/cache/$COMMIT_SHA1.list, and is completely\npack-independent.\n\nUsing this for object enumeration shaves almost 1 minute off server\npacking time; the clone dropped from 3m28s to 2m29s.  That is close to\nwhat I was getting with the cached pack idea, but the network transfer\nstayed the small 376 MiB.  I think this supports your pack v4 work...\nif we can speed up object enumeration to be this simple (scan down a\nlist of objects with their types declared inline, or implied by\nlocation), we can cut a full minute of CPU time off the server side.\n\n-- \nShawn.\n"},{"id":"160019","messageId":"alpine.LFD.2.00.1101282055190.8580@xanadu.home","threadId":"26353","inReplyTo":"AANLkTi=U7qRRij=BQXC1Goqa9toDFfaVKT=+-8zYxCcc@mail.gmail.com","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2011-01-29T04:08:23Z","receivedAt":"2011-01-29T04:08:23Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 28 Jan 2011, Shawn Pearce wrote:\n\n> On Fri, Jan 28, 2011 at 13:09, Nicolas Pitre <nico@fluxnic.net> wrote:\n> > On Fri, 28 Jan 2011, Shawn Pearce wrote:\n> >\n> >> On Fri, Jan 28, 2011 at 10:46, Nicolas Pitre <nico@fluxnic.net> wrote:\n> >> > On Fri, 28 Jan 2011, Shawn Pearce wrote:\n> >> >\n> >> >> This started because I was looking for a way to speed up clones coming\n> >> >> from a JGit server.  Cloning the linux-2.6 repository is painful,\n> \n> Well, scratch the idea in this thread.  I think.\n> \n> I retested JGit vs. CGit on an identical linux-2.6 repository.  The\n> repository was fully packed, but had two pack files.  362M and 57M,\n> and was created by packing a 1 month old master, marking it .keep, and\n> then repacking -a -d to get most recent last month into another pack.\n> This results in some files that should be delta compressed together\n> being stored whole in the two packs (obviously).\n> \n> The two implementations take the same amount of time to generate the\n> clone.  3m28s / 3m22s for JGit, 3m23s for C Git.  The JGit created\n> pack is actually smaller 376.30 MiB vs. C Git's 380.59 MiB.  I point\n> out this data because improvements made to JGit may show similar\n> improvements to CGit given how close they are in running time.\n\nWhat are those improvements?\n\nNow, the fact that JGit is so close to CGit must be because the actual \ncost is outside of them such as within zlib, otherwise the C code should \nnormally always be faster, right?\n\nLooking at the profile for \"git rev-list --objects --all > /dev/null\" \nfor the object enumeration phase, we have:\n\n# Samples: 1814637\n#\n# Overhead          Command  Shared Object  Symbol\n# ........  ...............  .............  ......\n#\n    28.81%              git  /home/nico/bin/git  [.] lookup_object\n    12.21%              git  /lib64/libz.so.1.2.3  [.] inflate\n    10.49%              git  /lib64/libz.so.1.2.3  [.] inflate_fast\n     7.47%              git  /lib64/libz.so.1.2.3  [.] inflate_table\n     6.66%              git  /lib64/libc-2.11.2.so  [.] __GI_memcpy\n     5.66%              git  /home/nico/bin/git  [.] find_pack_entry_one\n     2.98%              git  /home/nico/bin/git  [.] decode_tree_entry\n     2.73%              git  /lib64/libc-2.11.2.so  [.] _int_malloc\n     2.71%              git  /lib64/libz.so.1.2.3  [.] adler32\n     2.63%              git  /home/nico/bin/git  [.] process_tree\n     1.58%              git  [kernel]       [k] 0xffffffff8112fc0c\n     1.44%              git  /lib64/libc-2.11.2.so  [.] __strlen_sse2\n     1.31%              git  /home/nico/bin/git  [.] tree_entry\n     1.10%              git  /lib64/libc-2.11.2.so  [.] _int_free\n     0.96%              git  /home/nico/bin/git  [.] patch_delta\n     0.92%              git  /lib64/libc-2.11.2.so  [.] malloc_consolidate\n     0.86%              git  /lib64/libc-2.11.2.so  [.] __GI_vfprintf\n     0.80%              git  /home/nico/bin/git  [.] create_object\n     0.80%              git  /home/nico/bin/git  [.] lookup_blob\n     0.63%              git  /home/nico/bin/git  [.] update_tree_entry\n[...]\n\nSo we've got lookup_object() clearly at the top.  I suspect the \nhashcmp() in there, which probably gets inlined, is responsible for most \ncycles.  There is certainly a better way here, and probably in JGit you \nrely on some optimized facility provided by the language/library to \nperform that lookup.  So there is probably some easy improvements that \ncan be made here.\n\nOtherwise it is at least 12.21 + 10.49 + 7.47 + 2.71 = 32.88% spent \ndirectly in the zlib code, making it the biggest cost.  This is rather \nunavoidable unless the data structure is changed.  And pack v4 would \nprobably move things such as find_pack_entry_one, decode_tree_entry, \nprocess_tree and tree_entry off the radar as well.\n\nThe object writeout phase should pretty much be network bound.\n\n> I fully implemented the reuse of a cached pack behind a thin pack idea\n> I was trying to describe in this thread.  It saved 1m7s off the JGit\n> running time, but increased the data transfer by 25 MiB.  I didn't\n> expect this much of an increase, I honestly expected the thin pack\n> portion to be well, thinner.  The issue is the thin pack cannot delta\n> against all of the history, its only delta compressing against the tip\n> of the cached pack.  So long-lived side branches that forked off an\n> older part of the history aren't delta compressing well, or at all,\n> and that is significantly bloating the thin pack.  (Its also why that\n> \"newer\" pack is 57M, but should be 14M if correctly combined with the\n> cached pack.)  If I were to consider all of the objects in the cached\n> pack as potential delta base candidates for the thin pack, the entire\n> benefit of the cached pack disappears.\n\nYeah... this sucks.\n\n> I'm not sure I like it so much anymore.  :-)\n> \n> The idea was half-baked, and came at the end of a long day, and after\n> putting my cranky infant son down to sleep way past his normal bed\n> time.  I claim I was a sleep deprived new parent who wasn't thinking\n> things through enough before writing an email to git@vger.\n\nWell, this is still valuable information to archive.\n\nAnd I wish I had been able to still write such quality emails when I was \na new parent.  ;-)\n\n> >> sendfile() call for the bulk of the content.  I think we can just hand\n> >> off the major streaming to the kernel.\n> >\n> > While this might look like a good idea in theory, did you actually\n> > profile it to see if that would make a noticeable difference?  The\n> > pkt-line framing allows for asynchronous messages to be sent over a\n> > sideband,\n> \n> No, of course not.  The pkt-line framing is pretty low overhead, but\n> copying kernel buffer to userspace back to kernel buffer sort of sucks\n> for 400 MiB of data.  sendfile() on 400 MiB to a network socket is\n> much easier when its all kernel space.\n\nOf course.  But still... If you save 0.5 second by avoiding the copy to \nand from user space of that 400 MiB (based on my machine which can do \n1670MB/s) that's pretty much insignificant compared to the total time \nfor the clone, and therefore the wrong thing to optimize given the \nrequired protocol changes.\n\n> I figured, if it all worked\n> out already to just dump the pack to the wire as-is, then we probably\n> should also try to go for broke and reduce the userspace copying.  It\n> might not matter to your desktop, but ask John Hawley (CC'd) about\n> kernel.org and the git traffic volume he is serving.  They are doing\n> more than 1 million git:// requests per day now.\n\nImpressive.  However I suspect that the vast majority of those requests \nare from clients making a connection just to realize they're up to date \nalready.  I don't think the user space copying is really a problem.\n\nOf course, if we could have used sendfile() freely in, say, \ncopy_pack_data() then we would have done so long ago.  But we are \nchecksuming the data we create on the fly with the data we reuse from \ndisk so this is not necessarily a gain.\n\n> >> Plus we can safely do byte range requests for resumable clone within\n> >> the cached pack part of the stream.\n> >\n> > That part I'm not sure of.  We are still facing the same old issues\n> > here, as some mirrors might have the same commit edges for a cache pack\n> > but not necessarily the same packing result, etc.  So I'd keep that out\n> > of the picture for now.\n> \n> I don't think its that hard.  If we modify the transfer protocol to\n> allow the server to denote boundaries between packs, the server can\n> send the pack name (as in pack-$name.pack) and the pack SHA-1 trailer\n> to the client.  A client asking for resume of a cached pack presents\n> its original want list, these two SHA-1s, and the byte offset he wants\n> to restart from.  The server validates the want set is still\n> reachable, that the cached pack exists, and that the cached pack tips\n> are reachable from current refs.  If all of that is true, it validates\n> the trailing SHA-1 in the pack matches what the client gave it.  If\n> that matches, it should be OK to resume transfer from where the client\n> asked for.\n\nThis is still an half solution.  If your network connection drops after \nthe first 52 MiB of transfer given the scenario you provided then you're \nstill screwed.\n\n\nNicolas\n"},{"id":"160020","messageId":"AANLkTik01sVOKD+gW_BembUDN-LRbqJsPCCBBVyJJ11T@mail.gmail.com","threadId":"26353","inReplyTo":"alpine.LFD.2.00.1101282055190.8580@xanadu.home","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-01-29T04:35:48Z","receivedAt":"2011-01-29T04:35:48Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Fri, Jan 28, 2011 at 20:08, Nicolas Pitre <nico@fluxnic.net> wrote:\n>> pack is actually smaller 376.30 MiB vs. C Git's 380.59 MiB.  I point\n>> out this data because improvements made to JGit may show similar\n>> improvements to CGit given how close they are in running time.\n>\n> What are those improvements?\n\nNone right now.  JGit is similar to CGit algorithm-wise.  (Actually it\nlooks like JGit has a faster diff implementation, but that's a\ndifferent email.)\n\nIf you are asking about why JGit created a slightly smaller pack\nfile... it splits the delta window during threaded delta search\ndifferently than CGit does, and we align our blocks slightly\ndifferently when comparing two objects to generate a delta sequence\nfor them.  These two variations mean JGit produces different deltas\nthan CGit does.  Sometimes we are smaller, sometimes we are larger.\nBut its a small difference, on the order of 1-4 MiB for something like\nlinux-2.6.  I don't think its worthwhile trying to analyze the\nspecific differences in implementations and retrofit those differences\ninto the other one.\n\nWhat I was trying to say was, _if_ we made a change to JGit and it\ndropped the running time, that same change in CGit should have _at\nleast_ the same running time improvement, if not better.  I was\npointing out that this cached-pack change dropped the running time by\n1 minute, so CGit should also see a similar improvement (if not\nbetter).  I would prefer to test against CGit for this sort of thing,\nbut its been too long since I last poked pack-objects.c and the\nrevision code in CGit, while the JGit equivalents are really fresh in\nmy head.\n\n> Now, the fact that JGit is so close to CGit must be because the actual\n> cost is outside of them such as within zlib, otherwise the C code should\n> normally always be faster, right?\n\nYup, I mostly agree with this statement.  CGit does a lot of\nmalloc/free activity when reading objects in.  JGit does too, but we\noften fit into the young generation for the GC, which sometimes can be\nfaster to clean and recycle memory in.  We're not too far off from C\ncode.\n\nBut yes... our profile looks like this too:\n\n> Looking at the profile for \"git rev-list --objects --all > /dev/null\"\n> for the object enumeration phase, we have:\n>\n> # Samples: 1814637\n> #\n> # Overhead          Command  Shared Object  Symbol\n> # ........  ...............  .............  ......\n> #\n>    28.81%              git  /home/nico/bin/git  [.] lookup_object\n>    12.21%              git  /lib64/libz.so.1.2.3  [.] inflate\n>    10.49%              git  /lib64/libz.so.1.2.3  [.] inflate_fast\n>     7.47%              git  /lib64/libz.so.1.2.3  [.] inflate_table\n>     6.66%              git  /lib64/libc-2.11.2.so  [.] __GI_memcpy\n>     5.66%              git  /home/nico/bin/git  [.] find_pack_entry_one\n>     2.98%              git  /home/nico/bin/git  [.] decode_tree_entry\n> [...]\n>\n> So we've got lookup_object() clearly at the top.\n\nIsn't this the hash table lookup inside the revision pool, to see if\nthe object has already been visited?  That seems horrible, 28% of the\nCPU is going to probing that table.\n\n>  I suspect the\n> hashcmp() in there, which probably gets inlined, is responsible for most\n> cycles.\n\nProbably true.  I know our hashcmp() is inlined, its actually written\nby hand as 5 word compares, and is marked final, so the JIT is rather\nlikely to inline it.\n\n>  There is certainly a better way here, and probably in JGit you\n> rely on some optimized facility provided by the language/library to\n> perform that lookup.  So there is probably some easy improvements that\n> can be made here.\n\nNope.  Actually we have to bend over backwards and work against the\nlanguage to get anything even reasonably sane for performance.  Our\n\"solution\" in JGit has actually been used by Rob Pike to promote his\nGo programming language and why Java sucks as a language.  Its a great\nquote of mine that someone dragged up off the git@vger mailing list\nand started using to promote Go.\n\nAt least once I week I envy how easy it is to use hashcmp() and\nhashcpy() inside of CGit.  JGit's management of hashes is sh*t because\nwe have to bend so hard around the language.\n\n> Otherwise it is at least 12.21 + 10.49 + 7.47 + 2.71 = 32.88% spent\n> directly in the zlib code, making it the biggest cost.\n\nYea, that's what we have too, about 33% inside of zlib code... which\nis the same implementation that CGit uses.\n\n>  This is rather\n> unavoidable unless the data structure is changed.\n\nWe already knew this from our pack v4 experiments years ago.\n\n>  And pack v4 would\n> probably move things such as find_pack_entry_one, decode_tree_entry,\n> process_tree and tree_entry off the radar as well.\n\nThis is hard to do inside of CGit if I recall... but yes, changing the\nway trees are handled would really improve things.\n\n> The object writeout phase should pretty much be network bound.\n\nYes.\n\n>> I fully implemented the reuse of a cached pack behind a thin pack idea\n>> I was trying to describe in this thread.  It saved 1m7s off the JGit\n>> running time, but increased the data transfer by 25 MiB.\n>\n> Yeah... this sucks.\n\nVery much.  :-(\n\nBut this is a fundamental issue with our incremental fetch support\nanyway.  In this exact case if the client was at that 1 month old\ncommit, and fetched current master, he would pull 25 MiB of data.. but\nonly needed about 4-6 MiB worth of deltas if it was properly delta\ncompressed against the content we know he already has.  Our server\nside optimization of only pushing the immediate \"have\" list of the\nclient into the delta search window limits how much we can compress\nthe data we are sending.  If we were willing to push more in on the\nserver side, we could shrink the incremental fetch more.  But that's a\nCPU problem on the server.\n\n-- \nShawn.\n"},{"id":"160045","messageId":"7voc6yc2au.fsf@alter.siamese.dyndns.org","threadId":"26353","inReplyTo":"AANLkTi=U7qRRij=BQXC1Goqa9toDFfaVKT=+-8zYxCcc@mail.gmail.com","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-01-30T06:51:05Z","receivedAt":"2011-01-30T06:51:05Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Shawn Pearce <spearce@spearce.org> writes:\n\n> I fully implemented the reuse of a cached pack behind a thin pack idea\n> I was trying to describe in this thread.  It saved 1m7s off the JGit\n> running time, but increased the data transfer by 25 MiB.  I didn't\n> expect this much of an increase, I honestly expected the thin pack\n> portion to be well, thinner.  The issue is the thin pack cannot delta\n> against all of the history, its only delta compressing against the tip\n> of the cached pack.  So long-lived side branches that forked off an\n> older part of the history aren't delta compressing well, or at all,\n> and that is significantly bloating the thin pack.  (Its also why that\n> \"newer\" pack is 57M, but should be 14M if correctly combined with the\n> cached pack.)  If I were to consider all of the objects in the cached\n> pack as potential delta base candidates for the thin pack, the entire\n> benefit of the cached pack disappears.\n\nWhat if you instead use the cached pack this way?\n\n 0. You perform the proposed pre-traversal until you hit the tip of cached\n    pack(s), and realize that you will end up sending everything.\n\n 1. Instead of sending the new part of the history first and then sending\n    the cached pack(s), you send the contents of cached pack(s), but also\n    note what objects you sent;\n\n 2. Then you send the new part of the history, taking full advantage of\n    what you have already sent, perhaps doing only half of the reuse-delta\n    logic (i.e. you reuse what you can reuse, but you do _not_ punt on an\n    object that is not a delta in an existing pack).\n"},{"id":"160046","messageId":"7vk4hmbyuo.fsf@alter.siamese.dyndns.org","threadId":"26353","inReplyTo":"AANLkTimuW-7D4YA2jeF+y4DPE=CdqtL713MQK+1Gtp-d@mail.gmail.com","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-01-30T08:05:35Z","receivedAt":"2011-01-30T08:05:35Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Shawn Pearce <spearce@spearce.org> writes:\n\n> Using this for object enumeration shaves almost 1 minute off server\n> packing time; the clone dropped from 3m28s to 2m29s.  That is close to\n> what I was getting with the cached pack idea, but the network transfer\n> stayed the small 376 MiB.\n\nI like this result.\n\nThe amount of transfer being that small was something I didn't quite\nexpect, though.  Doesn't it indicate that our pathname based object\nclustering heuristics is not as effective as we hoped?\n"},{"id":"160055","messageId":"alpine.LFD.2.00.1101301208270.8580@xanadu.home","threadId":"26353","inReplyTo":"7voc6yc2au.fsf@alter.siamese.dyndns.org","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2011-01-30T17:14:47Z","receivedAt":"2011-01-30T17:14:47Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sat, 29 Jan 2011, Junio C Hamano wrote:\n\n> Shawn Pearce <spearce@spearce.org> writes:\n> \n> > I fully implemented the reuse of a cached pack behind a thin pack idea\n> > I was trying to describe in this thread.  It saved 1m7s off the JGit\n> > running time, but increased the data transfer by 25 MiB.  I didn't\n> > expect this much of an increase, I honestly expected the thin pack\n> > portion to be well, thinner.  The issue is the thin pack cannot delta\n> > against all of the history, its only delta compressing against the tip\n> > of the cached pack.  So long-lived side branches that forked off an\n> > older part of the history aren't delta compressing well, or at all,\n> > and that is significantly bloating the thin pack.  (Its also why that\n> > \"newer\" pack is 57M, but should be 14M if correctly combined with the\n> > cached pack.)  If I were to consider all of the objects in the cached\n> > pack as potential delta base candidates for the thin pack, the entire\n> > benefit of the cached pack disappears.\n> \n> What if you instead use the cached pack this way?\n> \n>  0. You perform the proposed pre-traversal until you hit the tip of cached\n>     pack(s), and realize that you will end up sending everything.\n> \n>  1. Instead of sending the new part of the history first and then sending\n>     the cached pack(s), you send the contents of cached pack(s), but also\n>     note what objects you sent;\n> \n>  2. Then you send the new part of the history, taking full advantage of\n>     what you have already sent, perhaps doing only half of the reuse-delta\n>     logic (i.e. you reuse what you can reuse, but you do _not_ punt on an\n>     object that is not a delta in an existing pack).\n\nThe problem is to determine the best base object to delta against.  If \nyou end up listing all the already sent objects and perform delta \nattempts against them for the remaining non delta objects to find the \nbest match then you might end up taking more CPU time than the current \nenumeration phase.\n\n\nNicolas\n"},{"id":"160056","messageId":"4D45A2BA.8040908@gmail.com","threadId":"26353","inReplyTo":"alpine.LFD.2.00.1101301208270.8580@xanadu.home","subject":"Re: [RFC] Add --create-cache to repack","fromName":"A Large Angry SCM","fromEmail":"gitzilla@gmail.com","sentAt":"2011-01-30T17:41:14Z","receivedAt":"2011-01-30T17:41:14Z","isPatch":false,"sender":{"key":"gitzilla@gmail.com","avatar":"https://gravatar.com/avatar/354625c442439908ff3dd99757dee330e29e9df7847472384faf7a00add247fb?d=mp&s=160"},"body":"On 01/30/2011 12:14 PM, Nicolas Pitre wrote:\n> On Sat, 29 Jan 2011, Junio C Hamano wrote:\n>\n>> Shawn Pearce<spearce@spearce.org>  writes:\n>>\n>>> I fully implemented the reuse of a cached pack behind a thin pack idea\n>>> I was trying to describe in this thread.  It saved 1m7s off the JGit\n>>> running time, but increased the data transfer by 25 MiB.  I didn't\n>>> expect this much of an increase, I honestly expected the thin pack\n>>> portion to be well, thinner.  The issue is the thin pack cannot delta\n>>> against all of the history, its only delta compressing against the tip\n>>> of the cached pack.  So long-lived side branches that forked off an\n>>> older part of the history aren't delta compressing well, or at all,\n>>> and that is significantly bloating the thin pack.  (Its also why that\n>>> \"newer\" pack is 57M, but should be 14M if correctly combined with the\n>>> cached pack.)  If I were to consider all of the objects in the cached\n>>> pack as potential delta base candidates for the thin pack, the entire\n>>> benefit of the cached pack disappears.\n>>\n>> What if you instead use the cached pack this way?\n>>\n>>   0. You perform the proposed pre-traversal until you hit the tip of cached\n>>      pack(s), and realize that you will end up sending everything.\n>>\n>>   1. Instead of sending the new part of the history first and then sending\n>>      the cached pack(s), you send the contents of cached pack(s), but also\n>>      note what objects you sent;\n>>\n>>   2. Then you send the new part of the history, taking full advantage of\n>>      what you have already sent, perhaps doing only half of the reuse-delta\n>>      logic (i.e. you reuse what you can reuse, but you do _not_ punt on an\n>>      object that is not a delta in an existing pack).\n>\n> The problem is to determine the best base object to delta against.  If\n> you end up listing all the already sent objects and perform delta\n> attempts against them for the remaining non delta objects to find the\n> best match then you might end up taking more CPU time than the current\n> enumeration phase.\n\nWhy worry about best here? Just add the object (or one of the objects) \nwith the same path from the commit you found in step 0, above, to the \ndelta base search for each object to pack.\n"},{"id":"160057","messageId":"AANLkTin_7jHRAx4yDw3DirLX7Lbdv3s3iz+-Q-QXCZqq@mail.gmail.com","threadId":"26353","inReplyTo":"7voc6yc2au.fsf@alter.siamese.dyndns.org","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-01-30T19:29:49Z","receivedAt":"2011-01-30T19:29:49Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Sat, Jan 29, 2011 at 22:51, Junio C Hamano <gitster@pobox.com> wrote:\n> Shawn Pearce <spearce@spearce.org> writes:\n>\n>> I fully implemented the reuse of a cached pack behind a thin pack idea\n>> I was trying to describe in this thread.  It saved 1m7s off the JGit\n>> running time, but increased the data transfer by 25 MiB.  I didn't\n>> expect this much of an increase, I honestly expected the thin pack\n>> portion to be well, thinner.  The issue is the thin pack cannot delta\n>> against all of the history, its only delta compressing against the tip\n>> of the cached pack.  So long-lived side branches that forked off an\n>> older part of the history aren't delta compressing well, or at all,\n>> and that is significantly bloating the thin pack.  (Its also why that\n>> \"newer\" pack is 57M, but should be 14M if correctly combined with the\n>> cached pack.)  If I were to consider all of the objects in the cached\n>> pack as potential delta base candidates for the thin pack, the entire\n>> benefit of the cached pack disappears.\n>\n> What if you instead use the cached pack this way?\n>\n>  0. You perform the proposed pre-traversal until you hit the tip of cached\n>    pack(s), and realize that you will end up sending everything.\n>\n>  1. Instead of sending the new part of the history first and then sending\n>    the cached pack(s), you send the contents of cached pack(s), but also\n>    note what objects you sent;\n\nThis is the part I was trying to avoid.  Making this list of objects\nfrom the cached pack(s) costs working set inside of the pack-objects\nprocess.  I had hoped that the cached packs would let me skip this\nstep.\n\nBut lets say that's acceptable cost.  We cannot efficiently make a\nuseful list of objects from the pack.  Scanning the .idx file only\ntells us the SHA-1.  It does not tell us the type, nor does it tell us\nwhat the path hash code would be for the object if it were a tree or\nblob.  So we cannot efficiently use this pack listing to construct the\ndelta window.\n\n>  2. Then you send the new part of the history, taking full advantage of\n>    what you have already sent, perhaps doing only half of the reuse-delta\n>    logic (i.e. you reuse what you can reuse, but you do _not_ punt on an\n>    object that is not a delta in an existing pack).\n\nWell, I guess we could go half-way.  We could try to use only\nnon-delta objects from the cached pack as potential delta bases for\nthis delta search.\n\nTo do that we would build the reverse index for the cached pack, then\ncheck each object's type code just before we send that part of the\ncached pack.  If its non-delta, we can get its SHA-1 from the reverse\nindex, toss the object into the delta search list, and copy out the\nlength of the object until the next object starts.\n\nHowever... I suspect our delta results would be the same as the thin\npack before cached pack test I did earlier.  The objects that are\nnon-delta in the cached pack are (in theory) approximately the objects\nimmediately reachable from the cached pack's tip.  That was already\nput into the delta window as the base candidates for the thin pack.\nThis may be a faster way to find that thin pack edge, but the data\ntransfer will still be sub-optimal because we cannot consider deltas\nas bases.\n\n-- \nShawn.\n"},{"id":"160059","messageId":"AANLkTi=mbeBsR5tr4J7kQCL6YqiGfttK01VUN016aapC@mail.gmail.com","threadId":"26353","inReplyTo":"7vk4hmbyuo.fsf@alter.siamese.dyndns.org","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-01-30T19:43:22Z","receivedAt":"2011-01-30T19:43:22Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Sun, Jan 30, 2011 at 00:05, Junio C Hamano <gitster@pobox.com> wrote:\n> Shawn Pearce <spearce@spearce.org> writes:\n>\n>> Using this for object enumeration shaves almost 1 minute off server\n>> packing time; the clone dropped from 3m28s to 2m29s.  That is close to\n>> what I was getting with the cached pack idea, but the network transfer\n>> stayed the small 376 MiB.\n>\n> I like this result.\n\nI'm really leaning towards putting this cached object list into JGit.\n\nI need to shave that 1 minute off server CPU time. I can afford the 41\nMiB disk (and kernel buffer cache), but I cannot really continue to\npay the 1 minute of CPU on each clone request for large repositories.\nThe object list of what is reachable from commit X isn't ever going to\nchange, and the path hash function is reasonably stable.  With a\nversion code in the file we can desupport old files if the path hash\nfunction changes.  10% more disk/kernel memory is cheap for some of my\nservers compared to 1 minute of CPU, and some explicit cache\nmanagement by the server administrator to construct the file.\n\n> The amount of transfer being that small was something I didn't quite\n> expect, though.  Doesn't it indicate that our pathname based object\n> clustering heuristics is not as effective as we hoped?\n\nI'm not sure I follow your question.\n\nI think the problem here is old side branches that got recently\nmerged.  Their _best_ delta base was some old revision, possibly close\nto where they branched off from.  Using a newer version of the file\nfor the delta base created a much larger delta.  E.g. consider a file\nwhere in more recent revisions a function was completely rewritten.\nIf you have to delta compress against that new version, but you use\nthe older definition of the function, you need to use insert\ninstructions\nfor the entire content of that old function.  But if you can delta\ncompress against the version you branched from (or one much closer to\nit in time), your delta would be very small as that function is\nhandled by the smaller copy instruction.\n\nOur clustering heuristics work fine.\n\nOur thin-pack selection of potential delta base candidates is not.  We\nare not very aggressive in loading the delta base window with\npotential candidates, which means we miss some really good compression\nopportunities.\n\n\nOoooh.\n\nI think my test was flawed.  I injected the cached pack's tip as the\nedge for the new stuff to delta compress against.  I should have\ninjected all of the merge bases between the cached pack's tip and the\nnew stuff.  Although the cached pack tip is one of the merge bases,\nits not all of them.  If we inject all of the merge bases, we can find\nthe revision that this old side branch is based on, and possibly get a\nbetter delta candidate for it.\n\nIIRC, upload-pack would have walked backwards further and found the\nmerge base for that side branch, and it would have been part of the\ndelta base candidates.  I think I need to re-do my cached pack test.\nGood thing I have history of my source code saved in this fancy\nrevision control thingy called \"git\".  :-)\n\n-- \nShawn.\n"},{"id":"160062","messageId":"7vaaiib1n1.fsf@alter.siamese.dyndns.org","threadId":"26353","inReplyTo":"AANLkTi=mbeBsR5tr4J7kQCL6YqiGfttK01VUN016aapC@mail.gmail.com","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-01-30T20:02:58Z","receivedAt":"2011-01-30T20:02:58Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Shawn Pearce <spearce@spearce.org> writes:\n\n>> The amount of transfer being that small was something I didn't quite\n>> expect, though.  Doesn't it indicate that our pathname based object\n>> clustering heuristics is not as effective as we hoped?\n>\n> I'm not sure I follow your question.\n\nI didn't see path information in your cachefile that contains C commits, T\ntrees, etc. that sped up the object enumeration, but you didn't observe\nmuch transfer inflation over the stock git.\n\n> Ooooh.\n>\n> I think my test was flawed.  I injected the cached pack's tip as the\n> edge for the new stuff to delta compress against.\n\nThat is one of the things I was wondering.  I manually created a thin pack\nwith only the 1-month-old tip as boundary, and another with all the\nboundaries that can be found by rev-list.  I didn't find much difference\nin the result, though, as \"rev-list --boundary --all --not $onemontholdtip\"\nhad only a few boundary entries in my test.\n"},{"id":"160064","messageId":"AANLkTi=4svGhw+L0StH=PKdVWZtbSTSKe+a8dqSmkSE0@mail.gmail.com","threadId":"26353","inReplyTo":"7vaaiib1n1.fsf@alter.siamese.dyndns.org","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-01-30T20:20:31Z","receivedAt":"2011-01-30T20:20:31Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Sun, Jan 30, 2011 at 12:02, Junio C Hamano <gitster@pobox.com> wrote:\n> Shawn Pearce <spearce@spearce.org> writes:\n>\n>>> The amount of transfer being that small was something I didn't quite\n>>> expect, though.  Doesn't it indicate that our pathname based object\n>>> clustering heuristics is not as effective as we hoped?\n>>\n>> I'm not sure I follow your question.\n>\n> I didn't see path information in your cachefile that contains C commits, T\n> trees, etc. that sped up the object enumeration, but you didn't observe\n> much transfer inflation over the stock git.\n\nI didn't store the path itself, I stored the path hash as a 4 byte\nint.  Its smaller, but still helps to schedule the object into the\nright position in the delta search.\n\n-- \nShawn.\n"},{"id":"160071","messageId":"AANLkTin6dcfMAD1tg+ROujy-_Dvi6KG-+-nfgG44u=Mh@mail.gmail.com","threadId":"26353","inReplyTo":"AANLkTi=U7qRRij=BQXC1Goqa9toDFfaVKT=+-8zYxCcc@mail.gmail.com","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-01-30T22:13:45Z","receivedAt":"2011-01-30T22:13:45Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Fri, Jan 28, 2011 at 17:32, Shawn Pearce <spearce@spearce.org> wrote:\n>\n> I fully implemented the reuse of a cached pack behind a thin pack idea\n> I was trying to describe in this thread.  It saved 1m7s off the JGit\n> running time, but increased the data transfer by 25 MiB.  I didn't\n> expect this much of an increase, I honestly expected the thin pack\n> portion to be well, thinner.\n\nJGit's thin pack creation is crap.  For example, this is the same fetch:\n\n$ git fetch ../tmp_linux26\nremote: Counting objects: 61521, done.\nremote: Compressing objects: 100% (12096/12096), done.\nremote: Total 50275 (delta 42578), reused 45220 (delta 37524)\nReceiving objects: 100% (50275/50275), 11.13 MiB | 7.29 MiB/s, done.\nResolving deltas: 100% (42578/42578), completed with 4968 local objects.\n\n$ git fetch git://localhost/tmp_linux26\nremote: Counting objects: 144190, done\nremote: Finding sources: 100% (50275/50275)\nremote: Compressing objects: 100% (106568/106568)\nremote: Compressing objects: 100% (12750/12750)\nReceiving objects: 100% (50275/50275), 24.66 MiB | 10.93 MiB/s, done.\nResolving deltas: 100% (40345/40345), completed with 2218 local objects.\n\n\nJGit produced an extra 13.53 MiB for this pack, because it missed\nabout 2,233 delta opportunities.  It turns out we are too aggressive\nat pushing objects from the edges into the delta windows.  JGit pushes\n*everything* in the edge commits, rather than only the paths that are\nactually used by the objects we need to send.  This floods the delta\nsearch window with garbage, and makes it less likely that an object to\nbe sent will find a relevant delta base in the search window.\n\n-- \nShawn.\n"},{"id":"160072","messageId":"alpine.LFD.2.00.1101301716400.8580@xanadu.home","threadId":"26353","inReplyTo":"AANLkTi=mbeBsR5tr4J7kQCL6YqiGfttK01VUN016aapC@mail.gmail.com","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2011-01-30T22:26:40Z","receivedAt":"2011-01-30T22:26:40Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sun, 30 Jan 2011, Shawn Pearce wrote:\n\n> On Sun, Jan 30, 2011 at 00:05, Junio C Hamano <gitster@pobox.com> wrote:\n> > Shawn Pearce <spearce@spearce.org> writes:\n> >\n> >> Using this for object enumeration shaves almost 1 minute off server\n> >> packing time; the clone dropped from 3m28s to 2m29s.  That is close to\n> >> what I was getting with the cached pack idea, but the network transfer\n> >> stayed the small 376 MiB.\n> >\n> > I like this result.\n> \n> I'm really leaning towards putting this cached object list into JGit.\n> \n> I need to shave that 1 minute off server CPU time. I can afford the 41\n> MiB disk (and kernel buffer cache), but I cannot really continue to\n> pay the 1 minute of CPU on each clone request for large repositories.\n> The object list of what is reachable from commit X isn't ever going to\n> change, and the path hash function is reasonably stable.  With a\n> version code in the file we can desupport old files if the path hash\n> function changes.  10% more disk/kernel memory is cheap for some of my\n> servers compared to 1 minute of CPU, and some explicit cache\n> management by the server administrator to construct the file.\n\nYep, I think this is probably the best short term solution.  Just walk \nthe commit graph as usual, and whenever the commit tip from the cache is \nmatched then just shove the entire cache content in the object list.\n\nAnd let's hope that eventually some future developments will make this \ncache redundant and obsolete.\n\n\nNicolas\n"},{"id":"160102","messageId":"AANLkTimW=fuKrhw6+ZDipEtSGX_oR4LbTZzyAxZ8Pry1@mail.gmail.com","threadId":"26353","inReplyTo":"AANLkTi=U7qRRij=BQXC1Goqa9toDFfaVKT=+-8zYxCcc@mail.gmail.com","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-01-31T18:47:34Z","receivedAt":"2011-01-31T18:47:34Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Fri, Jan 28, 2011 at 17:32, Shawn Pearce <spearce@spearce.org> wrote:\n>>> >\n>>> >> This started because I was looking for a way to speed up clones coming\n>>> >> from a JGit server.  Cloning the linux-2.6 repository is painful,\n>\n> Well, scratch the idea in this thread.  I think.\n\nNope, I'm back in favor with this after fixing JGit's thin pack\ngeneration.  Here's why.\n\nTake linux-2.6.git as of Jan 12th, with the cache root as of Dec 28th:\n\n  $ git update-ref HEAD f878133bf022717b880d0e0995b8f91436fd605c\n  $ git-repack.sh --create-cache \\\n      --cache-root=b52e2a6d6d05421dea6b6a94582126af8cd5cca2 \\\n      --cache-include=v2.6.11-tree\n  $ git repack -a -d\n\n  $ ls -lh objects/pack/\n  total 456M\n  1.4M pack-74af5edca80797736fe4de7279b2a81af98470a5.idx\n  38M pack-74af5edca80797736fe4de7279b2a81af98470a5.pack\n\n  49M pack-d3e77c8b3045c7f54fa2fb6bbfd4dceca1e2b9fa.idx\n  89 pack-d3e77c8b3045c7f54fa2fb6bbfd4dceca1e2b9fa.keep\n  368M pack-d3e77c8b3045c7f54fa2fb6bbfd4dceca1e2b9fa.pack\n\nOur \"recent history\" is 38M, and our \"cached pack\" is 368M.  Its a bit\nmore disk than is strictly necessary, this should be ~380M.  Call it\n~26M of wasted disk.  The \"cached object list\" I proposed elsewhere in\nthis thread would cost about 41M of disk and is utterly useless except\nfor initial clones.  Here we are wasting about 26M of disk to have\nslightly shorter delta chains in the cached pack (otherwise known as\nour ancient history).  So its a slightly smaller waste, and we get\nsome (minor) benefit.\n\n\nClone without pack caching:\n\n  $ time git clone --bare git://localhost/tmp_linux26_withTag tmp_in.git\n  Cloning into bare repository tmp_in.git...\n  remote: Counting objects: 1861830, done\n  remote: Finding sources: 100% (1861830/1861830)\n  remote: Getting sizes: 100% (88243/88243)\n  remote: Compressing objects: 100% (88184/88184)\n  Receiving objects: 100% (1861830/1861830), 376.01 MiB | 19.01 MiB/s, done.\n  remote: Total 1861830 (delta 4706), reused 1851053 (delta 1553844)\n  Resolving deltas: 100% (1564621/1564621), done.\n\n  real\t3m19.005s\n  user\t1m36.250s\n  sys\t0m10.290s\n\n\nClone with pack caching:\n\n  $ time git clone --bare git://localhost/tmp_linux26_withTag tmp_in.git\n  Cloning into bare repository tmp_in.git...\n  remote: Counting objects: 1601, done\n  remote: Counting objects: 1828460, done\n  remote: Finding sources: 100% (50475/50475)\n  remote: Getting sizes: 100% (18843/18843)\n  remote: Compressing objects: 100% (7585/7585)\n  remote: Total 1861830 (delta 2407), reused 1856197 (delta 37510)\n  Receiving objects: 100% (1861830/1861830), 378.40 MiB | 31.31 MiB/s, done.\n  Resolving deltas: 100% (1559477/1559477), done.\n\n  real\t2m2.938s\n  user\t1m35.890s\n  sys\t0m9.830s\n\n\nUsing the cached pack increased our total data transfer by 2.39 MiB,\nbut saved 1m17s on server computation time.  If we go back and look at\nour cached pack size (368M), the leading thin-pack should be about\n10.4 MiB (378.40M - 368M = 10.4M).  If I modify the tmp_in.git client\nto have only the cached pack's tip and fetch using CGit, we see the\nthin pack to bring ourselves current is 11.07 MiB (JGit does this in\n10.96 MiB):\n\n  $ cd tmp_in.git\n  $ git update-ref HEAD b52e2a6d6d05421dea6b6a94582126af8cd5cca2\n  $ git repack -a -d  ; # yay we are at ~1 month ago\n\n  $ time git fetch ../tmp_linux26_withTag\n  remote: Counting objects: 60570, done.\n  remote: Compressing objects: 100% (11924/11924), done.\n  remote: Total 49804 (delta 42196), reused 44837 (delta 37231)\n  Receiving objects: 100% (49804/49804), 11.07 MiB | 7.37 MiB/s, done.\n  Resolving deltas: 100% (42196/42196), completed with 4956 local objects.\n  From ../tmp_linux26_withTag\n   * branch            HEAD       -> FETCH_HEAD\n\n  real\t0m35.083s\n  user\t0m25.710s\n  sys\t0m1.190s\n\n\nThe pack caching feature is *no worse* in transfer size than if the\nclient copied the pack from 1 month ago, and then did an incremental\nfetch to bring themselves current.  Compared to the naive clone, it\nsaves an incredible amount of working set space and CPU time.  The\nserver only needs to keep track of the incremental thin pack, and can\ncompletely ignore the ancient history objects.  Its a great\nalternative for projects that want users to rsync/http dumb transport\ndown a large stable repository, then incremental fetch themselves\ncurrent.  Or busy mirror sites that are willing to trade some small\nbandwidth for server CPU and memory.\n\nIn this particular example, there is ~11 MiB of data that cannot be\nsafely resumed, or the first 2.9%.  At 56 KiB/s, a client needs to get\nthrough the first 3 minutes of transfer before they can reach the\nresumable checkpoint (where the thin pack ends, and the cached pack\nstarts).  It would be better if we could resume anywhere in the\nstream, but being able to resume the last 97% is infinitely better\nthan being able to resume nothing.  If someone wants to really go\ncrazy, this is where a \"gittorrent\" client could start up and handle\nthe remaining 97% of the transfer.  :-)\n\n\nI think this is worthwhile.  If we are afraid of the extra 2.39 MiB\ndata transfer this forces on the client when the repository owner\nenables the feature, we should go back and improve our thin-pack code.\n Transferring 11 MiB to catch up a kernel from Dec 28th to Jan 12th\nsounds like a lot of data, and any improvements in the general\nthin-pack code would shrink the leading thin-pack, possibly getting us\nthat 2.39 MiB back.\n\n-- \nShawn.\n"},{"id":"160122","messageId":"alpine.LFD.2.00.1101311421320.8580@xanadu.home","threadId":"26353","inReplyTo":"AANLkTimW=fuKrhw6+ZDipEtSGX_oR4LbTZzyAxZ8Pry1@mail.gmail.com","subject":"Re: [RFC] Add --create-cache to repack","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2011-01-31T21:48:08Z","receivedAt":"2011-01-31T21:48:08Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 31 Jan 2011, Shawn Pearce wrote:\n\n> On Fri, Jan 28, 2011 at 17:32, Shawn Pearce <spearce@spearce.org> wrote:\n> >>> >\n> >>> >> This started because I was looking for a way to speed up clones coming\n> >>> >> from a JGit server.  Cloning the linux-2.6 repository is painful,\n> >\n> > Well, scratch the idea in this thread.  I think.\n> \n> Nope, I'm back in favor with this after fixing JGit's thin pack\n> generation.  Here's why.\n> \n> Take linux-2.6.git as of Jan 12th, with the cache root as of Dec 28th:\n> \n>   $ git update-ref HEAD f878133bf022717b880d0e0995b8f91436fd605c\n>   $ git-repack.sh --create-cache \\\n>       --cache-root=b52e2a6d6d05421dea6b6a94582126af8cd5cca2 \\\n>       --cache-include=v2.6.11-tree\n>   $ git repack -a -d\n> \n>   $ ls -lh objects/pack/\n>   total 456M\n>   1.4M pack-74af5edca80797736fe4de7279b2a81af98470a5.idx\n>   38M pack-74af5edca80797736fe4de7279b2a81af98470a5.pack\n> \n>   49M pack-d3e77c8b3045c7f54fa2fb6bbfd4dceca1e2b9fa.idx\n>   89 pack-d3e77c8b3045c7f54fa2fb6bbfd4dceca1e2b9fa.keep\n>   368M pack-d3e77c8b3045c7f54fa2fb6bbfd4dceca1e2b9fa.pack\n> \n> Our \"recent history\" is 38M, and our \"cached pack\" is 368M.  Its a bit\n> more disk than is strictly necessary, this should be ~380M.  Call it\n> ~26M of wasted disk. \n\nThis is fine.  When doing an incremental fetch, the thin pack does \nminimize the transfer size, but it does increase the stored pack size by \nappending a bunch of non delta objects to make the pack complete.\n\nWhat happens though, is that when gc kicks in, the wasted space is \ncollected back.  Here with a single pack we wouldn't claim that space \nback as our current euristics is to reuse delta (non) pairing by \ndefault.  Maybe in that case we could simply not reuse deltas if they're \nof the REF_DELTA type.\n\n> The \"cached object list\" I proposed elsewhere in\n> this thread would cost about 41M of disk and is utterly useless except\n> for initial clones.  Here we are wasting about 26M of disk to have\n> slightly shorter delta chains in the cached pack (otherwise known as\n> our ancient history).  So its a slightly smaller waste, and we get\n> some (minor) benefit.\n\nWell, of course the ancient history you're willing to keep stable for a \nwhile could be repacked even more aggressively than usual.\n\n> Using the cached pack increased our total data transfer by 2.39 MiB,\n\nThat's more than acceptable IMHO. That's less than 1% of the total \ntransfer.\n\n> I think this is worthwhile.  If we are afraid of the extra 2.39 MiB\n> data transfer this forces on the client when the repository owner\n> enables the feature, we should go back and improve our thin-pack code.\n>  Transferring 11 MiB to catch up a kernel from Dec 28th to Jan 12th\n> sounds like a lot of data, \n\nWell, your timing for this test corresponds with the 2.6.38 merge window \nwhich is a high activity peak for this repository.  Still, that would \nprobably fit the usage scenario in practice pretty well where the cache \npack would be produced on a tagged release which happens right before \nthe merge window.\n\n\n> and any improvements in the general\n> thin-pack code would shrink the leading thin-pack, possibly getting us\n> that 2.39 MiB back.\n\nAny improvement to the thin pack would require more CPU cycles, possibly \nlot more.  So given this transfer overhead is less than 1% already I \ndon't think we need to bother.\n\n\nNicolas\n"}]}