{"thread":{"id":"18724","subject":"Performance issue: initial git clone causes massive repack","startedAt":"2009-04-04T22:07:43Z","lastAt":"2009-04-23T19:30:58Z","messageCount":97,"participants":["Robin H. Johnson","Nicolas Sebrecht","Shawn O. Pearce","Jeff King","david@lang.hm","Sverre Rabbelier","Robin Rosenberg","Nguyen Thai Ngoc Duy","Nicolas Pitre","Junio C Hamano","Matthieu Moy","Jon Smirl","Björn Steinbrink","Jakub Narebski","Martin Langhoff","Linus Torvalds","Mike Hommey","Mark Levedahl","Johannes Schindelin","Sam Vilain","Mike Ralphson","Pieter de Bie","Andreas Ericsson","Christian Couder"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"110349","messageId":"20090404220743.GA869@curie-int","threadId":"18724","inReplyTo":null,"subject":"Performance issue: initial git clone causes massive repack","fromName":"Robin H. Johnson","fromEmail":"robbat2@gentoo.org","sentAt":"2009-04-04T22:07:43Z","receivedAt":"2009-04-04T22:07:43Z","isPatch":false,"sender":{"key":"robbat2@gentoo.org","avatar":"https://avatars.githubusercontent.com/u/373898?v=4"},"body":"Hi,\n\nThis is a first in my series of mails over the next few days, on issues\nthat we've run into planning a potential migration for Gentoo's\nrepository into Git.\n\nOur full repository conversion is large, even after tuning the\nrepacking, the packed repository is between 1.4 and 1.6GiB. As of Feburary\n4th, 2009, it contained 4886949 objects. It is not suitable for\nsplitting into submodules either unfortunately - we have a lot of\ndirectory moves that would cause submodule bloat.\n\nDuring an initial clone, I see that git-upload-pack invokes\npack-objects, despite the ENTIRE repository already being packed - no\nloose objects whatsoever. git-upload-pack then seems to buffer in\nmemory.\n\nIn a small repository, this wouldn't be a problem, as the entire\nrepository can fit in memory very easily. However, with our large\nrepository, git-upload-pack and git-pack-objects grows in memory to well\nmore than the size of the packed repository, and are usually killed by\nthe OOM.\n\nDuring 'remote: Counting objects: 4886949, done.', git-upload-pack peaks at\n2474216KB VSZ and 1143048KB RSS. \nShortly thereafter, we get 'remote: Compressing objects:   0%\n(1328/1994284)', git-pack-objects with ~2.8GB VSZ and ~1.8GB RSS. Here,\nthe CPU burn also starts. On our test server machine (w/ git 1.6.0.6),\nit takes about 200 minutes walltime to finish the pack, IFF the OOM\ndoesn't kick in.\n\nGiven that the repo is entirely packed already, I see no point in doing\nthis.\n\nFor the initial clone, can the git-upload-pack algorithm please send\nexisting packs, and only generate a pack containing the non-packed\nitems?\n\n-- \nRobin Hugh Johnson\nGentoo Linux Developer & Infra Guy\nE-Mail     : robbat2@gentoo.org\nGnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n"},{"id":"110357","messageId":"20090405000536.GA12927@vidovic","threadId":"18724","inReplyTo":"20090404220743.GA869@curie-int","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Sebrecht","fromEmail":"nicolas.s-dev@laposte.net","sentAt":"2009-04-05T00:05:36Z","receivedAt":"2009-04-05T00:05:36Z","isPatch":false,"sender":{"key":"nicolas.s-dev@laposte.net","avatar":null},"body":"On Sat, Apr 04, 2009 at 03:07:43PM -0700, Robin H. Johnson wrote:\n\n> Our full repository conversion is large, even after tuning the\n> repacking, the packed repository is between 1.4 and 1.6GiB. As of Feburary\n> 4th, 2009, it contained 4886949 objects. It is not suitable for\n> splitting into submodules either unfortunately - we have a lot of\n> directory moves that would cause submodule bloat.\n\nActually, I'm not sure that a full portage tree repository would be the\nbest thing to do. It would not be suitable in the long term and working\non the repository/history would be a big mess. Why provide a such repo ?\nOr at least, why provide a such readable repo ?\n\nIMHO, you should provide a repository per upstream package on the main\nserver.\n\n\nPS: what about cc'ing gentoo-scm list ?\n\n-- \nNicolas Sebrecht\n"},{"id":"110360","messageId":"20090405T001239Z@curie.orbis-terrarum.net","threadId":"18724","inReplyTo":"20090405000536.GA12927@vidovic","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Robin H. Johnson","fromEmail":"robbat2@gentoo.org","sentAt":"2009-04-05T00:37:53Z","receivedAt":"2009-04-05T00:37:53Z","isPatch":false,"sender":{"key":"robbat2@gentoo.org","avatar":"https://avatars.githubusercontent.com/u/373898?v=4"},"body":"On Sun, Apr 05, 2009 at 02:05:36AM +0200, Nicolas Sebrecht wrote:\n> > Our full repository conversion is large, even after tuning the\n> > repacking, the packed repository is between 1.4 and 1.6GiB. As of Feburary\n> > 4th, 2009, it contained 4886949 objects. It is not suitable for\n> > splitting into submodules either unfortunately - we have a lot of\n> > directory moves that would cause submodule bloat.\n> Actually, I'm not sure that a full portage tree repository would be the\n> best thing to do. It would not be suitable in the long term and working\n> on the repository/history would be a big mess. Why provide a such repo ?\n> Or at least, why provide a such readable repo ?\n> \n> IMHO, you should provide a repository per upstream package on the main\n> server.\nThat causes incredibly bloat unfortunately.\n\nI'll summarize why here for the git mailing list. Most our developers\nhave the entire tree checked out, and in informal surveys, would like to\ncontinue to do so. There are ~13500 packages right now (I'm excluding\neclasses/, profiles/, scripts/), and growing by 15-25 new packages/week.\n(~45% of packages also have a files/ directory).\n\nFor each package, the .git directory, assuming in a single pack,\nconsumes at least 36 inodes.  Tail-packing is limited to Reiserfs3 and\nJFS, and isn't widely used other than that, so assuming 4KiB inodes,\nthat's an overhead of at least 144KiB per package. Multiple by the\nnumber of packages, and we get an overhead of 2GiB, before we've added\nANY content.\n\nWithout tail packing, the Gentoo tree is presently around 520MiB (you\ncan fit it into ~190MiB with tail packing). This means that\nrepo-per-package would have an overhead in the range of 400%.\n\nAdditionally, there's a lot of commonality between ebuilds and packages,\nand having repo-per-package means that the compression algorithms can't\nmake use of it - dictionary algorithms are effective at compression for\na reason.\n\nOverhead is the reason that we refused to migrate to SVN as well.\n- CVS, per each directory of data, has a constant overhead of 4 inodes\n  (CVS/ CVS/Root CVS/Repository CVS/Entries)\n- SVN, for each data directory, has another complete copy of the data,\n  plus a minimum of 10 other inodes.\n- Git costs a minimum 36 inodes per repository. In a fully packed repo,\n  the number of inodes tends to stay below 50 in all cases.\n\n> PS: what about cc'ing gentoo-scm list ?\nIt's not an open-posting list, so anybody here on the git list simply\nreplying would not get their post on there. The issue has been raised\nthere, and this mainly meant to find a resolution to that problem.\n\n-- \nRobin Hugh Johnson\nGentoo Linux Developer & Infra Guy\nE-Mail     : robbat2@gentoo.org\nGnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n"},{"id":"110379","messageId":"20090405035453.GB12927@vidovic","threadId":"18724","inReplyTo":"20090405T001239Z@curie.orbis-terrarum.net","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Sebrecht","fromEmail":"nicolas.s-dev@laposte.net","sentAt":"2009-04-05T03:54:53Z","receivedAt":"2009-04-05T03:54:53Z","isPatch":false,"sender":{"key":"nicolas.s-dev@laposte.net","avatar":null},"body":"On Sat, Apr 04, 2009 at 05:37:53PM -0700, Robin H. Johnson wrote:\n\n> That causes incredibly bloat unfortunately.\n> \n> I'll summarize why here for the git mailing list. Most our developers\n> have the entire tree checked out, and in informal surveys, would like to\n> continue to do so. There are ~13500 packages right now \n\nEach developer doesn't work on so many packages, right ? From my point\nof view, checkin'out the entire tree is the wrong way on how to do\nthings.\n\nAlso, you could keep an entire tree repo assuming it's _not_\n\"fetch-able\".\n\n> For each package, the .git directory, assuming in a single pack,\n> consumes at least 36 inodes.  Tail-packing is limited to Reiserfs3 and\n> JFS, and isn't widely used other than that, so assuming 4KiB inodes,\n> that's an overhead of at least 144KiB per package. Multiple by the\n> number of packages, and we get an overhead of 2GiB, before we've added\n> ANY content.\n\n> Without tail packing, the Gentoo tree is presently around 520MiB (you\n> can fit it into ~190MiB with tail packing). This means that\n> repo-per-package would have an overhead in the range of 400%.\n\nDon't know about the business for Gentoo, but HDD is cheap. Also, I'd\nlike to know how much space you will gain with the CVS to Git migration.\nHow bigger is a CVS repo against a Git one ?\n\nOne repo per category could be a good compromise assuming one seperate\nbranch per ebuild, then.\n\n> Additionally, there's a lot of commonality between ebuilds and packages,\n> and having repo-per-package means that the compression algorithms can't\n> make use of it - dictionary algorithms are effective at compression for\n> a reason.\n\nPlease, no. We are in the long term issues. Compression will be\nefficient. It's all about the content of the files and dictionary\nalgorithms certainly will do a good job over the ebuilds revisions.\n\n-- \nNicolas Sebrecht\n"},{"id":"110381","messageId":"20090405040831.GA18892@vidovic","threadId":"18724","inReplyTo":"20090405035453.GB12927@vidovic","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Sebrecht","fromEmail":"nicolas.s-dev@laposte.net","sentAt":"2009-04-05T04:08:31Z","receivedAt":"2009-04-05T04:08:31Z","isPatch":false,"sender":{"key":"nicolas.s-dev@laposte.net","avatar":null},"body":"On Sun, Apr 05, 2009 at 05:54:53AM +0200, Nicolas Sebrecht wrote:\n\n> One repo per category could be a good compromise assuming one seperate\n> branch per ebuild, then.\n\ns/ebuild/package/\n\n-- \nNicolas Sebrecht\n"},{"id":"110385","messageId":"20090405070412.GB869@curie-int","threadId":"18724","inReplyTo":"20090405035453.GB12927@vidovic","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Robin H. Johnson","fromEmail":"robbat2@gentoo.org","sentAt":"2009-04-05T07:04:12Z","receivedAt":"2009-04-05T07:04:12Z","isPatch":false,"sender":{"key":"robbat2@gentoo.org","avatar":"https://avatars.githubusercontent.com/u/373898?v=4"},"body":"Before I answer the rest of your post, I'd like to note that the matter\nof which choice between single-repo, repo-per-package, repo-per-category\nhas been flogged to death within Gentoo.\n\nI did not come to the Git mailing list to rehash those choices. I came\nhere to find a solution to the performance problem. While it shows up\nwith our repo, I'm certain that we're not the only people with the\nproblem. The GSoC 2009 ideas contain a potential project for caching the\ngenerated packs, which, while having value in itself, could be partially\navoided by sending suitable pre-built packs (if they exist) without any\nrepacking.\n\nOn Sun, Apr 05, 2009 at 05:54:53AM +0200, Nicolas Sebrecht wrote:\n> > That causes incredibly bloat unfortunately.\n> > \n> > I'll summarize why here for the git mailing list. Most our developers\n> > have the entire tree checked out, and in informal surveys, would like to\n> > continue to do so. There are ~13500 packages right now \n> Each developer doesn't work on so many packages, right ? From my point\n> of view, checkin'out the entire tree is the wrong way on how to do\n> things.\nAlso, I should note that working on the tree isn't the only reason to\nhave the tree checked out. While the great majority of Gentoo users have\ntheir trees purely from rsync, there is nothing stopping you from using\na tree from CVS (anonCVS for the users, master CVS server for the\ndevelopers).\n\nA quick bit of stats run show that while some developers only touch a\nfew packages, there are at least 200 developers that have done a major\nchange to 100 or more packages.\n\n> > Without tail packing, the Gentoo tree is presently around 520MiB (you\n> > can fit it into ~190MiB with tail packing). This means that\n> > repo-per-package would have an overhead in the range of 400%.\n> Don't know about the business for Gentoo, but HDD is cheap.\nThere's no reason to have bloat just for the layout to change.\n\n> Also, I'd like to know how much space you will gain with the CVS to Git >\n> migration.  How bigger is a CVS repo against a Git one ?\nFor the CVS checkouts right now: \n- ~410MiB of content (w/ 4kb inodes)\n- ~240MiB of CVS overhead (w/ 4kb inodes)\n(sorry about the earlier 520MiB number, I forgot to exclude a local dir\nof stats data on my box when I ran du quickly).\n\nOur experimental Git, with only a single repo for gentoo-x86:\n- ~410MiB of content (w/ 4kb inodes)\n- 80MiB - 1.6GiB of Git total overhead.\n\n80MiB of overhead is the total overhead with a shallow clone at depth 1.\n1.6GiB is with the full history.\n\nAnd per-package numbers, because we DID do an experimental conversion,\nlast year, although the packs might not have been optimal:\n- ~410MiB of content (w/ 4kb inodes)\n- 4.7GiB of Git total overhead, with a breakdown:\n  - 1.9GiB in inode waste\n  - 2.8GiB in packs\n\n> One repo per category could be a good compromise assuming one seperate\n> branch per package, then.\nOther downsides to repo-per-category and repo-per-package:\n- Raises difficulty in adding a new package/category. \n  You cannot just do 'mkdir && vi ... && git add && git commit' anymore.\n- The name of the directory for both of the category AND the package are not\n  specified in the ebuild, as such, unless they are checked out to the right\n  location, you will get breakage (definitely in the package name, and\n  about 10% of the time with categories).\n- You cannot use git-cvsserver with them cleanly and have the correct\n  behavior (we DO have developers that want to use the CVS emulation\n  layer) - adding a category or a package would NOT trigger the\n  addition of a new repo on the server when needed.\n- Does NOT present a good base for anybody wanting to branch the entire\n  tree themselves.\n  \n\n> > Additionally, there's a lot of commonality between ebuilds and packages,\n> > and having repo-per-package means that the compression algorithms can't\n> > make use of it - dictionary algorithms are effective at compression for\n> > a reason.\n> Please, no. We are in the long term issues. Compression will be\n> efficient. It's all about the content of the files and dictionary\n> algorithms certainly will do a good job over the ebuilds revisions.\nWe're already on track to drop the CVS $Header$, and thereafter, some of the\nebuilds are already on track to be smaller. Here's our prototype dev-perl/Sub-Name-0.04.\n====\n# Copyright 1999-2009 Gentoo Foundation\n# Distributed under the terms of the GNU General Public License v2\nMODULE_AUTHOR=XMATH\ninherit perl-module\nDESCRIPTION=\"(re)name a sub\"\nLICENSE=\"|| ( Artistic GPL-2 )\"\nSLOT=\"0\"\nKEYWORDS=\"~amd64 ~x86\"\nIUSE=\"\"\nSRC_TEST=do\n====\n\nWe can have all the CPAN packages from CPAN author XMATH, with changing\nonly the DESCRIPTION string. KEYWORDS then just changes over the package\nlifespan.\n\n-- \nRobin Hugh Johnson\nGentoo Linux Developer & Infra Guy\nE-Mail     : robbat2@gentoo.org\nGnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n"},{"id":"110428","messageId":"20090405190213.GA12929@vidovic","threadId":"18724","inReplyTo":"20090405070412.GB869@curie-int","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Sebrecht","fromEmail":"nicolas.s-dev@laposte.net","sentAt":"2009-04-05T19:02:13Z","receivedAt":"2009-04-05T19:02:13Z","isPatch":false,"sender":{"key":"nicolas.s-dev@laposte.net","avatar":null},"body":"On Sun, Apr 05, 2009 at 12:04:12AM -0700, Robin H. Johnson wrote:\n\n> Before I answer the rest of your post, I'd like to note that the matter\n> of which choice between single-repo, repo-per-package, repo-per-category\n> has been flogged to death within Gentoo.\n> \n> I did not come to the Git mailing list to rehash those choices. I came\n> here to find a solution to the performance problem.\n\nI understand. I know two ways to resolve this:\n- by resolving the performance problem itself,\n- by changing the workflow to something more accurate and more suitable\n  against the facts.\n\nMy point is that going from a centralized to a decentralized SCM\ninvolves breacking strongly how developers and maintainers work. What\nyou're currently suggesting is a way to work with Git in a centralized\nway. This sucks. To get the things right with Git I would avoid shared\nand global repositories. Gnome is doing it this way:\nhttp://gitorious.org/projects/gnome-svn-hooks/repos/mainline/trees/master\n\n>          The GSoC 2009 ideas contain a potential project for caching the\n> generated packs, which, while having value in itself, could be partially\n> avoided by sending suitable pre-built packs (if they exist) without any\n> repacking.\n\nRight. It could be an option to wait and see if the GSoC gives\nsomething.\n\n> Also, I should note that working on the tree isn't the only reason to\n> have the tree checked out. While the great majority of Gentoo users have\n> their trees purely from rsync, there is nothing stopping you from using\n> a tree from CVS (anonCVS for the users, master CVS server for the\n> developers).\n> \n> A quick bit of stats run show that while some developers only touch a\n> few packages, there are at least 200 developers that have done a major\n> change to 100 or more packages.\n\nThat's a point that has to be reconsidered. Not the fact that at least\n200 developers work on over 100 packages (this is really not an issue)¹\nbut the fact that they do that directly on the main repo/server. The\ngood way to achieve this is to send his work to the maintainer². The main\nissue is a better code reviewing.\n\n1. Some or all repo-per-category can be tracked with a simple script.\n2. Maintainers could be - or not be - the same developers as today.\nAdding a layer of maintainers in charge of EAPI review (for example) up\nto the packages-maintainers could help in fixing a lot of portage issues\nand would avoid \"simple developers\" to do crap on the main repo(s) that\nusers download.\n\n> And per-package numbers, because we DID do an experimental conversion,\n> last year, although the packs might not have been optimal:\n> - ~410MiB of content (w/ 4kb inodes)\n> - 4.7GiB of Git total overhead, with a breakdown:\n>   - 1.9GiB in inode waste\n>   - 2.8GiB in packs\n\nOk.\n\n> > One repo per category could be a good compromise assuming one seperate\n> > branch per package, then.\n> Other downsides to repo-per-category and repo-per-package:\n\nLet's forget a repo-per-package.\n\n> - Raises difficulty in adding a new package/category. \n>   You cannot just do 'mkdir && vi ... && git add && git commit' anymore.\n\nRight, but categories are not evolving that much.\n\n> - The name of the directory for both of the category AND the package are not\n>   specified in the ebuild, as such, unless they are checked out to the right\n>   location, you will get breakage (definitely in the package name, and\n>   about 10% of the time with categories).\n\nOf course. Quite franckly, it's recoverable without pain.\n\nA repo-per-category local workflow would be:\n$ git branch\n  master\n* next\n  package_one\n  package_two\n  [...]\n$ tree -a\n|-- .git\n|   |-- [...]\n|   [...]\n|-- package_one\n|   |-- ChangeLog\n|   |-- Manifest\n|   |-- metadata.xml\n|   |-- package_one-0.4.ebuild\n|   `-- package_one-0.5.ebuild\n|-- package_two\n|   |-- ChangeLog\n|   |-- Manifest\n|   |-- files\n|   |   |-- package_two.confd\n|   |   `-- package_two.rc\n|   |-- metadata.xml\n|   `-- package_two-0.7-r3.ebuild\n[...]\n\n$ git checkout package_one\n$ tree -a\n|-- .git\n|   |-- [...]\n|   [...]\n`-- package_one\n    |-- ChangeLog\n    |-- Manifest\n    |-- metadata.xml\n    |-- package_one-0.4.ebuild\n    `-- package_one-0.5.ebuild\n$ <hack, hack, hack>\n$ git checkout next\n$ git merge package_one \n\n> - Does NOT present a good base for anybody wanting to branch the entire\n>   tree themselves.\n\nScriptable.\n\n> We're already on track to drop the CVS $Header$, and thereafter, some of the\n> ebuilds are already on track to be smaller. Here's our prototype dev-perl/Sub-Name-0.04.\n> ====\n> # Copyright 1999-2009 Gentoo Foundation\n> # Distributed under the terms of the GNU General Public License v2\n> MODULE_AUTHOR=XMATH\n> inherit perl-module\n> DESCRIPTION=\"(re)name a sub\"\n> LICENSE=\"|| ( Artistic GPL-2 )\"\n> SLOT=\"0\"\n> KEYWORDS=\"~amd64 ~x86\"\n> IUSE=\"\"\n> SRC_TEST=do\n> ====\n> \n> We can have all the CPAN packages from CPAN author XMATH, with changing\n> only the DESCRIPTION string. KEYWORDS then just changes over the package\n> lifespan.\n\nSounds good.\n\n-- \nNicolas Sebrecht\n"},{"id":"110431","messageId":"20090405191703.GJ23521@spearce.org","threadId":"18724","inReplyTo":"20090405190213.GA12929@vidovic","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2009-04-05T19:17:03Z","receivedAt":"2009-04-05T19:17:03Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Nicolas Sebrecht <nicolas.s-dev@laposte.net> wrote:\n> On Sun, Apr 05, 2009 at 12:04:12AM -0700, Robin H. Johnson wrote:\n> >          The GSoC 2009 ideas contain a potential project for caching the\n> > generated packs, which, while having value in itself, could be partially\n> > avoided by sending suitable pre-built packs (if they exist) without any\n> > repacking.\n> \n> Right. It could be an option to wait and see if the GSoC gives\n> something.\n\nAnother option is to use rsync:// for initial clones.\n \nTell new developers that their initial command sequence to\n(efficiently) get the base tree is:\n\n  git clone rsync://git.gentoo.org/tree.git\n  cd tree\n  git config remote.origin.url git://git.gentoo.org/tree.git\n\nrsync should be more efficient at dragging 1.6GiB over the network,\nas its only streaming the files.  But it may fall over if the server\nhas a lot of loose objects; many more small files to create.\n\nOne way around that would be to use two repositories on the server;\na historical repository that is fully packed and contains the full\nhistory, and a bleeding edge repository that users would normally\nwork against:\n\n  git clone rsync://git.gentoo.org/fully-packed-tree.git tree\n  cd tree\n  git config remote.origin.url git://git.gentoo.org/tree.git\n  git pull\n\nThen every so often (e.g. once a Gentoo release cycle, so once\na year) pull the bleeding edge repository into the fully packed\nrepository.  That will introduce a single new pack file, so the\nfully packed repository grows at a rate of 2 inodes/year, and is\nstill very efficient to rsync on initial clones.\n\n\nThat caching GSoC project may help, but didn't I see earlier in\nthis thread that you have >4.8 million objects in your repository?\nAny proposals on that project would still have Git malloc()'ing\ndata per object; its ~80 bytes per object needed so that's a data\nsegment of 384+ MiB, per concurrent clone client.\n\n-- \nShawn.\n"},{"id":"110440","messageId":"20090405195714.GA4716@coredump.intra.peff.net","threadId":"18724","inReplyTo":"20090404220743.GA869@curie-int","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-04-05T19:57:14Z","receivedAt":"2009-04-05T19:57:14Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Apr 04, 2009 at 03:07:43PM -0700, Robin H. Johnson wrote:\n\n> During an initial clone, I see that git-upload-pack invokes\n> pack-objects, despite the ENTIRE repository already being packed - no\n> loose objects whatsoever. git-upload-pack then seems to buffer in\n> memory.\n\nWe need to run pack-objects even if the repo is fully packed because we\ndon't know what's _in_ the existing pack (or packs). In particular we\nwant to:\n\n  - combine multiple packs into a single pack; this is more efficient on\n    the network, because you can find more deltas, and I believe is\n    required because the protocol sends only a single pack.\n\n  - cull any objects which are not actually part of the reachability\n    chain from the refs we are sending\n\nIf no work needs to be done for either case, then pack-objects should\nbasically just figure that out and then send the existing pack (the\nexpensive bit is doing deltas, and we don't consider objects in the same\npack for deltas, as we know we have already considered that during the\nlast repack). It does mmap the whole pack, so you will see your virtual\nmemory jump, but nothing should require the whole pack being in memory\nat once.\n\npack-objects streams the output to upload-pack, which should only ever\nhave an 8K buffer of it in memory at any given time.\n\nAt least that is how it is all supposed to work, according to my\nunderstanding. So if you are seeing very high memory usage, I wonder if\nthere is a bug in pack-objects or upload-pack that can be fixed.\n\nMaybe somebody more knowledgeable than me about packing can comment.\n\n> During 'remote: Counting objects: 4886949, done.', git-upload-pack peaks at\n> 2474216KB VSZ and 1143048KB RSS. \n> Shortly thereafter, we get 'remote: Compressing objects:   0%\n> (1328/1994284)', git-pack-objects with ~2.8GB VSZ and ~1.8GB RSS. Here,\n> the CPU burn also starts. On our test server machine (w/ git 1.6.0.6),\n> it takes about 200 minutes walltime to finish the pack, IFF the OOM\n> doesn't kick in.\n\nHave you tried with a more recent git to see if it is any better? There\nhave been a number of changes since 1.6.0.6, although it looks like\nmostly dealing with better recovery from corrupted packs.\n\n> Given that the repo is entirely packed already, I see no point in doing\n> this.\n> \n> For the initial clone, can the git-upload-pack algorithm please send\n> existing packs, and only generate a pack containing the non-packed\n> items?\n\nI believe that would require a change to the protocol to allow multiple\npacks. However, it may be possible to munge the pack header in such a\nway that you basically concatenate multiple packs. You would still want\nto peek in the big pack to try deltas from the non-packed items, though.\n\nI think all of this falls into the realm of the GSOC pack caching project.\nThere have been other discussions on the list, so you might want to look\nthrough those for something useful.\n\n-Peff\n"},{"id":"110455","messageId":"20090405204325.GA31344@curie-int","threadId":"18724","inReplyTo":"20090405190213.GA12929@vidovic","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Robin H. Johnson","fromEmail":"robbat2@gentoo.org","sentAt":"2009-04-05T20:43:25Z","receivedAt":"2009-04-05T20:43:25Z","isPatch":false,"sender":{"key":"robbat2@gentoo.org","avatar":"https://avatars.githubusercontent.com/u/373898?v=4"},"body":"On Sun, Apr 05, 2009 at 09:02:13PM +0200, Nicolas Sebrecht wrote:\n> > Before I answer the rest of your post, I'd like to note that the matter\n> > of which choice between single-repo, repo-per-package, repo-per-category\n> > has been flogged to death within Gentoo.\n> > \n> > I did not come to the Git mailing list to rehash those choices. I came\n> > here to find a solution to the performance problem.\n> I understand. I know two ways to resolve this:\n> - by resolving the performance problem itself,\n> - by changing the workflow to something more accurate and more suitable\n>   against the facts.\n> \n> My point is that going from a centralized to a decentralized SCM\n> involves breacking strongly how developers and maintainers work. What\n> you're currently suggesting is a way to work with Git in a centralized\n> way. This sucks. To get the things right with Git I would avoid shared\n> and global repositories. Gnome is doing it this way:\n> http://gitorious.org/projects/gnome-svn-hooks/repos/mainline/trees/master\nThe entire matter of splitting the repository comes down to what should\nbe considered an atomic unit. For GNOME, KDE and all of the other large\nGit consumers that I'm aware of, there atomic units are individual\npackages - specifically because they make sense to be consumed without\nhaving all the rest of the packages. For the gentoo tree, it is an\natomic unit in itself. Changes to the profiles/ directory (for package\nmasks, USE keys are frequently related and need to be always committed\nand received atomically with changes to one or more packages.\n\n> >          The GSoC 2009 ideas contain a potential project for caching the\n> > generated packs, which, while having value in itself, could be partially\n> > avoided by sending suitable pre-built packs (if they exist) without any\n> > repacking.\n> Right. It could be an option to wait and see if the GSoC gives\n> something.\nHow hard is it to just look at the git-upload-pack code and make it\nrealize that it doesn't need to repack at all for this case.\n\n> > A quick bit of stats run show that while some developers only touch a\n> > few packages, there are at least 200 developers that have done a major\n> > change to 100 or more packages.\n> That's a point that has to be reconsidered. Not the fact that at least\n> 200 developers work on over 100 packages (this is really not an issue)¹\n> but the fact that they do that directly on the main repo/server. The\n> good way to achieve this is to send his work to the maintainer². The main\n> issue is a better code reviewing.\nThis has been shot down by our developer base. One of the grounds is\nthat there is no developer with sufficient time to take a merge-master\nrole on a regular basis like that.\n\n> 1. Some or all repo-per-category can be tracked with a simple script.\n> 2. Maintainers could be - or not be - the same developers as today.\n> Adding a layer of maintainers in charge of EAPI review (for example) up\n> to the packages-maintainers could help in fixing a lot of portage issues\n> and would avoid \"simple developers\" to do crap on the main repo(s) that\n> users download.\nYou imply that there is a problem in that field already, which I\ndisagree with.\n\n> > > One repo per category could be a good compromise assuming one seperate\n> > > branch per package, then.\n> > Other downsides to repo-per-category and repo-per-package:\n> Let's forget a repo-per-package.\nOne downside unique to repo-per-category is that when a package moves\ncross-category, you end up with it consuming space in packs on both\nsides.\n\n> > - Raises difficulty in adding a new package/category. \n> >   You cannot just do 'mkdir && vi ... && git add && git commit' anymore.\n> Right, but categories are not evolving that much.\nThere's demand to evolve them, but bulk package moves are painful with\nCVS, so it's been waiting for Git.\n\n> A repo-per-category local workflow would be:\n> [...]\n> $ git checkout package_one\n> $ tree -a\n> |-- .git\n> |   |-- [...]\n> |   [...]\n> `-- package_one\n>     |-- ChangeLog\n>     |-- Manifest\n>     |-- metadata.xml\n>     |-- package_one-0.4.ebuild\n>     `-- package_one-0.5.ebuild\nUmm, why does package_two not exist in the other branch?\nIf package_one depends on package_two, and you're in for a world of fail\nthe moment it you changes branches here.\n\n> > - Does NOT present a good base for anybody wanting to branch the entire\n> >   tree themselves.\n> Scriptable.\nYou dropped my cvsserver list item.\n\n-- \nRobin Hugh Johnson\nGentoo Linux Developer & Infra Guy\nE-Mail     : robbat2@gentoo.org\nGnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n"},{"id":"110458","messageId":"20090405210828.GB23604@spearce.org","threadId":"18724","inReplyTo":"20090405204325.GA31344@curie-int","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2009-04-05T21:08:28Z","receivedAt":"2009-04-05T21:08:28Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"Robin H. Johnson\" <robbat2@gentoo.org> wrote:\n> > >          The GSoC 2009 ideas contain a potential project for caching the\n> > > generated packs, which, while having value in itself, could be partially\n> > > avoided by sending suitable pre-built packs (if they exist) without any\n> > > repacking.\n> > Right. It could be an option to wait and see if the GSoC gives\n> > something.\n>\n> How hard is it to just look at the git-upload-pack code and make it\n> realize that it doesn't need to repack at all for this case.\n\nI don't need to go look.  I know that code.\n\nIts harder than you think.\n\nI'll tell you what, *you* go look at the git-upload-pack code and\ncome back with a patch that doesn't need to repack at all for this\ncase, *and* which Junio will actually apply.  If its any good,\nJunio would apply it pretty quickly.\n\nNobody else has managed to create such a patch just yet.  Because its\nextremely non-trivial.  Its pretty much never the case that an active\nrepository is fully repacked, so we always have to enumerate some\nnumber of loose objects and put them into a single outgoing pack\nfor the network.  Its also considered to be a security feature of\nGit that we only transmit reachable objects.\n\n-- \nShawn.\n"},{"id":"110463","messageId":"alpine.DEB.1.10.0904051419490.6245@asgard.lang.hm","threadId":"18724","inReplyTo":"20090405190213.GA12929@vidovic","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"","fromEmail":"david@lang.hm","sentAt":"2009-04-05T21:28:35Z","receivedAt":"2009-04-05T21:28:35Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Sun, 5 Apr 2009, Nicolas Sebrecht wrote:\n\n> On Sun, Apr 05, 2009 at 12:04:12AM -0700, Robin H. Johnson wrote:\n>\n>> Before I answer the rest of your post, I'd like to note that the matter\n>> of which choice between single-repo, repo-per-package, repo-per-category\n>> has been flogged to death within Gentoo.\n>>\n>> I did not come to the Git mailing list to rehash those choices. I came\n>> here to find a solution to the performance problem.\n>\n> I understand. I know two ways to resolve this:\n> - by resolving the performance problem itself,\n> - by changing the workflow to something more accurate and more suitable\n>  against the facts.\n>\n> My point is that going from a centralized to a decentralized SCM\n> involves breacking strongly how developers and maintainers work. What\n> you're currently suggesting is a way to work with Git in a centralized\n> way. This sucks. To get the things right with Git I would avoid shared\n> and global repositories. Gnome is doing it this way:\n> http://gitorious.org/projects/gnome-svn-hooks/repos/mainline/trees/master\n\nguys, back off a little on telling the gentoo people to change. the kernel \ndevelopers don't split th kernel into 'core' 'drivers' etc pieces just \nbecause some people only work on one area. I see the gentoo desire to keep \nthings in one repo as being something very similar.\n\nthe problem here is a real one, if you have a large repo, git send-pack \nwill always generate a new pack, even if it doesn't need to (with the \nextreme case being the the repo is fully packed)\n\n>>          The GSoC 2009 ideas contain a potential project for caching the\n>> generated packs, which, while having value in itself, could be partially\n>> avoided by sending suitable pre-built packs (if they exist) without any\n>> repacking.\n>\n> Right. It could be an option to wait and see if the GSoC gives\n> something.\n\nthe GSOC project is not the same thing. in this case the packs are already \n'cached' (they are stored on disk), what is needed is some option to let \ngit send existing pack(s) if they exist rather then taking the time to \ntry and generate an 'optimal' pack.\n\nI'm actually aurprised that this is happening, I thought that the \nrecommendation was that the public repository should do a very agressive \npack (that takes a lot of resources) for the old content so that people \ncloning from it get the advantage of the tight packing without having to \ndo it themselves.\n\nif the server _always_ re-generates the pack from scratch then this is a \nwaste of time (except for people who clone via the dumb, unsafe \nmechanisms)\n\nDavid Lang\n"},{"id":"110466","messageId":"fabb9a1e0904051436i1dc9c1bdhe86a23e470c756f9@mail.gmail.com","threadId":"18724","inReplyTo":"alpine.DEB.1.10.0904051419490.6245@asgard.lang.hm","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Sverre Rabbelier","fromEmail":"srabbelier@gmail.com","sentAt":"2009-04-05T21:36:19Z","receivedAt":"2009-04-05T21:36:19Z","isPatch":false,"sender":{"key":"srabbelier@gmail.com","avatar":"https://avatars.githubusercontent.com/u/3098?v=4"},"body":"Heya,\n\nOn Sun, Apr 5, 2009 at 23:28,  <david@lang.hm> wrote:\n> Guys, back off a little on telling the gentoo people to change.\n\nI agree here, we should either say \"look, we don't really support big\nrepositories because [explanation here], unless you [workarounds\nhere]\" OR we should work to improve the support we do have. Of course,\nthe latter option does not magically create developer time to work on\nthat, but if we do go that way we should at least tell people that we\nare aware of the problems and that it's on the global TODO list (not\nnecessarily on anyone's personal TODO list though).\nOf course, the problem can sometimes be solved by splitting the\nrepository, but I think it is important to have an official policy\nhere, do we want Git to support huge repositories, or do we not?\n\n-- \nCheers,\n\nSverre Rabbelier\n"},{"id":"110480","messageId":"20090405225954.GA18730@vidovic","threadId":"18724","inReplyTo":"alpine.DEB.1.10.0904051419490.6245@asgard.lang.hm","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Sebrecht","fromEmail":"nicolas.s-dev@laposte.net","sentAt":"2009-04-05T22:59:54Z","receivedAt":"2009-04-05T22:59:54Z","isPatch":false,"sender":{"key":"nicolas.s-dev@laposte.net","avatar":null},"body":"On Sun, Apr 05, 2009 at 02:28:35PM -0700, david@lang.hm wrote:\n\n> guys, back off a little on telling the gentoo people to change.\n\nDon't blame Git people, please. I currently am the only one here to\ndiscuss that way and see a painful work coming at Gentoo.\nGit people didn't discuss around thoses issues.\n\n>                                                                 the \n> kernel developers don't split th kernel into 'core' 'drivers' etc pieces \n> just because some people only work on one area.\n\nAnd you might notice that they don't provide a CVS access and actually\ndon't work around an unique shared repo. Also, you might notice that\nkeeping the history clean to assure the work on the kernel easier is not\nan elementary issue.\n\n> just because some people only work on one area. I see the gentoo desire \n> to keep things in one repo as being something very similar.\n\nThat's why I think the gentoo desire is not very clean (don't be\naffected). What I see is that in one hand you want a DSCM and on the\nother hand you want to keep a central shared repo.\n\n> the problem here is a real one, if you have a large repo, git send-pack  \n> will always generate a new pack, even if it doesn't need to (with the  \n> extreme case being the the repo is fully packed)\n\nWhat about the rsync solution given in this thread?\n\n-- \nNicolas Sebrecht\n"},{"id":"110481","messageId":"20090405230219.GB31344@curie-int","threadId":"18724","inReplyTo":"20090405191703.GJ23521@spearce.org","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Robin H. Johnson","fromEmail":"robbat2@gentoo.org","sentAt":"2009-04-05T23:02:19Z","receivedAt":"2009-04-05T23:02:19Z","isPatch":false,"sender":{"key":"robbat2@gentoo.org","avatar":"https://avatars.githubusercontent.com/u/373898?v=4"},"body":"On Sun, Apr 05, 2009 at 12:17:03PM -0700, Shawn O. Pearce wrote:\n> Another option is to use rsync:// for initial clones.\n>   git clone rsync://git.gentoo.org/tree.git\n> rsync should be more efficient at dragging 1.6GiB over the network,\n> as its only streaming the files.  But it may fall over if the server\n> has a lot of loose objects; many more small files to create.\nI just tried this, and ran into a segfault.\n\nOriginal command:\n# git clone rsync://git.overlays.gentoo.org/vcs-public-gitroot/exp/gentoo-x86.git\n\nIt looks at a glance like the linked list has a null value it hits during the\ninternal while loop, not checking 'list' before using 'list->next'.\n\ngdb> bt\n#0  strcmp () at ../sysdeps/x86_64/strcmp.S:30\n#1  0x000000000049474c in get_refs_via_rsync (transport=<value optimized out>, for_push=<value optimized out>) at transport.c:123\n#2  0x000000000049234c in transport_get_remote_refs (transport=0x725fc9) at transport.c:1045\n#3  0x000000000041620a in cmd_clone (argc=<value optimized out>, argv=0x7fff908c8550, prefix=<value optimized out>) at builtin-clone.c:487\n#4  0x0000000000404f59 in handle_internal_command (argc=0x2, argv=0x7fff908c8550) at git.c:244\n#5  0x0000000000405167 in main (argc=0x2, argv=0x7fff908c8550) at git.c:434\ngdb> up\n#1  0x000000000049474c in get_refs_via_rsync (transport=<value optimized out>, for_push=<value optimized out>) at transport.c:123\n123\t\t\t\t\t(cmp = strcmp(buffer + 41,\ngdb> print list\n$1 = {nr = 0x0, alloc = 0x0, name = 0x0}\n\nIf I go into the repo thereafter and manually run git-fetch again, it does work\nfine.\n\n> One way around that would be to use two repositories on the server;\n> a historical repository that is fully packed and contains the full\n> history, and a bleeding edge repository that users would normally\n> work against:\nYup, we've been considering similar. We do have one specific need with that\nhowever: to prevent resource abuse, we would like to DENY the ability to do the\ninitial clone with git:// then - just so that nobody tries to DoS our servers\nby doing a couple of hungry initial clones at once.\n\n> That caching GSoC project may help, but didn't I see earlier in\n> this thread that you have >4.8 million objects in your repository?\n> Any proposals on that project would still have Git malloc()'ing\n> data per object; its ~80 bytes per object needed so that's a data\n> segment of 384+ MiB, per concurrent clone client.\n384MiB or even 512MiB I can cover. It's the 200+ wallclock minutes of cpu burn\nwith no download that aren't acceptable.\n\nP.S.\nThe -v output of the rsync-mode git-fetch is very devoid of output. Can we\nmaybe pipe the rsync progress back?\n\n\n-- \nRobin Hugh Johnson\nGentoo Linux Developer & Infra Guy\nE-Mail     : robbat2@gentoo.org\nGnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n"},{"id":"110484","messageId":"alpine.DEB.1.10.0904051613420.6245@asgard.lang.hm","threadId":"18724","inReplyTo":"20090405225954.GA18730@vidovic","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"","fromEmail":"david@lang.hm","sentAt":"2009-04-05T23:20:07Z","receivedAt":"2009-04-05T23:20:07Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Mon, 6 Apr 2009, Nicolas Sebrecht wrote:\n\n> On Sun, Apr 05, 2009 at 02:28:35PM -0700, david@lang.hm wrote:\n>\n>> guys, back off a little on telling the gentoo people to change.\n>\n> Don't blame Git people, please. I currently am the only one here to\n> discuss that way and see a painful work coming at Gentoo.\n> Git people didn't discuss around thoses issues.\n>\n>>                                                                 the\n>> kernel developers don't split th kernel into 'core' 'drivers' etc pieces\n>> just because some people only work on one area.\n>\n> And you might notice that they don't provide a CVS access and actually\n> don't work around an unique shared repo. Also, you might notice that\n> keeping the history clean to assure the work on the kernel easier is not\n> an elementary issue.\n\nthese issues are completely seperate from the issue that the initial \nposter asked about, which is that when someone tries to do a clone of the \nrepository the system wastes a lot of time creating a new pack.\n\nthe kernel has a central public repo, they could run the cvs server on \nthat and still keep the rest of the kernel development exactly the way it \nis.\n\nif they are currently planning for one central repo with everyone pushing \nto it, I expect that they will change their workflow as they get used to \ngit, but that isn't going to address the problem in the tool.\n\n>> just because some people only work on one area. I see the gentoo desire\n>> to keep things in one repo as being something very similar.\n>\n> That's why I think the gentoo desire is not very clean (don't be\n> affected). What I see is that in one hand you want a DSCM and on the\n> other hand you want to keep a central shared repo.\n\ndon't worry about this part of things, worry about why the server wastes \nso many resources.\n\nif this is really what's happening, other projects will suffer as well \n(including the kernel, which has a very distributed workflow)\n\n>> the problem here is a real one, if you have a large repo, git send-pack\n>> will always generate a new pack, even if it doesn't need to (with the\n>> extreme case being the the repo is fully packed)\n>\n> What about the rsync solution given in this thread?\n\nthat may be a work-around for a situation where git just doesn't work, but \nhow do they prevent users from killing their server by trying to do a \nnormal git clone?\n\nDaivd Lang\n"},{"id":"110487","messageId":"200904060128.42095.robin.rosenberg.lists@dewire.com","threadId":"18724","inReplyTo":"alpine.DEB.1.10.0904051613420.6245@asgard.lang.hm","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg.lists@dewire.com","sentAt":"2009-04-05T23:28:41Z","receivedAt":"2009-04-05T23:28:41Z","isPatch":false,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"måndag 06 april 2009 01:20:07 skrev david@lang.hm:\n> >> the problem here is a real one, if you have a large repo, git send-pack\n> >> will always generate a new pack, even if it doesn't need to (with the\n> >> extreme case being the the repo is fully packed)\n> >\n> > What about the rsync solution given in this thread?\n> \n> that may be a work-around for a situation where git just doesn't work, but \n> how do they prevent users from killing their server by trying to do a \n> normal git clone?\n\nIs there no way of telling git not work so hard on packing? \n\nIf not, you could try JGit and compare. It's still too stupid to pack much, so it shouldn't spend much CPU time (for that reason at least). I haven't tried JGit's deamon for large amounts of data yet.\n\n-- robin\n"},{"id":"110488","messageId":"20090405T230552Z@curie.orbis-terrarum.net","threadId":"18724","inReplyTo":"20090405195714.GA4716@coredump.intra.peff.net","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Robin H. Johnson","fromEmail":"robbat2@gentoo.org","sentAt":"2009-04-05T23:38:31Z","receivedAt":"2009-04-05T23:38:31Z","isPatch":false,"sender":{"key":"robbat2@gentoo.org","avatar":"https://avatars.githubusercontent.com/u/373898?v=4"},"body":"On Sun, Apr 05, 2009 at 03:57:14PM -0400, Jeff King wrote:\n> > During an initial clone, I see that git-upload-pack invokes\n> > pack-objects, despite the ENTIRE repository already being packed - no\n> > loose objects whatsoever. git-upload-pack then seems to buffer in\n> > memory.\n> We need to run pack-objects even if the repo is fully packed because we\n> don't know what's _in_ the existing pack (or packs). In particular we\n> want to:\n>   - combine multiple packs into a single pack; this is more efficient on\n>     the network, because you can find more deltas, and I believe is\n>     required because the protocol sends only a single pack.\n> \n>   - cull any objects which are not actually part of the reachability\n>     chain from the refs we are sending\n> \n> If no work needs to be done for either case, then pack-objects should\n> basically just figure that out and then send the existing pack (the\n> expensive bit is doing deltas, and we don't consider objects in the same\n> pack for deltas, as we know we have already considered that during the\n> last repack). It does mmap the whole pack, so you will see your virtual\n> memory jump, but nothing should require the whole pack being in memory\n> at once.\nWhile my current pack setup has multiple packs of not more than 100MiB\neach, that was simply for ease of resume with rsync+http tests. Even\nwhen I already had a single pack, with every object reachable,\npack-objects was redoing the packing.\n\n> pack-objects streams the output to upload-pack, which should only ever\n> have an 8K buffer of it in memory at any given time.\n> \n> At least that is how it is all supposed to work, according to my\n> understanding. So if you are seeing very high memory usage, I wonder if\n> there is a bug in pack-objects or upload-pack that can be fixed.\n> \n> Maybe somebody more knowledgeable than me about packing can comment.\nLooking at the source, I agree that it should be buffering, however top and ps\nseem to disagree. 3GiB VSZ and 2.5GiB RSS here now.\n\n%CPU %MEM     VSZ     RSS STAT START   TIME COMMAND\n 0.0  0.0  140932    1040 Ss   16:09   0:00 \\_ git-upload-pack /code/gentoo/gentoo-git/gentoo-x86.git \n32.2  0.0       0       0 Z    16:09   1:50     \\_ [git-upload-pack] <defunct>\n80.8 44.2 3018484 2545700 Sl   16:09   4:36     \\_ git pack-objects --stdout --progress --delta-base-offset \n\nAlso, I did another trace, using some other hardware, in a LAN setting, and\nnoticed that git-upload-pack/pack-objects only seems to start output to the\nnetwork after it reaches 100% in 'remote: Compressing objects:'.\n\nRelatedly, throwing more RAM (6GiB total, vs. the previous 2GiB) at the server\nin this case cut the 200 wallclock minutes before any sending too place down to\n5 minutes.\n\n\n> > During 'remote: Counting objects: 4886949, done.', git-upload-pack peaks at\n> > 2474216KB VSZ and 1143048KB RSS. \n> > Shortly thereafter, we get 'remote: Compressing objects:   0%\n> > (1328/1994284)', git-pack-objects with ~2.8GB VSZ and ~1.8GB RSS. Here,\n> > the CPU burn also starts. On our test server machine (w/ git 1.6.0.6),\n> > it takes about 200 minutes walltime to finish the pack, IFF the OOM\n> > doesn't kick in.\n> Have you tried with a more recent git to see if it is any better? There\n> have been a number of changes since 1.6.0.6, although it looks like\n> mostly dealing with better recovery from corrupted packs.\nTesting right now, the above on the LAN setup was w/ current git HEAD.\n\n> > For the initial clone, can the git-upload-pack algorithm please send\n> > existing packs, and only generate a pack containing the non-packed\n> > items?\n> \n> I believe that would require a change to the protocol to allow multiple\n> packs. However, it may be possible to munge the pack header in such a\n> way that you basically concatenate multiple packs. You would still want\n> to peek in the big pack to try deltas from the non-packed items, though.\n> \n> I think all of this falls into the realm of the GSOC pack caching project.\n> There have been other discussions on the list, so you might want to look\n> through those for something useful.\nYes, both changing the protocol, and recognizing that existing packs may be\nsuitable to send could be considered as part of the caching project, as they\nfall under the aegis of making good use of what's stored in the cache already\nto send.\n\n-- \nRobin Hugh Johnson\nGentoo Linux Developer & Infra Guy\nE-Mail     : robbat2@gentoo.org\nGnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n"},{"id":"110489","messageId":"20090405T234147Z@curie.orbis-terrarum.net","threadId":"18724","inReplyTo":"20090405T230552Z@curie.orbis-terrarum.net","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Robin H. Johnson","fromEmail":"robbat2@gentoo.org","sentAt":"2009-04-05T23:42:35Z","receivedAt":"2009-04-05T23:42:35Z","isPatch":false,"sender":{"key":"robbat2@gentoo.org","avatar":"https://avatars.githubusercontent.com/u/373898?v=4"},"body":"On Sun, Apr 05, 2009 at 04:38:31PM -0700, Robin H. Johnson wrote:\n> > Have you tried with a more recent git to see if it is any better? There\n> > have been a number of changes since 1.6.0.6, although it looks like\n> > mostly dealing with better recovery from corrupted packs.\n> Testing right now, the above on the LAN setup was w/ current git HEAD.\nJust following up, there seems to be no significant change in results\nwith v1.6.2.2 over v1.6.0.6.\n\n-- \nRobin Hugh Johnson\nGentoo Linux Developer & Infra Guy\nE-Mail     : robbat2@gentoo.org\nGnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n"},{"id":"110494","messageId":"20090406T002445Z@curie.orbis-terrarum.net","threadId":"18724","inReplyTo":"0015174c150e49b5740466d7d2c2@google.com","subject":"Re: Re: Performance issue: initial git clone causes massive repack","fromName":"Robin H. Johnson","fromEmail":"robbat2@gentoo.org","sentAt":"2009-04-06T00:29:09Z","receivedAt":"2009-04-06T00:29:09Z","isPatch":false,"sender":{"key":"robbat2@gentoo.org","avatar":"https://avatars.githubusercontent.com/u/373898?v=4"},"body":"On Mon, Apr 06, 2009 at 12:17:18AM +0000, SRabbelier@gmail.com wrote:\n> Heya,\n>\n> On Mon, Apr 6, 2009 at 01:38, Robin H. Johnson robbat2@gentoo.org> wrote:\n>> Relatedly, throwing more RAM (6GiB total, vs. the previous 2GiB) at the \n>> server\n>> in this case cut the 200 wallclock minutes before any sending too place \n>> down to\n>> 5 minutes.\n> I'm curious what kind of hardware changes you made to achieve such an \n> enormous effect? Was it just added RAM on the same machine?\nNo, see the paragraph previous to that, showing it was a different machine that\njust happened to have 6GiB of RAM.\n\nThe key difference is that having 6GiB of RAM was enough to stop the\nswap/OOM-killing of git-pack-objects/git-upload-pack that happened on the slow\nserver, which I considered to be entirely unwarranted since the pack was\nalready generated and perfect for use.\n\n\"Slow\" server:\n- deadline scheduler\n- AMD Opteron 1210, single socket, 2 cores @ 1.8GHZ\n- 2GB Reg ECC RAM\n- 2x ST3250620AS, RAID1\n- 100Mbit internet feed, co-located in Texas.\n\n\"Fast\" server:\n- deadline scheduler\n- Intel Core2 Q6600, single socket, 4 cores @ 2.4GHz\n- ~5.7GiB of cheap RAM (6GiB in the box, 256 not usable due to BIOS MTRR brokeness)\n- 7x ST3320620AS, RAID-5.\n- Sitting on my home LAN, with a very crappy upload bandwidth to the Internet.\n\n-- \nRobin Hugh Johnson\nGentoo Linux Developer & Infra Guy\nE-Mail     : robbat2@gentoo.org\nGnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n"},{"id":"110504","messageId":"fcaeb9bf0904052010p34e3246bwd7e1f5297acf37e2@mail.gmail.com","threadId":"18724","inReplyTo":"20090405T230552Z@curie.orbis-terrarum.net","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-04-06T03:10:57Z","receivedAt":"2009-04-06T03:10:57Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Apr 6, 2009 at 9:38 AM, Robin H. Johnson <robbat2@gentoo.org> wrote:\n> Looking at the source, I agree that it should be buffering, however top and ps\n> seem to disagree. 3GiB VSZ and 2.5GiB RSS here now.\n>\n> %CPU %MEM     VSZ     RSS STAT START   TIME COMMAND\n>  0.0  0.0  140932    1040 Ss   16:09   0:00 \\_ git-upload-pack /code/gentoo/gentoo-git/gentoo-x86.git\n> 32.2  0.0       0       0 Z    16:09   1:50     \\_ [git-upload-pack] <defunct>\n> 80.8 44.2 3018484 2545700 Sl   16:09   4:36     \\_ git pack-objects --stdout --progress --delta-base-offset\n>\n> Also, I did another trace, using some other hardware, in a LAN setting, and\n> noticed that git-upload-pack/pack-objects only seems to start output to the\n> network after it reaches 100% in 'remote: Compressing objects:'.\n>\n> Relatedly, throwing more RAM (6GiB total, vs. the previous 2GiB) at the server\n> in this case cut the 200 wallclock minutes before any sending too place down to\n> 5 minutes.\n\nSearching back the archive, there was memory fragmentation issue with\ngcc repo. I wonder if it happens again. Maybe you should try Google\nallocator. BTW, did you try to turn off THREADED_DELTA_SEARCH?\n-- \nDuy\n"},{"id":"110507","messageId":"alpine.LFD.2.00.0904052315210.6741@xanadu.home","threadId":"18724","inReplyTo":"fabb9a1e0904051436i1dc9c1bdhe86a23e470c756f9@mail.gmail.com","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-06T03:24:27Z","receivedAt":"2009-04-06T03:24:27Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sun, 5 Apr 2009, Sverre Rabbelier wrote:\n\n> Heya,\n> \n> On Sun, Apr 5, 2009 at 23:28,  <david@lang.hm> wrote:\n> > Guys, back off a little on telling the gentoo people to change.\n> \n> I agree here, we should either say \"look, we don't really support big\n> repositories because [explanation here], unless you [workarounds\n> here]\" OR we should work to improve the support we do have. Of course,\n> the latter option does not magically create developer time to work on\n> that, but if we do go that way we should at least tell people that we\n> are aware of the problems and that it's on the global TODO list (not\n> necessarily on anyone's personal TODO list though).\n\nFor the record... I at least am aware of the problem and it is indeed on \nmy personal git todo list.  Not that I have a clear solution yet (I've \nbeen pondering on some git packing issues for almost 4 years now).\n\nStill, in this particular case, the problem appears to be unclear to me, \nlike \"this shouldn't be so bad\".\n\n> Of course, the problem can sometimes be solved by splitting the\n> repository, but I think it is important to have an official policy\n> here, do we want Git to support huge repositories, or do we not?\n\nI do.\n\n\nNicolas\n"},{"id":"110508","messageId":"alpine.LFD.2.00.0904052326090.6741@xanadu.home","threadId":"18724","inReplyTo":"alpine.DEB.1.10.0904051613420.6245@asgard.lang.hm","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-06T03:34:32Z","receivedAt":"2009-04-06T03:34:32Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sun, 5 Apr 2009, david@lang.hm wrote:\n\n> On Mon, 6 Apr 2009, Nicolas Sebrecht wrote:\n> \n> > On Sun, Apr 05, 2009 at 02:28:35PM -0700, david@lang.hm wrote:\n> > \n> > > guys, back off a little on telling the gentoo people to change.\n> > \n> > Don't blame Git people, please. I currently am the only one here to\n> > discuss that way and see a painful work coming at Gentoo.\n> > Git people didn't discuss around thoses issues.\n> > \n> > >                                                                 the\n> > > kernel developers don't split th kernel into 'core' 'drivers' etc pieces\n> > > just because some people only work on one area.\n> > \n> > And you might notice that they don't provide a CVS access and actually\n> > don't work around an unique shared repo. Also, you might notice that\n> > keeping the history clean to assure the work on the kernel easier is not\n> > an elementary issue.\n> \n> these issues are completely seperate from the issue that the initial poster\n> asked about, which is that when someone tries to do a clone of the repository\n> the system wastes a lot of time creating a new pack.\n\nAnd this shouldn't be, by design.  Especially if your repo serving clone \nrequests is already well packed.\n\nWhat git-pack-objects does in this case is not a full repack.  It \ninstead _reuse_ as much of the existing packs as possible, and only does \nthe heavy packing processing for loose objects and/or inter pack \nboundaryes when gluing everything together for streaming over the net.  \nIf for example you have a single pack because your repo is already fully \npacked, then the \"packing operation\" involved during a clone should \nmerely copy the existing pack over with no further attempt at delta \ncompression.\n\n> don't worry about this part of things, worry about why the server wastes so\n> many resources.\n\nIndeed.  And since a significant amount of code involved happens to be \nmine, I do wonder.\n\n\nNicolas\n"},{"id":"110509","messageId":"alpine.LFD.2.00.0904052336260.6741@xanadu.home","threadId":"18724","inReplyTo":"20090405T230552Z@curie.orbis-terrarum.net","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-06T04:06:00Z","receivedAt":"2009-04-06T04:06:00Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sun, 5 Apr 2009, Robin H. Johnson wrote:\n\n> On Sun, Apr 05, 2009 at 03:57:14PM -0400, Jeff King wrote:\n> > > During an initial clone, I see that git-upload-pack invokes\n> > > pack-objects, despite the ENTIRE repository already being packed - no\n> > > loose objects whatsoever. git-upload-pack then seems to buffer in\n> > > memory.\n> > We need to run pack-objects even if the repo is fully packed because we\n> > don't know what's _in_ the existing pack (or packs). In particular we\n> > want to:\n> >   - combine multiple packs into a single pack; this is more efficient on\n> >     the network, because you can find more deltas, and I believe is\n> >     required because the protocol sends only a single pack.\n> > \n> >   - cull any objects which are not actually part of the reachability\n> >     chain from the refs we are sending\n> > \n> > If no work needs to be done for either case, then pack-objects should\n> > basically just figure that out and then send the existing pack (the\n> > expensive bit is doing deltas, and we don't consider objects in the same\n> > pack for deltas, as we know we have already considered that during the\n> > last repack). It does mmap the whole pack, so you will see your virtual\n> > memory jump, but nothing should require the whole pack being in memory\n> > at once.\n\nActually the pack is mapped with a (configurable) window.  See the\ncore.packedGitWindowSize and core.packedGitLimit config options for \ndetails.\n\n> While my current pack setup has multiple packs of not more than 100MiB\n> each, that was simply for ease of resume with rsync+http tests. Even\n> when I already had a single pack, with every object reachable,\n> pack-objects was redoing the packing.\n\nIn that case it shouldn't have.\n\n> > pack-objects streams the output to upload-pack, which should only ever\n> > have an 8K buffer of it in memory at any given time.\n> > \n> > At least that is how it is all supposed to work, according to my\n> > understanding. So if you are seeing very high memory usage, I wonder if\n> > there is a bug in pack-objects or upload-pack that can be fixed.\n> > \n> > Maybe somebody more knowledgeable than me about packing can comment.\n> Looking at the source, I agree that it should be buffering, however top and ps\n> seem to disagree. 3GiB VSZ and 2.5GiB RSS here now.\n> \n> %CPU %MEM     VSZ     RSS STAT START   TIME COMMAND\n>  0.0  0.0  140932    1040 Ss   16:09   0:00 \\_ git-upload-pack /code/gentoo/gentoo-git/gentoo-x86.git \n> 32.2  0.0       0       0 Z    16:09   1:50     \\_ [git-upload-pack] <defunct>\n> 80.8 44.2 3018484 2545700 Sl   16:09   4:36     \\_ git pack-objects --stdout --progress --delta-base-offset \n> \n> Also, I did another trace, using some other hardware, in a LAN setting, and\n> noticed that git-upload-pack/pack-objects only seems to start output to the\n> network after it reaches 100% in 'remote: Compressing objects:'.\n\nThat's to be expected.  Delta compression matches objects which are not \nin the stream order at all.  Therefore it is not possible to start \noutputting pack data until this pass is done.  Still, this pass should \nnot be invoked if your repository is already fully packed into one pack.  \nCan you confirm this is actually the case?\n\n> Relatedly, throwing more RAM (6GiB total, vs. the previous 2GiB) at \n> the server in this case cut the 200 wallclock minutes before any \n> sending too place down to 5 minutes.\n\nWell... here's a wild guess.  In the source repository serving clone \nrequests, please do:\n\n\tgit config pack.deltaCacheSize 1\n\tgit config pack.deltaCacheLimit 0\n\nand try cloning again with a fully packed repository.\n\n> > > For the initial clone, can the git-upload-pack algorithm please send\n> > > existing packs, and only generate a pack containing the non-packed\n> > > items?\n> > \n> > I believe that would require a change to the protocol to allow multiple\n> > packs. However, it may be possible to munge the pack header in such a\n> > way that you basically concatenate multiple packs. You would still want\n> > to peek in the big pack to try deltas from the non-packed items, though.\n\nAs explained already, even if the protocol requires a single pack to be \ncreated, it is still made up of unmodified data segments from existing \npacks as much as possible.  So you should see it more or less as the \nconcatenation of those packs already, plus some munging over the edges.\n\n> > I think all of this falls into the realm of the GSOC pack caching project.\n> > There have been other discussions on the list, so you might want to look\n> > through those for something useful.\n> Yes, both changing the protocol, and recognizing that existing packs may be\n> suitable to send could be considered as part of the caching project, as they\n> fall under the aegis of making good use of what's stored in the cache already\n> to send.\n\nThe caching pack project is to address a different issue: mainly to \nbypass the object enumeration cost.  In other words, it could allow for \nskipping the \"Counting objects\" pass, and a tiny bit more.  At least in \ntheory that's about the main difference.  This has many drawbacks as \nwell though.\n\n\nNicolas\n"},{"id":"110510","messageId":"alpine.LFD.2.00.0904060006410.6741@xanadu.home","threadId":"18724","inReplyTo":"fcaeb9bf0904052010p34e3246bwd7e1f5297acf37e2@mail.gmail.com","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-06T04:09:17Z","receivedAt":"2009-04-06T04:09:17Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 6 Apr 2009, Nguyen Thai Ngoc Duy wrote:\n\n> On Mon, Apr 6, 2009 at 9:38 AM, Robin H. Johnson <robbat2@gentoo.org> wrote:\n> > Looking at the source, I agree that it should be buffering, however top and ps\n> > seem to disagree. 3GiB VSZ and 2.5GiB RSS here now.\n> >\n> > %CPU %MEM     VSZ     RSS STAT START   TIME COMMAND\n> >  0.0  0.0  140932    1040 Ss   16:09   0:00 \\_ git-upload-pack /code/gentoo/gentoo-git/gentoo-x86.git\n> > 32.2  0.0       0       0 Z    16:09   1:50     \\_ [git-upload-pack] <defunct>\n> > 80.8 44.2 3018484 2545700 Sl   16:09   4:36     \\_ git pack-objects --stdout --progress --delta-base-offset\n> >\n> > Also, I did another trace, using some other hardware, in a LAN setting, and\n> > noticed that git-upload-pack/pack-objects only seems to start output to the\n> > network after it reaches 100% in 'remote: Compressing objects:'.\n> >\n> > Relatedly, throwing more RAM (6GiB total, vs. the previous 2GiB) at the server\n> > in this case cut the 200 wallclock minutes before any sending too place down to\n> > 5 minutes.\n> \n> Searching back the archive, there was memory fragmentation issue with\n> gcc repo. I wonder if it happens again. Maybe you should try Google\n> allocator. BTW, did you try to turn off THREADED_DELTA_SEARCH?\n\nThat was for a _full_ repack, i.e. 'git repack -a -f'.  This is never \nthe case on a fetch/clone, like in this case, unless you have all your \nobjects in loose form.\n\n\nNicolas\n"},{"id":"110512","messageId":"7vab6ue520.fsf@gitster.siamese.dyndns.org","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904052326090.6741@xanadu.home","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-04-06T05:15:19Z","receivedAt":"2009-04-06T05:15:19Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nicolas Pitre <nico@cam.org> writes:\n\n> What git-pack-objects does in this case is not a full repack.  It \n> instead _reuse_ as much of the existing packs as possible, and only does \n> the heavy packing processing for loose objects and/or inter pack \n> boundaryes when gluing everything together for streaming over the net.  \n> If for example you have a single pack because your repo is already fully \n> packed, then the \"packing operation\" involved during a clone should \n> merely copy the existing pack over with no further attempt at delta \n> compression.\n\nOne possibile scenario that you still need to spend memory and cycle is if\nthe cloned repository was packed to an excessive depth to cause many of\nits objects to be in deltified form on insanely deep chains, while cloning\nsend-pack uses a depth that is more reasonable.  Then pack-objects invoked\nby send-pack is not allowed to reuse most of the objects and would end up\nredoing the delta on them.\n"},{"id":"110555","messageId":"vpq3acm6n7p.fsf@bauges.imag.fr","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904052326090.6741@xanadu.home","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2009-04-06T11:22:34Z","receivedAt":"2009-04-06T11:22:34Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"Nicolas Pitre <nico@cam.org> writes:\n\n> If for example you have a single pack because your repo is already fully \n> packed, then the \"packing operation\" involved during a clone should \n> merely copy the existing pack over with no further attempt at delta \n> compression.\n\nThere's still the question if your repository has too many objects\n(for example, a branch that you deleted without garbage-collecting\nit). Then, sending the whole pack sends data that one may have\nconsidered as \"secret\".\n\nTo me, this is a non-issue (if the content of these objects are\nsecret, then why are they here at all on a public server?), but I\nthink there were discussions here about it (can't find the right\nkeywords to dig the archives though), and other people may think\ndifferently.\n\nJeff King's answer in <20090405195714.GA4716@coredump.intra.peff.net>\ntackles this problem too.\n\n-- \nMatthieu\n"},{"id":"110568","messageId":"alpine.LFD.2.00.0904060901250.6741@xanadu.home","threadId":"18724","inReplyTo":"7vab6ue520.fsf@gitster.siamese.dyndns.org","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-06T13:12:22Z","receivedAt":"2009-04-06T13:12:22Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sun, 5 Apr 2009, Junio C Hamano wrote:\n\n> Nicolas Pitre <nico@cam.org> writes:\n> \n> > What git-pack-objects does in this case is not a full repack.  It \n> > instead _reuse_ as much of the existing packs as possible, and only does \n> > the heavy packing processing for loose objects and/or inter pack \n> > boundaryes when gluing everything together for streaming over the net.  \n> > If for example you have a single pack because your repo is already fully \n> > packed, then the \"packing operation\" involved during a clone should \n> > merely copy the existing pack over with no further attempt at delta \n> > compression.\n> \n> One possibile scenario that you still need to spend memory and cycle is if\n> the cloned repository was packed to an excessive depth to cause many of\n> its objects to be in deltified form on insanely deep chains, while cloning\n> send-pack uses a depth that is more reasonable.  Then pack-objects invoked\n> by send-pack is not allowed to reuse most of the objects and would end up\n> redoing the delta on them.\n\nNope.  When pack data is reused, there is simply no consideration what \nso ever for the actual delta depth limit.  Only when an object already \nbeing used as a delta base for reused deltas is itself subject to delta \ncompression does the real depth of the concerned delta chain is \nevaluated in order to not purposely bust the specified delta depth limit \n(otherwise a delta chain could grow unbounded).\n\n\nNicolas\n"},{"id":"110572","messageId":"alpine.LFD.2.00.0904060912530.6741@xanadu.home","threadId":"18724","inReplyTo":"vpq3acm6n7p.fsf@bauges.imag.fr","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-06T13:29:04Z","receivedAt":"2009-04-06T13:29:04Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 6 Apr 2009, Matthieu Moy wrote:\n\n> Nicolas Pitre <nico@cam.org> writes:\n> \n> > If for example you have a single pack because your repo is already fully \n> > packed, then the \"packing operation\" involved during a clone should \n> > merely copy the existing pack over with no further attempt at delta \n> > compression.\n> \n> There's still the question if your repository has too many objects\n> (for example, a branch that you deleted without garbage-collecting\n> it). Then, sending the whole pack sends data that one may have\n> considered as \"secret\".\n\nI said \"merely copy\", which is not a straight copy.  In this case, only \nthe relevant objects from the existing pack will be copied into the \nstreamed pack, and objects from the unused branch will be left behind.  \nIn that case, deltas which base object is left behind will automatically \nbe considered for alternative delta matching of course, but that is \nnormally a relatively small set of objects.  And if that set gets really \nbig, that means that an even bigger set of objects was left behind, \nmaking the actual repacking smaller in scope.\n\n> To me, this is a non-issue (if the content of these objects are\n> secret, then why are they here at all on a public server?), but I\n> think there were discussions here about it (can't find the right\n> keywords to dig the archives though), and other people may think\n> differently.\n\nGuess who was involved in that discussion...\n\nI may allow you to pull certain branches directly from my own PC through \nthe git native protocol.  That doesn't mean you have direct access to \nthe whole of any of the packs I have on my disk.\n\n\nNicolas\n"},{"id":"110580","messageId":"9e4733910904060652t6c0f37d9t246b7394e3aad350@mail.gmail.com","threadId":"18724","inReplyTo":"7vab6ue520.fsf@gitster.siamese.dyndns.org","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2009-04-06T13:52:56Z","receivedAt":"2009-04-06T13:52:56Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On Mon, Apr 6, 2009 at 1:15 AM, Junio C Hamano <gitster@pobox.com> wrote:\n> Nicolas Pitre <nico@cam.org> writes:\n>\n>> What git-pack-objects does in this case is not a full repack.  It\n>> instead _reuse_ as much of the existing packs as possible, and only does\n>> the heavy packing processing for loose objects and/or inter pack\n>> boundaryes when gluing everything together for streaming over the net.\n>> If for example you have a single pack because your repo is already fully\n>> packed, then the \"packing operation\" involved during a clone should\n>> merely copy the existing pack over with no further attempt at delta\n>> compression.\n>\n> One possibile scenario that you still need to spend memory and cycle is if\n> the cloned repository was packed to an excessive depth to cause many of\n> its objects to be in deltified form on insanely deep chains, while cloning\n> send-pack uses a depth that is more reasonable.  Then pack-objects invoked\n> by send-pack is not allowed to reuse most of the objects and would end up\n> redoing the delta on them.\n\nThat seems broken. You went through all of the trouble to make the\npack file smaller to reduce transmission time, and then clone undoes\nthe work.\n\nWhat about making a very simple special case for an initial clone?\nFirst thing an initial clone does is copy all of the pack files from\nthe server to the client without even looking at them. Some of these\npacks will probably be marked 'keep' because they are old history and\nhave been densely packed. Once the packs are down, start over and do a\nfetch taking these packs into account.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"110582","messageId":"20090406T140124Z@curie.orbis-terrarum.net","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904060912530.6741@xanadu.home","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Robin H. Johnson","fromEmail":"robbat2@gentoo.org","sentAt":"2009-04-06T14:03:20Z","receivedAt":"2009-04-06T14:03:20Z","isPatch":false,"sender":{"key":"robbat2@gentoo.org","avatar":"https://avatars.githubusercontent.com/u/373898?v=4"},"body":"I haven't read all this morning submissions to the thread yet, but I\nwanted to make two posts before I leave on a trip (in ~20 minutes), and\nI'll be back late on Thursday.\n\nOn Mon, Apr 06, 2009 at 09:29:04AM -0400, Nicolas Pitre wrote:\n> > To me, this is a non-issue (if the content of these objects are\n> > secret, then why are they here at all on a public server?), but I\n> > think there were discussions here about it (can't find the right\n> > keywords to dig the archives though), and other people may think\n> > differently.\n> Guess who was involved in that discussion...\n> I may allow you to pull certain branches directly from my own PC through \n> the git native protocol.  That doesn't mean you have direct access to \n> the whole of any of the packs I have on my disk.\nIf the native rsync protocol is allowed to the repo, then that argument\nis moot.\n\n-- \nRobin Hugh Johnson\nGentoo Linux Developer & Infra Guy\nE-Mail     : robbat2@gentoo.org\nGnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n"},{"id":"110583","messageId":"alpine.LFD.2.00.0904061011460.6741@xanadu.home","threadId":"18724","inReplyTo":"20090406T140124Z@curie.orbis-terrarum.net","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-06T14:14:45Z","receivedAt":"2009-04-06T14:14:45Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 6 Apr 2009, Robin H. Johnson wrote:\n\n> I haven't read all this morning submissions to the thread yet, but I\n> wanted to make two posts before I leave on a trip (in ~20 minutes), and\n> I'll be back late on Thursday.\n> \n> On Mon, Apr 06, 2009 at 09:29:04AM -0400, Nicolas Pitre wrote:\n> > > To me, this is a non-issue (if the content of these objects are\n> > > secret, then why are they here at all on a public server?), but I\n> > > think there were discussions here about it (can't find the right\n> > > keywords to dig the archives though), and other people may think\n> > > differently.\n> > Guess who was involved in that discussion...\n> > I may allow you to pull certain branches directly from my own PC through \n> > the git native protocol.  That doesn't mean you have direct access to \n> > the whole of any of the packs I have on my disk.\n> If the native rsync protocol is allowed to the repo, then that argument\n> is moot.\n\nThe rsync protocol is _not_ the native git protocol.  And I personally \ndon't encourage its usage either, except as a _temporary_ workaround for \nunresolved issues.  You will never see this protocol available from any \ngit server I maintain.\n\n\nNicolas\n"},{"id":"110584","messageId":"alpine.LFD.2.00.0904060959250.6741@xanadu.home","threadId":"18724","inReplyTo":"9e4733910904060652t6c0f37d9t246b7394e3aad350@mail.gmail.com","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-06T14:19:20Z","receivedAt":"2009-04-06T14:19:20Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 6 Apr 2009, Jon Smirl wrote:\n\n> On Mon, Apr 6, 2009 at 1:15 AM, Junio C Hamano <gitster@pobox.com> wrote:\n> > Nicolas Pitre <nico@cam.org> writes:\n> >\n> >> What git-pack-objects does in this case is not a full repack.  It\n> >> instead _reuse_ as much of the existing packs as possible, and only does\n> >> the heavy packing processing for loose objects and/or inter pack\n> >> boundaryes when gluing everything together for streaming over the net.\n> >> If for example you have a single pack because your repo is already fully\n> >> packed, then the \"packing operation\" involved during a clone should\n> >> merely copy the existing pack over with no further attempt at delta\n> >> compression.\n> >\n> > One possibile scenario that you still need to spend memory and cycle is if\n> > the cloned repository was packed to an excessive depth to cause many of\n> > its objects to be in deltified form on insanely deep chains, while cloning\n> > send-pack uses a depth that is more reasonable.  Then pack-objects invoked\n> > by send-pack is not allowed to reuse most of the objects and would end up\n> > redoing the delta on them.\n> \n> That seems broken. You went through all of the trouble to make the\n> pack file smaller to reduce transmission time, and then clone undoes\n> the work.\n\nAnd as I already explained, this is indeed not what happens.\n\n> What about making a very simple special case for an initial clone?\n\nThere should not be any need for initial clone hacks.\n\n> First thing an initial clone does is copy all of the pack files from\n> the server to the client without even looking at them.\n\nThis is a no go for reasons already stated many times.  There are \nsecurity implications (those packs might contain stuff that you didn't \nintend to be publically accessible) and there might be efficiency \nreasons as well (you might have a shared object store with lots of stuff \nunrelated to the particular clone).\n\nThe biggest cost right now when cloning a big packed repo is object \nenumeration.  Any other issues related to memory costs in the GB range \nsimply has no reason for it, and is mostly due to misconfigurations or \nbugs that have to be fixed.  Trying to work around the issue by all \nsorts of hacks is simply counter productive.\n\nIn the case that started this very thread, I suspect that a small \nmisfeature of some delta caching might be the culprit.  I asked Robin H. \nJohnson to perform a really simple config addition to his repo and \nretest, for which we still haven't seen any results yet.\n\n\nNicolas\n"},{"id":"110585","messageId":"20090406T140441Z@curie.orbis-terrarum.net","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904052336260.6741@xanadu.home","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Robin H. Johnson","fromEmail":"robbat2@gentoo.org","sentAt":"2009-04-06T14:20:55Z","receivedAt":"2009-04-06T14:20:55Z","isPatch":false,"sender":{"key":"robbat2@gentoo.org","avatar":"https://avatars.githubusercontent.com/u/373898?v=4"},"body":"Again, I'm about to leave on a trip for a few days (back late Thursday),\nbut just wanted to comment in on the thread.\n\nOn Mon, Apr 06, 2009 at 12:06:00AM -0400, Nicolas Pitre wrote:\n> > While my current pack setup has multiple packs of not more than 100MiB\n> > each, that was simply for ease of resume with rsync+http tests. Even\n> > when I already had a single pack, with every object reachable,\n> > pack-objects was redoing the packing.\n> In that case it shouldn't have.\nI'll retest that part on my return, but I'm pretty sure I did see the\nsame excess cputime usage.\n\n> > Also, I did another trace, using some other hardware, in a LAN setting, and\n> > noticed that git-upload-pack/pack-objects only seems to start output to the\n> > network after it reaches 100% in 'remote: Compressing objects:'.\n> That's to be expected.  Delta compression matches objects which are not \n> in the stream order at all.  Therefore it is not possible to start \n> outputting pack data until this pass is done.  Still, this pass should \n> not be invoked if your repository is already fully packed into one pack.  \nSo it's seeking around the existing packs before sending?\n\n> Can you confirm this is actually the case?\nThe most recent tests were with the 15(+ one partial) packs limited to a\nmax of 100MiB each, because that made resume for rsync/http during the\ntests much cleaner.\n\n> > Relatedly, throwing more RAM (6GiB total, vs. the previous 2GiB) at \n> > the server in this case cut the 200 wallclock minutes before any \n> > sending too place down to 5 minutes.\n> Well... here's a wild guess.  In the source repository serving clone \n> requests, please do:\n> \tgit config pack.deltaCacheSize 1\n> \tgit config pack.deltaCacheLimit 0\n> and try cloning again with a fully packed repository.\nI did the multiple pack case quickly, and found that it does still take\na long time in the low memory case. I'll do the test with a single pack\non my return.\n\n> The caching pack project is to address a different issue: mainly to \n> bypass the object enumeration cost.  In other words, it could allow for \n> skipping the \"Counting objects\" pass, and a tiny bit more.  At least in \n> theory that's about the main difference.  This has many drawbacks as \n> well though.\nRelatedly, would it be possible to keep a cache of enumerated objects\nthat was trivially updatable during pushes?\n\n-- \nRobin Hugh Johnson\nGentoo Linux Developer & Infra Guy\nE-Mail     : robbat2@gentoo.org\nGnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n"},{"id":"110586","messageId":"9e4733910904060737k3d1c082fk785cd98cdeb6d73d@mail.gmail.com","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904060959250.6741@xanadu.home","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2009-04-06T14:37:44Z","receivedAt":"2009-04-06T14:37:44Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On Mon, Apr 6, 2009 at 10:19 AM, Nicolas Pitre <nico@cam.org> wrote:\n> On Mon, 6 Apr 2009, Jon Smirl wrote:\n>\n>> On Mon, Apr 6, 2009 at 1:15 AM, Junio C Hamano <gitster@pobox.com> wrote:\n>> > Nicolas Pitre <nico@cam.org> writes:\n>> >\n>> >> What git-pack-objects does in this case is not a full repack.  It\n>> >> instead _reuse_ as much of the existing packs as possible, and only does\n>> >> the heavy packing processing for loose objects and/or inter pack\n>> >> boundaryes when gluing everything together for streaming over the net.\n>> >> If for example you have a single pack because your repo is already fully\n>> >> packed, then the \"packing operation\" involved during a clone should\n>> >> merely copy the existing pack over with no further attempt at delta\n>> >> compression.\n>> >\n>> > One possibile scenario that you still need to spend memory and cycle is if\n>> > the cloned repository was packed to an excessive depth to cause many of\n>> > its objects to be in deltified form on insanely deep chains, while cloning\n>> > send-pack uses a depth that is more reasonable.  Then pack-objects invoked\n>> > by send-pack is not allowed to reuse most of the objects and would end up\n>> > redoing the delta on them.\n>>\n>> That seems broken. You went through all of the trouble to make the\n>> pack file smaller to reduce transmission time, and then clone undoes\n>> the work.\n>\n> And as I already explained, this is indeed not what happens.\n>\n>> What about making a very simple special case for an initial clone?\n>\n> There should not be any need for initial clone hacks.\n>\n>> First thing an initial clone does is copy all of the pack files from\n>> the server to the client without even looking at them.\n>\n> This is a no go for reasons already stated many times.  There are\n> security implications (those packs might contain stuff that you didn't\n> intend to be publically accessible) and there might be efficiency\n> reasons as well (you might have a shared object store with lots of stuff\n> unrelated to the particular clone).\n\nHow do you deal with dense history packs? These packs take many hours\nto make (on a server class machine) and can be half the size of a\nregular pack. Shouldn't there be a way to copy these packs intact on\nan initial clone? It's ok if these packs are specially marked as being\nok to copy.\n\n>\n> The biggest cost right now when cloning a big packed repo is object\n> enumeration.  Any other issues related to memory costs in the GB range\n> simply has no reason for it, and is mostly due to misconfigurations or\n> bugs that have to be fixed.  Trying to work around the issue by all\n> sorts of hacks is simply counter productive.\n>\n> In the case that started this very thread, I suspect that a small\n> misfeature of some delta caching might be the culprit.  I asked Robin H.\n> Johnson to perform a really simple config addition to his repo and\n> retest, for which we still haven't seen any results yet.\n>\n>\n> Nicolas\n>\n\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"110590","messageId":"20090406144829.GF23604@spearce.org","threadId":"18724","inReplyTo":"9e4733910904060737k3d1c082fk785cd98cdeb6d73d@mail.gmail.com","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2009-04-06T14:48:29Z","receivedAt":"2009-04-06T14:48:29Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Jon Smirl <jonsmirl@gmail.com> wrote:\n> \n> How do you deal with dense history packs? These packs take many hours\n> to make (on a server class machine) and can be half the size of a\n> regular pack. Shouldn't there be a way to copy these packs intact on\n> an initial clone? It's ok if these packs are specially marked as being\n> ok to copy.\n\nThese should be copied as-is.\n\nBasically, object enumeration lists every reachable object, which\nshould include every object in this pack if its a \"dense history\npack\".  We then start to write out each object.  As each object\nis written we look to see if it already exists in a pack.  It does\n(in your dense history pack), so we then look to see if its delta\nbase is also in the output list (it is), so we send the data as-is.\n\n\nOne of the bigger costs with such clones is building that huge list\nof objects needed to send.  The primary cost appears to be unpacking\nthe trees from the \"dense history pack\", where delta chains are\nusually quite long.  The GSoC 2009 pack caching project idea is\nbased on the theory that we should be able to save a list of objects\nthat are reachable from some fixed point (e.g. a very well known,\nstable tag), and avoid needing to read these ancient trees.\n\nBut its just a theory.  Caching always costs you management\noverheads.  And it may not save us that much time .  And most of\nthe theory here is based on JGit's performance during packing,\n*not* git-core.\n\nI came up with the object list caching idea because JGit's object\nenumeration is just pitiful.  (Its Java, what do you want, if you\nwanted fast, you'd use portable assembler... like git-core does.)\nWhether or not its worth applying to git-core is another story\nentirely.\n\n-- \nShawn.\n"},{"id":"110594","messageId":"alpine.LFD.2.00.0904061042300.6741@xanadu.home","threadId":"18724","inReplyTo":"9e4733910904060737k3d1c082fk785cd98cdeb6d73d@mail.gmail.com","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-06T15:14:58Z","receivedAt":"2009-04-06T15:14:58Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 6 Apr 2009, Jon Smirl wrote:\n\n> On Mon, Apr 6, 2009 at 10:19 AM, Nicolas Pitre <nico@cam.org> wrote:\n> > On Mon, 6 Apr 2009, Jon Smirl wrote:\n> >\n> >> First thing an initial clone does is copy all of the pack files from\n> >> the server to the client without even looking at them.\n> >\n> > This is a no go for reasons already stated many times.  There are\n> > security implications (those packs might contain stuff that you didn't\n> > intend to be publically accessible) and there might be efficiency\n> > reasons as well (you might have a shared object store with lots of stuff\n> > unrelated to the particular clone).\n> \n> How do you deal with dense history packs? These packs take many hours\n> to make (on a server class machine) and can be half the size of a\n> regular pack. Shouldn't there be a way to copy these packs intact on\n> an initial clone? It's ok if these packs are specially marked as being\n> ok to copy.\n\n[sigh]\n\nLet me explain it all again.\n\nThere is basically two ways to create a new pack: the intelligent way, \nand the bruteforce way.\n\nWhen creating a new pack the intelligent way, what we do is to enumerate \nall the needed object and look them up in the object store.  When a \nparticular object is found, we create a record for that object and note \nin which pack it is located, at what offset in that pack, how much space \nit occupies in its compressed form within that pack, , and if whether it \nis a delta or not.  When that object is indeed a delta (the majority of \nobjects usually are) then we also keep a pointer on the record for the \nbase object for that delta.\n\nNext, for all objects in delta form which base object is also part of \nthe object enumeration and obviously part of the same pack, we simply \nflag those objects as directly reusable without any further processing.  \nThis means that, when those objects are about to be stored in the new \npack, their raw data is simply copied straight from the original pack \nusing the offset and size noted above.  In other words, those objects \nare simply never redeltified nor redeflated at all, and all the work \nthat was previously done to find the best delta match is preserved with \nno extra cost.\n\nOf course, when your repository is tightly packed into a single pack, \nthen all enumerated objects fall into the reusable category and \ntherefore a copy of the original pack is indeed sent over the wire.  \nOne exception is with older git clients which don't support the delta \nbase offset encoding, in which case the delta reference encoding is \nsubstituted on the fly with almost no cost (this is btw another reason \nwhy a dumb copy of existing pack may not work universally either).  But \nin the common case, you might see the above as just the same as if git \ndid copy the pack file because it really only reads some data from a \npack and immediately writes that data out.\n\nThe bruteforce repacking is different because it simply doesn't concern \nitself with existing deltas at all.  It instead start everything from \nscratch and perform the whole delta search all over for all objects.  \nThis is what takes lots of resources and CPU cycles, and as you may \nguess, is never used for fetch/clone requests.\n\n\nNicolas\n"},{"id":"110595","messageId":"9e4733910904060828m414dfe7v66b19f7b4c5b670e@mail.gmail.com","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904061042300.6741@xanadu.home","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2009-04-06T15:28:39Z","receivedAt":"2009-04-06T15:28:39Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On Mon, Apr 6, 2009 at 11:14 AM, Nicolas Pitre <nico@cam.org> wrote:\n> On Mon, 6 Apr 2009, Jon Smirl wrote:\n>\n>> On Mon, Apr 6, 2009 at 10:19 AM, Nicolas Pitre <nico@cam.org> wrote:\n>> > On Mon, 6 Apr 2009, Jon Smirl wrote:\n>> >\n>> >> First thing an initial clone does is copy all of the pack files from\n>> >> the server to the client without even looking at them.\n>> >\n>> > This is a no go for reasons already stated many times.  There are\n>> > security implications (those packs might contain stuff that you didn't\n>> > intend to be publically accessible) and there might be efficiency\n>> > reasons as well (you might have a shared object store with lots of stuff\n>> > unrelated to the particular clone).\n>>\n>> How do you deal with dense history packs? These packs take many hours\n>> to make (on a server class machine) and can be half the size of a\n>> regular pack. Shouldn't there be a way to copy these packs intact on\n>> an initial clone? It's ok if these packs are specially marked as being\n>> ok to copy.\n>\n> [sigh]\n>\n> Let me explain it all again.\n>\n> There is basically two ways to create a new pack: the intelligent way,\n> and the bruteforce way.\n>\n> When creating a new pack the intelligent way, what we do is to enumerate\n> all the needed object and look them up in the object store.  When a\n> particular object is found, we create a record for that object and note\n> in which pack it is located, at what offset in that pack, how much space\n> it occupies in its compressed form within that pack, , and if whether it\n> is a delta or not.  When that object is indeed a delta (the majority of\n> objects usually are) then we also keep a pointer on the record for the\n> base object for that delta.\n>\n> Next, for all objects in delta form which base object is also part of\n> the object enumeration and obviously part of the same pack, we simply\n> flag those objects as directly reusable without any further processing.\n> This means that, when those objects are about to be stored in the new\n> pack, their raw data is simply copied straight from the original pack\n> using the offset and size noted above.  In other words, those objects\n> are simply never redeltified nor redeflated at all, and all the work\n> that was previously done to find the best delta match is preserved with\n> no extra cost.\n\nDoes this process cause random reads all over a 2GB pack file? Busy\nservers can't keep a 2GB pack in memory.\nsendfile() the 2GB pack to client is way more efficient. (assuming the\npack is marked as being ok to send).\n\n>\n> Of course, when your repository is tightly packed into a single pack,\n> then all enumerated objects fall into the reusable category and\n> therefore a copy of the original pack is indeed sent over the wire.\n> One exception is with older git clients which don't support the delta\n> base offset encoding, in which case the delta reference encoding is\n> substituted on the fly with almost no cost (this is btw another reason\n> why a dumb copy of existing pack may not work universally either).  But\n> in the common case, you might see the above as just the same as if git\n> did copy the pack file because it really only reads some data from a\n> pack and immediately writes that data out.\n>\n> The bruteforce repacking is different because it simply doesn't concern\n> itself with existing deltas at all.  It instead start everything from\n> scratch and perform the whole delta search all over for all objects.\n> This is what takes lots of resources and CPU cycles, and as you may\n> guess, is never used for fetch/clone requests.\n>\n>\n> Nicolas\n>\n\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"110597","messageId":"alpine.LFD.2.00.0904061207110.6741@xanadu.home","threadId":"18724","inReplyTo":"9e4733910904060828m414dfe7v66b19f7b4c5b670e@mail.gmail.com","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-06T16:14:11Z","receivedAt":"2009-04-06T16:14:11Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 6 Apr 2009, Jon Smirl wrote:\n\n> On Mon, Apr 6, 2009 at 11:14 AM, Nicolas Pitre <nico@cam.org> wrote:\n> > This means that, when those objects are about to be stored in the new\n> > pack, their raw data is simply copied straight from the original pack\n> > using the offset and size noted above.  In other words, those objects\n> > are simply never redeltified nor redeflated at all, and all the work\n> > that was previously done to find the best delta match is preserved with\n> > no extra cost.\n> \n> Does this process cause random reads all over a 2GB pack file? Busy\n> servers can't keep a 2GB pack in memory.\n\nThe creation of a new pack follows the same object recency rule as the \nones it copies from, so the various reads should be perfectly \nsequential.\n\n> sendfile() the 2GB pack to client is way more efficient. (assuming the\n> pack is marked as being ok to send).\n\nGit is not a FTP server.  Otherwise we would have stayed with the rsync \nprotocol.\n\n\nNicolas\n"},{"id":"110678","messageId":"20090407081019.GK20356@atjola.homenet","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904052315210.6741@xanadu.home","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-07T08:10:19Z","receivedAt":"2009-04-07T08:10:19Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.05 23:24:27 -0400, Nicolas Pitre wrote:\n> On Sun, 5 Apr 2009, Sverre Rabbelier wrote:\n> \n> > Heya,\n> > \n> > On Sun, Apr 5, 2009 at 23:28,  <david@lang.hm> wrote:\n> > > Guys, back off a little on telling the gentoo people to change.\n> > \n> > I agree here, we should either say \"look, we don't really support big\n> > repositories because [explanation here], unless you [workarounds\n> > here]\" OR we should work to improve the support we do have. Of course,\n> > the latter option does not magically create developer time to work on\n> > that, but if we do go that way we should at least tell people that we\n> > are aware of the problems and that it's on the global TODO list (not\n> > necessarily on anyone's personal TODO list though).\n> \n> For the record... I at least am aware of the problem and it is indeed on \n> my personal git todo list.  Not that I have a clear solution yet (I've \n> been pondering on some git packing issues for almost 4 years now).\n> \n> Still, in this particular case, the problem appears to be unclear to me, \n> like \"this shouldn't be so bad\".\n\nIt's not primarily pack-objects, I think. It's the rev-list that's run\nby upload-pack.  Running \"git rev-list --objects --all\" on that repo\neats about 2G RSS, easily killing the system's cache on a small box,\nleading to swapping and a painful time reading the packfile contents\nafterwards to send them to the client.\n\nBjörn\n"},{"id":"110690","messageId":"m3tz5023rq.fsf@localhost.localdomain","threadId":"18724","inReplyTo":"20090407081019.GK20356@atjola.homenet","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-04-07T09:45:41Z","receivedAt":"2009-04-07T09:45:41Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Björn Steinbrink <B.Steinbrink@gmx.de> writes:\n> On 2009.04.05 23:24:27 -0400, Nicolas Pitre wrote:\n> > On Sun, 5 Apr 2009, Sverre Rabbelier wrote:\n> > > \n> > > I agree here, we should either say \"look, we don't really support big\n> > > repositories because [explanation here], unless you [workarounds\n> > > here]\" OR we should work to improve the support we do have. Of course,\n> > > the latter option does not magically create developer time to work on\n> > > that, but if we do go that way we should at least tell people that we\n> > > are aware of the problems and that it's on the global TODO list (not\n> > > necessarily on anyone's personal TODO list though).\n> > \n> > For the record... I at least am aware of the problem and it is indeed on \n> > my personal git todo list.  Not that I have a clear solution yet (I've \n> > been pondering on some git packing issues for almost 4 years now).\n> > \n> > Still, in this particular case, the problem appears to be unclear to me, \n> > like \"this shouldn't be so bad\".\n> \n> It's not primarily pack-objects, I think. It's the rev-list that's run\n> by upload-pack.  Running \"git rev-list --objects --all\" on that repo\n> eats about 2G RSS, easily killing the system's cache on a small box,\n> leading to swapping and a painful time reading the packfile contents\n> afterwards to send them to the client.\n\nThan I think that \"packfile caching\" GSoC project (which is IIRC\n\"object enumeration caching\", or at least includes it) should help\nhere.  You would, from what I understand, run \"git rev-list -objects\n--all --not <tops of cache>\" + sequential read of object enumeration\ncache...\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"110692","messageId":"46a038f90904070311n6606e678y43f716cbc05a397f@mail.gmail.com","threadId":"18724","inReplyTo":"20090405225954.GA18730@vidovic","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2009-04-07T10:11:15Z","receivedAt":"2009-04-07T10:11:15Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Mon, Apr 6, 2009 at 12:59 AM, Nicolas Sebrecht\n<nicolas.s-dev@laposte.net> wrote:\n> What about the rsync solution given in this thread?\n\nAlso, HTTP is excellent for initial clones, possibly better than rsync\nin some cases.\n\nThe Gentoo team has good reasons to do things their way, and it's IMHO\na wart in git that initial clones of large repos. But we do have valid\nworkarounds (as above) so they can use them .\n\ncheers,\n\n\n\nmartin\n-- \n martin.langhoff@gmail.com\n martin@laptop.org -- School Server Architect\n - ask interesting questions\n - don't get distracted with shiny stuff  - working code first\n - http://wiki.laptop.org/go/User:Martinlanghoff\n"},{"id":"110703","messageId":"alpine.LFD.2.00.0904070903020.6741@xanadu.home","threadId":"18724","inReplyTo":"m3tz5023rq.fsf@localhost.localdomain","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-07T13:13:45Z","receivedAt":"2009-04-07T13:13:45Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 7 Apr 2009, Jakub Narebski wrote:\n\n> Björn Steinbrink <B.Steinbrink@gmx.de> writes:\n> > On 2009.04.05 23:24:27 -0400, Nicolas Pitre wrote:\n> > > On Sun, 5 Apr 2009, Sverre Rabbelier wrote:\n> > > > \n> > > > I agree here, we should either say \"look, we don't really support big\n> > > > repositories because [explanation here], unless you [workarounds\n> > > > here]\" OR we should work to improve the support we do have. Of course,\n> > > > the latter option does not magically create developer time to work on\n> > > > that, but if we do go that way we should at least tell people that we\n> > > > are aware of the problems and that it's on the global TODO list (not\n> > > > necessarily on anyone's personal TODO list though).\n> > > \n> > > For the record... I at least am aware of the problem and it is indeed on \n> > > my personal git todo list.  Not that I have a clear solution yet (I've \n> > > been pondering on some git packing issues for almost 4 years now).\n> > > \n> > > Still, in this particular case, the problem appears to be unclear to me, \n> > > like \"this shouldn't be so bad\".\n> > \n> > It's not primarily pack-objects, I think. It's the rev-list that's run\n> > by upload-pack.  Running \"git rev-list --objects --all\" on that repo\n> > eats about 2G RSS, easily killing the system's cache on a small box,\n> > leading to swapping and a painful time reading the packfile contents\n> > afterwards to send them to the client.\n> \n> Than I think that \"packfile caching\" GSoC project (which is IIRC\n> \"object enumeration caching\", or at least includes it) should help\n> here.\n\nNO!\n\nPlease people stop being so creative with all sort of ways to simply \navoid the real issue and focussing on a real fix.  Git has not become \nwhat it is today by the accumulation of workarounds and ignorance of \nfundamental issues.\n\nHaving git-rev-list consume about 2G RSS for the enumeration of 4M \nobjects is simply inacceptable, period.  This is the equivalent of 500 \nbytes per object pinned in memory on average, just for listing object, \nwhich is completely silly. We ought to do better than that.\n\n\nNicolas\n"},{"id":"110704","messageId":"200904071537.04225.jnareb@gmail.com","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904070903020.6741@xanadu.home","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-04-07T13:37:03Z","receivedAt":"2009-04-07T13:37:03Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Tue, 7 Apr 2009, Nicolas Pitre wrote:\n> On Tue, 7 Apr 2009, Jakub Narebski wrote:\n>> Björn Steinbrink <B.Steinbrink@gmx.de> writes:\n\n[...]\n>>> It's not primarily pack-objects, I think. It's the rev-list that's run\n>>> by upload-pack.  Running \"git rev-list --objects --all\" on that repo\n>>> eats about 2G RSS, easily killing the system's cache on a small box,\n>>> leading to swapping and a painful time reading the packfile contents\n>>> afterwards to send them to the client.\n>> \n>> Than I think that \"packfile caching\" GSoC project (which is IIRC\n>> \"object enumeration caching\", or at least includes it) should help\n>> here.\n> \n> NO!\n> \n> Please people stop being so creative with all sort of ways to simply \n> avoid the real issue and focussing on a real fix.  Git has not become \n> what it is today by the accumulation of workarounds and ignorance of \n> fundamental issues.\n> \n> Having git-rev-list consume about 2G RSS for the enumeration of 4M \n> objects is simply inacceptable, period.  This is the equivalent of 500 \n> bytes per object pinned in memory on average, just for listing object, \n> which is completely silly. We ought to do better than that.\n\nI have thought that the large amount of memory consumed by git-rev-list\nwas caused by not-so-sequential access to very large packfile (1.5GB+ if\nI remember correctly), which I thought causes the whole packfile to be\nmmapped and not only window, plus large amount of objects in 300MB+ mem\nrange or something; those both would account for around 2GB.\n\nBesides even if git-rev-list wouldn't take so much memory, object\nenumeration caching would still help with CPU load... admittedly less.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"110706","messageId":"9e4733910904070703w22887bd6l7358ac8ec8b95c97@mail.gmail.com","threadId":"18724","inReplyTo":"200904071537.04225.jnareb@gmail.com","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2009-04-07T14:03:48Z","receivedAt":"2009-04-07T14:03:48Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"2009/4/7 Jakub Narebski <jnareb@gmail.com>:\n> On Tue, 7 Apr 2009, Nicolas Pitre wrote:\n>> On Tue, 7 Apr 2009, Jakub Narebski wrote:\n>>> Björn Steinbrink <B.Steinbrink@gmx.de> writes:\n>\n> [...]\n>>>> It's not primarily pack-objects, I think. It's the rev-list that's run\n>>>> by upload-pack.  Running \"git rev-list --objects --all\" on that repo\n>>>> eats about 2G RSS, easily killing the system's cache on a small box,\n>>>> leading to swapping and a painful time reading the packfile contents\n>>>> afterwards to send them to the client.\n>>>\n>>> Than I think that \"packfile caching\" GSoC project (which is IIRC\n>>> \"object enumeration caching\", or at least includes it) should help\n>>> here.\n>>\n>> NO!\n>>\n>> Please people stop being so creative with all sort of ways to simply\n>> avoid the real issue and focussing on a real fix.  Git has not become\n>> what it is today by the accumulation of workarounds and ignorance of\n>> fundamental issues.\n>>\n>> Having git-rev-list consume about 2G RSS for the enumeration of 4M\n>> objects is simply inacceptable, period.  This is the equivalent of 500\n>> bytes per object pinned in memory on average, just for listing object,\n>> which is completely silly. We ought to do better than that.\n>\n> I have thought that the large amount of memory consumed by git-rev-list\n> was caused by not-so-sequential access to very large packfile (1.5GB+ if\n> I remember correctly), which I thought causes the whole packfile to be\n> mmapped and not only window, plus large amount of objects in 300MB+ mem\n> range or something; those both would account for around 2GB.\n\nI don't know all of the finer details of chasing revision lists, but\nwould it help if pack files recorded the root IDs of their object\ntrees at creation time and stored it in the front of the pack?\n\n\n>\n> Besides even if git-rev-list wouldn't take so much memory, object\n> enumeration caching would still help with CPU load... admittedly less.\n>\n> --\n> Jakub Narebski\n> Poland\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"110708","messageId":"20090407142147.GA4413@atjola.homenet","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904070903020.6741@xanadu.home","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-07T14:21:47Z","receivedAt":"2009-04-07T14:21:47Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.07 09:13:45 -0400, Nicolas Pitre wrote:\n> On Tue, 7 Apr 2009, Jakub Narebski wrote:\n> \n> > Björn Steinbrink <B.Steinbrink@gmx.de> writes:\n> > > On 2009.04.05 23:24:27 -0400, Nicolas Pitre wrote:\n> > > > On Sun, 5 Apr 2009, Sverre Rabbelier wrote:\n> > > > > \n> > > > > I agree here, we should either say \"look, we don't really support big\n> > > > > repositories because [explanation here], unless you [workarounds\n> > > > > here]\" OR we should work to improve the support we do have. Of course,\n> > > > > the latter option does not magically create developer time to work on\n> > > > > that, but if we do go that way we should at least tell people that we\n> > > > > are aware of the problems and that it's on the global TODO list (not\n> > > > > necessarily on anyone's personal TODO list though).\n> > > > \n> > > > For the record... I at least am aware of the problem and it is indeed on \n> > > > my personal git todo list.  Not that I have a clear solution yet (I've \n> > > > been pondering on some git packing issues for almost 4 years now).\n> > > > \n> > > > Still, in this particular case, the problem appears to be unclear to me, \n> > > > like \"this shouldn't be so bad\".\n> > > \n> > > It's not primarily pack-objects, I think. It's the rev-list that's run\n> > > by upload-pack.  Running \"git rev-list --objects --all\" on that repo\n> > > eats about 2G RSS, easily killing the system's cache on a small box,\n> > > leading to swapping and a painful time reading the packfile contents\n> > > afterwards to send them to the client.\n> > \n> > Than I think that \"packfile caching\" GSoC project (which is IIRC\n> > \"object enumeration caching\", or at least includes it) should help\n> > here.\n> \n> NO!\n> \n> Please people stop being so creative with all sort of ways to simply \n> avoid the real issue and focussing on a real fix.  Git has not become \n> what it is today by the accumulation of workarounds and ignorance of \n> fundamental issues.\n> \n> Having git-rev-list consume about 2G RSS for the enumeration of 4M \n> objects is simply inacceptable, period.  This is the equivalent of 500 \n> bytes per object pinned in memory on average, just for listing object, \n> which is completely silly. We ought to do better than that.\n\nAh, crap, I might have been fooled by \"ps aux\", top actually shows about\n1.3G being shared, likely the mmapped pack files. And that will be\nreused, assuming the box has enough memory to keep all that stuff.\n\nBut that's still 700MB or about 150 bytes per object on average.\n\nA \"struct tree\" is 40 bytes here, adding the average path length (19 in\nthis repo) that's 59 byte, leaving about 90 bytes of \"overhead\" per\nobject, as end the end we seem to care only about the sha1 and the path\nname.\n\nAnd in the upload-pack case, there's also pack-objects running\nconcurrently, already going up to 950M RSS/100M shared _while_ the\nrev-list is still running. So that's 3G of memory usage (2G if you\nignore the shared stuff) before the \"Compressing objects\" part even\nstarts. And of course, pack-objects will apparently start to mmap the\npack files only after the rev-list finished, so a \"smart\" OS might have\nremoved a lot of the mmapped stuff from memory again, causing it to be\nre-read. :-/\n\nBjörn\n"},{"id":"110718","messageId":"alpine.LFD.2.00.0904071321520.6741@xanadu.home","threadId":"18724","inReplyTo":"20090407142147.GA4413@atjola.homenet","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-07T17:48:02Z","receivedAt":"2009-04-07T17:48:02Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 7 Apr 2009, Björn Steinbrink wrote:\n\n> On 2009.04.07 09:13:45 -0400, Nicolas Pitre wrote:\n> > Having git-rev-list consume about 2G RSS for the enumeration of 4M \n> > objects is simply inacceptable, period.  This is the equivalent of 500 \n> > bytes per object pinned in memory on average, just for listing object, \n> > which is completely silly. We ought to do better than that.\n> \n> Ah, crap, I might have been fooled by \"ps aux\", top actually shows about\n> 1.3G being shared, likely the mmapped pack files. And that will be\n> reused, assuming the box has enough memory to keep all that stuff.\n\nRight.  And since the pack is mapped read-only, it can be paged out \neasily by the OS.  And if that doesn't help, we already have \ncore.packedGitWindowSize and core.packedGitLimit config options to play \nwith.\n\n> But that's still 700MB or about 150 bytes per object on average.\n> \n> A \"struct tree\" is 40 bytes here, adding the average path length (19 in\n> this repo) that's 59 byte, leaving about 90 bytes of \"overhead\" per\n> object, as end the end we seem to care only about the sha1 and the path\n> name.\n\nI'm starting to think more seriously about pack v4 again, where each \npath components are indexed in a table.  Because most tree objects are \ndifferent revisions of the same path, this could represent a significant \nsaving in memory as well.\n\n> And in the upload-pack case, there's also pack-objects running\n> concurrently, already going up to 950M RSS/100M shared _while_ the\n> rev-list is still running. So that's 3G of memory usage (2G if you\n> ignore the shared stuff) before the \"Compressing objects\" part even\n> starts. And of course, pack-objects will apparently start to mmap the\n> pack files only after the rev-list finished, so a \"smart\" OS might have\n> removed a lot of the mmapped stuff from memory again, causing it to be\n> re-read. :-/\n\nThe first low hanging fruit to help this case is to make upload-pack use \nthe --revs argument with pack-object to let it do the object enumeration \nitself directly, instead of relying on the rev-list output through a \npipe.  This is what 'git repack' does already.  pack-objects has to \naccess the pack anyway, so this would eliminate an extra access from a \ndifferent process.\n\n\nNicolas\n"},{"id":"110721","messageId":"alpine.LFD.2.00.0904071348420.6741@xanadu.home","threadId":"18724","inReplyTo":"200904071537.04225.jnareb@gmail.com","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-07T17:59:50Z","receivedAt":"2009-04-07T17:59:50Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 7 Apr 2009, Jakub Narebski wrote:\n\n> On Tue, 7 Apr 2009, Nicolas Pitre wrote:\n> > Having git-rev-list consume about 2G RSS for the enumeration of 4M \n> > objects is simply inacceptable, period.  This is the equivalent of 500 \n> > bytes per object pinned in memory on average, just for listing object, \n> > which is completely silly. We ought to do better than that.\n> \n> I have thought that the large amount of memory consumed by git-rev-list\n> was caused by not-so-sequential access to very large packfile (1.5GB+ if\n> I remember correctly), which I thought causes the whole packfile to be\n> mmapped and not only window, plus large amount of objects in 300MB+ mem\n> range or something; those both would account for around 2GB.\n\nThe pack has not to be mapped all at once.  At least on 32-bit machines \nthe total pack mappings cannot exceed 256MB total by default.  On 64-bit \nmachines the default is 8GB which might not work very well if total \namount of RAM is lower than that.\n\nAnother consideration is the object layout in a pack.  Currently we have \ntree and blob objects mixed together so to have sequential pack access \nwhen performing a checkout.  Maybe having trees packed together would \nhelp a lot with object enumeration as the blobs have not to be mapped at \nall.  Remains to see how that might impact other operations though.\n\n> Besides even if git-rev-list wouldn't take so much memory, object\n> enumeration caching would still help with CPU load... admittedly less.\n\nYes, but let's not lose sight of all the inconvenients associated with \nextra caching.  If we can get away without it then all the better.\n\n\nNicolas\n"},{"id":"110724","messageId":"20090407181259.GB4413@atjola.homenet","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904071321520.6741@xanadu.home","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-07T18:12:59Z","receivedAt":"2009-04-07T18:12:59Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.07 13:48:02 -0400, Nicolas Pitre wrote:\n> On Tue, 7 Apr 2009, Björn Steinbrink wrote:\n> > And in the upload-pack case, there's also pack-objects running\n> > concurrently, already going up to 950M RSS/100M shared _while_ the\n> > rev-list is still running. So that's 3G of memory usage (2G if you\n> > ignore the shared stuff) before the \"Compressing objects\" part even\n> > starts. And of course, pack-objects will apparently start to mmap the\n> > pack files only after the rev-list finished, so a \"smart\" OS might have\n> > removed a lot of the mmapped stuff from memory again, causing it to be\n> > re-read. :-/\n> \n> The first low hanging fruit to help this case is to make upload-pack use \n> the --revs argument with pack-object to let it do the object enumeration \n> itself directly, instead of relying on the rev-list output through a \n> pipe.  This is what 'git repack' does already.  pack-objects has to \n> access the pack anyway, so this would eliminate an extra access from a \n> different process.\n\nHm, for an initial clone that would end up as:\ngit pack-objects --stdout --all\nright?\n\nIf so, that doesn't look it it's going to work out as easily as one\nwould hope. Robin said that both processes, git-upload-pack (which does\nthe rev-list) and pack-objects peaked at ~2GB of RSS (which probably\nincludes the mmapped packs). But the above pack-objects with --all peaks\nat 3.1G here, so it basically seems to keep all the stuff in memory that\nthe individual processes had. But this way, it's all at once, not 2G\nfirst and then 2G in a second process, after the first one exitted.\n\nBjörn\n"},{"id":"110727","messageId":"alpine.LFD.2.00.0904071454250.6741@xanadu.home","threadId":"18724","inReplyTo":"20090407181259.GB4413@atjola.homenet","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-07T18:56:41Z","receivedAt":"2009-04-07T18:56:41Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 7 Apr 2009, Björn Steinbrink wrote:\n\n> On 2009.04.07 13:48:02 -0400, Nicolas Pitre wrote:\n> > The first low hanging fruit to help this case is to make upload-pack use \n> > the --revs argument with pack-object to let it do the object enumeration \n> > itself directly, instead of relying on the rev-list output through a \n> > pipe.  This is what 'git repack' does already.  pack-objects has to \n> > access the pack anyway, so this would eliminate an extra access from a \n> > different process.\n> \n> Hm, for an initial clone that would end up as:\n> git pack-objects --stdout --all\n> right?\n> \n> If so, that doesn't look it it's going to work out as easily as one\n> would hope. Robin said that both processes, git-upload-pack (which does\n> the rev-list) and pack-objects peaked at ~2GB of RSS (which probably\n> includes the mmapped packs). But the above pack-objects with --all peaks\n> at 3.1G here, so it basically seems to keep all the stuff in memory that\n> the individual processes had. But this way, it's all at once, not 2G\n> first and then 2G in a second process, after the first one exitted.\n\nRight, and it is probably faster too.\n\nCan I get a copy of that repository somewhere?\n\n\nNicolas\n"},{"id":"110734","messageId":"20090407202725.GC4413@atjola.homenet","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904071454250.6741@xanadu.home","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-07T20:27:25Z","receivedAt":"2009-04-07T20:27:25Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.07 14:56:41 -0400, Nicolas Pitre wrote:\n> On Tue, 7 Apr 2009, Björn Steinbrink wrote:\n> \n> > On 2009.04.07 13:48:02 -0400, Nicolas Pitre wrote:\n> > > The first low hanging fruit to help this case is to make upload-pack use \n> > > the --revs argument with pack-object to let it do the object enumeration \n> > > itself directly, instead of relying on the rev-list output through a \n> > > pipe.  This is what 'git repack' does already.  pack-objects has to \n> > > access the pack anyway, so this would eliminate an extra access from a \n> > > different process.\n> > \n> > Hm, for an initial clone that would end up as:\n> > git pack-objects --stdout --all\n> > right?\n> > \n> > If so, that doesn't look it it's going to work out as easily as one\n> > would hope. Robin said that both processes, git-upload-pack (which does\n> > the rev-list) and pack-objects peaked at ~2GB of RSS (which probably\n> > includes the mmapped packs). But the above pack-objects with --all peaks\n> > at 3.1G here, so it basically seems to keep all the stuff in memory that\n> > the individual processes had. But this way, it's all at once, not 2G\n> > first and then 2G in a second process, after the first one exitted.\n> \n> Right, and it is probably faster too.\n> \n> Can I get a copy of that repository somewhere?\n\nhttp://git.overlays.gentoo.org/gitweb/?p=exp/gentoo-x86.git;a=summary\n\nAt least that's what I cloned ;-) I hope it's the right one, but it fits\nthe description...\n\nBjörn\n"},{"id":"110735","messageId":"20090407202954.GA13501@coredump.intra.peff.net","threadId":"18724","inReplyTo":"20090407181259.GB4413@atjola.homenet","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-04-07T20:29:54Z","receivedAt":"2009-04-07T20:29:54Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Apr 07, 2009 at 08:12:59PM +0200, Björn Steinbrink wrote:\n\n> If so, that doesn't look it it's going to work out as easily as one\n> would hope. Robin said that both processes, git-upload-pack (which does\n> the rev-list) and pack-objects peaked at ~2GB of RSS (which probably\n> includes the mmapped packs). But the above pack-objects with --all peaks\n\nI thought he said this, too, but look at the ps output he posted here:\n\n  http://article.gmane.org/gmane.comp.version-control.git/115739\n\nIt clearly shows upload-pack with a tiny RSS, and pack-objects doing all\nof the damage.\n\n-Peff\n"},{"id":"110736","messageId":"20090407203536.GD4413@atjola.homenet","threadId":"18724","inReplyTo":"20090407202954.GA13501@coredump.intra.peff.net","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-07T20:35:36Z","receivedAt":"2009-04-07T20:35:36Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.07 16:29:54 -0400, Jeff King wrote:\n> On Tue, Apr 07, 2009 at 08:12:59PM +0200, Björn Steinbrink wrote:\n> \n> > If so, that doesn't look it it's going to work out as easily as one\n> > would hope. Robin said that both processes, git-upload-pack (which does\n> > the rev-list) and pack-objects peaked at ~2GB of RSS (which probably\n> > includes the mmapped packs). But the above pack-objects with --all peaks\n> \n> I thought he said this, too, but look at the ps output he posted here:\n> \n>   http://article.gmane.org/gmane.comp.version-control.git/115739\n> \n> It clearly shows upload-pack with a tiny RSS, and pack-objects doing all\n> of the damage.\n\nThat second git-upload-pack is the interesting one. upload-pack forks to\ndo the rev-list stuff, without changing its process name, so it keeps\nbeing listed as upload-pack. And as the process already died, its\nRSS/VZS dropped to zero.\n\nBjörn\n"},{"id":"110781","messageId":"alpine.LFD.2.00.0904080041240.6741@xanadu.home","threadId":"18724","inReplyTo":"20090407202725.GC4413@atjola.homenet","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-08T04:52:54Z","receivedAt":"2009-04-08T04:52:54Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 7 Apr 2009, Björn Steinbrink wrote:\n\n> On 2009.04.07 14:56:41 -0400, Nicolas Pitre wrote:\n> > On Tue, 7 Apr 2009, Björn Steinbrink wrote:\n> > \n> > > On 2009.04.07 13:48:02 -0400, Nicolas Pitre wrote:\n> > > > The first low hanging fruit to help this case is to make upload-pack use \n> > > > the --revs argument with pack-object to let it do the object enumeration \n> > > > itself directly, instead of relying on the rev-list output through a \n> > > > pipe.  This is what 'git repack' does already.  pack-objects has to \n> > > > access the pack anyway, so this would eliminate an extra access from a \n> > > > different process.\n> > > \n> > > Hm, for an initial clone that would end up as:\n> > > git pack-objects --stdout --all\n> > > right?\n> > > \n> > > If so, that doesn't look it it's going to work out as easily as one\n> > > would hope. Robin said that both processes, git-upload-pack (which does\n> > > the rev-list) and pack-objects peaked at ~2GB of RSS (which probably\n> > > includes the mmapped packs). But the above pack-objects with --all peaks\n> > > at 3.1G here, so it basically seems to keep all the stuff in memory that\n> > > the individual processes had. But this way, it's all at once, not 2G\n> > > first and then 2G in a second process, after the first one exitted.\n> > \n> > Right, and it is probably faster too.\n> > \n> > Can I get a copy of that repository somewhere?\n> \n> http://git.overlays.gentoo.org/gitweb/?p=exp/gentoo-x86.git;a=summary\n> \n> At least that's what I cloned ;-) I hope it's the right one, but it fits\n> the description...\n\nOK.  FWIW, I repacked it with --window=250 --depth=250 and obtained a \n725MB pack file.  So that's about half the originally reported size.\n\n\nNicolas\n"},{"id":"110829","messageId":"20090408112854.GA8624@atjola.homenet","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904070903020.6741@xanadu.home","subject":"[PATCH] process_{tree,blob}: Remove useless xstrdup calls","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-08T11:28:54Z","receivedAt":"2009-04-08T11:28:54Z","isPatch":true,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"The name of the processed object was duplicated for passing it to\nadd_object(), but that already calls path_name, which allocates a new\nstring anyway. So the memory allocated by the xstrdup calls just went\nnowhere, leaking memory.\n\nSigned-off-by: Björn Steinbrink <B.Steinbrink@gmx.de>\n---\nThis reduces the RSS usage for a \"rev-list --all --objects\" by about 10% on\nthe gentoo repo (fully packed) as well as linux-2.6.git:\n\ngentoo:\n\t\t| old\t\t| new\t\t\n----------------|-------------------------------\nRSS\t\t|\t1537284 |\t1388408\nVSZ\t\t|\t1816852 |\t1667952\ntime elapsed\t|\t1:49.62 |\t1:48.99\nmin. page faults|\t 417178 |\t 379919\n\nlinux-2.6.git:\n\t\t| old\t\t| new\t\t\n----------------|-------------------------------\nRSS\t\t|\t 324452 |\t 292996\nVSZ\t\t|\t 491792 |\t 460376\ntime elapsed\t|\t0:14.53 |\t0:14.28\nmin. page faults|\t  89360 |\t  81613\n\n list-objects.c |    2 --\n reachable.c    |    1 -\n 2 files changed, 0 insertions(+), 3 deletions(-)\n\ndiff --git a/list-objects.c b/list-objects.c\nindex c8b8375..dd243c7 100644\n--- a/list-objects.c\n+++ b/list-objects.c\n@@ -23,7 +23,6 @@ static void process_blob(struct rev_info *revs,\n \tif (obj->flags & (UNINTERESTING | SEEN))\n \t\treturn;\n \tobj->flags |= SEEN;\n-\tname = xstrdup(name);\n \tadd_object(obj, p, path, name);\n }\n \n@@ -78,7 +77,6 @@ static void process_tree(struct rev_info *revs,\n \tif (parse_tree(tree) < 0)\n \t\tdie(\"bad tree object %s\", sha1_to_hex(obj->sha1));\n \tobj->flags |= SEEN;\n-\tname = xstrdup(name);\n \tadd_object(obj, p, path, name);\n \tme.up = path;\n \tme.elem = name;\ndiff --git a/reachable.c b/reachable.c\nindex 3b1c18f..b515fa2 100644\n--- a/reachable.c\n+++ b/reachable.c\n@@ -48,7 +48,6 @@ static void process_tree(struct tree *tree,\n \tobj->flags |= SEEN;\n \tif (parse_tree(tree) < 0)\n \t\tdie(\"bad tree object %s\", sha1_to_hex(obj->sha1));\n-\tname = xstrdup(name);\n \tadd_object(obj, p, path, name);\n \tme.up = path;\n \tme.elem = name;\n-- \n1.6.2.2.446.gfbdc0.dirty\n"},{"id":"111029","messageId":"20090410T203405Z@curie.orbis-terrarum.net","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904080041240.6741@xanadu.home","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Robin H. Johnson","fromEmail":"robbat2@gentoo.org","sentAt":"2009-04-10T20:38:46Z","receivedAt":"2009-04-10T20:38:46Z","isPatch":false,"sender":{"key":"robbat2@gentoo.org","avatar":"https://avatars.githubusercontent.com/u/373898?v=4"},"body":"On Wed, Apr 08, 2009 at 12:52:54AM -0400, Nicolas Pitre wrote:\n> > http://git.overlays.gentoo.org/gitweb/?p=exp/gentoo-x86.git;a=summary\n> > At least that's what I cloned ;-) I hope it's the right one, but it fits\n> > the description...\n> OK.  FWIW, I repacked it with --window=250 --depth=250 and obtained a \n> 725MB pack file.  So that's about half the originally reported size.\nThe one problem with having the single large packfile is that Git\ndoesn't have a trivial way to resume downloading it when the git://\nprotocol is used.\n\nFor our developers cursed with bad internet connections (a fair number\nof firewalls that don't seem to respect keepalive properly), I suppose\nI can probably just maintain a separate repo for their initial clones,\nwhich leaves a large overall download, but more chances to resume.\n\nPS #1: B.Steinbrink's memory improvement patch seems to work nicely too,\nbut more memory improvements in that realm are still needed.\n\nPS #2: We finally got some newer hardware to run the large repo, I'm\nworking on the install now, but until the memory issue is better\nresolved, I'm still worried we might run short if there are too many\nconcurrent clones.\n\n-- \nRobin Hugh Johnson\nGentoo Linux Developer & Infra Guy\nE-Mail     : robbat2@gentoo.org\nGnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n"},{"id":"111030","messageId":"alpine.LFD.2.00.0904101517520.4583@localhost.localdomain","threadId":"18724","inReplyTo":"20090408112854.GA8624@atjola.homenet","subject":"Re: [PATCH] process_{tree,blob}: Remove useless xstrdup calls","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-04-10T22:20:18Z","receivedAt":"2009-04-10T22:20:18Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 8 Apr 2009, Björn Steinbrink wrote:\n>\n> The name of the processed object was duplicated for passing it to\n> add_object(), but that already calls path_name, which allocates a new\n> string anyway. So the memory allocated by the xstrdup calls just went\n> nowhere, leaking memory.\n\nAck, ack.\n\nThere's another easy 5% or so for the built-in object walker: once we've \ncreated the hash from the name, the name isn't interesting any more, and \nso something trivial like this can help a bit.\n\nDoes it matter? Probably not on its own. But a few more memory saving \ntricks and it might all make a difference.\n\n\t\tLinus\n\n---\n builtin-pack-objects.c |    2 ++\n 1 files changed, 2 insertions(+), 0 deletions(-)\n\ndiff --git a/builtin-pack-objects.c b/builtin-pack-objects.c\nindex 9fc3b35..d00eabe 100644\n--- a/builtin-pack-objects.c\n+++ b/builtin-pack-objects.c\n@@ -1912,6 +1912,8 @@ static void show_object(struct object_array_entry *p)\n \tadd_preferred_base_object(p->name);\n \tadd_object_entry(p->item->sha1, p->item->type, p->name, 0);\n \tp->item->flags |= OBJECT_ADDED;\n+\tfree(p->name);\n+\tp->name = NULL;\n }\n \n static void show_edge(struct commit *commit)\n"},{"id":"111031","messageId":"alpine.LFD.2.00.0904101714420.4583@localhost.localdomain","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904101517520.4583@localhost.localdomain","subject":"Re: [PATCH] process_{tree,blob}: Remove useless xstrdup calls","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-04-11T00:27:58Z","receivedAt":"2009-04-11T00:27:58Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 10 Apr 2009, Linus Torvalds wrote:\n> \n> There's another easy 5% or so for the built-in object walker: once we've \n> created the hash from the name, the name isn't interesting any more, and \n> so something trivial like this can help a bit.\n\nHmm.\n\nHere's a less trivial thing, and slightly more dubious one.\n\nI was looking at that \"struct object_array objects\", and wondering why we \ndo that. I have honestly totally forgotten. Why not just call the \"show()\" \nfunction as we encounter the objects? Rather than add the objects to the \nobject_array, and then at the very end going through the array and doing a \n'show' on all, just do things more incrementally.\n\nNow, there are possible downsides to this:\n\n - the \"buffer using object_array\" _can_ in theory result in at least \n   better I-cache usage (two tight loops rather than one more spread out \n   one). I don't think this is a real issue, but in theory..\n\n - this _does_ change the order of the objects printed. Instead of doing a \n   \"process_tree(revs, commit->tree, &objects, NULL, \"\");\" in the loop \n   over the commits (which puts all the root trees _first_ in the object \n   list, this patch just adds them to the list of pending objects, and \n   then we'll traverse them in that order (and thus show each root tree \n   object together with the objects we discover under it)\n\n   I _think_ the new ordering actually makes more sense, but the object \n   ordering is actually a subtle thing when it comes to packing \n   efficiency, so any change in order is going to have implications for \n   packing. Good or bad, I dunno.\n\n - There may be some reason why we did it that odd way with the object \n   array, that I have simply forgotten.\n\nAnyway, this includes the \"free(name)\" in builtin-pack-objects.c: \nshow_object() logic, and now that we don't buffer up the objects before \nshowing them that may actually result in lower memory usage during that \nwhole traverse_commit_list() phase.\n\nThis is seriously not very deeply tested. It makes sense to me, it seems \nto pass all the tests, it looks ok, but...\n\nDoes anybody remember why we did that \"object_array\" thing? It used to be \nan \"object_list\" a long long time ago, but got changed into the array due \nto better memory usage patterns (those linked lists of obejcts are \nhorrible from a memory allocation standpoint). But I wonder why we didn't \ndo this back then. Maybe there's a reason for it.\n\nOr maybe there _used_ to be a reason, and no longer is. \n\n\t\t\tLinus\n\n---\n builtin-pack-objects.c |   14 ++++++++++----\n builtin-rev-list.c     |   20 ++++++++++----------\n list-objects.c         |   35 ++++++++++++++++++-----------------\n list-objects.h         |    2 +-\n revision.c             |    2 +-\n revision.h             |    2 ++\n upload-pack.c          |   12 ++++++------\n 7 files changed, 48 insertions(+), 39 deletions(-)\n\ndiff --git a/builtin-pack-objects.c b/builtin-pack-objects.c\nindex 9fc3b35..e028a02 100644\n--- a/builtin-pack-objects.c\n+++ b/builtin-pack-objects.c\n@@ -1907,11 +1907,17 @@ static void show_commit(struct commit *commit)\n \tcommit->object.flags |= OBJECT_ADDED;\n }\n \n-static void show_object(struct object_array_entry *p)\n+static void show_object(struct object *obj, const char *name)\n {\n-\tadd_preferred_base_object(p->name);\n-\tadd_object_entry(p->item->sha1, p->item->type, p->name, 0);\n-\tp->item->flags |= OBJECT_ADDED;\n+\tadd_preferred_base_object(name);\n+\tadd_object_entry(obj->sha1, obj->type, name, 0);\n+\tobj->flags |= OBJECT_ADDED;\n+\n+\t/*\n+\t * We will have generated the hash from the name,\n+\t * but not saved a pointer to it - we can free it\n+\t */\n+\tfree(name);\n }\n \n static void show_edge(struct commit *commit)\ndiff --git a/builtin-rev-list.c b/builtin-rev-list.c\nindex 40d5fcb..0815cf3 100644\n--- a/builtin-rev-list.c\n+++ b/builtin-rev-list.c\n@@ -168,27 +168,27 @@ static void finish_commit(struct commit *commit)\n \tcommit->buffer = NULL;\n }\n \n-static void finish_object(struct object_array_entry *p)\n+static void finish_object(struct object *obj, const char *name)\n {\n-\tif (p->item->type == OBJ_BLOB && !has_sha1_file(p->item->sha1))\n-\t\tdie(\"missing blob object '%s'\", sha1_to_hex(p->item->sha1));\n+\tif (obj->type == OBJ_BLOB && !has_sha1_file(obj->sha1))\n+\t\tdie(\"missing blob object '%s'\", sha1_to_hex(obj->sha1));\n }\n \n-static void show_object(struct object_array_entry *p)\n+static void show_object(struct object *obj, const char *name)\n {\n \t/* An object with name \"foo\\n0000000...\" can be used to\n \t * confuse downstream \"git pack-objects\" very badly.\n \t */\n-\tconst char *ep = strchr(p->name, '\\n');\n+\tconst char *ep = strchr(name, '\\n');\n \n-\tfinish_object(p);\n+\tfinish_object(obj, name);\n \tif (ep) {\n-\t\tprintf(\"%s %.*s\\n\", sha1_to_hex(p->item->sha1),\n-\t\t       (int) (ep - p->name),\n-\t\t       p->name);\n+\t\tprintf(\"%s %.*s\\n\", sha1_to_hex(obj->sha1),\n+\t\t       (int) (ep - name),\n+\t\t       name);\n \t}\n \telse\n-\t\tprintf(\"%s %s\\n\", sha1_to_hex(p->item->sha1), p->name);\n+\t\tprintf(\"%s %s\\n\", sha1_to_hex(obj->sha1), name);\n }\n \n static void show_edge(struct commit *commit)\ndiff --git a/list-objects.c b/list-objects.c\nindex dd243c7..5a4af62 100644\n--- a/list-objects.c\n+++ b/list-objects.c\n@@ -10,7 +10,7 @@\n \n static void process_blob(struct rev_info *revs,\n \t\t\t struct blob *blob,\n-\t\t\t struct object_array *p,\n+\t\t\t show_object_fn show,\n \t\t\t struct name_path *path,\n \t\t\t const char *name)\n {\n@@ -23,7 +23,7 @@ static void process_blob(struct rev_info *revs,\n \tif (obj->flags & (UNINTERESTING | SEEN))\n \t\treturn;\n \tobj->flags |= SEEN;\n-\tadd_object(obj, p, path, name);\n+\tshow(obj, path_name(path, name));\n }\n \n /*\n@@ -50,7 +50,7 @@ static void process_blob(struct rev_info *revs,\n  */\n static void process_gitlink(struct rev_info *revs,\n \t\t\t    const unsigned char *sha1,\n-\t\t\t    struct object_array *p,\n+\t\t\t    show_object_fn show,\n \t\t\t    struct name_path *path,\n \t\t\t    const char *name)\n {\n@@ -59,7 +59,7 @@ static void process_gitlink(struct rev_info *revs,\n \n static void process_tree(struct rev_info *revs,\n \t\t\t struct tree *tree,\n-\t\t\t struct object_array *p,\n+\t\t\t show_object_fn show,\n \t\t\t struct name_path *path,\n \t\t\t const char *name)\n {\n@@ -77,7 +77,7 @@ static void process_tree(struct rev_info *revs,\n \tif (parse_tree(tree) < 0)\n \t\tdie(\"bad tree object %s\", sha1_to_hex(obj->sha1));\n \tobj->flags |= SEEN;\n-\tadd_object(obj, p, path, name);\n+\tshow(obj, path_name(path, name));\n \tme.up = path;\n \tme.elem = name;\n \tme.elem_len = strlen(name);\n@@ -88,14 +88,14 @@ static void process_tree(struct rev_info *revs,\n \t\tif (S_ISDIR(entry.mode))\n \t\t\tprocess_tree(revs,\n \t\t\t\t     lookup_tree(entry.sha1),\n-\t\t\t\t     p, &me, entry.path);\n+\t\t\t\t     show, &me, entry.path);\n \t\telse if (S_ISGITLINK(entry.mode))\n \t\t\tprocess_gitlink(revs, entry.sha1,\n-\t\t\t\t\tp, &me, entry.path);\n+\t\t\t\t\tshow, &me, entry.path);\n \t\telse\n \t\t\tprocess_blob(revs,\n \t\t\t\t     lookup_blob(entry.sha1),\n-\t\t\t\t     p, &me, entry.path);\n+\t\t\t\t     show, &me, entry.path);\n \t}\n \tfree(tree->buffer);\n \ttree->buffer = NULL;\n@@ -134,16 +134,20 @@ void mark_edges_uninteresting(struct commit_list *list,\n \t}\n }\n \n+static void add_pending_tree(struct rev_info *revs, struct tree *tree)\n+{\n+\tadd_pending_object(revs, &tree->object, \"\");\n+}\n+\n void traverse_commit_list(struct rev_info *revs,\n \t\t\t  void (*show_commit)(struct commit *),\n-\t\t\t  void (*show_object)(struct object_array_entry *))\n+\t\t\t  void (*show_object)(struct object *, const char *))\n {\n \tint i;\n \tstruct commit *commit;\n-\tstruct object_array objects = { 0, 0, NULL };\n \n \twhile ((commit = get_revision(revs)) != NULL) {\n-\t\tprocess_tree(revs, commit->tree, &objects, NULL, \"\");\n+\t\tadd_pending_tree(revs, commit->tree);\n \t\tshow_commit(commit);\n \t}\n \tfor (i = 0; i < revs->pending.nr; i++) {\n@@ -154,25 +158,22 @@ void traverse_commit_list(struct rev_info *revs,\n \t\t\tcontinue;\n \t\tif (obj->type == OBJ_TAG) {\n \t\t\tobj->flags |= SEEN;\n-\t\t\tadd_object_array(obj, name, &objects);\n+\t\t\tshow_object(obj, name);\n \t\t\tcontinue;\n \t\t}\n \t\tif (obj->type == OBJ_TREE) {\n-\t\t\tprocess_tree(revs, (struct tree *)obj, &objects,\n+\t\t\tprocess_tree(revs, (struct tree *)obj, show_object,\n \t\t\t\t     NULL, name);\n \t\t\tcontinue;\n \t\t}\n \t\tif (obj->type == OBJ_BLOB) {\n-\t\t\tprocess_blob(revs, (struct blob *)obj, &objects,\n+\t\t\tprocess_blob(revs, (struct blob *)obj, show_object,\n \t\t\t\t     NULL, name);\n \t\t\tcontinue;\n \t\t}\n \t\tdie(\"unknown pending object %s (%s)\",\n \t\t    sha1_to_hex(obj->sha1), name);\n \t}\n-\tfor (i = 0; i < objects.nr; i++)\n-\t\tshow_object(&objects.objects[i]);\n-\tfree(objects.objects);\n \tif (revs->pending.nr) {\n \t\tfree(revs->pending.objects);\n \t\trevs->pending.nr = 0;\ndiff --git a/list-objects.h b/list-objects.h\nindex 0f41391..13b0dd9 100644\n--- a/list-objects.h\n+++ b/list-objects.h\n@@ -2,7 +2,7 @@\n #define LIST_OBJECTS_H\n \n typedef void (*show_commit_fn)(struct commit *);\n-typedef void (*show_object_fn)(struct object_array_entry *);\n+typedef void (*show_object_fn)(struct object *, const char *);\n typedef void (*show_edge_fn)(struct commit *);\n \n void traverse_commit_list(struct rev_info *revs, show_commit_fn, show_object_fn);\ndiff --git a/revision.c b/revision.c\nindex b6215cc..44a9ce2 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -15,7 +15,7 @@\n \n volatile show_early_output_fn_t show_early_output;\n \n-static char *path_name(struct name_path *path, const char *name)\n+char *path_name(struct name_path *path, const char *name)\n {\n \tstruct name_path *p;\n \tchar *n, *m;\ndiff --git a/revision.h b/revision.h\nindex 5adfc91..c89e8ff 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -146,6 +146,8 @@ struct name_path {\n \tconst char *elem;\n };\n \n+char *path_name(struct name_path *path, const char *name);\n+\n extern void add_object(struct object *obj,\n \t\t       struct object_array *p,\n \t\t       struct name_path *path,\ndiff --git a/upload-pack.c b/upload-pack.c\nindex a49d872..5524ac4 100644\n--- a/upload-pack.c\n+++ b/upload-pack.c\n@@ -78,20 +78,20 @@ static void show_commit(struct commit *commit)\n \tcommit->buffer = NULL;\n }\n \n-static void show_object(struct object_array_entry *p)\n+static void show_object(struct object *obj, const char *name)\n {\n \t/* An object with name \"foo\\n0000000...\" can be used to\n \t * confuse downstream git-pack-objects very badly.\n \t */\n-\tconst char *ep = strchr(p->name, '\\n');\n+\tconst char *ep = strchr(name, '\\n');\n \tif (ep) {\n-\t\tfprintf(pack_pipe, \"%s %.*s\\n\", sha1_to_hex(p->item->sha1),\n-\t\t       (int) (ep - p->name),\n-\t\t       p->name);\n+\t\tfprintf(pack_pipe, \"%s %.*s\\n\", sha1_to_hex(obj->sha1),\n+\t\t       (int) (ep - name),\n+\t\t       name);\n \t}\n \telse\n \t\tfprintf(pack_pipe, \"%s %s\\n\",\n-\t\t\t\tsha1_to_hex(p->item->sha1), p->name);\n+\t\t\t\tsha1_to_hex(obj->sha1), name);\n }\n \n static void show_edge(struct commit *commit)\n"},{"id":"111035","messageId":"alpine.LFD.2.00.0904101806340.4583@localhost.localdomain","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904101714420.4583@localhost.localdomain","subject":"Re: [PATCH] process_{tree,blob}: Remove useless xstrdup calls","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-04-11T01:15:26Z","receivedAt":"2009-04-11T01:15:26Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 10 Apr 2009, Linus Torvalds wrote:\n> \n> Here's a less trivial thing, and slightly more dubious one.\n\nI'm starting to like it more.\n\nIn particular, pushing the \"path_name()\" call _into_ the show() function \nwould seem to allow\n\n - more clarity into who \"owns\" the name (ie now when we free the name in \n   the show_object callback, it's because we generated it ourselves by \n   calling path_name())\n\n - not calling path_name() at all, either because we don't care about the \n   name in the first place, or because we are actually happy walking the \n   linked list of \"struct name_path *\" and the last component.\n\nNow, I didn't do that latter optimization, because it would require some \nmore coding, but especially looking at \"builtin-pack-objects.c\", we really \ndon't even want the whole pathname, we really would be better off with the \nlist of path components.\n\nWhy? We use that name for two things:\n - add_preferred_base_object(), which actually _wants_ to traverse the \n   path, and now does it by looking for '/' characters!\n - for 'name_hash()', which only cares about the last 16 characters of a \n   name, so again, generating the full name seems to be just unnecessary \n   work.\n\nAnyway, so I didn't look any closer at those things, but it did convince \nme that the \"show_object()\" calling convention was crazy, and we're \nactually better off doing _less_ in list-objects.c, and giving people \naccess to the internal data structures so that they can decide whether \nthey want to generate a path-name or not.\n\nThis patch does that, and then for people who did use the name (even if \nthey might do something more clever in the future), it just does the \nstraightforward \"name = path_name(path, component); .. free(name);\" thing.\n\nIt obviously goes on top of my previous patch.\n\n\n\t\tLinus\n\n---\n builtin-pack-objects.c |    9 +++------\n builtin-rev-list.c     |    8 +++++---\n list-objects.c         |   10 +++++-----\n list-objects.h         |    2 +-\n revision.c             |    4 ++--\n revision.h             |    2 +-\n upload-pack.c          |    4 +++-\n 7 files changed, 20 insertions(+), 19 deletions(-)\n\ndiff --git a/builtin-pack-objects.c b/builtin-pack-objects.c\nindex e028a02..d74d8a4 100644\n--- a/builtin-pack-objects.c\n+++ b/builtin-pack-objects.c\n@@ -1907,16 +1907,13 @@ static void show_commit(struct commit *commit)\n \tcommit->object.flags |= OBJECT_ADDED;\n }\n \n-static void show_object(struct object *obj, const char *name)\n+static void show_object(struct object *obj, const struct name_path *path, const char *last)\n {\n+\tchar *name = path_name(path, last);\n+\n \tadd_preferred_base_object(name);\n \tadd_object_entry(obj->sha1, obj->type, name, 0);\n \tobj->flags |= OBJECT_ADDED;\n-\n-\t/*\n-\t * We will have generated the hash from the name,\n-\t * but not saved a pointer to it - we can free it\n-\t */\n \tfree(name);\n }\n \ndiff --git a/builtin-rev-list.c b/builtin-rev-list.c\nindex 0815cf3..aba8a6f 100644\n--- a/builtin-rev-list.c\n+++ b/builtin-rev-list.c\n@@ -168,20 +168,21 @@ static void finish_commit(struct commit *commit)\n \tcommit->buffer = NULL;\n }\n \n-static void finish_object(struct object *obj, const char *name)\n+static void finish_object(struct object *obj, const struct name_path *path, const char *name)\n {\n \tif (obj->type == OBJ_BLOB && !has_sha1_file(obj->sha1))\n \t\tdie(\"missing blob object '%s'\", sha1_to_hex(obj->sha1));\n }\n \n-static void show_object(struct object *obj, const char *name)\n+static void show_object(struct object *obj, const struct name_path *path, const char *component)\n {\n+\tchar *name = path_name(path, component);\n \t/* An object with name \"foo\\n0000000...\" can be used to\n \t * confuse downstream \"git pack-objects\" very badly.\n \t */\n \tconst char *ep = strchr(name, '\\n');\n \n-\tfinish_object(obj, name);\n+\tfinish_object(obj, path, name);\n \tif (ep) {\n \t\tprintf(\"%s %.*s\\n\", sha1_to_hex(obj->sha1),\n \t\t       (int) (ep - name),\n@@ -189,6 +190,7 @@ static void show_object(struct object *obj, const char *name)\n \t}\n \telse\n \t\tprintf(\"%s %s\\n\", sha1_to_hex(obj->sha1), name);\n+\tfree(name);\n }\n \n static void show_edge(struct commit *commit)\ndiff --git a/list-objects.c b/list-objects.c\nindex 5a4af62..30ded3d 100644\n--- a/list-objects.c\n+++ b/list-objects.c\n@@ -23,7 +23,7 @@ static void process_blob(struct rev_info *revs,\n \tif (obj->flags & (UNINTERESTING | SEEN))\n \t\treturn;\n \tobj->flags |= SEEN;\n-\tshow(obj, path_name(path, name));\n+\tshow(obj, path, name);\n }\n \n /*\n@@ -77,7 +77,7 @@ static void process_tree(struct rev_info *revs,\n \tif (parse_tree(tree) < 0)\n \t\tdie(\"bad tree object %s\", sha1_to_hex(obj->sha1));\n \tobj->flags |= SEEN;\n-\tshow(obj, path_name(path, name));\n+\tshow(obj, path, name);\n \tme.up = path;\n \tme.elem = name;\n \tme.elem_len = strlen(name);\n@@ -140,8 +140,8 @@ static void add_pending_tree(struct rev_info *revs, struct tree *tree)\n }\n \n void traverse_commit_list(struct rev_info *revs,\n-\t\t\t  void (*show_commit)(struct commit *),\n-\t\t\t  void (*show_object)(struct object *, const char *))\n+\t\t\t  show_commit_fn show_commit,\n+\t\t\t  show_object_fn show_object)\n {\n \tint i;\n \tstruct commit *commit;\n@@ -158,7 +158,7 @@ void traverse_commit_list(struct rev_info *revs,\n \t\t\tcontinue;\n \t\tif (obj->type == OBJ_TAG) {\n \t\t\tobj->flags |= SEEN;\n-\t\t\tshow_object(obj, name);\n+\t\t\tshow_object(obj, NULL, name);\n \t\t\tcontinue;\n \t\t}\n \t\tif (obj->type == OBJ_TREE) {\ndiff --git a/list-objects.h b/list-objects.h\nindex 13b0dd9..0b2de64 100644\n--- a/list-objects.h\n+++ b/list-objects.h\n@@ -2,7 +2,7 @@\n #define LIST_OBJECTS_H\n \n typedef void (*show_commit_fn)(struct commit *);\n-typedef void (*show_object_fn)(struct object *, const char *);\n+typedef void (*show_object_fn)(struct object *, const struct name_path *, const char *);\n typedef void (*show_edge_fn)(struct commit *);\n \n void traverse_commit_list(struct rev_info *revs, show_commit_fn, show_object_fn);\ndiff --git a/revision.c b/revision.c\nindex 44a9ce2..bd0ea34 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -15,9 +15,9 @@\n \n volatile show_early_output_fn_t show_early_output;\n \n-char *path_name(struct name_path *path, const char *name)\n+char *path_name(const struct name_path *path, const char *name)\n {\n-\tstruct name_path *p;\n+\tconst struct name_path *p;\n \tchar *n, *m;\n \tint nlen = strlen(name);\n \tint len = nlen + 1;\ndiff --git a/revision.h b/revision.h\nindex c89e8ff..be39e7d 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -146,7 +146,7 @@ struct name_path {\n \tconst char *elem;\n };\n \n-char *path_name(struct name_path *path, const char *name);\n+char *path_name(const struct name_path *path, const char *name);\n \n extern void add_object(struct object *obj,\n \t\t       struct object_array *p,\ndiff --git a/upload-pack.c b/upload-pack.c\nindex 5524ac4..536efbb 100644\n--- a/upload-pack.c\n+++ b/upload-pack.c\n@@ -78,11 +78,12 @@ static void show_commit(struct commit *commit)\n \tcommit->buffer = NULL;\n }\n \n-static void show_object(struct object *obj, const char *name)\n+static void show_object(struct object *obj, const struct name_path *path, const char *component)\n {\n \t/* An object with name \"foo\\n0000000...\" can be used to\n \t * confuse downstream git-pack-objects very badly.\n \t */\n+\tconst char *name = path_name(path, component);\n \tconst char *ep = strchr(name, '\\n');\n \tif (ep) {\n \t\tfprintf(pack_pipe, \"%s %.*s\\n\", sha1_to_hex(obj->sha1),\n@@ -92,6 +93,7 @@ static void show_object(struct object *obj, const char *name)\n \telse\n \t\tfprintf(pack_pipe, \"%s %s\\n\",\n \t\t\t\tsha1_to_hex(obj->sha1), name);\n+\tfree(name);\n }\n \n static void show_edge(struct commit *commit)\n"},{"id":"111037","messageId":"alpine.LFD.2.00.0904102129510.6741@xanadu.home","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904101806340.4583@localhost.localdomain","subject":"Re: [PATCH] process_{tree,blob}: Remove useless xstrdup calls","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-11T01:34:12Z","receivedAt":"2009-04-11T01:34:12Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 10 Apr 2009, Linus Torvalds wrote:\n\n> \n> \n> On Fri, 10 Apr 2009, Linus Torvalds wrote:\n> > \n> > Here's a less trivial thing, and slightly more dubious one.\n> \n> I'm starting to like it more.\n> \n> In particular, pushing the \"path_name()\" call _into_ the show() function \n> would seem to allow\n> \n>  - more clarity into who \"owns\" the name (ie now when we free the name in \n>    the show_object callback, it's because we generated it ourselves by \n>    calling path_name())\n> \n>  - not calling path_name() at all, either because we don't care about the \n>    name in the first place, or because we are actually happy walking the \n>    linked list of \"struct name_path *\" and the last component.\n> \n> Now, I didn't do that latter optimization, because it would require some \n> more coding, but especially looking at \"builtin-pack-objects.c\", we really \n> don't even want the whole pathname, we really would be better off with the \n> list of path components.\n> \n> Why? We use that name for two things:\n>  - add_preferred_base_object(), which actually _wants_ to traverse the \n>    path, and now does it by looking for '/' characters!\n>  - for 'name_hash()', which only cares about the last 16 characters of a \n>    name, so again, generating the full name seems to be just unnecessary \n>    work.\n> \n> Anyway, so I didn't look any closer at those things, but it did convince \n> me that the \"show_object()\" calling convention was crazy, and we're \n> actually better off doing _less_ in list-objects.c, and giving people \n> access to the internal data structures so that they can decide whether \n> they want to generate a path-name or not.\n\nYES!\n\nI didn't look at the patch really closely, but this fits pretty well \nwith the philosophy behind pack v4 where path components are stored in a \nseparate table (instead of being duplicated in every tree objects for \nthe same path), hence generating path names on demand would be a real \nwin for those cases where it is not needed.\n\n\nNicolas\n"},{"id":"111038","messageId":"alpine.LFD.2.00.0904102147590.6741@xanadu.home","threadId":"18724","inReplyTo":"20090410T203405Z@curie.orbis-terrarum.net","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-11T01:58:11Z","receivedAt":"2009-04-11T01:58:11Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 10 Apr 2009, Robin H. Johnson wrote:\n\n> On Wed, Apr 08, 2009 at 12:52:54AM -0400, Nicolas Pitre wrote:\n> > > http://git.overlays.gentoo.org/gitweb/?p=exp/gentoo-x86.git;a=summary\n> > > At least that's what I cloned ;-) I hope it's the right one, but it fits\n> > > the description...\n> > OK.  FWIW, I repacked it with --window=250 --depth=250 and obtained a \n> > 725MB pack file.  So that's about half the originally reported size.\n> The one problem with having the single large packfile is that Git\n> doesn't have a trivial way to resume downloading it when the git://\n> protocol is used.\n\nHaving multiple packs won't help the git:// protocol at all in that \nregard.  In fact it'll just make it a bit harder on the server for all \ncases, which has to generate a single pack for streaming anyway by using \nmultiple source ones and perform extra work in attempting delta \ncompression across pack boundaries.\n\n> For our developers cursed with bad internet connections (a fair number\n> of firewalls that don't seem to respect keepalive properly), I suppose\n> I can probably just maintain a separate repo for their initial clones,\n> which leaves a large overall download, but more chances to resume.\n\nI don't know much about git's http protocol implementation, but I guess \nit should be able to resume the transfer of a pack file which might have \nbeen interrupted in the middle?  If no then this should be considered.\n\n> PS #1: B.Steinbrink's memory improvement patch seems to work nicely too,\n> but more memory improvements in that realm are still needed.\n\nGood.\n\n> PS #2: We finally got some newer hardware to run the large repo, I'm\n> working on the install now, but until the memory issue is better\n> resolved, I'm still worried we might run short if there are too many\n> concurrent clones.\n\nRight.\n\n\nNicolas (who wishes he could dedicate more time to git hacking)\n"},{"id":"111042","messageId":"20090411070605.GA29851@glandium.org","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904102147590.6741@xanadu.home","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2009-04-11T07:06:05Z","receivedAt":"2009-04-11T07:06:05Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Fri, Apr 10, 2009 at 09:58:11PM -0400, Nicolas Pitre wrote:\n> On Fri, 10 Apr 2009, Robin H. Johnson wrote:\n> \n> > On Wed, Apr 08, 2009 at 12:52:54AM -0400, Nicolas Pitre wrote:\n> > > > http://git.overlays.gentoo.org/gitweb/?p=exp/gentoo-x86.git;a=summary\n> > > > At least that's what I cloned ;-) I hope it's the right one, but it fits\n> > > > the description...\n> > > OK.  FWIW, I repacked it with --window=250 --depth=250 and obtained a \n> > > 725MB pack file.  So that's about half the originally reported size.\n> > The one problem with having the single large packfile is that Git\n> > doesn't have a trivial way to resume downloading it when the git://\n> > protocol is used.\n> \n> Having multiple packs won't help the git:// protocol at all in that \n> regard.  In fact it'll just make it a bit harder on the server for all \n> cases, which has to generate a single pack for streaming anyway by using \n> multiple source ones and perform extra work in attempting delta \n> compression across pack boundaries.\n> \n> > For our developers cursed with bad internet connections (a fair number\n> > of firewalls that don't seem to respect keepalive properly), I suppose\n> > I can probably just maintain a separate repo for their initial clones,\n> > which leaves a large overall download, but more chances to resume.\n> \n> I don't know much about git's http protocol implementation, but I guess \n> it should be able to resume the transfer of a pack file which might have \n> been interrupted in the middle?  If no then this should be considered.\n\nIt can.\n\nMike\n"},{"id":"111046","messageId":"20090411134112.GA1673@atjola.homenet","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904101806340.4583@localhost.localdomain","subject":"Re: [PATCH] process_{tree,blob}: Remove useless xstrdup calls","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-11T13:41:12Z","receivedAt":"2009-04-11T13:41:12Z","isPatch":true,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.10 18:15:26 -0700, Linus Torvalds wrote:\n> It obviously goes on top of my previous patch.\n\nGives some nice results for the \"rev-list --all --objects\" test on the\ngentoo repo says (with the old pack):\n\n     | With my patch | With your patch on top\n-----|---------------|-----------------------\nVSZ  |       1667952 | 1319324\nRSS  |       1388408 | 1126080\ntime |       1:48.99 | 1:42.24\n\nTesting a full repack, it feels slower during the \"Compressing objects\"\npart, but I don't have any hard numbers on that, and maybe I've just\nbeen more patient the last week, when I did the first repack on that\nrepo. I can just tell that it took about 13 minutes for the \"Compressing\nobjects\" part, and 18 minutes in total, on my Core 2 Quad 2.83GHz with\n4G of RAM.\n\nThe new pack is slightly worse than the old one (window=250, --depth=250):\nOld: 759662467\nNew: 759720234\n\nBut that's seems totally negligible, and at least the performance of the\n(stupid) rev-list test is not affected by the different pack layout.\n\nBjörn\n"},{"id":"111047","messageId":"20090411140756.GA15288@atjola.homenet","threadId":"18724","inReplyTo":"20090411134112.GA1673@atjola.homenet","subject":"Re: [PATCH] process_{tree,blob}: Remove useless xstrdup calls","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-11T14:07:56Z","receivedAt":"2009-04-11T14:07:56Z","isPatch":true,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.11 15:41:12 +0200, Björn Steinbrink wrote:\n> On 2009.04.10 18:15:26 -0700, Linus Torvalds wrote:\n> > It obviously goes on top of my previous patch.\n> \n> Gives some nice results for the \"rev-list --all --objects\" test on the\n> gentoo repo says (with the old pack):\n\nAnd for completeness, here are the results for linux-2.6.git\n\n     | With my patch | With your patch on top\n-----|---------------|-----------------------\nVSZ  |        460376 | 407900\nRSS  |        292996 | 239760\ntime |       0:14.28 | 0:14.66\n\nAnd again, the new pack is slightly worse than the old one\n (window=250, --depth=250).\nOld: 240238406\nNew: 240280452\n\nBut again, it's negligible.\n\nBjörn\n"},{"id":"111052","messageId":"grqjo1$at2$1@ger.gmane.org","threadId":"18724","inReplyTo":"20090404220743.GA869@curie-int","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Mark Levedahl","fromEmail":"mlevedahl@gmail.com","sentAt":"2009-04-11T17:24:14Z","receivedAt":"2009-04-11T17:24:14Z","isPatch":false,"sender":{"key":"mdl123@verizon.net","avatar":"https://avatars.githubusercontent.com/u/5302462?v=4"},"body":"Robin H. Johnson wrote:\n\n> Hi,\n> \n> This is a first in my series of mails over the next few days, on issues\n> that we've run into planning a potential migration for Gentoo's\n> repository into Git.\n> \n> Our full repository conversion is large, even after tuning the\n> repacking, the packed repository is between 1.4 and 1.6GiB. As of Feburary\n> 4th, 2009, it contained 4886949 objects. It is not suitable for\n> splitting into submodules either unfortunately - we have a lot of\n> directory moves that would cause submodule bloat.\n> \n> During an initial clone, I see that git-upload-pack invokes\n> pack-objects, despite the ENTIRE repository already being packed - no\n> loose objects whatsoever. git-upload-pack then seems to buffer in\n> memory.\n> \n\nHave you considered using a bundle as part of the initial clone process? The \nidea would be to periodically create a bundle\n\n\tgit bundle create <somename>.bundle [list of refs]\n\nand publish that on your website. A new user would then do\n\n\twget $uri-of-bundle\n\tgit clone <somename>.bundle\n\tcd $somename\n\tgit remote add origin $origin\n\tgit fetch\n\nand they have the current repo. As the bundle is a file, it can be \ndistributed by torrent or other method. The expense of creating the pack in \nthe bundle is paid exactly once when the bundle is created.\n\nMark\n"},{"id":"111055","messageId":"alpine.LFD.2.00.0904111055480.4583@localhost.localdomain","threadId":"18724","inReplyTo":"20090411140756.GA15288@atjola.homenet","subject":"Re: [PATCH] process_{tree,blob}: Remove useless xstrdup calls","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-04-11T18:06:17Z","receivedAt":"2009-04-11T18:06:17Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 11 Apr 2009, Björn Steinbrink wrote:\n> \n> And for completeness, here are the results for linux-2.6.git\n> \n>      | With my patch | With your patch on top\n> -----|---------------|-----------------------\n> VSZ  |        460376 | 407900\n> RSS  |        292996 | 239760\n> time |       0:14.28 | 0:14.66\n\nOk, it uses less memory, but more CPU time. That's reasonable - we \"waste\" \nCPU time on doing the extra free's, and since the memory use isn't a huge \nconstraining factor and cache behavior is bad anyway, it's then actually \nslightly slower.\n\n> And again, the new pack is slightly worse than the old one\n>  (window=250, --depth=250).\n> Old: 240238406\n> New: 240280452\n> \n> But again, it's negligible.\n\nWell, it's sad that it's consistently a bit worse, even if we're talking \njust small small fractions of a percent (looks like 0.02% bigger ;). \n\nAnd I think I can see why. The new code actually does a _better_ job of \nthe resulting list being in \"recency\" order, whereas the old code used to \noutput the root trees all together. Now they're spread out according to \nhow soon they are reached.\n\nThe object sorting code _should_ sort them by type, name and size (and \nthus the pack generation should generate the same deltas), but the name \nhashing is probably weak enough that it doesn't always do a perfect job, \nand then we likely get a slightly worse pack.\n\nBut it would be good to really understand that part. It's a _small_ \ndownside, but it's a downside.\n\nBut it's interesting to note how the bigger gentoo case actually improved \nin performance, probably because by then the denser memory use actually \nmeant that we had noticeably better cache and TLB behavior. So the patch \nhelps the bad case, at least.\n\n\t\t\tLinus\n"},{"id":"111056","messageId":"alpine.LFD.2.00.0904111115210.4583@localhost.localdomain","threadId":"18724","inReplyTo":"20090411140756.GA15288@atjola.homenet","subject":"Re: [PATCH] process_{tree,blob}: Remove useless xstrdup calls","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-04-11T18:19:01Z","receivedAt":"2009-04-11T18:19:01Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 11 Apr 2009, Björn Steinbrink wrote:\n\n> On 2009.04.11 15:41:12 +0200, Björn Steinbrink wrote:\n> > On 2009.04.10 18:15:26 -0700, Linus Torvalds wrote:\n> > > It obviously goes on top of my previous patch.\n> > \n> > Gives some nice results for the \"rev-list --all --objects\" test on the\n> > gentoo repo says (with the old pack):\n> >      | With my patch | With your patch on top\n> > -----|---------------|-----------------------\n> > VSZ  |       1667952 | 1319324\n> > RSS  |       1388408 | 1126080\n> \n> linux-2.6.git:\n> \n>      | With my patch | With your patch on top\n> -----|---------------|-----------------------\n> VSZ  |        460376 | 407900\n> RSS  |        292996 | 239760\n\nInteresting. That's a 18+% reduction in RSS in both cases. Much bigger \nthan I expected, or what I saw in my limited testing. Is this in 32-bit \nmode, where the pointers are cheaper, and thus the non-pointer data \nrelatively more expensive and a bigger percentage of the total? We really \nwasted a _lot_ of memory on those names.\n\n\t\tLinus\n"},{"id":"111057","messageId":"alpine.LFD.2.00.0904111119520.4583@localhost.localdomain","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904111055480.4583@localhost.localdomain","subject":"Re: [PATCH] process_{tree,blob}: Remove useless xstrdup calls","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-04-11T18:22:09Z","receivedAt":"2009-04-11T18:22:09Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 11 Apr 2009, Linus Torvalds wrote:\n> On Sat, 11 Apr 2009, Björn Steinbrink wrote:\n> \n> > And again, the new pack is slightly worse than the old one\n> >  (window=250, --depth=250).\n> > Old: 240238406\n> > New: 240280452\n> > \n> > But again, it's negligible.\n> \n> Well, it's sad that it's consistently a bit worse, even if we're talking \n> just small small fractions of a percent (looks like 0.02% bigger ;). \n\nOh, just wondering: that 0.02% is negligible, but did you use \"-f\" (or \n--no-reuse-delta if you're testing with 'git pack-objects') to see that \nit's actually re-computing the deltas?\n\nThe 0.02% difference might be just because of differences in pack layout. \nIf you force all deltas to be recomputed, maybe the difference is much \nbigger?\n\n\t\tLinus\n"},{"id":"111070","messageId":"20090411192231.GA21300@atjola.homenet","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904111119520.4583@localhost.localdomain","subject":"Re: [PATCH] process_{tree,blob}: Remove useless xstrdup calls","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-11T19:22:31Z","receivedAt":"2009-04-11T19:22:31Z","isPatch":true,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.11 11:22:09 -0700, Linus Torvalds wrote:\n> \n> \n> On Sat, 11 Apr 2009, Linus Torvalds wrote:\n> > On Sat, 11 Apr 2009, Björn Steinbrink wrote:\n> > \n> > > And again, the new pack is slightly worse than the old one\n> > >  (window=250, --depth=250).\n> > > Old: 240238406\n> > > New: 240280452\n> > > \n> > > But again, it's negligible.\n> > \n> > Well, it's sad that it's consistently a bit worse, even if we're talking \n> > just small small fractions of a percent (looks like 0.02% bigger ;). \n> \n> Oh, just wondering: that 0.02% is negligible, but did you use \"-f\" (or \n> --no-reuse-delta if you're testing with 'git pack-objects') to see that \n> it's actually re-computing the deltas?\n> \n> The 0.02% difference might be just because of differences in pack layout. \n> If you force all deltas to be recomputed, maybe the difference is much \n> bigger?\n\nYep, that was \"git repack -adf --window=250 --depth=250\".\n\nBjörn\n"},{"id":"111086","messageId":"20090411194000.GB21300@atjola.homenet","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904111115210.4583@localhost.localdomain","subject":"Re: [PATCH] process_{tree,blob}: Remove useless xstrdup calls","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-11T19:40:00Z","receivedAt":"2009-04-11T19:40:00Z","isPatch":true,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.11 11:19:01 -0700, Linus Torvalds wrote:\n> On Sat, 11 Apr 2009, Björn Steinbrink wrote:\n> > On 2009.04.11 15:41:12 +0200, Björn Steinbrink wrote:\n> > > On 2009.04.10 18:15:26 -0700, Linus Torvalds wrote:\n> > > > It obviously goes on top of my previous patch.\n> > > \n> > > Gives some nice results for the \"rev-list --all --objects\" test on the\n> > > gentoo repo says (with the old pack):\n> > >      | With my patch | With your patch on top\n> > > -----|---------------|-----------------------\n> > > VSZ  |       1667952 | 1319324\n> > > RSS  |       1388408 | 1126080\n> > \n> > linux-2.6.git:\n> > \n> >      | With my patch | With your patch on top\n> > -----|---------------|-----------------------\n> > VSZ  |        460376 | 407900\n> > RSS  |        292996 | 239760\n> \n> Interesting. That's a 18+% reduction in RSS in both cases. Much bigger \n> than I expected, or what I saw in my limited testing. Is this in 32-bit \n> mode, where the pointers are cheaper, and thus the non-pointer data \n> relatively more expensive and a bigger percentage of the total? We really \n> wasted a _lot_ of memory on those names.\n\nNo, this is x86-64, 8 byte pointers. But the savings are trivially\nexplained I think. The struct object_array things are 20 bytes here (per\nobject overhead!), so that's about 5M * 20 = 100M. And the average name\nlength for the objects was 19 bytes, which means about another 100M.\nBoth, the object_array stuff as well as the path names, were allocated\nand never freed. Your patch removed the object_array stuff, and it made\nthe memory allocations for the names temporary. Right?\n\nHad you moved just the path_name() calls, that would have meant that we\nhad needed to keep the name_path stuff around, which is also 20 bytes\nhere (two pointers, one int). And that would have meant that anything\nthat has a leading-up path shorter than 20 bytes (64 bit pointers) would\nhave seen increased memory usage (64bit pointers), but with 32 pointers,\nthe limit would have been 12 bytes.\n\nSo for the \"just move path_name() call\" solution, 32bit vs. 64bit would\nhave made a difference, but with your actual patches, you just turned\neverything into temporary allocations. So the 4byte overhead on 64bit\nplatforms is just once linear with the directory-depth of the current\nobject, instead of with the number of objects in total.\n\nRight?\n\nBjörn\n"},{"id":"111091","messageId":"alpine.LFD.2.00.0904111255530.4583@localhost.localdomain","threadId":"18724","inReplyTo":"20090411194000.GB21300@atjola.homenet","subject":"Re: [PATCH] process_{tree,blob}: Remove useless xstrdup calls","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-04-11T19:58:50Z","receivedAt":"2009-04-11T19:58:50Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 11 Apr 2009, Björn Steinbrink wrote:\n> \n> No, this is x86-64, 8 byte pointers. But the savings are trivially\n> explained I think. The struct object_array things are 20 bytes here (per\n> object overhead!), so that's about 5M * 20 = 100M. And the average name\n> length for the objects was 19 bytes, which means about another 100M.\n> Both, the object_array stuff as well as the path names, were allocated\n> and never freed. Your patch removed the object_array stuff, and it made\n> the memory allocations for the names temporary. Right?\n\nRight.\n\nMy original one-liner patch just did the name freeing part, but it did so \nonly at the _end_ (when actually calling show_object()), so it probably \ndidn't help RSS very much - because you still had one point in time where \nyou had all the names allocated. It probably helped packing (since it \nallocates more _afterwards_), but likely didn't make much of a difference \nfor just 'git rev-list\".\n\nSo that was the impetus for trying to just avoid the \"keep all objects \naround on the 'object_array' thing\" patch, and then cleaning up the \nshow_object() call semantics.\n\n\t\tLinus\n"},{"id":"111096","messageId":"20090411205044.GA21673@atjola.homenet","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904111055480.4583@localhost.localdomain","subject":"Re: [PATCH] process_{tree,blob}: Remove useless xstrdup calls","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-11T20:50:44Z","receivedAt":"2009-04-11T20:50:44Z","isPatch":true,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.11 11:06:17 -0700, Linus Torvalds wrote:\n> \n> \n> On Sat, 11 Apr 2009, Björn Steinbrink wrote:\n> > \n> > And for completeness, here are the results for linux-2.6.git\n> > \n> >      | With my patch | With your patch on top\n> > -----|---------------|-----------------------\n> > VSZ  |        460376 | 407900\n> > RSS  |        292996 | 239760\n> > time |       0:14.28 | 0:14.66\n> \n> Ok, it uses less memory, but more CPU time. That's reasonable - we \"waste\" \n> CPU time on doing the extra free's, and since the memory use isn't a huge \n> constraining factor and cache behavior is bad anyway, it's then actually \n> slightly slower.\n> \n> > And again, the new pack is slightly worse than the old one\n> >  (window=250, --depth=250).\n> > Old: 240238406\n> > New: 240280452\n> > \n> > But again, it's negligible.\n> \n> Well, it's sad that it's consistently a bit worse, even if we're talking \n> just small small fractions of a percent (looks like 0.02% bigger ;). \n> \n> And I think I can see why. The new code actually does a _better_ job of \n> the resulting list being in \"recency\" order, whereas the old code used to \n> output the root trees all together. Now they're spread out according to \n> how soon they are reached.\n\nHm, I don't think that was the case. When iterating over the commits,\nprocess_tree was called with commit->tree, and that added the root tree\nto the objects array as well as walking it to add all referenced objects.\n\nAnd yep, the 'old' \"rev-list --all-objects\" shows for example:\nebace34d059216b3573cd67a83068d2eafe2f2e7 read-cache.c\na869cb0789d8ad87f04d28dd9b703f3ff343a4a7 \n497a05b8fa8e9aa3a5db9b42e5c50392f352d2b4 cache.h\n91b2628e3c18e7f75e477c24197d9ef2eca14125 read-cache.c\n6862d1012681cd6812ab9bfe1a866446f92a7c91 read-tree.c\n\na869cb0 being a root tree, inbetween two blobs.\n\nBjörn\n"},{"id":"111103","messageId":"alpine.LFD.2.00.0904111441240.4583@localhost.localdomain","threadId":"18724","inReplyTo":"20090411205044.GA21673@atjola.homenet","subject":"Re: [PATCH] process_{tree,blob}: Remove useless xstrdup calls","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-04-11T21:43:14Z","receivedAt":"2009-04-11T21:43:14Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 11 Apr 2009, Björn Steinbrink wrote:\n> > \n> > And I think I can see why. The new code actually does a _better_ job of \n> > the resulting list being in \"recency\" order, whereas the old code used to \n> > output the root trees all together. Now they're spread out according to \n> > how soon they are reached.\n> \n> Hm, I don't think that was the case. When iterating over the commits,\n> process_tree was called with commit->tree, and that added the root tree\n> to the objects array as well as walking it to add all referenced objects.\n\nOh, you're right. We actually ended up walking the trees at that point, \nso recency should be the same. \n\nHmm. Where does the difference in ordering come from, then? \n\n\t\t\tLinus\n"},{"id":"111108","messageId":"20090411232431.GA22747@atjola.homenet","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904111441240.4583@localhost.localdomain","subject":"Re: [PATCH] process_{tree,blob}: Remove useless xstrdup calls","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-11T23:24:31Z","receivedAt":"2009-04-11T23:24:31Z","isPatch":true,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.11 14:43:14 -0700, Linus Torvalds wrote:\n> On Sat, 11 Apr 2009, Björn Steinbrink wrote:\n> > > \n> > > And I think I can see why. The new code actually does a _better_ job of \n> > > the resulting list being in \"recency\" order, whereas the old code used to \n> > > output the root trees all together. Now they're spread out according to \n> > > how soon they are reached.\n> > \n> > Hm, I don't think that was the case. When iterating over the commits,\n> > process_tree was called with commit->tree, and that added the root tree\n> > to the objects array as well as walking it to add all referenced objects.\n> \n> Oh, you're right. We actually ended up walking the trees at that point, \n> so recency should be the same. \n> \n> Hmm. Where does the difference in ordering come from, then? \n\nAh! The tag objects. Previously, they were added to the end of the\nobjects array, after all the objects from the process_tree() calls. But\nnow, the pending array is directly processed, causing the tags to show\nup earlier. The same is of course true for any other pending object, but\nverifying that for the tag objects was easier :-)\n\nBjörn\n"},{"id":"111290","messageId":"alpine.DEB.1.00.0904141749330.10279@pacific.mpi-cbg.de","threadId":"18724","inReplyTo":"20090410T203405Z@curie.orbis-terrarum.net","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-04-14T15:52:53Z","receivedAt":"2009-04-14T15:52:53Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 10 Apr 2009, Robin H. Johnson wrote:\n\n> On Wed, Apr 08, 2009 at 12:52:54AM -0400, Nicolas Pitre wrote:\n> > > http://git.overlays.gentoo.org/gitweb/?p=exp/gentoo-x86.git;a=summary\n> > > At least that's what I cloned ;-) I hope it's the right one, but it fits\n> > > the description...\n> > OK.  FWIW, I repacked it with --window=250 --depth=250 and obtained a \n> > 725MB pack file.  So that's about half the originally reported size.\n> The one problem with having the single large packfile is that Git\n> doesn't have a trivial way to resume downloading it when the git://\n> protocol is used.\n> \n> For our developers cursed with bad internet connections (a fair number\n> of firewalls that don't seem to respect keepalive properly), I suppose\n> I can probably just maintain a separate repo for their initial clones,\n> which leaves a large overall download, but more chances to resume.\n\nIMO the best we could do under these circumstances is to use fsck \n--lost-found to find those commits which have a complete history (i.e. no \n\"broken links\") -- this probably needs to be implemented as a special mode \nof --lost-found -- and store them in a temporary to-be-removed \nnamespace, say refs/heads/incomplete-refs/$number, which will be sent to \nthe server when fetching the next time.  (Might need some iterations to \nget everything, though.)\n\nCiao,\nDscho\n"},{"id":"111312","messageId":"alpine.LFD.2.00.0904141542161.6741@xanadu.home","threadId":"18724","inReplyTo":"alpine.DEB.1.00.0904141749330.10279@pacific.mpi-cbg.de","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-14T20:17:55Z","receivedAt":"2009-04-14T20:17:55Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 14 Apr 2009, Johannes Schindelin wrote:\n\n> Hi,\n> \n> On Fri, 10 Apr 2009, Robin H. Johnson wrote:\n> \n> > On Wed, Apr 08, 2009 at 12:52:54AM -0400, Nicolas Pitre wrote:\n> > > > http://git.overlays.gentoo.org/gitweb/?p=exp/gentoo-x86.git;a=summary\n> > > > At least that's what I cloned ;-) I hope it's the right one, but it fits\n> > > > the description...\n> > > OK.  FWIW, I repacked it with --window=250 --depth=250 and obtained a \n> > > 725MB pack file.  So that's about half the originally reported size.\n> > The one problem with having the single large packfile is that Git\n> > doesn't have a trivial way to resume downloading it when the git://\n> > protocol is used.\n> > \n> > For our developers cursed with bad internet connections (a fair number\n> > of firewalls that don't seem to respect keepalive properly), I suppose\n> > I can probably just maintain a separate repo for their initial clones,\n> > which leaves a large overall download, but more chances to resume.\n> \n> IMO the best we could do under these circumstances is to use fsck \n> --lost-found to find those commits which have a complete history (i.e. no \n> \"broken links\") -- this probably needs to be implemented as a special mode \n> of --lost-found -- and store them in a temporary to-be-removed \n> namespace, say refs/heads/incomplete-refs/$number, which will be sent to \n> the server when fetching the next time.  (Might need some iterations to \n> get everything, though.)\n\nWell, although this might seem a good idea, this would help only in \nthose cases where there is at least one complete revision available, \ni.e. no delta needed. This is usually true for the top commit after a \nrepack which objects are all stored at the front of the pack and serve \nas base objects for deltas from subsequent (older) commits.  Thing is, \nthat first revision is likely to occupy a significant portion of the \nwhole pack, like no less than the size of the equivalent .tar.gz for the \ncontent of that commit.  To see what this represents, just try a shallow \nclone with depth=1.  For the Linux kernel, this is more than 80MB while \nthe whole repo is in the 200MB range.  So if your connection isn't \nreliable enough to transfer at least that amount then you're screwed \nanyway.\n\nIndependently from this, I think there is quite a lot of confusion here.  \nAccording to Robin, the reason for splitting the large Gentoo repo into \nmultiple packs is apparently to help with the resuming of a clone.  We \nknow that the git:// protocol is currently not resumable, and having \nmultiple packs on the remote server won't change the outcome in any way \nas the client still receives a single big pack anyway.\n\nWRT the HTTP protocol, I was questioning git's ability to resume the \ntransfer of a pack in the middle if such transfer is interrupted without \nredownloading it all. And Mike Hommey says this is actually the case.\n\nMeaning there is simply no reason to split a big pack into multiple \nones.  If anything, it'll only make a clone over the native git protocol \nmore costly for the server which has to pack everything back together.\n\n\nNicolas\n"},{"id":"111316","messageId":"20090414T202206Z@curie.orbis-terrarum.net","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904141542161.6741@xanadu.home","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Robin H. Johnson","fromEmail":"robbat2@gentoo.org","sentAt":"2009-04-14T20:27:51Z","receivedAt":"2009-04-14T20:27:51Z","isPatch":false,"sender":{"key":"robbat2@gentoo.org","avatar":"https://avatars.githubusercontent.com/u/373898?v=4"},"body":"On Tue, Apr 14, 2009 at 04:17:55PM -0400, Nicolas Pitre wrote:\n> WRT the HTTP protocol, I was questioning git's ability to resume the \n> transfer of a pack in the middle if such transfer is interrupted without \n> redownloading it all. And Mike Hommey says this is actually the case.\nWith rsync:// it was helpful to split the pack, and resume there worked\nreasonably (see my other mail about the segfault that turns up\nsometimes).\n\nMore recent discussions raised the possibility of using git-bundle to\nprovide a more ideal initial download that they CAN resume easily, as\nwell as being able to move on from it.\n\nSo, from the Gentoo side right now, we're looking at this:\n1. Setup git-bundle for initial downloads.\n2. Disallow initial clones over git:// (allow updates ONLY)\n3. Disallow git-over-http, git-over-rsync.\n\nThis also avoids the wait time with the initial clone. Just grab the\nbundle with your choice of rsync or http, check it's integrity, throw it\ninto your repo, and update to the latest tree.\n\n-- \nRobin Hugh Johnson\nGentoo Linux Developer & Infra Guy\nE-Mail     : robbat2@gentoo.org\nGnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n"},{"id":"111315","messageId":"alpine.DEB.1.00.0904142229280.10279@pacific.mpi-cbg.de","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904141542161.6741@xanadu.home","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-04-14T20:30:24Z","receivedAt":"2009-04-14T20:30:24Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 14 Apr 2009, Nicolas Pitre wrote:\n\n> On Tue, 14 Apr 2009, Johannes Schindelin wrote:\n> \n> > IMO the best we could do under these circumstances [unreliable \n> > network] is to use fsck --lost-found to find those commits which have \n> > a complete history (i.e. no \"broken links\") -- this probably needs to \n> > be implemented as a special mode of --lost-found -- and store them in \n> > a temporary to-be-removed namespace, say \n> > refs/heads/incomplete-refs/$number, which will be sent to the server \n> > when fetching the next time.  (Might need some iterations to get \n> > everything, though.)\n> \n> Well, although this might seem a good idea, this would help only in \n> those cases where there is at least one complete revision available, \n> i.e. no delta needed. This is usually true for the top commit after a \n> repack which objects are all stored at the front of the pack and serve \n> as base objects for deltas from subsequent (older) commits.  Thing is, \n> that first revision is likely to occupy a significant portion of the \n> whole pack, like no less than the size of the equivalent .tar.gz for the \n> content of that commit.  To see what this represents, just try a shallow \n> clone with depth=1.  For the Linux kernel, this is more than 80MB while \n> the whole repo is in the 200MB range.  So if your connection isn't \n> reliable enough to transfer at least that amount then you're screwed \n> anyway.\n\nGood point.\n\nSorry for not thinking it through,\nDscho\n"},{"id":"111319","messageId":"alpine.LFD.2.00.0904141648370.6741@xanadu.home","threadId":"18724","inReplyTo":"20090414T202206Z@curie.orbis-terrarum.net","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-14T21:02:24Z","receivedAt":"2009-04-14T21:02:24Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 14 Apr 2009, Robin H. Johnson wrote:\n\n> More recent discussions raised the possibility of using git-bundle to\n> provide a more ideal initial download that they CAN resume easily, as\n> well as being able to move on from it.\n> \n> So, from the Gentoo side right now, we're looking at this:\n> 1. Setup git-bundle for initial downloads.\n> 2. Disallow initial clones over git:// (allow updates ONLY)\n> 3. Disallow git-over-http, git-over-rsync.\n> \n> This also avoids the wait time with the initial clone. Just grab the\n> bundle with your choice of rsync or http, check it's integrity, throw it\n> into your repo, and update to the latest tree.\n\nThis certainly makes lots of sense until we overcome the current clone \nbothleneck.  You should tightly repack your repository first, like with\n\"git repack -a -f -d --depth=100 --window=500\".  Use a fast machine with \nenough ram of course.  Then you'll have a nice and small bundle.\n\nOf course any git pack/bundle has full self-integrity built in.  So you \nshould not need to do a separate check.\n\nAnd don't forget to delete the bundle once it has been fetched into a \nfull repository, otherwise it'll only wastes disk space.\n\n\nNicolas\n"},{"id":"111334","messageId":"fcaeb9bf0904142009w5a21e483v7e98f91e5e35b14a@mail.gmail.com","threadId":"18724","inReplyTo":"20090414T202206Z@curie.orbis-terrarum.net","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-04-15T03:09:43Z","receivedAt":"2009-04-15T03:09:43Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Apr 15, 2009 at 6:27 AM, Robin H. Johnson <robbat2@gentoo.org> wrote:\n> So, from the Gentoo side right now, we're looking at this:\n> 1. Setup git-bundle for initial downloads.\n> 2. Disallow initial clones over git:// (allow updates ONLY)\n\nHow can you do that? If I understand git protocol correctly, there is\nno difference between a fetch request and a clone one.\n\n> 3. Disallow git-over-http, git-over-rsync.\n-- \nDuy\n"},{"id":"111336","messageId":"20090415T044931Z@curie.orbis-terrarum.net","threadId":"18724","inReplyTo":"fcaeb9bf0904142009w5a21e483v7e98f91e5e35b14a@mail.gmail.com","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Robin H. Johnson","fromEmail":"robbat2@gentoo.org","sentAt":"2009-04-15T05:53:17Z","receivedAt":"2009-04-15T05:53:17Z","isPatch":false,"sender":{"key":"robbat2@gentoo.org","avatar":"https://avatars.githubusercontent.com/u/373898?v=4"},"body":"On Wed, Apr 15, 2009 at 01:09:43PM +1000, Nguyen Thai Ngoc Duy wrote:\n> On Wed, Apr 15, 2009 at 6:27 AM, Robin H. Johnson <robbat2@gentoo.org> wrote:\n> > So, from the Gentoo side right now, we're looking at this:\n> > 1. Setup git-bundle for initial downloads.\n> > 2. Disallow initial clones over git:// (allow updates ONLY)\n> How can you do that? If I understand git protocol correctly, there is\n> no difference between a fetch request and a clone one.\nI'm planning on adding a new hook, in upload-pack.\nInputs: want_obj, have_obj\n\nNot sure of the best way to pass them yet, probably stdin, 'want ....',\n'have ....'.\n\nProbably best to run right before git-rev-list.\n\nFor the Gentoo-specific content of the hook, I'm after this design:\n- you don't send ANY have => you get the error\n- you have is too old => you get the error\n- you ask for something non-existent => you get the error\n\nThe error will be a message instructing you to use the bundle, and\npointing to a URL with detailed instructions.\n\nThe 'too old' case is to able better DoS prevention, stopping somebody\nmalicious from finding the first commit in the bundle, and pretending\nthey have it, asking for a pack from that to the HEAD.\n\n-- \nRobin Hugh Johnson\nGentoo Linux Developer & Infra Guy\nE-Mail     : robbat2@gentoo.org\nGnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n"},{"id":"111337","messageId":"7vljq2zckw.fsf@gitster.siamese.dyndns.org","threadId":"18724","inReplyTo":"fcaeb9bf0904142009w5a21e483v7e98f91e5e35b14a@mail.gmail.com","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-04-15T05:54:23Z","receivedAt":"2009-04-15T05:54:23Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyen Thai Ngoc Duy <pclouds@gmail.com> writes:\n\n>> 2. Disallow initial clones over git:// (allow updates ONLY)\n>\n> How can you do that? If I understand git protocol correctly, there is\n> no difference between a fetch request and a clone one.\n\nAt the protocol level, you can tell a clone request by noticing that the\ndownloading side does not have any \"have\" lines, but it is a different\nmatter what the software does out of the box.\n\nYou can patch upload-pack to reject such requests.  I am sure gentoo folks\nare capable of doing that ;-)\n\nAlso a rogue client can send a bogus \"have\" to fool that logic, and that\nis the primary reason why we do not have such a patch to upload-pack.  It\nis not worth it as a protection against determined people who want to DoS.\n"},{"id":"111354","messageId":"alpine.LFD.2.00.0904150738340.6741@xanadu.home","threadId":"18724","inReplyTo":"7vljq2zckw.fsf@gitster.siamese.dyndns.org","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-15T11:51:22Z","receivedAt":"2009-04-15T11:51:22Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 14 Apr 2009, Junio C Hamano wrote:\n\n> Nguyen Thai Ngoc Duy <pclouds@gmail.com> writes:\n> \n> > How can you do that? If I understand git protocol correctly, there is\n> > no difference between a fetch request and a clone one.\n> \n> At the protocol level, you can tell a clone request by noticing that the\n> downloading side does not have any \"have\" lines, but it is a different\n> matter what the software does out of the box.\n> \n> You can patch upload-pack to reject such requests.  I am sure gentoo folks\n> are capable of doing that ;-)\n> \n> Also a rogue client can send a bogus \"have\" to fool that logic, and that\n> is the primary reason why we do not have such a patch to upload-pack.  It\n> is not worth it as a protection against determined people who want to DoS.\n\nImplementing a minimum treshold with merge-base to ensure that the \nclient has at least commit X should be easy to do.  Unfortunately we \ndon't have any hook for such a purpose yet.\n\n\nNicolas\n"},{"id":"111913","messageId":"1240362948.22240.76.camel@maia.lan","threadId":"18724","inReplyTo":"20090414T202206Z@curie.orbis-terrarum.net","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2009-04-22T01:15:48Z","receivedAt":"2009-04-22T01:15:48Z","isPatch":false,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"On Tue, 2009-04-14 at 13:27 -0700, Robin H. Johnson wrote:\n> On Tue, Apr 14, 2009 at 04:17:55PM -0400, Nicolas Pitre wrote:\n> > WRT the HTTP protocol, I was questioning git's ability to resume the \n> > transfer of a pack in the middle if such transfer is interrupted without \n> > redownloading it all. And Mike Hommey says this is actually the case.\n> With rsync:// it was helpful to split the pack, and resume there worked\n> reasonably (see my other mail about the segfault that turns up\n> sometimes).\n> \n> More recent discussions raised the possibility of using git-bundle to\n> provide a more ideal initial download that they CAN resume easily, as\n> well as being able to move on from it.\n\nHey Robin,\n\nNow that the GSoC projects have been announced I can give you the good\nnews that one of our two projects is to optimise this stage in\ngit-daemon; I'm hoping we can get it down to being almost as cheap as\nthe workaround you described in your post.  I'll certainly be using your\nrepository as a test case :-)\n\nSo stay tuned!\nSam.\n"},{"id":"111933","messageId":"e2b179460904220255v58986bd5q7c22eb3ab8486157@mail.gmail.com","threadId":"18724","inReplyTo":"1240362948.22240.76.camel@maia.lan","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Mike Ralphson","fromEmail":"mike.ralphson@gmail.com","sentAt":"2009-04-22T09:55:45Z","receivedAt":"2009-04-22T09:55:45Z","isPatch":false,"sender":{"key":"mike.ralphson@gmail.com","avatar":"https://avatars.githubusercontent.com/u/21603?v=4"},"body":"2009/4/22 Sam Vilain <sam@vilain.net>\n> Now that the GSoC projects have been announced I can give you the good\n> news that one of our two projects...\n\nIt's sort of three, really...\n\nhttp://socghop.appspot.com/student_project/show/google/gsoc2009/mono/t124022708105\n\nMike\n"},{"id":"111941","messageId":"B2EB379C-72D6-4318-8873-E82B799F0437@ai.rug.nl","threadId":"18724","inReplyTo":"e2b179460904220255v58986bd5q7c22eb3ab8486157@mail.gmail.com","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Pieter de Bie","fromEmail":"pdebie@ai.rug.nl","sentAt":"2009-04-22T11:24:21Z","receivedAt":"2009-04-22T11:24:21Z","isPatch":false,"sender":{"key":"pdebie@ai.rug.nl","avatar":null},"body":"\nOn 22 apr 2009, at 10:55, Mike Ralphson wrote:\n\n> It's sort of three, really...\n>\n> http://socghop.appspot.com/student_project/show/google/gsoc2009/mono/t124022708105\n\nThat same project was also done by two(!) students last\nyear, but I don't think that worked out. I wonder how it'll\nplay out this year.\n\n- Pieter\n"},{"id":"111944","messageId":"alpine.DEB.1.00.0904221516250.14221@intel-tinevez-2-302","threadId":"18724","inReplyTo":"e2b179460904220255v58986bd5q7c22eb3ab8486157@mail.gmail.com","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-04-22T13:19:32Z","receivedAt":"2009-04-22T13:19:32Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 22 Apr 2009, Mike Ralphson wrote:\n\n> 2009/4/22 Sam Vilain <sam@vilain.net>\n> > Now that the GSoC projects have been announced I can give you the good\n> > news that one of our two projects...\n> \n> It's sort of three, really...\n> \n> http://socghop.appspot.com/student_project/show/google/gsoc2009/mono/t124022708105\n\nOMG!  That's the third time they are wasting Google's money: AFAICT they \nhaven't learnt from the past two years' failures.  At least I am not aware \nof any of them Mono guys trying to collaborate with us.\n\nOh well, maybe I should drop them a mail that they may get valuable input \nhere _iff_ they just ask.\n\nCiao,\nDscho\n"},{"id":"111947","messageId":"alpine.LFD.2.00.0904221011340.6741@xanadu.home","threadId":"18724","inReplyTo":"1240362948.22240.76.camel@maia.lan","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-22T14:14:48Z","receivedAt":"2009-04-22T14:14:48Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 22 Apr 2009, Sam Vilain wrote:\n\n> On Tue, 2009-04-14 at 13:27 -0700, Robin H. Johnson wrote:\n> > On Tue, Apr 14, 2009 at 04:17:55PM -0400, Nicolas Pitre wrote:\n> > > WRT the HTTP protocol, I was questioning git's ability to resume the \n> > > transfer of a pack in the middle if such transfer is interrupted without \n> > > redownloading it all. And Mike Hommey says this is actually the case.\n> > With rsync:// it was helpful to split the pack, and resume there worked\n> > reasonably (see my other mail about the segfault that turns up\n> > sometimes).\n> > \n> > More recent discussions raised the possibility of using git-bundle to\n> > provide a more ideal initial download that they CAN resume easily, as\n> > well as being able to move on from it.\n> \n> Hey Robin,\n> \n> Now that the GSoC projects have been announced I can give you the good\n> news that one of our two projects is to optimise this stage in\n> git-daemon; I'm hoping we can get it down to being almost as cheap as\n> the workaround you described in your post.  I'll certainly be using your\n> repository as a test case :-)\n\nPlease keep me in the loop as much as possible.  I'd prefer we're not in \ndisagreement over the implementation only after final patches are posted \nto the list.\n\n\nNicolas\n"},{"id":"111948","messageId":"20090422143503.GG23604@spearce.org","threadId":"18724","inReplyTo":"alpine.DEB.1.00.0904221516250.14221@intel-tinevez-2-302","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2009-04-22T14:35:03Z","receivedAt":"2009-04-22T14:35:03Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> On Wed, 22 Apr 2009, Mike Ralphson wrote:\n> \n> > 2009/4/22 Sam Vilain <sam@vilain.net>\n> > > Now that the GSoC projects have been announced I can give you the good\n> > > news that one of our two projects...\n> > \n> > It's sort of three, really...\n> > \n> > http://socghop.appspot.com/student_project/show/google/gsoc2009/mono/t124022708105\n> \n> OMG!  That's the third time they are wasting Google's money: AFAICT they \n> haven't learnt from the past two years' failures.  At least I am not aware \n> of any of them Mono guys trying to collaborate with us.\n\nYikes!\n\nWearing my Google hat, I have to cry a little.  I think its such\na waste.  But we don't tell the orgs what projects they should or\nshould not do, its at each org's individual discretion.  Clearly the\nMono folks feel they should \"spend\" a *fourth* slot on this project.\nThat or, Mono was granted one too many slots in the program.  *sigh*\n\nWearing my JGit maintainer hat, I have to cry a little.  The mentor\nfor this project should realize... we've spent over 3 years now\non JGit (it turned 3 on Mar 6 2009) and it *still* doesn't provide\na full replacement for git.git.\n\nI'd like to think that I'm not a moron, and that it really does\ntake 3 years of R&D work to find a suitable implementation of Git\nin a sandboxed language like Java.  Or, maybe I am a moron.  Linus,\nJunio and crew had git.git implemented in less time.\n \n> Oh well, maybe I should drop them a mail that they may get valuable input \n> here _iff_ they just ask.\n\nI've tried that in the past two years.  I've given up.  The JGit\ncode is available.  Its license is quite liberal.  They can look\nat it if they want.  My guess is, they won't.\n\n-- \nShawn.\n"},{"id":"111964","messageId":"49EF4867.8060002@op5.se","threadId":"18724","inReplyTo":"20090422143503.GG23604@spearce.org","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2009-04-22T16:40:07Z","receivedAt":"2009-04-22T16:40:07Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Shawn O. Pearce wrote:\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n>> On Wed, 22 Apr 2009, Mike Ralphson wrote:\n>>\n>>> 2009/4/22 Sam Vilain <sam@vilain.net>\n>>>> Now that the GSoC projects have been announced I can give you the good\n>>>> news that one of our two projects...\n>>> It's sort of three, really...\n>>>\n>>> http://socghop.appspot.com/student_project/show/google/gsoc2009/mono/t124022708105\n>> OMG!  That's the third time they are wasting Google's money: AFAICT they \n>> haven't learnt from the past two years' failures.  At least I am not aware \n>> of any of them Mono guys trying to collaborate with us.\n> \n\nI offered to assist with reviewing patches or explaining technicalia to them\nlast year (while I was learning a bit of C# myself), but got no patches or\nrequests from them at all.\n\n>> Oh well, maybe I should drop them a mail that they may get valuable input \n>> here _iff_ they just ask.\n> \n> I've tried that in the past two years.  I've given up.  The JGit\n> code is available.  Its license is quite liberal.  They can look\n> at it if they want.  My guess is, they won't.\n> \n\nI'm with Shawn here. They refuse to look at unmanaged code (that is, non-C#\ncode), and since there is none yet, they're in a sort of catch-22 when it\ncomes to reference implementations. Ah well. I'll join mono-develop mailing\nlist again and see what I can do to help.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n\nConsidering the successes of the wars on alcohol, poverty, drugs and\nterror, I think we should give some serious thought to declaring war\non peace.\n"},{"id":"111967","messageId":"alpine.DEB.1.00.0904221905560.7282@intel-tinevez-2-302","threadId":"18724","inReplyTo":"49EF4867.8060002@op5.se","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-04-22T17:06:27Z","receivedAt":"2009-04-22T17:06:27Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 22 Apr 2009, Andreas Ericsson wrote:\n\n> Ah well. I'll join mono-develop mailing list again and see what I can do \n> to help.\n\nThanks.  I think these guys are in serious need of help, not only in terms \nof Git, but also in managing GSoC.\n\nCiao,\nDscho\n"},{"id":"112021","messageId":"49EF93CA.20207@vilain.net","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904221011340.6741@xanadu.home","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2009-04-22T22:01:46Z","receivedAt":"2009-04-22T22:01:46Z","isPatch":false,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"Nicolas Pitre wrote:\n>> Now that the GSoC projects have been announced I can give you the good\n>> news that one of our two projects is to optimise this stage in\n>> git-daemon; I'm hoping we can get it down to being almost as cheap as\n>> the workaround you described in your post. I'll certainly be using your\n>> repository as a test case :-)\n>\n> Please keep me in the loop as much as possible. I'd prefer we're not in\n> disagreement over the implementation only after final patches are posted\n> to the list.\n\nThanks Nico, given your close working knowledge of the pack-objects\ncode this will be very much appreciated. Perhaps you can first help\nout by telling me what you have to say about moving object enumeration\nfrom upload-pack to pack-objects?\n\nCheers!\nSam.\n"},{"id":"112025","messageId":"20090422225019.GA17039@atjola.homenet","threadId":"18724","inReplyTo":"49EF93CA.20207@vilain.net","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-04-22T22:50:19Z","receivedAt":"2009-04-22T22:50:19Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.04.23 10:01:46 +1200, Sam Vilain wrote:\n> Nicolas Pitre wrote:\n> >> Now that the GSoC projects have been announced I can give you the good\n> >> news that one of our two projects is to optimise this stage in\n> >> git-daemon; I'm hoping we can get it down to being almost as cheap as\n> >> the workaround you described in your post. I'll certainly be using your\n> >> repository as a test case :-)\n> >\n> > Please keep me in the loop as much as possible. I'd prefer we're not in\n> > disagreement over the implementation only after final patches are posted\n> > to the list.\n> \n> Thanks Nico, given your close working knowledge of the pack-objects\n> code this will be very much appreciated. Perhaps you can first help\n> out by telling me what you have to say about moving object enumeration\n> from upload-pack to pack-objects?\n\nHere's a bit about that:\nhttp://article.gmane.org/gmane.comp.version-control.git/116032\n\nNote that my RSS measurement should be invalid by now. Linus's\npatches(*) should have improved the memory usage for that scenario by\nquite a bit, since we used to keep a lot of the stuff that the revision\nenumeration required in memory, even after that processed finished,\nwhich should no longer be the case IIRC. And the peak memory usage for\nthat process was also improved on its own, as the whole buffering is\ngone.\n\nBjörn\n\n(*) These commits:\n8d2dfc49b1     process_{tree,blob}: show objects without buffering\ncf2ab916af     show_object(): push path_name() call further down\n213152688c     process_{tree,blob}: Remove useless xstrdup calls\n"},{"id":"112028","messageId":"alpine.LFD.2.00.0904221858370.6741@xanadu.home","threadId":"18724","inReplyTo":"49EF93CA.20207@vilain.net","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-22T23:07:42Z","receivedAt":"2009-04-22T23:07:42Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 23 Apr 2009, Sam Vilain wrote:\n\n> Nicolas Pitre wrote:\n> >> Now that the GSoC projects have been announced I can give you the good\n> >> news that one of our two projects is to optimise this stage in\n> >> git-daemon; I'm hoping we can get it down to being almost as cheap as\n> >> the workaround you described in your post. I'll certainly be using your\n> >> repository as a test case :-)\n> >\n> > Please keep me in the loop as much as possible. I'd prefer we're not in\n> > disagreement over the implementation only after final patches are posted\n> > to the list.\n> \n> Thanks Nico, given your close working knowledge of the pack-objects\n> code this will be very much appreciated. Perhaps you can first help\n> out by telling me what you have to say about moving object enumeration\n> from upload-pack to pack-objects?\n\nIt is like a 25-line patch or so.  I did it once, although the shalow \nclone support was missing from it.  And somehow I managed to lose the \npatch while doing some reshuffling of unrelated bigger changes.\n\nBasically, you can pass the revision arguments to pack-objects directly \ninstead of passing them to rev-list and piping rev-list's output to \npack-objects.\n\nNicolas\n"},{"id":"112030","messageId":"alpine.DEB.1.00.0904230129290.10279@pacific.mpi-cbg.de","threadId":"18724","inReplyTo":"alpine.LFD.2.00.0904221858370.6741@xanadu.home","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-04-22T23:30:29Z","receivedAt":"2009-04-22T23:30:29Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 22 Apr 2009, Nicolas Pitre wrote:\n\n> On Thu, 23 Apr 2009, Sam Vilain wrote:\n> \n> > Nicolas Pitre wrote:\n> > >> Now that the GSoC projects have been announced I can give you the good\n> > >> news that one of our two projects is to optimise this stage in\n> > >> git-daemon; I'm hoping we can get it down to being almost as cheap as\n> > >> the workaround you described in your post. I'll certainly be using your\n> > >> repository as a test case :-)\n> > >\n> > > Please keep me in the loop as much as possible. I'd prefer we're not in\n> > > disagreement over the implementation only after final patches are posted\n> > > to the list.\n> > \n> > Thanks Nico, given your close working knowledge of the pack-objects\n> > code this will be very much appreciated. Perhaps you can first help\n> > out by telling me what you have to say about moving object enumeration\n> > from upload-pack to pack-objects?\n> \n> It is like a 25-line patch or so.  I did it once, although the shalow \n> clone support was missing from it.  And somehow I managed to lose the \n> patch while doing some reshuffling of unrelated bigger changes.\n> \n> Basically, you can pass the revision arguments to pack-objects directly \n> instead of passing them to rev-list and piping rev-list's output to \n> pack-objects.\n\nI seem to remember that somebody sent a patch within the last two weeks \nimplementing that, and if my memory does not fail me, in response to one \nof your mails mentioning this wish.\n\nCiao,\nDscho\n"},{"id":"112032","messageId":"alpine.LFD.2.00.0904222312130.6741@xanadu.home","threadId":"18724","inReplyTo":"alpine.DEB.1.00.0904230129290.10279@pacific.mpi-cbg.de","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-04-23T03:16:20Z","receivedAt":"2009-04-23T03:16:20Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 23 Apr 2009, Johannes Schindelin wrote:\n\n> Hi,\n> \n> On Wed, 22 Apr 2009, Nicolas Pitre wrote:\n> \n> > On Thu, 23 Apr 2009, Sam Vilain wrote:\n> > \n> > > Perhaps you can first help out by telling me what you have to say \n> > > about moving object enumeration from upload-pack to pack-objects?\n> > \n> > It is like a 25-line patch or so.  I did it once, although the shalow \n> > clone support was missing from it.  And somehow I managed to lose the \n> > patch while doing some reshuffling of unrelated bigger changes.\n> > \n> > Basically, you can pass the revision arguments to pack-objects directly \n> > instead of passing them to rev-list and piping rev-list's output to \n> > pack-objects.\n> \n> I seem to remember that somebody sent a patch within the last two weeks \n> implementing that, and if my memory does not fail me, in response to one \n> of your mails mentioning this wish.\n\nWell, if so I wasn't CC'd on the post, and my periodic scan of the git \nlist missed it as well.\n\n\nNicolas\n"},{"id":"112119","messageId":"200904232130.59010.chriscool@tuxfamily.org","threadId":"18724","inReplyTo":"e2b179460904220255v58986bd5q7c22eb3ab8486157@mail.gmail.com","subject":"Re: Performance issue: initial git clone causes massive repack","fromName":"Christian Couder","fromEmail":"chriscool@tuxfamily.org","sentAt":"2009-04-23T19:30:58Z","receivedAt":"2009-04-23T19:30:58Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"Le mercredi 22 avril 2009, Mike Ralphson a écrit :\n> 2009/4/22 Sam Vilain <sam@vilain.net>\n>\n> > Now that the GSoC projects have been announced I can give you the good\n> > news that one of our two projects...\n>\n> It's sort of three, really...\n>\n> http://socghop.appspot.com/student_project/show/google/gsoc2009/mono/t124\n>022708105\n\nThere is also this one:\n\nhttp://socghop.appspot.com/student_project/show/google/gsoc2009/hg/t124022472367\n\nRegards,\nChristian.\n"}]}