{"thread":{"id":"66083","subject":"Clones with fetch.bundleURI slower than standard, full clone?","startedAt":"2026-07-28T23:42:07Z","lastAt":"2026-07-29T00:41:11Z","messageCount":3,"participants":["Knop, Ryszard","brian m. carlson"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"549181","messageId":"008c6f4742d8e20124ed21d191178ce6db29aaa5.camel@intel.com","threadId":"66083","inReplyTo":null,"subject":"Clones with fetch.bundleURI slower than standard, full clone?","fromName":"Knop, Ryszard","fromEmail":"ryszard.knop@intel.com","sentAt":"2026-07-28T23:42:03Z","receivedAt":"2026-07-28T23:42:07Z","isPatch":false,"body":"Hey all,\n\nI'm working on a Linux kernel-related CI system where we need to\nperform full clones of the kernel repo in most jobs. Because of some\njobs in the pipeline, it usually cannot be a shallow clone :( Since the\nkernel repo is large and slow to clone, and I don't want to put undue\nstress on the remote host, I used git bundles, where CI clones the repo\nover a weekend, packages that as a bundle, then in jobs it gets used\nlike this (weird, but works for <REF> being a branch, tag or a specific\ncommit hash):\n\ngit init\ngit remote add origin <REPO-URL>\ngit config set fetch.bundleURI <BUNDLE-URL>\ngit fetch origin <REF>\ngit reset --hard FETCH_HEAD\n\nCloning a repo this way takes 4-5mins. Unpacking a bundle appears to be\nsuper slow. Not even faster than just running a full, normal clone from\nthe remote server, actually (~3-4mins for a single branch).\n\nOn one of the build VMs, with Git 2.53 (stock Ubuntu 26.04), GIT_TRACE\nsuggests most of the time is spent in some variation of `/usr/lib/git-\ncore/git index-pack --stdin -v --fix-thin '--keep=fetch-pack 39430 on\nbuild-server' --check-self-contained-and-connected`, and indeed that\nprocess burns 100% of its single thread for most of that time.\n\nIs it expected that doing it this way is so slow? The alternative is to\njust package and work with the whole bare repo, but bundles appear to\nbe an elegant way of dealing with exactly this scenario.\n\nThanks, Ryszard\n"},{"id":"549182","messageId":"amlF-ZepjtCZz1YE@fruit.crustytoothpaste.net","threadId":"66083","inReplyTo":"008c6f4742d8e20124ed21d191178ce6db29aaa5.camel@intel.com","subject":"Re: Clones with fetch.bundleURI slower than standard, full clone?","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-07-29T00:14:50Z","receivedAt":"2026-07-29T00:14:52Z","isPatch":false,"body":"On 2026-07-28 at 23:42:03, Knop, Ryszard wrote:\n> Hey all,\n> \n> I'm working on a Linux kernel-related CI system where we need to\n> perform full clones of the kernel repo in most jobs. Because of some\n> jobs in the pipeline, it usually cannot be a shallow clone :( Since the\n> kernel repo is large and slow to clone, and I don't want to put undue\n> stress on the remote host, I used git bundles, where CI clones the repo\n> over a weekend, packages that as a bundle, then in jobs it gets used\n> like this (weird, but works for <REF> being a branch, tag or a specific\n> commit hash):\n> \n> git init\n> git remote add origin <REPO-URL>\n> git config set fetch.bundleURI <BUNDLE-URL>\n> git fetch origin <REF>\n> git reset --hard FETCH_HEAD\n> \n> Cloning a repo this way takes 4-5mins. Unpacking a bundle appears to be\n> super slow. Not even faster than just running a full, normal clone from\n> the remote server, actually (~3-4mins for a single branch).\n\nYou're comparing apples to oranges here.  A clone of a single branch\nincludes only that line of history and only those objects, but when you\nuse a bundle with multiple refs, Git has to handle all of the objects in\nthe bundle's entire pack, not just the ref you've specified.  In order\nto compare adequately, you'd have to compare a bundle containing only\nthat one ref with the single-branch clone or a regular clone of the full\nrepository with your full bundles.\n\n> On one of the build VMs, with Git 2.53 (stock Ubuntu 26.04), GIT_TRACE\n> suggests most of the time is spent in some variation of `/usr/lib/git-\n> core/git index-pack --stdin -v --fix-thin '--keep=fetch-pack 39430 on\n> build-server' --check-self-contained-and-connected`, and indeed that\n> process burns 100% of its single thread for most of that time.\n\nAs far as I can tell, the unbundling code just calls index-pack, so it\nshould honour pack.threads.  Setting that value to 0 causes this code to\nbe executed:\n\n\t\t/*\n\t\t * Experiments show that going above 20 threads doesn't help,\n\t\t * no matter how many cores you have. Below that, we tend to\n\t\t * max at half the number of online_cpus(), presumably because\n\t\t * half of those are hyperthreads rather than full cores. We'll\n\t\t * never reduce the level below \"3\", though, to match a\n\t\t * historical value that nobody complained about.\n\t\t */\n\t\tif (nr_threads < 4)\n\t\t\t; /* too few cores to consider capping */\n\t\telse if (nr_threads < 6)\n\t\t\tnr_threads = 3; /* historic cap */\n\t\telse if (nr_threads < 40)\n\t\t\tnr_threads /= 2;\n\t\telse\n\t\t\tnr_threads = 20; /* hard cap */\n\nSo I would expect this to not be single threaded unless Git was compiled\nwithout pthreads, run on a machine with few cores, or configured to use\nonly a single thread.  If you can get threading to work here, I expect\nit will perform better, although I don't have any experience with bundle\nURIs so I can't really say for certain.\n\n> Is it expected that doing it this way is so slow? The alternative is to\n> just package and work with the whole bare repo, but bundles appear to\n> be an elegant way of dealing with exactly this scenario.\n\nTo be clear, it is insecure to do anything with an untrusted repo except\nclone or fetch from it, so you will almost certainly not want to\ndistribute bare repos, since that will encourage people to use them\nas-is (which is insecure).  Even if these are internal users who can\ntrust you, it encourages an anti-pattern which has security problems in\nthe general case.\n-- \nbrian m. carlson (they/them)\nToronto, Ontario, CA\n"},{"id":"549183","messageId":"c7a623675de5f209876336f29d9983610a906ddd.camel@intel.com","threadId":"66083","inReplyTo":"amlF-ZepjtCZz1YE@fruit.crustytoothpaste.net","subject":"Re: Clones with fetch.bundleURI slower than standard, full clone?","fromName":"Knop, Ryszard","fromEmail":"ryszard.knop@intel.com","sentAt":"2026-07-29T00:41:07Z","receivedAt":"2026-07-29T00:41:11Z","isPatch":false,"body":"On Wed, 2026-07-29 at 00:14 +0000, brian m. carlson wrote:\n> On 2026-07-28 at 23:42:03, Knop, Ryszard wrote:\n> > Hey all,\n> > \n> > I'm working on a Linux kernel-related CI system where we need to\n> > perform full clones of the kernel repo in most jobs. Because of some\n> > jobs in the pipeline, it usually cannot be a shallow clone :( Since the\n> > kernel repo is large and slow to clone, and I don't want to put undue\n> > stress on the remote host, I used git bundles, where CI clones the repo\n> > over a weekend, packages that as a bundle, then in jobs it gets used\n> > like this (weird, but works for <REF> being a branch, tag or a specific\n> > commit hash):\n> > \n> > git init\n> > git remote add origin <REPO-URL>\n> > git config set fetch.bundleURI <BUNDLE-URL>\n> > git fetch origin <REF>\n> > git reset --hard FETCH_HEAD\n> > \n> > Cloning a repo this way takes 4-5mins. Unpacking a bundle appears to be\n> > super slow. Not even faster than just running a full, normal clone from\n> > the remote server, actually (~3-4mins for a single branch).\n> \n> You're comparing apples to oranges here.  A clone of a single branch\n> includes only that line of history and only those objects, but when you\n> use a bundle with multiple refs, Git has to handle all of the objects in\n> the bundle's entire pack, not just the ref you've specified.  In order\n> to compare adequately, you'd have to compare a bundle containing only\n> that one ref with the single-branch clone or a regular clone of the full\n> repository with your full bundles.\n\nI thought that due to changes mentioned in this post, Git 2.50 and\nnewer should just take in all the bundle changes and there should not\nbe much of a difference during index packs:\n\nhttps://blog.gitbutler.com/going-down-the-rabbit-hole-of-gits-new-bundle-uri\n\n> \n> > On one of the build VMs, with Git 2.53 (stock Ubuntu 26.04), GIT_TRACE\n> > suggests most of the time is spent in some variation of `/usr/lib/git-\n> > core/git index-pack --stdin -v --fix-thin '--keep=fetch-pack 39430 on\n> > build-server' --check-self-contained-and-connected`, and indeed that\n> > process burns 100% of its single thread for most of that time.\n> \n> As far as I can tell, the unbundling code just calls index-pack, so it\n> should honour pack.threads.  Setting that value to 0 causes this code to\n> be executed:\n> \n> \t\t/*\n> \t\t * Experiments show that going above 20 threads doesn't help,\n> \t\t * no matter how many cores you have. Below that, we tend to\n> \t\t * max at half the number of online_cpus(), presumably because\n> \t\t * half of those are hyperthreads rather than full cores. We'll\n> \t\t * never reduce the level below \"3\", though, to match a\n> \t\t * historical value that nobody complained about.\n> \t\t */\n> \t\tif (nr_threads < 4)\n> \t\t\t; /* too few cores to consider capping */\n> \t\telse if (nr_threads < 6)\n> \t\t\tnr_threads = 3; /* historic cap */\n> \t\telse if (nr_threads < 40)\n> \t\t\tnr_threads /= 2;\n> \t\telse\n> \t\t\tnr_threads = 20; /* hard cap */\n> \n> So I would expect this to not be single threaded unless Git was compiled\n> without pthreads, run on a machine with few cores, or configured to use\n> only a single thread.  If you can get threading to work here, I expect\n> it will perform better, although I don't have any experience with bundle\n> URIs so I can't really say for certain.\n\nOhhh, this is way better, that halved the checkout time. Thank you! It\nseems like the process still has some single-threaded sections during\npacking, but it's a problem for another day.\n\n> > Is it expected that doing it this way is so slow? The alternative is to\n> > just package and work with the whole bare repo, but bundles appear to\n> > be an elegant way of dealing with exactly this scenario.\n> \n> To be clear, it is insecure to do anything with an untrusted repo except\n> clone or fetch from it, so you will almost certainly not want to\n> distribute bare repos, since that will encourage people to use them\n> as-is (which is insecure).  Even if these are internal users who can\n> trust you, it encourages an anti-pattern which has security problems in\n> the general case.\n\nI know it's not ideal, but our reference repos are trusted and the\ncache files are used in CI only, so it would probably be fine(tm). With\nthat config change above I don't have to fall back to this though.\n\nThanks again, Ryszard\n"}]}