{"thread":{"id":"65156","subject":"git-fetch takes forever on a slow network link. Can parallel mode help?","startedAt":"2026-03-06T20:20:01Z","lastAt":"2026-03-11T18:12:35Z","messageCount":9,"participants":["R. Diez","brian m. carlson"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"538105","messageId":"5c7c975e-2541-47e1-b789-fee1fdb77d2a@rd10.de","threadId":"65156","inReplyTo":null,"subject":"git-fetch takes forever on a slow network link. Can parallel mode help?","fromName":"R. Diez","fromEmail":"rdiez-2006@rd10.de","sentAt":"2026-03-06T20:13:58Z","receivedAt":"2026-03-06T20:20:01Z","isPatch":false,"sender":{"key":"rdiez-2006@rd10.de","avatar":null},"body":"Hi all:\n\nI have an SMB/CIFS connection to a file server over a slow link of about 1 Mbps download, and a faster upload of about 10 Mbps.\n\nMy smallish Git repository has its single origin on that file server. Unfortunately, I cannot set up any sort of Git server on the remote host.\n\ngit fetch takes a long time. If the repository is up to date, it takes about 25 seconds to realise that there is nothing to do.\n\nIf there are changes to download, it can take half an hour, even if the new commit history is rather small.\n\nThe network link is slow, but not that slow. I wonder what may be causing the long delays.\n\nThe first question is: how come it takes so long to determine that nothing has changed? Does git-fetch need to download a biggish file every time?\n\nPerhaps latency is more of an issue than bandwidth. I saw that git-fetch can work in parallel with --jobs=n . Doing parallel requests may help against round trip latency.\n\nHowever, the git-fetch documentation does not clearly state whether the parallel mode only helps if you have multiple remotes and/or multiple submodules. In my case, I just have a single repository with a single origin and no submodules.\n\nAdding --jobs=10 does not help in the 25-second case with no new commits to download.\n\nDoes anybody have any ideas about how to improve performance in this scenario?\n\nThanks in advance,\n   rdiez\n"},{"id":"538110","messageId":"aas--JZ-CCWN-o7O@fruit.crustytoothpaste.net","threadId":"65156","inReplyTo":"5c7c975e-2541-47e1-b789-fee1fdb77d2a@rd10.de","subject":"Re: git-fetch takes forever on a slow network link. Can parallel mode help?","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-03-06T20:54:16Z","receivedAt":"2026-03-06T20:54:18Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2026-03-06 at 20:13:58, R. Diez wrote:\n> Hi all:\n\nHey,\n\n> I have an SMB/CIFS connection to a file server over a slow link of about 1 Mbps download, and a faster upload of about 10 Mbps.\n> \n> My smallish Git repository has its single origin on that file server. Unfortunately, I cannot set up any sort of Git server on the remote host.\n> \n> git fetch takes a long time. If the repository is up to date, it takes about 25 seconds to realise that there is nothing to do.\n> \n> If there are changes to download, it can take half an hour, even if the new commit history is rather small.\n> \n> The network link is slow, but not that slow. I wonder what may be causing the long delays.\n> \n> The first question is: how come it takes so long to determine that nothing has changed? Does git-fetch need to download a biggish file every time?\n\n1 Mbps is considered extremely slow for a modern disk.  A floppy disk\nwas 250 kbps[0], so your speed is about four times that of a floppy\ndisk.  Hard disks in 1998 were about 10 MB/s[1], so about 80 times that\nspeed.  That's definitely a big part of the problem.\n\nSince this is presumably a bare repository, Git will first read the\nremote references to determine what's available, so if you're using the\ndefault files backend, it will read each of the refs, which may involve\nmany small network requests.  This performance could be improved with\n`git pack-refs` or by converting to the reftable backend, which will\nopen fewer files.  reftable also uses some simple compression for ref\nnames, which will help as well, but it requires a relatively recent Git.\n`git refs migrate` can be used to convert to reftable if you like.\n\nOnce Git knows what the remote repository's refs are, it will need to\nwalk the history to find out what it does and doesn't have.  If there\nare many lines of development, then Git will do more work; if there is\njust one main branch to fetch, then there will be less.  This will\ninvolve opening every loose commit or tag object or reading every packed\ncommit or tag object in the history path to determine what needs to be\ncopied.  If there's nothing to copy, then Git can determine that from\nthe refs and won't walk any history or copy any objects.\n\nIf you _do_ have to transfer data, I'm not sure whether having the data\npacked or loose will be more efficient in your case due to the slow\nspeed.  You can try packing the repository with `git gc` and see how\nthat affects future transfers.  If latency is the cost, then packing\nwill almost certainly be more efficient.\n\nYou can also see how long various operations take by using\n`GIT_TRACE2=1`, which will give some detailed timing information that\nwill help you see what the expensive parts are.\n\nIf you have some trace output showing timings, we can advise on what you\nmight do to help us address performance.\n\n> However, the git-fetch documentation does not clearly state whether the parallel mode only helps if you have multiple remotes and/or multiple submodules. In my case, I just have a single repository with a single origin and no submodules.\n\nParallel mode does not help with a single remote.  All the data for a\nsingle remote comes in one job.\n\n[0] https://stackoverflow.com/questions/52841124/how-fast-could-you-read-write-to-floppy-disks-both-3-1-4-and-5-1-2\n[1] https://goughlui.com/the-hard-disk-corner/hard-drive-performance-over-the-years/\n-- \nbrian m. carlson (they/them)\nToronto, Ontario, CA\n"},{"id":"538180","messageId":"1d6a8eec-20b3-4d6e-83f1-d18b7a3c0145@rd10.de","threadId":"65156","inReplyTo":"aas--JZ-CCWN-o7O@fruit.crustytoothpaste.net","subject":"Re: git-fetch takes forever on a slow network link. Can parallel mode help?","fromName":"R. Diez","fromEmail":"rdiez-2006@rd10.de","sentAt":"2026-03-07T21:28:10Z","receivedAt":"2026-03-07T21:28:19Z","isPatch":false,"sender":{"key":"rdiez-2006@rd10.de","avatar":null},"body":"Hallo Brian:\n\nFirst of all, thanks for your quick feedback.\n\n\n> Since this is presumably a bare repository,\n\nYes, the remote repository is bare.\n\n\n> [...]\n> This performance could be improved with `git pack-refs`\n\nAfter looking around, it turns out that the documentation of \"git gc\" says that \"packing refs\" is one of the things it already does.\n\nI'll check when it was the last time I did a \"git gc\" on the remote bare repository, when I'm there again.\n\n\n> or by converting to the reftable backend, which will open fewer files.\n\nThe documentation states: \"reftable for the reftable format. This format is experimental and its internals are subject to change.\". I am not ready to risk it yet on my precious Git repository. 8-)\n\n\n> [...]\n> You can also see how long various operations take by using\n> `GIT_TRACE2=1`, which will give some detailed timing information that\n> will help you see what the expensive parts are.\n\nThat didn't help much. Most of the time (23.7 from 24 seconds) is spent in a single child process:\nchild_start[0] 'git-upload-pack '\\''/home/rdiez/MountPoints/blah/blah'\\'''\n\nThe log talks about \"upload pack\", but I gather this is actually a download operation. It wouldn't be the first confusing item in Git. Or have I got it wrong?\n\nI added \"export GIT_TRACE_PACKET=true\", and then I got a more useful breakdown:\n\nThis takes around 13 seconds:\n\n   pkt-line.c:85           packet:  upload-pack< 0000\n\nI don't know what 0000 means. All other similar \"upload-pack\" lines have a hash there.\n\nAbout 2 seconds are spent here:\n\n  pkt-line.c:85           packet:  upload-pack> [some hash]  HEAD symref-target:refs/heads/master\n  pkt-line.c:85           packet:  upload-pack> [some hash]  refs/heads/master\n\n7 seconds are spent with \"upload-pack\" and \"fetch\" operations, mainly for single \"refs/tags\". I'll check whether that improves after the next \"git gc\" on the server.\n\n\n>> However, the git-fetch documentation does not clearly state whether the parallel mode only helps if you have multiple remotes and/or multiple submodules. In my case, I just have a single repository with a single origin and no submodules.\n> \n> Parallel mode does not help with a single remote.  All the data for a single remote comes in one job.\n\nIs this due to a simple implementation in Git? Could Git download such \"refs/tags\" files in parallel?\n\nBest regards,\n   rdiez\n"},{"id":"538185","messageId":"aazUlMBj_IK41Ss2@fruit.crustytoothpaste.net","threadId":"65156","inReplyTo":"1d6a8eec-20b3-4d6e-83f1-d18b7a3c0145@rd10.de","subject":"Re: git-fetch takes forever on a slow network link. Can parallel mode help?","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-03-08T01:44:52Z","receivedAt":"2026-03-08T01:45:00Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2026-03-07 at 21:28:10, R. Diez wrote:\n> Hallo Brian:\n\nHey,\n\n> > This performance could be improved with `git pack-refs`\n> \n> After looking around, it turns out that the documentation of \"git gc\" says that \"packing refs\" is one of the things it already does.\n> \n> I'll check when it was the last time I did a \"git gc\" on the remote bare repository, when I'm there again.\n\nYes, this is part of a gc.  However, packing refs is much lighter than a\nfull GC and will therefore be much faster to complete.\n\n> > or by converting to the reftable backend, which will open fewer files.\n> \n> The documentation states: \"reftable for the reftable format. This format is experimental and its internals are subject to change.\". I am not ready to risk it yet on my precious Git repository. 8-)\n\nIt will be the default on Git 3.0 and it's in use on major forges.  I\nalso use it on several of my development repositories.  It's stable and\nfunctional.  I'll try to send a patch to fix that text.\n\nI would definitely recommend at the very least Git 2.51 for this and\nideally the latest stable version, 2.53.  Git has had a lot of work on\nthis format to improve performance and stability over the past few\nreleases.\n\n> That didn't help much. Most of the time (23.7 from 24 seconds) is spent in a single child process:\n> child_start[0] 'git-upload-pack '\\''/home/rdiez/MountPoints/blah/blah'\\'''\n> \n> The log talks about \"upload pack\", but I gather this is actually a download operation. It wouldn't be the first confusing item in Git. Or have I got it wrong?\n\nupload-pack refers to what's happening on the server.  If you contact a\nGit server over something like HTTPS or SSH, then it will use\ngit-upload-pack to send data to you (a fetch or clone from your\nperspective) or git-receive-pack to receive data from you (a push from\nyour perspective).\n\nWhen you perform a local fetch, upload-pack is spawned in the remote\nrepository to serve data.\n\n> I added \"export GIT_TRACE_PACKET=true\", and then I got a more useful breakdown:\n> \n> This takes around 13 seconds:\n> \n>   pkt-line.c:85           packet:  upload-pack< 0000\n\nIs it just that line that takes 13 seconds or is the listing of\nreferences altogether that takes 13 seconds?  That particular line\nshould not take 13 seconds because it's literally just writing and\nflushing 4 bytes.\n\nIt would be helpful if you can to include the entire trace output so we\ncan see and analyze it ourselves.  It's very hard to analyze data from\nthe different sections in isolation if one is not intimately familiar\nwith the protocol.\n\n> I don't know what 0000 means. All other similar \"upload-pack\" lines have a hash there.\n\nGit uses a pkt-line format where each line or chunk of data is preceded\nby the total length of the data (including the length itself) encoded as\nfour hex characters.  So a single byte of data with the value A plus a\nnewline would be `0006A\\n` (four bytes for the length, plus two bytes of\ndata).  The special code 0000 is a flush packet and means that the end\nof a command or a section has been reached.  That's how Git knows the\nadvertisement has finished.\n\n`GIT_TRACE_PACKET` does not normally print the pkt-line unless it's a\nflush (0000) packet or a delimiter (0001) packet, since it would just be\nnoise.\n\n> About 2 seconds are spent here:\n> \n>  pkt-line.c:85           packet:  upload-pack> [some hash]  HEAD symref-target:refs/heads/master\n>  pkt-line.c:85           packet:  upload-pack> [some hash]  refs/heads/master\n\nThat's sending references, which is expected.\n\n> 7 seconds are spent with \"upload-pack\" and \"fetch\" operations, mainly for single \"refs/tags\". I'll check whether that improves after the next \"git gc\" on the server.\n\nOkay, this is helpful.  You probably have the `peel` capability, which\nmeans that when you have a tag, you get a line like this:\n\n    4a76996b9c60ca3f21e644d78e1e5089a06c6fb3 refs/tags/v0.1.0 peeled:b4c993704e90881bec9c217749be813c70ae2bb6\n\nThat `peeled` directive tells us what object the tag points to, but it\nmeans that the tag object has to be opened and read, which makes things\nmuch more expensive.  Unfortunately, there's no way to turn that\ncapability off, since Git doesn't usually have capability control\noptions for the protocol.\n\n_However_, if you pack references with `git pack-refs` or you use\nreftable, then Git will store the references both peeled and unpeeled,\nso it doesn't need to compute that.  reftable is better because _all_\ntags are stored both peeled and unpeeled, but as long as you're writing\nnew references into a files-style repository, the new references are\nunpacked (and therefore contain no peeling information).  reftable is\nalso a binary format which means that it's smaller than a packed-refs\nfile and since your read speed is the limiting factor, that should make\nreads faster.\n\n> > > However, the git-fetch documentation does not clearly state whether the parallel mode only helps if you have multiple remotes and/or multiple submodules. In my case, I just have a single repository with a single origin and no submodules.\n> > \n> > Parallel mode does not help with a single remote.  All the data for a single remote comes in one job.\n> \n> Is this due to a simple implementation in Git? Could Git download such \"refs/tags\" files in parallel?\n\nGit is already downloading them as efficiently as possible.  The\nprotocol has both sides advertise the references (branches, tags, etc.)\nthat they have and then, in a fetch or clone, the client sends a list of\nwhat it has and what it wants, and the two sides negotiate to come to an\nagreement on what needs to be sent.  This shared understanding includes\n_all_ of the objects necessary for everything the client wants but\ndoesn't have, and then those are all sent as part of one pack.\n\nParallelization would not help here because the limiting factor is the\nspeed of the connection (and in your case, literally the speed of\nreading data off the file system).  A different design with\nparallelization might work if one had a very fast connection and the\nspeed of deltification and compression were slower than enough to max\nout the connection, but that point is around 50 MB/s in a typical\nsituation and that wouldn't matter here because the server component is\non the same file systems as well.\n-- \nbrian m. carlson (they/them)\nToronto, Ontario, CA\n"},{"id":"538215","messageId":"0ebf757b-eab5-424a-a58b-e654b1a2942e@rd10.de","threadId":"65156","inReplyTo":"aazUlMBj_IK41Ss2@fruit.crustytoothpaste.net","subject":"Re: git-fetch takes forever on a slow network link. Can parallel mode help?","fromName":"R. Diez","fromEmail":"rdiez-2006@rd10.de","sentAt":"2026-03-08T21:08:41Z","receivedAt":"2026-03-08T21:08:50Z","isPatch":false,"sender":{"key":"rdiez-2006@rd10.de","avatar":null},"body":"Hi again:\n\n>> The log talks about \"upload pack\", but I gather this is actually a download operation. It wouldn't be the first confusing item in Git. Or have I got it wrong?\n> \n> upload-pack refers to what's happening on the server.  If you contact a\n> Git server over something like HTTPS or SSH, then it will use\n> git-upload-pack to send data to you (a fetch or clone from your\n> perspective) or git-receive-pack to receive data from you (a push from\n> your perspective).\n> \n> When you perform a local fetch, upload-pack is spawned in the remote\n> repository to serve data.\n\nMy client computer has an SMB/CIFS connection to the remote file server. That means the client has mounted the file share with \"mount.cifs\", so in this scenario nothing is happening on the server, as the connection is not HTTPS or SSH. No process will be spawned on the remote server.\n\nThat is the reason why I am getting confused. From my point of view, my client computer is not \"uploading\" anything when doing a \"git pull\".\n\nBut I guess Git is designed for all scenarios and will probably not use the correct terminology in my case.\n\nIn case it helps, I am using Git version 2.53.0.\n\n\n>> I added \"export GIT_TRACE_PACKET=true\", and then I got a more useful breakdown:\n>>\n>> This takes around 13 seconds:\n>>\n>>    pkt-line.c:85           packet:  upload-pack< 0000\n> \n> Is it just that line that takes 13 seconds or is the listing of\n> references altogether that takes 13 seconds?  That particular line\n> should not take 13 seconds because it's literally just writing and\n> flushing 4 bytes.\n> \n> It would be helpful if you can to include the entire trace output so we\n> can see and analyze it ourselves.  It's very hard to analyze data from\n> the different sections in isolation if one is not intimately familiar\n> with the protocol.\n\nThe log does not really say which operation is taking how long. It does not say when the listing of references starts or finishes, which files it is reading and how many bytes it is reading from each file, or whether the files are read sequentially or in parallel.\n\nThanks for your feedback. I know it is hard to help without the whole log, but I would have to ask for permission to upload a log with file paths, hashes and tag names. Or clean them all manually.\n\n\n>> 7 seconds are spent with \"upload-pack\" and \"fetch\" operations, mainly for single \"refs/tags\". I'll check whether that improves after the next \"git gc\" on the server.\n> \n> Okay, this is helpful.  You probably have the `peel` capability, which\n> means that when you have a tag, you get a line like this:\n> \n>      4a76996b9c60ca3f21e644d78e1e5089a06c6fb3 refs/tags/v0.1.0 peeled:b4c993704e90881bec9c217749be813c70ae2bb6\n\nYes, that is the case.\n\n\n> That `peeled` directive tells us what object the tag points to, but it\n> means that the tag object has to be opened and read, which makes things\n> much more expensive.  Unfortunately, there's no way to turn that\n> capability off, since Git doesn't usually have capability control\n> options for the protocol.\n\nOK, but there is no protocol here, Git is accessing the files over the mount.\n\n\n> _However_, if you pack references with `git pack-refs` or you use\n> [...]\n\nOK, I'll try with \"git gc\" on the remote server the next time I can.\n\n\n> Git is already downloading them as efficiently as possible.  The\n> protocol has both sides advertise the references (branches, tags, etc.)\n> that they have and then, in a fetch or clone, the client sends a list of\n> what it has and what it wants, and the two sides negotiate to come to an\n> agreement on what needs to be sent.  This shared understanding includes\n> _all_ of the objects necessary for everything the client wants but\n> doesn't have, and then those are all sent as part of one pack.\n> \n> Parallelization would not help here because the limiting factor is the\n> speed of the connection (and in your case, literally the speed of\n> reading data off the file system).\n> [...]\n\nI don't think that is the case. Git is accessing the remote repository over a mount (a file share), so there is no protocol or negotiation, although I am guessing it is happening virtually with the current Git implementation.\n\nIf I understand it correctly, without \"packed references\", Git will have to access a number of small files on the remote server. Even with packet references, there will probably still be a few small files to access, in addition to some biggish packed references file.\n\nIn the past, on rotational hard disks, issuing many such read requests in parallel wasn't beneficial to performance, because of the disk head seek times. That is, jumping around would thrash the disk instead of increasing performance.\n\nBut that is not true anymore with SSDs, and especially with file mounts over a network connection with a high latency. In that scenario, issuing parallel requests (with multiple threads or async I/O) should actually increase performance.\n\nIs my reasoning correct?\n\n\nAnother question: Would it help if I only fetched the 'master' branch? Something like \"git fetch origin master\". Most of the time, I am only interested in the main branch.\n\nI am guessing that \"git fetch\" will download all other branches by default, because of this:\n\n[remote \"origin\"]\nfetch = +refs/heads/*:refs/remotes/origin/*\n\nI read the \"git fetch\" documentation, but I didn't understand whether it will fetch by default everything or just the current branch.\n\nThanks again,\n  rdiez\n"},{"id":"538217","messageId":"aa39obsSbk9R1mqu@fruit.crustytoothpaste.net","threadId":"65156","inReplyTo":"0ebf757b-eab5-424a-a58b-e654b1a2942e@rd10.de","subject":"Re: git-fetch takes forever on a slow network link. Can parallel mode help?","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-03-08T22:52:17Z","receivedAt":"2026-03-08T22:52:19Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2026-03-08 at 21:08:41, R. Diez wrote:\n> My client computer has an SMB/CIFS connection to the remote file server. That means the client has mounted the file share with \"mount.cifs\", so in this scenario nothing is happening on the server, as the connection is not HTTPS or SSH. No process will be spawned on the remote server.\n> \n> That is the reason why I am getting confused. From my point of view, my client computer is not \"uploading\" anything when doing a \"git pull\".\n> \n> But I guess Git is designed for all scenarios and will probably not use the correct terminology in my case.\n\nFor an initial clone on a local file system, Git may shortcut spawning\nan upload-pack helper and simply copy or hard link files, but otherwise,\nall fetches require the use of upload-pack.\n\nThere are a couple reasons for this.  First, upload-pack is specifically\ndesigned to deal with untrusted data without executing code or honouring\nconfiguration values, which is important for security reasons.  Second,\nwhen you're doing a fetch, Git wants to copy only the necessary objects\nand it can only do that with a helper that can read the objects.  Simply\ncopying every pack and loose object would lead to enormous bloating of\nyour client repository because you'd end up with several copies of each\nobject.\n\n> The log does not really say which operation is taking how long. It does not say when the listing of references starts or finishes, which files it is reading and how many bytes it is reading from each file, or whether the files are read sequentially or in parallel.\n\nThe log includes timestamps, which allow us to infer that information.\n\n> Thanks for your feedback. I know it is hard to help without the whole log, but I would have to ask for permission to upload a log with file paths, hashes and tag names. Or clean them all manually.\n\nI'm afraid that without more information, it's going to be difficult for\nme or anyone else to give you accurate answers about how to improve\nthis.  The trace data is specifically designed to allow us to\ntroubleshoot problems and most forges and Git-adjacent projects would\nrequire you to provide a full trace output before even investigating\nfurther.\n\n> OK, but there is no protocol here, Git is accessing the files over the mount.\n\nAs mentioned above, there is a protocol because Git always uses one for\nfetches.\n\n> I don't think that is the case. Git is accessing the remote repository over a mount (a file share), so there is no protocol or negotiation, although I am guessing it is happening virtually with the current Git implementation.\n\n`git fetch` from a remote repository on a file system spawns an\nupload-pack process in the remote repository to handle the transfer.\n`git fetch` then speaks to it over standard input and standard output.\nSo the normal protocol is being used.\n\n> If I understand it correctly, without \"packed references\", Git will have to access a number of small files on the remote server. Even with packet references, there will probably still be a few small files to access, in addition to some biggish packed references file.\n\nCorrect.\n\n> In the past, on rotational hard disks, issuing many such read requests in parallel wasn't beneficial to performance, because of the disk head seek times. That is, jumping around would thrash the disk instead of increasing performance.\n> \n> But that is not true anymore with SSDs, and especially with file mounts over a network connection with a high latency. In that scenario, issuing parallel requests (with multiple threads or async I/O) should actually increase performance.\n\nGit, like virtually every other Unix program, is not designed for high\nlatency file systems.  Yes, in theory it could be faster to issue\nmultiple requests, but that would increase the need to buffer large\namounts of data in memory, increasing memory usage, and in the general\ncase, the fact is that the file system is much lower latency and much\nfaster than the network connection over which data is being sent, so\nthat's the case that Git optimizes for.\n\nrsync would also perform poorly in your case because it's again\noptimized for sending less data over the network than it receives from\nthe file system.  Similarly with tar over a network pipe.\n\nSo it's certainly the case that Git could handle this case better, but\nit also optimizes for the common case like virtually every other modern\nUnix program.\n\nIf you think it might be faster, you could try rsyncing the remote\nrepository to a separate directory on your local machine and then\nfetching from that.  That does require that both directories are\ncompletely quiescent at the moment with no modification at all.\n\n> Another question: Would it help if I only fetched the 'master' branch? Something like \"git fetch origin master\". Most of the time, I am only interested in the main branch.\n\nThat would likely be faster.  You may also want `--no-tags`, which\nprevents downloading tags that would point into the main branch.\n\n> I am guessing that \"git fetch\" will download all other branches by default, because of this:\n> \n> [remote \"origin\"]\n> fetch = +refs/heads/*:refs/remotes/origin/*\n> \n> I read the \"git fetch\" documentation, but I didn't understand whether it will fetch by default everything or just the current branch.\n\nA `git fetch origin` with that configuration will fetch every branch and\nevery tag that points into one of those branches.\n-- \nbrian m. carlson (they/them)\nToronto, Ontario, CA\n"},{"id":"538320","messageId":"3ed2d803-5df9-44f4-9427-958d28aa1c46@rd10.de","threadId":"65156","inReplyTo":"aa39obsSbk9R1mqu@fruit.crustytoothpaste.net","subject":"Re: git-fetch takes forever on a slow network link. Can parallel mode help?","fromName":"R. Diez","fromEmail":"rdiez-2006@rd10.de","sentAt":"2026-03-09T21:08:31Z","receivedAt":"2026-03-09T21:08:40Z","isPatch":false,"sender":{"key":"rdiez-2006@rd10.de","avatar":null},"body":"\nFirst of all, thanks for the information about upload-pack etc.\n\n> [...]\n> the fact is that the file system is much lower latency and much> faster than the network connection over which data is being sent, so\n> that's the case that Git optimizes for.\n\nI wouldn't say that reading sequentially is \"optimising\". It is just the limitation of a simple implementation. Like I said, with modern SSDs, issuing requests in parallel will be faster even on a local filesystem. That would be a real optimisation then.\n\nSome elderly Unix tools like GNU Make realised long time ago that parallel operation is the way to go. Git itself has realised too, so that it can now work in parallel in certain cases (multiple remote repositories, multiple submodules). So old Unix tools don't count as an excuse!\n\nI think we should clearly point out this deficiency. Git must not be perfect, but I would rather know the limitations upfront. At the very least, that would help me make decisions faster, like investing in some sort of a Git server instead of trying to optimise the SMB/CIFS mount.\n\nAnd who knows, maybe someone will see this post in the future and decide to implement parallel file operations (async I/O) inside upload-pack and the like.\n\n\n> rsync would also perform poorly in your case because it's again\n> optimized for sending less data over the network than it receives from\n> the file system.  Similarly with tar over a network pipe.\n\nrsync would probably look at the file dates and sizes and not transfer everything. There are even some parallel rsync variants designed to overcome high network latencies.\n\nBut I don't think rsync is worth the effort for me. I'll just wait a while longer every now and then.\n\n\nThere is one more thing I am curious about. Git does not document how it uses SSH (or at least I couldn't find it in the standard end-user documentation). Git cannot launch a process on the target host over SSH, unless Git is already installed on the remote system. After all, the local system may have a different architecture (like AMD vs ARM), so you cannot copy a binary across. And I haven't seen the requirement that Git must be installed on the remote host when connecting over SSH. In that case, I would have probably seen somewhere a version compatibility table between client and server.\n\nSo Git must be accessing files over SSH using the standard SSH file transfer operations. I am guessing that the same latency problem will apply here too, because uploads and downloads over SSH will also be sequential. Is my reasoning correct?\n\nOr does Git attempt to find out whether there is a Git on the other side? What happens if there isn't then?\n\n\n> [...]\n> A `git fetch origin` with that configuration will fetch every branch and\n> every tag that points into one of those branches.\n\nOK, thanks. It turns out my repository has no branches at all, so that wouldn't help me anyway.\n\nBest regards,\n   rdiez\n\n"},{"id":"538539","messageId":"abCgUu3ZFSOIZwKu@fruit.crustytoothpaste.net","threadId":"65156","inReplyTo":"3ed2d803-5df9-44f4-9427-958d28aa1c46@rd10.de","subject":"Re: git-fetch takes forever on a slow network link. Can parallel mode help?","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-03-10T22:50:58Z","receivedAt":"2026-03-10T22:51:05Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2026-03-09 at 21:08:31, R. Diez wrote:\n> There is one more thing I am curious about. Git does not document how it uses SSH (or at least I couldn't find it in the standard end-user documentation). Git cannot launch a process on the target host over SSH, unless Git is already installed on the remote system. After all, the local system may have a different architecture (like AMD vs ARM), so you cannot copy a binary across. And I haven't seen the requirement that Git must be installed on the remote host when connecting over SSH. In that case, I would have probably seen somewhere a version compatibility table between client and server.\n> \n> So Git must be accessing files over SSH using the standard SSH file transfer operations. I am guessing that the same latency problem will apply here too, because uploads and downloads over SSH will also be sequential. Is my reasoning correct?\n\nGit doesn't use standard SSH file transfer operations.  That would be\nmuch slower and it also works poorly when the remote side doesn't grant\naccess to a file system, such as with a forge or gitolite.\n\nSSH allows multiple commands to be run over a single connection with\n`-oControlMaster`, which can improve performance.  The benefit to\nrunning a single command for each Git operation is that we can do\nauthentication once at the beginning of each command, whereas if we have\na long-running SSH connection and attempt to do SFTP, we might\ninterleave requests on different repositories, so each request would\nhave to perform authentication.  That's not a problem if you're using\nUnix permissions to control access, but it scales really poorly when\nyour Git data is actually spread across many different file servers and\nthe user is accessing multiple repositories, such as is common on\nforges.\n\nUsing SSH file transfer operations also would not work well because you\nwould effectively have to download every pack file and loose object to\nbe sure you got the data you need, instead of getting a pack with only a\nfew objects if that's all you need.\n\nHowever, you can of course mount a remote file system as SFTP with\n`sshfs` and use it as a local file system if you actually have a real\nfile system on the remote side.  That will send multiple requests over\nthe connection when reading or writing since the `sshfs` does queue\nthose.\n\n> Or does Git attempt to find out whether there is a Git on the other side? What happens if there isn't then?\n\nGit invokes git-upload-pack on the remote side and talks to it over\nstandard input and output.  If there isn't one, then the operation\nfails.\n\nHere's an example:\n\n----\n% GIT_TRACE=1 git ls-remote git@github.com:git/git.git\n22:41:36.731673 git.c:502               trace: built-in: git ls-remote git@github.com:git/git.git\n22:41:36.731937 run-command.c:673       trace: run_command: unset GIT_PREFIX; GIT_PROTOCOL=version=2 ssh -o SendEnv=GIT_PROTOCOL git@github.com 'git-upload-pack '\\''git/git.git'\\'''\n22:41:36.731952 run-command.c:765       trace: start_command: /usr/bin/ssh -o SendEnv=GIT_PROTOCOL git@github.com 'git-upload-pack '\\''git/git.git'\\'''\n----\n\nI don't have any systems without Git on them, so I can't demonstrate the\nfailure case.\n-- \nbrian m. carlson (they/them)\nToronto, Ontario, CA\n"},{"id":"538647","messageId":"d7b6defb-1614-40e6-b46b-a36d71388431@rd10.de","threadId":"65156","inReplyTo":"abCgUu3ZFSOIZwKu@fruit.crustytoothpaste.net","subject":"Re: git-fetch takes forever on a slow network link. Can parallel mode help?","fromName":"R. Diez","fromEmail":"rdiez-2006@rd10.de","sentAt":"2026-03-11T18:05:41Z","receivedAt":"2026-03-11T18:12:35Z","isPatch":false,"sender":{"key":"rdiez-2006@rd10.de","avatar":null},"body":"\n> Git doesn't use standard SSH file transfer operations.\n> [...]\n\nOK, thanks for the information.\n\nI have finally done a \"git gc\" on the server side, and now a \"git pull\" from the client with no new commits to download takes 4 seconds, a drastic reduction from the 25 seconds it took before.\n\nI turns out I hadn't done a \"git gc\" on the server for over 2 years, so that many new references weren't packed.\n\nTherefore, I think that having many small files to read versus one packed-refs file makes a huge difference if you have mounted a remote filesystem over a network with a relatively high latency.\n\nMy 1 Mbps connection does not actually have such a high latency (around 40 ms measured with ping), but latency seems to have a much greater impact than the low bandwidth, at least with a packed-refs file which only weighs 64 kB.\n\nBest regards,\n   rdiez\n"}]}