{"thread":{"id":"61605","subject":"With big repos and slower connections, git clone can be hard to work with","startedAt":"2024-06-07T23:28:08Z","lastAt":"2025-09-08T02:44:10Z","messageCount":43,"participants":["ellie","rsbecker@nexbridge.com","Jeff King","Junio C Hamano","Patrick Steinhardt","Emily Shaffer","Ivan Frade","Toon claes","Sitaram Chamarty","Konstantin Khomoutov","Emanuel Czirai","Ellie"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"496684","messageId":"fec6ebc7-efd7-4c86-9dcc-2b006bd82e47@horse64.org","threadId":"61605","inReplyTo":null,"subject":"With big repos and slower connections, git clone can be hard to work with","fromName":"ellie","fromEmail":"el@horse64.org","sentAt":"2024-06-07T23:28:05Z","receivedAt":"2024-06-07T23:28:08Z","isPatch":false,"sender":{"key":"el@horse64.org","avatar":null},"body":"Dear git team,\n\nI'm terribly sorry if this is the wrong place, but I'd like to suggest a \npotential issue with \"git clone\".\n\nThe problem is that any sort of interruption or connection issue, no \nmatter how brief, causes the clone to stop and leave nothing behind:\n\n$ git clone https://github.com/Nheko-Reborn/nheko\nCloning into 'nheko'...\nremote: Enumerating objects: 43991, done.\nremote: Counting objects: 100% (6535/6535), done.\nremote: Compressing objects: 100% (1449/1449), done.\nerror: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly: \nCANCEL (err 8)\nerror: 2771 bytes of body are still expected\nfetch-pack: unexpected disconnect while reading sideband packet\nfatal: early EOF\nfatal: fetch-pack: invalid index-pack output\n$ cd nheko\nbash: cd: nheko: No such file or director\n\nIn my experience, this can be really impactful with 1. big repositories \nand 2. unreliable internet - which I would argue isn't unheard of! E.g. \na developer may work via mobile connection on a business trip. The \nresult can even be that a repository is uncloneable for some users!\n\nThis has left me in the absurd situation where I was able to download a \ntarball via HTTPS from the git hoster just fine, even way larger binary \nrelease items, thanks to the browser's HTTPS resume. And yet a simple \ngit clone of the same project failed repeatedly.\n\nMy deepest apologies if I missed an option to fix or address this. But \nsummed up, please consider making git clone recover from hiccups.\n\nRegards,\n\nEllie\n\nPS: I've seen git hosters have apparent proxy bugs, like timing out \nslower git clone connections from the server side even if the transfer \nis ongoing. A git auto-resume would reduce the impact of that, too.\n\n\n\n"},{"id":"496685","messageId":"0be201dab933$17c02530$47406f90$@nexbridge.com","threadId":"61605","inReplyTo":"fec6ebc7-efd7-4c86-9dcc-2b006bd82e47@horse64.org","subject":"RE: With big repos and slower connections, git clone can be hard to work with","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2024-06-07T23:33:19Z","receivedAt":"2024-06-07T23:33:33Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On Friday, June 7, 2024 7:28 PM, ellie wrote:\n>I'm terribly sorry if this is the wrong place, but I'd like to suggest a potential issue\n>with \"git clone\".\n>\n>The problem is that any sort of interruption or connection issue, no matter how\n>brief, causes the clone to stop and leave nothing behind:\n>\n>$ git clone https://github.com/Nheko-Reborn/nheko\n>Cloning into 'nheko'...\n>remote: Enumerating objects: 43991, done.\n>remote: Counting objects: 100% (6535/6535), done.\n>remote: Compressing objects: 100% (1449/1449), done.\n>error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\n>CANCEL (err 8)\n>error: 2771 bytes of body are still expected\n>fetch-pack: unexpected disconnect while reading sideband packet\n>fatal: early EOF\n>fatal: fetch-pack: invalid index-pack output $ cd nheko\n>bash: cd: nheko: No such file or director\n>\n>In my experience, this can be really impactful with 1. big repositories and 2.\n>unreliable internet - which I would argue isn't unheard of! E.g.\n>a developer may work via mobile connection on a business trip. The result can even\n>be that a repository is uncloneable for some users!\n>\n>This has left me in the absurd situation where I was able to download a tarball via\n>HTTPS from the git hoster just fine, even way larger binary release items, thanks to\n>the browser's HTTPS resume. And yet a simple git clone of the same project failed\n>repeatedly.\n>\n>My deepest apologies if I missed an option to fix or address this. But summed up,\n>please consider making git clone recover from hiccups.\n>\n>Regards,\n>\n>Ellie\n>\n>PS: I've seen git hosters have apparent proxy bugs, like timing out slower git clone\n>connections from the server side even if the transfer is ongoing. A git auto-resume\n>would reduce the impact of that, too.\n\nI suggest that you look into two git topics: --depth, which controls how much history is obtained in a clone, and sparse-checkout, which describes the part of the repository you will retrieve. You can prune the contents of the repository so that clone is faster, if you do not need all of the history, or all of the files. This is typically done in complex large repositories, particularly those used for production support as release repositories.\n--Randall\n\n"},{"id":"496686","messageId":"fdb869ef-4ce9-4859-9e36-445fd9200776@horse64.org","threadId":"61605","inReplyTo":"0be201dab933$17c02530$47406f90$@nexbridge.com","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"ellie","fromEmail":"el@horse64.org","sentAt":"2024-06-08T00:03:15Z","receivedAt":"2024-06-08T00:03:17Z","isPatch":false,"sender":{"key":"el@horse64.org","avatar":null},"body":"Thanks, this is very helpful as an emergency workaround!\n\nNevertheless, I usually want the entire history, especially since I \nwouldn't mind waiting half an hour. But without resume, I've encountered \nit regularly that it just won't complete even if I give it the time, \nwhile way longer downloads in the browser would. The key problem here \nseems to be the lack of any resume.\n\nI hope this helps to understand why I made the suggestion.\n\nRegards,\n\nEllie\n\nOn 6/8/24 1:33 AM, rsbecker@nexbridge.com wrote:\n> On Friday, June 7, 2024 7:28 PM, ellie wrote:\n>> I'm terribly sorry if this is the wrong place, but I'd like to suggest a potential issue\n>> with \"git clone\".\n>>\n>> The problem is that any sort of interruption or connection issue, no matter how\n>> brief, causes the clone to stop and leave nothing behind:\n>>\n>> $ git clone https://github.com/Nheko-Reborn/nheko\n>> Cloning into 'nheko'...\n>> remote: Enumerating objects: 43991, done.\n>> remote: Counting objects: 100% (6535/6535), done.\n>> remote: Compressing objects: 100% (1449/1449), done.\n>> error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\n>> CANCEL (err 8)\n>> error: 2771 bytes of body are still expected\n>> fetch-pack: unexpected disconnect while reading sideband packet\n>> fatal: early EOF\n>> fatal: fetch-pack: invalid index-pack output $ cd nheko\n>> bash: cd: nheko: No such file or director\n>>\n>> In my experience, this can be really impactful with 1. big repositories and 2.\n>> unreliable internet - which I would argue isn't unheard of! E.g.\n>> a developer may work via mobile connection on a business trip. The result can even\n>> be that a repository is uncloneable for some users!\n>>\n>> This has left me in the absurd situation where I was able to download a tarball via\n>> HTTPS from the git hoster just fine, even way larger binary release items, thanks to\n>> the browser's HTTPS resume. And yet a simple git clone of the same project failed\n>> repeatedly.\n>>\n>> My deepest apologies if I missed an option to fix or address this. But summed up,\n>> please consider making git clone recover from hiccups.\n>>\n>> Regards,\n>>\n>> Ellie\n>>\n>> PS: I've seen git hosters have apparent proxy bugs, like timing out slower git clone\n>> connections from the server side even if the transfer is ongoing. A git auto-resume\n>> would reduce the impact of that, too.\n> \n> I suggest that you look into two git topics: --depth, which controls how much history is obtained in a clone, and sparse-checkout, which describes the part of the repository you will retrieve. You can prune the contents of the repository so that clone is faster, if you do not need all of the history, or all of the files. This is typically done in complex large repositories, particularly those used for production support as release repositories.\n> --Randall\n> \n"},{"id":"496687","messageId":"0beb01dab93b$c01dfa10$4059ee30$@nexbridge.com","threadId":"61605","inReplyTo":"fdb869ef-4ce9-4859-9e36-445fd9200776@horse64.org","subject":"RE: With big repos and slower connections, git clone can be hard to work with","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2024-06-08T00:35:18Z","receivedAt":"2024-06-08T00:35:26Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On Friday, June 7, 2024 8:03 PM, ellie wrote:\n>Subject: Re: With big repos and slower connections, git clone can be hard to work\n>with\n>\n>Thanks, this is very helpful as an emergency workaround!\n>\n>Nevertheless, I usually want the entire history, especially since I wouldn't mind\n>waiting half an hour. But without resume, I've encountered it regularly that it just\n>won't complete even if I give it the time, while way longer downloads in the\n>browser would. The key problem here seems to be the lack of any resume.\n>\n>I hope this helps to understand why I made the suggestion.\n>\n>Regards,\n>\n>Ellie\n>\n>On 6/8/24 1:33 AM, rsbecker@nexbridge.com wrote:\n>> On Friday, June 7, 2024 7:28 PM, ellie wrote:\n>>> I'm terribly sorry if this is the wrong place, but I'd like to\n>>> suggest a potential issue with \"git clone\".\n>>>\n>>> The problem is that any sort of interruption or connection issue, no\n>>> matter how brief, causes the clone to stop and leave nothing behind:\n>>>\n>>> $ git clone https://github.com/Nheko-Reborn/nheko\n>>> Cloning into 'nheko'...\n>>> remote: Enumerating objects: 43991, done.\n>>> remote: Counting objects: 100% (6535/6535), done.\n>>> remote: Compressing objects: 100% (1449/1449), done.\n>>> error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\n>>> CANCEL (err 8)\n>>> error: 2771 bytes of body are still expected\n>>> fetch-pack: unexpected disconnect while reading sideband packet\n>>> fatal: early EOF\n>>> fatal: fetch-pack: invalid index-pack output $ cd nheko\n>>> bash: cd: nheko: No such file or director\n>>>\n>>> In my experience, this can be really impactful with 1. big repositories and 2.\n>>> unreliable internet - which I would argue isn't unheard of! E.g.\n>>> a developer may work via mobile connection on a business trip. The\n>>> result can even be that a repository is uncloneable for some users!\n>>>\n>>> This has left me in the absurd situation where I was able to download\n>>> a tarball via HTTPS from the git hoster just fine, even way larger\n>>> binary release items, thanks to the browser's HTTPS resume. And yet a\n>>> simple git clone of the same project failed repeatedly.\n>>>\n>>> My deepest apologies if I missed an option to fix or address this.\n>>> But summed up, please consider making git clone recover from hiccups.\n>>>\n>>> Regards,\n>>>\n>>> Ellie\n>>>\n>>> PS: I've seen git hosters have apparent proxy bugs, like timing out\n>>> slower git clone connections from the server side even if the\n>>> transfer is ongoing. A git auto-resume would reduce the impact of that, too.\n>>\n>> I suggest that you look into two git topics: --depth, which controls how much\n>history is obtained in a clone, and sparse-checkout, which describes the part of the\n>repository you will retrieve. You can prune the contents of the repository so that\n>clone is faster, if you do not need all of the history, or all of the files. This is typically\n>done in complex large repositories, particularly those used for production support\n>as release repositories.\n\nConsider doing the clone with --depth=1 then using git fetch --depth=n as the resume. There are other options that effectively give you a resume, including --deepen=n.\n\nBuild automation, like Jenkins, uses this to speed up the clone/checkout.\n\n"},{"id":"496688","messageId":"200c3bd2-6aa9-4bb2-8eda-881bb62cd064@horse64.org","threadId":"61605","inReplyTo":"0beb01dab93b$c01dfa10$4059ee30$@nexbridge.com","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"ellie","fromEmail":"el@horse64.org","sentAt":"2024-06-08T00:46:38Z","receivedAt":"2024-06-08T00:46:41Z","isPatch":false,"sender":{"key":"el@horse64.org","avatar":null},"body":"The deepening worked perfectly, thank you so much! I hope a resume will \nstill be considered however, if even just to help out newcomers.\n\nRegards,\n\nEllie\n\nOn 6/8/24 2:35 AM, rsbecker@nexbridge.com wrote:\n> On Friday, June 7, 2024 8:03 PM, ellie wrote:\n>> Subject: Re: With big repos and slower connections, git clone can be hard to work\n>> with\n>>\n>> Thanks, this is very helpful as an emergency workaround!\n>>\n>> Nevertheless, I usually want the entire history, especially since I wouldn't mind\n>> waiting half an hour. But without resume, I've encountered it regularly that it just\n>> won't complete even if I give it the time, while way longer downloads in the\n>> browser would. The key problem here seems to be the lack of any resume.\n>>\n>> I hope this helps to understand why I made the suggestion.\n>>\n>> Regards,\n>>\n>> Ellie\n>>\n>> On 6/8/24 1:33 AM, rsbecker@nexbridge.com wrote:\n>>> On Friday, June 7, 2024 7:28 PM, ellie wrote:\n>>>> I'm terribly sorry if this is the wrong place, but I'd like to\n>>>> suggest a potential issue with \"git clone\".\n>>>>\n>>>> The problem is that any sort of interruption or connection issue, no\n>>>> matter how brief, causes the clone to stop and leave nothing behind:\n>>>>\n>>>> $ git clone https://github.com/Nheko-Reborn/nheko\n>>>> Cloning into 'nheko'...\n>>>> remote: Enumerating objects: 43991, done.\n>>>> remote: Counting objects: 100% (6535/6535), done.\n>>>> remote: Compressing objects: 100% (1449/1449), done.\n>>>> error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\n>>>> CANCEL (err 8)\n>>>> error: 2771 bytes of body are still expected\n>>>> fetch-pack: unexpected disconnect while reading sideband packet\n>>>> fatal: early EOF\n>>>> fatal: fetch-pack: invalid index-pack output $ cd nheko\n>>>> bash: cd: nheko: No such file or director\n>>>>\n>>>> In my experience, this can be really impactful with 1. big repositories and 2.\n>>>> unreliable internet - which I would argue isn't unheard of! E.g.\n>>>> a developer may work via mobile connection on a business trip. The\n>>>> result can even be that a repository is uncloneable for some users!\n>>>>\n>>>> This has left me in the absurd situation where I was able to download\n>>>> a tarball via HTTPS from the git hoster just fine, even way larger\n>>>> binary release items, thanks to the browser's HTTPS resume. And yet a\n>>>> simple git clone of the same project failed repeatedly.\n>>>>\n>>>> My deepest apologies if I missed an option to fix or address this.\n>>>> But summed up, please consider making git clone recover from hiccups.\n>>>>\n>>>> Regards,\n>>>>\n>>>> Ellie\n>>>>\n>>>> PS: I've seen git hosters have apparent proxy bugs, like timing out\n>>>> slower git clone connections from the server side even if the\n>>>> transfer is ongoing. A git auto-resume would reduce the impact of that, too.\n>>>\n>>> I suggest that you look into two git topics: --depth, which controls how much\n>> history is obtained in a clone, and sparse-checkout, which describes the part of the\n>> repository you will retrieve. You can prune the contents of the repository so that\n>> clone is faster, if you do not need all of the history, or all of the files. This is typically\n>> done in complex large repositories, particularly those used for production support\n>> as release repositories.\n> \n> Consider doing the clone with --depth=1 then using git fetch --depth=n as the resume. There are other options that effectively give you a resume, including --deepen=n.\n> \n> Build automation, like Jenkins, uses this to speed up the clone/checkout.\n> \n"},{"id":"496692","messageId":"20240608084323.GB2390433@coredump.intra.peff.net","threadId":"61605","inReplyTo":"200c3bd2-6aa9-4bb2-8eda-881bb62cd064@horse64.org","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2024-06-08T08:43:23Z","receivedAt":"2024-06-08T08:43:24Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Jun 08, 2024 at 02:46:38AM +0200, ellie wrote:\n\n> The deepening worked perfectly, thank you so much! I hope a resume will\n> still be considered however, if even just to help out newcomers.\n\nBecause the packfile to send the user is created on the fly, making a\nclone fully resumable is tricky (a second clone may get an equivalent\nbut slightly different pack due to new objects entering the repo, or\neven raciness between threads).\n\nOne strategy people have worked on is for servers to point clients at\nstatic packfiles (which _do_ remain byte-for-byte identical, and can be\nresumed) to get some of the objects. But it requires some scheme on the\nserver side to decide when and how to create those packfiles. So while\nthere is support inside Git itself for this idea (both on the server and\nclient side), I don't know of any servers where it is in active use.\n\n-Peff\n"},{"id":"496696","messageId":"bc030171-70fa-41cf-945a-2d20bf237372@horse64.org","threadId":"61605","inReplyTo":"20240608084323.GB2390433@coredump.intra.peff.net","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"ellie","fromEmail":"el@horse64.org","sentAt":"2024-06-08T09:40:47Z","receivedAt":"2024-06-08T09:40:51Z","isPatch":false,"sender":{"key":"el@horse64.org","avatar":null},"body":"Sorry if I'm misunderstanding, and I assume this is a naive suggestion \nthat may not work in some way: but couldn't git somehow retain all the \nobjects it already has fully downloaded cached? And then otherwise start \nover cleanly (and automatically), but just get the objects it already \nhas from the local cache? In practice, that might already be enough to \nget through a longer clone despite occasional hiccups.\n\nSorry, I'm really not qualified to make good suggestions, it's just that \nthe current situation feels frustrating as an outside user.\n\nRegards,\n\nEllie\n\nOn 6/8/24 10:43 AM, Jeff King wrote:\n> On Sat, Jun 08, 2024 at 02:46:38AM +0200, ellie wrote:\n> \n>> The deepening worked perfectly, thank you so much! I hope a resume will\n>> still be considered however, if even just to help out newcomers.\n> \n> Because the packfile to send the user is created on the fly, making a\n> clone fully resumable is tricky (a second clone may get an equivalent\n> but slightly different pack due to new objects entering the repo, or\n> even raciness between threads).\n> \n> One strategy people have worked on is for servers to point clients at\n> static packfiles (which _do_ remain byte-for-byte identical, and can be\n> resumed) to get some of the objects. But it requires some scheme on the\n> server side to decide when and how to create those packfiles. So while\n> there is support inside Git itself for this idea (both on the server and\n> client side), I don't know of any servers where it is in active use.\n> \n> -Peff\n"},{"id":"496698","messageId":"7de78f3d-f174-4bf6-837f-9c90bf935d21@horse64.org","threadId":"61605","inReplyTo":"bc030171-70fa-41cf-945a-2d20bf237372@horse64.org","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"ellie","fromEmail":"el@horse64.org","sentAt":"2024-06-08T09:44:09Z","receivedAt":"2024-06-08T09:44:11Z","isPatch":false,"sender":{"key":"el@horse64.org","avatar":null},"body":"Another idea that probably is silly in some way too: couldn't after the \nfirst error, git automatically start over and do this whole --depth=1 \nfollowed by --deepen... automatically? I feel like anything that \nwouldn't require knowing and manually doing that process would be an \nimprovement for people affected often by this.\n\nRegards,\n\nEllie\n\nOn 6/8/24 11:40 AM, ellie wrote:\n> Sorry if I'm misunderstanding, and I assume this is a naive suggestion \n> that may not work in some way: but couldn't git somehow retain all the \n> objects it already has fully downloaded cached? And then otherwise start \n> over cleanly (and automatically), but just get the objects it already \n> has from the local cache? In practice, that might already be enough to \n> get through a longer clone despite occasional hiccups.\n> \n> Sorry, I'm really not qualified to make good suggestions, it's just that \n> the current situation feels frustrating as an outside user.\n> \n> Regards,\n> \n> Ellie\n> \n> On 6/8/24 10:43 AM, Jeff King wrote:\n>> On Sat, Jun 08, 2024 at 02:46:38AM +0200, ellie wrote:\n>>\n>>> The deepening worked perfectly, thank you so much! I hope a resume will\n>>> still be considered however, if even just to help out newcomers.\n>>\n>> Because the packfile to send the user is created on the fly, making a\n>> clone fully resumable is tricky (a second clone may get an equivalent\n>> but slightly different pack due to new objects entering the repo, or\n>> even raciness between threads).\n>>\n>> One strategy people have worked on is for servers to point clients at\n>> static packfiles (which _do_ remain byte-for-byte identical, and can be\n>> resumed) to get some of the objects. But it requires some scheme on the\n>> server side to decide when and how to create those packfiles. So while\n>> there is support inside Git itself for this idea (both on the server and\n>> client side), I don't know of any servers where it is in active use.\n>>\n>> -Peff\n"},{"id":"496702","messageId":"20240608103533.GD2659849@coredump.intra.peff.net","threadId":"61605","inReplyTo":"bc030171-70fa-41cf-945a-2d20bf237372@horse64.org","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2024-06-08T10:35:33Z","receivedAt":"2024-06-08T10:35:34Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Jun 08, 2024 at 11:40:47AM +0200, ellie wrote:\n\n> Sorry if I'm misunderstanding, and I assume this is a naive suggestion that\n> may not work in some way: but couldn't git somehow retain all the objects it\n> already has fully downloaded cached? And then otherwise start over cleanly\n> (and automatically), but just get the objects it already has from the local\n> cache? In practice, that might already be enough to get through a longer\n> clone despite occasional hiccups.\n\nThe problem is that the client/server communication does not share an\nexplicit list of objects. Instead, the client tells the server some\npoints in the object graph that it wants (i.e., the tips of some\nbranches that it wants to fetch) and that it already has (existing\nbranches, or nothing in the case of a clone), and then the server can do\nits own graph traversal to figure out what needs to be sent.\n\nWhen you've got a partially completed clone, the client can figure out\nwhich objects it received. But it can't tell the server \"hey, I have\ncommit XYZ, don't send that\". Because the server would assume that\nhaving XYZ means that it has all of the objects reachable from there\n(parent commits, their trees and blobs, and so on). And the pack does\nnot come in that order.\n\nAnd even if there was a way to disable reachability analysis, and send a\n\"raw\" set of objects that we already have, it would be prohibitively\nlarge. The full set of sha1 hashes for linux.git is over 200MB. So\nnaively saying \"don't send object X, I have it\" would approach that\nsize.\n\nIt's possible the client could do some analysis to see if it has\ncomplete segments of history. In practice it won't, because of the way\nwe order packfiles (it's split by type, and then roughly\nreverse-chronological through history). If the server re-ordered its\nresponse to fill history from the bottom up, it would be possible. We\ndon't do that now because it's not really the optimal order for\naccessing objects in day-to-day use, and the packfile the server sends\nis stored directly on disk by the client.\n\n-Peff\n"},{"id":"496703","messageId":"20240608103832.GE2659849@coredump.intra.peff.net","threadId":"61605","inReplyTo":"7de78f3d-f174-4bf6-837f-9c90bf935d21@horse64.org","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2024-06-08T10:38:32Z","receivedAt":"2024-06-08T10:38:33Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Jun 08, 2024 at 11:44:09AM +0200, ellie wrote:\n\n> Another idea that probably is silly in some way too: couldn't after the\n> first error, git automatically start over and do this whole --depth=1\n> followed by --deepen... automatically? I feel like anything that wouldn't\n> require knowing and manually doing that process would be an improvement for\n> people affected often by this.\n\nI'm skeptical that shallow-cloning and deepening is a good strategy in\ngeneral. Serving shallow clones like this is expensive for the server,\nand there's more network overhead in the back-and-forth requests.\n\nIt also only slices up the repository in one dimension. There could be\na single tree that's really big, or even a single blob that you can\nnever get past.\n\nSo yes, it may work sometimes, but I don't think it's something we\nshould codify.\n\n-Peff\n"},{"id":"496708","messageId":"a424b43f-d477-46bf-a28e-9d7a87a0130e@horse64.org","threadId":"61605","inReplyTo":"20240608103533.GD2659849@coredump.intra.peff.net","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"ellie","fromEmail":"el@horse64.org","sentAt":"2024-06-08T11:05:47Z","receivedAt":"2024-06-08T11:05:50Z","isPatch":false,"sender":{"key":"el@horse64.org","avatar":null},"body":"\nI see! Unfortunate, but I'm thankful for your detailed explanation.\n\nThe \"shallow-cloning and deepening is [...] expensive for the server\" \nmakes me sadder about the current situation. I don't like that I need to \nmake the server's life hard just because my connection is shaky... :-|\n\n > It's possible the client could do some analysis to see if it has\n > complete segments of history. In practice it won't, because of the way\n > we order packfiles (it's split by type, and then roughly\n > reverse-chronological through history). If the server re-ordered its\n > response to fill history from the bottom up, it would be possible.\n\nI wonder if that would be the most feasible idea, if any at all...?\n\nMy main take-away is that I don't know enough to suggest a good way out, \nand that git is even more impressive and complex tech than I thought. \nThanks so much for the detailed responses, and I hope at least some of \nmy uninformed rambling was of any use.\n\nRegards,\n\nEllie\n\nOn 6/8/24 12:35 PM, Jeff King wrote:\n> On Sat, Jun 08, 2024 at 11:40:47AM +0200, ellie wrote:\n> \n>> Sorry if I'm misunderstanding, and I assume this is a naive suggestion that\n>> may not work in some way: but couldn't git somehow retain all the objects it\n>> already has fully downloaded cached? And then otherwise start over cleanly\n>> (and automatically), but just get the objects it already has from the local\n>> cache? In practice, that might already be enough to get through a longer\n>> clone despite occasional hiccups.\n> \n> The problem is that the client/server communication does not share an\n> explicit list of objects. Instead, the client tells the server some\n> points in the object graph that it wants (i.e., the tips of some\n> branches that it wants to fetch) and that it already has (existing\n> branches, or nothing in the case of a clone), and then the server can do\n> its own graph traversal to figure out what needs to be sent.\n> \n> When you've got a partially completed clone, the client can figure out\n> which objects it received. But it can't tell the server \"hey, I have\n> commit XYZ, don't send that\". Because the server would assume that\n> having XYZ means that it has all of the objects reachable from there\n> (parent commits, their trees and blobs, and so on). And the pack does\n> not come in that order.\n> \n> And even if there was a way to disable reachability analysis, and send a\n> \"raw\" set of objects that we already have, it would be prohibitively\n> large. The full set of sha1 hashes for linux.git is over 200MB. So\n> naively saying \"don't send object X, I have it\" would approach that\n> size.\n> \n> It's possible the client could do some analysis to see if it has\n> complete segments of history. In practice it won't, because of the way\n> we order packfiles (it's split by type, and then roughly\n> reverse-chronological through history). If the server re-ordered its\n> response to fill history from the bottom up, it would be possible. We\n> don't do that now because it's not really the optimal order for\n> accessing objects in day-to-day use, and the packfile the server sends\n> is stored directly on disk by the client.\n> \n> -Peff\n"},{"id":"496722","messageId":"xmqq8qzfcl98.fsf@gitster.g","threadId":"61605","inReplyTo":"20240608084323.GB2390433@coredump.intra.peff.net","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2024-06-08T19:00:51Z","receivedAt":"2024-06-08T19:00:54Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> One strategy people have worked on is for servers to point clients at\n> static packfiles (which _do_ remain byte-for-byte identical, and can be\n> resumed) to get some of the objects. But it requires some scheme on the\n> server side to decide when and how to create those packfiles. So while\n> there is support inside Git itself for this idea (both on the server and\n> client side), I don't know of any servers where it is in active use.\n\nDidn't the bundle URL work originate at GitHub?  I thought this use\ncase was a reasonable match to the mechanism.\n\n"},{"id":"496725","messageId":"15611436-01a4-4df5-a8a6-240e502128e4@horse64.org","threadId":"61605","inReplyTo":"xmqq8qzfcl98.fsf@gitster.g","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"ellie","fromEmail":"el@horse64.org","sentAt":"2024-06-08T20:16:21Z","receivedAt":"2024-06-08T20:16:26Z","isPatch":false,"sender":{"key":"el@horse64.org","avatar":null},"body":"(I'm probably not the person to answer fully. But I can say HTTPS Git \nclones from GitHub don't ever resume for me, if that's informative.)\n\nOn 6/8/24 9:00 PM, Junio C Hamano wrote:\n> Jeff King <peff@peff.net> writes:\n> \n>> One strategy people have worked on is for servers to point clients at\n>> static packfiles (which _do_ remain byte-for-byte identical, and can be\n>> resumed) to get some of the objects. But it requires some scheme on the\n>> server side to decide when and how to create those packfiles. So while\n>> there is support inside Git itself for this idea (both on the server and\n>> client side), I don't know of any servers where it is in active use.\n> \n> Didn't the bundle URL work originate at GitHub?  I thought this use\n> case was a reasonable match to the mechanism.\n> \n"},{"id":"496753","messageId":"ZmahU1Gv_abt5RGn@tanuki","threadId":"61605","inReplyTo":"20240608084323.GB2390433@coredump.intra.peff.net","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2024-06-10T06:46:43Z","receivedAt":"2024-06-10T06:46:49Z","isPatch":false,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Sat, Jun 08, 2024 at 04:43:23AM -0400, Jeff King wrote:\n> On Sat, Jun 08, 2024 at 02:46:38AM +0200, ellie wrote:\n> \n> > The deepening worked perfectly, thank you so much! I hope a resume will\n> > still be considered however, if even just to help out newcomers.\n> \n> Because the packfile to send the user is created on the fly, making a\n> clone fully resumable is tricky (a second clone may get an equivalent\n> but slightly different pack due to new objects entering the repo, or\n> even raciness between threads).\n> \n> One strategy people have worked on is for servers to point clients at\n> static packfiles (which _do_ remain byte-for-byte identical, and can be\n> resumed) to get some of the objects. But it requires some scheme on the\n> server side to decide when and how to create those packfiles. So while\n> there is support inside Git itself for this idea (both on the server and\n> client side), I don't know of any servers where it is in active use.\n\nAt GitLab, we have started to roll out use of bundle URIs so that we can\npregenerate them and thus reduce load. The next step to evaluate in this\ncontext is whether we can easily reuse that infrastructure to eventually\nenable resumable clones via such bundle URIs. I assume that it cannot be\nthat hard to make this work.\n\nThat of course wouldn't be a perfect solution, as the clone can only be\nresumed as long as such a pregenerated bundle continues to exist on the\nserver. But it should still be way better compared to the status quo.\n\nPatrick\n"},{"id":"496787","messageId":"CAJoAoZkP58ZM4J3ejemyiqkkbEaQdphoyGj_LmX9-xb_eMgb4A@mail.gmail.com","threadId":"61605","inReplyTo":"20240608084323.GB2390433@coredump.intra.peff.net","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"Emily Shaffer","fromEmail":"nasamuffin@google.com","sentAt":"2024-06-10T19:04:30Z","receivedAt":"2024-06-10T19:04:42Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Sat, Jun 8, 2024 at 1:43 AM Jeff King <peff@peff.net> wrote:\n>\n> On Sat, Jun 08, 2024 at 02:46:38AM +0200, ellie wrote:\n>\n> > The deepening worked perfectly, thank you so much! I hope a resume will\n> > still be considered however, if even just to help out newcomers.\n>\n> Because the packfile to send the user is created on the fly, making a\n> clone fully resumable is tricky (a second clone may get an equivalent\n> but slightly different pack due to new objects entering the repo, or\n> even raciness between threads).\n>\n> One strategy people have worked on is for servers to point clients at\n> static packfiles (which _do_ remain byte-for-byte identical, and can be\n> resumed) to get some of the objects. But it requires some scheme on the\n> server side to decide when and how to create those packfiles. So while\n> there is support inside Git itself for this idea (both on the server and\n> client side), I don't know of any servers where it is in active use.\n\nWe use packfile offloading heavily at Google (any repositories hosted\nat *.googlesource.com, as well as our internal-facing hosting). It\nworks quite well for us scaling large projects like Android and\nChrome; we've been using it for some time now and are happy with it.\n\nHowever, one thing that's missing is the resumable download Ellie is\ndescribing. With a clone which has been turned into a packfile fetch\nfrom a different data store, it *should* be resumable. But the client\ncurrently lacks the ability to do that. (This just came up for us\ninternally the other day, and we ended up moving an internal bug to\nhttps://git.g-issues.gerritcodereview.com/issues/345241684.) After a\nresumed clone like this, you may not necessarily have latest - for\nexample, you may lose connection with 90% of the clone finished, then\nnot get connection back for some days, after which point upstream has\nmoved as Peff described elsewhere in this thread. But it would still\nprobably be cheaper to resume that 10% of packfile fetch from the\noffloaded data store, then do an incremental fetch back to the server\nto get the couple days of updates on top, as compared to starting over\nfrom zero with the server.\n\nIt seems to me that packfile URIs and bundle URIs are similar enough\nthat we could work out similar logic for both, no? Or maybe there's\nsomething I'm missing about the way bundle offloading differs from\npackfiles.\n\n - Emily\n\n>\n> -Peff\n>\n"},{"id":"496800","messageId":"xmqq5xug1qrf.fsf@gitster.g","threadId":"61605","inReplyTo":"CAJoAoZkP58ZM4J3ejemyiqkkbEaQdphoyGj_LmX9-xb_eMgb4A@mail.gmail.com","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2024-06-10T20:34:12Z","receivedAt":"2024-06-10T20:34:17Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Emily Shaffer <nasamuffin@google.com> writes:\n\n> It seems to me that packfile URIs and bundle URIs are similar enough\n> that we could work out similar logic for both, no? Or maybe there's\n> something I'm missing about the way bundle offloading differs from\n> packfiles.\n\nProbably we can deprecate one and let the other one take over?  It\nseems that bundleURI have plenty of documentation, but the only hit\nfor packfile URI side I find in the output of\n\n    $ git grep -i 'pack.*file.*uri' Documentation\n\nis the description of how the designed protocol extension is\nsupposed to work in Documentation/technical/packfile-uri.txt and not\neven the configuration variable uploadpack.blobPackfileURI that\ncontrols the \"experimental\" feature is documented.\n\nPerhaps whoever was adding the feature to the public side stopped\nafter pushing out the absolute minimum and lost interest or\nsomething?  We should update the documentation to reflect the\ncurrent status (e.g. is it still experimental? what more work do we\nneed on top of it to make it no longer experimental?), add at least\nminimum description for server operators how to configure it on the\nserver side, etc. (I am assuming that the end-user does not have to\ndo anything to get the feature, as long as their version of Git is\nrecent enough).\n\nThanks.\n\n"},{"id":"496808","messageId":"f5c24dfc-8d35-4418-b8f6-0a03d70c0917@horse64.org","threadId":"61605","inReplyTo":"xmqq5xug1qrf.fsf@gitster.g","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"ellie","fromEmail":"el@horse64.org","sentAt":"2024-06-10T21:55:31Z","receivedAt":"2024-06-10T21:55:40Z","isPatch":false,"sender":{"key":"el@horse64.org","avatar":null},"body":"Sorry for again another total newcomer/outsider question: Is a bundle or \npack file something any regular git HTTPS instance would naturally \nprovide when setup the usual ways? Like, if resume relied on that, would \nthis work when following the standard smart HTTP setup procedure \nhttps://git-scm.com/book/en/v2/Git-on-the-Server-Smart-HTTP (sorry if I \ngot the wrong link) and then git cloning from that? That would result in \nthe best availability of such a resume feature, if it ever came to be.\n\nRegards,\n\nEllie\n\nOn 6/10/24 10:34 PM, Junio C Hamano wrote:\n> Emily Shaffer <nasamuffin@google.com> writes:\n> \n>> It seems to me that packfile URIs and bundle URIs are similar enough\n>> that we could work out similar logic for both, no? Or maybe there's\n>> something I'm missing about the way bundle offloading differs from\n>> packfiles.\n> \n> Probably we can deprecate one and let the other one take over?  It\n> seems that bundleURI have plenty of documentation, but the only hit\n> for packfile URI side I find in the output of\n> \n>      $ git grep -i 'pack.*file.*uri' Documentation\n> \n> is the description of how the designed protocol extension is\n> supposed to work in Documentation/technical/packfile-uri.txt and not\n> even the configuration variable uploadpack.blobPackfileURI that\n> controls the \"experimental\" feature is documented.\n> \n> Perhaps whoever was adding the feature to the public side stopped\n> after pushing out the absolute minimum and lost interest or\n> something?  We should update the documentation to reflect the\n> current status (e.g. is it still experimental? what more work do we\n> need on top of it to make it no longer experimental?), add at least\n> minimum description for server operators how to configure it on the\n> server side, etc. (I am assuming that the end-user does not have to\n> do anything to get the feature, as long as their version of Git is\n> recent enough).\n> \n> Thanks.\n> \n"},{"id":"496825","messageId":"20240611062623.GA3248245@coredump.intra.peff.net","threadId":"61605","inReplyTo":"CAJoAoZkP58ZM4J3ejemyiqkkbEaQdphoyGj_LmX9-xb_eMgb4A@mail.gmail.com","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2024-06-11T06:26:23Z","receivedAt":"2024-06-11T06:26:31Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jun 10, 2024 at 12:04:30PM -0700, Emily Shaffer wrote:\n\n> > One strategy people have worked on is for servers to point clients at\n> > static packfiles (which _do_ remain byte-for-byte identical, and can be\n> > resumed) to get some of the objects. But it requires some scheme on the\n> > server side to decide when and how to create those packfiles. So while\n> > there is support inside Git itself for this idea (both on the server and\n> > client side), I don't know of any servers where it is in active use.\n> \n> We use packfile offloading heavily at Google (any repositories hosted\n> at *.googlesource.com, as well as our internal-facing hosting). It\n> works quite well for us scaling large projects like Android and\n> Chrome; we've been using it for some time now and are happy with it.\n\nCool! I'm glad to hear it is in use.\n\nIt might be helpful for other potential users if you can share how you\ndecide when to create the off-loaded packfiles, what goes in them, and\nso on. IIRC the server-side config is mostly geared at stuffing a few\nlarge blobs into a pack (since each blob must have an individual config\nkey). Maybe JGit (which I'm assuming is what powers googlesource) has\nbetter options there.\n\n> However, one thing that's missing is the resumable download Ellie is\n> describing. With a clone which has been turned into a packfile fetch\n> from a different data store, it *should* be resumable. But the client\n> currently lacks the ability to do that. (This just came up for us\n> internally the other day, and we ended up moving an internal bug to\n> https://git.g-issues.gerritcodereview.com/issues/345241684.) After a\n> resumed clone like this, you may not necessarily have latest - for\n> example, you may lose connection with 90% of the clone finished, then\n> not get connection back for some days, after which point upstream has\n> moved as Peff described elsewhere in this thread. But it would still\n> probably be cheaper to resume that 10% of packfile fetch from the\n> offloaded data store, then do an incremental fetch back to the server\n> to get the couple days of updates on top, as compared to starting over\n> from zero with the server.\n\nI do agree that resuming the offloaded parts, even if it is a few days\nlater, will generally be beneficial.\n\nFor packfile offloading, I think the server has to be aware of what's in\nthe packfiles (since it has to know not to send you those objects). So\nif you got all of the server's response packfile, but didn't finish the\noffloaded packfiles, it's a no-brainer to finish downloading them,\ncompleting your old clone. And then you can fetch on top of that to get\nfully up to date.\n\nBut if you didn't get all of the server's response, then you have to\ncontact it again. If it points you to the same offloaded packfile, you\ncan resume that transfer. But if it has moved on and doesn't advertise\nthat packfile anymore, I don't think it's useful.\n\nWhereas with bundleURI offloading, I think the client could always\nresume grabbing the bundle. Whatever it got is going to be useful\nbecause it will tell the server what it already has in the usual way\n(packfile offloads can't do that because the individual packfiles don't\nenforce the usual reachability guarantees).\n\n> It seems to me that packfile URIs and bundle URIs are similar enough\n> that we could work out similar logic for both, no? Or maybe there's\n> something I'm missing about the way bundle offloading differs from\n> packfiles.\n\nThey are pretty similar, but I think the resume strategy would be a\nlittle different, based on what I wrote above.\n\nIn general I don't think packfile-uris are that useful for resuming,\ncompared to bundle URIs.\n\n-Peff\n"},{"id":"496826","messageId":"20240611063123.GB3248245@coredump.intra.peff.net","threadId":"61605","inReplyTo":"xmqq5xug1qrf.fsf@gitster.g","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2024-06-11T06:31:23Z","receivedAt":"2024-06-11T06:31:25Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jun 10, 2024 at 01:34:12PM -0700, Junio C Hamano wrote:\n\n> Emily Shaffer <nasamuffin@google.com> writes:\n> \n> > It seems to me that packfile URIs and bundle URIs are similar enough\n> > that we could work out similar logic for both, no? Or maybe there's\n> > something I'm missing about the way bundle offloading differs from\n> > packfiles.\n> \n> Probably we can deprecate one and let the other one take over?  It\n> seems that bundleURI have plenty of documentation, but the only hit\n> for packfile URI side I find in the output of\n> \n>     $ git grep -i 'pack.*file.*uri' Documentation\n> \n> is the description of how the designed protocol extension is\n> supposed to work in Documentation/technical/packfile-uri.txt and not\n> even the configuration variable uploadpack.blobPackfileURI that\n> controls the \"experimental\" feature is documented.\n\nI think they serve two different purposes. A packfile URI does not have\nany connectivity guarantees. So it lets a server say \"here's all the\nobjects, except for XYZ which you should fetch from this URL\". That's\ngood for offloading pieces of a clone, like single large objects.\n\nWhereas bundle URIs require very little cooperation from the server.\nWhile a server can advertise bundle URIs, it doesn't need to know about\nthe particular bundle a client grabbed. The client comes back with the\nusual have/want, just like any other fetching client.\n\nAt least that's my understanding. I have to admit I didn't follow the\nrecent bundleURI work all that closely.\n\n-Peff\n"},{"id":"496926","messageId":"xmqqh6dzy0mr.fsf@gitster.g","threadId":"61605","inReplyTo":"20240611063123.GB3248245@coredump.intra.peff.net","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2024-06-11T15:12:12Z","receivedAt":"2024-06-11T15:12:15Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> I think they serve two different purposes. A packfile URI does not have\n> any connectivity guarantees. So it lets a server say \"here's all the\n> objects, except for XYZ which you should fetch from this URL\". That's\n> good for offloading pieces of a clone, like single large objects.\n>\n> Whereas bundle URIs require very little cooperation from the server.\n> While a server can advertise bundle URIs, it doesn't need to know about\n> the particular bundle a client grabbed. The client comes back with the\n> usual have/want, just like any other fetching client.\n\nYes, a bundle being a self-contained \"object-store + tips\", it is\na much more suitable building block for offloading clone traffic.\n"},{"id":"496958","messageId":"CANQMx9W9dCoxbzwP5c3CJjiw8Q7hcZgxrUxajt8EAkShvbdm8Q@mail.gmail.com","threadId":"61605","inReplyTo":"20240611062623.GA3248245@coredump.intra.peff.net","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"Ivan Frade","fromEmail":"ifrade@google.com","sentAt":"2024-06-11T19:40:14Z","receivedAt":"2024-06-11T19:40:28Z","isPatch":false,"sender":{"key":"ifrade@google.com","avatar":"https://avatars.githubusercontent.com/u/58185630?v=4"},"body":"On Mon, Jun 10, 2024 at 11:27 PM Jeff King <peff@peff.net> wrote:\n>\n> On Mon, Jun 10, 2024 at 12:04:30PM -0700, Emily Shaffer wrote:\n>\n> > > One strategy people have worked on is for servers to point clients at\n> > > static packfiles (which _do_ remain byte-for-byte identical, and can be\n> > > resumed) to get some of the objects. But it requires some scheme on the\n> > > server side to decide when and how to create those packfiles. So while\n> > > there is support inside Git itself for this idea (both on the server and\n> > > client side), I don't know of any servers where it is in active use.\n> >\n> > We use packfile offloading heavily at Google (any repositories hosted\n> > at *.googlesource.com, as well as our internal-facing hosting). It\n> > works quite well for us scaling large projects like Android and\n> > Chrome; we've been using it for some time now and are happy with it.\n>\n> Cool! I'm glad to hear it is in use.\n>\n> It might be helpful for other potential users if you can share how you\n> decide when to create the off-loaded packfiles, what goes in them, and\n> so on. IIRC the server-side config is mostly geared at stuffing a few\n> large blobs into a pack (since each blob must have an individual config\n> key). Maybe JGit (which I'm assuming is what powers googlesource) has\n> better options there.\n\nIIRC the upstream conf was oriented to offload individual blobs. In\nJGit/Google we do the offloading at pack level. We write to storage\nand CDN when creating a pack and keep the offloaded location in the\npack metadata. We do this only in certain conditions (GC, above a\ncertain size,...).\n\nAt serving time, if we see that we need to send a pack \"as-is\" (all\nobjects inside are needed) and it has an offload, then we mark it to\nsend the URL instead of the contents. As the offload is just a copy of\nthe pack, we can use the pack bitmap to know what is there or not.\n\n> > However, one thing that's missing is the resumable download Ellie is\n> > describing.\n\nAnother thing missing in the offload story is supporting offloads in\nnon-http protocols. e.g. after cloning via my-protocol://, being able\nto fetch my-protocol://blah/blah urls.\n\nIvan\n"},{"id":"497086","messageId":"87msnpjgqc.fsf@iotcl.com","threadId":"61605","inReplyTo":"f5c24dfc-8d35-4418-b8f6-0a03d70c0917@horse64.org","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"Toon claes","fromEmail":"toon@iotcl.com","sentAt":"2024-06-13T10:10:19Z","receivedAt":"2024-06-13T10:10:35Z","isPatch":false,"sender":{"key":"toon@iotcl.com","avatar":"https://avatars.githubusercontent.com/u/121621?v=4"},"body":"ellie <el@horse64.org> writes:\n\n> Sorry for again another total newcomer/outsider question:\n\nDon't apologize for asking these questions, you're more than welcome.\n\n> Is a bundle or pack file something any regular git HTTPS instance\n> would naturally provide when setup the usual ways?\n\nYes and no. Bundle and packfile format can used in many places.\nPackfiles are used to transfer a bunch of objects, or store them locally\nin Git's object database. A bundle is a packfile, but with a leading\nheader describing refs. You can read about that at\nhttps://git-scm.com/docs/gitformat-bundle.\n\n> Like, if resume relied on that, would this work when following the\n> standard smart HTTP setup procedure\n> https://git-scm.com/book/en/v2/Git-on-the-Server-Smart-HTTP (sorry if\n> I got the wrong link) and then git cloning from that? That would\n> result in the best availability of such a resume feature, if it ever\n> came to be.\n\nAs mentioned elsewhere in the thread, on clone (and fetch) the client\nnegotiates with the server which objects to download. Because the state\nof the remote repository can change between clones, so will the result\nof this negotiation. This means the content of the packfile sent over\nmight differ, which is disruptive for caching these files.\n\nThat's why the proposal of bundle URI or packfile URI is suggested. In\ncase of bundle URI, it will tell the client to download a pre-made\nbundle before starting the negotiation. This bundle can be stored on a\nCDN or whatever static HTTP(s) server. But it requires the server to\ncreate it, store it, and tell the client about it. This is not something\nthat's builtin into Git itself at the moment.\n\nThis is not really related to the Smart HTTP protocol, because it can be\nused over SSH as well. But when such file is stored on a regular HTTP\nserver, we can rely on resumable downloads. Only after that bundle is\ndownloaded, the client will start the negotiation with the server to get\nmissing objects and refs (which should be a small subset when the bundle\nis recent).\n\n\n--\nToon\n"},{"id":"497832","messageId":"Zn9o_UCjtf9MuwvH@sita-dell","threadId":"61605","inReplyTo":"xmqqh6dzy0mr.fsf@gitster.g","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"Sitaram Chamarty","fromEmail":"sitaramc@gmail.com","sentAt":"2024-06-29T01:53:01Z","receivedAt":"2024-06-29T01:53:07Z","isPatch":false,"sender":{"key":"sitaramc@gmail.com","avatar":"https://avatars.githubusercontent.com/u/43316?v=4"},"body":"On Tue, Jun 11, 2024 at 08:12:12AM -0700, Junio C Hamano wrote:\n> Jeff King <peff@peff.net> writes:\n> \n> > I think they serve two different purposes. A packfile URI does not have\n> > any connectivity guarantees. So it lets a server say \"here's all the\n> > objects, except for XYZ which you should fetch from this URL\". That's\n> > good for offloading pieces of a clone, like single large objects.\n> >\n> > Whereas bundle URIs require very little cooperation from the server.\n> > While a server can advertise bundle URIs, it doesn't need to know about\n> > the particular bundle a client grabbed. The client comes back with the\n> > usual have/want, just like any other fetching client.\n> \n> Yes, a bundle being a self-contained \"object-store + tips\", it is\n> a much more suitable building block for offloading clone traffic.\n\n[Adding mricon to cc]\n\nApologies for jumping in so late...\n\nGitolite supports this out of the box.  Just a couple of lines\nchange to the rc file and users can just run `rsync` (still\nmediated and access controlled by gitolite) to get a bundle.\nAdmittedly the first call by someone may take some time but it\n*is* resumable.\n\nSee [1] for details.\n\n[1]: https://github.com/sitaramc/gitolite/blob/master/src/commands/rsync\n"},{"id":"498204","messageId":"d3b3c9bb-fa2a-422d-99a7-4add5f98326e@horse64.org","threadId":"61605","inReplyTo":"200c3bd2-6aa9-4bb2-8eda-881bb62cd064@horse64.org","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"ellie","fromEmail":"el@horse64.org","sentAt":"2024-07-07T23:42:01Z","receivedAt":"2024-07-07T23:50:22Z","isPatch":false,"sender":{"key":"el@horse64.org","avatar":null},"body":"I have now encountered a repository where even --deepen=1 is bound to be \nfailing because it pulls in something fairly large that takes a few \nminutes. (Possibly, the server proxy has a faulty timeout setting that \npunishes slow connections, but for connections unreliable on the client \nside the problem would be the same.)\n\nSo this workaround sadly doesn't seem to cover all cases of resume.\n\nRegards,\n\nEllie\n\nOn 6/8/24 2:46 AM, ellie wrote:\n> The deepening worked perfectly, thank you so much! I hope a resume will \n> still be considered however, if even just to help out newcomers.\n> \n> Regards,\n> \n> Ellie\n> \n> On 6/8/24 2:35 AM, rsbecker@nexbridge.com wrote:\n>> On Friday, June 7, 2024 8:03 PM, ellie wrote:\n>>> Subject: Re: With big repos and slower connections, git clone can be \n>>> hard to work\n>>> with\n>>>\n>>> Thanks, this is very helpful as an emergency workaround!\n>>>\n>>> Nevertheless, I usually want the entire history, especially since I \n>>> wouldn't mind\n>>> waiting half an hour. But without resume, I've encountered it \n>>> regularly that it just\n>>> won't complete even if I give it the time, while way longer downloads \n>>> in the\n>>> browser would. The key problem here seems to be the lack of any resume.\n>>>\n>>> I hope this helps to understand why I made the suggestion.\n>>>\n>>> Regards,\n>>>\n>>> Ellie\n>>>\n>>> On 6/8/24 1:33 AM, rsbecker@nexbridge.com wrote:\n>>>> On Friday, June 7, 2024 7:28 PM, ellie wrote:\n>>>>> I'm terribly sorry if this is the wrong place, but I'd like to\n>>>>> suggest a potential issue with \"git clone\".\n>>>>>\n>>>>> The problem is that any sort of interruption or connection issue, no\n>>>>> matter how brief, causes the clone to stop and leave nothing behind:\n>>>>>\n>>>>> $ git clone https://github.com/Nheko-Reborn/nheko\n>>>>> Cloning into 'nheko'...\n>>>>> remote: Enumerating objects: 43991, done.\n>>>>> remote: Counting objects: 100% (6535/6535), done.\n>>>>> remote: Compressing objects: 100% (1449/1449), done.\n>>>>> error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\n>>>>> CANCEL (err 8)\n>>>>> error: 2771 bytes of body are still expected\n>>>>> fetch-pack: unexpected disconnect while reading sideband packet\n>>>>> fatal: early EOF\n>>>>> fatal: fetch-pack: invalid index-pack output $ cd nheko\n>>>>> bash: cd: nheko: No such file or director\n>>>>>\n>>>>> In my experience, this can be really impactful with 1. big \n>>>>> repositories and 2.\n>>>>> unreliable internet - which I would argue isn't unheard of! E.g.\n>>>>> a developer may work via mobile connection on a business trip. The\n>>>>> result can even be that a repository is uncloneable for some users!\n>>>>>\n>>>>> This has left me in the absurd situation where I was able to download\n>>>>> a tarball via HTTPS from the git hoster just fine, even way larger\n>>>>> binary release items, thanks to the browser's HTTPS resume. And yet a\n>>>>> simple git clone of the same project failed repeatedly.\n>>>>>\n>>>>> My deepest apologies if I missed an option to fix or address this.\n>>>>> But summed up, please consider making git clone recover from hiccups.\n>>>>>\n>>>>> Regards,\n>>>>>\n>>>>> Ellie\n>>>>>\n>>>>> PS: I've seen git hosters have apparent proxy bugs, like timing out\n>>>>> slower git clone connections from the server side even if the\n>>>>> transfer is ongoing. A git auto-resume would reduce the impact of \n>>>>> that, too.\n>>>>\n>>>> I suggest that you look into two git topics: --depth, which controls \n>>>> how much\n>>> history is obtained in a clone, and sparse-checkout, which describes \n>>> the part of the\n>>> repository you will retrieve. You can prune the contents of the \n>>> repository so that\n>>> clone is faster, if you do not need all of the history, or all of the \n>>> files. This is typically\n>>> done in complex large repositories, particularly those used for \n>>> production support\n>>> as release repositories.\n>>\n>> Consider doing the clone with --depth=1 then using git fetch --depth=n \n>> as the resume. There are other options that effectively give you a \n>> resume, including --deepen=n.\n>>\n>> Build automation, like Jenkins, uses this to speed up the clone/checkout.\n>>\n"},{"id":"498206","messageId":"0a7401dad0d6$10d27e20$32777a60$@nexbridge.com","threadId":"61605","inReplyTo":"d3b3c9bb-fa2a-422d-99a7-4add5f98326e@horse64.org","subject":"RE: With big repos and slower connections, git clone can be hard to work with","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2024-07-08T01:27:52Z","receivedAt":"2024-07-08T01:28:01Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On Sunday, July 7, 2024 7:42 PM, ellie wrote:\n>I have now encountered a repository where even --deepen=1 is bound to be failing\n>because it pulls in something fairly large that takes a few minutes. (Possibly, the\n>server proxy has a faulty timeout setting that punishes slow connections, but for\n>connections unreliable on the client side the problem would be the same.)\n>\n>So this workaround sadly doesn't seem to cover all cases of resume.\n>\n>Regards,\n>\n>Ellie\n>\n>On 6/8/24 2:46 AM, ellie wrote:\n>> The deepening worked perfectly, thank you so much! I hope a resume\n>> will still be considered however, if even just to help out newcomers.\n>>\n>> Regards,\n>>\n>> Ellie\n>>\n>> On 6/8/24 2:35 AM, rsbecker@nexbridge.com wrote:\n>>> On Friday, June 7, 2024 8:03 PM, ellie wrote:\n>>>> Subject: Re: With big repos and slower connections, git clone can be\n>>>> hard to work with\n>>>>\n>>>> Thanks, this is very helpful as an emergency workaround!\n>>>>\n>>>> Nevertheless, I usually want the entire history, especially since I\n>>>> wouldn't mind waiting half an hour. But without resume, I've\n>>>> encountered it regularly that it just won't complete even if I give\n>>>> it the time, while way longer downloads in the browser would. The\n>>>> key problem here seems to be the lack of any resume.\n>>>>\n>>>> I hope this helps to understand why I made the suggestion.\n>>>>\n>>>> Regards,\n>>>>\n>>>> Ellie\n>>>>\n>>>> On 6/8/24 1:33 AM, rsbecker@nexbridge.com wrote:\n>>>>> On Friday, June 7, 2024 7:28 PM, ellie wrote:\n>>>>>> I'm terribly sorry if this is the wrong place, but I'd like to\n>>>>>> suggest a potential issue with \"git clone\".\n>>>>>>\n>>>>>> The problem is that any sort of interruption or connection issue,\n>>>>>> no matter how brief, causes the clone to stop and leave nothing behind:\n>>>>>>\n>>>>>> $ git clone https://github.com/Nheko-Reborn/nheko\n>>>>>> Cloning into 'nheko'...\n>>>>>> remote: Enumerating objects: 43991, done.\n>>>>>> remote: Counting objects: 100% (6535/6535), done.\n>>>>>> remote: Compressing objects: 100% (1449/1449), done.\n>>>>>> error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\n>>>>>> CANCEL (err 8)\n>>>>>> error: 2771 bytes of body are still expected\n>>>>>> fetch-pack: unexpected disconnect while reading sideband packet\n>>>>>> fatal: early EOF\n>>>>>> fatal: fetch-pack: invalid index-pack output $ cd nheko\n>>>>>> bash: cd: nheko: No such file or director\n>>>>>>\n>>>>>> In my experience, this can be really impactful with 1. big\n>>>>>> repositories and 2.\n>>>>>> unreliable internet - which I would argue isn't unheard of! E.g.\n>>>>>> a developer may work via mobile connection on a business trip. The\n>>>>>> result can even be that a repository is uncloneable for some users!\n>>>>>>\n>>>>>> This has left me in the absurd situation where I was able to\n>>>>>> download a tarball via HTTPS from the git hoster just fine, even\n>>>>>> way larger binary release items, thanks to the browser's HTTPS\n>>>>>> resume. And yet a simple git clone of the same project failed repeatedly.\n>>>>>>\n>>>>>> My deepest apologies if I missed an option to fix or address this.\n>>>>>> But summed up, please consider making git clone recover from hiccups.\n>>>>>>\n>>>>>> Regards,\n>>>>>>\n>>>>>> Ellie\n>>>>>>\n>>>>>> PS: I've seen git hosters have apparent proxy bugs, like timing\n>>>>>> out slower git clone connections from the server side even if the\n>>>>>> transfer is ongoing. A git auto-resume would reduce the impact of\n>>>>>> that, too.\n>>>>>\n>>>>> I suggest that you look into two git topics: --depth, which\n>>>>> controls how much\n>>>> history is obtained in a clone, and sparse-checkout, which describes\n>>>> the part of the repository you will retrieve. You can prune the\n>>>> contents of the repository so that clone is faster, if you do not\n>>>> need all of the history, or all of the files. This is typically done\n>>>> in complex large repositories, particularly those used for\n>>>> production support as release repositories.\n>>>\n>>> Consider doing the clone with --depth=1 then using git fetch\n>>> --depth=n as the resume. There are other options that effectively\n>>> give you a resume, including --deepen=n.\n>>>\n>>> Build automation, like Jenkins, uses this to speed up the clone/checkout.\n\nCan you please provide more details on this? It is difficult to understand your issue without knowing what situation is failing? What size file? Is this a large single pack file? Can you reproduce this with a script we can try?\n\n"},{"id":"498207","messageId":"15bb8955-8ef6-4d83-b10c-e8593f65790c@horse64.org","threadId":"61605","inReplyTo":"0a7401dad0d6$10d27e20$32777a60$@nexbridge.com","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"ellie","fromEmail":"el@horse64.org","sentAt":"2024-07-08T02:28:25Z","receivedAt":"2024-07-08T02:28:38Z","isPatch":false,"sender":{"key":"el@horse64.org","avatar":null},"body":"I was intending to suggest that depending on the largest object in the \nrepository, resume may remain a concern for lower end users. My \napologies for being unclear.\n\nAs for my concrete problem, I can only guess what's happening, maybe \ngithub's HTTPS proxy too eagerly discarding slow connections:\n\n$ git clone https://github.com/maliit/keyboard maliit-keyboard\nCloning into 'maliit-keyboard'...\nremote: Enumerating objects: 23243, done.\nremote: Counting objects: 100% (464/464), done.\nremote: Compressing objects: 100% (207/207), done.\nerror: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly: \nCANCEL (err 8)\nerror: 2507 bytes of body are still expected\nfetch-pack: unexpected disconnect while reading sideband packet\nfatal: early EOF\nfatal: fetch-pack: invalid index-pack output\n\nA deepen seems to fail for this repo since one deepen step already gets \nkilled off. Git HTTPS clones from any other hoster I tried, including \ngitlab.com, work fine, as do git SSH clones from github.com.\n\nSorry for the long tangent. Basically, my point was just that resume \nstill seems like a good idea even with deepen existing.\n\nRegards,\n\nEllie\n\nOn 7/8/24 3:27 AM, rsbecker@nexbridge.com wrote:\n> On Sunday, July 7, 2024 7:42 PM, ellie wrote:\n>> I have now encountered a repository where even --deepen=1 is bound to be failing\n>> because it pulls in something fairly large that takes a few minutes. (Possibly, the\n>> server proxy has a faulty timeout setting that punishes slow connections, but for\n>> connections unreliable on the client side the problem would be the same.)\n>>\n>> So this workaround sadly doesn't seem to cover all cases of resume.\n>>\n>> Regards,\n>>\n>> Ellie\n>>\n>> On 6/8/24 2:46 AM, ellie wrote:\n>>> The deepening worked perfectly, thank you so much! I hope a resume\n>>> will still be considered however, if even just to help out newcomers.\n>>>\n>>> Regards,\n>>>\n>>> Ellie\n>>>\n>>> On 6/8/24 2:35 AM, rsbecker@nexbridge.com wrote:\n>>>> On Friday, June 7, 2024 8:03 PM, ellie wrote:\n>>>>> Subject: Re: With big repos and slower connections, git clone can be\n>>>>> hard to work with\n>>>>>\n>>>>> Thanks, this is very helpful as an emergency workaround!\n>>>>>\n>>>>> Nevertheless, I usually want the entire history, especially since I\n>>>>> wouldn't mind waiting half an hour. But without resume, I've\n>>>>> encountered it regularly that it just won't complete even if I give\n>>>>> it the time, while way longer downloads in the browser would. The\n>>>>> key problem here seems to be the lack of any resume.\n>>>>>\n>>>>> I hope this helps to understand why I made the suggestion.\n>>>>>\n>>>>> Regards,\n>>>>>\n>>>>> Ellie\n>>>>>\n>>>>> On 6/8/24 1:33 AM, rsbecker@nexbridge.com wrote:\n>>>>>> On Friday, June 7, 2024 7:28 PM, ellie wrote:\n>>>>>>> I'm terribly sorry if this is the wrong place, but I'd like to\n>>>>>>> suggest a potential issue with \"git clone\".\n>>>>>>>\n>>>>>>> The problem is that any sort of interruption or connection issue,\n>>>>>>> no matter how brief, causes the clone to stop and leave nothing behind:\n>>>>>>>\n>>>>>>> $ git clone https://github.com/Nheko-Reborn/nheko\n>>>>>>> Cloning into 'nheko'...\n>>>>>>> remote: Enumerating objects: 43991, done.\n>>>>>>> remote: Counting objects: 100% (6535/6535), done.\n>>>>>>> remote: Compressing objects: 100% (1449/1449), done.\n>>>>>>> error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\n>>>>>>> CANCEL (err 8)\n>>>>>>> error: 2771 bytes of body are still expected\n>>>>>>> fetch-pack: unexpected disconnect while reading sideband packet\n>>>>>>> fatal: early EOF\n>>>>>>> fatal: fetch-pack: invalid index-pack output $ cd nheko\n>>>>>>> bash: cd: nheko: No such file or director\n>>>>>>>\n>>>>>>> In my experience, this can be really impactful with 1. big\n>>>>>>> repositories and 2.\n>>>>>>> unreliable internet - which I would argue isn't unheard of! E.g.\n>>>>>>> a developer may work via mobile connection on a business trip. The\n>>>>>>> result can even be that a repository is uncloneable for some users!\n>>>>>>>\n>>>>>>> This has left me in the absurd situation where I was able to\n>>>>>>> download a tarball via HTTPS from the git hoster just fine, even\n>>>>>>> way larger binary release items, thanks to the browser's HTTPS\n>>>>>>> resume. And yet a simple git clone of the same project failed repeatedly.\n>>>>>>>\n>>>>>>> My deepest apologies if I missed an option to fix or address this.\n>>>>>>> But summed up, please consider making git clone recover from hiccups.\n>>>>>>>\n>>>>>>> Regards,\n>>>>>>>\n>>>>>>> Ellie\n>>>>>>>\n>>>>>>> PS: I've seen git hosters have apparent proxy bugs, like timing\n>>>>>>> out slower git clone connections from the server side even if the\n>>>>>>> transfer is ongoing. A git auto-resume would reduce the impact of\n>>>>>>> that, too.\n>>>>>>\n>>>>>> I suggest that you look into two git topics: --depth, which\n>>>>>> controls how much\n>>>>> history is obtained in a clone, and sparse-checkout, which describes\n>>>>> the part of the repository you will retrieve. You can prune the\n>>>>> contents of the repository so that clone is faster, if you do not\n>>>>> need all of the history, or all of the files. This is typically done\n>>>>> in complex large repositories, particularly those used for\n>>>>> production support as release repositories.\n>>>>\n>>>> Consider doing the clone with --depth=1 then using git fetch\n>>>> --depth=n as the resume. There are other options that effectively\n>>>> give you a resume, including --deepen=n.\n>>>>\n>>>> Build automation, like Jenkins, uses this to speed up the clone/checkout.\n> \n> Can you please provide more details on this? It is difficult to understand your issue without knowing what situation is failing? What size file? Is this a large single pack file? Can you reproduce this with a script we can try?\n> \n"},{"id":"498240","messageId":"000001dad132$a9339170$fb9ab450$@nexbridge.com","threadId":"61605","inReplyTo":"15bb8955-8ef6-4d83-b10c-e8593f65790c@horse64.org","subject":"RE: With big repos and slower connections, git clone can be hard to work with","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2024-07-08T12:30:41Z","receivedAt":"2024-07-08T12:30:52Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On Sunday, July 7, 2024 10:28 PM, ellie wrote:\n>I was intending to suggest that depending on the largest object in the repository,\n>resume may remain a concern for lower end users. My apologies for being unclear.\n>\n>As for my concrete problem, I can only guess what's happening, maybe github's\n>HTTPS proxy too eagerly discarding slow connections:\n>\n>$ git clone https://github.com/maliit/keyboard maliit-keyboard Cloning into 'maliit-\n>keyboard'...\n>remote: Enumerating objects: 23243, done.\n>remote: Counting objects: 100% (464/464), done.\n>remote: Compressing objects: 100% (207/207), done.\n>error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\n>CANCEL (err 8)\n>error: 2507 bytes of body are still expected\n>fetch-pack: unexpected disconnect while reading sideband packet\n>fatal: early EOF\n>fatal: fetch-pack: invalid index-pack output\n>\n>A deepen seems to fail for this repo since one deepen step already gets killed off. Git\n>HTTPS clones from any other hoster I tried, including gitlab.com, work fine, as do git\n>SSH clones from github.com.\n>\n>Sorry for the long tangent. Basically, my point was just that resume still seems like a\n>good idea even with deepen existing.\n>\n>Regards,\n>\n>Ellie\n>\n>On 7/8/24 3:27 AM, rsbecker@nexbridge.com wrote:\n>> On Sunday, July 7, 2024 7:42 PM, ellie wrote:\n>>> I have now encountered a repository where even --deepen=1 is bound to\n>>> be failing because it pulls in something fairly large that takes a\n>>> few minutes. (Possibly, the server proxy has a faulty timeout setting\n>>> that punishes slow connections, but for connections unreliable on the\n>>> client side the problem would be the same.)\n>>>\n>>> So this workaround sadly doesn't seem to cover all cases of resume.\n>>>\n>>> Regards,\n>>>\n>>> Ellie\n>>>\n>>> On 6/8/24 2:46 AM, ellie wrote:\n>>>> The deepening worked perfectly, thank you so much! I hope a resume\n>>>> will still be considered however, if even just to help out newcomers.\n>>>>\n>>>> Regards,\n>>>>\n>>>> Ellie\n>>>>\n>>>> On 6/8/24 2:35 AM, rsbecker@nexbridge.com wrote:\n>>>>> On Friday, June 7, 2024 8:03 PM, ellie wrote:\n>>>>>> Subject: Re: With big repos and slower connections, git clone can\n>>>>>> be hard to work with\n>>>>>>\n>>>>>> Thanks, this is very helpful as an emergency workaround!\n>>>>>>\n>>>>>> Nevertheless, I usually want the entire history, especially since\n>>>>>> I wouldn't mind waiting half an hour. But without resume, I've\n>>>>>> encountered it regularly that it just won't complete even if I\n>>>>>> give it the time, while way longer downloads in the browser would.\n>>>>>> The key problem here seems to be the lack of any resume.\n>>>>>>\n>>>>>> I hope this helps to understand why I made the suggestion.\n>>>>>>\n>>>>>> Regards,\n>>>>>>\n>>>>>> Ellie\n>>>>>>\n>>>>>> On 6/8/24 1:33 AM, rsbecker@nexbridge.com wrote:\n>>>>>>> On Friday, June 7, 2024 7:28 PM, ellie wrote:\n>>>>>>>> I'm terribly sorry if this is the wrong place, but I'd like to\n>>>>>>>> suggest a potential issue with \"git clone\".\n>>>>>>>>\n>>>>>>>> The problem is that any sort of interruption or connection\n>>>>>>>> issue, no matter how brief, causes the clone to stop and leave nothing\n>behind:\n>>>>>>>>\n>>>>>>>> $ git clone https://github.com/Nheko-Reborn/nheko\n>>>>>>>> Cloning into 'nheko'...\n>>>>>>>> remote: Enumerating objects: 43991, done.\n>>>>>>>> remote: Counting objects: 100% (6535/6535), done.\n>>>>>>>> remote: Compressing objects: 100% (1449/1449), done.\n>>>>>>>> error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\n>>>>>>>> CANCEL (err 8)\n>>>>>>>> error: 2771 bytes of body are still expected\n>>>>>>>> fetch-pack: unexpected disconnect while reading sideband packet\n>>>>>>>> fatal: early EOF\n>>>>>>>> fatal: fetch-pack: invalid index-pack output $ cd nheko\n>>>>>>>> bash: cd: nheko: No such file or director\n>>>>>>>>\n>>>>>>>> In my experience, this can be really impactful with 1. big\n>>>>>>>> repositories and 2.\n>>>>>>>> unreliable internet - which I would argue isn't unheard of! E.g.\n>>>>>>>> a developer may work via mobile connection on a business trip.\n>>>>>>>> The result can even be that a repository is uncloneable for some users!\n>>>>>>>>\n>>>>>>>> This has left me in the absurd situation where I was able to\n>>>>>>>> download a tarball via HTTPS from the git hoster just fine, even\n>>>>>>>> way larger binary release items, thanks to the browser's HTTPS\n>>>>>>>> resume. And yet a simple git clone of the same project failed repeatedly.\n>>>>>>>>\n>>>>>>>> My deepest apologies if I missed an option to fix or address this.\n>>>>>>>> But summed up, please consider making git clone recover from hiccups.\n>>>>>>>>\n>>>>>>>> Regards,\n>>>>>>>>\n>>>>>>>> Ellie\n>>>>>>>>\n>>>>>>>> PS: I've seen git hosters have apparent proxy bugs, like timing\n>>>>>>>> out slower git clone connections from the server side even if\n>>>>>>>> the transfer is ongoing. A git auto-resume would reduce the\n>>>>>>>> impact of that, too.\n>>>>>>>\n>>>>>>> I suggest that you look into two git topics: --depth, which\n>>>>>>> controls how much\n>>>>>> history is obtained in a clone, and sparse-checkout, which\n>>>>>> describes the part of the repository you will retrieve. You can\n>>>>>> prune the contents of the repository so that clone is faster, if\n>>>>>> you do not need all of the history, or all of the files. This is\n>>>>>> typically done in complex large repositories, particularly those\n>>>>>> used for production support as release repositories.\n>>>>>\n>>>>> Consider doing the clone with --depth=1 then using git fetch\n>>>>> --depth=n as the resume. There are other options that effectively\n>>>>> give you a resume, including --deepen=n.\n>>>>>\n>>>>> Build automation, like Jenkins, uses this to speed up the clone/checkout.\n>>\n>> Can you please provide more details on this? It is difficult to understand your issue\n>without knowing what situation is failing? What size file? Is this a large single pack\n>file? Can you reproduce this with a script we can try?\n>>\n\nFirst, for this mailing list, please put your replies at the bottom.\n\nSecond, the full clone takes under 5 seconds on my system and does not experience any error that you are seeing. I suggest that your ISP may be throttling your account. I have seen this happen on some ISPs under SSH but few under HTTPS. It is likely a firewall or as you said, a proxy setting. GitHub has no proxy.\n\nMy suggestion is that this is more of a communication issue instead of than a large repo issue. 133Mb is a relatively small a repository and clones quickly. This might be something to take up on the GitHub support forums rather that for git - since it seems like something in the path outside of git is not working correctly. None of the files in this repository, including pack-files is larger than 100 blocks, so there is not much point with a mid-pack restart.\n\n"},{"id":"498241","messageId":"793a0c16-c2e5-4fbb-9e97-297c096fe42f@horse64.org","threadId":"61605","inReplyTo":"000001dad132$a9339170$fb9ab450$@nexbridge.com","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"ellie","fromEmail":"el@horse64.org","sentAt":"2024-07-08T12:41:48Z","receivedAt":"2024-07-08T12:41:53Z","isPatch":false,"sender":{"key":"el@horse64.org","avatar":null},"body":"\n\nOn 7/8/24 2:30 PM, rsbecker@nexbridge.com wrote:\n> On Sunday, July 7, 2024 10:28 PM, ellie wrote:\n>> I was intending to suggest that depending on the largest object in the repository,\n>> resume may remain a concern for lower end users. My apologies for being unclear.\n>>\n>> As for my concrete problem, I can only guess what's happening, maybe github's\n>> HTTPS proxy too eagerly discarding slow connections:\n>>\n>> $ git clone https://github.com/maliit/keyboard maliit-keyboard Cloning into 'maliit-\n>> keyboard'...\n>> remote: Enumerating objects: 23243, done.\n>> remote: Counting objects: 100% (464/464), done.\n>> remote: Compressing objects: 100% (207/207), done.\n>> error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\n>> CANCEL (err 8)\n>> error: 2507 bytes of body are still expected\n>> fetch-pack: unexpected disconnect while reading sideband packet\n>> fatal: early EOF\n>> fatal: fetch-pack: invalid index-pack output\n>>\n>> A deepen seems to fail for this repo since one deepen step already gets killed off. Git\n>> HTTPS clones from any other hoster I tried, including gitlab.com, work fine, as do git\n>> SSH clones from github.com.\n>>\n>> Sorry for the long tangent. Basically, my point was just that resume still seems like a\n>> good idea even with deepen existing.\n>>\n>> Regards,\n>>\n>> Ellie\n>>\n>> On 7/8/24 3:27 AM, rsbecker@nexbridge.com wrote:\n>>> On Sunday, July 7, 2024 7:42 PM, ellie wrote:\n>>>> I have now encountered a repository where even --deepen=1 is bound to\n>>>> be failing because it pulls in something fairly large that takes a\n>>>> few minutes. (Possibly, the server proxy has a faulty timeout setting\n>>>> that punishes slow connections, but for connections unreliable on the\n>>>> client side the problem would be the same.)\n>>>>\n>>>> So this workaround sadly doesn't seem to cover all cases of resume.\n>>>>\n>>>> Regards,\n>>>>\n>>>> Ellie\n>>>>\n>>>> On 6/8/24 2:46 AM, ellie wrote:\n>>>>> The deepening worked perfectly, thank you so much! I hope a resume\n>>>>> will still be considered however, if even just to help out newcomers.\n>>>>>\n>>>>> Regards,\n>>>>>\n>>>>> Ellie\n>>>>>\n>>>>> On 6/8/24 2:35 AM, rsbecker@nexbridge.com wrote:\n>>>>>> On Friday, June 7, 2024 8:03 PM, ellie wrote:\n>>>>>>> Subject: Re: With big repos and slower connections, git clone can\n>>>>>>> be hard to work with\n>>>>>>>\n>>>>>>> Thanks, this is very helpful as an emergency workaround!\n>>>>>>>\n>>>>>>> Nevertheless, I usually want the entire history, especially since\n>>>>>>> I wouldn't mind waiting half an hour. But without resume, I've\n>>>>>>> encountered it regularly that it just won't complete even if I\n>>>>>>> give it the time, while way longer downloads in the browser would.\n>>>>>>> The key problem here seems to be the lack of any resume.\n>>>>>>>\n>>>>>>> I hope this helps to understand why I made the suggestion.\n>>>>>>>\n>>>>>>> Regards,\n>>>>>>>\n>>>>>>> Ellie\n>>>>>>>\n>>>>>>> On 6/8/24 1:33 AM, rsbecker@nexbridge.com wrote:\n>>>>>>>> On Friday, June 7, 2024 7:28 PM, ellie wrote:\n>>>>>>>>> I'm terribly sorry if this is the wrong place, but I'd like to\n>>>>>>>>> suggest a potential issue with \"git clone\".\n>>>>>>>>>\n>>>>>>>>> The problem is that any sort of interruption or connection\n>>>>>>>>> issue, no matter how brief, causes the clone to stop and leave nothing\n>> behind:\n>>>>>>>>>\n>>>>>>>>> $ git clone https://github.com/Nheko-Reborn/nheko\n>>>>>>>>> Cloning into 'nheko'...\n>>>>>>>>> remote: Enumerating objects: 43991, done.\n>>>>>>>>> remote: Counting objects: 100% (6535/6535), done.\n>>>>>>>>> remote: Compressing objects: 100% (1449/1449), done.\n>>>>>>>>> error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\n>>>>>>>>> CANCEL (err 8)\n>>>>>>>>> error: 2771 bytes of body are still expected\n>>>>>>>>> fetch-pack: unexpected disconnect while reading sideband packet\n>>>>>>>>> fatal: early EOF\n>>>>>>>>> fatal: fetch-pack: invalid index-pack output $ cd nheko\n>>>>>>>>> bash: cd: nheko: No such file or director\n>>>>>>>>>\n>>>>>>>>> In my experience, this can be really impactful with 1. big\n>>>>>>>>> repositories and 2.\n>>>>>>>>> unreliable internet - which I would argue isn't unheard of! E.g.\n>>>>>>>>> a developer may work via mobile connection on a business trip.\n>>>>>>>>> The result can even be that a repository is uncloneable for some users!\n>>>>>>>>>\n>>>>>>>>> This has left me in the absurd situation where I was able to\n>>>>>>>>> download a tarball via HTTPS from the git hoster just fine, even\n>>>>>>>>> way larger binary release items, thanks to the browser's HTTPS\n>>>>>>>>> resume. And yet a simple git clone of the same project failed repeatedly.\n>>>>>>>>>\n>>>>>>>>> My deepest apologies if I missed an option to fix or address this.\n>>>>>>>>> But summed up, please consider making git clone recover from hiccups.\n>>>>>>>>>\n>>>>>>>>> Regards,\n>>>>>>>>>\n>>>>>>>>> Ellie\n>>>>>>>>>\n>>>>>>>>> PS: I've seen git hosters have apparent proxy bugs, like timing\n>>>>>>>>> out slower git clone connections from the server side even if\n>>>>>>>>> the transfer is ongoing. A git auto-resume would reduce the\n>>>>>>>>> impact of that, too.\n>>>>>>>>\n>>>>>>>> I suggest that you look into two git topics: --depth, which\n>>>>>>>> controls how much\n>>>>>>> history is obtained in a clone, and sparse-checkout, which\n>>>>>>> describes the part of the repository you will retrieve. You can\n>>>>>>> prune the contents of the repository so that clone is faster, if\n>>>>>>> you do not need all of the history, or all of the files. This is\n>>>>>>> typically done in complex large repositories, particularly those\n>>>>>>> used for production support as release repositories.\n>>>>>>\n>>>>>> Consider doing the clone with --depth=1 then using git fetch\n>>>>>> --depth=n as the resume. There are other options that effectively\n>>>>>> give you a resume, including --deepen=n.\n>>>>>>\n>>>>>> Build automation, like Jenkins, uses this to speed up the clone/checkout.\n>>>\n>>> Can you please provide more details on this? It is difficult to understand your issue\n>> without knowing what situation is failing? What size file? Is this a large single pack\n>> file? Can you reproduce this with a script we can try?\n>>>\n> \n> First, for this mailing list, please put your replies at the bottom.\n> \n> Second, the full clone takes under 5 seconds on my system and does not experience any error that you are seeing. I suggest that your ISP may be throttling your account. I have seen this happen on some ISPs under SSH but few under HTTPS. It is likely a firewall or as you said, a proxy setting. GitHub has no proxy.\n> \n> My suggestion is that this is more of a communication issue instead of than a large repo issue. 133Mb is a relatively small a repository and clones quickly. This might be something to take up on the GitHub support forums rather that for git - since it seems like something in the path outside of git is not working correctly. None of the files in this repository, including pack-files is larger than 100 blocks, so there is not much point with a mid-pack restart.\n> \n\nI apologize for not placing the responses where expected.\n\nIt seems extremely unlikely to me to be possibly an ISP issue, for which \nI already listed the reasons. An additional one is HTTPS downloads from \ngithub outside of git, e.g. from zip archives, for way larger files work \nfine as well.\n\nNevertheless, this irrelevant to my initial request. Since even if it's \nnot caused by a Github server side issue, a resume would still help.\n\nRegards,\n\nEllie\n"},{"id":"498265","messageId":"001201dad147$e9fdf9b0$bdf9ed10$@nexbridge.com","threadId":"61605","inReplyTo":"20240708143239.vq47dg7mgh33hykf@carbon","subject":"RE: With big repos and slower connections, git clone can be hard to work with","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2024-07-08T15:02:50Z","receivedAt":"2024-07-08T15:03:03Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On Monday, July 8, 2024 10:33 AM, Konstantin Khomoutov wrote:\n>On Mon, Jul 08, 2024 at 04:28:25AM +0200, ellie wrote:\n>\n>[...]\n>> error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\n>> CANCEL (err 8)\n>[...]\n>> It seems extremely unlikely to me to be possibly an ISP issue, for\n>> which I already listed the reasons. An additional one is HTTPS\n>> downloads from github outside of git, e.g. from zip archives, for way\n>> larger files work fine as well.\n>[...]\n>\n>What if you explicitly disable HTTP/2 when cloning?\n>\n>  git -c http.version=HTTP/1.1 clone ...\n>\n>should probably do this.\n\nI can verify that this works in my environment.\n\n"},{"id":"498267","messageId":"20240708143239.vq47dg7mgh33hykf@carbon","threadId":"61605","inReplyTo":"15bb8955-8ef6-4d83-b10c-e8593f65790c@horse64.org","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"Konstantin Khomoutov","fromEmail":"kostix@bswap.ru","sentAt":"2024-07-08T14:32:39Z","receivedAt":"2024-07-08T15:14:14Z","isPatch":false,"sender":{"key":"kostix@bswap.ru","avatar":null},"body":"On Mon, Jul 08, 2024 at 04:28:25AM +0200, ellie wrote:\n\n[...]\n> error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly: CANCEL\n> (err 8)\n[...]\n> It seems extremely unlikely to me to be possibly an ISP issue, for which I\n> already listed the reasons. An additional one is HTTPS downloads from github\n> outside of git, e.g. from zip archives, for way larger files work fine as\n> well.\n[...]\n\nWhat if you explicitly disable HTTP/2 when cloning?\n\n  git -c http.version=HTTP/1.1 clone ...\n\nshould probably do this.\n\n"},{"id":"498268","messageId":"2e10070f-2720-4d70-aa15-d4c008cc57bf@horse64.org","threadId":"61605","inReplyTo":"20240708143239.vq47dg7mgh33hykf@carbon","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"ellie","fromEmail":"el@horse64.org","sentAt":"2024-07-08T15:14:33Z","receivedAt":"2024-07-08T15:14:38Z","isPatch":false,"sender":{"key":"el@horse64.org","avatar":null},"body":"\n\nOn 7/8/24 4:32 PM, Konstantin Khomoutov wrote:\n> On Mon, Jul 08, 2024 at 04:28:25AM +0200, ellie wrote:\n> \n> [...]\n>> error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly: CANCEL\n>> (err 8)\n> [...]\n>> It seems extremely unlikely to me to be possibly an ISP issue, for which I\n>> already listed the reasons. An additional one is HTTPS downloads from github\n>> outside of git, e.g. from zip archives, for way larger files work fine as\n>> well.\n> [...]\n> \n> What if you explicitly disable HTTP/2 when cloning?\n> \n>    git -c http.version=HTTP/1.1 clone ...\n> \n> should probably do this.\n> \n\nThanks for the idea! I tested it:\n\n$  git -c http.version=HTTP/1.1 clone https://github.com/maliit/keyboard \nmaliit-keyboard\nCloning into 'maliit-keyboard'...\nremote: Enumerating objects: 23243, done.\nremote: Counting objects: 100% (464/464), done.\nremote: Compressing objects: 100% (207/207), done.\nerror: RPC failed; curl 18 transfer closed with outstanding read data \nremaining\nerror: 5361 bytes of body are still expected\nfetch-pack: unexpected disconnect while reading sideband packet\nfatal: early EOF\nfatal: fetch-pack: invalid index-pack output\n\nSadly, it seems like the error is only slightly different. It was still \nworth a try. I contacted GitHub support a while ago but it got stuck. If \nthere were resume available such hiccups wouldn't matter, I hope that \nexplains why I suggested that feature.\n\nRegards,\n\nEllie\n"},{"id":"498272","messageId":"001301dad14b$f8f0e460$ead2ad20$@nexbridge.com","threadId":"61605","inReplyTo":"2e10070f-2720-4d70-aa15-d4c008cc57bf@horse64.org","subject":"RE: With big repos and slower connections, git clone can be hard to work with","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2024-07-08T15:31:53Z","receivedAt":"2024-07-08T15:32:09Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On Monday, July 8, 2024 11:15 AM, ellie wrote:\n>On 7/8/24 4:32 PM, Konstantin Khomoutov wrote:\n>> On Mon, Jul 08, 2024 at 04:28:25AM +0200, ellie wrote:\n>>\n>> [...]\n>>> error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\n>>> CANCEL (err 8)\n>> [...]\n>>> It seems extremely unlikely to me to be possibly an ISP issue, for\n>>> which I already listed the reasons. An additional one is HTTPS\n>>> downloads from github outside of git, e.g. from zip archives, for way\n>>> larger files work fine as well.\n>> [...]\n>>\n>> What if you explicitly disable HTTP/2 when cloning?\n>>\n>>    git -c http.version=HTTP/1.1 clone ...\n>>\n>> should probably do this.\n>>\n>\n>Thanks for the idea! I tested it:\n>\n>$  git -c http.version=HTTP/1.1 clone https://github.com/maliit/keyboard\n>maliit-keyboard\n>Cloning into 'maliit-keyboard'...\n>remote: Enumerating objects: 23243, done.\n>remote: Counting objects: 100% (464/464), done.\n>remote: Compressing objects: 100% (207/207), done.\n>error: RPC failed; curl 18 transfer closed with outstanding read data remaining\n>error: 5361 bytes of body are still expected\n>fetch-pack: unexpected disconnect while reading sideband packet\n>fatal: early EOF\n>fatal: fetch-pack: invalid index-pack output\n>\n>Sadly, it seems like the error is only slightly different. It was still worth a try. I\n>contacted GitHub support a while ago but it got stuck. If there were resume\n>available such hiccups wouldn't matter, I hope that explains why I suggested that\n>feature.\n\nI don't really understand what \"it got stuck\" means. Is that a colloquialism? What got stuck? That case at GitHub?\n\nHave you tried git config --global http.postBuffer 524288000\n\nIt might help. The feature being requesting, even if possible, will probably not happen quickly, unless someone has a solid and simple design for this. That is why we are trying to figure out the root cause of your situation, which is not clear to me as to what exactly is failing (possibly a buffer size issue, if this is consistently failing). My experience, as I said before, on these symptoms, is a proxy (even a local one) that is in the way. If you have your linux instance on a VM, the hypervisor may not be configured correctly. Lack of further evidence (all we really have is the curl RPC failure) makes diagnosing this very difficult.\n\n"},{"id":"498275","messageId":"47799635-7832-4c89-b4d3-e992d49ad40c@horse64.org","threadId":"61605","inReplyTo":"001301dad14b$f8f0e460$ead2ad20$@nexbridge.com","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"ellie","fromEmail":"el@horse64.org","sentAt":"2024-07-08T15:48:56Z","receivedAt":"2024-07-08T15:49:01Z","isPatch":false,"sender":{"key":"el@horse64.org","avatar":null},"body":"\n\nOn 7/8/24 5:31 PM, rsbecker@nexbridge.com wrote:\n> On Monday, July 8, 2024 11:15 AM, ellie wrote:\n>> On 7/8/24 4:32 PM, Konstantin Khomoutov wrote:\n>>> On Mon, Jul 08, 2024 at 04:28:25AM +0200, ellie wrote:\n>>>\n>>> [...]\n>>>> error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\n>>>> CANCEL (err 8)\n>>> [...]\n>>>> It seems extremely unlikely to me to be possibly an ISP issue, for\n>>>> which I already listed the reasons. An additional one is HTTPS\n>>>> downloads from github outside of git, e.g. from zip archives, for way\n>>>> larger files work fine as well.\n>>> [...]\n>>>\n>>> What if you explicitly disable HTTP/2 when cloning?\n>>>\n>>>     git -c http.version=HTTP/1.1 clone ...\n>>>\n>>> should probably do this.\n>>>\n>>\n>> Thanks for the idea! I tested it:\n>>\n>> $  git -c http.version=HTTP/1.1 clone https://github.com/maliit/keyboard\n>> maliit-keyboard\n>> Cloning into 'maliit-keyboard'...\n>> remote: Enumerating objects: 23243, done.\n>> remote: Counting objects: 100% (464/464), done.\n>> remote: Compressing objects: 100% (207/207), done.\n>> error: RPC failed; curl 18 transfer closed with outstanding read data remaining\n>> error: 5361 bytes of body are still expected\n>> fetch-pack: unexpected disconnect while reading sideband packet\n>> fatal: early EOF\n>> fatal: fetch-pack: invalid index-pack output\n>>\n>> Sadly, it seems like the error is only slightly different. It was still worth a try. I\n>> contacted GitHub support a while ago but it got stuck. If there were resume\n>> available such hiccups wouldn't matter, I hope that explains why I suggested that\n>> feature.\n> \n> I don't really understand what \"it got stuck\" means. Is that a colloquialism? What got stuck? That case at GitHub?\n> \n> Have you tried git config --global http.postBuffer 524288000\n> \n> It might help. The feature being requesting, even if possible, will probably not happen quickly, unless someone has a solid and simple design for this. That is why we are trying to figure out the root cause of your situation, which is not clear to me as to what exactly is failing (possibly a buffer size issue, if this is consistently failing). My experience, as I said before, on these symptoms, is a proxy (even a local one) that is in the way. If you have your linux instance on a VM, the hypervisor may not be configured correctly. Lack of further evidence (all we really have is the curl RPC failure) makes diagnosing this very difficult.\n> \n\nThanks for your response, I appreciate it. I don't know what the hold up \nis for them, but I'm probably too unimportant, which I understand. I'm \nnot an enterprise user, and >99% of others have faster connections than \nme which is perhaps why they dodge this config(?) issue.\n\nAnd thanks for your suggestion, but sadly it seems to have no effect:\n\n$ git config --global http.postBuffer 524288000\n$ git -c http.version=HTTP/1.1 clone https://github.com/maliit/keyboard \nmaliit-keyboard\nCloning into 'maliit-keyboard'...\nremote: Enumerating objects: 23243, done.\nremote: Counting objects: 100% (464/464), done.\nremote: Compressing objects: 100% (207/207), done.\nerror: RPC failed; curl 18 transfer closed with outstanding read data \nremaining\nerror: 2444 bytes of body are still expected\nfetch-pack: unexpected disconnect while reading sideband packet\nfatal: early EOF\nfatal: fetch-pack: invalid index-pack output\n\nI'm doubtful this is solvable without either some resume or a fix from \nGithub's end. But I can use SSH clone so this isn't urgent.\n\nResume just seemed like an idea that would also help others, and it's \nwhat makes many other internet services work much better for me.\n\nRegards,\n\nEllie\n\n"},{"id":"498276","messageId":"CAFjaU5sGvRD+jXOgLhx9qjQ_McawEzVt035DE6b2nx7+rU188A@mail.gmail.com","threadId":"61605","inReplyTo":"001301dad14b$f8f0e460$ead2ad20$@nexbridge.com","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"Emanuel Czirai","fromEmail":"correabuscar+gitml@gmail.com","sentAt":"2024-07-08T16:09:24Z","receivedAt":"2024-07-08T16:09:34Z","isPatch":false,"sender":{"key":"correabuscar+gitml@gmail.com","avatar":null},"body":"Can try traffic shaping it, temporarily, just to can reproduce the\nissue on (presumably)anyone's linux machine, like:\n$ sudo tc qdisc change dev em1 root tbf rate 8kbit burst 8kbit latency 100ms\n(replace em1 with eth0 or whichever `ip a` reports as your LAN interface)\n\nLook at it:\n$ sudo tc qdisc show dev em1\nqdisc tbf 8001: root refcnt 2 rate 8Kbit burst 1Kb lat 100ms\n\n$ git clone https://github.com/maliit/keyboard\nCloning into 'keyboard'...\nremote: Enumerating objects: 23243, done. remote: Counting objects:\n100% (464/464), done. remote: Compressing objects: 100% (207/207),\ndone. error: 153 bytes of body are still expectedMiB | 1.14 MiB/s\nfetch-pack: unexpected disconnect while reading sideband packet\nfatal: early EOF\nfatal: fetch-pack: invalid index-pack output\n\n\nIt's different for me, but maybe this traffic shaping idea might still\nhelp if properly modified? (maybe it's too fast still? or not latent\nenough, I don't know)\n\nI tried it again: (seems different)\n$ git clone https://github.com/maliit/keyboard\nCloning into 'keyboard'...\nremote: Enumerating objects: 23243, done.\nremote: Counting objects: 100% (464/464), done.\nremote: Compressing objects: 100% (207/207), done.\nerror: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\nCANCEL (err 8)\nerror: 7932 bytes of body are still expected\nfetch-pack: unexpected disconnect while reading sideband packet\nfatal: early EOF\nfatal: fetch-pack: invalid index-pack output\n\nChange it: (use different values here for those 8 values and for the\n100, if needed, you get the picture)\n$ sudo tc qdisc change dev em1 root tbf rate 8kbit burst 8kbit latency 100ms\n\nor Delete it:(restore your unshaped traffic)\n$ sudo tc qdisc del dev em1 root\n\nLook at it after deletion:\n$ sudo tc qdisc show dev em1\nqdisc fq_codel 0: root refcnt 2 limit 10240p flows 1024 quantum 1514\ntarget 5ms interval 100ms memory_limit 32Mb ecn drop_batch 64\n\n/sbin/tc comes from package sys-apps/iproute2 6.9.0 on my Gentoo, ymmv.\nGood luck.\n\nOn Mon, Jul 8, 2024 at 5:32 PM <rsbecker@nexbridge.com> wrote:\n>\n> On Monday, July 8, 2024 11:15 AM, ellie wrote:\n> >On 7/8/24 4:32 PM, Konstantin Khomoutov wrote:\n> >> On Mon, Jul 08, 2024 at 04:28:25AM +0200, ellie wrote:\n> >>\n> >> [...]\n> >>> error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\n> >>> CANCEL (err 8)\n> >> [...]\n> >>> It seems extremely unlikely to me to be possibly an ISP issue, for\n> >>> which I already listed the reasons. An additional one is HTTPS\n> >>> downloads from github outside of git, e.g. from zip archives, for way\n> >>> larger files work fine as well.\n> >> [...]\n> >>\n> >> What if you explicitly disable HTTP/2 when cloning?\n> >>\n> >>    git -c http.version=HTTP/1.1 clone ...\n> >>\n> >> should probably do this.\n> >>\n> >\n> >Thanks for the idea! I tested it:\n> >\n> >$  git -c http.version=HTTP/1.1 clone https://github.com/maliit/keyboard\n> >maliit-keyboard\n> >Cloning into 'maliit-keyboard'...\n> >remote: Enumerating objects: 23243, done.\n> >remote: Counting objects: 100% (464/464), done.\n> >remote: Compressing objects: 100% (207/207), done.\n> >error: RPC failed; curl 18 transfer closed with outstanding read data remaining\n> >error: 5361 bytes of body are still expected\n> >fetch-pack: unexpected disconnect while reading sideband packet\n> >fatal: early EOF\n> >fatal: fetch-pack: invalid index-pack output\n> >\n> >Sadly, it seems like the error is only slightly different. It was still worth a try. I\n> >contacted GitHub support a while ago but it got stuck. If there were resume\n> >available such hiccups wouldn't matter, I hope that explains why I suggested that\n> >feature.\n>\n> I don't really understand what \"it got stuck\" means. Is that a colloquialism? What got stuck? That case at GitHub?\n>\n> Have you tried git config --global http.postBuffer 524288000\n>\n> It might help. The feature being requesting, even if possible, will probably not happen quickly, unless someone has a solid and simple design for this. That is why we are trying to figure out the root cause of your situation, which is not clear to me as to what exactly is failing (possibly a buffer size issue, if this is consistently failing). My experience, as I said before, on these symptoms, is a proxy (even a local one) that is in the way. If you have your linux instance on a VM, the hypervisor may not be configured correctly. Lack of further evidence (all we really have is the curl RPC failure) makes diagnosing this very difficult.\n>\n>\n"},{"id":"498279","messageId":"001a01dad153$271c2e60$75548b20$@nexbridge.com","threadId":"61605","inReplyTo":"47799635-7832-4c89-b4d3-e992d49ad40c@horse64.org","subject":"RE: With big repos and slower connections, git clone can be hard to work with","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2024-07-08T16:23:17Z","receivedAt":"2024-07-08T16:23:26Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On Monday, July 8, 2024 11:49 AM, ellie wrote:\n>On 7/8/24 5:31 PM, rsbecker@nexbridge.com wrote:\n>> On Monday, July 8, 2024 11:15 AM, ellie wrote:\n>>> On 7/8/24 4:32 PM, Konstantin Khomoutov wrote:\n>>>> On Mon, Jul 08, 2024 at 04:28:25AM +0200, ellie wrote:\n>>>>\n>>>> [...]\n>>>>> error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\n>>>>> CANCEL (err 8)\n>>>> [...]\n>>>>> It seems extremely unlikely to me to be possibly an ISP issue, for\n>>>>> which I already listed the reasons. An additional one is HTTPS\n>>>>> downloads from github outside of git, e.g. from zip archives, for\n>>>>> way larger files work fine as well.\n>>>> [...]\n>>>>\n>>>> What if you explicitly disable HTTP/2 when cloning?\n>>>>\n>>>>     git -c http.version=HTTP/1.1 clone ...\n>>>>\n>>>> should probably do this.\n>>>>\n>>>\n>>> Thanks for the idea! I tested it:\n>>>\n>>> $  git -c http.version=HTTP/1.1 clone\n>>> https://github.com/maliit/keyboard\n>>> maliit-keyboard\n>>> Cloning into 'maliit-keyboard'...\n>>> remote: Enumerating objects: 23243, done.\n>>> remote: Counting objects: 100% (464/464), done.\n>>> remote: Compressing objects: 100% (207/207), done.\n>>> error: RPC failed; curl 18 transfer closed with outstanding read data\n>>> remaining\n>>> error: 5361 bytes of body are still expected\n>>> fetch-pack: unexpected disconnect while reading sideband packet\n>>> fatal: early EOF\n>>> fatal: fetch-pack: invalid index-pack output\n>>>\n>>> Sadly, it seems like the error is only slightly different. It was\n>>> still worth a try. I contacted GitHub support a while ago but it got\n>>> stuck. If there were resume available such hiccups wouldn't matter, I\n>>> hope that explains why I suggested that feature.\n>>\n>> I don't really understand what \"it got stuck\" means. Is that a colloquialism? What\n>got stuck? That case at GitHub?\n>>\n>> Have you tried git config --global http.postBuffer 524288000\n>>\n>> It might help. The feature being requesting, even if possible, will probably not\n>happen quickly, unless someone has a solid and simple design for this. That is why\n>we are trying to figure out the root cause of your situation, which is not clear to me\n>as to what exactly is failing (possibly a buffer size issue, if this is consistently failing).\n>My experience, as I said before, on these symptoms, is a proxy (even a local one)\n>that is in the way. If you have your linux instance on a VM, the hypervisor may not\n>be configured correctly. Lack of further evidence (all we really have is the curl RPC\n>failure) makes diagnosing this very difficult.\n>>\n>\n>Thanks for your response, I appreciate it. I don't know what the hold up is for them,\n>but I'm probably too unimportant, which I understand. I'm not an enterprise user,\n>and >99% of others have faster connections than me which is perhaps why they\n>dodge this config(?) issue.\n>\n>And thanks for your suggestion, but sadly it seems to have no effect:\n>\n>$ git config --global http.postBuffer 524288000 $ git -c http.version=HTTP/1.1\n>clone https://github.com/maliit/keyboard\n>maliit-keyboard\n>Cloning into 'maliit-keyboard'...\n>remote: Enumerating objects: 23243, done.\n>remote: Counting objects: 100% (464/464), done.\n>remote: Compressing objects: 100% (207/207), done.\n>error: RPC failed; curl 18 transfer closed with outstanding read data remaining\n>error: 2444 bytes of body are still expected\n>fetch-pack: unexpected disconnect while reading sideband packet\n>fatal: early EOF\n>fatal: fetch-pack: invalid index-pack output\n>\n>I'm doubtful this is solvable without either some resume or a fix from Github's end.\n>But I can use SSH clone so this isn't urgent.\n>\n>Resume just seemed like an idea that would also help others, and it's what makes\n>many other internet services work much better for me.\n\nI do not know which pack file is having the issue - it may be the first one. Try running with the following environment variables GIT_TRACE=true and GIT_PACKET_TRACE=true. This will not correct the problem but might give additional helpful information. git uses libcurl to perform https transfers - which appears to be where the error is coming from. It is my opinion, given the issue is very likely in curl, that a restart capability will not help at all - at least not until we find the actual root cause (still mostly an unknown, although this error is widely discussed online in other non-git places). The failure appears to be transferring a single pack file (139824442 bytes) size may be an issue, but restarting in the middle of a pack file may not solve the problem (discussed in other threads) as the file is potentially built on demand (as I understand it from GitHub) and may not be the same on the next clone attempt. What we probably will find is that a restart will be stuck in the same spot and not move forward because the failure is not at a file boundary.\n\nIn addition to this, GitHub may have limits on the size of files that can be transferred, which you might be hitting (unlikely but possible). Check your plan options. I tried on a light plan, so this is unlikely but I want to exclude it.\n\n\n"},{"id":"498280","messageId":"20240708154457.jpt2aa5orzxy6kqh@carbon","threadId":"61605","inReplyTo":"2e10070f-2720-4d70-aa15-d4c008cc57bf@horse64.org","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"Konstantin Khomoutov","fromEmail":"kostix@bswap.ru","sentAt":"2024-07-08T15:44:57Z","receivedAt":"2024-07-08T16:24:42Z","isPatch":false,"sender":{"key":"kostix@bswap.ru","avatar":null},"body":"On Mon, Jul 08, 2024 at 05:14:33PM +0200, ellie wrote:\n\n[...]\n> > > error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly: CANCEL\n> > > (err 8)\n> > [...]\n> > > It seems extremely unlikely to me to be possibly an ISP issue, for which I\n> > > already listed the reasons. An additional one is HTTPS downloads from github\n> > > outside of git, e.g. from zip archives, for way larger files work fine as\n> > > well.\n> > [...]\n> > What if you explicitly disable HTTP/2 when cloning?\n[...]\n> Thanks for the idea! I tested it:\n> \n> $  git -c http.version=HTTP/1.1 clone https://github.com/maliit/keyboard\n\nOver there at SO people are trying all sorts of black magic to combat a\nproblem which manifests itself in a way very similar to yours [1]. I'm not\nsure anything from there could be of help but maybe worth trying anyway as you\ncan override any (or almost any) Git's configuration setting using that \"-c\"\ncommand-line option, so basically test round-trips should not be painstakingly\nlong.\n\n[...]\n> fetch-pack: unexpected disconnect while reading sideband packet\n[...]\n> Sadly, it seems like the error is only slightly different.\n\nI actually find it interesting that in each case a sideband packet is\nmentioned. But quite possibly it's a red herring anyway.\n\n 1. https://stackoverflow.com/questions/66366582\n\n"},{"id":"498281","messageId":"001b01dad153$ba880ca0$2f9825e0$@nexbridge.com","threadId":"61605","inReplyTo":"20240708154457.jpt2aa5orzxy6kqh@carbon","subject":"RE: With big repos and slower connections, git clone can be hard to work with","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2024-07-08T16:27:24Z","receivedAt":"2024-07-08T16:27:36Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On Monday, July 8, 2024 11:45 AM, Konstantin Khomoutov wrote:\n>On Mon, Jul 08, 2024 at 05:14:33PM +0200, ellie wrote:\n>\n>[...]\n>> > > error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\n>> > > CANCEL (err 8)\n>> > [...]\n>> > > It seems extremely unlikely to me to be possibly an ISP issue, for\n>> > > which I already listed the reasons. An additional one is HTTPS\n>> > > downloads from github outside of git, e.g. from zip archives, for\n>> > > way larger files work fine as well.\n>> > [...]\n>> > What if you explicitly disable HTTP/2 when cloning?\n>[...]\n>> Thanks for the idea! I tested it:\n>>\n>> $  git -c http.version=HTTP/1.1 clone\n>> https://github.com/maliit/keyboard\n>\n>Over there at SO people are trying all sorts of black magic to combat a\nproblem\n>which manifests itself in a way very similar to yours [1]. I'm not sure\nanything from\n>there could be of help but maybe worth trying anyway as you can override\nany (or\n>almost any) Git's configuration setting using that \"-c\"\n>command-line option, so basically test round-trips should not be\npainstakingly\n>long.\n>\n>[...]\n>> fetch-pack: unexpected disconnect while reading sideband packet\n>[...]\n>> Sadly, it seems like the error is only slightly different.\n>\n>I actually find it interesting that in each case a sideband packet is\nmentioned. But\n>quite possibly it's a red herring anyway.\n>\n> 1. https://stackoverflow.com/questions/66366582\n\nI have customers who hit this problem frequently setting up git. It is 99%\nof the time a firewall or proxy configuration issue, not specific to GitHub,\nand changes to those usually resolve the problem. The firewall and proxy can\nbe implemented in the ISP's modem if coming from a home network. That is why\nI really think the OP's issue is the network, not something that can\nreasonably fixed in git. I think the network speed is also a potential\nred-herring unless the speed issue relates to the ISP's configuration.\n\n"},{"id":"498291","messageId":"ff50859c-8e67-46ed-a8bd-ad7f836374c1@horse64.org","threadId":"61605","inReplyTo":"001a01dad153$271c2e60$75548b20$@nexbridge.com","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"ellie","fromEmail":"el@horse64.org","sentAt":"2024-07-08T17:06:12Z","receivedAt":"2024-07-08T17:06:15Z","isPatch":false,"sender":{"key":"el@horse64.org","avatar":null},"body":"\n\nOn 7/8/24 6:23 PM, rsbecker@nexbridge.com wrote:\n> On Monday, July 8, 2024 11:49 AM, ellie wrote:\n>> On 7/8/24 5:31 PM, rsbecker@nexbridge.com wrote:\n>>> On Monday, July 8, 2024 11:15 AM, ellie wrote:\n>>>> On 7/8/24 4:32 PM, Konstantin Khomoutov wrote:\n>>>>> On Mon, Jul 08, 2024 at 04:28:25AM +0200, ellie wrote:\n>>>>>\n>>>>> [...]\n>>>>>> error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\n>>>>>> CANCEL (err 8)\n>>>>> [...]\n>>>>>> It seems extremely unlikely to me to be possibly an ISP issue, for\n>>>>>> which I already listed the reasons. An additional one is HTTPS\n>>>>>> downloads from github outside of git, e.g. from zip archives, for\n>>>>>> way larger files work fine as well.\n>>>>> [...]\n>>>>>\n>>>>> What if you explicitly disable HTTP/2 when cloning?\n>>>>>\n>>>>>      git -c http.version=HTTP/1.1 clone ...\n>>>>>\n>>>>> should probably do this.\n>>>>>\n>>>>\n>>>> Thanks for the idea! I tested it:\n>>>>\n>>>> $  git -c http.version=HTTP/1.1 clone\n>>>> https://github.com/maliit/keyboard\n>>>> maliit-keyboard\n>>>> Cloning into 'maliit-keyboard'...\n>>>> remote: Enumerating objects: 23243, done.\n>>>> remote: Counting objects: 100% (464/464), done.\n>>>> remote: Compressing objects: 100% (207/207), done.\n>>>> error: RPC failed; curl 18 transfer closed with outstanding read data\n>>>> remaining\n>>>> error: 5361 bytes of body are still expected\n>>>> fetch-pack: unexpected disconnect while reading sideband packet\n>>>> fatal: early EOF\n>>>> fatal: fetch-pack: invalid index-pack output\n>>>>\n>>>> Sadly, it seems like the error is only slightly different. It was\n>>>> still worth a try. I contacted GitHub support a while ago but it got\n>>>> stuck. If there were resume available such hiccups wouldn't matter, I\n>>>> hope that explains why I suggested that feature.\n>>>\n>>> I don't really understand what \"it got stuck\" means. Is that a colloquialism? What\n>> got stuck? That case at GitHub?\n>>>\n>>> Have you tried git config --global http.postBuffer 524288000\n>>>\n>>> It might help. The feature being requesting, even if possible, will probably not\n>> happen quickly, unless someone has a solid and simple design for this. That is why\n>> we are trying to figure out the root cause of your situation, which is not clear to me\n>> as to what exactly is failing (possibly a buffer size issue, if this is consistently failing).\n>> My experience, as I said before, on these symptoms, is a proxy (even a local one)\n>> that is in the way. If you have your linux instance on a VM, the hypervisor may not\n>> be configured correctly. Lack of further evidence (all we really have is the curl RPC\n>> failure) makes diagnosing this very difficult.\n>>>\n>>\n>> Thanks for your response, I appreciate it. I don't know what the hold up is for them,\n>> but I'm probably too unimportant, which I understand. I'm not an enterprise user,\n>> and >99% of others have faster connections than me which is perhaps why they\n>> dodge this config(?) issue.\n>>\n>> And thanks for your suggestion, but sadly it seems to have no effect:\n>>\n>> $ git config --global http.postBuffer 524288000 $ git -c http.version=HTTP/1.1\n>> clone https://github.com/maliit/keyboard\n>> maliit-keyboard\n>> Cloning into 'maliit-keyboard'...\n>> remote: Enumerating objects: 23243, done.\n>> remote: Counting objects: 100% (464/464), done.\n>> remote: Compressing objects: 100% (207/207), done.\n>> error: RPC failed; curl 18 transfer closed with outstanding read data remaining\n>> error: 2444 bytes of body are still expected\n>> fetch-pack: unexpected disconnect while reading sideband packet\n>> fatal: early EOF\n>> fatal: fetch-pack: invalid index-pack output\n>>\n>> I'm doubtful this is solvable without either some resume or a fix from Github's end.\n>> But I can use SSH clone so this isn't urgent.\n>>\n>> Resume just seemed like an idea that would also help others, and it's what makes\n>> many other internet services work much better for me.\n> \n> I do not know which pack file is having the issue - it may be the first one. Try running with the following environment variables GIT_TRACE=true and GIT_PACKET_TRACE=true. This will not correct the problem but might give additional helpful information. git uses libcurl to perform https transfers - which appears to be where the error is coming from. It is my opinion, given the issue is very likely in curl, that a restart capability will not help at all - at least not until we find the actual root cause (still mostly an unknown, although this error is widely discussed online in other non-git places). The failure appears to be transferring a single pack file (139824442 bytes) size may be an issue, but restarting in the middle of a pack file may not solve the problem (discussed in other threads) as the file is potentially built on demand (as I understand it from GitHub) and may not be the same on the next clone attempt. What we probably will find is that a restart will be stuck in the same spot and not move forward because the failure is not at a file boundary.\n> \n> In addition to this, GitHub may have limits on the size of files that can be transferred, which you might be hitting (unlikely but possible). Check your plan options. I tried on a light plan, so this is unlikely but I want to exclude it.\n> \n> \nI attached the output of this command:\n\n$ GIT_TRACE=true GIT_PACKET_TRACE=true git -c http.version=HTTP/1.1 \nclone https://github.com/malii\nt/keyboard maliit-keyboard > log.txt 2>&1\n\nMy best guess is still that due to some unfortunate timeout choice, \nGithub's end simply becomes impatient and closes the connection.\n\nRegards,\n\nEllie\n\n\n18:44:33.182907 git.c:465               trace: built-in: git clone https://github.com/maliit/keyboard maliit-keyboard\nCloning into 'maliit-keyboard'...\n18:44:33.186926 run-command.c:657       trace: run_command: git remote-https origin https://github.com/maliit/keyboard\n18:44:33.188668 git.c:750               trace: exec: git-remote-https origin https://github.com/maliit/keyboard\n18:44:33.188728 run-command.c:657       trace: run_command: git-remote-https origin https://github.com/maliit/keyboard\n18:44:34.757740 run-command.c:657       trace: run_command: git index-pack --stdin --fix-thin '--keep=fetch-pack 14261 on elliedeck' --check-self-contained-and-connected\n18:44:34.759305 git.c:465               trace: built-in: git index-pack --stdin --fix-thin '--keep=fetch-pack 14261 on elliedeck' --check-self-contained-and-connected\nerror: RPC failed; curl 18 transfer closed with outstanding read data remaining\nerror: 5858 bytes of body are still expected\nfetch-pack: unexpected disconnect while reading sideband packet\nfatal: early EOF\nfatal: fetch-pack: invalid index-pack output\n"},{"id":"498299","messageId":"003d01dad15d$abd3f750$037be5f0$@nexbridge.com","threadId":"61605","inReplyTo":"ff50859c-8e67-46ed-a8bd-ad7f836374c1@horse64.org","subject":"RE: With big repos and slower connections, git clone can be hard to work with","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2024-07-08T17:38:34Z","receivedAt":"2024-07-08T17:38:44Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On Monday, July 8, 2024 1:06 PM, ellie wrote:\n>On 7/8/24 6:23 PM, rsbecker@nexbridge.com wrote:\n>> On Monday, July 8, 2024 11:49 AM, ellie wrote:\n>>> On 7/8/24 5:31 PM, rsbecker@nexbridge.com wrote:\n>>>> On Monday, July 8, 2024 11:15 AM, ellie wrote:\n>>>>> On 7/8/24 4:32 PM, Konstantin Khomoutov wrote:\n>>>>>> On Mon, Jul 08, 2024 at 04:28:25AM +0200, ellie wrote:\n>>>>>>\n>>>>>> [...]\n>>>>>>> error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\n>>>>>>> CANCEL (err 8)\n>>>>>> [...]\n>>>>>>> It seems extremely unlikely to me to be possibly an ISP issue,\n>>>>>>> for which I already listed the reasons. An additional one is\n>>>>>>> HTTPS downloads from github outside of git, e.g. from zip\n>>>>>>> archives, for way larger files work fine as well.\n>>>>>> [...]\n>>>>>>\n>>>>>> What if you explicitly disable HTTP/2 when cloning?\n>>>>>>\n>>>>>>      git -c http.version=HTTP/1.1 clone ...\n>>>>>>\n>>>>>> should probably do this.\n>>>>>>\n>>>>>\n>>>>> Thanks for the idea! I tested it:\n>>>>>\n>>>>> $  git -c http.version=HTTP/1.1 clone\n>>>>> https://github.com/maliit/keyboard\n>>>>> maliit-keyboard\n>>>>> Cloning into 'maliit-keyboard'...\n>>>>> remote: Enumerating objects: 23243, done.\n>>>>> remote: Counting objects: 100% (464/464), done.\n>>>>> remote: Compressing objects: 100% (207/207), done.\n>>>>> error: RPC failed; curl 18 transfer closed with outstanding read\n>>>>> data remaining\n>>>>> error: 5361 bytes of body are still expected\n>>>>> fetch-pack: unexpected disconnect while reading sideband packet\n>>>>> fatal: early EOF\n>>>>> fatal: fetch-pack: invalid index-pack output\n>>>>>\n>>>>> Sadly, it seems like the error is only slightly different. It was\n>>>>> still worth a try. I contacted GitHub support a while ago but it\n>>>>> got stuck. If there were resume available such hiccups wouldn't\n>>>>> matter, I hope that explains why I suggested that feature.\n>>>>\n>>>> I don't really understand what \"it got stuck\" means. Is that a\n>>>> colloquialism? What\n>>> got stuck? That case at GitHub?\n>>>>\n>>>> Have you tried git config --global http.postBuffer 524288000\n>>>>\n>>>> It might help. The feature being requesting, even if possible, will\n>>>> probably not\n>>> happen quickly, unless someone has a solid and simple design for\n>>> this. That is why we are trying to figure out the root cause of your\n>>> situation, which is not clear to me as to what exactly is failing (possibly a buffer\n>size issue, if this is consistently failing).\n>>> My experience, as I said before, on these symptoms, is a proxy (even\n>>> a local one) that is in the way. If you have your linux instance on a\n>>> VM, the hypervisor may not be configured correctly. Lack of further\n>>> evidence (all we really have is the curl RPC\n>>> failure) makes diagnosing this very difficult.\n>>>>\n>>>\n>>> Thanks for your response, I appreciate it. I don't know what the hold\n>>> up is for them, but I'm probably too unimportant, which I understand.\n>>> I'm not an enterprise user, and >99% of others have faster\n>>> connections than me which is perhaps why they dodge this config(?) issue.\n>>>\n>>> And thanks for your suggestion, but sadly it seems to have no effect:\n>>>\n>>> $ git config --global http.postBuffer 524288000 $ git -c\n>>> http.version=HTTP/1.1 clone https://github.com/maliit/keyboard\n>>> maliit-keyboard\n>>> Cloning into 'maliit-keyboard'...\n>>> remote: Enumerating objects: 23243, done.\n>>> remote: Counting objects: 100% (464/464), done.\n>>> remote: Compressing objects: 100% (207/207), done.\n>>> error: RPC failed; curl 18 transfer closed with outstanding read data\n>>> remaining\n>>> error: 2444 bytes of body are still expected\n>>> fetch-pack: unexpected disconnect while reading sideband packet\n>>> fatal: early EOF\n>>> fatal: fetch-pack: invalid index-pack output\n>>>\n>>> I'm doubtful this is solvable without either some resume or a fix from Github's\n>end.\n>>> But I can use SSH clone so this isn't urgent.\n>>>\n>>> Resume just seemed like an idea that would also help others, and it's\n>>> what makes many other internet services work much better for me.\n>>\n>> I do not know which pack file is having the issue - it may be the first one. Try\n>running with the following environment variables GIT_TRACE=true and\n>GIT_PACKET_TRACE=true. This will not correct the problem but might give\n>additional helpful information. git uses libcurl to perform https transfers - which\n>appears to be where the error is coming from. It is my opinion, given the issue is\n>very likely in curl, that a restart capability will not help at all - at least not until we\n>find the actual root cause (still mostly an unknown, although this error is widely\n>discussed online in other non-git places). The failure appears to be transferring a\n>single pack file (139824442 bytes) size may be an issue, but restarting in the middle\n>of a pack file may not solve the problem (discussed in other threads) as the file is\n>potentially built on demand (as I understand it from GitHub) and may not be the\n>same on the next clone attempt. What we probably will find is that a restart will be\n>stuck in the same spot and not move forward because the failure is not at a file\n>boundary.\n>>\n>> In addition to this, GitHub may have limits on the size of files that can be\n>transferred, which you might be hitting (unlikely but possible). Check your plan\n>options. I tried on a light plan, so this is unlikely but I want to exclude it.\n>>\n>>\n>I attached the output of this command:\n>\n>$ GIT_TRACE=true GIT_PACKET_TRACE=true git -c http.version=HTTP/1.1 clone\n>https://github.com/malii t/keyboard maliit-keyboard > log.txt 2>&1\n>\n>My best guess is still that due to some unfortunate timeout choice, Github's end\n>simply becomes impatient and closes the connection.\n\n18:44:33.182907 git.c:465               trace: built-in: git clone https://github.com/maliit/keyboard maliit-keyboard\nCloning into 'maliit-keyboard'...\n18:44:33.186926 run-command.c:657       trace: run_command: git remote-https origin https://github.com/maliit/keyboard\n18:44:33.188668 git.c:750               trace: exec: git-remote-https origin https://github.com/maliit/keyboard\n18:44:33.188728 run-command.c:657       trace: run_command: git-remote-https origin https://github.com/maliit/keyboard\n18:44:34.757740 run-command.c:657       trace: run_command: git index-pack --stdin --fix-thin '--keep=fetch-pack 14261 on elliedeck' --check-self-contained-and-connected\n18:44:34.759305 git.c:465               trace: built-in: git index-pack --stdin --fix-thin '--keep=fetch-pack 14261 on elliedeck' --check-self-contained-and-connected\nerror: RPC failed; curl 18 transfer closed with outstanding read data remaining\nerror: 5858 bytes of body are still expected\nfetch-pack: unexpected disconnect while reading sideband packet\nfatal: early EOF\nfatal: fetch-pack: invalid index-pack output\n\nFrom what I could tell from the log, the operation took less than 3 seconds. How long does it appear to take to you? This does not look like a timeout. In fact, it looks like the failure happened before git was able to process any content. From what I read from the log, libcurl encountered a failure and passed that up to git, which stopped the operation. You could try putting -v into your .curlrc file or otherwise getting some verbose information out of curl here the failure is occurring. I would also suggest passing this over to the curl team for examination. I am at a loss on resolving this, further, particularly if there are no intermediary components like firewalls and proxies - note that many ISPs build firewalls and proxies into their NAT routers. A curl verbose trace might show this. My home ISP in Canada has all kinds of stuff in their cable modems, which I had disabled by the tech who installed the box, and I have no issues cloning the above repo. They do have QoS limits but have not blocked https downloads.\n\n--Randall\n\n"},{"id":"498662","messageId":"a89fdb7c-745d-4e92-ac07-724cbe146116@horse64.org","threadId":"61605","inReplyTo":"001b01dad153$ba880ca0$2f9825e0$@nexbridge.com","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"ellie","fromEmail":"el@horse64.org","sentAt":"2024-07-14T12:00:53Z","receivedAt":"2024-07-14T12:01:10Z","isPatch":false,"sender":{"key":"el@horse64.org","avatar":null},"body":"\n\nOn 7/8/24 6:27 PM, rsbecker@nexbridge.com wrote:\n> On Monday, July 8, 2024 11:45 AM, Konstantin Khomoutov wrote:\n>> On Mon, Jul 08, 2024 at 05:14:33PM +0200, ellie wrote:\n>>\n>> [...]\n>>>>> error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\n>>>>> CANCEL (err 8)\n>>>> [...]\n>>>>> It seems extremely unlikely to me to be possibly an ISP issue, for\n>>>>> which I already listed the reasons. An additional one is HTTPS\n>>>>> downloads from github outside of git, e.g. from zip archives, for\n>>>>> way larger files work fine as well.\n>>>> [...]\n>>>> What if you explicitly disable HTTP/2 when cloning?\n>> [...]\n>>> Thanks for the idea! I tested it:\n>>>\n>>> $  git -c http.version=HTTP/1.1 clone\n>>> https://github.com/maliit/keyboard\n>>\n>> Over there at SO people are trying all sorts of black magic to combat a\n> problem\n>> which manifests itself in a way very similar to yours [1]. I'm not sure\n> anything from\n>> there could be of help but maybe worth trying anyway as you can override\n> any (or\n>> almost any) Git's configuration setting using that \"-c\"\n>> command-line option, so basically test round-trips should not be\n> painstakingly\n>> long.\n>>\n>> [...]\n>>> fetch-pack: unexpected disconnect while reading sideband packet\n>> [...]\n>>> Sadly, it seems like the error is only slightly different.\n>>\n>> I actually find it interesting that in each case a sideband packet is\n> mentioned. But\n>> quite possibly it's a red herring anyway.\n>>\n>> 1. https://stackoverflow.com/questions/66366582\n> \n> I have customers who hit this problem frequently setting up git. It is 99%\n> of the time a firewall or proxy configuration issue, not specific to GitHub,\n> and changes to those usually resolve the problem. The firewall and proxy can\n> be implemented in the ISP's modem if coming from a home network. That is why\n> I really think the OP's issue is the network, not something that can\n> reasonably fixed in git. I think the network speed is also a potential\n> red-herring unless the speed issue relates to the ISP's configuration.\n> \nFor what it's worth, it's definitely Github-specifc for me. Maybe one \nday Github support will respond, I can only hope.\n\nRegards,\n\nEllie\n\n"},{"id":"499224","messageId":"bd5ebf84-5d47-437c-b773-580188cfb001@horse64.org","threadId":"61605","inReplyTo":"001b01dad153$ba880ca0$2f9825e0$@nexbridge.com","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"ellie","fromEmail":"el@horse64.org","sentAt":"2024-07-24T06:42:54Z","receivedAt":"2024-07-24T06:50:22Z","isPatch":false,"sender":{"key":"el@horse64.org","avatar":null},"body":"For what it's worth, Github support now confirmed to me that it looks \nlike they might have a timeout problem on their side, but until more \npeople report it they likely won't address it. I appreciate their \nhonesty. But I think it shows the vulnerability of a process without \nresume well.\n\n(Sorry to harp on, I thought this extra info might be interesting.)\n\nRegards,\n\nEllie\n\nOn 7/8/24 6:27 PM, rsbecker@nexbridge.com wrote:\n> On Monday, July 8, 2024 11:45 AM, Konstantin Khomoutov wrote:\n>> On Mon, Jul 08, 2024 at 05:14:33PM +0200, ellie wrote:\n>>\n>> [...]\n>>>>> error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\n>>>>> CANCEL (err 8)\n>>>> [...]\n>>>>> It seems extremely unlikely to me to be possibly an ISP issue, for\n>>>>> which I already listed the reasons. An additional one is HTTPS\n>>>>> downloads from github outside of git, e.g. from zip archives, for\n>>>>> way larger files work fine as well.\n>>>> [...]\n>>>> What if you explicitly disable HTTP/2 when cloning?\n>> [...]\n>>> Thanks for the idea! I tested it:\n>>>\n>>> $  git -c http.version=HTTP/1.1 clone\n>>> https://github.com/maliit/keyboard\n>>\n>> Over there at SO people are trying all sorts of black magic to combat a\n> problem\n>> which manifests itself in a way very similar to yours [1]. I'm not sure\n> anything from\n>> there could be of help but maybe worth trying anyway as you can override\n> any (or\n>> almost any) Git's configuration setting using that \"-c\"\n>> command-line option, so basically test round-trips should not be\n> painstakingly\n>> long.\n>>\n>> [...]\n>>> fetch-pack: unexpected disconnect while reading sideband packet\n>> [...]\n>>> Sadly, it seems like the error is only slightly different.\n>>\n>> I actually find it interesting that in each case a sideband packet is\n> mentioned. But\n>> quite possibly it's a red herring anyway.\n>>\n>> 1. https://stackoverflow.com/questions/66366582\n> \n> I have customers who hit this problem frequently setting up git. It is 99%\n> of the time a firewall or proxy configuration issue, not specific to GitHub,\n> and changes to those usually resolve the problem. The firewall and proxy can\n> be implemented in the ISP's modem if coming from a home network. That is why\n> I really think the OP's issue is the network, not something that can\n> reasonably fixed in git. I think the network speed is also a potential\n> red-herring unless the speed issue relates to the ISP's configuration.\n> \n"},{"id":"503773","messageId":"55e7ddb5-327a-49f2-9f7c-e285ead69fb2@horse64.org","threadId":"61605","inReplyTo":"fec6ebc7-efd7-4c86-9dcc-2b006bd82e47@horse64.org","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"Ellie","fromEmail":"el@horse64.org","sentAt":"2024-09-30T21:01:53Z","receivedAt":"2024-09-30T21:08:15Z","isPatch":false,"sender":{"key":"el@horse64.org","avatar":null},"body":"My apologies for bringing this up again, but for what it's worth, this \ngit repository I can't even clone at depth 1:\n\n$ git clone --depth 1 https://github.com/alf632/terrain3dglitch\nCloning into 'terrain3dglitch'...\nremote: Enumerating objects: 697, done.\nremote: Counting objects: 100% (697/697), done.\nremote: Compressing objects: 100% (439/439), done.\nerror: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly: \nCANCEL (err 8)\nerror: 1754 bytes of body are still expected\nfetch-pack: unexpected disconnect while reading sideband packet\nfatal: early EOF\nfatal: fetch-pack: invalid index-pack output\n\nThe problem seems to be possibly amplified by a timeout config issue \nfrom github's side, but also made worse by depth 1 already being 100MB+. \nDownloading that amount without resume isn't feasible for everyone. I'm \nassuming if I need all files and sub dirs, there's no workaround here?\n\nI don't want to waste anybody's time, I'm just hoping to provide some \nfurther data points that in some edge cases, this can be impactful.\n\n(And sorry if I did something silly while cloning and didn't realize.)\n\nRegards,\n\nEllie\n\nOn 6/8/24 1:28 AM, ellie wrote:\n> Dear git team,\n> \n> I'm terribly sorry if this is the wrong place, but I'd like to suggest a \n> potential issue with \"git clone\".\n> \n> The problem is that any sort of interruption or connection issue, no \n> matter how brief, causes the clone to stop and leave nothing behind:\n> \n> $ git clone https://github.com/Nheko-Reborn/nheko\n> Cloning into 'nheko'...\n> remote: Enumerating objects: 43991, done.\n> remote: Counting objects: 100% (6535/6535), done.\n> remote: Compressing objects: 100% (1449/1449), done.\n> error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly: \n> CANCEL (err 8)\n> error: 2771 bytes of body are still expected\n> fetch-pack: unexpected disconnect while reading sideband packet\n> fatal: early EOF\n> fatal: fetch-pack: invalid index-pack output\n> $ cd nheko\n> bash: cd: nheko: No such file or director\n> \n> In my experience, this can be really impactful with 1. big repositories \n> and 2. unreliable internet - which I would argue isn't unheard of! E.g. \n> a developer may work via mobile connection on a business trip. The \n> result can even be that a repository is uncloneable for some users!\n> \n> This has left me in the absurd situation where I was able to download a \n> tarball via HTTPS from the git hoster just fine, even way larger binary \n> release items, thanks to the browser's HTTPS resume. And yet a simple \n> git clone of the same project failed repeatedly.\n> \n> My deepest apologies if I missed an option to fix or address this. But \n> summed up, please consider making git clone recover from hiccups.\n> \n> Regards,\n> \n> Ellie\n> \n> PS: I've seen git hosters have apparent proxy bugs, like timing out \n> slower git clone connections from the server side even if the transfer \n> is ongoing. A git auto-resume would reduce the impact of that, too.\n> \n> \n> \n\n"},{"id":"525746","messageId":"15eac16b-41b6-4bfc-91c7-4997d390cc5b@horse64.org","threadId":"61605","inReplyTo":"001b01dad153$ba880ca0$2f9825e0$@nexbridge.com","subject":"Re: With big repos and slower connections, git clone can be hard to work with","fromName":"Ellie","fromEmail":"el@horse64.org","sentAt":"2025-09-08T02:34:53Z","receivedAt":"2025-09-08T02:44:10Z","isPatch":false,"sender":{"key":"el@horse64.org","avatar":null},"body":"This has been addressed on Github's side by now, it seems to have been a \nGithub server config issue.\n\nNevertheless, the ability to resume a file transfer remains what some \nwould consider essential for internet software. I still hope it'll be \nadded one day.\n\nThank you for the lively debate.\n\nRegards,\n\nEllie\n\nOn 7/8/24 6:27 PM, rsbecker@nexbridge.com wrote:\n> On Monday, July 8, 2024 11:45 AM, Konstantin Khomoutov wrote:\n>> On Mon, Jul 08, 2024 at 05:14:33PM +0200, ellie wrote:\n>>\n>> [...]\n>>>>> error: RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly:\n>>>>> CANCEL (err 8)\n>>>> [...]\n>>>>> It seems extremely unlikely to me to be possibly an ISP issue, for\n>>>>> which I already listed the reasons. An additional one is HTTPS\n>>>>> downloads from github outside of git, e.g. from zip archives, for\n>>>>> way larger files work fine as well.\n>>>> [...]\n>>>> What if you explicitly disable HTTP/2 when cloning?\n>> [...]\n>>> Thanks for the idea! I tested it:\n>>>\n>>> $  git -c http.version=HTTP/1.1 clone\n>>> https://github.com/maliit/keyboard\n>>\n>> Over there at SO people are trying all sorts of black magic to combat a\n> problem\n>> which manifests itself in a way very similar to yours [1]. I'm not sure\n> anything from\n>> there could be of help but maybe worth trying anyway as you can override\n> any (or\n>> almost any) Git's configuration setting using that \"-c\"\n>> command-line option, so basically test round-trips should not be\n> painstakingly\n>> long.\n>>\n>> [...]\n>>> fetch-pack: unexpected disconnect while reading sideband packet\n>> [...]\n>>> Sadly, it seems like the error is only slightly different.\n>>\n>> I actually find it interesting that in each case a sideband packet is\n> mentioned. But\n>> quite possibly it's a red herring anyway.\n>>\n>> 1. https://stackoverflow.com/questions/66366582\n> \n> I have customers who hit this problem frequently setting up git. It is 99%\n> of the time a firewall or proxy configuration issue, not specific to GitHub,\n> and changes to those usually resolve the problem. The firewall and proxy can\n> be implemented in the ISP's modem if coming from a home network. That is why\n> I really think the OP's issue is the network, not something that can\n> reasonably fixed in git. I think the network speed is also a potential\n> red-herring unless the speed issue relates to the ISP's configuration.\n> \n\n"}]}