{"thread":{"id":"31755","subject":"Git ~unusable on slow lines :,'C","startedAt":"2012-10-08T18:27:54Z","lastAt":"2012-10-09T17:39:57Z","messageCount":7,"participants":["Marcel Partap","Carlos Martín Nieto","Shawn Pearce","Junio C Hamano"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"200774","messageId":"50731B2A.6040104@gmx.net","threadId":"31755","inReplyTo":null,"subject":"Git ~unusable on slow lines :,'C","fromName":"Marcel Partap","fromEmail":"mpartap@gmx.net","sentAt":"2012-10-08T18:27:54Z","receivedAt":"2012-10-08T18:27:54Z","isPatch":false,"sender":{"key":"mpartap@gmx.net","avatar":"https://gravatar.com/avatar/48334bf11d5314a55b0d9bb4e0eeb8f9ccd836ed2f978181850423d8ddf284ee?d=mp&s=160"},"body":"Dear Git Devs,\nI love GIT, but since a couple of months I'm on 3G and after my traffic\nlimit is transcended, things slow down to a feeble 8KiB/s. Juuuust like\nback then - things moved somewhat slower. And I'm fine with that - as\nlong as things just keep moving.\nUnfortunately, git does not scale down very well, so for ten more days I\nwill be unable to get the newest commits onto my machine. Which is very,\nvery sad :/\n> git fetch --verbose --all \n> Fetching origin\n> POST git-upload-pack (1023 bytes)\n> POST git-upload-pack (gzip 1123 to 614 bytes)\n> POST git-upload-pack (gzip 1973 to 1030 bytes)\n> POST git-upload-pack (gzip 5173 to 2639 bytes)\n> POST git-upload-pack (gzip 7978 to 4042 bytes)\n> remote: Counting objects: 24504, done.\n> remote: Compressing objects: 100% (10705/10705), done.\n> error: RPC failed; result=56, HTTP code = 200iB | 10 KiB/s       \n> fatal: The remote end hung up unexpectedly\n> fatal: early EOF\n> fatal: index-pack failed\n> error: Could not fetch origin\nBam, the server kicked me off after taking to long to sync my copy.\nMultiple potential points of action:\n- git fetch should show the total amount of data it is about to transfer!\n- when ab^H^Horting, the cursor should be moved down (tput cud1) to not\noverwrite previous output\n- would be nice to be able to tell git fetch to get the next chunk of\nsay 500 commits instead of trying to receive ALL commits, then b0rking\nafter umpteen percent on server timeout. Not?\n\n#Regards!Marcel c:\n"},{"id":"200799","messageId":"87lifgct3j.fsf@centaur.cmartin.tk","threadId":"31755","inReplyTo":"50731B2A.6040104@gmx.net","subject":"Re: Git ~unusable on slow lines :,'C","fromName":"Carlos Martín Nieto","fromEmail":"cmn@elego.de","sentAt":"2012-10-09T01:49:36Z","receivedAt":"2012-10-09T01:49:36Z","isPatch":false,"sender":{"key":"cmn@elego.de","avatar":"https://avatars.githubusercontent.com/u/335443?v=4"},"body":"Marcel Partap <mpartap@gmx.net> writes:\n\n> Dear Git Devs,\n> I love GIT, but since a couple of months I'm on 3G and after my traffic\n> limit is transcended, things slow down to a feeble 8KiB/s. Juuuust like\n> back then - things moved somewhat slower. And I'm fine with that - as\n> long as things just keep moving.\n> Unfortunately, git does not scale down very well, so for ten more days I\n> will be unable to get the newest commits onto my machine. Which is very,\n> very sad :/\n>> git fetch --verbose --all \n>> Fetching origin\n>> POST git-upload-pack (1023 bytes)\n>> POST git-upload-pack (gzip 1123 to 614 bytes)\n>> POST git-upload-pack (gzip 1973 to 1030 bytes)\n>> POST git-upload-pack (gzip 5173 to 2639 bytes)\n>> POST git-upload-pack (gzip 7978 to 4042 bytes)\n>> remote: Counting objects: 24504, done.\n>> remote: Compressing objects: 100% (10705/10705), done.\n>> error: RPC failed; result=56, HTTP code = 200iB | 10 KiB/s       \n>> fatal: The remote end hung up unexpectedly\n>> fatal: early EOF\n>> fatal: index-pack failed\n>> error: Could not fetch origin\n> Bam, the server kicked me off after taking to long to sync my copy.\n\nThis is unrelated to git. The HTTP server's configuration is too\nimpatient.\n\n> Multiple potential points of action:\n> - git fetch should show the total amount of data it is about to\n> transfer!\n\nIt can't, because it doesn't know.\n\n> - when ab^H^Horting, the cursor should be moved down (tput cud1) to not\n> overwrite previous output\n\nThe error message doesn't really know whether it is going to overwrite\nit (the CR comes from the server), though I suppose an extra LF wouldn't\nhurt there.\n\n> - would be nice to be able to tell git fetch to get the next chunk of\n> say 500 commits instead of trying to receive ALL commits, then b0rking\n> after umpteen percent on server timeout. Not?\n\nYou asked for the current state of the repository, and that's what its\ngiving you. The timeout has nothing to do with git, if you can't\nconvince the admins to increase it, you can try using another transport\nwhich doesn't suffer from HTTP, as it's most likely an anti-DoS measure.\n\nIf you want to download it bit by bit, you can tell fetch to download\nparticular tags. Doing this automatically for this would be working\naround a configuration issue for a particular server, which is generally\nbetter fixed in other ways.\n\n\n   cmn\n"},{"id":"200835","messageId":"50742F53.3050205@gmx.net","threadId":"31755","inReplyTo":"87lifgct3j.fsf@centaur.cmartin.tk","subject":"Re: Git ~unusable on slow lines :,'C","fromName":"Marcel Partap","fromEmail":"mpartap@gmx.net","sentAt":"2012-10-09T14:06:11Z","receivedAt":"2012-10-09T14:06:11Z","isPatch":false,"sender":{"key":"mpartap@gmx.net","avatar":"https://gravatar.com/avatar/48334bf11d5314a55b0d9bb4e0eeb8f9ccd836ed2f978181850423d8ddf284ee?d=mp&s=160"},"body":">> Bam, the server kicked me off after taking to long to sync my copy.\n> This is unrelated to git. The HTTP server's configuration is too\n> impatient.\nYes. How does that mean it is unrelated to git?\n\n>> - git fetch should show the total amount of data it is about to\n>> transfer!\n> It can't, because it doesn't know.\nThe server side doesn't know at how much the objects *it just repacked\nfor transfer* weigh in?\nIf that truly is the case, wouldn't it make sense to make git a little\nmore introspective? f.e.\n> # git info git://foo.org/bar.git\n> .. [server generating figures] ..\n> URL: git://foo.org/bar.git\n> Created/Earliest commit: ...\n> Last modified/Latest commit: ...\n> Total object count: .... (..commits, ..files, .. directories)\n> Total repository size (compressed): ... MiB\n> Branches:\n> [git branch -va] + branch size\n\n> The error message doesn't really know whether it is going to overwrite\n> it (the CR comes from the server), though I suppose an extra LF wouldn't\n> hurt there.\nDefinitely wouldn't hurt.\n\n>> - would be nice to be able to tell git fetch to get the next chunk of\n>> say 500 commits instead of trying to receive ALL commits, then b0rking\n>> after umpteen percent on server timeout. Not?\n> You asked for the current state of the repository, and that's what its\n> giving you.\nAnd instead, I would rather like to ask for the next 500 commits. No way\nto do it.\n\n> The timeout has nothing to do with git, if you can't\n> convince the admins to increase it, you can try using another transport\n> which doesn't suffer from HTTP, as it's most likely an anti-DoS measure.\nSee, I probably can't convince the admins to drop their anti-dos measures.\nAnd they (drupal.org admins) probably will not change their allowed\nprotocol policies.\nDespite that, i've had timeouts or simply stale connections dying down\nbefore with other repositories and various transport modes.\nThe easiest fix would be an option to tell git to not fetch everything...\n\n> If you want to download it bit by bit, you can tell fetch to download\n> particular tags.\n..without specifying specific commit tags.\nBrowsing gitweb sites to find a tag for which the fetch doesn't time out\nis hugely inconvenient, especially on a slow line.\n\n> Doing this automatically for this would be working\n> around a configuration issue for a particular server, which is generally\n> better fixed in other ways.\nIt is not only a configuration issue for one particular server. Git in\ngeneral is hardly usable on slow lines because\n- it doesn't show the volume of data that is to be downloaded!\n- it doesn't allow the user to sync up in steps the circumstances will\nallow to succeed.\n\n#Regards!Marcel.\n"},{"id":"200838","messageId":"CAJo=hJv+CtEcGFPhe2xPsfrPmdfOuakMovbk8-cJmFjOnwKWnQ@mail.gmail.com","threadId":"31755","inReplyTo":"50742F53.3050205@gmx.net","subject":"Re: Git ~unusable on slow lines :,'C","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2012-10-09T15:58:03Z","receivedAt":"2012-10-09T15:58:03Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Tue, Oct 9, 2012 at 7:06 AM, Marcel Partap <mpartap@gmx.net> wrote:\n>>> Bam, the server kicked me off after taking to long to sync my copy.\n>> This is unrelated to git. The HTTP server's configuration is too\n>> impatient.\n> Yes. How does that mean it is unrelated to git?\n\nIt means its out of our control, we cannot modify the HTTP server's\nconfiguration to have a longer timeout. We can recommend that the\ntimeout be increased, but as you point out the admins may not do that.\n\n>>> - git fetch should show the total amount of data it is about to\n>>> transfer!\n>> It can't, because it doesn't know.\n> The server side doesn't know at how much the objects *it just repacked\n> for transfer* weigh in?\n\nActually it does. Its just not used here. What value is that to you?\nYou asked for the repository. If you know its size is going to be ~105\nMiB you have two choices... continue to get the repository you asked\nfor, or disconnect and give up. Either way the size doesn't help you.\nIt would require a protocol modification to send a size estimate down\nto the client before the data in order to give the client a better\nprogress meter than the object count (allowing it instead to track by\nbytes received). But this has been seen as not very useful or\nworthwhile since it doesn't really help anyone do anything better. So\nwhy change the protocol?\n\n>> You asked for the current state of the repository, and that's what its\n>> giving you.\n> And instead, I would rather like to ask for the next 500 commits. No way\n> to do it.\n\nNo, there isn't. Git assumes that once it has commit X, all versions\nthat predate X are already on the local workstation. This is a\nfundamental assumption that the entire protocol relies on. It is not\ntrivial to change. We have been through this many times on the mailing\nlist, please search the archives for \"resumable clone\".\n\n>> The timeout has nothing to do with git, if you can't\n>> convince the admins to increase it, you can try using another transport\n>> which doesn't suffer from HTTP, as it's most likely an anti-DoS measure.\n> See, I probably can't convince the admins to drop their anti-dos measures.\n> And they (drupal.org admins) probably will not change their allowed\n> protocol policies.\n\nThen if they are hosting really big repositories that are hard for\ntheir contributors to obtain, they should take the time to write a\nscript that periodically creates a bundle file for each repository\nusing `git bundle create repo.bundle --all`. They can host these\nbundle files in any file transport service like HTTP or BitTorrent,\nand users can download and resume these using normal HTTP download\ntools. Once you have a bundle file locally, you can clone from it with\nmodern Git with `git clone $(pwd)/repo.bundle` to initialize the\nrepository.\n\nThis is currently the best way to support resumable clone. The repo\nwill be stale by whatever time has elapsed since the bundle file was\ncreated. But then Git can do an incremental fetch to catch up, and\nthis transfer size should be limited to the progress made since the\nbundle was made. If bundles are made once per month or after each\nmajor release its usually a manageable delta.\n\n> It is not only a configuration issue for one particular server. Git in\n> general is hardly usable on slow lines because\n> - it doesn't show the volume of data that is to be downloaded!\n\nIf it did show you, what would you do? Declare defeat before it even\nstarts to download and give up and start a thread about how Git\nrequires too much bandwidth?\n\nHave you tried to shallow clone the repository in question?\n\n> - it doesn't allow the user to sync up in steps the circumstances will\n> allow to succeed.\n\nSadly, this is quite true. :-(\n"},{"id":"200842","messageId":"7vd30rh9uo.fsf@alter.siamese.dyndns.org","threadId":"31755","inReplyTo":"87lifgct3j.fsf@centaur.cmartin.tk","subject":"Re: Git ~unusable on slow lines :,'C","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-10-09T16:46:23Z","receivedAt":"2012-10-09T16:46:23Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"cmn@elego.de (Carlos Martín Nieto) writes:\n\n> If you want to download it bit by bit, you can tell fetch to download\n> particular tags. Doing this automatically for this would be working\n> around a configuration issue for a particular server, which is generally\n> better fixed in other ways.\n\nAs part of an upcoming \"protocol update\" discussion, we may want to\ninclude allowing \"upload-pack\" to accept a request for commit that\nis not at the tip of any ref.\n\nE.g. \"want refs/heads/master~*0.1\" might ask \"I know your entire\nhistory is very big; please give me only the one tenth of the oldest\nhistory during this round.\" (this is not a suggestion on how to do\nthis at the UI level).\n"},{"id":"200844","messageId":"50745C89.7060707@gmx.net","threadId":"31755","inReplyTo":"CAJo=hJv+CtEcGFPhe2xPsfrPmdfOuakMovbk8-cJmFjOnwKWnQ@mail.gmail.com","subject":"Re: Git ~unusable on slow lines :,'C","fromName":"Marcel Partap","fromEmail":"mpartap@gmx.net","sentAt":"2012-10-09T17:19:05Z","receivedAt":"2012-10-09T17:19:05Z","isPatch":false,"sender":{"key":"mpartap@gmx.net","avatar":"https://gravatar.com/avatar/48334bf11d5314a55b0d9bb4e0eeb8f9ccd836ed2f978181850423d8ddf284ee?d=mp&s=160"},"body":">>>> - git fetch should show the total amount of data it is about to\n>>>> transfer!\n>>> It can't, because it doesn't know.\n>> The server side doesn't know at how much the objects *it just repacked\n>> for transfer* weigh in?\n> Actually it does.\nThen, please, make it display it.\n\n> What value is that to you?\nThe size that is to be transferred, and the total repository size.\n\n> You asked for the repository. If you know its size is going to be ~105\n> MiB you have two choices... continue to get the repository you asked\n> for, or disconnect and give up.\n> Either way the size doesn't help you.\nYes it does - when displayed, one could make an informed choice.\nBut it doesn't show this, just the object count.. and that is of low\nexpressiveness.\nIt so happened last week that I tried cloning a repository with a\nseemingly moderate amount of objects and small code base. However, full\nJava RE zips had been checked in and updated multiple times - suddenly\nmy monthly 3G traffic limit was exhausted. Needless to say without a\nclue how much more data would follow, I aborted the transfer - and was\nleft with a net result of *zilch* bytes of code, and a line cut down to\nridiculous speed. Now I can't even sync up my Drupal copy.\n\n> It would require a protocol modification to send a size estimate down\n> to the client before the data in order to give the client a better\n> progress meter than the object count (allowing it instead to track by\n> bytes received).\nWell, if it requires that, so be it. I fail to understand why this\nwasn't considered before.\n\n> But this has been seen as not very useful or worthwhile\n> since it doesn't really help anyone do anything better.\nHuh?\n\n> So why change the protocol?\nSanity? Usability of git with slow lines?\n\n\n> Git assumes that once it has commit X, all versions\n> that predate X are already on the local workstation.\nAnd that's true for all my repositories, since none of them was cloned\n--shallow.\n\n> This is a fundamental assumption that the entire protocol relies on.\nWhat about --shallow, --depth?\n\n> It is not trivial to change.\nMany changes for the better are not trivial. And still worth it.\n\n> We have been through this many times on the mailing\n> list, please search the archives for \"resumable clone\".\nOk - yet that probably doesn't invalidate all arguments in favor of it.\n\n> they should [...] host these bundle files [...]\n> and users can download and resume these\nThanks for the tip, I will forward it to the server administrators.\nHowever, this does not help to handle the huge amount of commits to\nfetch that pile up within a couple of months.\n\n> This is currently the best way to support resumable clone.\nI wasn't even mentioning that, but that'd be nice to have aswell^^...\n\n> If bundles are made once per month or after each\n> major release its usually a manageable delta.\nWhile downloading bundle delta files definitely is a plausible solution\n- isn't that quite far from user friendly?\n\n> If it did show you, what would you do?\nNot try to checkout a repository full of JRE zips blindfolded?\n\n> Declare defeat before it even\n> starts to download and give up and start a thread about how Git\n> requires too much bandwidth?\nKindly ask the author to locally rewrite his history and recreate the\nrepository with *LINKS* to JRE zips instead?\nNot for a second did I doubt the efficiency of git's packing and\ncompression algorithms! That's why I'm quite amazed about the shear\nexistence of these issues of not showing the repository size before\ndownloading (or, IIUC, *anywhere*) and a protocol that is incapable of\nresuming or partly fetching a repository, even though it obviously\nprovides means of negotiation between server and client.. Just boggles\nme that within 7+ years of development this hasn't been addressed\n(disclaimer: I do not claim to grok the protocol - not wanting to put\nblame on anyone here :).\n\n> Have you tried to shallow clone the repository in question?\nNo - would it allow me to fuse the two repositories afterwards? That'd\nactually be quite cool and a good idea to instantly solve my current\nproblem... gonna try that, thx :)\n\n#Regards!Marcel\n"},{"id":"200847","messageId":"87a9vvczo2.fsf@centaur.cmartin.tk","threadId":"31755","inReplyTo":"50742F53.3050205@gmx.net","subject":"Re: Git ~unusable on slow lines :,'C","fromName":"Carlos Martín Nieto","fromEmail":"cmn@dwim.me","sentAt":"2012-10-09T17:39:57Z","receivedAt":"2012-10-09T17:39:57Z","isPatch":false,"sender":{"key":"cmn@dwim.me","avatar":"https://avatars.githubusercontent.com/u/335443?v=4"},"body":"Marcel Partap <mpartap@gmx.net> writes:\n\n>>> Bam, the server kicked me off after taking to long to sync my copy.\n>> This is unrelated to git. The HTTP server's configuration is too\n>> impatient.\n> Yes. How does that mean it is unrelated to git?\n>\n>>> - git fetch should show the total amount of data it is about to\n>>> transfer!\n>> It can't, because it doesn't know.\n> The server side doesn't know at how much the objects *it just repacked\n> for transfer* weigh in?\n> If that truly is the case, wouldn't it make sense to make git a little\n> more introspective? f.e.\n\nIt sends you more objects than the ones it just repacked in the normal\ncase. It could tell you, but it would have to keep track of more\ninformation (which would make it take longer for the first bytes to get\nto you) for little gain. The only thing you'd be able to do is to\nabort the transfer immediately, but you can do that anyway, and waiting\nis only going to add history to download.\n\n>> # git info git://foo.org/bar.git\n>> .. [server generating figures] ..\n>> URL: git://foo.org/bar.git\n>> Created/Earliest commit: ...\n>> Last modified/Latest commit: ...\n>> Total object count: .... (..commits, ..files, .. directories)\n>> Total repository size (compressed): ... MiB\n>> Branches:\n>> [git branch -va] + branch size\n>\n>> The error message doesn't really know whether it is going to overwrite\n>> it (the CR comes from the server), though I suppose an extra LF wouldn't\n>> hurt there.\n> Definitely wouldn't hurt.\n>\n>>> - would be nice to be able to tell git fetch to get the next chunk of\n>>> say 500 commits instead of trying to receive ALL commits, then b0rking\n>>> after umpteen percent on server timeout. Not?\n>> You asked for the current state of the repository, and that's what its\n>> giving you.\n> And instead, I would rather like to ask for the next 500 commits. No way\n> to do it.\n\nDo you mean that there are no tags in between your current state and the\none you want to be at?\n\n>\n>> The timeout has nothing to do with git, if you can't\n>> convince the admins to increase it, you can try using another transport\n>> which doesn't suffer from HTTP, as it's most likely an anti-DoS measure.\n> See, I probably can't convince the admins to drop their anti-dos measures.\n> And they (drupal.org admins) probably will not change their allowed\n> protocol policies.\n\nSwitch to using the raw git protocol, which is much less likely to have\nthis sort of measure.\n\n> Despite that, i've had timeouts or simply stale connections dying down\n> before with other repositories and various transport modes.\n> The easiest fix would be an option to tell git to not fetch everything...\n>\n>> If you want to download it bit by bit, you can tell fetch to download\n>> particular tags.\n> ..without specifying specific commit tags.\n> Browsing gitweb sites to find a tag for which the fetch doesn't time out\n> is hugely inconvenient, especially on a slow line.\n\nDon't use the web then. Use ls-remote to see what's at the other end.\n\n>\n>> Doing this automatically for this would be working\n>> around a configuration issue for a particular server, which is generally\n>> better fixed in other ways.\n> It is not only a configuration issue for one particular server. Git in\n> general is hardly usable on slow lines because\n> - it doesn't show the volume of data that is to be downloaded!\n\nHow would showing the amount of data help your connection?\n\n> - it doesn't allow the user to sync up in steps the circumstances will\n> allow to succeed.\n\nThis is unfortunate is some circunstances, but you haven't shown that\nyours is one of these.\n\n\n   cmn\n"}]}