{"thread":{"id":"16745","subject":"git-clone --how-much-disk-space-will-this-cost-me? [--depth n]","startedAt":"2008-12-15T23:53:42Z","lastAt":"2008-12-17T16:56:45Z","messageCount":12,"participants":["jidanni@jidanni.org","Jean-Luc Herren","Jeff King","Nicolas Pitre","Shawn O. Pearce"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"98012","messageId":"878wqhxaex.fsf@jidanni.org","threadId":"16745","inReplyTo":null,"subject":"git-clone --how-much-disk-space-will-this-cost-me? [--depth n]","fromName":"","fromEmail":"jidanni@jidanni.org","sentAt":"2008-12-15T23:53:42Z","receivedAt":"2008-12-15T23:53:42Z","isPatch":false,"sender":{"key":"jidanni@jidanni.org","avatar":"https://gravatar.com/avatar/36568d4af4c8d3e71627ef3b8c8d00e39065b12f29676cccd38ced75e68fa2a6?d=mp&s=160"},"body":"The git-clone manpage should mention how to determine how much disk\nspace will be used.\n\nYou see we beginners (who haven't learned git yet, so no patches\nforthcoming, thank you) are often told \"Just do git-clone\ngit://git.example.org/bla/ to get started!\". Being smart, we read up on\n--depth 1 to limit potential disk occupation, but we still have no\nidea of how much disk space we will need. We cant just use HEAD(1)\nbecause this is not HTTP.\n\nTherefore the git-clone man page, one of the main entry points for the\nbeginner, should say how to determine how much disk space we will need\nfor git-clone or git-clone --depth 1 etc.\n\nAnd don't tell us to just figure it out from the progress messages\nafter the download begins, and hit ^C if we don't like it.\n\nLet's take a look at those messages while were at it,\n$ git-clone --depth 1 git://git.sv.gnu.org/coreutils/\nInitialized empty Git repository in /usr/local/src/jidanni/coreutils/.git/\nremote: Counting objects: 26240, done.\nremote: Compressing objects: 100% (14001/14001), done.\nremote: Total 26240 (delta 21577), reused 15354 (delta 12095)\nReceiving objects: 100% (26240/26240), 15.76 MiB | 26 KiB/s, done.\nResolving deltas: 100% (21577/21577), done.\n$ du -sh\n27M  .\nNope, nowhere does it directly say \"You Holmes, are in for 27\nMegabytes (on your piddly modem)\". There obviously is math involved to\nfigure it out... math!\n\nAlso add examples of how one first probes a remote tree one has been\ntold about, determines what parts of it he might want, and then\nfinally git-clones just those parts.\n\nAlso document what --depth 0 or even -1 will do.\n"},{"id":"98015","messageId":"4946F4D9.8050803@gmx.ch","threadId":"16745","inReplyTo":"878wqhxaex.fsf@jidanni.org","subject":"Re: git-clone --how-much-disk-space-will-this-cost-me? [--depth n]","fromName":"Jean-Luc Herren","fromEmail":"jlh@gmx.ch","sentAt":"2008-12-16T00:22:49Z","receivedAt":"2008-12-16T00:22:49Z","isPatch":false,"sender":{"key":"jlh@gmx.ch","avatar":null},"body":"Hi!\n\njidanni@jidanni.org wrote:\n> The git-clone manpage should mention how to determine how much disk\n> space will be used.\n> [...]\n> And don't tell us to just figure it out from the progress messages\n> after the download begins, and hit ^C if we don't like it.\n\nMaybe that's a dumb answer, but... why not?  This works pretty\nwell for me.\n\n> Nope, nowhere does it directly say \"You Holmes, are in for 27\n> Megabytes (on your piddly modem)\". There obviously is math involved to\n> figure it out... math!\n\nSo maybe what you really want is an ETA display during the cloning\nprocess?  Sounds like a good idea to me.\n\njlh\n"},{"id":"98016","messageId":"87zlixvtu9.fsf@jidanni.org","threadId":"16745","inReplyTo":"4946F4D9.8050803@gmx.ch","subject":"Re: git-clone --how-much-disk-space-will-this-cost-me? [--depth n]","fromName":"","fromEmail":"jidanni@jidanni.org","sentAt":"2008-12-16T00:37:02Z","receivedAt":"2008-12-16T00:37:02Z","isPatch":false,"sender":{"key":"jidanni@jidanni.org","avatar":"https://gravatar.com/avatar/36568d4af4c8d3e71627ef3b8c8d00e39065b12f29676cccd38ced75e68fa2a6?d=mp&s=160"},"body":">> And don't tell us to just figure it out from the progress messages\n>> after the download begins, and hit ^C if we don't like it.\n\nJH> Maybe that's a dumb answer, but... why not?  This works pretty\nJH> well for me.\n\nSounds like my last marriage. \"Just hit ^C if you don't like it\". How\ndo you think the in-laws will feel? Nope, plan ahead I now say.\n\nJH> So maybe what you really want is an ETA display during the cloning\nJH> process?  Sounds like a good idea to me.\n\nETA implies that git has an estimate of what is going to happen.\n\nThe key is to now allow the user to get such an estimate too, before\ndeciding to git-clone or not.\n"},{"id":"98017","messageId":"20081216004339.GA3679@coredump.intra.peff.net","threadId":"16745","inReplyTo":"878wqhxaex.fsf@jidanni.org","subject":"Re: git-clone --how-much-disk-space-will-this-cost-me? [--depth n]","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-12-16T00:43:40Z","receivedAt":"2008-12-16T00:43:40Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Dec 16, 2008 at 07:53:42AM +0800, jidanni@jidanni.org wrote:\n\n> The git-clone manpage should mention how to determine how much disk\n> space will be used.\n\nOK. Do you have a suggestion for how to figure that out?\n\n> Let's take a look at those messages while were at it,\n> $ git-clone --depth 1 git://git.sv.gnu.org/coreutils/\n> Initialized empty Git repository in /usr/local/src/jidanni/coreutils/.git/\n> remote: Counting objects: 26240, done.\n> remote: Compressing objects: 100% (14001/14001), done.\n> remote: Total 26240 (delta 21577), reused 15354 (delta 12095)\n> Receiving objects: 100% (26240/26240), 15.76 MiB | 26 KiB/s, done.\n> Resolving deltas: 100% (21577/21577), done.\n> $ du -sh\n> 27M  .\n> Nope, nowhere does it directly say \"You Holmes, are in for 27\n> Megabytes (on your piddly modem)\". There obviously is math involved to\n> figure it out... math!\n\nThat's because we don't know that it will be 27 megabytes. That progress\ncounter is counting the number of _objects_, not bytes. So you can make\na rough estimate, but only after receiving some objects, and even then\nit can be wildly off (because you are assuming the size of the objects\nstill to get averages the same as the size of the objects you have\nalready gotten).\n\nAFAIK, nowhere in the sent data is there an indication of how many bytes\nare in the resulting pack (and in many cases, the pack is generated on\nthe fly and the information not only is not sent, but is not available\nanywhere).\n\n-Peff\n"},{"id":"98028","messageId":"49470D65.40808@gmx.ch","threadId":"16745","inReplyTo":"87zlixvtu9.fsf@jidanni.org","subject":"Re: git-clone --how-much-disk-space-will-this-cost-me? [--depth n]","fromName":"Jean-Luc Herren","fromEmail":"jlh@gmx.ch","sentAt":"2008-12-16T02:07:33Z","receivedAt":"2008-12-16T02:07:33Z","isPatch":false,"sender":{"key":"jlh@gmx.ch","avatar":null},"body":"jidanni@jidanni.org wrote:\n> JH> So maybe what you really want is an ETA display during the cloning\n> JH> process?  Sounds like a good idea to me.\n> \n> ETA implies that git has an estimate of what is going to happen.\n\nAren't you implying this too from the beginning?  But reading\nJeff's reply, there seems to be a reason why there isn't an ETA\nalready.\n\nHowever, since some repositories get cloned in the same way very\noften, there could be some cache that keeps these size information\naround for any subsequent identical clones.  The server could then\nsend a hint about the expected amount of data at the beginning.\n\njlh\n"},{"id":"98042","messageId":"alpine.LFD.2.00.0812160039180.30035@xanadu.home","threadId":"16745","inReplyTo":"49470D65.40808@gmx.ch","subject":"Re: git-clone --how-much-disk-space-will-this-cost-me? [--depth n]","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-12-16T05:45:56Z","receivedAt":"2008-12-16T05:45:56Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 16 Dec 2008, Jean-Luc Herren wrote:\n\n> jidanni@jidanni.org wrote:\n> > JH> So maybe what you really want is an ETA display during the cloning\n> > JH> process?  Sounds like a good idea to me.\n> > \n> > ETA implies that git has an estimate of what is going to happen.\n> \n> Aren't you implying this too from the beginning?  But reading\n> Jeff's reply, there seems to be a reason why there isn't an ETA\n> already.\n> \n> However, since some repositories get cloned in the same way very\n> often, there could be some cache that keeps these size information\n> around for any subsequent identical clones.  The server could then\n> send a hint about the expected amount of data at the beginning.\n\nAnd then you'll end up being the unlucky bastard to be the first to \nclones the new latest revision of a repository, and ETA won't be \navailable, and you'll complain about the fact that sometimes it is there \nand sometimes it is not.\n\nThe fact is, fundamentally, we don't know how many bytes to push when \ngenerating a pack to answer the clone request.  Sometimes we _could_ but \nnot always.  It is therefore better to be consistent and let people know \nthat there is simply no ETA.\n\n\nNicolas\n"},{"id":"98147","messageId":"20081217154407.GZ32487@spearce.org","threadId":"16745","inReplyTo":"alpine.LFD.2.00.0812160039180.30035@xanadu.home","subject":"Re: git-clone --how-much-disk-space-will-this-cost-me? [--depth n]","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-12-17T15:44:07Z","receivedAt":"2008-12-17T15:44:07Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Nicolas Pitre <nico@cam.org> wrote:\n> On Tue, 16 Dec 2008, Jean-Luc Herren wrote:\n> > jidanni@jidanni.org wrote:\n> > > JH> So maybe what you really want is an ETA display during the cloning\n> > > JH> process?  Sounds like a good idea to me.\n> \n> And then you'll end up being the unlucky bastard to be the first to \n> clones the new latest revision of a repository, and ETA won't be \n> available, and you'll complain about the fact that sometimes it is there \n> and sometimes it is not.\n> \n> The fact is, fundamentally, we don't know how many bytes to push when \n> generating a pack to answer the clone request.  Sometimes we _could_ but \n> not always.  It is therefore better to be consistent and let people know \n> that there is simply no ETA.\n\nHmm.\n\nWhat if on an initial clone (no \"have\" lines received) we sum up\nthe sizes of the *.pack and all of the loose objects and sent\nthat as an initial size estimate.  Its going to be the upper bound\nof the final pack that we send.  At worst it over-estimates on the\nsize and download finishes faster.\n\nI'm willing to bet that most of the \"big\" repositories out there\ndon't have a lot of garbage in them.  Linus' kernel repository\ndoesn't rewind, so he has 0 garbage.  Anyone cloning from him would\nget a reasonable estimate.  Likewise with a Gentoo/KDE/WebKit/gcc\nsort of giant tree most of that is in a huge historical pack.\nThat one pack file alone is completely reachable and dominates the\ntransfer size.\n\nOn smaller trees where people may have a lot of rebase garbage or\neverything is loose the estimate will be quite a bit above what we\ntransfer, but how much so that it matters?\n\nYea, a single stray binary of some *.mpg or *.iso accidentally\nadded and then removed (and now unreachable) will vastly inflate\nthe numbers.  In which case the repository owner will be encouraged\nto prune when people won't clone his estimated 8 GiB download,\nwhich is actually only 1 MiB.\n\n-- \nShawn.\n"},{"id":"98150","messageId":"alpine.LFD.2.00.0812171104340.30035@xanadu.home","threadId":"16745","inReplyTo":"20081217154407.GZ32487@spearce.org","subject":"Re: git-clone --how-much-disk-space-will-this-cost-me? [--depth n]","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-12-17T16:15:56Z","receivedAt":"2008-12-17T16:15:56Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 17 Dec 2008, Shawn O. Pearce wrote:\n\n> Nicolas Pitre <nico@cam.org> wrote:\n> > The fact is, fundamentally, we don't know how many bytes to push when \n> > generating a pack to answer the clone request.  Sometimes we _could_ but \n> > not always.  It is therefore better to be consistent and let people know \n> > that there is simply no ETA.\n> \n> Hmm.\n> \n> What if on an initial clone (no \"have\" lines received) we sum up\n> the sizes of the *.pack and all of the loose objects and sent\n> that as an initial size estimate.  Its going to be the upper bound\n> of the final pack that we send.  At worst it over-estimates on the\n> size and download finishes faster.\n\nIt is a kludge.  It makes the system imprecise for little benefit. Once \nyou start adding kludges like that into your system, people will always \nask for more kludges, and in the end your system isn't as reliable.  We \nall know about some other operating system which was designed like that. \nI personally don't want to go there.\n\n> Yea, a single stray binary of some *.mpg or *.iso accidentally\n> added and then removed (and now unreachable) will vastly inflate\n> the numbers.  In which case the repository owner will be encouraged\n> to prune when people won't clone his estimated 8 GiB download,\n> which is actually only 1 MiB.\n\nAnd I consider any system doing such thing completely stupid.  Either \nyou consistently know the information or you don't.  When you don't, it \nis best to not create expectations for the user.  And so far I think \nthat 99.9% of git users are just fine with the progress display we \ncurrently provide.\n\n\nNicolas\n"},{"id":"98151","messageId":"20081217162127.GG32487@spearce.org","threadId":"16745","inReplyTo":"alpine.LFD.2.00.0812171104340.30035@xanadu.home","subject":"Re: git-clone --how-much-disk-space-will-this-cost-me? [--depth n]","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-12-17T16:21:27Z","receivedAt":"2008-12-17T16:21:27Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Nicolas Pitre <nico@cam.org> wrote:\n> \n> And I consider any system doing such thing completely stupid.  Either \n> you consistently know the information or you don't.  When you don't, it \n> is best to not create expectations for the user.  And so far I think \n> that 99.9% of git users are just fine with the progress display we \n> currently provide.\n\nCertainly true here; I never care how big the source I'm cloning is.\nBut then again I have pretty good network connectivity at work\nand at least cable modem service at home...  most things clone down\npretty fast.\n\nIts a quick hack to give a size upper bound.  I don't think its\nthat ugly.  Our network protocol is uglier with all of its hidden\nfields jammed behind that NUL in the first advertisement line.\nBut I digress.\n\nThe better feature is probably resumable clone anyway.  At least\nthen people can abort a \"long running\" clone and have a good chance\nthey can pick it up again in the near future.  Its also not easy to\nimplement, which is why we've only been talking about it for years\nand never actually seen a patch proposing to do it.\n\n-- \nShawn.\n"},{"id":"98152","messageId":"alpine.LFD.2.00.0812171136250.30035@xanadu.home","threadId":"16745","inReplyTo":"20081217162127.GG32487@spearce.org","subject":"Re: git-clone --how-much-disk-space-will-this-cost-me? [--depth n]","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-12-17T16:46:42Z","receivedAt":"2008-12-17T16:46:42Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 17 Dec 2008, Shawn O. Pearce wrote:\n\n> Nicolas Pitre <nico@cam.org> wrote:\n> > \n> > And I consider any system doing such thing completely stupid.  Either \n> > you consistently know the information or you don't.  When you don't, it \n> > is best to not create expectations for the user.  And so far I think \n> > that 99.9% of git users are just fine with the progress display we \n> > currently provide.\n> \n> Certainly true here; I never care how big the source I'm cloning is.\n> But then again I have pretty good network connectivity at work\n> and at least cable modem service at home...  most things clone down\n> pretty fast.\n> \n> Its a quick hack to give a size upper bound.  I don't think its\n> that ugly.  Our network protocol is uglier with all of its hidden\n> fields jammed behind that NUL in the first advertisement line.\n> But I digress.\n\nThe ugliness in the protocol is encapsulated away from user view, and we \ncould even seemlessly introduce a new protocol at any time with no \nissues if we wanted to.\n\nThis \"quick hack\" is imprecise, unreliable, and directly affect user \nperception.  This is way more dammageable as once users are used to it, \ngood or bad, it won't be possible to get rid of it.\n\n> The better feature is probably resumable clone anyway.  At least\n> then people can abort a \"long running\" clone and have a good chance\n> they can pick it up again in the near future.\n\nAbsolutely.\n\n> Its also not easy to\n> implement, which is why we've only been talking about it for years\n> and never actually seen a patch proposing to do it.\n\nA partial clone could possibly be turned into a shalow clone if at least \nthe top commit is complete ...\n\n\nNicolas\n"},{"id":"98153","messageId":"20081217164841.GH32487@spearce.org","threadId":"16745","inReplyTo":"alpine.LFD.2.00.0812171136250.30035@xanadu.home","subject":"Re: git-clone --how-much-disk-space-will-this-cost-me? [--depth n]","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-12-17T16:48:41Z","receivedAt":"2008-12-17T16:48:41Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Nicolas Pitre <nico@cam.org> wrote:\n> > Its also not easy to\n> > implement, which is why we've only been talking about it for years\n> > and never actually seen a patch proposing to do it.\n> \n> A partial clone could possibly be turned into a shalow clone if at least \n> the top commit is complete ...\n\nBut you of all people should know well that the top commit is also\na huge part of most clones.  Getting that top commit can be 30-60%\nof the repository itself.  :-|\n\n-- \nShawn.\n"},{"id":"98154","messageId":"alpine.LFD.2.00.0812171152490.30035@xanadu.home","threadId":"16745","inReplyTo":"20081217164841.GH32487@spearce.org","subject":"Re: git-clone --how-much-disk-space-will-this-cost-me? [--depth n]","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-12-17T16:56:45Z","receivedAt":"2008-12-17T16:56:45Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 17 Dec 2008, Shawn O. Pearce wrote:\n\n> Nicolas Pitre <nico@cam.org> wrote:\n> > > Its also not easy to\n> > > implement, which is why we've only been talking about it for years\n> > > and never actually seen a patch proposing to do it.\n> > \n> > A partial clone could possibly be turned into a shalow clone if at least \n> > the top commit is complete ...\n> \n> But you of all people should know well that the top commit is also\n> a huge part of most clones.  Getting that top commit can be 30-60%\n> of the repository itself.  :-|\n\nSure I know.  This is why I'm not pushing this solution really much.  ;)\n\nI have ideas about how to solve this in a really nice way, but that \nimplies pack V4.\n\n\nNicolas\n"}]}