{"thread":{"id":"41590","subject":"Resumable git clone?","startedAt":"2016-03-02T01:30:56Z","lastAt":"2016-03-24T21:06:37Z","messageCount":24,"participants":["Josh Triplett","Stefan Beller","Duy Nguyen","Al Viro","Junio C Hamano","Jeff King","Bhavik Bavishi","Konstantin Ryabitsev","Philip Oakley"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"280027","messageId":"20160302012922.GA17114@jtriplet-mobl2.jf.intel.com","threadId":"41590","inReplyTo":null,"subject":"Resumable git clone?","fromName":"Josh Triplett","fromEmail":"josh@joshtriplett.org","sentAt":"2016-03-02T01:30:56Z","receivedAt":"2016-03-02T01:30:56Z","isPatch":false,"sender":{"key":"josh@joshtriplett.org","avatar":"https://avatars.githubusercontent.com/u/162737?v=4"},"body":"If you clone a repository, and the connection drops, the next attempt\nwill have to start from scratch.  This can add significant time and\nexpense if you're on a low-bandwidth or metered connection trying to\nclone something like Linux.\n\nWould it be possible to make git clone resumable after a partial clone?\n(And, ideally, to make that the default?)\n\nIn a discussion elsewhere, Al Viro suggested taking the partial pack\nreceived so far, repairing any truncation, indexing the objects it\ncontains, and then re-running clone and not having to fetch those\nobjects.  This may also require extending receive-pack's protocol for\ndetermining objects the recipient already has, as the partial pack may\nnot have a consistent set of reachable objects.\n\nBefore starting down the path of developing patches for this, does the\napproach seem potentially reasonable?\n\n- Josh Triplett\n"},{"id":"280028","messageId":"CAGZ79kYjuaOiTCC-NnZDQs=XGbgXWhJe7gk576jod4QnV57eEg@mail.gmail.com","threadId":"41590","inReplyTo":"20160302012922.GA17114@jtriplet-mobl2.jf.intel.com","subject":"Re: Resumable git clone?","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2016-03-02T01:40:28Z","receivedAt":"2016-03-02T01:40:28Z","isPatch":false,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"+ Duy, who tried resumable clone a few days/weeks ago\n\nOn Tue, Mar 1, 2016 at 5:30 PM, Josh Triplett <josh@joshtriplett.org> wrote:\n> If you clone a repository, and the connection drops, the next attempt\n> will have to start from scratch.  This can add significant time and\n> expense if you're on a low-bandwidth or metered connection trying to\n> clone something like Linux.\n>\n> Would it be possible to make git clone resumable after a partial clone?\n> (And, ideally, to make that the default?)\n>\n> In a discussion elsewhere, Al Viro suggested taking the partial pack\n> received so far,\n\nok,\n\n> repairing any truncation,\n\nSo throwing away half finished stuff while keeping the front load?\n\n> indexing the objects it\n> contains, and then re-running clone and not having to fetch those\n> objects.\n\nThe pack is not deterministic for a given repository. When creating\nthe pack, you may encounter races between threads, such that the order\nin a pack differs.\n\n> This may also require extending receive-pack's protocol for\n> determining objects the recipient already has, as the partial pack may\n> not have a consistent set of reachable objects.\n>\n> Before starting down the path of developing patches for this, does the\n> approach seem potentially reasonable?\n\nI think that sounds reasonable on a high level, but I'd expect it blows up\nin complexity as in the receive-pack's protocol or in the code for having\nto handle partial stuff.\n\nThanks,\nStefan\n\n>\n> - Josh Triplett\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"280029","messageId":"CACsJy8B6_mRpw7ADyZ3H5vWq=JzEUX0yRHJM7pqQgCPQbvhOwA@mail.gmail.com","threadId":"41590","inReplyTo":"20160302012922.GA17114@jtriplet-mobl2.jf.intel.com","subject":"Re: Resumable git clone?","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2016-03-02T01:45:04Z","receivedAt":"2016-03-02T01:45:04Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Mar 2, 2016 at 8:30 AM, Josh Triplett <josh@joshtriplett.org> wrote:\n> If you clone a repository, and the connection drops, the next attempt\n> will have to start from scratch.  This can add significant time and\n> expense if you're on a low-bandwidth or metered connection trying to\n> clone something like Linux.\n>\n> Would it be possible to make git clone resumable after a partial clone?\n> (And, ideally, to make that the default?)\n>\n> In a discussion elsewhere, Al Viro suggested taking the partial pack\n> received so far, repairing any truncation, indexing the objects it\n> contains, and then re-running clone and not having to fetch those\n> objects.  This may also require extending receive-pack's protocol for\n> determining objects the recipient already has, as the partial pack may\n> not have a consistent set of reachable objects.\n>\n> Before starting down the path of developing patches for this, does the\n> approach seem potentially reasonable?\n\nThis topic came up recently (thanks Sarah!) and Shawn proposed a\ndifferent approach that (I think) is simpler and more effective for\nresume _clone_ case. I'm not sure if anybody is implementing it\nthough.\n\n[1] http://thread.gmane.org/gmane.comp.version-control.git/285921\n-- \nDuy\n"},{"id":"280030","messageId":"20160302023024.GG17997@ZenIV.linux.org.uk","threadId":"41590","inReplyTo":"CAGZ79kYjuaOiTCC-NnZDQs=XGbgXWhJe7gk576jod4QnV57eEg@mail.gmail.com","subject":"Re: Resumable git clone?","fromName":"Al Viro","fromEmail":"viro@zeniv.linux.org.uk","sentAt":"2016-03-02T02:30:24Z","receivedAt":"2016-03-02T02:30:24Z","isPatch":false,"sender":{"key":"viro@zeniv.linux.org.uk","avatar":null},"body":"On Tue, Mar 01, 2016 at 05:40:28PM -0800, Stefan Beller wrote:\n\n> So throwing away half finished stuff while keeping the front load?\n\nThrow away the object that got truncated and ones for which delta chain\ndoesn't resolve entirely in the transferred part.\n \n> > indexing the objects it\n> > contains, and then re-running clone and not having to fetch those\n> > objects.\n> \n> The pack is not deterministic for a given repository. When creating\n> the pack, you may encounter races between threads, such that the order\n> in a pack differs.\n\nFWIW, I wasn't proposing to recreate the remaining bits of that _pack_;\njust do the normal pull with one addition: start with sending the list\nof sha1 of objects you are about to send and let the recepient reply\nwith \"I already have <set of sha1>, don't bother with those\".  And exclude\nthose from the transfer.  Encoding for the set being available is an\ninteresting variable here - might be plain list of sha1, might be its\ncomplement (\"I want the following subset\"), might be \"145th to 1029th,\n1517th and 1890th to 1920th of the list you've sent\"; which form ends\nup more efficient needs to be found experimentally...\n\nIIRC, the objection had been that the organisation of the pack will lead\nto many cases when deltas are transferred *first*, with base object not\ngetting there prior to disconnect.  I suspect that fraction of the objects\ngetting through would still be worth it, but I hadn't experimented enough\nto be able to tell...\n\nI was more interested in resumable _pull_, with restarted clone treated as\nspecial case of that.\n"},{"id":"280036","messageId":"xmqq8u215r25.fsf@gitster.mtv.corp.google.com","threadId":"41590","inReplyTo":"20160302023024.GG17997@ZenIV.linux.org.uk","subject":"Re: Resumable git clone?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-03-02T06:31:30Z","receivedAt":"2016-03-02T06:31:30Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Al Viro <viro@ZenIV.linux.org.uk> writes:\n\n> FWIW, I wasn't proposing to recreate the remaining bits of that _pack_;\n> just do the normal pull with one addition: start with sending the list\n> of sha1 of objects you are about to send and let the recepient reply\n> with \"I already have <set of sha1>, don't bother with those\".  And exclude\n> those from the transfer.\n\nI did a quick-and-dirty unscientific experiment.\n\nI had a clone of Linus's repository that was about a week old, whose\ntip was at 4de8ebef (Merge tag 'trace-fixes-v4.5-rc5' of\ngit://git.kernel.org/pub/scm/linux/kernel/git/rostedt/linux-trace,\n2016-02-22).  To bring it up to date (i.e. a pull about a week's\nworth of progress) to f691b77b (Merge branch 'for-linus' of\ngit://git.kernel.org/pub/scm/linux/kernel/git/viro/vfs, 2016-03-01):\n\n    $ git rev-list --objects 4de8ebef..f691b77b1fc | wc -l\n    1396\n    $ git rev-parse 4de8ebef..f691b77b1fc |\n      git pack-objects --revs --delta-base-offset --stdout |\n      wc -c\n    2444127\n\nSo in order to salvage some transfer out of 2.4MB, the hypothetical\nAl protocol would first have the upload-pack give 20*1396 = 28kB\nobject names to fetch-pack; no matter how fetch-pack encodes its\npreference, its answer would be less than 28kB.  We would likely to\ndesign this part of the new protocol in line with the existing part\nand use textual object names, so let's round them up to 100kB.\n\nThat is quite small, even if you are on a crappy connection that you\nneed to retry 5 times, the additional overhead to negotiate the list\nof objects alone would be 0.5MB (or less than 20% of the real\ntransfer).\n\nThat is quite interesting [*1*].\n\nFor the approach to be practical, you would have to write a program\nthat reads from a truncated packfile and writes a new packfile,\nexcising deltas that lack their bases, to salvage objects from a\nhalf-transferred packfile; it is however unclear how involved the\ncode would get.\n\nIt is probably OK for a tiny pack that has only 1400 objects--we\ncould just pass the early part through unpack-objects and let it die\nwhen it hits EOF, but for a \"resumable clone\", I do not think you\ncan afford to unpack 4.6M objects in the kernel repository into\nloose objects.\n\nThe approach of course requires the server end to spend 5 times as\nmany cycles as usual in order to help a client that retries 5 times.\n\nOn the other hand, the resumable \"clone\" we were discussing by\nallowing the server to respond with a slightly older bundle or a\npack and then asking the client to fill the latest bits by a\nfollow-up fetch targets to reduce the load of the server side (the\n\"slightly older\" part can be offloaded to CDN).  It is a happy side\neffect that material offloaded to CDN can more easily obtained via\nHTTPS that is trivially resumable ;-)\n\nI think your \"I've got these already\" extention may be worth trying,\nand it is definitely better than the \"let's make sure the server end\ncreates byte-for-byte identical pack stream, and discard the early\npart without sending it to the network\", and it may help resuming a\nsmall incremental fetch, but I do not think it is advisable to use\nit for a full clone, given that it is very likely that we would be\nadding the \"offload 'clone' to CDN\" kind.  Even though I can foresee\nboth kinds to co-exist, I do not think it is practical to offer it\nfor resuming multi-hour cloning of the kernel repository (or worse,\nAndroid repositories) over a trans-Pacific link, for example.\n\n\n[Footnote]\n\n*1* To update v4.5-rc1 to today's HEAD involves 10809 objects, and\n    the pack data takes 14955728 bytes.  That translates to ~440kB\n    needed to advertise a list of textual object names to salvage\n    object transfer of 15MB.\n"},{"id":"280037","messageId":"CACsJy8DcNrOmrKKPibV6GuSqspovBmHzUv_mRB6fZyLjw5wWzQ@mail.gmail.com","threadId":"41590","inReplyTo":"xmqq8u215r25.fsf@gitster.mtv.corp.google.com","subject":"Re: Resumable git clone?","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2016-03-02T07:37:53Z","receivedAt":"2016-03-02T07:37:53Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Mar 2, 2016 at 1:31 PM, Junio C Hamano <gitster@pobox.com> wrote:\n> Al Viro <viro@ZenIV.linux.org.uk> writes:\n>\n>> FWIW, I wasn't proposing to recreate the remaining bits of that _pack_;\n>> just do the normal pull with one addition: start with sending the list\n>> of sha1 of objects you are about to send and let the recepient reply\n>> with \"I already have <set of sha1>, don't bother with those\".  And exclude\n>> those from the transfer.\n>\n> I did a quick-and-dirty unscientific experiment.\n>\n> I had a clone of Linus's repository that was about a week old, whose\n> tip was at 4de8ebef (Merge tag 'trace-fixes-v4.5-rc5' of\n> git://git.kernel.org/pub/scm/linux/kernel/git/rostedt/linux-trace,\n> 2016-02-22).  To bring it up to date (i.e. a pull about a week's\n> worth of progress) to f691b77b (Merge branch 'for-linus' of\n> git://git.kernel.org/pub/scm/linux/kernel/git/viro/vfs, 2016-03-01):\n>\n>     $ git rev-list --objects 4de8ebef..f691b77b1fc | wc -l\n>     1396\n>     $ git rev-parse 4de8ebef..f691b77b1fc |\n>       git pack-objects --revs --delta-base-offset --stdout |\n>       wc -c\n>     2444127\n>\n> So in order to salvage some transfer out of 2.4MB, the hypothetical\n> Al protocol would first have the upload-pack give 20*1396 = 28kB\n\nIt could be 10*1396 or less. If the server calculates the shortest\nunambiguous SHA-1 length (quite cheap on fully packed repo) and sends\nit to the client, the client can just sends short SHA-1 instead. It's\nracy though because objects are being added to the server and abbrev\nlength may go up. But we can check ambiguity for all SHA-1 sent by\nclient and ask for resend for ambiguous ones.\n\nOn my linux-2.6.git, 10 letters (so 5 bytes) are needed for\nunambiguous short SHA-1. But we can even go optimistic and ask the\nclient for shorter SHA-1 with hope that resend won't be many.\n\n> object names to fetch-pack; no matter how fetch-pack encodes its\n> preference, its answer would be less than 28kB.  We would likely to\n> design this part of the new protocol in line with the existing part\n> and use textual object names, so let's round them up to 100kB.\n-- \nDuy\n"},{"id":"280038","messageId":"CACsJy8AFXJBc8awQ6uNwgzMjOn9v_+yE9t+bR2Bv9f1kwGw0Yg@mail.gmail.com","threadId":"41590","inReplyTo":"CACsJy8DcNrOmrKKPibV6GuSqspovBmHzUv_mRB6fZyLjw5wWzQ@mail.gmail.com","subject":"Re: Resumable git clone?","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2016-03-02T07:44:00Z","receivedAt":"2016-03-02T07:44:00Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Mar 2, 2016 at 2:37 PM, Duy Nguyen <pclouds@gmail.com> wrote:\n>> So in order to salvage some transfer out of 2.4MB, the hypothetical\n>> Al protocol would first have the upload-pack give 20*1396 = 28kB\n>\n> It could be 10*1396 or less....\n\nOops somehow I read previous mails as client sends SHA-1 to server,\nnot the other way around that you and Al were talking about. But the\nsame principle applies to the other direction, I think.\n-- \nDuy\n"},{"id":"280039","messageId":"20160302075437.GA8024@x","threadId":"41590","inReplyTo":"CACsJy8DcNrOmrKKPibV6GuSqspovBmHzUv_mRB6fZyLjw5wWzQ@mail.gmail.com","subject":"Re: Resumable git clone?","fromName":"Josh Triplett","fromEmail":"josh@joshtriplett.org","sentAt":"2016-03-02T07:54:37Z","receivedAt":"2016-03-02T07:54:37Z","isPatch":false,"sender":{"key":"josh@joshtriplett.org","avatar":"https://avatars.githubusercontent.com/u/162737?v=4"},"body":"On Wed, Mar 02, 2016 at 02:37:53PM +0700, Duy Nguyen wrote:\n> On Wed, Mar 2, 2016 at 1:31 PM, Junio C Hamano <gitster@pobox.com> wrote:\n> > Al Viro <viro@ZenIV.linux.org.uk> writes:\n> >\n> >> FWIW, I wasn't proposing to recreate the remaining bits of that _pack_;\n> >> just do the normal pull with one addition: start with sending the list\n> >> of sha1 of objects you are about to send and let the recepient reply\n> >> with \"I already have <set of sha1>, don't bother with those\".  And exclude\n> >> those from the transfer.\n> >\n> > I did a quick-and-dirty unscientific experiment.\n> >\n> > I had a clone of Linus's repository that was about a week old, whose\n> > tip was at 4de8ebef (Merge tag 'trace-fixes-v4.5-rc5' of\n> > git://git.kernel.org/pub/scm/linux/kernel/git/rostedt/linux-trace,\n> > 2016-02-22).  To bring it up to date (i.e. a pull about a week's\n> > worth of progress) to f691b77b (Merge branch 'for-linus' of\n> > git://git.kernel.org/pub/scm/linux/kernel/git/viro/vfs, 2016-03-01):\n> >\n> >     $ git rev-list --objects 4de8ebef..f691b77b1fc | wc -l\n> >     1396\n> >     $ git rev-parse 4de8ebef..f691b77b1fc |\n> >       git pack-objects --revs --delta-base-offset --stdout |\n> >       wc -c\n> >     2444127\n> >\n> > So in order to salvage some transfer out of 2.4MB, the hypothetical\n> > Al protocol would first have the upload-pack give 20*1396 = 28kB\n> \n> It could be 10*1396 or less. If the server calculates the shortest\n> unambiguous SHA-1 length (quite cheap on fully packed repo) and sends\n> it to the client, the client can just sends short SHA-1 instead. It's\n> racy though because objects are being added to the server and abbrev\n> length may go up. But we can check ambiguity for all SHA-1 sent by\n> client and ask for resend for ambiguous ones.\n> \n> On my linux-2.6.git, 10 letters (so 5 bytes) are needed for\n> unambiguous short SHA-1. But we can even go optimistic and ask the\n> client for shorter SHA-1 with hope that resend won't be many.\n\nI don't think it's worth the trouble and ambiguity to send abbreviated\nobject names over the wire.  I think several simpler optimizations seem\npreferable, such as binary object names, and abbreviating complete\nobject sets (\"I have these commits/trees and everything they need\nrecursively; I also have this stack of random objects.\").\n\nThat would work especially well for resumable pull, or for the case of\noptimizing pull during the merge window.\n\n- Josh Triplett\n"},{"id":"280040","messageId":"20160302081344.GB8024@x","threadId":"41590","inReplyTo":"20160302023024.GG17997@ZenIV.linux.org.uk","subject":"Re: Resumable git clone?","fromName":"Josh Triplett","fromEmail":"josh@joshtriplett.org","sentAt":"2016-03-02T08:13:44Z","receivedAt":"2016-03-02T08:13:44Z","isPatch":false,"sender":{"key":"josh@joshtriplett.org","avatar":"https://avatars.githubusercontent.com/u/162737?v=4"},"body":"On Wed, Mar 02, 2016 at 02:30:24AM +0000, Al Viro wrote:\n> On Tue, Mar 01, 2016 at 05:40:28PM -0800, Stefan Beller wrote:\n> \n> > So throwing away half finished stuff while keeping the front load?\n> \n> Throw away the object that got truncated and ones for which delta chain\n> doesn't resolve entirely in the transferred part.\n>  \n> > > indexing the objects it\n> > > contains, and then re-running clone and not having to fetch those\n> > > objects.\n> > \n> > The pack is not deterministic for a given repository. When creating\n> > the pack, you may encounter races between threads, such that the order\n> > in a pack differs.\n> \n> FWIW, I wasn't proposing to recreate the remaining bits of that _pack_;\n> just do the normal pull with one addition: start with sending the list\n> of sha1 of objects you are about to send and let the recepient reply\n> with \"I already have <set of sha1>, don't bother with those\".  And exclude\n> those from the transfer.  Encoding for the set being available is an\n> interesting variable here - might be plain list of sha1, might be its\n> complement (\"I want the following subset\"), might be \"145th to 1029th,\n> 1517th and 1890th to 1920th of the list you've sent\"; which form ends\n> up more efficient needs to be found experimentally...\n\nAs a simple proposal, the server could send the list of hashes (in\napproximately the same order it would send the pack), the client could\nsend back a bitmap where '0' means \"send it\" and '1' means \"got that one\nalready\", and the client could compress that bitmap.  That gives you the\nRLE and similar without having to write it yourself.  That might not be\noptimal, but it would likely set a high bar with minimal effort.\n\nOne debatable optimization on top of that would rely on git object\nstructure to imply objects hashes without sending them: the message from\nthe server could have a list of commit/tree hashes that imply sending\nall objects reachable from those, without having to send all the implied\nhashes.  However, that would then make the message back from the client\nabout what it already has larger and more complicated; that might not\nmake it worthwhile.\n\nThis seems like a good case for doing the simplest possible thing first\n(complete hash list, compressed \"got it already\" bitmap), seeing how\nmuch benefit that provides, and creating a v2 protocol if some\nadditional optimization proves sufficiently worthwhile.\n\n- Josh Triplett\n"},{"id":"280041","messageId":"CACsJy8DSt8V1u2pWFQ6OcSWyKrVVnnnmWpvo5vk_u6QXkqkbcw@mail.gmail.com","threadId":"41590","inReplyTo":"20160302023024.GG17997@ZenIV.linux.org.uk","subject":"Re: Resumable git clone?","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2016-03-02T08:14:00Z","receivedAt":"2016-03-02T08:14:00Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Mar 2, 2016 at 9:30 AM, Al Viro <viro@zeniv.linux.org.uk> wrote:\n> IIRC, the objection had been that the organisation of the pack will lead\n> to many cases when deltas are transferred *first*, with base object not\n> getting there prior to disconnect.  I suspect that fraction of the objects\n> getting through would still be worth it, but I hadn't experimented enough\n> to be able to tell...\n\nNo. If deltas refer to  base objects by offset,  the (unsigned) offset\nis negated before use. So base objects must always sent first. If\ndeltas refer to base objects by full SHA-1 then base objects can\nappear anywhere in the pack in theory. But I think we only use full\nSHA-1 references for out-of-thin-pack objects, never to an existing\nobject in the pack.\n-- \nDuy\n"},{"id":"280042","messageId":"CACsJy8CBBk4bgz6Gn0QvCwWtOsqcQZBYgOBQTd=4Y+2YKs44Qg@mail.gmail.com","threadId":"41590","inReplyTo":"20160302081344.GB8024@x","subject":"Re: Resumable git clone?","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2016-03-02T08:22:17Z","receivedAt":"2016-03-02T08:22:17Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Mar 2, 2016 at 3:13 PM, Josh Triplett <josh@joshtriplett.org> wrote:\n> On Wed, Mar 02, 2016 at 02:30:24AM +0000, Al Viro wrote:\n>> On Tue, Mar 01, 2016 at 05:40:28PM -0800, Stefan Beller wrote:\n>>\n>> > So throwing away half finished stuff while keeping the front load?\n>>\n>> Throw away the object that got truncated and ones for which delta chain\n>> doesn't resolve entirely in the transferred part.\n>>\n>> > > indexing the objects it\n>> > > contains, and then re-running clone and not having to fetch those\n>> > > objects.\n>> >\n>> > The pack is not deterministic for a given repository. When creating\n>> > the pack, you may encounter races between threads, such that the order\n>> > in a pack differs.\n>>\n>> FWIW, I wasn't proposing to recreate the remaining bits of that _pack_;\n>> just do the normal pull with one addition: start with sending the list\n>> of sha1 of objects you are about to send and let the recepient reply\n>> with \"I already have <set of sha1>, don't bother with those\".  And exclude\n>> those from the transfer.  Encoding for the set being available is an\n>> interesting variable here - might be plain list of sha1, might be its\n>> complement (\"I want the following subset\"), might be \"145th to 1029th,\n>> 1517th and 1890th to 1920th of the list you've sent\"; which form ends\n>> up more efficient needs to be found experimentally...\n>\n> As a simple proposal, the server could send the list of hashes (in\n> approximately the same order it would send the pack), the client could\n> send back a bitmap where '0' means \"send it\" and '1' means \"got that one\n> already\", and the client could compress that bitmap.  That gives you the\n> RLE and similar without having to write it yourself.  That might not be\n> optimal, but it would likely set a high bar with minimal effort.\n\nWe have an implementation of EWAH bitmap compression, so compressing\nis not a problem.\n\nBut I still don't see why it's more efficient to have the server send\nthe hash list to the client. Assume you need to transfer N objects.\nThat direction makes you always send N hashes. But if the client sends\nthe list of already fetched objects, M, then M <= N. And we won't need\nto send the bitmap. What did I miss?\n-- \nDuy\n"},{"id":"280044","messageId":"xmqq4mcp5lij.fsf@gitster.mtv.corp.google.com","threadId":"41590","inReplyTo":"20160302075437.GA8024@x","subject":"Re: Resumable git clone?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-03-02T08:31:16Z","receivedAt":"2016-03-02T08:31:16Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Josh Triplett <josh@joshtriplett.org> writes:\n\n> I don't think it's worth the trouble and ambiguity to send abbreviated\n> object names over the wire.  \n\nYup.  My unscientific experiment was to show that the list would be\nfar smaller than the actual transfer and between full binary and\nfull textual object name representations there would not be much\nmeaningful difference--you seem to have a better design sense to\ngrasp that point ;-)\n\n> I think several simpler optimizations seem\n> preferable, such as binary object names, and abbreviating complete\n> object sets (\"I have these commits/trees and everything they need\n> recursively; I also have this stack of random objects.\").\n\nGiven the way pack stream is organized (i.e. commits first and then\ntrees and blobs that belong to the same delta chain together), and\nour assumed goal being to salvage objects from an interrupted\ntransfer of a packfile, you are unlikely to ever see \"I have these\ncommits/trees and everything they need\" that are salvaged from such\na failed transfer.  So I doubt such an optimization is worth doing.\n\nBesides it is very expensive to compute (the computation is done on\nthe client side, so the cycles burned and the time the user has to\nwait is of much less concern, though); you'd essentially be doing\n\"git fsck\" to find the \"dangling\" objects.\n\nThe list of what would be transferred needs to come in full from the\nserver end, as the list names objects that the receiving end may not\nhave seen, but the response by the client could be encoded much\ntightly.  For the full list of N objects from the server, we can\nthink of your response to be a bitstream of N bits, each on-bit in\nwhich signals an unwanted object in the list.  You can optimize this\ntransfer by RLE compressing the bitstream, for example.\n\nAs git-over-HTTP is stateless, however, you cannot assume that the\nserver side remembers what it sent to the client (instead, the\nclient side needs to re-post what it heard from the server in the\nprevious exchange to allow the server side to use it after\nvalidating).  So \"objects at these indices in your list\" kind of\noptimization may not work very well in that environment.  I'd\nimagine that an exchange of \"Here are the list of objects\", \"Give me\nthese objects\" done naively in full 40-hex object names would work\nOK there, though.\n"},{"id":"280045","messageId":"20160302083227.GA30065@sigill.intra.peff.net","threadId":"41590","inReplyTo":"CACsJy8CBBk4bgz6Gn0QvCwWtOsqcQZBYgOBQTd=4Y+2YKs44Qg@mail.gmail.com","subject":"Re: Resumable git clone?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-03-02T08:32:27Z","receivedAt":"2016-03-02T08:32:27Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Mar 02, 2016 at 03:22:17PM +0700, Duy Nguyen wrote:\n\n> > As a simple proposal, the server could send the list of hashes (in\n> > approximately the same order it would send the pack), the client could\n> > send back a bitmap where '0' means \"send it\" and '1' means \"got that one\n> > already\", and the client could compress that bitmap.  That gives you the\n> > RLE and similar without having to write it yourself.  That might not be\n> > optimal, but it would likely set a high bar with minimal effort.\n> \n> We have an implementation of EWAH bitmap compression, so compressing\n> is not a problem.\n> \n> But I still don't see why it's more efficient to have the server send\n> the hash list to the client. Assume you need to transfer N objects.\n> That direction makes you always send N hashes. But if the client sends\n> the list of already fetched objects, M, then M <= N. And we won't need\n> to send the bitmap. What did I miss?\n\nRight, I don't see what the point is in compressing the bitmap. The sha1\nlist for a clone of linux.git is 87 megabytes. The return bitmap, even\nnaively, is 500K. Unless you are trying to optimize for wildly\nasymmetric links.\n\nIf the client just naively sends \"here's what I have\", then we know it\ncan never be _more_ than 87 megabytes. And as a bonus, the longer the\nlist is, the more we are saving (so at the moment you are sending 82MB,\nit's really worth it, because you do have 95% of the pack, which is\nworth amortizing).\n\nI'm still a little dubious that anything involving \"send all the hashes\"\nis going to be useful in practice, especially for something like the\nkernel (where you have tons of huge small objects that delta well). It\nwould work better when you have gigantic objects that don't delta (so\nthe cost of a sha1 versus the object size is way better), but then I\nthink we'd do better to transfer all of the normal-sized bits up front,\nand then allow fetching the large stuff separately.\n\n-Peff\n"},{"id":"280046","messageId":"xmqqziuh46hb.fsf@gitster.mtv.corp.google.com","threadId":"41590","inReplyTo":"20160302012922.GA17114@jtriplet-mobl2.jf.intel.com","subject":"Re: Resumable git clone?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-03-02T08:41:20Z","receivedAt":"2016-03-02T08:41:20Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Josh Triplett <josh@joshtriplett.org> writes:\n\n> If you clone a repository, and the connection drops, the next attempt\n> will have to start from scratch.  This can add significant time and\n> expense if you're on a low-bandwidth or metered connection trying to\n> clone something like Linux.\n\nFor this particular issue, your friendly k.org administrator already\nhas a solution.  Torvalds/linux.git is made into a bundle weekly\nwith\n\n    $ git bundle create clone.bundle --all\n\nand the result placed on k.org CDN.  So low-bandwidth cloners can\ngrab it over resumable http, clone from the bundle, and then fill\nthe most recent part by fetching from k.org already.\n\nThe tooling to allow this kind of \"bundle\" (and possibly other forms\nof \"CDN offload\" material) transparently used by \"git clone\" was the\nproposal by Shawn Pearce mentioned elsewhere in this thread.\n"},{"id":"280056","messageId":"CACsJy8D69ieHSKTFC=0hsz1Ss+bgajxXmcyf4Ma7mrjWrp_NXA@mail.gmail.com","threadId":"41590","inReplyTo":"xmqq4mcp5lij.fsf@gitster.mtv.corp.google.com","subject":"Re: Resumable git clone?","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2016-03-02T09:28:59Z","receivedAt":"2016-03-02T09:28:59Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Mar 2, 2016 at 3:31 PM, Junio C Hamano <gitster@pobox.com> wrote:\n> Josh Triplett <josh@joshtriplett.org> writes:\n>\n>> I don't think it's worth the trouble and ambiguity to send abbreviated\n>> object names over the wire.\n>\n> Yup.  My unscientific experiment was to show that the list would be\n> far smaller than the actual transfer and between full binary and\n> full textual object name representations there would not be much\n> meaningful difference--you seem to have a better design sense to\n> grasp that point ;-)\n\nIt may matter, depending on your user target. In order to progress a\nfetch/pull, I need to get at least one object before my connection\ngoes down. Picking a random blob in the \"large file\" range in\nlinux-2.6, fs/nls/nls_cp950.c, 500kb. Let's assume the worst case that\nthe blob is transferred gzipped, not deltified, that's about 100k.\nAssume again I'm a lazy linux lurker who only fetches after every\nrelease, the rev-list output between v4.2 and v4.3 is 6M. Even if we\ntransfer this list over http with compression, the list is 2.9M, way\nbigger than one blob transfer. Which raises the bar to my successful\nfetch.\n-- \nDuy\n"},{"id":"280059","messageId":"56D6C4B0.5080003@gmail.com","threadId":"41590","inReplyTo":"20160302083227.GA30065@sigill.intra.peff.net","subject":"Re: Resumable git clone?","fromName":"Bhavik Bavishi","fromEmail":"bhavikdbavishi@gmail.com","sentAt":"2016-03-02T10:47:12Z","receivedAt":"2016-03-02T10:47:12Z","isPatch":false,"sender":{"key":"bhavikdbavishi@gmail.com","avatar":"https://gravatar.com/avatar/aaefd93ba9f5c7aa32b90ece242f4eb00076a893ef70d9cf9939000ad1c5352a?d=mp&s=160"},"body":"On 3/2/16 2:02 PM, Jeff King wrote:\n> On Wed, Mar 02, 2016 at 03:22:17PM +0700, Duy Nguyen wrote:\n>\n>>> As a simple proposal, the server could send the list of hashes (in\n>>> approximately the same order it would send the pack), the client could\n>>> send back a bitmap where '0' means \"send it\" and '1' means \"got that one\n>>> already\", and the client could compress that bitmap.  That gives you the\n>>> RLE and similar without having to write it yourself.  That might not be\n>>> optimal, but it would likely set a high bar with minimal effort.\n>>\n>> We have an implementation of EWAH bitmap compression, so compressing\n>> is not a problem.\n>>\n>> But I still don't see why it's more efficient to have the server send\n>> the hash list to the client. Assume you need to transfer N objects.\n>> That direction makes you always send N hashes. But if the client sends\n>> the list of already fetched objects, M, then M <= N. And we won't need\n>> to send the bitmap. What did I miss?\n>\n> Right, I don't see what the point is in compressing the bitmap. The sha1\n> list for a clone of linux.git is 87 megabytes. The return bitmap, even\n> naively, is 500K. Unless you are trying to optimize for wildly\n> asymmetric links.\n>\n> If the client just naively sends \"here's what I have\", then we know it\n> can never be _more_ than 87 megabytes. And as a bonus, the longer the\n> list is, the more we are saving (so at the moment you are sending 82MB,\n> it's really worth it, because you do have 95% of the pack, which is\n> worth amortizing).\n>\n> I'm still a little dubious that anything involving \"send all the hashes\"\n> is going to be useful in practice, especially for something like the\n> kernel (where you have tons of huge small objects that delta well). It\n> would work better when you have gigantic objects that don't delta (so\n> the cost of a sha1 versus the object size is way better), but then I\n> think we'd do better to transfer all of the normal-sized bits up front,\n> and then allow fetching the large stuff separately.\n>\n> -Peff\n>\n\n\nIn case if we can have object-lookup-db like provisioning with stored \ninformation like SHA-1, type of object, parent if any, size of that \nobject, as in entire hierarchy tree without data like commit message, \ntag name. This implementation may be look as bit duplication of existing \ninformation.\n\nAt initial clone time server sends object-lookup-db to client and then, \nby reading object-lookup-db client sends SHA1 to server to get/fecth \nobjects, it can be got in parallel, as well. This process may not be \ntransfer efficient but it can be resumable, as client knows what got \nsync and what's remain and which SHA1 refers to what object type.\n"},{"id":"280064","messageId":"20160302155119.GB8064@gmail.com","threadId":"41590","inReplyTo":"xmqqziuh46hb.fsf@gitster.mtv.corp.google.com","subject":"Re: Resumable git clone?","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2016-03-02T15:51:19Z","receivedAt":"2016-03-02T15:51:19Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On Wed, Mar 02, 2016 at 12:41:20AM -0800, Junio C Hamano wrote:\n> Josh Triplett <josh@joshtriplett.org> writes:\n> \n> > If you clone a repository, and the connection drops, the next attempt\n> > will have to start from scratch.  This can add significant time and\n> > expense if you're on a low-bandwidth or metered connection trying to\n> > clone something like Linux.\n> \n> For this particular issue, your friendly k.org administrator already\n> has a solution.  Torvalds/linux.git is made into a bundle weekly\n> with\n> \n>     $ git bundle create clone.bundle --all\n> \n> and the result placed on k.org CDN.  So low-bandwidth cloners can\n> grab it over resumable http, clone from the bundle, and then fill\n> the most recent part by fetching from k.org already.\n\nI finally got around to documenting this here:\nhttps://kernel.org/cloning-linux-from-a-bundle.html\n\n> The tooling to allow this kind of \"bundle\" (and possibly other forms\n> of \"CDN offload\" material) transparently used by \"git clone\" was the\n> proposal by Shawn Pearce mentioned elsewhere in this thread.\n\nTo reiterate, I believe that would be an awesome feature.\n\nRegards,\n-- \nKonstantin Ryabitsev\nLinux Foundation Collab Projects\nMontréal, Québec\n"},{"id":"280065","messageId":"20160302164006.GA13790@x","threadId":"41590","inReplyTo":"CACsJy8CBBk4bgz6Gn0QvCwWtOsqcQZBYgOBQTd=4Y+2YKs44Qg@mail.gmail.com","subject":"Re: Resumable git clone?","fromName":"Josh Triplett","fromEmail":"josh@joshtriplett.org","sentAt":"2016-03-02T16:40:06Z","receivedAt":"2016-03-02T16:40:06Z","isPatch":false,"sender":{"key":"josh@joshtriplett.org","avatar":"https://avatars.githubusercontent.com/u/162737?v=4"},"body":"On Wed, Mar 02, 2016 at 03:22:17PM +0700, Duy Nguyen wrote:\n> On Wed, Mar 2, 2016 at 3:13 PM, Josh Triplett <josh@joshtriplett.org> wrote:\n> > On Wed, Mar 02, 2016 at 02:30:24AM +0000, Al Viro wrote:\n> >> On Tue, Mar 01, 2016 at 05:40:28PM -0800, Stefan Beller wrote:\n> >>\n> >> > So throwing away half finished stuff while keeping the front load?\n> >>\n> >> Throw away the object that got truncated and ones for which delta chain\n> >> doesn't resolve entirely in the transferred part.\n> >>\n> >> > > indexing the objects it\n> >> > > contains, and then re-running clone and not having to fetch those\n> >> > > objects.\n> >> >\n> >> > The pack is not deterministic for a given repository. When creating\n> >> > the pack, you may encounter races between threads, such that the order\n> >> > in a pack differs.\n> >>\n> >> FWIW, I wasn't proposing to recreate the remaining bits of that _pack_;\n> >> just do the normal pull with one addition: start with sending the list\n> >> of sha1 of objects you are about to send and let the recepient reply\n> >> with \"I already have <set of sha1>, don't bother with those\".  And exclude\n> >> those from the transfer.  Encoding for the set being available is an\n> >> interesting variable here - might be plain list of sha1, might be its\n> >> complement (\"I want the following subset\"), might be \"145th to 1029th,\n> >> 1517th and 1890th to 1920th of the list you've sent\"; which form ends\n> >> up more efficient needs to be found experimentally...\n> >\n> > As a simple proposal, the server could send the list of hashes (in\n> > approximately the same order it would send the pack), the client could\n> > send back a bitmap where '0' means \"send it\" and '1' means \"got that one\n> > already\", and the client could compress that bitmap.  That gives you the\n> > RLE and similar without having to write it yourself.  That might not be\n> > optimal, but it would likely set a high bar with minimal effort.\n> \n> We have an implementation of EWAH bitmap compression, so compressing\n> is not a problem.\n> \n> But I still don't see why it's more efficient to have the server send\n> the hash list to the client. Assume you need to transfer N objects.\n> That direction makes you always send N hashes. But if the client sends\n> the list of already fetched objects, M, then M <= N. And we won't need\n> to send the bitmap. What did I miss?\n\nM can potentially be larger than N if you have many remotes and branches\nin your local repository that the server doesn't have.  However, that\ncertainly wouldn't be the common case, and in that case heuristics on\nthe client side could help there in determining a subset to send.\n\nI can't think of any good argument for the server's hash list; a\nclient-sent list does seem reasonable.\n\n- Josh Triplett\n"},{"id":"280066","messageId":"20160302164118.GA13732@x","threadId":"41590","inReplyTo":"xmqq4mcp5lij.fsf@gitster.mtv.corp.google.com","subject":"Re: Resumable git clone?","fromName":"Josh Triplett","fromEmail":"josh@joshtriplett.org","sentAt":"2016-03-02T16:41:18Z","receivedAt":"2016-03-02T16:41:18Z","isPatch":false,"sender":{"key":"josh@joshtriplett.org","avatar":"https://avatars.githubusercontent.com/u/162737?v=4"},"body":"On Wed, Mar 02, 2016 at 12:31:16AM -0800, Junio C Hamano wrote:\n> Josh Triplett <josh@joshtriplett.org> writes:\n> > I think several simpler optimizations seem\n> > preferable, such as binary object names, and abbreviating complete\n> > object sets (\"I have these commits/trees and everything they need\n> > recursively; I also have this stack of random objects.\").\n> \n> Given the way pack stream is organized (i.e. commits first and then\n> trees and blobs that belong to the same delta chain together), and\n> our assumed goal being to salvage objects from an interrupted\n> transfer of a packfile, you are unlikely to ever see \"I have these\n> commits/trees and everything they need\" that are salvaged from such\n> a failed transfer.  So I doubt such an optimization is worth doing.\n\nTrue for the resumable clone case.  For that optimization, I was\nthinking of the \"pull during the merge window\" case that Al Viro was\nalso interested in optimizing.\n\n> Besides it is very expensive to compute (the computation is done on\n> the client side, so the cycles burned and the time the user has to\n> wait is of much less concern, though); you'd essentially be doing\n> \"git fsck\" to find the \"dangling\" objects.\n\nTrading client-side computation for bandwidth can potentially be\nworthwhile if you have plenty of local compute but a slow and metered\nlink.\n\n> The list of what would be transferred needs to come in full from the\n> server end, as the list names objects that the receiving end may not\n> have seen, but the response by the client could be encoded much\n> tightly.  For the full list of N objects from the server, we can\n> think of your response to be a bitstream of N bits, each on-bit in\n> which signals an unwanted object in the list.  You can optimize this\n> transfer by RLE compressing the bitstream, for example.\n> \n> As git-over-HTTP is stateless, however, you cannot assume that the\n> server side remembers what it sent to the client (instead, the\n> client side needs to re-post what it heard from the server in the\n> previous exchange to allow the server side to use it after\n> validating).  So \"objects at these indices in your list\" kind of\n> optimization may not work very well in that environment.  I'd\n> imagine that an exchange of \"Here are the list of objects\", \"Give me\n> these objects\" done naively in full 40-hex object names would work\n> OK there, though.\n\nGood point.  Between statelessness and Duy's point about the client list\nusually being smaller than the server list, perhaps it would make sense\nto not have the server send a list at all, and just have the client send\nits own list.\n\n- Josh Triplett\n"},{"id":"280067","messageId":"20160302164906.GB13732@x","threadId":"41590","inReplyTo":"xmqqziuh46hb.fsf@gitster.mtv.corp.google.com","subject":"Re: Resumable git clone?","fromName":"Josh Triplett","fromEmail":"josh@joshtriplett.org","sentAt":"2016-03-02T16:49:06Z","receivedAt":"2016-03-02T16:49:06Z","isPatch":false,"sender":{"key":"josh@joshtriplett.org","avatar":"https://avatars.githubusercontent.com/u/162737?v=4"},"body":"On Wed, Mar 02, 2016 at 12:41:20AM -0800, Junio C Hamano wrote:\n> Josh Triplett <josh@joshtriplett.org> writes:\n> > If you clone a repository, and the connection drops, the next attempt\n> > will have to start from scratch.  This can add significant time and\n> > expense if you're on a low-bandwidth or metered connection trying to\n> > clone something like Linux.\n> \n> For this particular issue, your friendly k.org administrator already\n> has a solution.  Torvalds/linux.git is made into a bundle weekly\n> with\n> \n>     $ git bundle create clone.bundle --all\n> \n> and the result placed on k.org CDN.  So low-bandwidth cloners can\n> grab it over resumable http, clone from the bundle, and then fill\n> the most recent part by fetching from k.org already.\n> \n> The tooling to allow this kind of \"bundle\" (and possibly other forms\n> of \"CDN offload\" material) transparently used by \"git clone\" was the\n> proposal by Shawn Pearce mentioned elsewhere in this thread.\n\nThat does help in the case of cloning torvalds/linux.git from\nkernel.org, and I'd love to see it used transparently.\n\nHowever, even with that, I still also see value in a resumable git clone\n(or git pull) for many other repositories elsewhere, with a somewhat\nlower pull-to-push ratio than kernel.org.  Supporting resumption based\non objects, without the repository needing to generate and keep around a\nbundle, seems preferable for such repositories.\n"},{"id":"280077","messageId":"xmqqk2lk4vb5.fsf@gitster.mtv.corp.google.com","threadId":"41590","inReplyTo":"20160302164906.GB13732@x","subject":"Re: Resumable git clone?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-03-02T17:57:18Z","receivedAt":"2016-03-02T17:57:18Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Josh Triplett <josh@joshtriplett.org> writes:\n\n> That does help in the case of cloning torvalds/linux.git from\n> kernel.org, and I'd love to see it used transparently.\n>\n> However, even with that, I still also see value in a resumable git clone\n> (or git pull) for many other repositories elsewhere,...\n\nBy \"transparently\" the statement you are responding to meant many\nthings.\n\n\"git clone\" of course need to be updated on the client side, but\nthings like \"git repack\" that is run on the server end may start\nproducing extra files in the repository, and updated \"git daemon\"\nand/or \"git upload-pack\" would take these extra files as a signal\nthat the material produced during the last repack is usable for\nbootstrapping a new clone with \"wget -c\" equivalent.  So even if you\nare not yet automatically offloading to CDN, such a set of updates\non the server side would \"transparently\" enable the resumable clone\nfor all repositories elsewhere when deployed and enabled ;-)\n"},{"id":"281646","messageId":"C59B0CDA60BC402B900305A9D62D815B@PhilipOakley","threadId":"41590","inReplyTo":"xmqqziuh46hb.fsf@gitster.mtv.corp.google.com","subject":"Re: Resumable git clone?","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.org","sentAt":null,"receivedAt":"2016-03-24T07:44:04Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"From: \"Junio C Hamano\" <gitster@pobox.com>\nSent: Wednesday, March 02, 2016 8:41 AM\n> Josh Triplett <josh@joshtriplett.org> writes:\n>\n>> If you clone a repository, and the connection drops, the next attempt\n>> will have to start from scratch.  This can add significant time and\n>> expense if you're on a low-bandwidth or metered connection trying to\n>> clone something like Linux.\n>\n> For this particular issue, your friendly k.org administrator already\n> has a solution.  Torvalds/linux.git is made into a bundle weekly\n> with\n>\n>    $ git bundle create clone.bundle --all\n>\n\nIsn't this use of '--all' a bit of oversharing? I had proposed a doc patch\nto the bundle manpage way back (see $gmane/205897) to give the\nuser that example, but it wasn't accepted as it was thought wrong.\n\n\" I also think \"--all\" is a bad advice for another reason.  Doesn't it\nshove refs from refs/remotes/* hierarchy in the resulting bundle?\nIt is fine for archiving purposes, but it does not seem to be a good\nadvice to create a bundle to clone from.\"\n\nPerhaps the '--clone-bundle' (or maybe'--bundle-clone') option from \n$gmane/288222  [PATCH] index-pack: --clone-bundle option 2016-03-03 maybe a \nsuitable new <rev-list-arg> to get just the right content?\n\n> and the result placed on k.org CDN.  So low-bandwidth cloners can\n> grab it over resumable http, clone from the bundle, and then fill\n> the most recent part by fetching from k.org already.\n>\n> The tooling to allow this kind of \"bundle\" (and possibly other forms\n> of \"CDN offload\" material) transparently used by \"git clone\" was the\n> proposal by Shawn Pearce mentioned elsewhere in this thread.\n>\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n"},{"id":"281677","messageId":"xmqqy497an4a.fsf@gitster.mtv.corp.google.com","threadId":"41590","inReplyTo":"C59B0CDA60BC402B900305A9D62D815B@PhilipOakley","subject":"Re: Resumable git clone?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-03-24T15:53:25Z","receivedAt":"2016-03-24T15:53:25Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Philip Oakley\" <philipoakley@iee.org> writes:\n\n> From: \"Junio C Hamano\" <gitster@pobox.com>\n>>\n>>> If you clone a repository, and the connection drops, the next attempt\n>>> will have to start from scratch.  This can add significant time and\n>>> expense if you're on a low-bandwidth or metered connection trying to\n>>> clone something like Linux.\n>>\n>> For this particular issue, your friendly k.org administrator already\n>> has a solution.  Torvalds/linux.git is made into a bundle weekly\n>> with\n>>\n>>    $ git bundle create clone.bundle --all\n>>\n>> and the result placed on k.org CDN.  So low-bandwidth cloners can\n>> grab it over resumable http, clone from the bundle, and then fill\n>> the most recent part by fetching from k.org already.\n>\n> Isn't this use of '--all' a bit of oversharing?\n\nNot for the exact use case mentioned; k.org administrator knows what\nis in Linus's repository and is aware that there is no remote-tracking\nbranches or secret branches that may make the resulting bundle unsuitable\nfor priming a clone.\n\n> \" I also think \"--all\" is a bad advice for another reason.\n\nI do not think it is a good advice for everybody, but the thing is,\nwhat you are responding is not an advice.  It is just a statement of\na fact, what is already done, one of the existing practices that an\napproach to \"resumable clone\" may want to help.\n"},{"id":"281722","messageId":"211C0A1532414ED79933D75741AB26A0@PhilipOakley","threadId":"41590","inReplyTo":"xmqqy497an4a.fsf@gitster.mtv.corp.google.com","subject":"Re: Resumable git clone?","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.org","sentAt":null,"receivedAt":"2016-03-24T21:06:37Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"From: \"Junio C Hamano\" <gitster@pobox.com>\n> \"Philip Oakley\" <philipoakley@iee.org> writes:\n>\n>> From: \"Junio C Hamano\" <gitster@pobox.com>\n>>>\n>>>> If you clone a repository, and the connection drops, the next attempt\n>>>> will have to start from scratch.  This can add significant time and\n>>>> expense if you're on a low-bandwidth or metered connection trying to\n>>>> clone something like Linux.\n>>>\n>>> For this particular issue, your friendly k.org administrator already\n>>> has a solution.  Torvalds/linux.git is made into a bundle weekly\n>>> with\n>>>\n>>>    $ git bundle create clone.bundle --all\n>>>\n>>> and the result placed on k.org CDN.  So low-bandwidth cloners can\n>>> grab it over resumable http, clone from the bundle, and then fill\n>>> the most recent part by fetching from k.org already.\n>>\n>> Isn't this use of '--all' a bit of oversharing?\n>\n> Not for the exact use case mentioned; k.org administrator knows what\n> is in Linus's repository and is aware that there is no remote-tracking\n> branches or secret branches that may make the resulting bundle unsuitable\n> for priming a clone.\n\nOK\n>\n>> \" I also think \"--all\" is a bad advice for another reason.\n>\n> I do not think it is a good advice for everybody, but the thing is,\n> what you are responding is not an advice.  It is just a statement of\n> a fact, what is already done, one of the existing practices that an\n> approach to \"resumable clone\" may want to help.\n>\nI was picking up on the need, for others who maybe generating clone bundles, \nthat '--all' may not be the right thing for them, and that somewhere we \nshould record whatever is deemed the equivalent of the current clone \ncommand. This would get away from the web examples which show '--all' as a \nquick solution for bundling (I'm one of the offenders there).\n\nIf I understand the clone code, the  rev-list-args would be \n\"HEAD --branches\". But I could well be wrong.\n--\nPhilip \n"}]}