{"thread":{"id":"35417","subject":"How to resume broke clone ?","startedAt":"2013-11-28T03:13:54Z","lastAt":"2013-12-16T19:37:06Z","messageCount":34,"participants":["zhifeng hu","Trần Ngọc Quân","Duy Nguyen","Karsten Blees","Tay Ray Chuan","Jeff King","Shawn Pearce","Jakub Narebski","Michael Haggerty","Junio C Hamano","Jonathan Nieder"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"231221","messageId":"AAA12788-A242-41B8-B47D-1A0A52F33FC1@ancientrocklab.com","threadId":"35417","inReplyTo":null,"subject":"How to resume broke clone ?","fromName":"zhifeng hu","fromEmail":"zf@ancientrocklab.com","sentAt":"2013-11-28T03:13:54Z","receivedAt":"2013-11-28T03:13:54Z","isPatch":false,"sender":{"key":"zf@ancientrocklab.com","avatar":"https://gravatar.com/avatar/caeb8758bbe57865b250b2020f9b03edf729ea63c9f0ce48a2f85c4a797c2595?d=mp&s=160"},"body":"Hello all:\nToday i want to clone the Linux Kernel git repository.\ngit://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git\n\nI am in china. our bandwidth is very limitation. Less than 50Kb/s.\n\nThe clone progress is very slow, and broken times and time.\nI am very unhappy. \nBecause i could not easily to clone kernel.\n\nI had do some research about resume clone , but no good plan how to resolve this problem .\n\n\nWould it be possible add resume transfer clone repository after the transfer broken?\n\nsuch as bittorrent  download. or what ever.\n\nzhifeng hu \n"},{"id":"231223","messageId":"5296F343.6050506@gmail.com","threadId":"35417","inReplyTo":"AAA12788-A242-41B8-B47D-1A0A52F33FC1@ancientrocklab.com","subject":"Re: How to resume broke clone ?","fromName":"Trần Ngọc Quân","fromEmail":"vnwildman@gmail.com","sentAt":"2013-11-28T07:39:47Z","receivedAt":"2013-11-28T07:39:47Z","isPatch":false,"sender":{"key":"vnwildman@gmail.com","avatar":"https://avatars.githubusercontent.com/u/508758?v=4"},"body":"On 28/11/2013 10:13, zhifeng hu wrote:\n> Hello all:\n> Today i want to clone the Linux Kernel git repository.\n> git://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git\n>\n> I am in china. our bandwidth is very limitation. Less than 50Kb/s.\nThis repo is really too big.\nYou may consider using --depth option if you don't want full history, or\nclone from somewhere have better bandwidth\n$ git clone --depth=1\ngit://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git\nyou may chose other mirror (github.com) for example\nsee git-clone(1)\n\n-- \nTrần Ngọc Quân.\n"},{"id":"231224","messageId":"560807D9-CE82-4CF6-A1CC-54E7CCA624F9@ancientrocklab.com","threadId":"35417","inReplyTo":"5296F343.6050506@gmail.com","subject":"Re: How to resume broke clone ?","fromName":"zhifeng hu","fromEmail":"zf@ancientrocklab.com","sentAt":"2013-11-28T07:41:43Z","receivedAt":"2013-11-28T07:41:43Z","isPatch":false,"sender":{"key":"zf@ancientrocklab.com","avatar":"https://gravatar.com/avatar/caeb8758bbe57865b250b2020f9b03edf729ea63c9f0ce48a2f85c4a797c2595?d=mp&s=160"},"body":"Thanks for reply, But I am developer, I want to clone full repository, I need to view code since very early.\n\nzhifeng hu \n\n\n\nOn Nov 28, 2013, at 3:39 PM, Trần Ngọc Quân <vnwildman@gmail.com> wrote:\n\n> On 28/11/2013 10:13, zhifeng hu wrote:\n>> Hello all:\n>> Today i want to clone the Linux Kernel git repository.\n>> git://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git\n>> \n>> I am in china. our bandwidth is very limitation. Less than 50Kb/s.\n> This repo is really too big.\n> You may consider using --depth option if you don't want full history, or\n> clone from somewhere have better bandwidth\n> $ git clone --depth=1\n> git://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git\n> you may chose other mirror (github.com) for example\n> see git-clone(1)\n> \n> -- \n> Trần Ngọc Quân.\n> \n"},{"id":"231225","messageId":"CACsJy8DbJZmBCnfzNqfmEnRpqVcc42Q_-jz3r=sYVRPhsCkS5A@mail.gmail.com","threadId":"35417","inReplyTo":"560807D9-CE82-4CF6-A1CC-54E7CCA624F9@ancientrocklab.com","subject":"Re: How to resume broke clone ?","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-11-28T08:14:29Z","receivedAt":"2013-11-28T08:14:29Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, Nov 28, 2013 at 2:41 PM, zhifeng hu <zf@ancientrocklab.com> wrote:\n> Thanks for reply, But I am developer, I want to clone full repository, I need to view code since very early.\n\nif it works with --depth =1, you can incrementally run \"fetch\n--depth=N\" with N larger and larger.\n\nBut it may be easier to ask kernel.org admin, or any dev with a public\nweb server, to provide you a git bundle you can download via http.\nThen you can fetch on top.\n-- \nDuy\n"},{"id":"231227","messageId":"5297004F.4090003@gmail.com","threadId":"35417","inReplyTo":"CACsJy8DbJZmBCnfzNqfmEnRpqVcc42Q_-jz3r=sYVRPhsCkS5A@mail.gmail.com","subject":"Re: How to resume broke clone ?","fromName":"Karsten Blees","fromEmail":"karsten.blees@gmail.com","sentAt":"2013-11-28T08:35:27Z","receivedAt":"2013-11-28T08:35:27Z","isPatch":false,"sender":{"key":"karsten.blees@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1111200?v=4"},"body":"Am 28.11.2013 09:14, schrieb Duy Nguyen:\n> On Thu, Nov 28, 2013 at 2:41 PM, zhifeng hu <zf@ancientrocklab.com> wrote:\n>> Thanks for reply, But I am developer, I want to clone full repository, I need to view code since very early.\n> \n> if it works with --depth =1, you can incrementally run \"fetch\n> --depth=N\" with N larger and larger.\n> \n> But it may be easier to ask kernel.org admin, or any dev with a public\n> web server, to provide you a git bundle you can download via http.\n> Then you can fetch on top.\n> \n\nOr simply download the individual files (via ftp/http) and clone locally:\n\n> wget -r ftp://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/\n> git clone git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git\n> cd linux\n> git remote set-url origin git://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git\n"},{"id":"231230","messageId":"CACsJy8AdOAPT-RfD0NfZj_cQPBSUrVKn8yS7JRe=-4k8C8TvQg@mail.gmail.com","threadId":"35417","inReplyTo":"5297004F.4090003@gmail.com","subject":"Re: How to resume broke clone ?","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-11-28T08:50:02Z","receivedAt":"2013-11-28T08:50:02Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, Nov 28, 2013 at 3:35 PM, Karsten Blees <karsten.blees@gmail.com> wrote:\n> Or simply download the individual files (via ftp/http) and clone locally:\n>\n>> wget -r ftp://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/\n>> git clone git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git\n>> cd linux\n>> git remote set-url origin git://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git\n\nYeah I didn't realize it is published over dumb http too. You may need\nto be careful with this though because it's not atomic and you may get\nrefs that point nowhere because you're already done with \"pack\"\ndirectory when you come to fetcing \"refs\" and did not see new packs...\nIf dumb commit walker supports resume (I don't know) then it'll be\nsafer to do\n\ngit clone http://git.kernel.org/....\n\nIf it does not support resume, I don't think it's hard to do.\n-- \nDuy\n"},{"id":"231231","messageId":"211D44CB-64A2-4FCA-B4A7-40845B97E9A1@ancientrocklab.com","threadId":"35417","inReplyTo":"CACsJy8AdOAPT-RfD0NfZj_cQPBSUrVKn8yS7JRe=-4k8C8TvQg@mail.gmail.com","subject":"Re: How to resume broke clone ?","fromName":"zhifeng hu","fromEmail":"zf@ancientrocklab.com","sentAt":"2013-11-28T08:55:17Z","receivedAt":"2013-11-28T08:55:17Z","isPatch":false,"sender":{"key":"zf@ancientrocklab.com","avatar":"https://gravatar.com/avatar/caeb8758bbe57865b250b2020f9b03edf729ea63c9f0ce48a2f85c4a797c2595?d=mp&s=160"},"body":"The repository growing fast, things get harder . Now the size reach several GB, it may possible be TB, YB.\nWhen then, How do we handle this?\nIf the transfer broken, and it can not be resume transfer, waste time and waste bandwidth.\n\nGit should be better support resume transfer.\nIt now seems not doing better it’s job.\nShare code, manage code, transfer code, what would it be a VCS we imagine it ?\n \nzhifeng hu \n\n\n\nOn Nov 28, 2013, at 4:50 PM, Duy Nguyen <pclouds@gmail.com> wrote:\n\n> On Thu, Nov 28, 2013 at 3:35 PM, Karsten Blees <karsten.blees@gmail.com> wrote:\n>> Or simply download the individual files (via ftp/http) and clone locally:\n>> \n>>> wget -r ftp://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/\n>>> git clone git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git\n>>> cd linux\n>>> git remote set-url origin git://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git\n> \n> Yeah I didn't realize it is published over dumb http too. You may need\n> to be careful with this though because it's not atomic and you may get\n> refs that point nowhere because you're already done with \"pack\"\n> directory when you come to fetcing \"refs\" and did not see new packs...\n> If dumb commit walker supports resume (I don't know) then it'll be\n> safer to do\n> \n> git clone http://git.kernel.org/....\n> \n> If it does not support resume, I don't think it's hard to do.\n> -- \n> Duy\n"},{"id":"231232","messageId":"CACsJy8AOVWF2HssWNeYkVvYdmAXJOQ8HOehxJ0wpBFchA87ZWw@mail.gmail.com","threadId":"35417","inReplyTo":"211D44CB-64A2-4FCA-B4A7-40845B97E9A1@ancientrocklab.com","subject":"Re: How to resume broke clone ?","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-11-28T09:09:18Z","receivedAt":"2013-11-28T09:09:18Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, Nov 28, 2013 at 3:55 PM, zhifeng hu <zf@ancientrocklab.com> wrote:\n> The repository growing fast, things get harder . Now the size reach several GB, it may possible be TB, YB.\n> When then, How do we handle this?\n> If the transfer broken, and it can not be resume transfer, waste time and waste bandwidth.\n>\n> Git should be better support resume transfer.\n> It now seems not doing better it’s job.\n> Share code, manage code, transfer code, what would it be a VCS we imagine it ?\n\nYou're welcome to step up and do it. On top of my head  there are a few options:\n\n - better integration with git bundles, provide a way to seamlessly\ncreate/fetch/resume the bundles with \"git clone\" and \"git fetch\"\n - shallow/narrow clone. the idea is get a small part of the repo, one\ndepth, a few paths, then get more and more over many iterations so if\nwe fail one iteration we don't lose everything\n - stablize pack order so we can resume downloading a pack\n - remote alternates, the repo will ask for more and more objects as\nyou need them (so goodbye to distributed model)\n-- \nDuy\n"},{"id":"231234","messageId":"CALUzUxrEvuKuN+v-hJLQd5KoV-fzxVYvg5pj7XoLBVap7mgA=Q@mail.gmail.com","threadId":"35417","inReplyTo":"CACsJy8DbJZmBCnfzNqfmEnRpqVcc42Q_-jz3r=sYVRPhsCkS5A@mail.gmail.com","subject":"Re: How to resume broke clone ?","fromName":"Tay Ray Chuan","fromEmail":"rctay89@gmail.com","sentAt":"2013-11-28T09:20:36Z","receivedAt":"2013-11-28T09:20:36Z","isPatch":false,"sender":{"key":"rctay89@gmail.com","avatar":"https://avatars.githubusercontent.com/u/61553?v=4"},"body":"On Thu, Nov 28, 2013 at 4:14 PM, Duy Nguyen <pclouds@gmail.com> wrote:\n> On Thu, Nov 28, 2013 at 2:41 PM, zhifeng hu <zf@ancientrocklab.com> wrote:\n>> Thanks for reply, But I am developer, I want to clone full repository, I need to view code since very early.\n>\n> if it works with --depth =1, you can incrementally run \"fetch\n> --depth=N\" with N larger and larger.\n\nI second Duy Nguyen's and Trần Ngọc Quân's suggestion to 1) initially\ncreate a \"shallow\" clone then 2) incrementally deepen your clone.\n\nZhifeng, in the course of your research into resumable cloning, you\nmight have learnt that while it's a really valuable feature, it's also\na pretty hard problem at the same time. So it's not because git\ndoesn't want to have this feature.\n\n-- \nCheers,\nRay Chuan\n"},{"id":"231235","messageId":"20131128092935.GC11444@sigill.intra.peff.net","threadId":"35417","inReplyTo":"CACsJy8AOVWF2HssWNeYkVvYdmAXJOQ8HOehxJ0wpBFchA87ZWw@mail.gmail.com","subject":"Re: How to resume broke clone ?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-11-28T09:29:35Z","receivedAt":"2013-11-28T09:29:35Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Nov 28, 2013 at 04:09:18PM +0700, Duy Nguyen wrote:\n\n> > Git should be better support resume transfer.\n> > It now seems not doing better it’s job.\n> > Share code, manage code, transfer code, what would it be a VCS we imagine it ?\n> \n> You're welcome to step up and do it. On top of my head  there are a few options:\n> \n>  - better integration with git bundles, provide a way to seamlessly\n> create/fetch/resume the bundles with \"git clone\" and \"git fetch\"\n\nI posted patches for this last year. One of the things that I got hung\nup on was that I spooled the bundle to disk, and then cloned from it.\nWhich meant that you needed twice the disk space for a moment. I wanted\nto teach index-pack to \"--fix-thin\" a pack that was already on disk, so\nthat we could spool to disk, and then finalize it without making another\ncopy.\n\nOne of the downsides of this approach is that it requires the repo\nprovider (or somebody else) to provide the bundle. I think that is\nsomething that a big site like GitHub would do (and probably push the\nbundles out to a CDN, too, to make getting them faster). But it's not a\nuniversal solution.\n\n>  - stablize pack order so we can resume downloading a pack\n\nI think stabilizing in all cases (e.g., including ones where the content\nhas changed) is hard, but I wonder if it would be enough to handle the\neasy cases, where nothing has changed. If the server does not use\nmultiple threads for delta computation, it should generate the same pack\nfrom the same on-disk deterministically. We just need a way for the\nclient to indicate that it has the same partial pack.\n\nI'm thinking that the server would report some opaque hash representing\nthe current pack. The client would record that, along with the number of\npack bytes it received. If the transfer is interrupted, the client comes\nback with the hash/bytes pair. The server starts to generate the pack,\nchecks whether the hash matches, and if so, says \"here is the same pack,\nresuming at byte X\".\n\nWhat would need to go into such a hash? It would need to represent the\nexact bytes that will go into the pack, but without actually generating\nthose bytes. Perhaps a sha1 over the sequence of <object sha1, type,\nbase (if applicable), length> for each object would be enough. We should\nknow that after calling compute_write_order. If the client has a match,\nwe should be able to skip ahead to the correct byte.\n\n>  - remote alternates, the repo will ask for more and more objects as\n> you need them (so goodbye to distributed model)\n\nThis is also something I've been playing with, but just for very large\nobjects (so to support something like git-media, but below the object\ngraph layer). I don't think it would apply here, as the kernel has a lot\nof small objects, and getting them in the tight delta'd pack format\nincreases efficiency a lot.\n\n-Peff\n"},{"id":"231236","messageId":"F569EBDF-D8B5-47D5-8C2F-DA3A0F6C207E@ancientrocklab.com","threadId":"35417","inReplyTo":"CALUzUxrEvuKuN+v-hJLQd5KoV-fzxVYvg5pj7XoLBVap7mgA=Q@mail.gmail.com","subject":"Re: How to resume broke clone ?","fromName":"zhifeng hu","fromEmail":"zf@ancientrocklab.com","sentAt":"2013-11-28T09:29:49Z","receivedAt":"2013-11-28T09:29:49Z","isPatch":false,"sender":{"key":"zf@ancientrocklab.com","avatar":"https://gravatar.com/avatar/caeb8758bbe57865b250b2020f9b03edf729ea63c9f0ce48a2f85c4a797c2595?d=mp&s=160"},"body":"Once using git clone —depth or git fetch —depth,\nWhile you want to move backward.\nyou may face problem\n\n git fetch --depth=105\nerror: Could not read 483bbf41ca5beb7e38b3b01f21149c56a1154b7a\nerror: Could not read aacb82de3ff8ae7b0a9e4cfec16c1807b6c315ef\nerror: Could not read 5a1758710d06ce9ddef754a8ee79408277032d8b\nerror: Could not read a7d5629fe0580bd3e154206388371f5b8fc832db\nerror: Could not read 073291c476b4edb4d10bbada1e64b471ba153b6b\n\n\nzhifeng hu \n\n\n\nOn Nov 28, 2013, at 5:20 PM, Tay Ray Chuan <rctay89@gmail.com> wrote:\n\n> On Thu, Nov 28, 2013 at 4:14 PM, Duy Nguyen <pclouds@gmail.com> wrote:\n>> On Thu, Nov 28, 2013 at 2:41 PM, zhifeng hu <zf@ancientrocklab.com> wrote:\n>>> Thanks for reply, But I am developer, I want to clone full repository, I need to view code since very early.\n>> \n>> if it works with --depth =1, you can incrementally run \"fetch\n>> --depth=N\" with N larger and larger.\n> \n> I second Duy Nguyen's and Trần Ngọc Quân's suggestion to 1) initially\n> create a \"shallow\" clone then 2) incrementally deepen your clone.\n> \n> Zhifeng, in the course of your research into resumable cloning, you\n> might have learnt that while it's a really valuable feature, it's also\n> a pretty hard problem at the same time. So it's not because git\n> doesn't want to have this feature.\n> \n> -- \n> Cheers,\n> Ray Chuan\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"231238","messageId":"CACsJy8Cx_u_w+JoQnf9XA98tt4QVNzJ2zATBpWgk4_N3+dgCrg@mail.gmail.com","threadId":"35417","inReplyTo":"20131128092935.GC11444@sigill.intra.peff.net","subject":"Re: How to resume broke clone ?","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-11-28T10:17:30Z","receivedAt":"2013-11-28T10:17:30Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, Nov 28, 2013 at 4:29 PM, Jeff King <peff@peff.net> wrote:\n>>  - stablize pack order so we can resume downloading a pack\n>\n> I think stabilizing in all cases (e.g., including ones where the content\n> has changed) is hard, but I wonder if it would be enough to handle the\n> easy cases, where nothing has changed. If the server does not use\n> multiple threads for delta computation, it should generate the same pack\n> from the same on-disk deterministically. We just need a way for the\n> client to indicate that it has the same partial pack.\n>\n> I'm thinking that the server would report some opaque hash representing\n> the current pack. The client would record that, along with the number of\n> pack bytes it received. If the transfer is interrupted, the client comes\n> back with the hash/bytes pair. The server starts to generate the pack,\n> checks whether the hash matches, and if so, says \"here is the same pack,\n> resuming at byte X\".\n>\n> What would need to go into such a hash? It would need to represent the\n> exact bytes that will go into the pack, but without actually generating\n> those bytes. Perhaps a sha1 over the sequence of <object sha1, type,\n> base (if applicable), length> for each object would be enough. We should\n> know that after calling compute_write_order. If the client has a match,\n> we should be able to skip ahead to the correct byte.\n\nExactly. The hash would include the list of sha-1 and object source,\nthe git version (so changes in code or default values are covered),\nthe list of config keys/values that may impact pack generation\nalgorithm (like window size..), .git/shallow, refs/replace,\n.git/graft, all or most of command line options. If we audit the code\ncarefully I think we can cover all input that influences pack\ngeneration. From then on it's just a matter of protocol extension. It\nalso opens an opportunity for optional server side caching, just save\nthe pack and associate it with the hash. Next time the client asks to\nresume, the server has everything ready.\n-- \nDuy\n"},{"id":"231251","messageId":"CAJo=hJuBTjGfF2PvaCn_v4hy4qDfFyB=FXbY0=Oz3hcE0L=L4Q@mail.gmail.com","threadId":"35417","inReplyTo":"20131128092935.GC11444@sigill.intra.peff.net","subject":"Re: How to resume broke clone ?","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2013-11-28T19:15:27Z","receivedAt":"2013-11-28T19:15:27Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Thu, Nov 28, 2013 at 1:29 AM, Jeff King <peff@peff.net> wrote:\n> On Thu, Nov 28, 2013 at 04:09:18PM +0700, Duy Nguyen wrote:\n>\n>> > Git should be better support resume transfer.\n>> > It now seems not doing better it’s job.\n>> > Share code, manage code, transfer code, what would it be a VCS we imagine it ?\n>>\n>> You're welcome to step up and do it. On top of my head  there are a few options:\n>>\n>>  - better integration with git bundles, provide a way to seamlessly\n>> create/fetch/resume the bundles with \"git clone\" and \"git fetch\"\n\nWe have been thinking about formalizing the /clone.bundle hack used by\nrepo on Android. If the server has the bundle, add a capability in the\nrefs advertisement saying its available, and the clone client can\nfirst fetch $URL/clone.bundle.\n\nFor most Git repositories the bundle can be constructed by saving the\nbundle reference header into a file, e.g.\n$GIT_DIR/objects/pack/pack-$NAME.bh at the same time the pack is\ncreated. The bundle can be served by combining the .bh and .pack\nstreams onto the network. It is very little additional disk overhead\nfor the origin server, but allows resumable clone, provided the server\nhas not done a GC.\n\n> I posted patches for this last year. One of the things that I got hung\n> up on was that I spooled the bundle to disk, and then cloned from it.\n> Which meant that you needed twice the disk space for a moment.\n\nI don't think this is a huge concern. In many cases the checked out\ncopy of the repository approaches a sizable fraction of the .pack\nitself. If you don't have 2x .pack disk available at clone time you\nmay be in trouble anyway as you try to work with the repository post\nclone.\n\n> I wanted\n> to teach index-pack to \"--fix-thin\" a pack that was already on disk, so\n> that we could spool to disk, and then finalize it without making another\n> copy.\n\nDon't you need to separate the bundle header from the pack data before\nyou do this? If the bundle is only used at clone time there is no\n--fix-thin step.\n\n> One of the downsides of this approach is that it requires the repo\n> provider (or somebody else) to provide the bundle. I think that is\n> something that a big site like GitHub would do (and probably push the\n> bundles out to a CDN, too, to make getting them faster). But it's not a\n> universal solution.\n\nSee above, I think you can reasonably do the /clone.bundle\nautomatically on any HTTP server. Big sites might choose to have\n/clone.bundle do a redirect into a caching CDN that fills itself by\ngoing to the application servers to obtain the current data. This is\nwhat we do for Android.\n\n>>  - stablize pack order so we can resume downloading a pack\n>\n> I think stabilizing in all cases (e.g., including ones where the content\n> has changed) is hard, but I wonder if it would be enough to handle the\n> easy cases, where nothing has changed. If the server does not use\n> multiple threads for delta computation, it should generate the same pack\n> from the same on-disk deterministically. We just need a way for the\n> client to indicate that it has the same partial pack.\n>\n> I'm thinking that the server would report some opaque hash representing\n> the current pack. The client would record that, along with the number of\n> pack bytes it received. If the transfer is interrupted, the client comes\n> back with the hash/bytes pair. The server starts to generate the pack,\n> checks whether the hash matches, and if so, says \"here is the same pack,\n> resuming at byte X\".\n\nAn important part of this is the want set must be identical to the\nprior request. It is entirely possible the branch tips have advanced\nsince the prior packing attempt started.\n\n> What would need to go into such a hash? It would need to represent the\n> exact bytes that will go into the pack, but without actually generating\n> those bytes. Perhaps a sha1 over the sequence of <object sha1, type,\n> base (if applicable), length> for each object would be enough. We should\n> know that after calling compute_write_order. If the client has a match,\n> we should be able to skip ahead to the correct byte.\n\nI don't think Length is sufficient.\n\nThe repository could have recompressed an object with the same length\nbut different libz encoding. I wonder if loose object recompression is\nreliable enough about libz encoding to resume in the middle of an\nobject? Is it just based on libz version?\n\nYou may need to do include information about the source of the object,\ne.g. the trailing 20 byte hash in the source pack file.\n"},{"id":"231252","messageId":"CAJo=hJvk0WJJM1jn_XiqpJM791pu=Wmh7CObpnAS60TFQFOfeQ@mail.gmail.com","threadId":"35417","inReplyTo":"F569EBDF-D8B5-47D5-8C2F-DA3A0F6C207E@ancientrocklab.com","subject":"Re: How to resume broke clone ?","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2013-11-28T19:35:56Z","receivedAt":"2013-11-28T19:35:56Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Thu, Nov 28, 2013 at 1:29 AM, zhifeng hu <zf@ancientrocklab.com> wrote:\n> Once using git clone —depth or git fetch —depth,\n> While you want to move backward.\n> you may face problem\n>\n>  git fetch --depth=105\n> error: Could not read 483bbf41ca5beb7e38b3b01f21149c56a1154b7a\n> error: Could not read aacb82de3ff8ae7b0a9e4cfec16c1807b6c315ef\n> error: Could not read 5a1758710d06ce9ddef754a8ee79408277032d8b\n> error: Could not read a7d5629fe0580bd3e154206388371f5b8fc832db\n> error: Could not read 073291c476b4edb4d10bbada1e64b471ba153b6b\n\nWe now have a resumable bundle available through our kernel.org\nmirror. The bundle is 658M.\n\n  mkdir linux\n  cd linux\n  git init\n\n  wget https://kernel.googlesource.com/pub/scm/linux/kernel/git/torvalds/linux/clone.bundle\n\n  sha1sum clone.bundle\n  96831de0b81713333e5ebba94edb31e37e70e1df  clone.bundle\n\n  git fetch -u ./clone.bundle refs/*:refs/*\n  git reset --hard\n\nYou can also use our mirror as an upstream, as we have servers in Asia\nthat lag no more than 5 or 6 minutes behind kernel.org:\n\n  git remote add origin\nhttps://kernel.googlesource.com/pub/scm/linux/kernel/git/torvalds/linux/\n"},{"id":"231264","messageId":"loom.20131128T225302-989@post.gmane.org","threadId":"35417","inReplyTo":"F569EBDF-D8B5-47D5-8C2F-DA3A0F6C207E@ancientrocklab.com","subject":"Re: How to resume broke clone ?","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2013-11-28T21:54:16Z","receivedAt":"2013-11-28T21:54:16Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"zhifeng hu <zf <at> ancientrocklab.com> writes:\n\n> \n> Once using git clone —depth or git fetch —depth,\n> While you want to move backward.\n> you may face problem\n> \n>  git fetch --depth=105\n> error: Could not read 483bbf41ca5beb7e38b3b01f21149c56a1154b7a\n> error: Could not read aacb82de3ff8ae7b0a9e4cfec16c1807b6c315ef\n> error: Could not read 5a1758710d06ce9ddef754a8ee79408277032d8b\n> error: Could not read a7d5629fe0580bd3e154206388371f5b8fc832db\n> error: Could not read 073291c476b4edb4d10bbada1e64b471ba153b6b\n\nBTW. there was (is?) a bundler service at http://bundler.caurea.org/\nbut I don't know if it can create Linux-size bundle.\n\n-- \nJakub Narębski\n"},{"id":"231535","messageId":"20131204200850.GB16603@sigill.intra.peff.net","threadId":"35417","inReplyTo":"CAJo=hJuBTjGfF2PvaCn_v4hy4qDfFyB=FXbY0=Oz3hcE0L=L4Q@mail.gmail.com","subject":"Re: How to resume broke clone ?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-12-04T20:08:50Z","receivedAt":"2013-12-04T20:08:50Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Nov 28, 2013 at 11:15:27AM -0800, Shawn Pearce wrote:\n\n> >>  - better integration with git bundles, provide a way to seamlessly\n> >> create/fetch/resume the bundles with \"git clone\" and \"git fetch\"\n> \n> We have been thinking about formalizing the /clone.bundle hack used by\n> repo on Android. If the server has the bundle, add a capability in the\n> refs advertisement saying its available, and the clone client can\n> first fetch $URL/clone.bundle.\n\nYes, that was going to be my next step after getting the bundle fetch\nsupport in. If we are going to do this, though, I'd really love for it\nto not be \"hey, fetch .../clone.bundle from me\", but a full-fledged\n\"here are full URLs of my mirrors\".\n\nThen you can redirect a non-http cloner to http to grab the bundle. Or\nredirect them to a CDN. Or even somebody else's server entirely (e.g.,\n\"go fetch from Linus first, my piddly server cannot feed you the whole\nkernel\"). Some of the redirects you can do by issuing an http redirect\nto \"/clone.bundle\", but the cross-protocol ones are tricky.\n\nIf we advertise it as a blob in a specialized ref (e.g., \"refs/mirrors\")\nit does not add much overhead over a simple capability. There are a few\nextra round trips to actually fetch the blob (client sends a want and no\nhaves, then server sends the pack), but I think that's negligible when\nwe are talking about redirecting a full clone. In either case, we have\nto hang up the original connection, fetch the mirror, and then come\nback.\n\n> For most Git repositories the bundle can be constructed by saving the\n> bundle reference header into a file, e.g.\n> $GIT_DIR/objects/pack/pack-$NAME.bh at the same time the pack is\n> created. The bundle can be served by combining the .bh and .pack\n> streams onto the network. It is very little additional disk overhead\n> for the origin server,\n\nThat's clever. It does not work out of the box if you are using\nalternates, but I think it could be adapted in certain situations. E.g.,\nif you layer the pack so that one \"base\" repo always has its full pack\nat the start, which is something we're already doing at GitHub.\n\n> but allows resumable clone, provided the server has not done a GC.\n\nAs an aside, the current transfer-resuming code in http.c is\nquestionable.  It does not use etags or any sort of invalidation\nmechanism, but just assumes hitting the same URL will give the same\nbytes. That _usually_ works for dumb fetching of objects and packfiles,\nthough it is possible for a pack to change representation without\nchanging name.\n\nMy bundle patches inherited the same flaw, but it is much worse there,\nbecause your URL may very well just be \"clone.bundle\" that gets updated\nperiodically.\n\n> > I posted patches for this last year. One of the things that I got hung\n> > up on was that I spooled the bundle to disk, and then cloned from it.\n> > Which meant that you needed twice the disk space for a moment.\n> \n> I don't think this is a huge concern. In many cases the checked out\n> copy of the repository approaches a sizable fraction of the .pack\n> itself. If you don't have 2x .pack disk available at clone time you\n> may be in trouble anyway as you try to work with the repository post\n> clone.\n\nYeah, in retrospect I was being stupid to let that hold it up. I'll\nrevisit the patches (I've rebased them forward over the past year, so it\nshouldn't be too bad).\n\n> > I wanted\n> > to teach index-pack to \"--fix-thin\" a pack that was already on disk, so\n> > that we could spool to disk, and then finalize it without making another\n> > copy.\n> \n> Don't you need to separate the bundle header from the pack data before\n> you do this?\n\nYes, though it isn't hard. We have to fetch part of the bundle header\ninto memory during discover_refs(), since that is when we realize we are\ngetting a bundle and not just the refs. From there you can spool the\nbundle header to disk, and then the packfile separately.\n\nMy original implementation did that, though I don't remember if that one\ngot posted to the list (after realizing that I couldn't just\n\"--fix-thin\" directly, I simplified it to just spool the whole thing to\na single file).\n\n> If the bundle is only used at clone time there is no\n> --fix-thin step.\n\nYes, for the particular use case of a clone-mirror, you wouldn't need to\n--fix-thin. But I think \"git fetch https://example.com/foo.bundle\"\nshould work in the general case (and it does with my patches).\n\n> See above, I think you can reasonably do the /clone.bundle\n> automatically on any HTTP server.\n\nYeah, the \".bh\" trick you mentioned is low enough impact to the server\nthat we could just unconditionally make it part of the repack.\n\n> > What would need to go into such a hash? It would need to represent the\n> > exact bytes that will go into the pack, but without actually generating\n> > those bytes. Perhaps a sha1 over the sequence of <object sha1, type,\n> > base (if applicable), length> for each object would be enough. We should\n> > know that after calling compute_write_order. If the client has a match,\n> > we should be able to skip ahead to the correct byte.\n> \n> I don't think Length is sufficient.\n> \n> The repository could have recompressed an object with the same length\n> but different libz encoding. I wonder if loose object recompression is\n> reliable enough about libz encoding to resume in the middle of an\n> object? Is it just based on libz version?\n> \n> You may need to do include information about the source of the object,\n> e.g. the trailing 20 byte hash in the source pack file.\n\nYeah, I think you're right that it's too flaky without recording the\nsource. At any rate, I think I prefer the bundle approach you mentioned\nabove. It solves the same problem, and is a lot more flexible (e.g., for\noffloading to other servers).\n\n-Peff\n"},{"id":"231573","messageId":"CAJo=hJuRz9Qc8ztQATkEs8huDfiANMA6gZEOapoofVdoY82k4g@mail.gmail.com","threadId":"35417","inReplyTo":"20131204200850.GB16603@sigill.intra.peff.net","subject":"Re: How to resume broke clone ?","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2013-12-05T06:50:27Z","receivedAt":"2013-12-05T06:50:27Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Wed, Dec 4, 2013 at 12:08 PM, Jeff King <peff@peff.net> wrote:\n> On Thu, Nov 28, 2013 at 11:15:27AM -0800, Shawn Pearce wrote:\n>\n>> >>  - better integration with git bundles, provide a way to seamlessly\n>> >> create/fetch/resume the bundles with \"git clone\" and \"git fetch\"\n>>\n>> We have been thinking about formalizing the /clone.bundle hack used by\n>> repo on Android. If the server has the bundle, add a capability in the\n>> refs advertisement saying its available, and the clone client can\n>> first fetch $URL/clone.bundle.\n>\n> Yes, that was going to be my next step after getting the bundle fetch\n> support in.\n\nYay!\n\n> If we are going to do this, though, I'd really love for it\n> to not be \"hey, fetch .../clone.bundle from me\", but a full-fledged\n> \"here are full URLs of my mirrors\".\n\nAck. I agree completely.\n\n> Then you can redirect a non-http cloner to http to grab the bundle. Or\n> redirect them to a CDN. Or even somebody else's server entirely (e.g.,\n> \"go fetch from Linus first, my piddly server cannot feed you the whole\n> kernel\"). Some of the redirects you can do by issuing an http redirect\n> to \"/clone.bundle\", but the cross-protocol ones are tricky.\n\nAck. My thoughts exactly. Especially the part of \"my piddly server\nshouldn't have to serve you a clone of Linus' tree when there are many\npublic hosts mirroring his code available to anyone\". It is simply not\nfair to clone Linus' tree off some guy's home ADSL connection, his\nuplink probably sucks. But it is reasonable to fetch his incremental\ndelta after cloning from some other well known and well connected\nsource.\n\n> If we advertise it as a blob in a specialized ref (e.g., \"refs/mirrors\")\n> it does not add much overhead over a simple capability. There are a few\n> extra round trips to actually fetch the blob (client sends a want and no\n> haves, then server sends the pack), but I think that's negligible when\n> we are talking about redirecting a full clone. In either case, we have\n> to hang up the original connection, fetch the mirror, and then come\n> back.\n\nI wasn't thinking about using a \"well known blob\" for this.\n\nJonathan, Dave, Colby and I were kicking this idea around on Monday\nduring lunch. If the initial ref advertisement included a \"mirrors\"\ncapability the client could respond with \"want mirrors\" instead of the\nusual want/have negotiation. The server could then return the mirror\nURLs as pkt-lines, one per pkt. Its one extra RTT, but this is trivial\ncompared to the cost to really clone the repository.\n\nThese pkt-lines need to be a bit more than just URL. Or we need a new\nURL like \"bundle:http://....\" to denote a resumable bundle over HTTP\nvs. a normal HTTP URL that might not be a bundle file, and is just a\nbetter connected server.\n\n\nThe mirror URLs could be stored in $GIT_DIR/config as a simple\nmulti-value variable. Unfortunately that isn't easily remotely\neditable. But I am not sure I care?\n\nGitHub doesn't let you edit $GIT_DIR/config, but it doesn't need to.\nFor most repositories hosted at GitHub, GitHub is probably the best\nconnected server for that repository. For repositories that are\nincredibly high traffic GitHub might out of its own interest want to\nconfigure mirror URLs on some sort of CDN to distribute the network\ntraffic closer to the edges. Repository owners just shouldn't have to\nworry about these sorts of details. It should be managed by the\nhosting service.\n\nIn my case for android.googlesource.com we want bundles on the CDN\nnear the network edges, and our repository owners don't care to know\nthe details of that. They just want our server software to make it all\nhappen, and our servers already manage $GIT_DIR/config for them. It\nalso mostly manages /clone.bundle on the CDN. And /clone.bundle is an\nugly, limited hack.\n\nFor the average home user sharing their working repository over git://\nfrom their home ADSL or cable connection, editing .git/config is\neasier than a blob in refs/mirrors. They already know how to edit\n.git/config to manage remotes. Heck, remote.origin.url might already\nbe a good mirror address to advertise, especially if the client isn't\non the same /24 as the server and the remote.origin.url is something\nlike \"git.kernel.org\". :-)\n\n>> For most Git repositories the bundle can be constructed by saving the\n>> bundle reference header into a file, e.g.\n>> $GIT_DIR/objects/pack/pack-$NAME.bh at the same time the pack is\n>> created. The bundle can be served by combining the .bh and .pack\n>> streams onto the network. It is very little additional disk overhead\n>> for the origin server,\n>\n> That's clever. It does not work out of the box if you are using\n> alternates, but I think it could be adapted in certain situations. E.g.,\n> if you layer the pack so that one \"base\" repo always has its full pack\n> at the start, which is something we're already doing at GitHub.\n\nYes, well, I was assuming the pack was a fully connected repack.\nAlternates always creates a partial pack. But if you have an\nalternate, that alternate maybe should be given as a mirror URL? And\nallow the client to recurse the alternate mirror URL list too?\n\nBy listing the alternate as a mirror a client could maybe discover the\nresumable clone bundle in the alternate, grab that first to bootstrap,\nreducing the amount it has to obtain in a non-resumable way. Or... the\ndescendant repository could offer its own bundle with the \"must have\"\nassertions from the alternate at the time it repacked. So the .bh file\nwould have a number of ^ lines and the bundle was built with a \"--not\n...\" list.\n\n>> but allows resumable clone, provided the server has not done a GC.\n>\n> As an aside, the current transfer-resuming code in http.c is\n> questionable.  It does not use etags or any sort of invalidation\n> mechanism, but just assumes hitting the same URL will give the same\n> bytes.\n\nYea, our lunch conversation eventually reached this part too. repo's\n/clone.bundle hack is equally stupid and assumes a resume will get the\ncorrect data, with no validation. If you resume with the wrong data\nwhile inside of the pack stream the pack will be invalid; the SHA-1\ntrailer won't match. But you won't know until you have downloaded the\nentire useless file. Resuming a 700M download after the first 10M only\nto find out the first 10M is mismatched sucks.\n\nWhat really got us worried was the bundle header has no checksums, and\na resume in the bundle header from the wrong version could be\ninteresting.\n\n> That _usually_ works for dumb fetching of objects and packfiles,\n> though it is possible for a pack to change representation without\n> changing name.\n\nYes. And this is why the packfile name algorithm is horribly flawed. I\nkeep saying we should change it to name the pack using the last 20\nbytes of the file but ... nobody has written the patch for that?  :-)\n\n> My bundle patches inherited the same flaw, but it is much worse there,\n> because your URL may very well just be \"clone.bundle\" that gets updated\n> periodically.\n\nYup, you followed the same thing we did in repo, which is horribly wrong.\n\nWe should try to use ETag if available to safely resume, and we should\ntry to encourage people to use stronger names when pointing to URLs\nthat are resumable, like a bundle on a CDN. If the URL is offered by\nthe server in pkt-lines after the advertisement its easy for the\nserver to return the current CDN URL, and easy for the server to\nimplement enforcement of the URLs being unique. Especially if you\nmanage the CDN automatically; e.g. Android uses tools to build the CDN\nfiles and push them out. Its easy for us to ensure these have unique\nURLs on every push. A bundling server bundling once a day or once a\nweek could simply date stamp each run.\n\n>> > I posted patches for this last year. One of the things that I got hung\n>> > up on was that I spooled the bundle to disk, and then cloned from it.\n>> > Which meant that you needed twice the disk space for a moment.\n>>\n>> I don't think this is a huge concern. In many cases the checked out\n>> copy of the repository approaches a sizable fraction of the .pack\n>> itself. If you don't have 2x .pack disk available at clone time you\n>> may be in trouble anyway as you try to work with the repository post\n>> clone.\n>\n> Yeah, in retrospect I was being stupid to let that hold it up. I'll\n> revisit the patches (I've rebased them forward over the past year, so it\n> shouldn't be too bad).\n\nI keep prodding Jonathan to work on this too, because I'd really like\nto get this out of repo and just have it be something git knows how to\ndo. And bigger mirrors like git.kernel.org could do a quick\ngrep/sort/uniq -c through their access logs and periodically bundle up\na few repositories that are cloned often. E.g. we all know\ngit.kernel.org should just bundle Linus' repository.\n"},{"id":"231613","messageId":"52A07DC5.5090508@alum.mit.edu","threadId":"35417","inReplyTo":"CAJo=hJuRz9Qc8ztQATkEs8huDfiANMA6gZEOapoofVdoY82k4g@mail.gmail.com","subject":"Re: How to resume broke clone ?","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2013-12-05T13:21:09Z","receivedAt":"2013-12-05T13:21:09Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"This discussion has mostly been about letting small Git servers delegate\nthe work of an initial clone to a beefier server.  I haven't seen any\nexplicit mention of the inverse:\n\nSuppose a company has a central Git server that is meant to be the\n\"single source of truth\", but has worldwide offices and wants to locate\nbootstrap mirrors in each office.  The end users would not even want to\nknow that there are multiple servers.  Hosters like GitHub might also\nencourage their big customers to set up bootstrap mirror(s) in-house to\nmake cloning faster for their users while reducing internet traffic and\nthe burden on their own infrastructure.  The goal would be to make the\nsystem transparent to users and easily reconfigurable as circumstances\nchange.\n\nOne alternative would be to ask users to clone from their local mirror.\n The local mirror would give them whatever it has, then do the\nequivalent of a permanent redirect to tell the client \"from now on, use\nthe central server\" to get the rest of the initial clone and for future\nfetches.  But this would require users to know which mirror is \"local\".\n\nA better alternative would be to ask users to clone from the central\nserver.  In this case, the central server would want to tell the clients\nto grab what they can from their local bootstrap mirror and then come\nback to the central server for any remainders.  The trick is that which\nbootstrap mirror is \"local\" would vary from client to client.\n\nI suppose that this could be implemented using what you have discussed\nby having the central server direct the client to a URL that resolves\ndifferently for different clients, CDN-like.  Alternatively, the central\nGit server could itself look where a request is coming from and use some\nintelligence to redirect the client to the closest bootstrap mirror from\nits own list.  Or the server could pass the client a list of known\nmirrors, and the client could try to determine which one is closest (and\nreachable!).\n\nI'm not sure that this idea is interesting, but I just wanted to throw\nit out there as a related use case that seems a bit different than what\nyou have been discussing.\n\nMichael\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"231614","messageId":"CAJo=hJuRTHEO6Mpaaa7b=_ozDq3UfVNEx6iabKP4xRVoRZLcUw@mail.gmail.com","threadId":"35417","inReplyTo":"52A07DC5.5090508@alum.mit.edu","subject":"Re: How to resume broke clone ?","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2013-12-05T15:11:15Z","receivedAt":"2013-12-05T15:11:15Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Thu, Dec 5, 2013 at 5:21 AM, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> This discussion has mostly been about letting small Git servers delegate\n> the work of an initial clone to a beefier server.  I haven't seen any\n> explicit mention of the inverse:\n>\n> Suppose a company has a central Git server that is meant to be the\n> \"single source of truth\", but has worldwide offices and wants to locate\n> bootstrap mirrors in each office.  The end users would not even want to\n> know that there are multiple servers.  Hosters like GitHub might also\n> encourage their big customers to set up bootstrap mirror(s) in-house to\n> make cloning faster for their users while reducing internet traffic and\n> the burden on their own infrastructure.  The goal would be to make the\n> system transparent to users and easily reconfigurable as circumstances\n> change.\n\nI think there is a different way to do that.\n\nBuild a caching Git proxy server. And teach Git clients to use it.\n\n\nOne idea we had at $DAY_JOB a couple of years ago was to build a\ndaemon that sat in the background and continuously fetched content\nfrom repository upstreams. We made it efficient by modifying the Git\nprotocol to use a hanging network socket, and the upstream server\nwould broadcast push pack files down these hanging streams as pushes\nwere received.\n\nThe original intent was for an Android developer to be able to have\nhis working tree forest of 500 repositories subscribe to our internal\nserver's broadcast stream. We figured if the server knows exactly\nwhich refs every client has, because they all have the same ones, and\ntheir streams are all still open and active, then the server can make\nexactly one incremental thin pack and send the same copy to every\nclient. Its \"just\" a socket write problem. Instead of packing the same\nstuff 100x for 100x clients its packed once and sent 100x.\n\nThen we realized remote offices could also install this software on a\nlocal server, and use this as a fan-out distributor within the LAN. We\nwere originally thinking about some remote offices on small Internet\nconnections, where delivery of 10 MiB x 20 was a lot but delivery of\n10 MiB once and local fan-out on the Ethernet was easy.\n\nThe JGit patches for this work are still pending[1].\n\n\nIf clients had a local Git-aware cache server in their office and\n~/.gitconfig had the address of it, your problem becomes simple.\n\nClients clone from the public URL e.g. GitHub, but the local cache\nserver first gives the client a URL to clone from itself. After that\nis complete then the client can fetch from the upstream. The cache\nserver can be self-maintaining, watching its requests to see what is\naccessed often-ish, and keep those repositories current-ish locally by\nrunning git fetch itself in the background.\n\nIts easy to do this with bundles on \"CDN\" like HTTP. Just use the\noffice's caching HTTP proxy server. Assuming its cache is big enough\nfor those large Git bundle payloads, and the viral cat videos. But you\nare at the mercy of the upstream bundler rebuilding the bundles. And\nrefetching them in whole. Neither of which is great.\n\nA simple self-contained server that doesn't accept pushes, but knows\nhow to clone repositories, fetch them periodically, and run `git gc`,\nworks well. And the mirror URL extension we have been discussing in\nthis thread would work fine here. The cache server can return URLs\nthat point to itself. Or flat out proxy the Git transaction with the\norigin server.\n\n\n[1] https://git.eclipse.org/r/#/q/owner:wetherbeei%2540google.com+status:open,n,z\n"},{"id":"231615","messageId":"20131205160418.GA27869@sigill.intra.peff.net","threadId":"35417","inReplyTo":"CAJo=hJuRz9Qc8ztQATkEs8huDfiANMA6gZEOapoofVdoY82k4g@mail.gmail.com","subject":"Re: How to resume broke clone ?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-12-05T16:04:18Z","receivedAt":"2013-12-05T16:04:18Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Dec 04, 2013 at 10:50:27PM -0800, Shawn Pearce wrote:\n\n> I wasn't thinking about using a \"well known blob\" for this.\n> \n> Jonathan, Dave, Colby and I were kicking this idea around on Monday\n> during lunch. If the initial ref advertisement included a \"mirrors\"\n> capability the client could respond with \"want mirrors\" instead of the\n> usual want/have negotiation. The server could then return the mirror\n> URLs as pkt-lines, one per pkt. Its one extra RTT, but this is trivial\n> compared to the cost to really clone the repository.\n\nI don't think this is any more or less efficient than the blob scheme.\nIn both cases, the client sends a single \"want\" line and no \"have\"\nlines, and then the server responds with the output (either pkt-lines,\nor a single-blob pack).\n\nWhat I like about the blob approach is:\n\n  1. It requires zero extra code on the server. This makes\n     implementation simple, but also means you can deploy it\n     on existing servers (or even on non-pkt-line servers like\n     dumb http).\n\n  2. It's very debuggable from the client side. You can fetch the blob,\n     look at it, and decide which mirror you want outside of git if you\n     want to (true, you can teach the git client to dump the pkt-line\n     URLs, too, but that's extra code). You could even do this with an\n     existing git client that has not yet learned about the mirror\n     redirect.\n\n  3. It removes any size or structure limits that the protocol imposes\n     (I was planning to use git-config format for the blob itself). The\n     URLs themselves aren't big, but we may want to annotate them with\n     metadata.\n\n     You mentioned \"this is a bundle\" versus \"this is a regular http\n     server\" below. You might also want to provide network location\n     information (e.g., \"this is a good mirror if you are in Asia\"),\n     though for the most part I'd expect that to happen magically via\n     CDN.\n\n     When we discussed this before, the concept came up of offering not\n     just a clone bundle, but \"slices\" of history (as thin-pack\n     bundles), so that a fetch could grab a sequence of resumable\n     slices, starting with what they have, and then topping off with a\n     true fetch. You would want to provide the start and end points of\n     each slice.\n\n  4. You can manage it remotely via the git protocol (more discussion\n     below).\n\n  5. A clone done with \"--mirror\" will actually propagate the mirror\n     file automatically.\n\nWhat are the advantages of the pkt-line approach? The biggest one I can\nthink of is that it does not pollute the refs namespace. While (5) is\nconvenient in some cases, it would make it more of a pain if you are\ntrying to keep a clone mirror up to date, but do _not_ want to pass\nalong upstream's mirror file.\n\nYou may want to have a server implementation that offers a dynamic\nmirror, rather than a true object we have in the ODB. That is possible\nwith a mirror blob, but is slightly harder (you have to fake the object\nrather than just dumping a line).\n\n> These pkt-lines need to be a bit more than just URL. Or we need a new\n> URL like \"bundle:http://....\" to denote a resumable bundle over HTTP\n> vs. a normal HTTP URL that might not be a bundle file, and is just a\n> better connected server.\n\nRight, I think that's the most critical one (though you could also just\nuse the convention of \".bundle\" in the URL). I think we may want to\nleave room for more metadata, though.\n\n> The mirror URLs could be stored in $GIT_DIR/config as a simple\n> multi-value variable. Unfortunately that isn't easily remotely\n> editable. But I am not sure I care?\n\nFor big sites that manage the bundles on behalf of the user, I don't\nthink it is an issue. For somebody running their own small site, I think\nit is a useful way of moving the data to the server.\n\n> For the average home user sharing their working repository over git://\n> from their home ADSL or cable connection, editing .git/config is\n> easier than a blob in refs/mirrors. They already know how to edit\n> .git/config to manage remotes.\n\nYes, but it's editing .git/config on the server, not on the client,\nwhich may be slightly harder for some people. I do think we'd want\nsome tool support on the client side. git-config recently learned to\nread from a blob. The next step is:\n\n  git config --blob=refs/mirrors --edit\n\nor\n\n  git config --blob=refs/mirrors mirror.ko.url git://git.kernel.org/...\n  git config --blob=refs/mirrors mirror.ko.bundle true\n\nWe can't add tool support for editing .git/config on the server side,\nbecause the method for doing so isn't standard.\n\n> Heck, remote.origin.url might already\n> be a good mirror address to advertise, especially if the client isn't\n> on the same /24 as the server and the remote.origin.url is something\n> like \"git.kernel.org\". :-)\n\nYou could have a \"git-advertise-upstream\" that generates a mirror blob\nfrom your remotes config and pushes it to your publishing point. That\nmay be overkill, but I don't think it's possible with a\n.git/config-based solution.\n\n> > That's clever. It does not work out of the box if you are using\n> > alternates, but I think it could be adapted in certain situations. E.g.,\n> > if you layer the pack so that one \"base\" repo always has its full pack\n> > at the start, which is something we're already doing at GitHub.\n> \n> Yes, well, I was assuming the pack was a fully connected repack.\n> Alternates always creates a partial pack. But if you have an\n> alternate, that alternate maybe should be given as a mirror URL? And\n> allow the client to recurse the alternate mirror URL list too?\n\nThe problem for us is not that we have a partial pack, but that the\nalternates pack has a lot of other junk in it. A linux.git clone is\n650MB or so. The packfile for all of the linux.git forks together on\nGitHub is several gigabytes.\n\n> What really got us worried was the bundle header has no checksums, and\n> a resume in the bundle header from the wrong version could be\n> interesting.\n\nThe bundle header is small enough that you should just throw it away if\nyou didn't get the whole thing (IIRC, that is what my patches do,\nbecause it does not do _anything_ until we receive the whole ref\nadvertisement, at which point we decide if it is smart, dumb, or a\nbundle).\n\n> Yes. And this is why the packfile name algorithm is horribly flawed. I\n> keep saying we should change it to name the pack using the last 20\n> bytes of the file but ... nobody has written the patch for that?  :-)\n\nTotally agree. I think we could also get rid of the horrible hacks in\nrepack where we pack to a tempfile, then have to do another tempfile\ndance (which is not atomic!) to move the same-named packfile out of the\nway. If the name were based on the content, we could just throw away our\nnew pack if one of the same name is already there (just like we do for\nloose objects).\n\nI haven't looked at making such a patch, but I think it shouldn't be too\ncomplicated. My big worry would be weird fallouts from some hidden part\nof the code that we don't realize is depending on the current naming\nscheme. :)\n\n-Peff\n"},{"id":"231616","messageId":"20131205161217.GB27869@sigill.intra.peff.net","threadId":"35417","inReplyTo":"52A07DC5.5090508@alum.mit.edu","subject":"Re: How to resume broke clone ?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-12-05T16:12:17Z","receivedAt":"2013-12-05T16:12:17Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Dec 05, 2013 at 02:21:09PM +0100, Michael Haggerty wrote:\n\n> A better alternative would be to ask users to clone from the central\n> server.  In this case, the central server would want to tell the clients\n> to grab what they can from their local bootstrap mirror and then come\n> back to the central server for any remainders.  The trick is that which\n> bootstrap mirror is \"local\" would vary from client to client.\n>\n> I suppose that this could be implemented using what you have discussed\n> by having the central server direct the client to a URL that resolves\n> differently for different clients, CDN-like.  Alternatively, the central\n> Git server could itself look where a request is coming from and use some\n> intelligence to redirect the client to the closest bootstrap mirror from\n> its own list.  Or the server could pass the client a list of known\n> mirrors, and the client could try to determine which one is closest (and\n> reachable!).\n\nExactly. I think this will mostly happen via CDN, but I had also\nenvisioned that the server could add metadata to a list of possible\nmirrors, like:\n\n       [mirror \"ko-us\"]\n       url = http://git.us.kernel.org/...\n       zone = us\n\n       [mirror \"ko-cn\"]\n       url = http://git.cn.kernel.org/...\n       zone = cn\n\nIf the \"zone\" keys follow a micro-format convention, then the client\nknows that it prefers \"cn\" over \"us\" (either on the command line, or a\nlocal config option in ~/.gitconfig).\n\nThe biggest problem with all of this is that the server has to know\nabout the mirrors. If you want to set up an in-house mirror for\nsomething hosted on GitHub, but its only available to people in your\ncompany, then GitHub would not want to advertise it. You need some way\nto tell your clients about the mirror (and that is the inverse-mirror\n\"fetch from the mirror, which tells you it is just a bootstrap and to\nnow switch to the real repo\" scheme that I think you were describing\nearlier).\n\n-Peff\n"},{"id":"231617","messageId":"xmqqeh5ri3d3.fsf@gitster.dls.corp.google.com","threadId":"35417","inReplyTo":"20131205160418.GA27869@sigill.intra.peff.net","subject":"Re: How to resume broke clone ?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-12-05T18:01:28Z","receivedAt":"2013-12-05T18:01:28Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> Right, I think that's the most critical one (though you could also just\n> use the convention of \".bundle\" in the URL). I think we may want to\n> leave room for more metadata, though.\n\nGood. I like this line of thinking.\n\n>> Heck, remote.origin.url might already\n>> be a good mirror address to advertise, especially if the client isn't\n>> on the same /24 as the server and the remote.origin.url is something\n>> like \"git.kernel.org\". :-)\n>\n> You could have a \"git-advertise-upstream\" that generates a mirror blob\n> from your remotes config and pushes it to your publishing point. That\n> may be overkill, but I don't think it's possible with a\n> .git/config-based solution.\n\nI do not think I follow.  The upload-pack service could be taught to\npay attention to the uploadpack.advertiseUpstream config at runtime,\nadvertise 'mirror' capability, and then respond with the list of\nremote.*.url it uses when asked (if we go with the pkt-line based\napproach).  Alternatively, it could also be taught to pay attention\nto the same config at runtime, create an blob to advertise the list\nof remote.*.url it uses and store it in refs/mirror (or do this\npurely in-core without actually writing to the refs/ namespace), and\nemit an entry for refs/mirror using that blob object name in the\nls-remote part of the response (if we go with the magic blob based\napproach).\n\n>> Yes. And this is why the packfile name algorithm is horribly flawed. I\n>> keep saying we should change it to name the pack using the last 20\n>> bytes of the file but ... nobody has written the patch for that?  :-)\n>\n> Totally agree. I think we could also get rid of the horrible hacks in\n> repack where we pack to a tempfile, then have to do another tempfile\n> dance (which is not atomic!) to move the same-named packfile out of the\n> way. If the name were based on the content, we could just throw away our\n> new pack if one of the same name is already there (just like we do for\n> loose objects).\n\nYay.\n"},{"id":"231622","messageId":"20131205190824.GA19039@sigill.intra.peff.net","threadId":"35417","inReplyTo":"xmqqeh5ri3d3.fsf@gitster.dls.corp.google.com","subject":"Re: How to resume broke clone ?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-12-05T19:08:24Z","receivedAt":"2013-12-05T19:08:24Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Dec 05, 2013 at 10:01:28AM -0800, Junio C Hamano wrote:\n\n> > You could have a \"git-advertise-upstream\" that generates a mirror blob\n> > from your remotes config and pushes it to your publishing point. That\n> > may be overkill, but I don't think it's possible with a\n> > .git/config-based solution.\n> \n> I do not think I follow.  The upload-pack service could be taught to\n> pay attention to the uploadpack.advertiseUpstream config at runtime,\n> advertise 'mirror' capability, and then respond with the list of\n> remote.*.url it uses when asked (if we go with the pkt-line based\n> approach).\n\nI was assuming a triangular workflow, where your publishing point (that\nother people will fetch from) does not know anything about the upstream.\nLike:\n\n  $ git clone git://git.kernel.org/pub/scm/git/git.git\n  $ hack hack hack; commit commit commit\n  $ git remote add me myserver:/var/git/git.git\n  $ git push me\n  $ git advertise-upstream origin me\n\nIf your publishing point is already fetching from another upstream, then\nyeah, I'd agree that dynamically generating it from the config is fine.\n\n> Alternatively, it could also be taught to pay attention\n> to the same config at runtime, create an blob to advertise the list\n> of remote.*.url it uses and store it in refs/mirror (or do this\n> purely in-core without actually writing to the refs/ namespace), and\n> emit an entry for refs/mirror using that blob object name in the\n> ls-remote part of the response (if we go with the magic blob based\n> approach).\n\nYes. The pkt-line versus refs distinction is purely a protocol issue.\nYou can do anything you want on the backend with either of them,\nincluding faking the ref (you can also accept fake pushes to\nrefs/mirror, too, if you really want people to be able to upload that\nway).\n\nBut it is worth considering what implementation difficulties we would\nrun across in either case. Producing a fake refs/mirror blob that\nresponds like a normal ref is more work than just dumping the lines. If\nwe're always just going to generate it dynamically anyway, then we can\nsave ourselves some effort.\n\n-Peff\n"},{"id":"231635","messageId":"20131205202807.GA19042@sigill.intra.peff.net","threadId":"35417","inReplyTo":"20131205160418.GA27869@sigill.intra.peff.net","subject":"[PATCH] pack-objects: name pack files after trailer hash","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-12-05T20:28:07Z","receivedAt":"2013-12-05T20:28:07Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Dec 05, 2013 at 11:04:18AM -0500, Jeff King wrote:\n\n> > Yes. And this is why the packfile name algorithm is horribly flawed. I\n> > keep saying we should change it to name the pack using the last 20\n> > bytes of the file but ... nobody has written the patch for that?  :-)\n> \n> Totally agree. I think we could also get rid of the horrible hacks in\n> repack where we pack to a tempfile, then have to do another tempfile\n> dance (which is not atomic!) to move the same-named packfile out of the\n> way. If the name were based on the content, we could just throw away our\n> new pack if one of the same name is already there (just like we do for\n> loose objects).\n> \n> I haven't looked at making such a patch, but I think it shouldn't be too\n> complicated. My big worry would be weird fallouts from some hidden part\n> of the code that we don't realize is depending on the current naming\n> scheme. :)\n\nSo here's the first part, that actually changes the name. It passes the\ntest suite, so it must be good, right? And just look at that diffstat.\n\nThis actually applies on top of e74435a (sha1write: make buffer\nconst-correct, 2013-10-24), which is on another topic. Since the sha1\nparameter to write_idx_file is no longer used as both an in- and out-\nparameter (yuck), we can make it const.\n\nThe second half would be to simplify git-repack. The current behavior is\nto replace the old packfile with a tricky rename dance. Which is still\ncorrect, but overly complicated. We should be able to just drop the new\npackfile, since we know the bytes are identical (or rename the new one\nover the old, though I think keeping the old is probably kinder to the\ndisk cache, especially if another process already has it mmap'd).\n\n-- >8 --\nSubject: pack-objects: name pack files after trailer hash\n\nOur current scheme for naming packfiles is to calculate the\nsha1 hash of the sorted list of objects contained in the\npackfile. This gives us a unique name, so we are reasonably\nsure that two packs with the same name will contain the same\nobjects.\n\nIt does not, however, tell us that two such packs have the\nexact same bytes. This makes things awkward if we repack the\nsame set of objects. Due to run-to-run variations, the bytes\nmay not be identical (e.g., changed zlib or git versions,\ndifferent source object reuse due to new packs in the\nrepository, or even different deltas due to races during a\nmulti-threaded delta search).\n\nIn theory, this could be helpful to a program that cares\nthat the packfile contains a certain set of objects, but\ndoes not care about the particular representation. In\npractice, no part of git makes use of that, and in many\ncases it is potentially harmful. For example, if a dumb http\nclient fetches the .idx file, it must be sure to get the\nexact .pack that matches it. Similarly, a partial transfer\nof a .pack file cannot be safely resumed, as the actual\nbytes may have changed.  This could also affect a local\nclient which opened the .idx and .pack files, closes the\n.pack file (due to memory or file descriptor limits), and\nthen re-opens a changed packfile.\n\nIn all of these cases, git can detect the problem, as we\nhave the sha1 of the bytes themselves in the pack trailer\n(which we verify on transfer), and the .idx file references\nthe trailer from the matching packfile. But it would be\nsimpler and more efficient to actually get the correct\nbytes, rather than noticing the problem and having to\nrestart the operation.\n\nThis patch simply uses the pack trailer sha1 as the pack\nname. It should be similarly unique, but covers the exact\nrepresentation of the objects. Other parts of git should not\ncare, as the pack name is returned by pack-objects and is\nessentially opaque.\n\nOne test needs to be updated, because it actually corrupts a\npack and expects that re-packing the corrupted bytes will\nuse the same name. It won't anymore, but we can easily just\nuse the name that pack-objects hands back.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n pack-write.c          | 8 +-------\n pack.h                | 2 +-\n t/t5302-pack-index.sh | 4 ++--\n 3 files changed, 4 insertions(+), 10 deletions(-)\n\ndiff --git a/pack-write.c b/pack-write.c\nindex ca9e63b..ddc174e 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -44,14 +44,13 @@ static int need_large_offset(off_t offset, const struct pack_idx_option *opts)\n  */\n const char *write_idx_file(const char *index_name, struct pack_idx_entry **objects,\n \t\t\t   int nr_objects, const struct pack_idx_option *opts,\n-\t\t\t   unsigned char *sha1)\n+\t\t\t   const unsigned char *sha1)\n {\n \tstruct sha1file *f;\n \tstruct pack_idx_entry **sorted_by_sha, **list, **last;\n \toff_t last_obj_offset = 0;\n \tuint32_t array[256];\n \tint i, fd;\n-\tgit_SHA_CTX ctx;\n \tuint32_t index_version;\n \n \tif (nr_objects) {\n@@ -114,9 +113,6 @@ const char *write_idx_file(const char *index_name, struct pack_idx_entry **objec\n \t}\n \tsha1write(f, array, 256 * 4);\n \n-\t/* compute the SHA1 hash of sorted object names. */\n-\tgit_SHA1_Init(&ctx);\n-\n \t/*\n \t * Write the actual SHA1 entries..\n \t */\n@@ -128,7 +124,6 @@ const char *write_idx_file(const char *index_name, struct pack_idx_entry **objec\n \t\t\tsha1write(f, &offset, 4);\n \t\t}\n \t\tsha1write(f, obj->sha1, 20);\n-\t\tgit_SHA1_Update(&ctx, obj->sha1, 20);\n \t\tif ((opts->flags & WRITE_IDX_STRICT) &&\n \t\t    (i && !hashcmp(list[-2]->sha1, obj->sha1)))\n \t\t\tdie(\"The same object %s appears twice in the pack\",\n@@ -178,7 +173,6 @@ const char *write_idx_file(const char *index_name, struct pack_idx_entry **objec\n \tsha1write(f, sha1, 20);\n \tsha1close(f, NULL, ((opts->flags & WRITE_IDX_VERIFY)\n \t\t\t    ? CSUM_CLOSE : CSUM_FSYNC));\n-\tgit_SHA1_Final(sha1, &ctx);\n \treturn index_name;\n }\n \ndiff --git a/pack.h b/pack.h\nindex aa6ee7d..12d9516 100644\n--- a/pack.h\n+++ b/pack.h\n@@ -76,7 +76,7 @@ struct pack_idx_entry {\n struct progress;\n typedef int (*verify_fn)(const unsigned char*, enum object_type, unsigned long, void*, int*);\n \n-extern const char *write_idx_file(const char *index_name, struct pack_idx_entry **objects, int nr_objects, const struct pack_idx_option *, unsigned char *sha1);\n+extern const char *write_idx_file(const char *index_name, struct pack_idx_entry **objects, int nr_objects, const struct pack_idx_option *, const unsigned char *sha1);\n extern int check_pack_crc(struct packed_git *p, struct pack_window **w_curs, off_t offset, off_t len, unsigned int nr);\n extern int verify_pack_index(struct packed_git *);\n extern int verify_pack(struct packed_git *, verify_fn fn, struct progress *, uint32_t);\ndiff --git a/t/t5302-pack-index.sh b/t/t5302-pack-index.sh\nindex fe82025..4bbb718 100755\n--- a/t/t5302-pack-index.sh\n+++ b/t/t5302-pack-index.sh\n@@ -174,11 +174,11 @@ test_expect_success \\\n test_expect_success \\\n     '[index v1] 5) pack-objects happily reuses corrupted data' \\\n     'pack4=$(git pack-objects test-4 <obj-list) &&\n-     test -f \"test-4-${pack1}.pack\"'\n+     test -f \"test-4-${pack4}.pack\"'\n \n test_expect_success \\\n     '[index v1] 6) newly created pack is BAD !' \\\n-    'test_must_fail git verify-pack -v \"test-4-${pack1}.pack\"'\n+    'test_must_fail git verify-pack -v \"test-4-${pack4}.pack\"'\n \n test_expect_success \\\n     '[index v2] 1) stream pack to repository' \\\n-- \n1.8.5.524.g6743da6\n"},{"id":"231640","messageId":"CAJo=hJtSppKYGSG9RS74AjDC_OfNy+EWWf+V7BETO0gASJS9gg@mail.gmail.com","threadId":"35417","inReplyTo":"20131205202807.GA19042@sigill.intra.peff.net","subject":"Re: [PATCH] pack-objects: name pack files after trailer hash","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2013-12-05T21:56:03Z","receivedAt":"2013-12-05T21:56:03Z","isPatch":true,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Thu, Dec 5, 2013 at 12:28 PM, Jeff King <peff@peff.net> wrote:\n> Subject: pack-objects: name pack files after trailer hash\n>\n> Our current scheme for naming packfiles is to calculate the\n> sha1 hash of the sorted list of objects contained in the\n> packfile. This gives us a unique name, so we are reasonably\n> sure that two packs with the same name will contain the same\n> objects.\n\nYay-by: Shawn Pearce <spearce@spearce.org>\n\n> ---\n>  pack-write.c          | 8 +-------\n>  pack.h                | 2 +-\n>  t/t5302-pack-index.sh | 4 ++--\n>  3 files changed, 4 insertions(+), 10 deletions(-)\n\nObviously this is correct given the diffstat. :-)\n"},{"id":"231642","messageId":"xmqq4n6m52fy.fsf@gitster.dls.corp.google.com","threadId":"35417","inReplyTo":"20131205202807.GA19042@sigill.intra.peff.net","subject":"Re: [PATCH] pack-objects: name pack files after trailer hash","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-12-05T22:59:45Z","receivedAt":"2013-12-05T22:59:45Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> The second half would be to simplify git-repack. The current behavior is\n> to replace the old packfile with a tricky rename dance. Which is still\n> correct, but overly complicated. We should be able to just drop the new\n> packfile, since we know the bytes are identical (or rename the new one\n> over the old, though I think keeping the old is probably kinder to the\n> disk cache, especially if another process already has it mmap'd).\n\nConcurred.\n\n> One test needs to be updated, because it actually corrupts a\n> pack and expects that re-packing the corrupted bytes will\n> use the same name. It won't anymore, but we can easily just\n> use the name that pack-objects hands back.\n\nRe-reading the tests in that script, I am not sure if keeping these\ntests is even a sane thing to do, by the way.  It \"expects\" that\ncertain breakages are propagated, and anybody who breaks that\nexpectation by improving pack-objects etc. to catch such breakages\nwill be yelled at by breaking the test that used to pass.\n\nSeeing that the way the test scripts are line-wrapped follows the\nancient convention, I suspect that this may be because it predates\nour more recent best practice to document known breakages with\ntest_expect_failure.\n\n> diff --git a/t/t5302-pack-index.sh b/t/t5302-pack-index.sh\n> index fe82025..4bbb718 100755\n> --- a/t/t5302-pack-index.sh\n> +++ b/t/t5302-pack-index.sh\n> @@ -174,11 +174,11 @@ test_expect_success \\\n>  test_expect_success \\\n>      '[index v1] 5) pack-objects happily reuses corrupted data' \\\n>      'pack4=$(git pack-objects test-4 <obj-list) &&\n> -     test -f \"test-4-${pack1}.pack\"'\n> +     test -f \"test-4-${pack4}.pack\"'\n>  \n>  test_expect_success \\\n>      '[index v1] 6) newly created pack is BAD !' \\\n> -    'test_must_fail git verify-pack -v \"test-4-${pack1}.pack\"'\n> +    'test_must_fail git verify-pack -v \"test-4-${pack4}.pack\"'\n\nA good thing is that the above hunks are the right thing to do, even\nif we are to modernise these tests so that they document a known\nbreakage with expect-failure.\n\nThanks.\n"},{"id":"231683","messageId":"20131206221805.GE25620@sigill.intra.peff.net","threadId":"35417","inReplyTo":"xmqq4n6m52fy.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH] pack-objects: name pack files after trailer hash","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-12-06T22:18:06Z","receivedAt":"2013-12-06T22:18:06Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Dec 05, 2013 at 02:59:45PM -0800, Junio C Hamano wrote:\n\n> > One test needs to be updated, because it actually corrupts a\n> > pack and expects that re-packing the corrupted bytes will\n> > use the same name. It won't anymore, but we can easily just\n> > use the name that pack-objects hands back.\n> \n> Re-reading the tests in that script, I am not sure if keeping these\n> tests is even a sane thing to do, by the way.  It \"expects\" that\n> certain breakages are propagated, and anybody who breaks that\n> expectation by improving pack-objects etc. to catch such breakages\n> will be yelled at by breaking the test that used to pass.\n\nI had a similar thought, but I figured I would leave it for the person\nwho _does_ make that change. The yelling will be a good signal that\nthey've got it right, and they can clean up the test (either by dropping\nit, or modifying it to check the right thing) at that point.\n\n> Seeing that the way the test scripts are line-wrapped follows the\n> ancient convention, I suspect that this may be because it predates\n> our more recent best practice to document known breakages with\n> test_expect_failure.\n\nI read it more as \"make sure that the v1 index breaks, so when we are\ntesting v2 we know it is not an accident that we notice the breakage\".\n\nBut I also see your reason, and I think it would be fine to use\ntest_expect_failure.\n\n-Peff\n"},{"id":"232047","messageId":"52AEAEB2.6060203@alum.mit.edu","threadId":"35417","inReplyTo":"20131205202807.GA19042@sigill.intra.peff.net","subject":"Re: [PATCH] pack-objects: name pack files after trailer hash","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2013-12-16T07:41:38Z","receivedAt":"2013-12-16T07:41:38Z","isPatch":true,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 12/05/2013 09:28 PM, Jeff King wrote:\n> [...]\n> This patch simply uses the pack trailer sha1 as the pack\n> name. It should be similarly unique, but covers the exact\n> representation of the objects. Other parts of git should not\n> care, as the pack name is returned by pack-objects and is\n> essentially opaque.\n> [...]\n\nPeff,\n\nThe old naming scheme is documented in\nDocumentation/git-pack-objects.txt, under \"OPTIONS\" -> \"base-name\":\n\n> base-name::\n> \tWrite into a pair of files (.pack and .idx), using\n> \t<base-name> to determine the name of the created file.\n> \tWhen this option is used, the two files are written in\n> \t<base-name>-<SHA-1>.{pack,idx} files.  <SHA-1> is a hash\n> \tof the sorted object names to make the resulting filename\n> \tbased on the pack content, and written to the standard\n> \toutput of the command.\n\nThe documentation should either be updated or the description of the\nnaming scheme should be removed altogether.\n\nMichael\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"232061","messageId":"20131216190445.GB29324@sigill.intra.peff.net","threadId":"35417","inReplyTo":"52AEAEB2.6060203@alum.mit.edu","subject":"Re: [PATCH] pack-objects: name pack files after trailer hash","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-12-16T19:04:45Z","receivedAt":"2013-12-16T19:04:45Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Dec 16, 2013 at 08:41:38AM +0100, Michael Haggerty wrote:\n\n> The old naming scheme is documented in\n> Documentation/git-pack-objects.txt, under \"OPTIONS\" -> \"base-name\":\n> \n> > base-name::\n> > \tWrite into a pair of files (.pack and .idx), using\n> > \t<base-name> to determine the name of the created file.\n> > \tWhen this option is used, the two files are written in\n> > \t<base-name>-<SHA-1>.{pack,idx} files.  <SHA-1> is a hash\n> > \tof the sorted object names to make the resulting filename\n> > \tbased on the pack content, and written to the standard\n> > \toutput of the command.\n> \n> The documentation should either be updated or the description of the\n> naming scheme should be removed altogether.\n\nThanks. I looked in Documentation/technical for anything to update, but\ndidn't imagine we would be advertising the format in the user-facing\ndocumentation. :)\n\nThe original patch is in next, so here's one on top. I just updated the\ndescription. I was tempted to explicitly say something like \"this is\nopaque and meaningless to you, don't rely on it\", but I don't know that\nthere is any need.\n\n-- >8 --\nSubject: docs: update pack-objects \"base-name\" description\n\nAs of 1190a1a, the SHA-1 used to determine the filename is\nnow calculated differently. Update the documentation to\nreflect this.\n\nNoticed-by: Michael Haggerty <mhagger@alum.mit.edu>\nSigned-off-by: Jeff King <peff@peff.net>\n---\nOn top of jk/name-pack-after-byte-representations, naturally.\n\n Documentation/git-pack-objects.txt | 3 +--\n 1 file changed, 1 insertion(+), 2 deletions(-)\n\ndiff --git a/Documentation/git-pack-objects.txt b/Documentation/git-pack-objects.txt\nindex d94edcd..c69affc 100644\n--- a/Documentation/git-pack-objects.txt\n+++ b/Documentation/git-pack-objects.txt\n@@ -51,8 +51,7 @@ base-name::\n \t<base-name> to determine the name of the created file.\n \tWhen this option is used, the two files are written in\n \t<base-name>-<SHA-1>.{pack,idx} files.  <SHA-1> is a hash\n-\tof the sorted object names to make the resulting filename\n-\tbased on the pack content, and written to the standard\n+\tof the bytes of the packfile, and is written to the standard\n \toutput of the command.\n \n --stdout::\n-- \n1.8.5.524.g6743da6\n"},{"id":"232065","messageId":"20131216191933.GE2311@google.com","threadId":"35417","inReplyTo":"20131216190445.GB29324@sigill.intra.peff.net","subject":"Re: [PATCH] pack-objects: name pack files after trailer hash","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2013-12-16T19:19:33Z","receivedAt":"2013-12-16T19:19:33Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Jeff King wrote:\n\n> The original patch is in next, so here's one on top. I just updated the\n> description.\n\nThanks.\n\n>              I was tempted to explicitly say something like \"this is\n> opaque and meaningless to you, don't rely on it\", but I don't know that\n> there is any need.\n[...]\n> On top of jk/name-pack-after-byte-representations, naturally.\n\nI think there is --- if someone starts caring about the SHA-1 used,\nthey won't be able to act on old packfiles that were created before\nthis change.  How about something like the following instead?\n\n-- >8 --\nFrom: Jeff King <peff@peff.net>\nSubject: pack-objects doc: treat output filename as opaque\n\nAfter 1190a1a (pack-objects: name pack files after trailer hash,\n2013-12-05), the SHA-1 used to determine the filename is calculated\ndifferently.  Update the documentation to not guarantee anything more\nthan that the SHA-1 depends on the pack content somehow.\n\nHopefully this will discourage readers from depending on the old or\nthe new calculation.\n\nReported-by: Michael Haggerty <mhagger@alum.mit.edu>\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Jonathan Nieder <jrnieder@gmail.com>\n---\n Documentation/git-pack-objects.txt | 3 +--\n 1 file changed, 1 insertion(+), 2 deletions(-)\n\ndiff --git a/Documentation/git-pack-objects.txt b/Documentation/git-pack-objects.txt\nindex d94edcd..cdab9ed 100644\n--- a/Documentation/git-pack-objects.txt\n+++ b/Documentation/git-pack-objects.txt\n@@ -51,8 +51,7 @@ base-name::\n \t<base-name> to determine the name of the created file.\n \tWhen this option is used, the two files are written in\n \t<base-name>-<SHA-1>.{pack,idx} files.  <SHA-1> is a hash\n-\tof the sorted object names to make the resulting filename\n-\tbased on the pack content, and written to the standard\n+\tbased on the pack content and is written to the standard\n \toutput of the command.\n \n --stdout::\n-- \n1.8.5.1\n"},{"id":"232069","messageId":"20131216192830.GA30238@sigill.intra.peff.net","threadId":"35417","inReplyTo":"20131216191933.GE2311@google.com","subject":"Re: [PATCH] pack-objects: name pack files after trailer hash","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-12-16T19:28:30Z","receivedAt":"2013-12-16T19:28:30Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Dec 16, 2013 at 11:19:33AM -0800, Jonathan Nieder wrote:\n\n> >              I was tempted to explicitly say something like \"this is\n> > opaque and meaningless to you, don't rely on it\", but I don't know that\n> > there is any need.\n> [...]\n> > On top of jk/name-pack-after-byte-representations, naturally.\n> \n> I think there is --- if someone starts caring about the SHA-1 used,\n> they won't be able to act on old packfiles that were created before\n> this change.  How about something like the following instead?\n\nRight, my point was that I do not think anybody has ever cared, and I do\nnot see them starting now. But that is just my intuition.\n\n> diff --git a/Documentation/git-pack-objects.txt b/Documentation/git-pack-objects.txt\n> index d94edcd..cdab9ed 100644\n> --- a/Documentation/git-pack-objects.txt\n> +++ b/Documentation/git-pack-objects.txt\n> @@ -51,8 +51,7 @@ base-name::\n>  \t<base-name> to determine the name of the created file.\n>  \tWhen this option is used, the two files are written in\n>  \t<base-name>-<SHA-1>.{pack,idx} files.  <SHA-1> is a hash\n> -\tof the sorted object names to make the resulting filename\n> -\tbased on the pack content, and written to the standard\n> +\tbased on the pack content and is written to the standard\n\nI'm fine with that. I was worried it would get clunky, but the way you\nhave worded it is good.\n\n-Peff\n"},{"id":"232070","messageId":"xmqqzjo0oako.fsf@gitster.dls.corp.google.com","threadId":"35417","inReplyTo":"20131216190445.GB29324@sigill.intra.peff.net","subject":"Re: [PATCH] pack-objects: name pack files after trailer hash","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-12-16T19:33:11Z","receivedAt":"2013-12-16T19:33:11Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> I was tempted to explicitly say something like \"this is\n> opaque and meaningless to you, don't rely on it\", but I don't know that\n> there is any need.\n\nThanks.\n\nWhen we did the original naming, it was envisioned that we may use\nthe name for fsck to make sure that the pack contains what it\ncontains in the name, but it never materialized.  The most prominent\nand useful characteristic of the new naming scheme is that two\npackfiles with the same name must be identical, and we may want to\nstart using it some time later once everybody repacked their packs\nwith the updated pack-objects.\n\nBut until that time comes, some packs in existing repositories will\nhash to their names while others do not, so spelling out how the new\nnames are derived without saying older pack-objects used to name\ntheir output differently may add more confusion than it is worth.\n\n>  \t<base-name> to determine the name of the created file.\n>  \tWhen this option is used, the two files are written in\n>  \t<base-name>-<SHA-1>.{pack,idx} files.  <SHA-1> is a hash\n> +\tof the bytes of the packfile, and is written to the standard\n\n\"hash of the bytes of the packfile\" tempts one to do\n\n    $ sha1sum .git/objects/pack/pack-*.pack\n\nbut that is not what we expect. I wonder if there are better ways to\nphrase it (or alternatively perhaps we want to make that expectation\nhold by updating our code to hash)?\n"},{"id":"232071","messageId":"20131216193533.GA3488@sigill.intra.peff.net","threadId":"35417","inReplyTo":"xmqqzjo0oako.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH] pack-objects: name pack files after trailer hash","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-12-16T19:35:33Z","receivedAt":"2013-12-16T19:35:33Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Dec 16, 2013 at 11:33:11AM -0800, Junio C Hamano wrote:\n\n> >  \t<base-name> to determine the name of the created file.\n> >  \tWhen this option is used, the two files are written in\n> >  \t<base-name>-<SHA-1>.{pack,idx} files.  <SHA-1> is a hash\n> > +\tof the bytes of the packfile, and is written to the standard\n> \n> \"hash of the bytes of the packfile\" tempts one to do\n> \n>     $ sha1sum .git/objects/pack/pack-*.pack\n> \n> but that is not what we expect. I wonder if there are better ways to\n> phrase it (or alternatively perhaps we want to make that expectation\n> hold by updating our code to hash)?\n\nYeah, I wondered about that, but didn't think it was worth the verbosity\nto explain that the true derivation. I think Jonathan's suggestion takes\ncare of it, though.\n\n-Peff\n"},{"id":"232072","messageId":"xmqqsitsoae5.fsf@gitster.dls.corp.google.com","threadId":"35417","inReplyTo":"20131216192830.GA30238@sigill.intra.peff.net","subject":"Re: [PATCH] pack-objects: name pack files after trailer hash","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-12-16T19:37:06Z","receivedAt":"2013-12-16T19:37:06Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> On Mon, Dec 16, 2013 at 11:19:33AM -0800, Jonathan Nieder wrote:\n>\n>> >              I was tempted to explicitly say something like \"this is\n>> > opaque and meaningless to you, don't rely on it\", but I don't know that\n>> > there is any need.\n>> [...]\n>> > On top of jk/name-pack-after-byte-representations, naturally.\n>> \n>> I think there is --- if someone starts caring about the SHA-1 used,\n>> they won't be able to act on old packfiles that were created before\n>> this change.  How about something like the following instead?\n>\n> Right, my point was that I do not think anybody has ever cared, and I do\n> not see them starting now. But that is just my intuition.\n>\n>> diff --git a/Documentation/git-pack-objects.txt b/Documentation/git-pack-objects.txt\n>> index d94edcd..cdab9ed 100644\n>> --- a/Documentation/git-pack-objects.txt\n>> +++ b/Documentation/git-pack-objects.txt\n>> @@ -51,8 +51,7 @@ base-name::\n>>  \t<base-name> to determine the name of the created file.\n>>  \tWhen this option is used, the two files are written in\n>>  \t<base-name>-<SHA-1>.{pack,idx} files.  <SHA-1> is a hash\n>> -\tof the sorted object names to make the resulting filename\n>> -\tbased on the pack content, and written to the standard\n>> +\tbased on the pack content and is written to the standard\n>\n> I'm fine with that. I was worried it would get clunky, but the way you\n> have worded it is good.\n\nOur mails crossed; I think the above is good.\n\nThanks.\n"}]}