{"thread":{"id":"53469","subject":"Add a \"Flattened Cache\" to `git --clone`?","startedAt":"2020-05-14T14:34:23Z","lastAt":"2020-05-25T14:02:41Z","messageCount":17,"participants":["Caleb Gray","Konstantin Ryabitsev","Bryan Turner","Theodore Y. Ts'o","Eric Sunshine","Junio C Hamano","Eric Wong"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"397806","messageId":"CAGjfG9a-MSg7v6+wynR1gL0zoe+Kv8HZfR8oxe+a3r59cGhEeg@mail.gmail.com","threadId":"53469","inReplyTo":null,"subject":"Add a \"Flattened Cache\" to `git --clone`?","fromName":"Caleb Gray","fromEmail":"hey@calebgray.com","sentAt":"2020-05-14T14:34:08Z","receivedAt":"2020-05-14T14:34:23Z","isPatch":false,"sender":{"key":"hey@calebgray.com","avatar":null},"body":"I've done some searching around the Internet, mailing lists, and\nreached out in IRC a couple of days ago... and haven't found anyone\nelse asking about a long-brewed contribution idea that I'd finally\nlike to implement. First I wanted to run it by you guys, though, since\nthis is my first time reaching out.\n\nAssuming my idea doesn't contradict other best practices or standards\nalready in place,  I'd like to transform the typical `git clone` flow\nfrom:\n\n Cloning into 'linux'...\n remote: Enumerating objects: 4154, done.\n remote: Counting objects: 100% (4154/4154), done.\n remote: Compressing objects: 100% (2535/2535), done.\n remote: Total 7344127 (delta 2564), reused 2167 (delta 1612),\npack-reused 7339973\n Receiving objects: 100% (7344127/7344127), 1.22 GiB | 8.51 MiB/s, done.\n Resolving deltas: 100% (6180880/6180880), done.\n\nTo subsequent clones (until cache invalidated) using the \"flattened\ncache\" version (presumably built while fulfilling the first clone\nrequest above):\n\n Cloning into 'linux'...\n Receiving cache: 100% (7344127/7344127), 1.22 GiB | 8.51 MiB/s, done.\n\nI've always imagined that this feature would only apply to a \"vanilla\"\nclone (that is, one without any flags that change the end result)...\nbut that's only because I've never actually cracked open the `git`\ncodebase yet to validate/invalidated the complexity of this feature.\nI'm writing in hopes that someone else has thought about it... and\nmight share what they already know. :P\n\nThanks so much for your time!\n\nSincerely,\nCaleb\n"},{"id":"397823","messageId":"20200514203326.2aqxolq5u75jx64q@chatter.i7.local","threadId":"53469","inReplyTo":"CAGjfG9a-MSg7v6+wynR1gL0zoe+Kv8HZfR8oxe+a3r59cGhEeg@mail.gmail.com","subject":"Re: Add a \"Flattened Cache\" to `git --clone`?","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2020-05-14T20:33:26Z","receivedAt":"2020-05-14T20:33:32Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On Thu, May 14, 2020 at 07:34:08AM -0700, Caleb Gray wrote:\n> I've done some searching around the Internet, mailing lists, and\n> reached out in IRC a couple of days ago... and haven't found anyone\n> else asking about a long-brewed contribution idea that I'd finally\n> like to implement. First I wanted to run it by you guys, though, since\n> this is my first time reaching out.\n> \n> Assuming my idea doesn't contradict other best practices or standards\n> already in place,  I'd like to transform the typical `git clone` flow\n> from:\n> \n>  Cloning into 'linux'...\n>  remote: Enumerating objects: 4154, done.\n>  remote: Counting objects: 100% (4154/4154), done.\n>  remote: Compressing objects: 100% (2535/2535), done.\n>  remote: Total 7344127 (delta 2564), reused 2167 (delta 1612),\n> pack-reused 7339973\n>  Receiving objects: 100% (7344127/7344127), 1.22 GiB | 8.51 MiB/s, done.\n>  Resolving deltas: 100% (6180880/6180880), done.\n> \n> To subsequent clones (until cache invalidated) using the \"flattened\n> cache\" version (presumably built while fulfilling the first clone\n> request above):\n> \n>  Cloning into 'linux'...\n>  Receiving cache: 100% (7344127/7344127), 1.22 GiB | 8.51 MiB/s, done.\n\nI don't think it's a common workflow for someone to repeatedly clone \nlinux.git. Automated processes like CI would be doing it, but they tend \nto blow away the local disk between jobs, so they are unlikely to \nbenefit from any native git local cache for something like this (in \nfact, we recommend that people use clone.bundle files for their CI \nneeds, as described here: \nhttps://www.kernel.org/best-way-to-do-linux-clones-for-your-ci.html).\n\nI believe there's quite a bit of work being done by Gitlab folks to make \nit possible to offload more object fetching to lookaside-caches like \nCDN. Perhaps one of them can provide an update on how that is going.\n\n-K\n"},{"id":"397824","messageId":"CAGyf7-E6amUCOs7fZ_7Zfjfx5qTwytM+zROZQqaM6NML2Ci-Zw@mail.gmail.com","threadId":"53469","inReplyTo":"20200514203326.2aqxolq5u75jx64q@chatter.i7.local","subject":"Re: Add a \"Flattened Cache\" to `git --clone`?","fromName":"Bryan Turner","fromEmail":"bturner@atlassian.com","sentAt":"2020-05-14T20:54:00Z","receivedAt":"2020-05-14T20:54:13Z","isPatch":false,"sender":{"key":"bturner@atlassian.com","avatar":"https://gravatar.com/avatar/16bcf3167981c1ef7c804e502642366d888a35b0d0b0a4ca01fdc442aa1acb1e?d=mp&s=160"},"body":"On Thu, May 14, 2020 at 1:33 PM Konstantin Ryabitsev\n<konstantin@linuxfoundation.org> wrote:\n>\n> On Thu, May 14, 2020 at 07:34:08AM -0700, Caleb Gray wrote:\n> > I've done some searching around the Internet, mailing lists, and\n> > reached out in IRC a couple of days ago... and haven't found anyone\n> > else asking about a long-brewed contribution idea that I'd finally\n> > like to implement. First I wanted to run it by you guys, though, since\n> > this is my first time reaching out.\n> >\n> > Assuming my idea doesn't contradict other best practices or standards\n> > already in place,  I'd like to transform the typical `git clone` flow\n> > from:\n> >\n> >  Cloning into 'linux'...\n> >  remote: Enumerating objects: 4154, done.\n> >  remote: Counting objects: 100% (4154/4154), done.\n> >  remote: Compressing objects: 100% (2535/2535), done.\n> >  remote: Total 7344127 (delta 2564), reused 2167 (delta 1612),\n> > pack-reused 7339973\n> >  Receiving objects: 100% (7344127/7344127), 1.22 GiB | 8.51 MiB/s, done.\n> >  Resolving deltas: 100% (6180880/6180880), done.\n> >\n> > To subsequent clones (until cache invalidated) using the \"flattened\n> > cache\" version (presumably built while fulfilling the first clone\n> > request above):\n> >\n> >  Cloning into 'linux'...\n> >  Receiving cache: 100% (7344127/7344127), 1.22 GiB | 8.51 MiB/s, done.\n>\n> I don't think it's a common workflow for someone to repeatedly clone\n> linux.git. Automated processes like CI would be doing it, but they tend\n> to blow away the local disk between jobs, so they are unlikely to\n> benefit from any native git local cache for something like this (in\n> fact, we recommend that people use clone.bundle files for their CI\n> needs, as described here:\n> https://www.kernel.org/best-way-to-do-linux-clones-for-your-ci.html).\n>\n> I believe there's quite a bit of work being done by Gitlab folks to make\n> it possible to offload more object fetching to lookaside-caches like\n> CDN. Perhaps one of them can provide an update on how that is going.\n\nI can't speak for Gitlab, but Bitbucket Server (formerly Stash) has\ndone this for years, and I believe Github does as well. For Bitbucket\nServer, our caching doesn't change what the client sees (i.e. they\nstill see \"Counting objects\", \"Compressing objects\"), but the early\nsteps essentially jump straight to 100% (since that progress\ninformation is included in our cached data) and then the client starts\nreceiving the pack.\n\nI'm not sure how straightforward--or desirable--it would be for\nsomething like this to be done natively by Git itself. Certainly it\nwould make building hosting solutions simpler, which could be a win\nfor simpler setups that don't use something like Bitbucket Server,\nGitlab or Github, but I'm not sure that's a big \"win\". Effort on\nsomething like clonebundles (in Mercurial parlance) or similar seems\nlikely to offer a lot more bang for the buck than caching packs for\nspecific wants/haves.\n\nJust my 2 cents as someone who has directly worked on this sort of caching.\n\nBryan Turner\n"},{"id":"397825","messageId":"20200514210501.GY1596452@mit.edu","threadId":"53469","inReplyTo":"20200514203326.2aqxolq5u75jx64q@chatter.i7.local","subject":"Re: Add a \"Flattened Cache\" to `git --clone`?","fromName":"Theodore Y. Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2020-05-14T21:05:01Z","receivedAt":"2020-05-14T21:05:09Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Thu, May 14, 2020 at 04:33:26PM -0400, Konstantin Ryabitsev wrote:\n> > Assuming my idea doesn't contradict other best practices or standards\n> > already in place,  I'd like to transform the typical `git clone` flow\n> > from:\n> > \n> >  Cloning into 'linux'...\n> >  remote: Enumerating objects: 4154, done.\n> >  remote: Counting objects: 100% (4154/4154), done.\n> >  remote: Compressing objects: 100% (2535/2535), done.\n> >  remote: Total 7344127 (delta 2564), reused 2167 (delta 1612),\n> > pack-reused 7339973\n> >  Receiving objects: 100% (7344127/7344127), 1.22 GiB | 8.51 MiB/s, done.\n> >  Resolving deltas: 100% (6180880/6180880), done.\n> > \n> > To subsequent clones (until cache invalidated) using the \"flattened\n> > cache\" version (presumably built while fulfilling the first clone\n> > request above):\n> > \n> >  Cloning into 'linux'...\n> >  Receiving cache: 100% (7344127/7344127), 1.22 GiB | 8.51 MiB/s, done.\n> \n> I don't think it's a common workflow for someone to repeatedly clone \n> linux.git. Automated processes like CI would be doing it, but they tend \n> to blow away the local disk between jobs, so they are unlikely to \n> benefit from any native git local cache for something like this (in \n> fact, we recommend that people use clone.bundle files for their CI \n> needs, as described here: \n> https://www.kernel.org/best-way-to-do-linux-clones-for-your-ci.html).\n\nIf the goal is a git local cache, we have this today.  I'm not sure\nthis is what Caleb was asking for, though:\n\ngit clone --bare https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git base\ngit clone --reference base https://git.kernel.org/pub/scm/linux/kernel/git/tytso/ext4.git ext4\n\n\t\t\t\t\t\t\t- Ted\n"},{"id":"397827","messageId":"CAPig+cR4GBkd2==5G1d_514bcsAULYaFQatQg7odmO-ZmHHohg@mail.gmail.com","threadId":"53469","inReplyTo":"20200514210501.GY1596452@mit.edu","subject":"Re: Add a \"Flattened Cache\" to `git --clone`?","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2020-05-14T21:09:56Z","receivedAt":"2020-05-14T21:10:11Z","isPatch":false,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Thu, May 14, 2020 at 5:05 PM Theodore Y. Ts'o <tytso@mit.edu> wrote:\n> If the goal is a git local cache, we have this today.  I'm not sure\n> this is what Caleb was asking for, though:\n>\n> git clone --bare https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git base\n> git clone --reference base https://git.kernel.org/pub/scm/linux/kernel/git/tytso/ext4.git ext4\n\nFor that sort of use-case, git-worktree may also be a suitable solution.\n"},{"id":"397828","messageId":"20200514211040.a7hrirdzgkphx3la@chatter.i7.local","threadId":"53469","inReplyTo":"20200514210501.GY1596452@mit.edu","subject":"Re: Add a \"Flattened Cache\" to `git --clone`?","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2020-05-14T21:10:40Z","receivedAt":"2020-05-14T21:10:45Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On Thu, May 14, 2020 at 05:05:01PM -0400, Theodore Y. Ts'o wrote:\n> > \n> > I don't think it's a common workflow for someone to repeatedly clone \n> > linux.git. Automated processes like CI would be doing it, but they tend \n> > to blow away the local disk between jobs, so they are unlikely to \n> > benefit from any native git local cache for something like this (in \n> > fact, we recommend that people use clone.bundle files for their CI \n> > needs, as described here: \n> > https://www.kernel.org/best-way-to-do-linux-clones-for-your-ci.html).\n> \n> If the goal is a git local cache, we have this today.  I'm not sure\n> this is what Caleb was asking for, though:\n\nRight, I think I misunderstood his request -- I'd assumed we were \ntalking about local cache, whereas it's about server-side or proxy-side \ncache.\n\nI think something like git-caching-proxy would be a neat project, \nbecause it would significantly improve mirroring for CI deployments \nwithout requiring that each individual job implements clone.bundle \nprefetching.\n\n-K\n"},{"id":"397829","messageId":"xmqqv9kyp63p.fsf@gitster.c.googlers.com","threadId":"53469","inReplyTo":"20200514203326.2aqxolq5u75jx64q@chatter.i7.local","subject":"Re: Add a \"Flattened Cache\" to `git --clone`?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2020-05-14T21:19:22Z","receivedAt":"2020-05-14T21:19:30Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Konstantin Ryabitsev <konstantin@linuxfoundation.org> writes:\n\n> On Thu, May 14, 2020 at 07:34:08AM -0700, Caleb Gray wrote:\n>> ...\n>> To subsequent clones (until cache invalidated) using the \"flattened\n>> cache\" version (presumably built while fulfilling the first clone\n>> request above):\n>> \n>>  Cloning into 'linux'...\n>>  Receiving cache: 100% (7344127/7344127), 1.22 GiB | 8.51 MiB/s, done.\n>\n> I don't think it's a common workflow for someone to repeatedly clone \n> linux.git. Automated processes like CI would be doing it, but they tend \n> to blow away the local disk between jobs, so they are unlikely to \n> benefit from any native git local cache for something like this (in \n> fact, we recommend that people use clone.bundle files for their CI \n> needs, as described here: \n> https://www.kernel.org/best-way-to-do-linux-clones-for-your-ci.html).\n\nI have a feeling that the use case you are talking about is\ndifferent from what the original message assumes what use case needs\nto be helped (even though the original message lacks substance and\nit is hard to guess what idea is being proposed).  \n\nGiven the phrase like \"while fulfilling the first clone request\", I\ntook it to mean that a cache would sit on the source side, not on\nthe client side.  You seem to be talking about keeping a copy of\nwhat you earlier cloned to save incoming bandwidth on the client\nside.\n\n"},{"id":"397830","messageId":"xmqqr1vmp5wf.fsf@gitster.c.googlers.com","threadId":"53469","inReplyTo":"20200514211040.a7hrirdzgkphx3la@chatter.i7.local","subject":"Re: Add a \"Flattened Cache\" to `git --clone`?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2020-05-14T21:23:44Z","receivedAt":"2020-05-14T21:23:52Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Konstantin Ryabitsev <konstantin@linuxfoundation.org> writes:\n\n> I think something like git-caching-proxy would be a neat project, \n> because it would significantly improve mirroring for CI deployments \n> without requiring that each individual job implements clone.bundle \n> prefetching.\n\nWhat are we improving with such a proxy, though?\n\nNot bandwidth to the client, apparently.  I thought that with the\nreachability bitmap on the server side with reusing packed object,\nit was more or less a solved problem that the server end spends way\ntoo much time enumerating, deltifying and compressing the object\ndata?\n"},{"id":"397831","messageId":"CAGjfG9bsQh2C6WP242v4LoiaSdghZDPuqns0VO82Txe-V54_KA@mail.gmail.com","threadId":"53469","inReplyTo":"20200514210501.GY1596452@mit.edu","subject":"Re: Add a \"Flattened Cache\" to `git --clone`?","fromName":"Caleb Gray","fromEmail":"hey@calebgray.com","sentAt":"2020-05-14T21:33:06Z","receivedAt":"2020-05-14T21:33:23Z","isPatch":false,"sender":{"key":"hey@calebgray.com","avatar":null},"body":"To Clarify: I'm talking about a server-side only cache which behaves\nmuch like a `tar` file: it is a flat version of exactly(*) what ends\nup on the client's storage. When a client runs `git --clone` and\nthere's a valid cache on the other end, that's all that gets streamed.\n\nKonstantin's point that a repo like Linux is bound to see little/no\nbenefit (in fact, it'll just constantly invalidate/rewrite the ~1gb\ncache) is reasonable. This feature definitely targets the \"niche\"\naudience of repos with less-frequent-pushes-to-master-than-clones.\n\nBryan is exactly on the right track for what I'm referring to: the CDN\napproach did come to mind (and is superior in nearly every way).\n\nJunio nailed it: I'm not hoping for anything revolutionary here, just\nhoping to reduce the redundant steps in clone down to a single\n(presumably faster) step.\n\nIf the community agrees that there's little/no benefit to the\nlimitations of having a \"cache for master and that's all,\" I'm also\nmore than capable of designing a more useful/complex graph/reduce\nbased solution which could dynamically bundle the most statistically\nrelevant data for whatever context the code is working in, though-- I\ncan't commit to any sort of deadline for that sort of a contribution.\n\n\n\nOn Thu, May 14, 2020 at 2:05 PM Theodore Y. Ts'o <tytso@mit.edu> wrote:\n>\n> On Thu, May 14, 2020 at 04:33:26PM -0400, Konstantin Ryabitsev wrote:\n> > > Assuming my idea doesn't contradict other best practices or standards\n> > > already in place,  I'd like to transform the typical `git clone` flow\n> > > from:\n> > >\n> > >  Cloning into 'linux'...\n> > >  remote: Enumerating objects: 4154, done.\n> > >  remote: Counting objects: 100% (4154/4154), done.\n> > >  remote: Compressing objects: 100% (2535/2535), done.\n> > >  remote: Total 7344127 (delta 2564), reused 2167 (delta 1612),\n> > > pack-reused 7339973\n> > >  Receiving objects: 100% (7344127/7344127), 1.22 GiB | 8.51 MiB/s, done.\n> > >  Resolving deltas: 100% (6180880/6180880), done.\n> > >\n> > > To subsequent clones (until cache invalidated) using the \"flattened\n> > > cache\" version (presumably built while fulfilling the first clone\n> > > request above):\n> > >\n> > >  Cloning into 'linux'...\n> > >  Receiving cache: 100% (7344127/7344127), 1.22 GiB | 8.51 MiB/s, done.\n> >\n> > I don't think it's a common workflow for someone to repeatedly clone\n> > linux.git. Automated processes like CI would be doing it, but they tend\n> > to blow away the local disk between jobs, so they are unlikely to\n> > benefit from any native git local cache for something like this (in\n> > fact, we recommend that people use clone.bundle files for their CI\n> > needs, as described here:\n> > https://www.kernel.org/best-way-to-do-linux-clones-for-your-ci.html).\n>\n> If the goal is a git local cache, we have this today.  I'm not sure\n> this is what Caleb was asking for, though:\n>\n> git clone --bare https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git base\n> git clone --reference base https://git.kernel.org/pub/scm/linux/kernel/git/tytso/ext4.git ext4\n>\n>                                                         - Ted\n"},{"id":"397833","messageId":"20200514214404.bcbjskgi52bwedlh@chatter.i7.local","threadId":"53469","inReplyTo":"xmqqr1vmp5wf.fsf@gitster.c.googlers.com","subject":"Re: Add a \"Flattened Cache\" to `git --clone`?","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2020-05-14T21:44:04Z","receivedAt":"2020-05-14T21:44:09Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On Thu, May 14, 2020 at 02:23:44PM -0700, Junio C Hamano wrote:\n> > I think something like git-caching-proxy would be a neat project, \n> > because it would significantly improve mirroring for CI deployments \n> > without requiring that each individual job implements clone.bundle \n> > prefetching.\n> \n> What are we improving with such a proxy, though?\n> \n> Not bandwidth to the client, apparently. \n\nWell, if it sits in front of the CI subnet, then it *does* save \nbandwidth.\n\nHere's an example with the exact situation we have:\n\n- the Gerrit server is on the US West Coast\n- the CI builder is on the East Coast\n- each CI job does a full transfer of the multi-MB repo across the \n  continent, even when cloning shallow\n\nWe solve this by having a local mirror of the repository, but this \nrequires active mirroring to be pre-setup. A caching proxy that could:\n\n- receive a request for a repository\n- stream the response back to the client\n- cache objects locally\n- use local cache to construct future requests, so only missing objects \n  are fetched from the remote repo regardless of the haves on the actual \n  client...\n\n..now, that would be kinda neat, but I'm not sure how sane or fragile \nthat setup would be. :)\n\n> I thought that with the\n> reachability bitmap on the server side with reusing packed object,\n> it was more or less a solved problem that the server end spends way\n> too much time enumerating, deltifying and compressing the object\n> data?\n\nIndeed, it's not really solving anything for this case.\n\n-K\n"},{"id":"397835","messageId":"xmqqmu6ap4dw.fsf@gitster.c.googlers.com","threadId":"53469","inReplyTo":"CAGjfG9bsQh2C6WP242v4LoiaSdghZDPuqns0VO82Txe-V54_KA@mail.gmail.com","subject":"Re: Add a \"Flattened Cache\" to `git --clone`?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2020-05-14T21:56:27Z","receivedAt":"2020-05-14T21:56:33Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Caleb Gray <hey@calebgray.com> writes:\n\n> To Clarify: I'm talking about a server-side only cache which behaves\n> much like a `tar` file: it is a flat version of exactly(*) what ends\n> up on the client's storage. When a client runs `git --clone` and\n> there's a valid cache on the other end, that's all that gets streamed.\n\nSo this is to save server processing time only.  It does not save\nbandwidth (the \"cache\" is bit-for-bit idential replay of the clone\nrequest it served earlier), and it does not save client processing\ncycles (as the receiving end must validate the whole packdata it\nreceived before it can even know what objects it received).\n\nOK.\n\n\n\n\n"},{"id":"397837","messageId":"CAGjfG9akT+KG-tttRWEX_ZqxrqPoY_4Ed7Pymt4DkV5Rgc1CEA@mail.gmail.com","threadId":"53469","inReplyTo":"xmqqmu6ap4dw.fsf@gitster.c.googlers.com","subject":"Re: Add a \"Flattened Cache\" to `git --clone`?","fromName":"Caleb Gray","fromEmail":"hey@calebgray.com","sentAt":"2020-05-14T22:04:46Z","receivedAt":"2020-05-14T22:05:02Z","isPatch":false,"sender":{"key":"hey@calebgray.com","avatar":null},"body":"Actually those are the steps that I'm explicitly hoping can be\nskipped, both on server and client, after the first successful clone\nrequest transaction. The cache itself would be of the end resulting\n`.git` directory (client side)... unless I have misconceptions about\nthe complexity of reproducing what ends up on the client side from the\nserver side... I figured the shared library probably offers endpoints\nfor the information I'd need to achieve that.\n\n\nOn Thu, May 14, 2020 at 2:56 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Caleb Gray <hey@calebgray.com> writes:\n>\n> > To Clarify: I'm talking about a server-side only cache which behaves\n> > much like a `tar` file: it is a flat version of exactly(*) what ends\n> > up on the client's storage. When a client runs `git --clone` and\n> > there's a valid cache on the other end, that's all that gets streamed.\n>\n> So this is to save server processing time only.  It does not save\n> bandwidth (the \"cache\" is bit-for-bit idential replay of the clone\n> request it served earlier), and it does not save client processing\n> cycles (as the receiving end must validate the whole packdata it\n> received before it can even know what objects it received).\n>\n> OK.\n>\n>\n>\n>\n"},{"id":"397840","messageId":"xmqqeermp2sw.fsf@gitster.c.googlers.com","threadId":"53469","inReplyTo":"CAGjfG9akT+KG-tttRWEX_ZqxrqPoY_4Ed7Pymt4DkV5Rgc1CEA@mail.gmail.com","subject":"Re: Add a \"Flattened Cache\" to `git --clone`?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2020-05-14T22:30:39Z","receivedAt":"2020-05-14T22:30:44Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Caleb Gray <hey@calebgray.com> writes:\n\n> Actually those are the steps that I'm explicitly hoping can be\n> skipped, both on server and client, after the first successful clone\n> request transaction. The cache itself would be of the end resulting\n> `.git` directory (client side)... unless I have misconceptions about\n\nIf you look at .git/objects/pack/ directory in your repository that\nis a clone of somebody else, most likely you'd find even number of\nfiles in there, those whose filename ends with .pack and their\ncounterparts whose filename ends with .idx extension.  Both files\nmust exist to perform any local operation, but during the initial\ncloning, only the bits in the former are transferred, and the\ncontents of the latter must be constructed from the bits in the\nformer.\n\nYou can introduce a new protocol that copies the contents of the\n.idx, but the contents of that file MUST be validated on the\nreceiving end, which entails the same amount of computation as our\nclients currently spend to construct it out of .pack, so in the end,\nyou'd be wasting more bandwidth to transfer .idx which is redundant\ninformation without saving processing cycles.\n"},{"id":"397842","messageId":"CAGyf7-Fvyes2AwPhqAz=GhpDb=P664DLpK+4o3-ChKyR5KmJQw@mail.gmail.com","threadId":"53469","inReplyTo":"CAGjfG9akT+KG-tttRWEX_ZqxrqPoY_4Ed7Pymt4DkV5Rgc1CEA@mail.gmail.com","subject":"Re: Add a \"Flattened Cache\" to `git --clone`?","fromName":"Bryan Turner","fromEmail":"bturner@atlassian.com","sentAt":"2020-05-14T22:44:59Z","receivedAt":"2020-05-14T22:45:13Z","isPatch":false,"sender":{"key":"bturner@atlassian.com","avatar":"https://gravatar.com/avatar/16bcf3167981c1ef7c804e502642366d888a35b0d0b0a4ca01fdc442aa1acb1e?d=mp&s=160"},"body":"On Thu, May 14, 2020 at 3:05 PM Caleb Gray <hey@calebgray.com> wrote:\n>\n> Actually those are the steps that I'm explicitly hoping can be\n> skipped, both on server and client, after the first successful clone\n> request transaction. The cache itself would be of the end resulting\n> `.git` directory (client side)... unless I have misconceptions about\n> the complexity of reproducing what ends up on the client side from the\n> server side... I figured the shared library probably offers endpoints\n> for the information I'd need to achieve that.\n\nI don't know that such an approach would ever get accepted. At most it\ncould only be a partial replica of a `.get` directory. For example,\nincluding `.git/config` or `.git/hooks` carries some heavy security\nconsiderations that make it very unlikely such a change would get\naccepted. When you pare down the things from the `.git` directory that\ncan reasonably be included, I suspect you're pretty much going to be\nleft with `.git/objects/pack/pack-<something>.pack`, and perhaps\n`.git/packed-refs` (although the ref negotiations the client needs to\ndo in order to even request a pack means you're unlikely to actually\n_benefit_ from including `packed-refs`).\n\nFor such client-side caching, really you may be better of not trying\nto reinvent the wheel and instead, as others have suggested, simply\nuse `git clone --reference` (possibly plus `--dissociate` if you don't\nwant any long-term connection between clones) to allow `git clone` to\nreference all the objects you have available locally to skip most of\nthe pack transfer. If you do this, then `git clone` can _already_ make\nuse of local `idx` files, in addition to packs, to save work.\n\nBryan Turner\n"},{"id":"397942","messageId":"20200515214257.GA21855@dcvr","threadId":"53469","inReplyTo":"20200514214404.bcbjskgi52bwedlh@chatter.i7.local","subject":"Re: Add a \"Flattened Cache\" to `git --clone`?","fromName":"Eric Wong","fromEmail":"e@yhbt.net","sentAt":"2020-05-15T21:42:57Z","receivedAt":"2020-05-15T21:42:59Z","isPatch":false,"sender":{"key":"e@yhbt.net","avatar":null},"body":"Konstantin Ryabitsev <konstantin@linuxfoundation.org> wrote:\n> On Thu, May 14, 2020 at 02:23:44PM -0700, Junio C Hamano wrote:\n> > > I think something like git-caching-proxy would be a neat project, \n> > > because it would significantly improve mirroring for CI deployments \n> > > without requiring that each individual job implements clone.bundle \n> > > prefetching.\n> > \n> > What are we improving with such a proxy, though?\n> > \n> > Not bandwidth to the client, apparently. \n> \n> Well, if it sits in front of the CI subnet, then it *does* save \n> bandwidth.\n\nAgreed.\n\n> Here's an example with the exact situation we have:\n> \n> - the Gerrit server is on the US West Coast\n> - the CI builder is on the East Coast\n> - each CI job does a full transfer of the multi-MB repo across the \n>   continent, even when cloning shallow\n> \n> We solve this by having a local mirror of the repository, but this \n> requires active mirroring to be pre-setup. A caching proxy that could:\n> \n> - receive a request for a repository\n> - stream the response back to the client\n> - cache objects locally\n> - use local cache to construct future requests, so only missing objects \n>   are fetched from the remote repo regardless of the haves on the actual \n>   client...\n\nAn off-the-shelf HTTP caching proxy (e.g. polipo, Squid) could\ndo a good enough job with dumb HTTP clones (via GIT_SMART_HTTP=0\nenv).\n\nWith well-packed repos, the dumb HTTP transfer cost shouldn't be\ntoo high (and git 2.10+ got way faster on the client side with\npoorly-packed repos, thanks to the Linux kernel-derived list.h).\n\nThe occasional full repack on the source git server will\ninvalidate caches and result in a giant download; but it's\nbetter than no caching at all and doing giant cross-country\ntransfers all day long.\n\nThat said, I'm not sure if any client-side caching proxies can\nMITM HTTPS and save bandwidth with HTTPS everywhere, nowadays.\nI seem to recall polipo being abandoned because of HTTPS.\nMaybe there's a caching HTTPS MITM proxy out there...\n"},{"id":"398032","messageId":"20200517221213.l3q2creiddpylbpm@chatter.i7.local","threadId":"53469","inReplyTo":"20200515214257.GA21855@dcvr","subject":"Re: Add a \"Flattened Cache\" to `git --clone`?","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2020-05-17T22:12:13Z","receivedAt":"2020-05-17T22:12:19Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On Fri, May 15, 2020 at 09:42:57PM +0000, Eric Wong wrote:\n> That said, I'm not sure if any client-side caching proxies can\n> MITM HTTPS and save bandwidth with HTTPS everywhere, nowadays.\n> I seem to recall polipo being abandoned because of HTTPS.\n> Maybe there's a caching HTTPS MITM proxy out there...\n\nRight, this can't operate as a transparent proxy. However, it could work \nin combination with insteadOf on the client, e.g., if the repo URL is \nhttps://example.com/foo/bar.git, the CI builder could set a global \ninsteadOf in /etc/gitconfig before kicking off the job:\n\n[url http://local.proxy]\n  insteadOf = https://example.com\n\nThis way CI job maintainers could continue to use canonical repo URLs, \nbut actual requests would go out to the local proxy and be cached.\n\n-K\n"},{"id":"398488","messageId":"CAGjfG9YXQ5fkwY6DS8iOymHHXm-r+Rs4yah1afngwkGxNOE9gA@mail.gmail.com","threadId":"53469","inReplyTo":"1061511589863147@mail.yandex.ru","subject":"Re: Add a \"Flattened Cache\" to `git --clone`?","fromName":"Caleb Gray","fromEmail":"hey@calebgray.com","sentAt":"2020-05-25T14:02:21Z","receivedAt":"2020-05-25T14:02:41Z","isPatch":false,"sender":{"key":"hey@calebgray.com","avatar":null},"body":"For a repo like git itself, the assertions regarding the way git\ncurrently builds its data (in fact, including the `checkout` portion)\ndoes compete directly with the \"cached result\" methodology! Holy shit\nguys, I'm impressed as hell.\n\ntl;dr: The way I read the raw numbers, `git` ends up being as-fast-as\n(or faster) than a \"cache\" of the .git folder. Without doing further\nresearch, I'm inclined to agree with the previously mentioned bitmap\nmethod already being effectively as efficient as (more efficient\nthan!?) a cache.\n\n\nMethodology/Reasoning:\nvirtualized: verified zero network chatter on eth0 before and after each test.\ntcpflow: to gather the bits for the entire transaction... from just\nbefore the execution of `git clone` was started, and closing the\nlistener just after execution ended. (not worrying about\nprotocols/overhead)\ntar: to compare the size of the repository on disk with the tcpflow\nresults. (not worrying about compensating for\nheaders/metadata/overhead)\ngzip: to theoretically, I haven't checked anything, compensate for\nseemingly arbitrary size differences when downloading over HTTPS.\ntime: (really) rough measure of execution time.\n\n\nCommands used to generate files:\n*.tcpflow: `sudo tcpflow -p -c -i eth0 > $filename.tcpflow`\n*.tar: `tar cf $filename.tar .git`\n*.gz: `gzip -9 $filename.tar`\n\n\nResults:\n\n75M kernelorg.tar\n72M kernelorg.tar.gz\n69M kernelorg_git.tcpflow\n69M kernelorg_https.tcpflow\n\n145M github.tar\n143M github.tar.gz\n143M github_git.tcpflow\n142M github_https.tcpflow\n\n\nOther Tests (sanity checks):\n\nCloned a gitea mirror of kernel.org's git:\n69M gitea_git.tcpflow\n69M gitea_https.tcpflow\n\nCloned a bitbucket mirror of kernel.org's git:\n69M bitbucket_git.tcpflow\n69M bitbucket_https.tcpflow\n\n$ time git clone git://git.kernel.org/pub/scm/git/git.git\nCloning into 'git'...\nremote: Enumerating objects: 15475, done.\nremote: Counting objects: 100% (15475/15475), done.\nremote: Compressing objects: 100% (861/861), done.\nremote: Total 287977 (delta 14910), reused 14907 (delta 14610),\npack-reused 272502\nReceiving objects: 100% (287977/287977), 66.09 MiB | 4.87 MiB/s, done.\nResolving deltas: 100% (217420/217420), done.\n\nreal    0m20.000s\nuser    0m15.414s\nsys     0m1.606s\n\n$ time wget https://calebgray.com/public/kernelorg.tar.gz\n--2020-05-25 06:11:29--  https://calebgray.com/public/kernelorg.tar.gz\nResolving calebgray.com (calebgray.com)... 192.3.203.78\nConnecting to calebgray.com (calebgray.com)|192.3.203.78|:443... connected.\nHTTP request sent, awaiting response... 200 OK\nLength: 74593708 (71M) [application/octet-stream]\nSaving to: ‘kernelorg.tar.gz’\n\nkernelorg.tar.gz\n100%[========================================================================================>]\n 71.14M  4.81MB/s    in 19s\n\n2020-05-25 06:11:48 (3.79 MB/s) - ‘kernelorg.tar.gz’ saved [74593708/74593708]\n\nreal 0m19.420s\nuser 0m0.030s\nsys 0m0.280s\n\n\nThanks everyone for your input and time! I love git, you guys do great work!\n\nP.S. I ran a few other benchmarks outside of these, and the timing\nalways worked out to be more/less the same between the reported\ntransfer rate (as told by my router, as well) and the \"real\" time it\ntook to download (for both `git` and `wget`).\n\nP.P.S. I haven't investigated the reason for the github repo being\nnearly twice the size as the kernel.org hosted copy. That one stands\nout as potentially part of the proxy discussion, or there's actually a\ndifference in the repo's data. Curiosity will likely get the best of\nme eventually.\n\n\n\n\nOn Mon, May 18, 2020 at 9:40 PM Konstantin Tokarev <annulen@yandex.ru> wrote:\n>\n>\n>\n> 18.05.2020, 01:12, \"Konstantin Ryabitsev\" <konstantin@linuxfoundation.org>:\n> > On Fri, May 15, 2020 at 09:42:57PM +0000, Eric Wong wrote:\n> >>  That said, I'm not sure if any client-side caching proxies can\n> >>  MITM HTTPS and save bandwidth with HTTPS everywhere, nowadays.\n> >>  I seem to recall polipo being abandoned because of HTTPS.\n> >>  Maybe there's a caching HTTPS MITM proxy out there...\n> >\n> > Right, this can't operate as a transparent proxy.\n>\n> AFAIK, Squid can do MITM, caching and operate transparently.\n> In the past it was done via ssl_bump directive, but seems like syntax changed a bit\n> in modern versions.\n"}]}