{"thread":{"id":"53116","subject":"Inefficiency of partial shallow clone vs shallow clone + \"old-style\" sparse checkout","startedAt":"2020-03-27T21:08:22Z","lastAt":"2020-04-01T19:25:36Z","messageCount":13,"participants":["Konstantin Tokarev","Jeff King","Derrick Stolee","Taylor Blau"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"394197","messageId":"2814631585342072@sas8-da6d7485e0c7.qloud-c.yandex.net","threadId":"53116","inReplyTo":null,"subject":"Inefficiency of partial shallow clone vs shallow clone + \"old-style\" sparse checkout","fromName":"Konstantin Tokarev","fromEmail":"annulen@yandex.ru","sentAt":"2020-03-27T21:08:17Z","receivedAt":"2020-03-27T21:08:22Z","isPatch":false,"sender":{"key":"annulen@yandex.ru","avatar":null},"body":"Hello,\n\nIs it a known thing that addition of --filter=blob:none to workflow with shalow clone (e.g. --depth=1)\nand following sparse checkout may significantly slow down process and result in much larger\n.git repository?\n\nIn case anyone is interested, I've posted my measurements at [1].\n\nI understand this may have something to do with GitHub's server side implementation, but AFAIK\nthere are some GitHub folks here as well.\n\n[1] https://gist.github.com/annulen/835ac561e22bedd7138d13392a7a53be\n\n-- \nRegards,\nKonstantin\n\n"},{"id":"394229","messageId":"20200328144023.GB1198080@coredump.intra.peff.net","threadId":"53116","inReplyTo":"2814631585342072@sas8-da6d7485e0c7.qloud-c.yandex.net","subject":"Re: Inefficiency of partial shallow clone vs shallow clone + \"old-style\" sparse checkout","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-03-28T14:40:23Z","receivedAt":"2020-03-28T14:40:28Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Mar 28, 2020 at 12:08:17AM +0300, Konstantin Tokarev wrote:\n\n> Is it a known thing that addition of --filter=blob:none to workflow\n> with shalow clone (e.g. --depth=1) and following sparse checkout may\n> significantly slow down process and result in much larger .git\n> repository?\n> \n> In case anyone is interested, I've posted my measurements at [1].\n> \n> I understand this may have something to do with GitHub's server side\n> implementation, but AFAIK there are some GitHub folks here as well.\n\nI think the problem is on the client side. Just with a local git.git\nclone, try this:\n\n  $ git config uploadpack.allowfilter true\n  $ git clone --no-local --bare --depth=1 --filter=blob:none . both.git\n  Cloning into bare repository 'both.git'...\n  remote: Enumerating objects: 197, done.\n  remote: Counting objects: 100% (197/197), done.\n  remote: Compressing objects: 100% (153/153), done.\n  remote: Total 197 (delta 3), reused 171 (delta 3), pack-reused 0\n  Receiving objects: 100% (197/197), 113.63 KiB | 28.41 MiB/s, done.\n  Resolving deltas: 100% (3/3), done.\n  remote: Enumerating objects: 1871, done.\n  remote: Counting objects: 100% (1871/1871), done.\n  remote: Compressing objects: 100% (870/870), done.\n  remote: Total 1871 (delta 1001), reused 1855 (delta 994), pack-reused 0\n  Receiving objects: 100% (1871/1871), 384.93 KiB | 38.49 MiB/s, done.\n  Resolving deltas: 100% (1001/1001), done.\n  remote: Enumerating objects: 1878, done.\n  remote: Counting objects: 100% (1878/1878), done.\n  remote: Compressing objects: 100% (872/872), done.\n  remote: Total 1878 (delta 1004), reused 1864 (delta 999), pack-reused 0\n  Receiving objects: 100% (1878/1878), 386.41 KiB | 25.76 MiB/s, done.\n  Resolving deltas: 100% (1004/1004), done.\n  remote: Enumerating objects: 1903, done.\n  remote: Counting objects: 100% (1903/1903), done.\n  remote: Compressing objects: 100% (882/882), done.\n  remote: Total 1903 (delta 1020), reused 1890 (delta 1014), pack-reused 0\n  Receiving objects: 100% (1903/1903), 391.05 KiB | 16.29 MiB/s, done.\n  Resolving deltas: 100% (1020/1020), done.\n  remote: Enumerating objects: 1975, done.\n  remote: Counting objects: 100% (1975/1975), done.\n  remote: Compressing objects: 100% (915/915), done.\n  remote: Total 1975 (delta 1059), reused 1959 (delta 1052), pack-reused 0\n  Receiving objects: 100% (1975/1975), 405.58 KiB | 16.90 MiB/s, done.\n  Resolving deltas: 100% (1059/1059), done.\n  [...and so on...]\n\nOops. The backtrace for the clone during this process looks like:\n\n  [...]\n  #11 0x000055b980be01dc in fetch_objects (remote_name=0x55b981607620 \"origin\", oids=0x55b9816217a8, oid_nr=1)\n      at promisor-remote.c:47\n  #12 0x000055b980be0812 in promisor_remote_get_direct (repo=0x55b980dcab00 <the_repo>, oids=0x55b9816217a8, oid_nr=1)\n      at promisor-remote.c:247\n  #13 0x000055b980c3e475 in do_oid_object_info_extended (r=0x55b980dcab00 <the_repo>, oid=0x55b9816217a8, \n      oi=0x55b980dcaec0 <blank_oi>, flags=0) at sha1-file.c:1511\n  #14 0x000055b980c3e579 in oid_object_info_extended (r=0x55b980dcab00 <the_repo>, oid=0x55b9816217a8, oi=0x0, flags=0)\n      at sha1-file.c:1544\n  #15 0x000055b980c3f7bc in repo_has_object_file_with_flags (r=0x55b980dcab00 <the_repo>, oid=0x55b9816217a8, flags=0)\n      at sha1-file.c:1980\n  #16 0x000055b980c3f7ee in repo_has_object_file (r=0x55b980dcab00 <the_repo>, oid=0x55b9816217a8) at sha1-file.c:1986\n  #17 0x000055b980a54533 in write_followtags (refs=0x55b981610900, \n      msg=0x55b981601230 \"clone: from /home/peff/compile/git/.\") at builtin/clone.c:646\n  #18 0x000055b980a54723 in update_remote_refs (refs=0x55b981610900, mapped_refs=0x55b98160fe20, \n      remote_head_points_at=0x0, branch_top=0x55b981601130 \"refs/heads/\", \n      msg=0x55b981601230 \"clone: from /home/peff/compile/git/.\", transport=0x55b98160da90, check_connectivity=1, \n      check_refs_are_promisor_objects_only=1) at builtin/clone.c:699\n  #19 0x000055b980a5625b in cmd_clone (argc=2, argv=0x7fff5e0a1e70, prefix=0x0) at builtin/clone.c:1280\n  [...]\n\nSo I guess the problem is not with shallow clones specifically, but they\nlead us to not having fetched the commits pointed to by tags, which\nleads to us trying to fault in those commits (and their trees) rather\nthan realizing that we weren't meant to have them. And the size of the\nlocal repo balloons because you're fetching all those commits one by\none, and not getting the benefit of the deltas you would when you do a\nsingle --filter=blob:none fetch.\n\nI guess we need something like this:\n\ndiff --git a/builtin/clone.c b/builtin/clone.c\nindex 488bdb0741..a1879994f5 100644\n--- a/builtin/clone.c\n+++ b/builtin/clone.c\n@@ -643,7 +643,8 @@ static void write_followtags(const struct ref *refs, const char *msg)\n \t\t\tcontinue;\n \t\tif (ends_with(ref->name, \"^{}\"))\n \t\t\tcontinue;\n-\t\tif (!has_object_file(&ref->old_oid))\n+\t\tif (!has_object_file_with_flags(&ref->old_oid,\n+\t\t\t\t\t\tOBJECT_INFO_SKIP_FETCH_OBJECT))\n \t\t\tcontinue;\n \t\tupdate_ref(msg, ref->name, &ref->old_oid, NULL, 0,\n \t\t\t   UPDATE_REFS_DIE_ON_ERR);\n\nwhich seems to produce the desired result.\n\n-Peff\n"},{"id":"394240","messageId":"decf87bb-dffc-e44e-912e-fe51bc2514c3@gmail.com","threadId":"53116","inReplyTo":"20200328144023.GB1198080@coredump.intra.peff.net","subject":"Re: Inefficiency of partial shallow clone vs shallow clone + \"old-style\" sparse checkout","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-03-28T16:58:41Z","receivedAt":"2020-03-28T16:58:48Z","isPatch":false,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 3/28/2020 10:40 AM, Jeff King wrote:\n> On Sat, Mar 28, 2020 at 12:08:17AM +0300, Konstantin Tokarev wrote:\n> \n>> Is it a known thing that addition of --filter=blob:none to workflow\n>> with shalow clone (e.g. --depth=1) and following sparse checkout may\n>> significantly slow down process and result in much larger .git\n>> repository?\n\nIn general, I would recommend not using shallow clones in conjunction\nwith partial clone. The blob:none filter will get you what you really\nwant from shallow clone without any of the downsides of shallow clone.\n\nYou do point out a bug that happens when these features are combined,\nwhich is helpful. I'm just recommending that you do not combine these\nfeatures as you'll have a better experience (in my opinion).\n\n>> In case anyone is interested, I've posted my measurements at [1].\n>>\n>> I understand this may have something to do with GitHub's server side\n>> implementation, but AFAIK there are some GitHub folks here as well.\n> \n> I think the problem is on the client side. Just with a local git.git\n> clone, try this:\n> \n>   $ git config uploadpack.allowfilter true\n>   $ git clone --no-local --bare --depth=1 --filter=blob:none . both.git\n>   Cloning into bare repository 'both.git'...\n>   remote: Enumerating objects: 197, done.\n>   remote: Counting objects: 100% (197/197), done.\n>   remote: Compressing objects: 100% (153/153), done.\n>   remote: Total 197 (delta 3), reused 171 (delta 3), pack-reused 0\n>   Receiving objects: 100% (197/197), 113.63 KiB | 28.41 MiB/s, done.\n>   Resolving deltas: 100% (3/3), done.\n>   remote: Enumerating objects: 1871, done.\n>   remote: Counting objects: 100% (1871/1871), done.\n>   remote: Compressing objects: 100% (870/870), done.\n>   remote: Total 1871 (delta 1001), reused 1855 (delta 994), pack-reused 0\n>   Receiving objects: 100% (1871/1871), 384.93 KiB | 38.49 MiB/s, done.\n>   Resolving deltas: 100% (1001/1001), done.\n>   remote: Enumerating objects: 1878, done.\n>   remote: Counting objects: 100% (1878/1878), done.\n>   remote: Compressing objects: 100% (872/872), done.\n>   remote: Total 1878 (delta 1004), reused 1864 (delta 999), pack-reused 0\n>   Receiving objects: 100% (1878/1878), 386.41 KiB | 25.76 MiB/s, done.\n>   Resolving deltas: 100% (1004/1004), done.\n>   remote: Enumerating objects: 1903, done.\n>   remote: Counting objects: 100% (1903/1903), done.\n>   remote: Compressing objects: 100% (882/882), done.\n>   remote: Total 1903 (delta 1020), reused 1890 (delta 1014), pack-reused 0\n>   Receiving objects: 100% (1903/1903), 391.05 KiB | 16.29 MiB/s, done.\n>   Resolving deltas: 100% (1020/1020), done.\n>   remote: Enumerating objects: 1975, done.\n>   remote: Counting objects: 100% (1975/1975), done.\n>   remote: Compressing objects: 100% (915/915), done.\n>   remote: Total 1975 (delta 1059), reused 1959 (delta 1052), pack-reused 0\n>   Receiving objects: 100% (1975/1975), 405.58 KiB | 16.90 MiB/s, done.\n>   Resolving deltas: 100% (1059/1059), done.\n>   [...and so on...]\n> \n> Oops. The backtrace for the clone during this process looks like:\n> \n>   [...]\n>   #11 0x000055b980be01dc in fetch_objects (remote_name=0x55b981607620 \"origin\", oids=0x55b9816217a8, oid_nr=1)\n>       at promisor-remote.c:47\n>   #12 0x000055b980be0812 in promisor_remote_get_direct (repo=0x55b980dcab00 <the_repo>, oids=0x55b9816217a8, oid_nr=1)\n>       at promisor-remote.c:247\n>   #13 0x000055b980c3e475 in do_oid_object_info_extended (r=0x55b980dcab00 <the_repo>, oid=0x55b9816217a8, \n>       oi=0x55b980dcaec0 <blank_oi>, flags=0) at sha1-file.c:1511\n>   #14 0x000055b980c3e579 in oid_object_info_extended (r=0x55b980dcab00 <the_repo>, oid=0x55b9816217a8, oi=0x0, flags=0)\n>       at sha1-file.c:1544\n>   #15 0x000055b980c3f7bc in repo_has_object_file_with_flags (r=0x55b980dcab00 <the_repo>, oid=0x55b9816217a8, flags=0)\n>       at sha1-file.c:1980\n>   #16 0x000055b980c3f7ee in repo_has_object_file (r=0x55b980dcab00 <the_repo>, oid=0x55b9816217a8) at sha1-file.c:1986\n>   #17 0x000055b980a54533 in write_followtags (refs=0x55b981610900, \n>       msg=0x55b981601230 \"clone: from /home/peff/compile/git/.\") at builtin/clone.c:646\n>   #18 0x000055b980a54723 in update_remote_refs (refs=0x55b981610900, mapped_refs=0x55b98160fe20, \n>       remote_head_points_at=0x0, branch_top=0x55b981601130 \"refs/heads/\", \n>       msg=0x55b981601230 \"clone: from /home/peff/compile/git/.\", transport=0x55b98160da90, check_connectivity=1, \n>       check_refs_are_promisor_objects_only=1) at builtin/clone.c:699\n>   #19 0x000055b980a5625b in cmd_clone (argc=2, argv=0x7fff5e0a1e70, prefix=0x0) at builtin/clone.c:1280\n>   [...]\n> \n> So I guess the problem is not with shallow clones specifically, but they\n> lead us to not having fetched the commits pointed to by tags, which\n> leads to us trying to fault in those commits (and their trees) rather\n> than realizing that we weren't meant to have them. And the size of the\n> local repo balloons because you're fetching all those commits one by\n> one, and not getting the benefit of the deltas you would when you do a\n> single --filter=blob:none fetch.\n> \n> I guess we need something like this:\n> \n> diff --git a/builtin/clone.c b/builtin/clone.c\n> index 488bdb0741..a1879994f5 100644\n> --- a/builtin/clone.c\n> +++ b/builtin/clone.c\n> @@ -643,7 +643,8 @@ static void write_followtags(const struct ref *refs, const char *msg)\n>  \t\t\tcontinue;\n>  \t\tif (ends_with(ref->name, \"^{}\"))\n>  \t\t\tcontinue;\n> -\t\tif (!has_object_file(&ref->old_oid))\n> +\t\tif (!has_object_file_with_flags(&ref->old_oid,\n> +\t\t\t\t\t\tOBJECT_INFO_SKIP_FETCH_OBJECT))\n>  \t\t\tcontinue;\n>  \t\tupdate_ref(msg, ref->name, &ref->old_oid, NULL, 0,\n>  \t\t\t   UPDATE_REFS_DIE_ON_ERR);\n> \n> which seems to produce the desired result.\n\nThis is a good find, and I expect we will find more \"opportunities\"\nto insert OBJECT_INFO_SKIP_FETCH_OBJECT like this.\n\n-Stolee\n\n"},{"id":"394439","messageId":"20200331214653.GA95875@syl.local","threadId":"53116","inReplyTo":"decf87bb-dffc-e44e-912e-fe51bc2514c3@gmail.com","subject":"Re: Inefficiency of partial shallow clone vs shallow clone + \"old-style\" sparse checkout","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-03-31T21:46:53Z","receivedAt":"2020-03-31T21:47:00Z","isPatch":false,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Sat, Mar 28, 2020 at 12:58:41PM -0400, Derrick Stolee wrote:\n> On 3/28/2020 10:40 AM, Jeff King wrote:\n> > On Sat, Mar 28, 2020 at 12:08:17AM +0300, Konstantin Tokarev wrote:\n> >\n> >> Is it a known thing that addition of --filter=blob:none to workflow\n> >> with shalow clone (e.g. --depth=1) and following sparse checkout may\n> >> significantly slow down process and result in much larger .git\n> >> repository?\n>\n> In general, I would recommend not using shallow clones in conjunction\n> with partial clone. The blob:none filter will get you what you really\n> want from shallow clone without any of the downsides of shallow clone.\n>\n> You do point out a bug that happens when these features are combined,\n> which is helpful. I'm just recommending that you do not combine these\n> features as you'll have a better experience (in my opinion).\n>\n> >> In case anyone is interested, I've posted my measurements at [1].\n> >>\n> >> I understand this may have something to do with GitHub's server side\n> >> implementation, but AFAIK there are some GitHub folks here as well.\n> >\n> > I think the problem is on the client side. Just with a local git.git\n> > clone, try this:\n> >\n> >   $ git config uploadpack.allowfilter true\n> >   $ git clone --no-local --bare --depth=1 --filter=blob:none . both.git\n> >   Cloning into bare repository 'both.git'...\n> >   remote: Enumerating objects: 197, done.\n> >   remote: Counting objects: 100% (197/197), done.\n> >   remote: Compressing objects: 100% (153/153), done.\n> >   remote: Total 197 (delta 3), reused 171 (delta 3), pack-reused 0\n> >   Receiving objects: 100% (197/197), 113.63 KiB | 28.41 MiB/s, done.\n> >   Resolving deltas: 100% (3/3), done.\n> >   remote: Enumerating objects: 1871, done.\n> >   remote: Counting objects: 100% (1871/1871), done.\n> >   remote: Compressing objects: 100% (870/870), done.\n> >   remote: Total 1871 (delta 1001), reused 1855 (delta 994), pack-reused 0\n> >   Receiving objects: 100% (1871/1871), 384.93 KiB | 38.49 MiB/s, done.\n> >   Resolving deltas: 100% (1001/1001), done.\n> >   remote: Enumerating objects: 1878, done.\n> >   remote: Counting objects: 100% (1878/1878), done.\n> >   remote: Compressing objects: 100% (872/872), done.\n> >   remote: Total 1878 (delta 1004), reused 1864 (delta 999), pack-reused 0\n> >   Receiving objects: 100% (1878/1878), 386.41 KiB | 25.76 MiB/s, done.\n> >   Resolving deltas: 100% (1004/1004), done.\n> >   remote: Enumerating objects: 1903, done.\n> >   remote: Counting objects: 100% (1903/1903), done.\n> >   remote: Compressing objects: 100% (882/882), done.\n> >   remote: Total 1903 (delta 1020), reused 1890 (delta 1014), pack-reused 0\n> >   Receiving objects: 100% (1903/1903), 391.05 KiB | 16.29 MiB/s, done.\n> >   Resolving deltas: 100% (1020/1020), done.\n> >   remote: Enumerating objects: 1975, done.\n> >   remote: Counting objects: 100% (1975/1975), done.\n> >   remote: Compressing objects: 100% (915/915), done.\n> >   remote: Total 1975 (delta 1059), reused 1959 (delta 1052), pack-reused 0\n> >   Receiving objects: 100% (1975/1975), 405.58 KiB | 16.90 MiB/s, done.\n> >   Resolving deltas: 100% (1059/1059), done.\n> >   [...and so on...]\n> >\n> > Oops. The backtrace for the clone during this process looks like:\n> >\n> >   [...]\n> >   #11 0x000055b980be01dc in fetch_objects (remote_name=0x55b981607620 \"origin\", oids=0x55b9816217a8, oid_nr=1)\n> >       at promisor-remote.c:47\n> >   #12 0x000055b980be0812 in promisor_remote_get_direct (repo=0x55b980dcab00 <the_repo>, oids=0x55b9816217a8, oid_nr=1)\n> >       at promisor-remote.c:247\n> >   #13 0x000055b980c3e475 in do_oid_object_info_extended (r=0x55b980dcab00 <the_repo>, oid=0x55b9816217a8,\n> >       oi=0x55b980dcaec0 <blank_oi>, flags=0) at sha1-file.c:1511\n> >   #14 0x000055b980c3e579 in oid_object_info_extended (r=0x55b980dcab00 <the_repo>, oid=0x55b9816217a8, oi=0x0, flags=0)\n> >       at sha1-file.c:1544\n> >   #15 0x000055b980c3f7bc in repo_has_object_file_with_flags (r=0x55b980dcab00 <the_repo>, oid=0x55b9816217a8, flags=0)\n> >       at sha1-file.c:1980\n> >   #16 0x000055b980c3f7ee in repo_has_object_file (r=0x55b980dcab00 <the_repo>, oid=0x55b9816217a8) at sha1-file.c:1986\n> >   #17 0x000055b980a54533 in write_followtags (refs=0x55b981610900,\n> >       msg=0x55b981601230 \"clone: from /home/peff/compile/git/.\") at builtin/clone.c:646\n> >   #18 0x000055b980a54723 in update_remote_refs (refs=0x55b981610900, mapped_refs=0x55b98160fe20,\n> >       remote_head_points_at=0x0, branch_top=0x55b981601130 \"refs/heads/\",\n> >       msg=0x55b981601230 \"clone: from /home/peff/compile/git/.\", transport=0x55b98160da90, check_connectivity=1,\n> >       check_refs_are_promisor_objects_only=1) at builtin/clone.c:699\n> >   #19 0x000055b980a5625b in cmd_clone (argc=2, argv=0x7fff5e0a1e70, prefix=0x0) at builtin/clone.c:1280\n> >   [...]\n> >\n> > So I guess the problem is not with shallow clones specifically, but they\n> > lead us to not having fetched the commits pointed to by tags, which\n> > leads to us trying to fault in those commits (and their trees) rather\n> > than realizing that we weren't meant to have them. And the size of the\n> > local repo balloons because you're fetching all those commits one by\n> > one, and not getting the benefit of the deltas you would when you do a\n> > single --filter=blob:none fetch.\n> >\n> > I guess we need something like this:\n> >\n> > diff --git a/builtin/clone.c b/builtin/clone.c\n> > index 488bdb0741..a1879994f5 100644\n> > --- a/builtin/clone.c\n> > +++ b/builtin/clone.c\n> > @@ -643,7 +643,8 @@ static void write_followtags(const struct ref *refs, const char *msg)\n> >  \t\t\tcontinue;\n> >  \t\tif (ends_with(ref->name, \"^{}\"))\n> >  \t\t\tcontinue;\n> > -\t\tif (!has_object_file(&ref->old_oid))\n> > +\t\tif (!has_object_file_with_flags(&ref->old_oid,\n> > +\t\t\t\t\t\tOBJECT_INFO_SKIP_FETCH_OBJECT))\n> >  \t\t\tcontinue;\n> >  \t\tupdate_ref(msg, ref->name, &ref->old_oid, NULL, 0,\n> >  \t\t\t   UPDATE_REFS_DIE_ON_ERR);\n> >\n> > which seems to produce the desired result.\n>\n> This is a good find, and I expect we will find more \"opportunities\"\n> to insert OBJECT_INFO_SKIP_FETCH_OBJECT like this.\n\nShould we turn this into a proper patch and have it reviewed? It seems\nto be helping the situation, and after thinking about it (only briefly,\nbut more than not ;-)), this seems like the right direction.\n\nThere's an argument to be had about bundling a number of these up\ninstead of having a slow drip of patches that sprinkle\n'SKIP_FETCH_OBJECT' everywhere, but I don't think that we want perfect\nto be the enemy of good here.\n\nPeff, if you don't feel like doing this, or have a backlog that is too\nlong, I'd be happy to polish this for you.\n\n> -Stolee\n\nThanks,\nTaylor\n"},{"id":"394442","messageId":"8919571585692069@vla5-e043431e7e8d.qloud-c.yandex.net","threadId":"53116","inReplyTo":"decf87bb-dffc-e44e-912e-fe51bc2514c3@gmail.com","subject":"Re: Inefficiency of partial shallow clone vs shallow clone + \"old-style\" sparse checkout","fromName":"Konstantin Tokarev","fromEmail":"annulen@yandex.ru","sentAt":"2020-03-31T22:10:41Z","receivedAt":"2020-03-31T22:10:49Z","isPatch":false,"sender":{"key":"annulen@yandex.ru","avatar":null},"body":"\n\n28.03.2020, 19:58, \"Derrick Stolee\" <stolee@gmail.com>:\n> On 3/28/2020 10:40 AM, Jeff King wrote:\n>>  On Sat, Mar 28, 2020 at 12:08:17AM +0300, Konstantin Tokarev wrote:\n>>\n>>>  Is it a known thing that addition of --filter=blob:none to workflow\n>>>  with shalow clone (e.g. --depth=1) and following sparse checkout may\n>>>  significantly slow down process and result in much larger .git\n>>>  repository?\n>\n> In general, I would recommend not using shallow clones in conjunction\n> with partial clone. The blob:none filter will get you what you really\n> want from shallow clone without any of the downsides of shallow clone.\n\nIs it really so?\n\nAs you can see from my measurements [1], in my case simple shallow clone (1)\nruns faster than simple partial clone (2) and produces slightly smaller .git,\nfrom which I can infer that (2) downloads some data which is not downloaded\nin (1).\n\nTo be clear, use case which I'm interested right now is checking out sources in\ncloud CI system like GitHub Actions for one shot build. Right now checkout usually\ntakes 1-2 minutes and my hope was that someday in the future it would be possible\\\nto make it faster.\n\n[1] https://gist.github.com/annulen/835ac561e22bedd7138d13392a7a53be\n\n-- \nRegards,\nKonstantin\n\n"},{"id":"394443","messageId":"4872731585693023@vla5-c5051da8689e.qloud-c.yandex.net","threadId":"53116","inReplyTo":"8919571585692069@vla5-e043431e7e8d.qloud-c.yandex.net","subject":"Re: Inefficiency of partial shallow clone vs shallow clone + \"old-style\" sparse checkout","fromName":"Konstantin Tokarev","fromEmail":"annulen@yandex.ru","sentAt":"2020-03-31T22:23:28Z","receivedAt":"2020-03-31T22:23:33Z","isPatch":false,"sender":{"key":"annulen@yandex.ru","avatar":null},"body":"\n\n01.04.2020, 01:10, \"Konstantin Tokarev\" <annulen@yandex.ru>:\n> 28.03.2020, 19:58, \"Derrick Stolee\" <stolee@gmail.com>:\n>>  On 3/28/2020 10:40 AM, Jeff King wrote:\n>>>   On Sat, Mar 28, 2020 at 12:08:17AM +0300, Konstantin Tokarev wrote:\n>>>\n>>>>   Is it a known thing that addition of --filter=blob:none to workflow\n>>>>   with shalow clone (e.g. --depth=1) and following sparse checkout may\n>>>>   significantly slow down process and result in much larger .git\n>>>>   repository?\n>>\n>>  In general, I would recommend not using shallow clones in conjunction\n>>  with partial clone. The blob:none filter will get you what you really\n>>  want from shallow clone without any of the downsides of shallow clone.\n>\n> Is it really so?\n>\n> As you can see from my measurements [1], in my case simple shallow clone (1)\n> runs faster than simple partial clone (2) and produces slightly smaller .git,\n> from which I can infer that (2) downloads some data which is not downloaded\n> in (1).\n\nActually, as I have full git logs for all these cases, there is no need to be guessing:\n    (1) downloads 295085 git objects of total size 1.00 GiB\n    (2) downloads 1949129 git objects of total size 1.01 GiB\nTotal sizes are very close, but (2) downloads much more objects, and also it uses\n3 passes to download them which leads to less efficient use of network bandwidth.\n\n>\n> To be clear, use case which I'm interested right now is checking out sources in\n> cloud CI system like GitHub Actions for one shot build. Right now checkout usually\n> takes 1-2 minutes and my hope was that someday in the future it would be possible\\\n> to make it faster.\n>\n> [1] https://gist.github.com/annulen/835ac561e22bedd7138d13392a7a53be\n>\n> --\n> Regards,\n> Konstantin\n\n-- \nRegards,\nKonstantin\n\n"},{"id":"394449","messageId":"0bf763ad-5f1d-65e2-bf3a-a4b7d5a7b3e3@gmail.com","threadId":"53116","inReplyTo":"4872731585693023@vla5-c5051da8689e.qloud-c.yandex.net","subject":"Re: Inefficiency of partial shallow clone vs shallow clone + \"old-style\" sparse checkout","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-04-01T00:09:26Z","receivedAt":"2020-04-01T00:09:32Z","isPatch":false,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 3/31/2020 6:23 PM, Konstantin Tokarev wrote:\n> 01.04.2020, 01:10, \"Konstantin Tokarev\" <annulen@yandex.ru>:\n>> 28.03.2020, 19:58, \"Derrick Stolee\" <stolee@gmail.com>:\n>>>  On 3/28/2020 10:40 AM, Jeff King wrote:\n>>>>   On Sat, Mar 28, 2020 at 12:08:17AM +0300, Konstantin Tokarev wrote:\n>>>>\n>>>>>   Is it a known thing that addition of --filter=blob:none to workflow\n>>>>>   with shalow clone (e.g. --depth=1) and following sparse checkout may\n>>>>>   significantly slow down process and result in much larger .git\n>>>>>   repository?\n>>>\n>>>  In general, I would recommend not using shallow clones in conjunction\n>>>  with partial clone. The blob:none filter will get you what you really\n>>>  want from shallow clone without any of the downsides of shallow clone.\n>>\n>> Is it really so?\n>>\n>> As you can see from my measurements [1], in my case simple shallow clone (1)\n>> runs faster than simple partial clone (2) and produces slightly smaller .git,\n>> from which I can infer that (2) downloads some data which is not downloaded\n>> in (1).\n> \n> Actually, as I have full git logs for all these cases, there is no need to be guessing:\n>     (1) downloads 295085 git objects of total size 1.00 GiB\n>     (2) downloads 1949129 git objects of total size 1.01 GiB\n\nIt is worth pointing out that these sizes are very close. The number of objects\nmay be part of why the timing is so different as the client needs to parse all\ndeltas to verify the object contents.\n\nRe-running the test with GIT_TRACE2_PERF=1 might reveal some interesting info\nabout which regions are slower than others.\n\n> Total sizes are very close, but (2) downloads much more objects, and also it uses\n> 3 passes to download them which leads to less efficient use of network bandwidth.\n\nThree passes, being:\n\n1. Download commits and trees.\n2. Initialize sparse-checkout with blobs at root.\n3. Expand sparse-checkout.\n\nIs that right? You could group 1 & 2 by setting your sparse-checkout patterns\nbefore initializing a checkout (if you clone with --no-checkout). Your link\nsays you did this:\n\n\tgit clone <mode> --no-checkout <url> <dir>\n\tgit sparse-checkout init\n\tgit sparse-checkout set '/*' '!LayoutTests'\n\nTry doing it this way instead:\n\n\tgit clone <mode> --no-checkout <url> <dir>\n\tgit config core.sparseCheckout true\n\tgit sparse-checkout set '/*' '!LayoutTests'\n\nBy doing it this way, you skip the step where the 'init' subcommand looks\nfor all blobs at root and does a network call for them. Should remove some\noverhead.\n\nLess efficient use of network bandwidth is one thing, but shallow clones are\nalso more CPU-intensive with the \"counting objects\" phase on the server. Your\nlink shares the following end-to-end timings:\n\n* Shallow-clone: 234s\n* Partial clone: 286s\n* Both(???): 1023s\n\nThe data implies that by asking for both you actually got a full clone (4.1 GB).\n\nThe 234s to 286s difference is meaningful. Almost a minute.\n\n>> To be clear, use case which I'm interested right now is checking out sources in\n>> cloud CI system like GitHub Actions for one shot build. Right now checkout usually\n>> takes 1-2 minutes and my hope was that someday in the future it would be possible\\\n>> to make it faster.\n\nAs long as you delete the shallow clone every time, then you also remove the\ndownsides of a shallow clone related to a later fetch or attempts to push.\n\nIf possible, a repo this size would benefit from persistent build agents that\nyou control. They can keep a copy of the repo around and do incremental fetches\nthat are much faster. It's a larger investment to run your own build lab, though.\nBut sometimes making builds faster is expensive. It depends on how \"expensive\" those\nfour minute clones per build are in terms of your team waiting.\n\nThanks,\n-Stolee\n"},{"id":"394450","messageId":"8268671585700012@iva3-58091f505f14.qloud-c.yandex.net","threadId":"53116","inReplyTo":"0bf763ad-5f1d-65e2-bf3a-a4b7d5a7b3e3@gmail.com","subject":"Re: Inefficiency of partial shallow clone vs shallow clone + \"old-style\" sparse checkout","fromName":"Konstantin Tokarev","fromEmail":"annulen@yandex.ru","sentAt":"2020-04-01T01:49:20Z","receivedAt":"2020-04-01T01:49:29Z","isPatch":false,"sender":{"key":"annulen@yandex.ru","avatar":null},"body":"\n\n01.04.2020, 03:09, \"Derrick Stolee\" <stolee@gmail.com>:\n> On 3/31/2020 6:23 PM, Konstantin Tokarev wrote:\n>>  01.04.2020, 01:10, \"Konstantin Tokarev\" <annulen@yandex.ru>:\n>>>  28.03.2020, 19:58, \"Derrick Stolee\" <stolee@gmail.com>:\n>>>>   On 3/28/2020 10:40 AM, Jeff King wrote:\n>>>>>    On Sat, Mar 28, 2020 at 12:08:17AM +0300, Konstantin Tokarev wrote:\n>>>>>\n>>>>>>    Is it a known thing that addition of --filter=blob:none to workflow\n>>>>>>    with shalow clone (e.g. --depth=1) and following sparse checkout may\n>>>>>>    significantly slow down process and result in much larger .git\n>>>>>>    repository?\n>>>>\n>>>>   In general, I would recommend not using shallow clones in conjunction\n>>>>   with partial clone. The blob:none filter will get you what you really\n>>>>   want from shallow clone without any of the downsides of shallow clone.\n>>>\n>>>  Is it really so?\n>>>\n>>>  As you can see from my measurements [1], in my case simple shallow clone (1)\n>>>  runs faster than simple partial clone (2) and produces slightly smaller .git,\n>>>  from which I can infer that (2) downloads some data which is not downloaded\n>>>  in (1).\n>>\n>>  Actually, as I have full git logs for all these cases, there is no need to be guessing:\n>>      (1) downloads 295085 git objects of total size 1.00 GiB\n>>      (2) downloads 1949129 git objects of total size 1.01 GiB\n>\n> It is worth pointing out that these sizes are very close. The number of objects\n> may be part of why the timing is so different as the client needs to parse all\n> deltas to verify the object contents.\n>\n> Re-running the test with GIT_TRACE2_PERF=1 might reveal some interesting info\n> about which regions are slower than others.\n\nHere are trace results for (1) with fix discussed below:\nhttps://gist.github.com/annulen/58b868e35e992105e7028946a8370795\n\nHere are trace results for (2) with fix discussed below:\nhttps://gist.github.com/annulen/fa1ef1b5d1056e6dede815e9ebf85c03\n\n>\n>>  Total sizes are very close, but (2) downloads much more objects, and also it uses\n>>  3 passes to download them which leads to less efficient use of network bandwidth.\n>\n> Three passes, being:\n>\n> 1. Download commits and trees.\n> 2. Initialize sparse-checkout with blobs at root.\n> 3. Expand sparse-checkout.\n>\n> Is that right? You could group 1 & 2 by setting your sparse-checkout patterns\n> before initializing a checkout (if you clone with --no-checkout). Your link\n> says you did this:\n>\n>         git clone <mode> --no-checkout <url> <dir>\n>         git sparse-checkout init\n>         git sparse-checkout set '/*' '!LayoutTests'\n>\n> Try doing it this way instead:\n>\n>         git clone <mode> --no-checkout <url> <dir>\n>         git config core.sparseCheckout true\n>         git sparse-checkout set '/*' '!LayoutTests'\n>\n> By doing it this way, you skip the step where the 'init' subcommand looks\n> for all blobs at root and does a network call for them. Should remove some\n> overhead.\n\nThanks, that helped. Now git downloads object only two times.\n\nFrom reading man page I assumed that `git sparse-checkout init` should do\nthe same as `git config core.sparseCheckout true`, unless `--cone` argument\nis specified.\n\n>\n> Less efficient use of network bandwidth is one thing, but shallow clones are\n> also more CPU-intensive with the \"counting objects\" phase on the server. Your\n> link shares the following end-to-end timings:\n>\n> * Shallow-clone: 234s\n> * Partial clone: 286s\n> * Both(???): 1023s\n>\n> The data implies that by asking for both you actually got a full clone (4.1 GB).\n\nNo, this is still a partial clone, full clone takes more than 6 GB\n\n>\n> The 234s to 286s difference is meaningful. Almost a minute.\n>\n>>>  To be clear, use case which I'm interested right now is checking out sources in\n>>>  cloud CI system like GitHub Actions for one shot build. Right now checkout usually\n>>>  takes 1-2 minutes and my hope was that someday in the future it would be possible\\\n>>>  to make it faster.\n>\n> As long as you delete the shallow clone every time, then you also remove the\n> downsides of a shallow clone related to a later fetch or attempts to push.\n>\n> If possible, a repo this size would benefit from persistent build agents that\n> you control. They can keep a copy of the repo around and do incremental fetches\n> that are much faster. It's a larger investment to run your own build lab, though.\n> But sometimes making builds faster is expensive. It depends on how \"expensive\" those\n> four minute clones per build are in terms of your team waiting.\n\nNo, current checkout times for shallow clone + sparseCheckout are quite acceptable.\n(FWIW, initially I used shallow clone without sparseCheckout, as the latter is not supported\nby GitHub Actions out of the box, and those times were NOT acceptable, as depending on\nserver load checkout could take 16 minutes or even more)\n\nFor now it's just my curioisty and desire to provide info which could make Git better.\nIt just seemed logical to me initially that if we limit both required paths in worktree and history\ndepth on object download stage, it should be more efficient than limiting only  history depth\n(or, at least, have the same efficiency).\n\nBTW, I did more measurements and results seem to be highly dependent on server side.\nOnce partial clone (2) even worked faster than shallow clone + sparseCheckout (1). Still, most of\nthe time (1) is faster than (2).\n\n-- \nRegards,\nKonstantin\n\n"},{"id":"394488","messageId":"20200401114446.GA1589184@coredump.intra.peff.net","threadId":"53116","inReplyTo":"8268671585700012@iva3-58091f505f14.qloud-c.yandex.net","subject":"Re: Inefficiency of partial shallow clone vs shallow clone + \"old-style\" sparse checkout","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-04-01T11:44:46Z","receivedAt":"2020-04-01T11:44:48Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Apr 01, 2020 at 04:49:20AM +0300, Konstantin Tokarev wrote:\n\n> > Less efficient use of network bandwidth is one thing, but shallow clones are\n> > also more CPU-intensive with the \"counting objects\" phase on the server. Your\n> > link shares the following end-to-end timings:\n> >\n> > * Shallow-clone: 234s\n> > * Partial clone: 286s\n> > * Both(???): 1023s\n> >\n> > The data implies that by asking for both you actually got a full clone (4.1 GB).\n> \n> No, this is still a partial clone, full clone takes more than 6 GB\n\nI think that 4GB number is just because of the bug, though. With the fix\nI showed earlier, doing clones of linux.git from a local repo yields:\n\n  type       objects (in passes)      bytes  time\n  ----       -----------------------  -----  ----\n  shallow      71447 (  71447+  n/a)  188MB   23s\n  blob:none  5260567 (5193557+67010)  870MB   99s\n  both         71447 (   4437+67010)  188MB   37s\n\nThe object counts and sizes make sense. blob:none is still going to get\nthe whole history of commits and trees, which are substantial. The sizes\nfor \"shallow\" and \"both\" are the same, because the checkout is going to\ngrab all of the blobs from the tip commit, which were included in the\noriginal \"shallow\" anyway. It does take longer, because they come in a\nsecond followup fetch (though I'm surprised it's so _much_ slower).\n\nSo to me that implies that shallow is strictly better than partial if\nyou're just going to check out the full tip commit. But doing both\ntogether opens up the possibility of narrowing the sparse checkout.\nDoing:\n\n  $ git clone --no-local --no-checkout --filter=blob:none --depth=1 \\\n      /path/to/linux sparse\n  $ cd sparse\n  $ git sparse-checkout set arch\n\nfetches 20795 objects (4437+16357+1), consuming only 27MB.\n\n-Peff\n"},{"id":"394489","messageId":"20200401121537.GA1916590@coredump.intra.peff.net","threadId":"53116","inReplyTo":"20200328144023.GB1198080@coredump.intra.peff.net","subject":"[PATCH] clone: use \"quick\" lookup while following tags","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-04-01T12:15:37Z","receivedAt":"2020-04-01T12:15:46Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Mar 28, 2020 at 10:40:23AM -0400, Jeff King wrote:\n\n> So I guess the problem is not with shallow clones specifically, but they\n> lead us to not having fetched the commits pointed to by tags, which\n> leads to us trying to fault in those commits (and their trees) rather\n> than realizing that we weren't meant to have them. And the size of the\n> local repo balloons because you're fetching all those commits one by\n> one, and not getting the benefit of the deltas you would when you do a\n> single --filter=blob:none fetch.\n> \n> I guess we need something like this:\n\nThe issue is actually with --single-branch, which is implied by --depth.\nBut the fix is the same either way.\n\nHere it is with a commit message and test.\n\n-- >8 --\nSubject: [PATCH] clone: use \"quick\" lookup while following tags\n\nWhen cloning with --single-branch, we implement git-fetch's usual\ntag-following behavior, grabbing any tag objects that point to objects\nwe have locally.\n\nWhen we're a partial clone, though, our has_object_file() check will\nactually lazy-fetch each tag. That not only defeats the purpose of\n--single-branch, but it does it incredibly slowly, potentially kicking\noff a new fetch for each tag. This is even worse for a shallow clone,\nwhich implies --single-branch, because even tags which are supersets of\neach other will be fetched individually.\n\nWe can fix this by passing OBJECT_INFO_SKIP_FETCH_OBJECT to the call,\nwhich is what git-fetch does in this case.\n\nLikewise, let's include OBJECT_INFO_QUICK, as that's what git-fetch\ndoes. The rationale is discussed in 5827a03545 (fetch: use \"quick\"\nhas_sha1_file for tag following, 2016-10-13), but here the tradeoff\nwould apply even more so because clone is very unlikely to be racing\nwith another process repacking our newly-created repository.\n\nThis may provide a very small speedup even in the non-partial case case,\nas we'd avoid calling reprepare_packed_git() for each tag (though in\npractice, we'd only have a single packfile, so that reprepare should be\nquite cheap).\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n builtin/clone.c          | 4 +++-\n t/t5616-partial-clone.sh | 8 ++++++++\n 2 files changed, 11 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/clone.c b/builtin/clone.c\nindex d8b1f413aa..9da6459f1d 100644\n--- a/builtin/clone.c\n+++ b/builtin/clone.c\n@@ -643,7 +643,9 @@ static void write_followtags(const struct ref *refs, const char *msg)\n \t\t\tcontinue;\n \t\tif (ends_with(ref->name, \"^{}\"))\n \t\t\tcontinue;\n-\t\tif (!has_object_file(&ref->old_oid))\n+\t\tif (!has_object_file_with_flags(&ref->old_oid,\n+\t\t\t\t\t\tOBJECT_INFO_QUICK |\n+\t\t\t\t\t\tOBJECT_INFO_SKIP_FETCH_OBJECT))\n \t\t\tcontinue;\n \t\tupdate_ref(msg, ref->name, &ref->old_oid, NULL, 0,\n \t\t\t   UPDATE_REFS_DIE_ON_ERR);\ndiff --git a/t/t5616-partial-clone.sh b/t/t5616-partial-clone.sh\nindex 77bb91e976..8f0d81a27e 100755\n--- a/t/t5616-partial-clone.sh\n+++ b/t/t5616-partial-clone.sh\n@@ -415,6 +415,14 @@ test_expect_success 'verify fetch downloads only one pack when updating refs' '\n \ttest_line_count = 3 pack-list\n '\n \n+test_expect_success 'single-branch tag following respects partial clone' '\n+\tgit clone --single-branch -b B --filter=blob:none \\\n+\t\t\"file://$(pwd)/srv.bare\" single &&\n+\tgit -C single rev-parse --verify refs/tags/B &&\n+\tgit -C single rev-parse --verify refs/tags/A &&\n+\ttest_must_fail git -C single rev-parse --verify refs/tags/C\n+'\n+\n . \"$TEST_DIRECTORY\"/lib-httpd.sh\n start_httpd\n \n-- \n2.26.0.408.gebd8a4413c\n\n"},{"id":"394490","messageId":"20200401121846.GB1916590@coredump.intra.peff.net","threadId":"53116","inReplyTo":"20200331214653.GA95875@syl.local","subject":"Re: Inefficiency of partial shallow clone vs shallow clone + \"old-style\" sparse checkout","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-04-01T12:18:46Z","receivedAt":"2020-04-01T12:18:49Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Mar 31, 2020 at 03:46:53PM -0600, Taylor Blau wrote:\n\n> On Sat, Mar 28, 2020 at 12:58:41PM -0400, Derrick Stolee wrote:\n>\n> > This is a good find, and I expect we will find more \"opportunities\"\n> > to insert OBJECT_INFO_SKIP_FETCH_OBJECT like this.\n\nThe worst part is that we _did_ find this in git-fetch, and fixed it in\n3e96c66805 (partial-clone: avoid fetching when looking for objects,\n2020-02-21). But as usual git-clone has its own kind-of-the-same\nimplementation of the same feature.\n\n> Should we turn this into a proper patch and have it reviewed? It seems\n> to be helping the situation, and after thinking about it (only briefly,\n> but more than not ;-)), this seems like the right direction.\n\nYes, I wanted to simplify the test and double-check the addition of\nQUICK. See the patch I just posted in:\n\n  https://lore.kernel.org/git/20200401121537.GA1916590@coredump.intra.peff.net/\n\n-Peff\n"},{"id":"394510","messageId":"111951585768302@iva8-bad53723c646.qloud-c.yandex.net","threadId":"53116","inReplyTo":"20200401121537.GA1916590@coredump.intra.peff.net","subject":"Re: [PATCH] clone: use \"quick\" lookup while following tags","fromName":"Konstantin Tokarev","fromEmail":"annulen@yandex.ru","sentAt":"2020-04-01T19:12:30Z","receivedAt":"2020-04-01T19:12:37Z","isPatch":true,"sender":{"key":"annulen@yandex.ru","avatar":null},"body":"\n\n01.04.2020, 15:15, \"Jeff King\" <peff@peff.net>:\n> On Sat, Mar 28, 2020 at 10:40:23AM -0400, Jeff King wrote:\n>\n>>  So I guess the problem is not with shallow clones specifically, but they\n>>  lead us to not having fetched the commits pointed to by tags, which\n>>  leads to us trying to fault in those commits (and their trees) rather\n>>  than realizing that we weren't meant to have them. And the size of the\n>>  local repo balloons because you're fetching all those commits one by\n>>  one, and not getting the benefit of the deltas you would when you do a\n>>  single --filter=blob:none fetch.\n>>\n>>  I guess we need something like this:\n>\n> The issue is actually with --single-branch, which is implied by --depth.\n> But the fix is the same either way.\n>\n> Here it is with a commit message and test.\n>\n> -- >8 --\n> Subject: [PATCH] clone: use \"quick\" lookup while following tags\n>\n> When cloning with --single-branch, we implement git-fetch's usual\n> tag-following behavior, grabbing any tag objects that point to objects\n> we have locally.\n>\n> When we're a partial clone, though, our has_object_file() check will\n> actually lazy-fetch each tag. That not only defeats the purpose of\n> --single-branch, but it does it incredibly slowly, potentially kicking\n> off a new fetch for each tag. This is even worse for a shallow clone,\n> which implies --single-branch, because even tags which are supersets of\n> each other will be fetched individually.\n>\n> We can fix this by passing OBJECT_INFO_SKIP_FETCH_OBJECT to the call,\n> which is what git-fetch does in this case.\n>\n> Likewise, let's include OBJECT_INFO_QUICK, as that's what git-fetch\n> does. The rationale is discussed in 5827a03545 (fetch: use \"quick\"\n> has_sha1_file for tag following, 2016-10-13), but here the tradeoff\n> would apply even more so because clone is very unlikely to be racing\n> with another process repacking our newly-created repository.\n>\n> This may provide a very small speedup even in the non-partial case case,\n> as we'd avoid calling reprepare_packed_git() for each tag (though in\n> practice, we'd only have a single packfile, so that reprepare should be\n> quite cheap).\n>\n> Signed-off-by: Jeff King <peff@peff.net>\n> ---\n>  builtin/clone.c | 4 +++-\n>  t/t5616-partial-clone.sh | 8 ++++++++\n>  2 files changed, 11 insertions(+), 1 deletion(-)\n>\n> diff --git a/builtin/clone.c b/builtin/clone.c\n> index d8b1f413aa..9da6459f1d 100644\n> --- a/builtin/clone.c\n> +++ b/builtin/clone.c\n> @@ -643,7 +643,9 @@ static void write_followtags(const struct ref *refs, const char *msg)\n>                          continue;\n>                  if (ends_with(ref->name, \"^{}\"))\n>                          continue;\n> - if (!has_object_file(&ref->old_oid))\n> + if (!has_object_file_with_flags(&ref->old_oid,\n> + OBJECT_INFO_QUICK |\n> + OBJECT_INFO_SKIP_FETCH_OBJECT))\n>                          continue;\n>                  update_ref(msg, ref->name, &ref->old_oid, NULL, 0,\n>                             UPDATE_REFS_DIE_ON_ERR);\n> diff --git a/t/t5616-partial-clone.sh b/t/t5616-partial-clone.sh\n> index 77bb91e976..8f0d81a27e 100755\n> --- a/t/t5616-partial-clone.sh\n> +++ b/t/t5616-partial-clone.sh\n> @@ -415,6 +415,14 @@ test_expect_success 'verify fetch downloads only one pack when updating refs' '\n>          test_line_count = 3 pack-list\n>  '\n>\n> +test_expect_success 'single-branch tag following respects partial clone' '\n> + git clone --single-branch -b B --filter=blob:none \\\n> + \"file://$(pwd)/srv.bare\" single &&\n> + git -C single rev-parse --verify refs/tags/B &&\n> + git -C single rev-parse --verify refs/tags/A &&\n> + test_must_fail git -C single rev-parse --verify refs/tags/C\n> +'\n> +\n>  . \"$TEST_DIRECTORY\"/lib-httpd.sh\n>  start_httpd\n>\n> --\n> 2.26.0.408.gebd8a4413c\n\nWould you recommend using this patch in production, or is it better to wait for reviews?\n\n-- \nRegards,\nKonstantin\n\n"},{"id":"394511","messageId":"20200401192533.GA3057718@coredump.intra.peff.net","threadId":"53116","inReplyTo":"111951585768302@iva8-bad53723c646.qloud-c.yandex.net","subject":"Re: [PATCH] clone: use \"quick\" lookup while following tags","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-04-01T19:25:33Z","receivedAt":"2020-04-01T19:25:36Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Apr 01, 2020 at 10:12:30PM +0300, Konstantin Tokarev wrote:\n\n> Would you recommend using this patch in production, or is it better to wait for reviews?\n\nOnly you know the risk tolerance of your production systems, so I make\nno guarantees.  But I would personally think it is quite safe.\n\n-Peff\n"}]}