{"thread":{"id":"54470","subject":"Questions about partial clone with '--filter=tree:0'","startedAt":"2020-10-20T17:17:38Z","lastAt":"2020-10-26T20:08:59Z","messageCount":9,"participants":["Alexandr Miloslavskiy","Taylor Blau","Jonathan Tan"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"407992","messageId":"aa7b89ee-08aa-7943-6a00-28dcf344426e@syntevo.com","threadId":"54470","inReplyTo":null,"subject":"Questions about partial clone with '--filter=tree:0'","fromName":"Alexandr Miloslavskiy","fromEmail":"alexandr.miloslavskiy@syntevo.com","sentAt":"2020-10-20T17:09:36Z","receivedAt":"2020-10-20T17:17:38Z","isPatch":false,"sender":{"key":"alexandr.miloslavskiy@syntevo.com","avatar":null},"body":"This is a edited copy of message I sent 2 weeks ago, which unfortunately\ndidn't receive any replies. I tried to make make it shorter this time :)\n\n----\n\nWe are implementing a git UI. One interesting case is the repository\ncloned with '--filter=tree:0', because it makes it a lot harder to\nrun basic git operations such as file log and blame.\n\nThe problems and potential solutions are outlined below. We should be\nable to make patches for (2) and (3) if it makes sense to patch these.\n\n(1) Is it even considered a realistic use case?\n-----------------------------------------------\nSummary: is '--filter=tree:0' a realistic or \"crazy\" scenario that is\nnot considered worthy of supporting?\n\nI decided to use Linux repo, which is reasonably large, and it seems\nthat '--filter=tree:0' could be desired because it helps with disk\nspace (~0.66gb) and network (~0.54gb):\n\nhttps://github.com/torvalds/linux.git\n   951025 commits total.\n\n   git clone --bare <url>\n\t7'624'042 objects\n\t   2.86gb network\n\t   3.10gb disk\n   git clone --bare --filter=blob:none <url>\n\t5'484'714 (71.9%) objects\n\t   1.01gb (35.3%) network\n\t   1.16gb (37.4%) disk\n   git clone --bare --filter=tree:0 <url>\n\t  951'693 (12.5%) objects\n\t   0.47gb (16.4%) network\n\t   0.50gb (16.1%) disk\n   git clone --bare --depth 1 --branch master <url>\n\t   74'380 ( 0.9%) objects\n\t   0.19gb ( 6.6%) network\n\t   0.19gb ( 6.1%) disk\n\n(2) A command to enrich repo with trees\n---------------------------------------\nThere is no good way to \"un-partial\" repository that was cloned with\n'--filter=tree:0' to have all trees, but no blobs.\n\nThere seems to be a dirty way of doing that by abusing 'fetch --deepen'\nwhich happens to skip \"ref tip already present locally\" check, but\nit will also re-download all commits, which means extra ~0.5gb network\nin case of Linux repo.\n\n(3) A command to download ALL trees and/or blobs for a subpath\n-----------------------------------------------\nSummary: Running a Blame or file log in '--filter=tree:0' repo is\ncurrently very inefficient, up to a point where it can be discussed\nas not really working.\n\nThe suggested command will be able to accept a path and download ALL\ntrees and/or blobs that match it.\n\nThis will solve many problems at once:\n* Solve (2)\n* Make it possible to prepare for efficient blame and file log\n* Make a new experience with super-mono-repos, where user will now\n   be able to only download a part of it by path.\n\nCurrently '--filter=sparse:oid' is there to support that, but it is\nvery hard to use on client side, because it requires paths to be\nalready present in a commit on server.\n\nFor a possible solution, it sounds reasonable to have such filter:\n   --filter=sparse:pathlist=/1/2'\nPath list could be delimited with some special character, and paths\nthemselves could be escaped.\n"},{"id":"408030","messageId":"20201020222934.GB93217@nand.local","threadId":"54470","inReplyTo":"aa7b89ee-08aa-7943-6a00-28dcf344426e@syntevo.com","subject":"Re: Questions about partial clone with '--filter=tree:0'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-10-20T22:29:34Z","receivedAt":"2020-10-20T22:29:40Z","isPatch":false,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Hi Alexandr,\n\nOn Tue, Oct 20, 2020 at 07:09:36PM +0200, Alexandr Miloslavskiy wrote:\n> This is a edited copy of message I sent 2 weeks ago, which unfortunately\n> didn't receive any replies. I tried to make make it shorter this time :)\n\nOops. That can happen sometimes, but thanks for re-sending. I'll try to\nanswer the basic points below.\n\n> ----\n>\n> We are implementing a git UI. One interesting case is the repository\n> cloned with '--filter=tree:0', because it makes it a lot harder to\n> run basic git operations such as file log and blame.\n>\n> The problems and potential solutions are outlined below. We should be\n> able to make patches for (2) and (3) if it makes sense to patch these.\n>\n> (1) Is it even considered a realistic use case?\n> -----------------------------------------------\n> Summary: is '--filter=tree:0' a realistic or \"crazy\" scenario that is\n> not considered worthy of supporting?\n\nIt's not an unrealistic scenario, but it might be for what you're trying\nto build. If your UI needs to run, say, 'git log --patch' to show a\nhistorical revision, then you're going to need to fault in a lot of\nmissing objects.\n\nIf that's not something that you need to do often or ever, then having\n'--filter=tree:0' is a good way to get the least amount of data possible\nwhen using a partial clone. But if you're going to be performing\noperations that need those missing objects, you're probably better eat\nthe network/storage cost of it all at once, rather than making the user\nwait for Git to fault in the set of missing objects that it happens to\nneed.\n\n> (2) A command to enrich repo with trees\n> ---------------------------------------\n> There is no good way to \"un-partial\" repository that was cloned with\n> '--filter=tree:0' to have all trees, but no blobs.\n\nThere is no command to do that directly, but it is something that Git is\ncapable of.\n\nIt would look something like:\n\n  $ git config remote.origin.partialclonefilter 'blob:none'\n\nNow your repository is in a state where it has no blobs or trees, but\nthe filter does not prohibit it from getting the trees, so you can ask\nit to grab everything you're missing with:\n\n  $ git fetch origin\n\nThis should even be a pretty fast operation for repositories that have\nbitmaps due to some topics that Peff and I sent to the list a while ago.\nIf it isn't, please let me know.\n\n> There seems to be a dirty way of doing that by abusing 'fetch --deepen'\n> which happens to skip \"ref tip already present locally\" check, but\n> it will also re-download all commits, which means extra ~0.5gb network\n> in case of Linux repo.\n\nMmm, this is probably not what you're looking for. You may be confusing\nshallow clones (of which --deepen is relevant) with partial clones\n(to which --deepen is irrelevant).\n\n> (3) A command to download ALL trees and/or blobs for a subpath\n> -----------------------------------------------\n> Summary: Running a Blame or file log in '--filter=tree:0' repo is\n> currently very inefficient, up to a point where it can be discussed\n> as not really working.\n\nThis may be a \"don't hold it that way\" kind of response, but I don't\nthink that this is quite what you want. Recall that cloning a\nrepository with an object filter happens in two steps: first, an initial\ndownload of all of the objects that it thinks you need, and then\n(second) a follow-up fetch requesting the objects that you need to\npopulate your checkout.\n\nI think what you probably want is a step 1.5 to tell Git \"I'm not going\nto ask for or care about the entirety of my working copy, I really just\nwant objects in path...\", and you can do that with sparse checkouts. See\nhttps://git-scm.com/docs/git-sparse-checkout for more.\n\nThe flow might be something like:\n\n  $ git clone --sparse --filter=tree:0 git@yourhost.com:repo.git\n\nand then:\n\n  $ cd repo\n  $ git sparse-checkout add foo bar baz\n  $ git checkout .\n\nThanks,\nTaylor\n"},{"id":"408069","messageId":"a4a20c67-4ee3-77b2-8d57-f30843572aa4@syntevo.com","threadId":"54470","inReplyTo":"20201020222934.GB93217@nand.local","subject":"Re: Questions about partial clone with '--filter=tree:0'","fromName":"Alexandr Miloslavskiy","fromEmail":"alexandr.miloslavskiy@syntevo.com","sentAt":"2020-10-21T17:10:02Z","receivedAt":"2020-10-21T17:10:11Z","isPatch":false,"sender":{"key":"alexandr.miloslavskiy@syntevo.com","avatar":null},"body":"On 21.10.2020 0:29, Taylor Blau wrote:\n> Oops. That can happen sometimes, but thanks for re-sending. I'll try to\n> answer the basic points below.\n\nThanks for stepping in!\n\n>> (1) Is it even considered a realistic use case?\n>> -----------------------------------------------\n>> Summary: is '--filter=tree:0' a realistic or \"crazy\" scenario that is\n>> not considered worthy of supporting?\n>\n> It's not an unrealistic scenario, but it might be for what you're trying\n> to build. If your UI needs to run, say, 'git log --patch' to show a\n> historical revision, then you're going to need to fault in a lot of\n> missing objects.\n>\n> If that's not something that you need to do often or ever, then having\n> '--filter=tree:0' is a good way to get the least amount of data possible\n> when using a partial clone. But if you're going to be performing\n> operations that need those missing objects, you're probably better eat\n> the network/storage cost of it all at once, rather than making the user\n> wait for Git to fault in the set of missing objects that it happens to\n> need.\n\nWe currently do not intend to use '--filter=tree:0' ourself, but we are \ntrying to support all kinds of user repositories with our UI. So we \nbasically have these choices:\n\nA) Declare '--filter=tree:0' repos as completely wrong and unsupported\n    in out UI, also giving an option to \"un-partial\" them.\n\nB) Support '--filter=tree:0' repos, but don't support operations such\n    as blame and file log\n\nC) Use some magic to efficiently download objects that will be needed\n    for a command such as Blame, while keeping the rest of the repository\n    partial. This is where the command described in (3) will help a lot.\n\nWe would of course prefer (C) if it's reasonably possible.\n\n>> (2) A command to enrich repo with trees\n>> ---------------------------------------\n>> There is no good way to \"un-partial\" repository that was cloned with\n>> '--filter=tree:0' to have all trees, but no blobs.\n>\n> There is no command to do that directly, but it is something that Git is\n> capable of.\n>\n> It would look something like:\n>\n>    $ git config remote.origin.partialclonefilter 'blob:none'\n>\n> Now your repository is in a state where it has no blobs or trees, but\n> the filter does not prohibit it from getting the trees, so you can ask\n> it to grab everything you're missing with:\n>\n>    $ git fetch origin\n>\n> This should even be a pretty fast operation for repositories that have\n> bitmaps due to some topics that Peff and I sent to the list a while ago.\n> If it isn't, please let me know.\n\nUnfortunately this does not work as expected. Try the following steps:\n\nA) Clone repo with '--filter=tree:0'\n    $ git clone --bare --filter=tree:0 --branch master \nhttps://github.com/git/git.git\n\nB) Change filter to 'blob:none'\n    $ cd git.git\n    $ git config remote.origin.partialclonefilter 'blob:none'\n\nC) fetch\n    $ git fetch origin\n    Note that there is no 'Receiving objects:' output.\n\nD) Verify that trees were downloaded\n    $ git cat-file -p HEAD | grep tree\n      tree ee5b5b41305cda618862beebc9c94859ae276e5a\n    $ git cat-file -t ee5b5b41305cda618862beebc9c94859ae276e5a\n      Note that 1 object gets downloaded. This confirms that (C) didn't\n      achieve the goal.\n\nIt happens due to 'check_exist_and_connected()' test in 'fetch_refs()'.\nSince the tip of the ref is already available locally (even though it\nis missing all trees), nothing is downloaded.\n\n>> There seems to be a dirty way of doing that by abusing 'fetch --deepen'\n>> which happens to skip \"ref tip already present locally\" check, but\n>> it will also re-download all commits, which means extra ~0.5gb network\n>> in case of Linux repo.\n>\n> Mmm, this is probably not what you're looking for. You may be confusing\n> shallow clones (of which --deepen is relevant) with partial clones\n> (to which --deepen is irrelevant).\n\nYes, '--deepen' is intended for shallow clones. But abusing it for\npartial clones allows to skip 'check_exist_and_connected()' test.\nHowever, I did more testing today, and in many cases server itself\nrefuses to send objects, probably due to sent 'HAVE' or something\nelse. So even '--deepen' doesn't really help.\n\n> I think what you probably want is a step 1.5 to tell Git \"I'm not going\n> to ask for or care about the entirety of my working copy, I really just\n> want objects in path...\", and you can do that with sparse checkouts. See\n> https://git-scm.com/docs/git-sparse-checkout for more.\n\nFor simplicity of discussion, let's focus on the problem of running\nBlame efficiently in a repo that was cloned with '--filter=tree:0'. In\norder to blame file '/1/2/Foo.txt', we will need the following:\n\n* Trees '/1'\n* Trees '/1/2'\n* Blobs '/1/2/Foo.txt'\n\nAll of these will be needed to unknown commit depth. For simplicity,\nthe proposed command will download these for all commits. Specifying\na range of revisions could be nice, but I feel that it's not worth the\ncomplexity.\n\nCorrect me if I'm wrong: I think that sparse checkout will not help to\nachieve the goal?\n\nThis is why I suggest a command that will accept paths and send\nrequested objects, also forcing server to assume that all of them are\nmissing in client's repository.\n"},{"id":"408073","messageId":"20201021173153.GC1237181@nand.local","threadId":"54470","inReplyTo":"a4a20c67-4ee3-77b2-8d57-f30843572aa4@syntevo.com","subject":"Re: Questions about partial clone with '--filter=tree:0'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-10-21T17:31:53Z","receivedAt":"2020-10-21T17:32:00Z","isPatch":false,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Oct 21, 2020 at 07:10:02PM +0200, Alexandr Miloslavskiy wrote:\n> We currently do not intend to use '--filter=tree:0' ourself, but we are\n> trying to support all kinds of user repositories with our UI. So we\n> basically have these choices:\n>\n> A) Declare '--filter=tree:0' repos as completely wrong and unsupported\n>    in out UI, also giving an option to \"un-partial\" them.\n>\n> B) Support '--filter=tree:0' repos, but don't support operations such\n>    as blame and file log\n>\n> C) Use some magic to efficiently download objects that will be needed\n>    for a command such as Blame, while keeping the rest of the repository\n>    partial. This is where the command described in (3) will help a lot.\n>\n> We would of course prefer (C) if it's reasonably possible.\n\n(C) is probably the most reasonable. If you have a promisor remote which\nis missing objects, running 'git blame' etc. will transparently download\nwhatever objects it is missing.\n\n> Unfortunately this does not work as expected. Try the following steps:\n>\n> A) Clone repo with '--filter=tree:0'\n>    $ git clone --bare --filter=tree:0 --branch master\n> https://github.com/git/git.git\n>\n> B) Change filter to 'blob:none'\n>    $ cd git.git\n>    $ git config remote.origin.partialclonefilter 'blob:none'\n>\n> C) fetch\n>    $ git fetch origin\n>    Note that there is no 'Receiving objects:' output.\n\nAh; I would have thought that the server would have sent objects, even\nthough we have lots of 'have' lines, since we are treating the server as\na promisor remote and might not have the full reachability closure over\nthe haves.\n\nJonathan Tan knows better than I do here. Maybe he could chime in.\n\n> > I think what you probably want is a step 1.5 to tell Git \"I'm not going\n> > to ask for or care about the entirety of my working copy, I really just\n> > want objects in path...\", and you can do that with sparse checkouts. See\n> > https://git-scm.com/docs/git-sparse-checkout for more.\n>\n> For simplicity of discussion, let's focus on the problem of running\n> Blame efficiently in a repo that was cloned with '--filter=tree:0'. In\n> order to blame file '/1/2/Foo.txt', we will need the following:\n>\n> * Trees '/1'\n> * Trees '/1/2'\n> * Blobs '/1/2/Foo.txt'\n>\n> All of these will be needed to unknown commit depth. For simplicity,\n> the proposed command will download these for all commits. Specifying\n> a range of revisions could be nice, but I feel that it's not worth the\n> complexity.\n>\n> Correct me if I'm wrong: I think that sparse checkout will not help to\n> achieve the goal?\n\nI see what you're saying. Here sparse-checkout and partial clones\nconfusingly diverge: what you really want is to say \"I want all of the\nobjects that I need to construct this directory at any point in history\"\nso that you can run \"git blame\" on some path within that directory\nwithout the need for a follow-up fetch.\n\n> This is why I suggest a command that will accept paths and send\n> requested objects, also forcing server to assume that all of them are\n> missing in client's repository.\n\nIn any case the '--filter=sparse:<oid>' bit is not recommended for use,\nbut perhaps this is a convincing use-case. I didn't follow the partial\nclone development close enough to know whether this has already been\ndiscussed, but I'm sure that it has.\n\nThanks,\nTaylor\n"},{"id":"408219","messageId":"67a5edc8-d2f3-735a-fc7c-0fb1b93a7cf3@syntevo.com","threadId":"54470","inReplyTo":"20201021173153.GC1237181@nand.local","subject":"Re: Questions about partial clone with '--filter=tree:0'","fromName":"Alexandr Miloslavskiy","fromEmail":"alexandr.miloslavskiy@syntevo.com","sentAt":"2020-10-21T17:46:09Z","receivedAt":"2020-10-23T10:30:59Z","isPatch":false,"sender":{"key":"alexandr.miloslavskiy@syntevo.com","avatar":null},"body":"On 21.10.2020 19:31, Taylor Blau wrote:\n\n > If you have a promisor remote which is missing objects, running\n > 'git blame' etc. will transparently download whatever objects it\n > is missing.\n\nThis is correct, but it downloads things one at a time, which in\ncase of larger repo such as Linux could take weeks to complete. And\ndownloading more things at once isn't easy without the suggested\ncommand.\n\nIt is possible to traverse commit graph, requesting all discovered\nobjects at once, but again in case of Linux, that would mean sending\nmultiple requests with lists of 1 million+ oids. And the number of\nrequests is around the maximum tree depth. Doesn't sound nice.\n\n> Jonathan Tan knows better than I do here. Maybe he could chime in.\n\nI already CC'ed him, I hope he finds time to reply.\n\n> I see what you're saying. Here sparse-checkout and partial clones\n> confusingly diverge: what you really want is to say \"I want all of the\n> objects that I need to construct this directory at any point in history\"\n> so that you can run \"git blame\" on some path within that directory\n> without the need for a follow-up fetch.\n\nRight.\n\n> In any case the '--filter=sparse:<oid>' bit is not recommended for use,\n> but perhaps this is a convincing use-case. I didn't follow the partial\n> clone development close enough to know whether this has already been\n> discussed, but I'm sure that it has.\n\nUnfortunately '--filter=sparse:<oid>' requires the list to be already\ncommitted on server, which limits the usefulness of it a lot.\n"},{"id":"408433","messageId":"20201026182417.2105954-1-jonathantanmy@google.com","threadId":"54470","inReplyTo":"aa7b89ee-08aa-7943-6a00-28dcf344426e@syntevo.com","subject":"Re: Questions about partial clone with '--filter=tree:0'","fromName":"Jonathan Tan","fromEmail":"jonathantanmy@google.com","sentAt":"2020-10-26T18:24:17Z","receivedAt":"2020-10-26T18:24:23Z","isPatch":false,"sender":{"key":"jonathantanmy@fastmail.com","avatar":null},"body":"> (1) Is it even considered a realistic use case?\n> -----------------------------------------------\n> Summary: is '--filter=tree:0' a realistic or \"crazy\" scenario that is\n> not considered worthy of supporting?\n> \n> I decided to use Linux repo, which is reasonably large, and it seems\n> that '--filter=tree:0' could be desired because it helps with disk\n> space (~0.66gb) and network (~0.54gb):\n\nSorry for the late reply - I have been out of office for a while.\n\nAs Taylor said in another email, it's good for some use cases but\nperhaps not for the \"blame\" one that you describe later.\n\n> (2) A command to enrich repo with trees\n> ---------------------------------------\n> There is no good way to \"un-partial\" repository that was cloned with\n> '--filter=tree:0' to have all trees, but no blobs.\n> \n> There seems to be a dirty way of doing that by abusing 'fetch --deepen'\n> which happens to skip \"ref tip already present locally\" check, but\n> it will also re-download all commits, which means extra ~0.5gb network\n> in case of Linux repo.\n\nThat's true. I made some progress with cbe566a071 (\"negotiator/noop: add\nnoop fetch negotiator\", 2020-08-18) (which adds a no-op negotiatior, so\nthe client never reports its own commits as \"have\") but as you said in\nanother email, we still run into the problem that if we have the commit\nthat we're fetching, we still won't fetch it.\n\n> (3) A command to download ALL trees and/or blobs for a subpath\n> -----------------------------------------------\n> Summary: Running a Blame or file log in '--filter=tree:0' repo is\n> currently very inefficient, up to a point where it can be discussed\n> as not really working.\n> \n> The suggested command will be able to accept a path and download ALL\n> trees and/or blobs that match it.\n> \n> This will solve many problems at once:\n> * Solve (2)\n> * Make it possible to prepare for efficient blame and file log\n> * Make a new experience with super-mono-repos, where user will now\n>    be able to only download a part of it by path.\n\nTo clarify: we partially support the last point - \"git clone\" now\nsupports \"--sparse\". When used with \"--filter\", only the blobs in the\nsparse checkout specification will be fetched, so users are already able\nto download only the objects in a specific path. Having said that, I\nthink you also want the histories of these objects, so admittedly this\nis not complete for your use case.\n\n> Currently '--filter=sparse:oid' is there to support that, but it is\n> very hard to use on client side, because it requires paths to be\n> already present in a commit on server.\n> \n> For a possible solution, it sounds reasonable to have such filter:\n>    --filter=sparse:pathlist=/1/2'\n> Path list could be delimited with some special character, and paths\n> themselves could be escaped.\n\nHaving such an option (and teaching \"blame\" to use it to prefetch) would\nindeed speed up \"blame\". But if we implement this, what would happen if\nthe user ran \"blame\" on the same file twice? I can't think of a way of\npreventing the same fetch from happening twice except by checking the\nexistence of, say, the last 10 OIDs corresponding to that path. But if\nwe have the list of those 10 OIDs, we could just prefetch those 10 OIDs\nwithout needing a new filter.\n\nAnother issue (but a smaller one) is this does not fetch all objects\nnecessary if the file being \"blame\"d has been renamed, but that is\nprobably solvable - we can just refetch with the old name.\n\nAnother possible solution that has been discussed before (but a much\nmore involved one) is to teach Git to be able to serve results of\ncomputations, and then have \"blame\" be able to stitch that with local\ndata. (For example, \"blame\" could check the history of a certain path to\nfind the commit(s) that the remote has information of, query the remote\nfor those commits, and then stitch the results together with local\nhistory.) This scheme would work not only for \"blame\" but for things\nlike \"grep\" (with history) and \"log -S\", whereas\n\"--filter=sparse:parthlist\" would only work with \"blame\". But\nadmittedly, this solution is more involved.\n"},{"id":"408436","messageId":"2f04c074-3eee-766c-bedb-2e3cc0a91528@syntevo.com","threadId":"54470","inReplyTo":"20201026182417.2105954-1-jonathantanmy@google.com","subject":"Re: Questions about partial clone with '--filter=tree:0'","fromName":"Alexandr Miloslavskiy","fromEmail":"alexandr.miloslavskiy@syntevo.com","sentAt":"2020-10-26T18:44:27Z","receivedAt":"2020-10-26T19:19:01Z","isPatch":false,"sender":{"key":"alexandr.miloslavskiy@syntevo.com","avatar":null},"body":"On 26.10.2020 19:24, Jonathan Tan wrote:\n> Sorry for the late reply - I have been out of office for a while.\n\nI'm quite happy to get the replies at all, even if later. Thanks!\n\n> As Taylor said in another email, it's good for some use cases but\n> perhaps not for the \"blame\" one that you describe later.\n\nOK, so our expectations seem to match your expectations, that's good.\n\n> That's true. I made some progress with cbe566a071 (\"negotiator/noop: add\n> noop fetch negotiator\", 2020-08-18) (which adds a no-op negotiatior, so\n> the client never reports its own commits as \"have\") but as you said in\n> another email, we still run into the problem that if we have the commit\n> that we're fetching, we still won't fetch it.\n\nRight, I already discovered 'fetch.negotiationAlgorithm=noop' and gave \nit a quick try, but it didn't seem to help at all.\n\n> To clarify: we partially support the last point - \"git clone\" now\n> supports \"--sparse\". When used with \"--filter\", only the blobs in the\n> sparse checkout specification will be fetched, so users are already able\n> to download only the objects in a specific path.\n\nI see. Still, it seems that two other problems will be solved.\n\n> Having said that, I\n> think you also want the histories of these objects, so admittedly this\n> is not complete for your use case.\n\nRight.\n\n> Having such an option (and teaching \"blame\" to use it to prefetch) would\n> indeed speed up \"blame\". But if we implement this, what would happen if\n> the user ran \"blame\" on the same file twice? I can't think of a way of\n> preventing the same fetch from happening twice except by checking the\n> existence of, say, the last 10 OIDs corresponding to that path. But if\n> we have the list of those 10 OIDs, we could just prefetch those 10 OIDs\n> without needing a new filter.\n\nI must admit that I didn't notice this problem. Still, it seems easy \nenough to solve with this approach:\n\n1) Estimate number of missing things\n2) If \"many\", just download everything for <path> as described before\n    and consider it done.\n3) If \"not so many\", assemble a list of OIDs on the boundary of unknown\n    (for example, all root tree OIDs for commits that are missing any\n    trees) and use the usual fetch to download all OIDs in one go.\n4) Repeat step 3 multiple times. Only N=<maximum tree depth> requests\n    are needed, regardless of the number of commits.\n\n> Another issue (but a smaller one) is this does not fetch all objects\n> necessary if the file being \"blame\"d has been renamed, but that is\n> probably solvable - we can just refetch with the old name.\n\nRight, we also discussed this and figured that we'd just query more\nthings as needed. Maybe also individual other blobs for rename detection.\n\n> Another possible solution that has been discussed before (but a much\n> more involved one) is to teach Git to be able to serve results of\n> computations, and then have \"blame\" be able to stitch that with local\n> data. (For example, \"blame\" could check the history of a certain path to\n> find the commit(s) that the remote has information of, query the remote\n> for those commits, and then stitch the results together with local\n> history.) This scheme would work not only for \"blame\" but for things\n> like \"grep\" (with history) and \"log -S\", whereas\n> \"--filter=sparse:parthlist\" would only work with \"blame\". But\n> admittedly, this solution is more involved.\n\nI understand that you're basically talking about implementing \nprefetching in git itself? To my understanding, this will still need \neither the command I suggested, or implement graph walking with massive \nOID requests as described above in 1)2)3)4). The latter will not require \nprotocol changes, but will involve sending quite a bit of OIDs around.\n"},{"id":"408450","messageId":"20201026194635.2119420-1-jonathantanmy@google.com","threadId":"54470","inReplyTo":"2f04c074-3eee-766c-bedb-2e3cc0a91528@syntevo.com","subject":"Re: Questions about partial clone with '--filter=tree:0'","fromName":"Jonathan Tan","fromEmail":"jonathantanmy@google.com","sentAt":"2020-10-26T19:46:35Z","receivedAt":"2020-10-26T19:46:42Z","isPatch":false,"sender":{"key":"jonathantanmy@fastmail.com","avatar":null},"body":"> > Having such an option (and teaching \"blame\" to use it to prefetch) would\n> > indeed speed up \"blame\". But if we implement this, what would happen if\n> > the user ran \"blame\" on the same file twice? I can't think of a way of\n> > preventing the same fetch from happening twice except by checking the\n> > existence of, say, the last 10 OIDs corresponding to that path. But if\n> > we have the list of those 10 OIDs, we could just prefetch those 10 OIDs\n> > without needing a new filter.\n> \n> I must admit that I didn't notice this problem. Still, it seems easy \n> enough to solve with this approach:\n> \n> 1) Estimate number of missing things\n> 2) If \"many\", just download everything for <path> as described before\n>     and consider it done.\n> 3) If \"not so many\", assemble a list of OIDs on the boundary of unknown\n>     (for example, all root tree OIDs for commits that are missing any\n>     trees) and use the usual fetch to download all OIDs in one go.\n> 4) Repeat step 3 multiple times. Only N=<maximum tree depth> requests\n>     are needed, regardless of the number of commits.\n\nMy point was that if you can estimate it (\"have the list of those 10\nOIDs\"), then you can just fetch it. This does send \"quite a bit of\nOIDs\", as you said below - I'll address it below.\n\n> > Another possible solution that has been discussed before (but a much\n> > more involved one) is to teach Git to be able to serve results of\n> > computations, and then have \"blame\" be able to stitch that with local\n> > data. (For example, \"blame\" could check the history of a certain path to\n> > find the commit(s) that the remote has information of, query the remote\n> > for those commits, and then stitch the results together with local\n> > history.) This scheme would work not only for \"blame\" but for things\n> > like \"grep\" (with history) and \"log -S\", whereas\n> > \"--filter=sparse:parthlist\" would only work with \"blame\". But\n> > admittedly, this solution is more involved.\n> \n> I understand that you're basically talking about implementing \n> prefetching in git itself?\n\nNo - I did talk about prefetching earlier, but here I mean having Git on\nthe server perform the \"blame\" computation itself.\n\nFor example, let's say I want to run \"blame\" on foo.txt at HEAD. HEAD\nand HEAD^ are commits that only the local client has, whereas HEAD^^ was\nfetched from the remote. By comparing HEAD, HEAD^, and HEAD^^, Git knows\nwhich lines come from HEAD and HEAD^. For the rest, Git would make a\nrequest to the server, passing the commit ID and the path, and would get\nback a list of line numbers and commits.\n\n> To my understanding, this will still need \n> either the command I suggested, or implement graph walking with massive \n> OID requests as described above in 1)2)3)4). The latter will not require \n> protocol changes, but will involve sending quite a bit of OIDs around.\n\nYes, prefetching will require graph walking with large OID requests but\nwill not require protocol changes, as you say. I'm not too worried about\nthe large numbers of OIDs - Git servers already have to support\nrelatively large numbers of OIDs to support the bulk prefetch we do\nduring things like checkout and diff.\n"},{"id":"408452","messageId":"6a09c0cd-8e88-1f53-72ca-bc6f9182b517@syntevo.com","threadId":"54470","inReplyTo":"20201026194635.2119420-1-jonathantanmy@google.com","subject":"Re: Questions about partial clone with '--filter=tree:0'","fromName":"Alexandr Miloslavskiy","fromEmail":"alexandr.miloslavskiy@syntevo.com","sentAt":"2020-10-26T20:08:54Z","receivedAt":"2020-10-26T20:08:59Z","isPatch":false,"sender":{"key":"alexandr.miloslavskiy@syntevo.com","avatar":null},"body":"On 26.10.2020 20:46, Jonathan Tan wrote:\n > No - I did talk about prefetching earlier, but here I mean having\n > Git on the server perform the \"blame\" computation itself.\n\nOh! That's an interesting twist. Unfortunately for us, we are\nimplementing our own Blame logic. Thinking of which, I'm now becoming\nmore convinced that graph walking could be the best solution for us,\nbecause it allows any logic, including custom file rename detection.\n\n > For example, let's say I want to run \"blame\" on foo.txt at HEAD. HEAD\n > and HEAD^ are commits that only the local client has, whereas HEAD^^ was\n > fetched from the remote. By comparing HEAD, HEAD^, and HEAD^^, Git knows\n > which lines come from HEAD and HEAD^. For the rest, Git would make a\n > request to the server, passing the commit ID and the path, and would get\n > back a list of line numbers and commits.\n\nSounds quite involved indeed! It's curious how git kind of shifts\ntowards classic server-side VCS such as SVN. When partial clones are\ninvolved, that is.\n\n > Yes, prefetching will require graph walking with large OID requests but\n > will not require protocol changes, as you say. I'm not too worried about\n > the large numbers of OIDs - Git servers already have to support\n > relatively large numbers of OIDs to support the bulk prefetch we do\n > during things like checkout and diff.\n\nHmm, let's talk about Linux repository for the sake of the numbers.\nThe number of commits is ~1M. For a typical Blame (without rename\ndetection), every request will traverse the trees one level deeper, and\nfor just one file blamed, that would mean 1 or 0 trees per commit \n(depending on whether the tree was modified by the commit). The first\nrequest to discover root trees is going to be the largest, and will\nrequest (1*numCommits) OIDs. That makes 1M OIDs in worst case, with\nsubsequent requests probably at ~0.1M, and there will be 1 request per\nevery path component in blamed path.\n\nSo the question is, will git server (or git hosting) become upset\nabout requests for 1M OIDs? Never really tried what is the cost of such\nrequest, what do you think?\n"}]}