{"thread":{"id":"58077","subject":"How to reduce pickaxe times for a particular repo?","startedAt":"2022-06-28T10:51:03Z","lastAt":"2022-07-01T18:21:54Z","messageCount":6,"participants":["Pavel Rappo","Ævar Arnfjörð Bjarmason","Derrick Stolee","Jeff King"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"458030","messageId":"CAChcVumN66OxOjag9gPqgLq7gQrgdaEkZAJabusE-gGC7LLVyw@mail.gmail.com","threadId":"58077","inReplyTo":null,"subject":"How to reduce pickaxe times for a particular repo?","fromName":"Pavel Rappo","fromEmail":"pavel.rappo@gmail.com","sentAt":"2022-06-28T10:50:49Z","receivedAt":"2022-06-28T10:51:03Z","isPatch":false,"sender":{"key":"pavel.rappo@gmail.com","avatar":null},"body":"I have a repo of the following characteristics:\n\n  * 1 branch\n  * 100,000 commits\n  * 1TB in size\n  * The tip of the branch has 55,000 files\n  * No new commits are expected: the repo is abandoned and kept for\narchaeological purposes.\n\nTypically, a `git log -S/-G` lookup takes around a minute to complete.\nI would like to significantly reduce that time. How can I do that? I\ncan spend up to 10x more disk space, if required. The machine has 10\ncores and 32GB of RAM.\n\nThanks,\n-Pavel\n"},{"id":"458031","messageId":"220628.86bkudf19g.gmgdl@evledraar.gmail.com","threadId":"58077","inReplyTo":"CAChcVumN66OxOjag9gPqgLq7gQrgdaEkZAJabusE-gGC7LLVyw@mail.gmail.com","subject":"Re: How to reduce pickaxe times for a particular repo?","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-06-28T11:35:19Z","receivedAt":"2022-06-28T11:58:35Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Jun 28 2022, Pavel Rappo wrote:\n\n> I have a repo of the following characteristics:\n>\n>   * 1 branch\n>   * 100,000 commits\n>   * 1TB in size\n>   * The tip of the branch has 55,000 files\n>   * No new commits are expected: the repo is abandoned and kept for\n> archaeological purposes.\n>\n> Typically, a `git log -S/-G` lookup takes around a minute to complete.\n> I would like to significantly reduce that time. How can I do that? I\n> can spend up to 10x more disk space, if required. The machine has 10\n> cores and 32GB of RAM.\n\nIn git as it stands now the main thing you can do is to limit your seach\nby paths, and if you use the commit-graph and have a git that's using\n\"commitGraph.readChangedPaths\" (defaults to true) doing e.g.:\n\n    git log -p -G<rx> -- tests/\n\nCan really help, or any other filter, such as --author or whatever.\n\nBut eventually you'll simply run into the regex engine being slow, if\nyou're feeling very adventurous I have a very WIP branch to make this a\nlot faster by making -S and -G use PCREv2 as a backend:\nhttp://github.com/avar/git/tree/avar/pcre2-conversion-of-diffcore-pickaxe\n\nBench mark results (made sometime last year) were:\n\n    Test                                                                      origin/next       HEAD\n    ------------------------------------------------------------------------------------------------------------------\n    4209.1: git log -S'int main' <limit-rev>..                                0.38(0.36+0.01)   0.37(0.33+0.04) -2.6%\n    4209.2: git log -S'æ' <limit-rev>..                                       0.51(0.47+0.04)   0.32(0.27+0.05) -37.3%\n    4209.3: git log --pickaxe-regex -S'(int|void|null)' <limit-rev>..         0.72(0.68+0.03)   0.57(0.54+0.03) -20.8%\n    4209.4: git log --pickaxe-regex -S'if *\\([^ ]+ & ' <limit-rev>..          0.60(0.55+0.02)   0.39(0.34+0.05) -35.0%\n    4209.5: git log --pickaxe-regex -S'[àáâãäåæñøùúûüýþ]' <limit-rev>..       0.43(0.40+0.03)   0.50(0.44+0.06) +16.3%\n    4209.6: git log -G'(int|void|null)' <limit-rev>..                         0.64(0.55+0.09)   0.63(0.56+0.05) -1.6%\n    4209.7: git log -G'if *\\([^ ]+ & ' <limit-rev>..                          0.64(0.59+0.05)   0.63(0.56+0.06) -1.6%\n    4209.8: git log -G'[àáâãäåæñøùúûüýþ]' <limit-rev>..                       0.63(0.54+0.08)   0.62(0.55+0.06) -1.6%\n    4209.9: git log -i -S'int main' <limit-rev>..                             0.39(0.35+0.03)   0.38(0.35+0.02) -2.6%\n    4209.10: git log -i -S'æ' <limit-rev>..                                   0.39(0.33+0.06)   0.32(0.28+0.04) -17.9%\n    4209.11: git log -i --pickaxe-regex -S'(int|void|null)' <limit-rev>..     0.90(0.84+0.05)   0.58(0.53+0.04) -35.6%\n    4209.12: git log -i --pickaxe-regex -S'if *\\([^ ]+ & ' <limit-rev>..      0.71(0.64+0.06)   0.40(0.37+0.03) -43.7%\n    4209.13: git log -i --pickaxe-regex -S'[àáâãäåæñøùúûüýþ]' <limit-rev>..   0.43(0.40+0.03)   0.50(0.46+0.04) +16.3%\n    4209.14: git log -i -G'(int|void|null)' <limit-rev>..                     0.64(0.57+0.06)   0.62(0.56+0.05) -3.1%\n    4209.15: git log -i -G'if *\\([^ ]+ & ' <limit-rev>..                      0.65(0.59+0.06)   0.63(0.54+0.08) -3.1%\n    4209.16: git log -i -G'[àáâãäåæñøùúûüýþ]' <limit-rev>..                   0.63(0.55+0.08)   0.62(0.56+0.05) -1.6%\n\nSo it's much faster on some queries in particular, I don't think that\ncode is ready for git.git in its current form, but if you're desperate\nfor performance and need to run ad-hoc queries...\n\nI don't know the full shape of your repo but 1TB in size probably means\nsome very big files? I think you might want to experiment with e.g. a\nfiltered repo to filter out big blobs or something else you may be\nneedlessly searching though (binaries?).\n\nI.e. I think you're probably getting a lot of OS cache churn, where we\ncan't have the working data in memory for your whole search, so you're\nmainly I/O bound.\n\nI did want to (as a future infinite time project) create a search index\nfor regexes in git for -S and -G, i.e. we'd store something like\ntrigrams of potentially matchable content, so we could skip commits &\ntrees quickly if the diff e.g. didn't. contain the fixed string \"int\" or\nwhatever.\n\nBut that's a much bigger project...\n\nIf you're really desperate for performance & willing to hack on\nsomtething custom you could emulate that with a hacky solution, e.g.:\n\n 1. Create a COMMIT=DIFF pair for all commits in your repo, or e.g.\n    PATH=DIFF (so one concat'd diff with all modifications ever to a\n    given path)\n\n 2. Stick that into Lucene with trigram indexing, e.g. ElasticSearch\n    might make this easy. Make sure not to \"store documents\" in the\n    index, you just want the reverse index from say \"int\" to \"documents\"\n    that contain it.\n\n 3. Do a two-step search, where a search like \"foo.*bar\" is first\n    against tha index, where you find say all commits that have \"foo\" in\n    the diff OR \"bar\" in the diff, ditto changed paths.\n\n 4. Feed that list into the \"real\" git log -S or -G search, either\n    limiting by commits, or by paths (taking advantage of the\n    commit-graph path index).\n\nFor someone familiar with the tools involved that should be about a day\nto get to a rough hacky solution, it's mostly gluing existing OTS\nsoftware together.\n\nYou should be able to get your searches down to the tens of millisecond\nrange with that if also carefully manage which parts are in cache, but\nit depends a lot on the exact shape of data in your repo, how much\nmemory you have etc.\n"},{"id":"458032","messageId":"CAChcVu=w8mxFtXHukZkf-VswchH_sRppCm=0XZbwh=9-Y4P8cg@mail.gmail.com","threadId":"58077","inReplyTo":"220628.86bkudf19g.gmgdl@evledraar.gmail.com","subject":"Re: How to reduce pickaxe times for a particular repo?","fromName":"Pavel Rappo","fromEmail":"pavel.rappo@gmail.com","sentAt":"2022-06-28T12:35:55Z","receivedAt":"2022-06-28T12:36:10Z","isPatch":false,"sender":{"key":"pavel.rappo@gmail.com","avatar":null},"body":"On Tue, Jun 28, 2022 at 12:58 PM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n\n<snip>\n\n> But eventually you'll simply run into the regex engine being slow\n\nSince I know very little about git internals, I was under a naive\nimpression that a significant, if not comparable to that of regex,\nportion of pickaxe's time is spent on computing diffs between\nrevisions. So I assumed that there was a way to pre-compute those\ndiffs.\n\n<snip>\n\n>  2. Stick that into Lucene with trigram indexing, e.g. ElasticSearch\n>     might make this easy.\n\n<snip>\n\n> For someone familiar with the tools involved that should be about a day\n> to get to a rough hacky solution, it's mostly gluing existing OTS\n> software together.\n\n<snip>\n\nI'll see what I can do with external systems. You see, I initially\ncame from a similar repository exposed through OpenGrok. But I think\nthat something was wrong with the index or query syntax because I\ncouldn't find the things that I knew were there. I was able to secure\na git repo that was close to that of OpenGrok as I found pickaxe to be\nrobust albeit slow alternative for my searches.\n\nThanks for the suggestion.\n"},{"id":"458033","messageId":"6439e948-ff79-9e10-97f5-378806e25b5b@github.com","threadId":"58077","inReplyTo":"CAChcVumN66OxOjag9gPqgLq7gQrgdaEkZAJabusE-gGC7LLVyw@mail.gmail.com","subject":"Re: How to reduce pickaxe times for a particular repo?","fromName":"Derrick Stolee","fromEmail":"derrickstolee@github.com","sentAt":"2022-06-28T13:01:17Z","receivedAt":"2022-06-28T13:01:26Z","isPatch":false,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 6/28/2022 6:50 AM, Pavel Rappo wrote:\n\nHi Pavel! Welcome.\n\n> I have a repo of the following characteristics:\n> \n>   * 1 branch\n>   * 100,000 commits\n\nThis is not too large.\n\n>   * 1TB in size\n\nThis _is_ large.\n\n>   * The tip of the branch has 55,000 files\n\nAnd again, this is not large.\n\nThis means you have some very large files in your repo, perhaps\neven binary files that you don't intend to search.\n\n>   * No new commits are expected: the repo is abandoned and kept for\n> archaeological purposes.\n> \n> Typically, a `git log -S/-G` lookup takes around a minute to complete.\n> I would like to significantly reduce that time. How can I do that? I\n> can spend up to 10x more disk space, if required. The machine has 10\n> cores and 32GB of RAM.\n\nYou are using -S<string> or -G<regex> to see which commits change the\nnumber of matches of that <string> or <regex>. If you don't provide a\npathspec, then Git will search every changed file, including those\nvery large binary files.\n\nPerhaps you'd like to start by providing a pathspec that limits the\nsearch to only the meaningful code files?\n\nAs far as I know, Git doesn't have any data structures that can speed\nup content-based matches like this. The commit-graph's content-changed\nBloom filters only help Git with questions like \"did this specific file\nchange?\" which is not going to be a critical code path in what you're\ndescribing.\n\nI'm not sure what you're actually trying to ask with -S or -G, so maybe\nit is worth considering other types of queries, such as -L<n>,<m>:<file>\nor something. This is just a shot in the dark, as you might be doing the\nonly thing you _can_ do to solve your problem.\n\nThanks,\n-Stolee\n"},{"id":"458089","messageId":"220629.8635fnfxnz.gmgdl@evledraar.gmail.com","threadId":"58077","inReplyTo":"CAChcVu=w8mxFtXHukZkf-VswchH_sRppCm=0XZbwh=9-Y4P8cg@mail.gmail.com","subject":"Re: How to reduce pickaxe times for a particular repo?","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-06-29T12:31:15Z","receivedAt":"2022-06-29T12:43:02Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Jun 28 2022, Pavel Rappo wrote:\n\n> On Tue, Jun 28, 2022 at 12:58 PM Ævar Arnfjörð Bjarmason\n> <avarab@gmail.com> wrote:\n>\n> <snip>\n>\n>> But eventually you'll simply run into the regex engine being slow\n>\n> Since I know very little about git internals, I was under a naive\n> impression that a significant, if not comparable to that of regex,\n> portion of pickaxe's time is spent on computing diffs between\n> revisions. So I assumed that there was a way to pre-compute those\n> diffs.\n\nYes and no, maybe sort of :)\n\nFirstly, -S doesn't involve a diff, it's comparing the raw pre-post\nimage, and seeing how many times we match.\n\n-G does involve computing the diff.\n\nOne the one hand we're fast at making diffs, but that really shouldn't\nbe significant compared to the speed of a regex engine.\n\nThe other side of this is that we're really stupid about how we invoke\nthe regex engine, historical reasons, backwards compatibility & all\nthat, but we:\n\n * Aren't compiling the regex once, and using it N times in some cases\n   (I have some local patches to fix this)\n * Are computing matches one line at a time, when we could e.g. point\n   PCRE to an entire diff with the right line-split options.\n * Are often doing needless work, e.g. in v2.33 I solved an issue with\n   us continuing to create diffs when we could abort early (see\n   f97fe358576 (pickaxe -G: don't special-case create/delete,\n   2021-04-12)), which resulted in some speed-up.q\n\nSome of these are tricky to fix.\n> <snip>\n>\n>>  2. Stick that into Lucene with trigram indexing, e.g. ElasticSearch\n>>     might make this easy.\n>\n> <snip>\n>\n>> For someone familiar with the tools involved that should be about a day\n>> to get to a rough hacky solution, it's mostly gluing existing OTS\n>> software together.\n>\n> <snip>\n>\n> I'll see what I can do with external systems. You see, I initially\n> came from a similar repository exposed through OpenGrok. But I think\n> that something was wrong with the index or query syntax because I\n> couldn't find the things that I knew were there. I was able to secure\n> a git repo that was close to that of OpenGrok as I found pickaxe to be\n> robust albeit slow alternative for my searches.\n\nThis is the first time I hear about OpenGrok, so no idea, sorry.\n\nOne common pitfall with search indexes is that they tend to have a\nblacklist of words, e.g. Lucene will have \"for\", \"or\" and other common\nEnglish words as part of its defaults, so if you're trying to e.g. find\nwhen you altered a for-loop you might silently be getting no results.\n"},{"id":"458403","messageId":"Yr87Pt8/l2Tte/Gd@coredump.intra.peff.net","threadId":"58077","inReplyTo":"6439e948-ff79-9e10-97f5-378806e25b5b@github.com","subject":"Re: How to reduce pickaxe times for a particular repo?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2022-07-01T18:21:50Z","receivedAt":"2022-07-01T18:21:54Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jun 28, 2022 at 09:01:17AM -0400, Derrick Stolee wrote:\n\n> > Typically, a `git log -S/-G` lookup takes around a minute to complete.\n> > I would like to significantly reduce that time. How can I do that? I\n> > can spend up to 10x more disk space, if required. The machine has 10\n> > cores and 32GB of RAM.\n> \n> You are using -S<string> or -G<regex> to see which commits change the\n> number of matches of that <string> or <regex>. If you don't provide a\n> pathspec, then Git will search every changed file, including those\n> very large binary files.\n> \n> Perhaps you'd like to start by providing a pathspec that limits the\n> search to only the meaningful code files?\n\nI think \"-S\" will search every file, since it's just counting instances\nof the token in each file. But \"-G\" does a diff first, so it skips\nbinary files. So you could probably speed it up in general with a\n.gitattributes that mark large binary files as such. Sort of the same\nconcept as your pathspec suggestion (which is a good one), but you don't\nhave to remember to add it to each invocation. :)\n\n-Peff\n"}]}