{"thread":{"id":"46964","subject":"git-clone causes out of memory","startedAt":"2017-10-13T09:59:36Z","lastAt":"2017-10-14T02:43:38Z","messageCount":23,"participants":["Constantine","Mike Hommey","Christian Couder","Junio C Hamano","Jeff King","Derrick Stolee"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"330341","messageId":"515b1400-4053-70b0-18e2-1f61ebc3b2d7@yandex.ru","threadId":"46964","inReplyTo":null,"subject":"git-clone causes out of memory","fromName":"Constantine","fromEmail":"hi-angel@yandex.ru","sentAt":"2017-10-13T09:51:58Z","receivedAt":"2017-10-13T09:59:36Z","isPatch":false,"sender":{"key":"hi-angel@yandex.ru","avatar":null},"body":"There's a gitbomb on github. It is undoubtedly creative and funny, but \nsince this is a bug in git, I thought it'd be nice to report. The command:\n\n\t$ git clone https://github.com/x0rz/ShadowBrokersFiles\n\nquickly fills out the RAM (e.g. 4GB of free memory for me). To recover, \ncall oom-killer through Alt+SysRq+f, then switch to a TTY and back.\n\nGit version: 2.14.2\n\nP.S.: I am not subscribed to the ML, acc. to \nhttps://git-scm.com/community I will be added to CC.\n"},{"id":"330342","messageId":"20171013100603.5eed26sjjigph2il@glandium.org","threadId":"46964","inReplyTo":"515b1400-4053-70b0-18e2-1f61ebc3b2d7@yandex.ru","subject":"Re: git-clone causes out of memory","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2017-10-13T10:06:03Z","receivedAt":"2017-10-13T10:06:12Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Fri, Oct 13, 2017 at 12:51:58PM +0300, Constantine wrote:\n> There's a gitbomb on github. It is undoubtedly creative and funny, but since\n> this is a bug in git, I thought it'd be nice to report. The command:\n> \n> \t$ git clone https://github.com/x0rz/ShadowBrokersFiles\n\nWhat fills memory is actually the checkout part of the command. git\nclone -n doesn't fail.\n\nCredit should go where it's due: https://kate.io/blog/git-bomb/\n(with the bonus that it comes with explanations)\n\nMike\n"},{"id":"330343","messageId":"CAP8UFD1KuBdUCo=x_q4__s1kW15CWMH1jJkKzXqmf3=T3jcrng@mail.gmail.com","threadId":"46964","inReplyTo":"20171013100603.5eed26sjjigph2il@glandium.org","subject":"Re: git-clone causes out of memory","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2017-10-13T10:26:46Z","receivedAt":"2017-10-13T10:26:53Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"On Fri, Oct 13, 2017 at 12:06 PM, Mike Hommey <mh@glandium.org> wrote:\n> On Fri, Oct 13, 2017 at 12:51:58PM +0300, Constantine wrote:\n>> There's a gitbomb on github. It is undoubtedly creative and funny, but since\n>> this is a bug in git, I thought it'd be nice to report. The command:\n>>\n>>       $ git clone https://github.com/x0rz/ShadowBrokersFiles\n>\n> What fills memory is actually the checkout part of the command. git\n> clone -n doesn't fail.\n>\n> Credit should go where it's due: https://kate.io/blog/git-bomb/\n> (with the bonus that it comes with explanations)\n\nYeah, there is a thread on Hacker News about this too:\n\nhttps://news.ycombinator.com/item?id=15457076\n\nThe original repo on GitHub is:\n\nhttps://github.com/Katee/git-bomb.git\n\nAfter cloning it with -n, there is the following \"funny\" situation:\n\n$ time git rev-list HEAD\n7af99c9e7d4768fa681f4fe4ff61259794cf719b\n18ed56cbc5012117e24a603e7c072cf65d36d469\n45546f17e5801791d4bc5968b91253a2f4b0db72\n\nreal    0m0.004s\nuser    0m0.000s\nsys     0m0.004s\n$ time git rev-list HEAD -- d0/d0/d0/d0/d0/d0/d0/d0/d0/d0/f0\n\nreal    0m0.004s\nuser    0m0.000s\nsys     0m0.000s\n$ time git rev-list HEAD -- d0/d0/d0/d0/d0/d0/d0/d0/d0/d0\n\nreal    0m0.004s\nuser    0m0.000s\nsys     0m0.000s\n$ time git rev-list HEAD -- d0/d0/d0/d0/d0/d0/d0/d0/\n45546f17e5801791d4bc5968b91253a2f4b0db72\n\nreal    0m0.005s\nuser    0m0.008s\nsys     0m0.000s\n$ time git rev-list HEAD -- d0/d0/d0/d0/d0/\n45546f17e5801791d4bc5968b91253a2f4b0db72\n\nreal    0m0.203s\nuser    0m0.112s\nsys     0m0.088s\n$ time git rev-list HEAD -- d0/d0/d0/d0/\n45546f17e5801791d4bc5968b91253a2f4b0db72\n\nreal    0m1.305s\nuser    0m0.720s\nsys     0m0.580s\n$ time git rev-list HEAD -- d0/d0/d0/\n45546f17e5801791d4bc5968b91253a2f4b0db72\n\nreal    0m12.135s\nuser    0m6.700s\nsys     0m5.412s\n\nSo `git rev-list` becomes exponentially more expensive when you run it\non a shorter directory path, though it is fast if you run it without a\npath.\n"},{"id":"330344","messageId":"20171013103722.rvr7536mu2hoo4wb@glandium.org","threadId":"46964","inReplyTo":"CAP8UFD1KuBdUCo=x_q4__s1kW15CWMH1jJkKzXqmf3=T3jcrng@mail.gmail.com","subject":"Re: git-clone causes out of memory","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2017-10-13T10:37:22Z","receivedAt":"2017-10-13T10:37:41Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Fri, Oct 13, 2017 at 12:26:46PM +0200, Christian Couder wrote:\n> On Fri, Oct 13, 2017 at 12:06 PM, Mike Hommey <mh@glandium.org> wrote:\n> > On Fri, Oct 13, 2017 at 12:51:58PM +0300, Constantine wrote:\n> >> There's a gitbomb on github. It is undoubtedly creative and funny, but since\n> >> this is a bug in git, I thought it'd be nice to report. The command:\n> >>\n> >>       $ git clone https://github.com/x0rz/ShadowBrokersFiles\n> >\n> > What fills memory is actually the checkout part of the command. git\n> > clone -n doesn't fail.\n> >\n> > Credit should go where it's due: https://kate.io/blog/git-bomb/\n> > (with the bonus that it comes with explanations)\n> \n> Yeah, there is a thread on Hacker News about this too:\n> \n> https://news.ycombinator.com/item?id=15457076\n> \n> The original repo on GitHub is:\n> \n> https://github.com/Katee/git-bomb.git\n> \n> After cloning it with -n, there is the following \"funny\" situation:\n> \n> $ time git rev-list HEAD\n> 7af99c9e7d4768fa681f4fe4ff61259794cf719b\n> 18ed56cbc5012117e24a603e7c072cf65d36d469\n> 45546f17e5801791d4bc5968b91253a2f4b0db72\n> \n> real    0m0.004s\n> user    0m0.000s\n> sys     0m0.004s\n> $ time git rev-list HEAD -- d0/d0/d0/d0/d0/d0/d0/d0/d0/d0/f0\n> \n> real    0m0.004s\n> user    0m0.000s\n> sys     0m0.000s\n> $ time git rev-list HEAD -- d0/d0/d0/d0/d0/d0/d0/d0/d0/d0\n> \n> real    0m0.004s\n> user    0m0.000s\n> sys     0m0.000s\n> $ time git rev-list HEAD -- d0/d0/d0/d0/d0/d0/d0/d0/\n> 45546f17e5801791d4bc5968b91253a2f4b0db72\n> \n> real    0m0.005s\n> user    0m0.008s\n> sys     0m0.000s\n> $ time git rev-list HEAD -- d0/d0/d0/d0/d0/\n> 45546f17e5801791d4bc5968b91253a2f4b0db72\n> \n> real    0m0.203s\n> user    0m0.112s\n> sys     0m0.088s\n> $ time git rev-list HEAD -- d0/d0/d0/d0/\n> 45546f17e5801791d4bc5968b91253a2f4b0db72\n> \n> real    0m1.305s\n> user    0m0.720s\n> sys     0m0.580s\n> $ time git rev-list HEAD -- d0/d0/d0/\n> 45546f17e5801791d4bc5968b91253a2f4b0db72\n> \n> real    0m12.135s\n> user    0m6.700s\n> sys     0m5.412s\n> \n> So `git rev-list` becomes exponentially more expensive when you run it\n> on a shorter directory path, though it is fast if you run it without a\n> path.\n\nThat's because there are 10^7 files under d0/d0/d0, 10^6 under\nd0/d0/d0/d0/, 10^5 under d0/d0/d0/d0/d0/ etc.\n\nSo really, this is all about things being slower when there's a crazy\nnumber of files. Picture me surprised.\n\nWhat makes it kind of special is that the repository contains a lot of\npaths/files, but very few objects, because it's duplicating everything.\n\nAll the 10^10 blobs have the same content, all the 10^9 trees that point\nto them have the same content, all the 10^8 trees that point to those\ntrees have the same content, etc.\n\nIf git wasn't effectively deduplicating identical content, the repository\nwould be multiple gigabytes large.\n\nMike\n"},{"id":"330345","messageId":"CAP8UFD3vniWZQ9Wb1oMo-bbj8n7CTjTHUNhBRwg6jN9x0+ApAQ@mail.gmail.com","threadId":"46964","inReplyTo":"20171013103722.rvr7536mu2hoo4wb@glandium.org","subject":"Re: git-clone causes out of memory","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2017-10-13T10:44:52Z","receivedAt":"2017-10-13T10:44:59Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"On Fri, Oct 13, 2017 at 12:37 PM, Mike Hommey <mh@glandium.org> wrote:\n> On Fri, Oct 13, 2017 at 12:26:46PM +0200, Christian Couder wrote:\n>>\n>> After cloning it with -n, there is the following \"funny\" situation:\n>>\n>> $ time git rev-list HEAD\n>> 7af99c9e7d4768fa681f4fe4ff61259794cf719b\n>> 18ed56cbc5012117e24a603e7c072cf65d36d469\n>> 45546f17e5801791d4bc5968b91253a2f4b0db72\n>>\n>> real    0m0.004s\n>> user    0m0.000s\n>> sys     0m0.004s\n>> $ time git rev-list HEAD -- d0/d0/d0/d0/d0/d0/d0/d0/d0/d0/f0\n>>\n>> real    0m0.004s\n>> user    0m0.000s\n>> sys     0m0.000s\n>> $ time git rev-list HEAD -- d0/d0/d0/d0/d0/d0/d0/d0/d0/d0\n>>\n>> real    0m0.004s\n>> user    0m0.000s\n>> sys     0m0.000s\n>> $ time git rev-list HEAD -- d0/d0/d0/d0/d0/d0/d0/d0/\n>> 45546f17e5801791d4bc5968b91253a2f4b0db72\n>>\n>> real    0m0.005s\n>> user    0m0.008s\n>> sys     0m0.000s\n>> $ time git rev-list HEAD -- d0/d0/d0/d0/d0/\n>> 45546f17e5801791d4bc5968b91253a2f4b0db72\n>>\n>> real    0m0.203s\n>> user    0m0.112s\n>> sys     0m0.088s\n>> $ time git rev-list HEAD -- d0/d0/d0/d0/\n>> 45546f17e5801791d4bc5968b91253a2f4b0db72\n>>\n>> real    0m1.305s\n>> user    0m0.720s\n>> sys     0m0.580s\n>> $ time git rev-list HEAD -- d0/d0/d0/\n>> 45546f17e5801791d4bc5968b91253a2f4b0db72\n>>\n>> real    0m12.135s\n>> user    0m6.700s\n>> sys     0m5.412s\n>>\n>> So `git rev-list` becomes exponentially more expensive when you run it\n>> on a shorter directory path, though it is fast if you run it without a\n>> path.\n>\n> That's because there are 10^7 files under d0/d0/d0, 10^6 under\n> d0/d0/d0/d0/, 10^5 under d0/d0/d0/d0/d0/ etc.\n>\n> So really, this is all about things being slower when there's a crazy\n> number of files. Picture me surprised.\n>\n> What makes it kind of special is that the repository contains a lot of\n> paths/files, but very few objects, because it's duplicating everything.\n>\n> All the 10^10 blobs have the same content, all the 10^9 trees that point\n> to them have the same content, all the 10^8 trees that point to those\n> trees have the same content, etc.\n>\n> If git wasn't effectively deduplicating identical content, the repository\n> would be multiple gigabytes large.\n\nYeah, but perhaps Git could be smarter when rev-listing too and avoid\nprocessing files or directories it has already seen?\n"},{"id":"330348","messageId":"xmqqr2u7uuc8.fsf@gitster.mtv.corp.google.com","threadId":"46964","inReplyTo":"CAP8UFD3vniWZQ9Wb1oMo-bbj8n7CTjTHUNhBRwg6jN9x0+ApAQ@mail.gmail.com","subject":"Re: git-clone causes out of memory","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-10-13T12:04:55Z","receivedAt":"2017-10-13T12:06:03Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Christian Couder <christian.couder@gmail.com> writes:\n\n> Yeah, but perhaps Git could be smarter when rev-listing too and avoid\n> processing files or directories it has already seen?\n\nAren't you suggesting to optimize for a wrong case?\n"},{"id":"330349","messageId":"2f9b8856-dacc-768d-32c2-985f5f145ba7@yandex.ru","threadId":"46964","inReplyTo":"xmqqr2u7uuc8.fsf@gitster.mtv.corp.google.com","subject":"Re: git-clone causes out of memory","fromName":"Constantine","fromEmail":"hi-angel@yandex.ru","sentAt":"2017-10-13T12:12:43Z","receivedAt":"2017-10-13T12:20:20Z","isPatch":false,"sender":{"key":"hi-angel@yandex.ru","avatar":null},"body":"On 13.10.2017 15:04, Junio C Hamano wrote:\n> Christian Couder <christian.couder@gmail.com> writes:\n> \n>> Yeah, but perhaps Git could be smarter when rev-listing too and avoid\n>> processing files or directories it has already seen?\n> \n> Aren't you suggesting to optimize for a wrong case?\n> \n\nAnything that is possible with a software should be considered as a \npossible usecase. It's in fact a DoS attack. Imagine there's a server \nthat using git to process something, and now there's a way to knock down \nthis server. It's also bad from a promotional stand point.\n"},{"id":"330350","messageId":"20171013123521.hop5hrfsyagu7znl@sigill.intra.peff.net","threadId":"46964","inReplyTo":"20171013100603.5eed26sjjigph2il@glandium.org","subject":"Re: git-clone causes out of memory","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-10-13T12:35:21Z","receivedAt":"2017-10-13T12:35:28Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Oct 13, 2017 at 07:06:03PM +0900, Mike Hommey wrote:\n\n> On Fri, Oct 13, 2017 at 12:51:58PM +0300, Constantine wrote:\n> > There's a gitbomb on github. It is undoubtedly creative and funny, but since\n> > this is a bug in git, I thought it'd be nice to report. The command:\n> > \n> > \t$ git clone https://github.com/x0rz/ShadowBrokersFiles\n> \n> What fills memory is actually the checkout part of the command. git\n> clone -n doesn't fail.\n> \n> Credit should go where it's due: https://kate.io/blog/git-bomb/\n> (with the bonus that it comes with explanations)\n\nThat is a nice explanation, and I think the dates point to the kate.io\npost as the parent there.\n\nI certainly have come across this type of \"bomb\" repo before, but I\ncouldn't come up with any public references (somebody uploaded such a\nrepository to GitHub in 2014). So I think this may be the first public\ndiscussion of the idea.\n\n-Peff\n"},{"id":"330351","messageId":"20171013124456.qsbaol7txdgdb6wq@sigill.intra.peff.net","threadId":"46964","inReplyTo":"2f9b8856-dacc-768d-32c2-985f5f145ba7@yandex.ru","subject":"Re: git-clone causes out of memory","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-10-13T12:44:56Z","receivedAt":"2017-10-13T12:45:03Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Oct 13, 2017 at 03:12:43PM +0300, Constantine wrote:\n\n> On 13.10.2017 15:04, Junio C Hamano wrote:\n> > Christian Couder <christian.couder@gmail.com> writes:\n> > \n> > > Yeah, but perhaps Git could be smarter when rev-listing too and avoid\n> > > processing files or directories it has already seen?\n> > \n> > Aren't you suggesting to optimize for a wrong case?\n> > \n> \n> Anything that is possible with a software should be considered as a possible\n> usecase. It's in fact a DoS attack. Imagine there's a server that using git\n> to process something, and now there's a way to knock down this server. It's\n> also bad from a promotional stand point.\n\nBut the point is that you'd have the same problem with any repository\nthat had 10^7 files in it. Yes, it's convenient for the attacker that\nthere are only 9 objects, but fundamentally it's pretty easy for an\nattacker to construct repositories that have large trees (or very deep\ntrees -- that's what causes stack exhaustion in some cases).\n\nNote too that this attack almost always comes down to the diff code\n(which is why it kicks in for pathspec limiting) which has to actually\nexpand the tree. Most \"normal\" server-side operations (like accepting\npushes or serving fetches) operate only on the object graph and _do_\navoid processing already-seen objects.\n\nAs soon as servers start trying to checkout or diff, though, the attack\nsurface gets quite large. And you really need to start thinking about\nhaving resource limits and quotas for CPU and memory use of each process\n(and group by requesting user, IP, repo, etc).\n\nI think the main thing Git could be doing here is to limit the size of\nthe tree (both width and depth). But arbitrary limits like that have a\nway of being annoying, and I think it just pushes resource-exhaustion\nattacks off a little (e.g., can you construct a blob that behaves badly\nwith the \"--patch\"?).\n\n-Peff\n"},{"id":"330355","messageId":"f35d03b5-a525-87b3-a426-bd892edf0c36@gmail.com","threadId":"46964","inReplyTo":"20171013124456.qsbaol7txdgdb6wq@sigill.intra.peff.net","subject":"Re: git-clone causes out of memory","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2017-10-13T13:15:53Z","receivedAt":"2017-10-13T13:16:07Z","isPatch":false,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 10/13/2017 8:44 AM, Jeff King wrote:\n> On Fri, Oct 13, 2017 at 03:12:43PM +0300, Constantine wrote:\n>\n>> On 13.10.2017 15:04, Junio C Hamano wrote:\n>>> Christian Couder <christian.couder@gmail.com> writes:\n>>>\n>>>> Yeah, but perhaps Git could be smarter when rev-listing too and avoid\n>>>> processing files or directories it has already seen?\n>>> Aren't you suggesting to optimize for a wrong case?\n>>>\n>> Anything that is possible with a software should be considered as a possible\n>> usecase. It's in fact a DoS attack. Imagine there's a server that using git\n>> to process something, and now there's a way to knock down this server. It's\n>> also bad from a promotional stand point.\n> But the point is that you'd have the same problem with any repository\n> that had 10^7 files in it. Yes, it's convenient for the attacker that\n> there are only 9 objects, but fundamentally it's pretty easy for an\n> attacker to construct repositories that have large trees (or very deep\n> trees -- that's what causes stack exhaustion in some cases).\n>\n> Note too that this attack almost always comes down to the diff code\n> (which is why it kicks in for pathspec limiting) which has to actually\n> expand the tree. Most \"normal\" server-side operations (like accepting\n> pushes or serving fetches) operate only on the object graph and _do_\n> avoid processing already-seen objects.\n>\n> As soon as servers start trying to checkout or diff, though, the attack\n> surface gets quite large. And you really need to start thinking about\n> having resource limits and quotas for CPU and memory use of each process\n> (and group by requesting user, IP, repo, etc).\n>\n> I think the main thing Git could be doing here is to limit the size of\n> the tree (both width and depth). But arbitrary limits like that have a\n> way of being annoying, and I think it just pushes resource-exhaustion\n> attacks off a little (e.g., can you construct a blob that behaves badly\n> with the \"--patch\"?).\n>\n> -Peff\n\nI'm particularly interested in why `git rev-list HEAD -- [path]` gets \nslower in this case, because I wrote the history algorithm used by VSTS. \nIn our algorithm, we only walk the list of objects from commit to the \ntree containing the path item. For example, in the path d0/d0/d0, we \nwould only walk:\n\n     commit --root--> tree --d0--> tree --d0--> tree [parse oid for d0 \nentry]\n\n From this, we can determine the TREESAME relationship by parsing four \nobjects without parsing all contents below d0/d0/d0.\n\nThe reason we have exponential behavior in `git rev-list` is because we \nare calling diff_tree_oid() in tree-diff.c recursively without \nshort-circuiting on equal OIDs.\n\nI will prepare a patch that adds this OID-equal short-circuit to avoid \nthis exponential behavior. I'll model my patch against a similar patch \nin master:\n\n     commit d12a8cf0af18804c2000efc7a0224da631e04cd1 unpack-trees: avoid \nduplicate ODB lookups during checkout\n\nIt will also significantly speed up rev-list calls for short paths in \ndeep repositories. It will not be very measurable in the git or Linux \nrepos because their shallow folder structure.\n\nThanks,\n-Stolee\n"},{"id":"330356","messageId":"a4ebf552-35d4-d55f-6f08-731afa2cd2de@gmail.com","threadId":"46964","inReplyTo":"f35d03b5-a525-87b3-a426-bd892edf0c36@gmail.com","subject":"Re: git-clone causes out of memory","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2017-10-13T13:39:14Z","receivedAt":"2017-10-13T13:39:24Z","isPatch":false,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 10/13/2017 9:15 AM, Derrick Stolee wrote:\n> On 10/13/2017 8:44 AM, Jeff King wrote:\n>> On Fri, Oct 13, 2017 at 03:12:43PM +0300, Constantine wrote:\n>>\n>>> On 13.10.2017 15:04, Junio C Hamano wrote:\n>>>> Christian Couder <christian.couder@gmail.com> writes:\n>>>>\n>>>>> Yeah, but perhaps Git could be smarter when rev-listing too and avoid\n>>>>> processing files or directories it has already seen?\n>>>> Aren't you suggesting to optimize for a wrong case?\n>>>>\n>>> Anything that is possible with a software should be considered as a \n>>> possible\n>>> usecase. It's in fact a DoS attack. Imagine there's a server that \n>>> using git\n>>> to process something, and now there's a way to knock down this \n>>> server. It's\n>>> also bad from a promotional stand point.\n>> But the point is that you'd have the same problem with any repository\n>> that had 10^7 files in it. Yes, it's convenient for the attacker that\n>> there are only 9 objects, but fundamentally it's pretty easy for an\n>> attacker to construct repositories that have large trees (or very deep\n>> trees -- that's what causes stack exhaustion in some cases).\n>>\n>> Note too that this attack almost always comes down to the diff code\n>> (which is why it kicks in for pathspec limiting) which has to actually\n>> expand the tree. Most \"normal\" server-side operations (like accepting\n>> pushes or serving fetches) operate only on the object graph and _do_\n>> avoid processing already-seen objects.\n>>\n>> As soon as servers start trying to checkout or diff, though, the attack\n>> surface gets quite large. And you really need to start thinking about\n>> having resource limits and quotas for CPU and memory use of each process\n>> (and group by requesting user, IP, repo, etc).\n>>\n>> I think the main thing Git could be doing here is to limit the size of\n>> the tree (both width and depth). But arbitrary limits like that have a\n>> way of being annoying, and I think it just pushes resource-exhaustion\n>> attacks off a little (e.g., can you construct a blob that behaves badly\n>> with the \"--patch\"?).\n>>\n>> -Peff\n>\n> I'm particularly interested in why `git rev-list HEAD -- [path]` gets \n> slower in this case, because I wrote the history algorithm used by \n> VSTS. In our algorithm, we only walk the list of objects from commit \n> to the tree containing the path item. For example, in the path \n> d0/d0/d0, we would only walk:\n>\n>     commit --root--> tree --d0--> tree --d0--> tree [parse oid for d0 \n> entry]\n>\n> From this, we can determine the TREESAME relationship by parsing four \n> objects without parsing all contents below d0/d0/d0.\n>\n> The reason we have exponential behavior in `git rev-list` is because \n> we are calling diff_tree_oid() in tree-diff.c recursively without \n> short-circuiting on equal OIDs.\n>\n> I will prepare a patch that adds this OID-equal short-circuit to avoid \n> this exponential behavior. I'll model my patch against a similar patch \n> in master:\n>\n>     commit d12a8cf0af18804c2000efc7a0224da631e04cd1 unpack-trees: \n> avoid duplicate ODB lookups during checkout\n>\n> It will also significantly speed up rev-list calls for short paths in \n> deep repositories. It will not be very measurable in the git or Linux \n> repos because their shallow folder structure.\n>\n> Thanks,\n> -Stolee\n\nSince I don't understand enough about the consumers to diff_tree_oid() \n(and the fact that the recursive behavior may be wanted in some cases), \nI think we can fix this in builtin/rev-list.c with this simple diff:\n\n---\n\ndiff --git a/builtin/rev-list.c b/builtin/rev-list.c\nindex ded1577424..b2e8e02cc8 100644\n--- a/builtin/rev-list.c\n+++ b/builtin/rev-list.c\n@@ -285,6 +285,9 @@ int cmd_rev_list(int argc, const char **argv, const \nchar *prefix)\n\n         git_config(git_default_config, NULL);\n         init_revisions(&revs, prefix);\n+\n+       revs.pruning.flags = revs.pruning.flags & ~DIFF_OPT_RECURSIVE;\n+\n         revs.abbrev = DEFAULT_ABBREV;\n         revs.commit_format = CMIT_FMT_UNSPECIFIED;\n         argc = setup_revisions(argc, argv, &revs, NULL);\n\n"},{"id":"330357","messageId":"20171013135058.q7vhufdtin42ddic@sigill.intra.peff.net","threadId":"46964","inReplyTo":"a4ebf552-35d4-d55f-6f08-731afa2cd2de@gmail.com","subject":"Re: git-clone causes out of memory","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-10-13T13:50:58Z","receivedAt":"2017-10-13T13:51:05Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Oct 13, 2017 at 09:39:14AM -0400, Derrick Stolee wrote:\n\n> Since I don't understand enough about the consumers to diff_tree_oid() (and\n> the fact that the recursive behavior may be wanted in some cases), I think\n> we can fix this in builtin/rev-list.c with this simple diff:\n> \n> ---\n> \n> diff --git a/builtin/rev-list.c b/builtin/rev-list.c\n> index ded1577424..b2e8e02cc8 100644\n> --- a/builtin/rev-list.c\n> +++ b/builtin/rev-list.c\n> @@ -285,6 +285,9 @@ int cmd_rev_list(int argc, const char **argv, const char\n> *prefix)\n> \n>         git_config(git_default_config, NULL);\n>         init_revisions(&revs, prefix);\n> +\n> +       revs.pruning.flags = revs.pruning.flags & ~DIFF_OPT_RECURSIVE;\n> +\n\nHmm, this feels wrong, because we _do_ want to recurse down and follow\nthe pathspec to see if there are real changes.\n\nWe should be comparing an empty tree and d0/d0/d0/d0 (or however deep\nyour pathspec goes). We should be able to see immediately that the entry\nis not present between the two and not bother descending. After all,\nwe've set the QUICK flag in init_revisions(). So the real question is\nwhy QUICK is not kicking in.\n\n-Peff\n"},{"id":"330358","messageId":"53f98311-3c5f-9863-5f6c-bc4f25fad317@gmail.com","threadId":"46964","inReplyTo":"20171013135058.q7vhufdtin42ddic@sigill.intra.peff.net","subject":"Re: git-clone causes out of memory","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2017-10-13T13:55:15Z","receivedAt":"2017-10-13T13:55:23Z","isPatch":false,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 10/13/2017 9:50 AM, Jeff King wrote:\n> On Fri, Oct 13, 2017 at 09:39:14AM -0400, Derrick Stolee wrote:\n>\n>> Since I don't understand enough about the consumers to diff_tree_oid() (and\n>> the fact that the recursive behavior may be wanted in some cases), I think\n>> we can fix this in builtin/rev-list.c with this simple diff:\n>>\n>> ---\n>>\n>> diff --git a/builtin/rev-list.c b/builtin/rev-list.c\n>> index ded1577424..b2e8e02cc8 100644\n>> --- a/builtin/rev-list.c\n>> +++ b/builtin/rev-list.c\n>> @@ -285,6 +285,9 @@ int cmd_rev_list(int argc, const char **argv, const char\n>> *prefix)\n>>\n>>          git_config(git_default_config, NULL);\n>>          init_revisions(&revs, prefix);\n>> +\n>> +       revs.pruning.flags = revs.pruning.flags & ~DIFF_OPT_RECURSIVE;\n>> +\n\n(Note: I'm running tests now and see that this breaks behavior. \nDefinitely not the solution we want.)\n\n> Hmm, this feels wrong, because we _do_ want to recurse down and follow\n> the pathspec to see if there are real changes.\n>\n> We should be comparing an empty tree and d0/d0/d0/d0 (or however deep\n> your pathspec goes). We should be able to see immediately that the entry\n> is not present between the two and not bother descending. After all,\n> we've set the QUICK flag in init_revisions(). So the real question is\n> why QUICK is not kicking in.\n>\n> -Peff\n\nI'm struggling to understand your meaning. We want to walk from root to \nd0/d0/d0/d0, but there is no reason to walk beyond that tree. But maybe \nthat's what the QUICK flag is supposed to do.\n\nThanks,\n-Stolee\n"},{"id":"330359","messageId":"20171013135636.o2vhktt7aqx6luuy@sigill.intra.peff.net","threadId":"46964","inReplyTo":"53f98311-3c5f-9863-5f6c-bc4f25fad317@gmail.com","subject":"Re: git-clone causes out of memory","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-10-13T13:56:37Z","receivedAt":"2017-10-13T13:56:44Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Oct 13, 2017 at 09:55:15AM -0400, Derrick Stolee wrote:\n\n> > We should be comparing an empty tree and d0/d0/d0/d0 (or however deep\n> > your pathspec goes). We should be able to see immediately that the entry\n> > is not present between the two and not bother descending. After all,\n> > we've set the QUICK flag in init_revisions(). So the real question is\n> > why QUICK is not kicking in.\n> \n> I'm struggling to understand your meaning. We want to walk from root to\n> d0/d0/d0/d0, but there is no reason to walk beyond that tree. But maybe\n> that's what the QUICK flag is supposed to do.\n\nYes, that's exactly what it is for. When we see the first difference we\nshould say \"aha, the caller only wanted to know whether there was a\ndifference, not what it was\" and return immediately. See\ndiff_can_quit_early().\n\n-Peff\n"},{"id":"330360","messageId":"20171013141018.62zvezivkkhloc5d@sigill.intra.peff.net","threadId":"46964","inReplyTo":"20171013135636.o2vhktt7aqx6luuy@sigill.intra.peff.net","subject":"Re: git-clone causes out of memory","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-10-13T14:10:18Z","receivedAt":"2017-10-13T14:10:39Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Oct 13, 2017 at 09:56:36AM -0400, Jeff King wrote:\n\n> On Fri, Oct 13, 2017 at 09:55:15AM -0400, Derrick Stolee wrote:\n> \n> > > We should be comparing an empty tree and d0/d0/d0/d0 (or however deep\n> > > your pathspec goes). We should be able to see immediately that the entry\n> > > is not present between the two and not bother descending. After all,\n> > > we've set the QUICK flag in init_revisions(). So the real question is\n> > > why QUICK is not kicking in.\n> > \n> > I'm struggling to understand your meaning. We want to walk from root to\n> > d0/d0/d0/d0, but there is no reason to walk beyond that tree. But maybe\n> > that's what the QUICK flag is supposed to do.\n> \n> Yes, that's exactly what it is for. When we see the first difference we\n> should say \"aha, the caller only wanted to know whether there was a\n> difference, not what it was\" and return immediately. See\n> diff_can_quit_early().\n\nHmm. So this patch makes it go fast:\n\ndiff --git a/revision.c b/revision.c\nindex d167223e69..b52ea4e9d8 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -409,7 +409,7 @@ static void file_add_remove(struct diff_options *options,\n \tint diff = addremove == '+' ? REV_TREE_NEW : REV_TREE_OLD;\n \n \ttree_difference |= diff;\n-\tif (tree_difference == REV_TREE_DIFFERENT)\n+\tif (tree_difference & REV_TREE_DIFFERENT)\n \t\tDIFF_OPT_SET(options, HAS_CHANGES);\n }\n \n\nBut that essentially makes the conditional a noop (since we know we set\neither NEW or OLD above and DIFFERENT is the union of those flags).\n\nI'm not sure I understand why file_add_remove() would ever want to avoid\nsetting HAS_CHANGES (certainly its companion file_change() always does).\nThis goes back to Junio's dd47aa3133 (try-to-simplify-commit: use\ndiff-tree --quiet machinery., 2007-03-14).\n\nMaybe I am missing something, but AFAICT this was always buggy. But\nsince it only affects adds and deletes, maybe nobody noticed? I'm also\nnot sure if it only causes a slowdown, or if this could cause us to\nerroneously mark something as TREESAME which isn't (I _do_ think people\nwould have noticed that).\n\n-Peff\n"},{"id":"330361","messageId":"20171013142004.ocxpdkkbcxpi52yv@sigill.intra.peff.net","threadId":"46964","inReplyTo":"20171013141018.62zvezivkkhloc5d@sigill.intra.peff.net","subject":"Re: git-clone causes out of memory","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-10-13T14:20:05Z","receivedAt":"2017-10-13T14:20:11Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Oct 13, 2017 at 10:10:18AM -0400, Jeff King wrote:\n\n> Hmm. So this patch makes it go fast:\n> \n> diff --git a/revision.c b/revision.c\n> index d167223e69..b52ea4e9d8 100644\n> --- a/revision.c\n> +++ b/revision.c\n> @@ -409,7 +409,7 @@ static void file_add_remove(struct diff_options *options,\n>  \tint diff = addremove == '+' ? REV_TREE_NEW : REV_TREE_OLD;\n>  \n>  \ttree_difference |= diff;\n> -\tif (tree_difference == REV_TREE_DIFFERENT)\n> +\tif (tree_difference & REV_TREE_DIFFERENT)\n>  \t\tDIFF_OPT_SET(options, HAS_CHANGES);\n>  }\n>  \n> \n> But that essentially makes the conditional a noop (since we know we set\n> either NEW or OLD above and DIFFERENT is the union of those flags).\n> \n> I'm not sure I understand why file_add_remove() would ever want to avoid\n> setting HAS_CHANGES (certainly its companion file_change() always does).\n> This goes back to Junio's dd47aa3133 (try-to-simplify-commit: use\n> diff-tree --quiet machinery., 2007-03-14).\n> \n> Maybe I am missing something, but AFAICT this was always buggy. But\n> since it only affects adds and deletes, maybe nobody noticed? I'm also\n> not sure if it only causes a slowdown, or if this could cause us to\n> erroneously mark something as TREESAME which isn't (I _do_ think people\n> would have noticed that).\n\nAnswering my own question a little, there is a hint in the comment\na few lines above:\n\n  /*\n   * The goal is to get REV_TREE_NEW as the result only if the\n   * diff consists of all '+' (and no other changes), REV_TREE_OLD\n   * if the whole diff is removal of old data, and otherwise\n   * REV_TREE_DIFFERENT (of course if the trees are the same we\n   * want REV_TREE_SAME).\n   * That means that once we get to REV_TREE_DIFFERENT, we do not\n   * have to look any further.\n   */\n\nSo my patch above is breaking that. But it's not clear from that comment\nwhy we care about knowing the different between NEW, OLD, and DIFFERENT.\n\nGrepping around for REV_TREE_NEW and REV_TREE_OLD, I think the answer is\nin try_to_simplify_commit():\n\n     case REV_TREE_NEW:\n              if (revs->remove_empty_trees &&\n                  rev_same_tree_as_empty(revs, p)) {\n                      /* We are adding all the specified\n                       * paths from this parent, so the\n                       * history beyond this parent is not\n                       * interesting.  Remove its parents\n                       * (they are grandparents for us).\n                       * IOW, we pretend this parent is a\n                       * \"root\" commit.\n                       */\n                      if (parse_commit(p) < 0)\n                              die(\"cannot simplify commit %s (invalid %s)\",\n                                  oid_to_hex(&commit->object.oid),\n                                  oid_to_hex(&p->object.oid));\n                      p->parents = NULL;\n              }\n    /* fallthrough */\n    case REV_TREE_OLD:\n    case REV_TREE_DIFFERENT:\n\nSo when --remove-empty is not in effect (and it's not by default), we\ndon't care about OLD vs NEW, and we should be able to optimize further.\n\n-Peff\n"},{"id":"330362","messageId":"42cbcb4f-7f9d-df69-f55e-0ba42b931957@gmail.com","threadId":"46964","inReplyTo":"20171013142004.ocxpdkkbcxpi52yv@sigill.intra.peff.net","subject":"Re: git-clone causes out of memory","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2017-10-13T14:25:10Z","receivedAt":"2017-10-13T14:25:17Z","isPatch":false,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 10/13/2017 10:20 AM, Jeff King wrote:\n> On Fri, Oct 13, 2017 at 10:10:18AM -0400, Jeff King wrote:\n>\n>> Hmm. So this patch makes it go fast:\n>>\n>> diff --git a/revision.c b/revision.c\n>> index d167223e69..b52ea4e9d8 100644\n>> --- a/revision.c\n>> +++ b/revision.c\n>> @@ -409,7 +409,7 @@ static void file_add_remove(struct diff_options *options,\n>>   \tint diff = addremove == '+' ? REV_TREE_NEW : REV_TREE_OLD;\n>>   \n>>   \ttree_difference |= diff;\n>> -\tif (tree_difference == REV_TREE_DIFFERENT)\n>> +\tif (tree_difference & REV_TREE_DIFFERENT)\n>>   \t\tDIFF_OPT_SET(options, HAS_CHANGES);\n>>   }\n>>   \n>>\n>> But that essentially makes the conditional a noop (since we know we set\n>> either NEW or OLD above and DIFFERENT is the union of those flags).\n>>\n>> I'm not sure I understand why file_add_remove() would ever want to avoid\n>> setting HAS_CHANGES (certainly its companion file_change() always does).\n>> This goes back to Junio's dd47aa3133 (try-to-simplify-commit: use\n>> diff-tree --quiet machinery., 2007-03-14).\n>>\n>> Maybe I am missing something, but AFAICT this was always buggy. But\n>> since it only affects adds and deletes, maybe nobody noticed? I'm also\n>> not sure if it only causes a slowdown, or if this could cause us to\n>> erroneously mark something as TREESAME which isn't (I _do_ think people\n>> would have noticed that).\n> Answering my own question a little, there is a hint in the comment\n> a few lines above:\n>\n>    /*\n>     * The goal is to get REV_TREE_NEW as the result only if the\n>     * diff consists of all '+' (and no other changes), REV_TREE_OLD\n>     * if the whole diff is removal of old data, and otherwise\n>     * REV_TREE_DIFFERENT (of course if the trees are the same we\n>     * want REV_TREE_SAME).\n>     * That means that once we get to REV_TREE_DIFFERENT, we do not\n>     * have to look any further.\n>     */\n>\n> So my patch above is breaking that. But it's not clear from that comment\n> why we care about knowing the different between NEW, OLD, and DIFFERENT.\n>\n> Grepping around for REV_TREE_NEW and REV_TREE_OLD, I think the answer is\n> in try_to_simplify_commit():\n>\n>       case REV_TREE_NEW:\n>                if (revs->remove_empty_trees &&\n>                    rev_same_tree_as_empty(revs, p)) {\n>                        /* We are adding all the specified\n>                         * paths from this parent, so the\n>                         * history beyond this parent is not\n>                         * interesting.  Remove its parents\n>                         * (they are grandparents for us).\n>                         * IOW, we pretend this parent is a\n>                         * \"root\" commit.\n>                         */\n>                        if (parse_commit(p) < 0)\n>                                die(\"cannot simplify commit %s (invalid %s)\",\n>                                    oid_to_hex(&commit->object.oid),\n>                                    oid_to_hex(&p->object.oid));\n>                        p->parents = NULL;\n>                }\n>      /* fallthrough */\n>      case REV_TREE_OLD:\n>      case REV_TREE_DIFFERENT:\n>\n> So when --remove-empty is not in effect (and it's not by default), we\n> don't care about OLD vs NEW, and we should be able to optimize further.\n>\n> -Peff\n\nThis does appear to be the problem. The missing DIFF_OPT_HAS_CHANGES is \ncausing diff_can_quit_early() to return false. Due to the corner-case of \nthe bug it seems it will not be a huge performance improvement in most \ncases. Still worth fixing and I'm looking at your suggestions to try and \nlearn this area better.\n\nIt will speed up natural cases of someone adding or renaming a folder \nwith a lot of contents, including someone initializing a repository from \nan existing codebase. This git bomb case happens to add all of the \nfolders at once, which is why the performance is so noticeable. I use a \nversion of this \"exponential file growth\" for testing that increases the \nfolder-depth one commit at a time, so the contents are rather small in \nthe first \"add\", hence not noticing this issue.\n\nThanks,\n-Stolee\n"},{"id":"330363","messageId":"20171013142646.evapso5uxzvh2r2p@sigill.intra.peff.net","threadId":"46964","inReplyTo":"42cbcb4f-7f9d-df69-f55e-0ba42b931957@gmail.com","subject":"Re: git-clone causes out of memory","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-10-13T14:26:46Z","receivedAt":"2017-10-13T14:26:54Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Oct 13, 2017 at 10:25:10AM -0400, Derrick Stolee wrote:\n\n> This does appear to be the problem. The missing DIFF_OPT_HAS_CHANGES is\n> causing diff_can_quit_early() to return false. Due to the corner-case of the\n> bug it seems it will not be a huge performance improvement in most cases.\n> Still worth fixing and I'm looking at your suggestions to try and learn this\n> area better.\n\nYeah, I just timed some pathspec limits on linux.git, and it makes at\nbest a fraction of a percent improvement (but any improvement is well\nwithin run-to-run noise). Which is not surprising.\n\nI agree it's worth fixing, though.\n\n-Peff\n"},{"id":"330364","messageId":"c4b88c7e-7be3-6acb-71f9-fb0185dc0b7d@gmail.com","threadId":"46964","inReplyTo":"20171013142646.evapso5uxzvh2r2p@sigill.intra.peff.net","subject":"Re: git-clone causes out of memory","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2017-10-13T14:30:48Z","receivedAt":"2017-10-13T14:30:57Z","isPatch":false,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 10/13/2017 10:26 AM, Jeff King wrote:\n> On Fri, Oct 13, 2017 at 10:25:10AM -0400, Derrick Stolee wrote:\n>\n>> This does appear to be the problem. The missing DIFF_OPT_HAS_CHANGES is\n>> causing diff_can_quit_early() to return false. Due to the corner-case of the\n>> bug it seems it will not be a huge performance improvement in most cases.\n>> Still worth fixing and I'm looking at your suggestions to try and learn this\n>> area better.\n> Yeah, I just timed some pathspec limits on linux.git, and it makes at\n> best a fraction of a percent improvement (but any improvement is well\n> within run-to-run noise). Which is not surprising.\n>\n> I agree it's worth fixing, though.\n>\n> -Peff\n\nI just ran a first-level folder history in the Windows repository and \n`git rev-list HEAD -- windows >/dev/null` went from 0.43s to 0.04s. So, \na good percentage improvement but even on this enormous repo the bug \ndoesn't present an issue to human users.\n\nThanks,\n-Stolee\n"},{"id":"330365","messageId":"20171013152745.cgqt3qgvcngyr5ew@sigill.intra.peff.net","threadId":"46964","inReplyTo":"20171013142646.evapso5uxzvh2r2p@sigill.intra.peff.net","subject":"[PATCH] revision: quit pruning diff more quickly when possible","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-10-13T15:27:45Z","receivedAt":"2017-10-13T15:27:58Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Oct 13, 2017 at 10:26:46AM -0400, Jeff King wrote:\n\n> On Fri, Oct 13, 2017 at 10:25:10AM -0400, Derrick Stolee wrote:\n> \n> > This does appear to be the problem. The missing DIFF_OPT_HAS_CHANGES is\n> > causing diff_can_quit_early() to return false. Due to the corner-case of the\n> > bug it seems it will not be a huge performance improvement in most cases.\n> > Still worth fixing and I'm looking at your suggestions to try and learn this\n> > area better.\n> \n> Yeah, I just timed some pathspec limits on linux.git, and it makes at\n> best a fraction of a percent improvement (but any improvement is well\n> within run-to-run noise). Which is not surprising.\n> \n> I agree it's worth fixing, though.\n\nHere it is cleaned up and with a commit message. There's another case\nthat can be optimized, too: --remove-empty with an all-deletions commit.\nThat's probably even more obscure and pathological, but it was easy to\ncover in the same breath.\n\nI didn't bother making a perf script, since this really isn't indicative\nof real-world performance. If we wanted to do perf regression tests\nhere, I think the best path forward would be:\n\n  1. Make sure there the perf tests cover pathspecs (maybe in p0001?).\n\n  2. Make it easy to run the whole perf suite against a \"bomb\" repo.\n     This surely isn't the only slow thing of interest.\n\n-- >8 --\nSubject: revision: quit pruning diff more quickly when possible\n\nWhen the revision traversal machinery is given a pathspec,\nwe must compute the parent-diff for each commit to determine\nwhich ones are TREESAME. We set the QUICK diff flag to avoid\nlooking at more entries than we need; we really just care\nwhether there are any changes at all.\n\nBut there is one case where we want to know a bit more: if\n--remove-empty is set, we care about finding cases where the\nchange consists only of added entries (in which case we may\nprune the parent in try_to_simplify_commit()). To cover that\ncase, our file_add_remove() callback does not quit the diff\nupon seeing an added entry; it keeps looking for other types\nof entries.\n\nBut this means when --remove-empty is not set (and it is not\nby default), we compute more of the diff than is necessary.\nYou can see this in a pathological case where a commit adds\na very large number of entries, and we limit based on a\nbroad pathspec. E.g.:\n\n  perl -e '\n    chomp(my $blob = `git hash-object -w --stdin </dev/null`);\n    for my $a (1..1000) {\n      for my $b (1..1000) {\n        print \"100644 $blob\\t$a/$b\\n\";\n      }\n    }\n  ' | git update-index --index-info\n  git commit -qm add\n\n  git rev-list HEAD -- .\n\nThis case takes about 100ms now, but after this patch only\nneeds 6ms. That's not a huge improvement, but it's easy to\nget and it protects us against even more pathological cases\n(e.g., going from 1 million to 10 million files would take\nten times as long with the current code, but not increase at\nall after this patch).\n\nThis is reported to minorly speed-up pathspec limiting in\nreal world repositories (like the 100-million-file Windows\nrepository), but probably won't make a noticeable difference\noutside of pathological setups.\n\nThis patch actually covers the case without --remove-empty,\nand the case where we see only deletions. See the in-code\ncomment for details.\n\nNote that we have to add a new member to the diff_options\nstruct so that our callback can see the value of\nrevs->remove_empty_trees. This callback parameter could be\npassed to the \"add_remove\" and \"change\" callbacks, but\nthere's not much point. They already receive the\ndiff_options struct, and doing it this way avoids having to\nupdate the function signature of the other callbacks\n(arguably the format_callback and output_prefix functions\ncould benefit from the same simplification).\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n diff.h     |  1 +\n revision.c | 16 +++++++++++++---\n 2 files changed, 14 insertions(+), 3 deletions(-)\n\ndiff --git a/diff.h b/diff.h\nindex 7dcfcfbef7..4a34d256f1 100644\n--- a/diff.h\n+++ b/diff.h\n@@ -180,6 +180,7 @@ struct diff_options {\n \tpathchange_fn_t pathchange;\n \tchange_fn_t change;\n \tadd_remove_fn_t add_remove;\n+\tvoid *change_fn_data;\n \tdiff_format_fn_t format_callback;\n \tvoid *format_callback_data;\n \tdiff_prefix_fn_t output_prefix;\ndiff --git a/revision.c b/revision.c\nindex 8fd222f3bf..a3f245e2cc 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -399,8 +399,16 @@ static struct commit *one_relevant_parent(const struct rev_info *revs,\n  * if the whole diff is removal of old data, and otherwise\n  * REV_TREE_DIFFERENT (of course if the trees are the same we\n  * want REV_TREE_SAME).\n- * That means that once we get to REV_TREE_DIFFERENT, we do not\n- * have to look any further.\n+ *\n+ * The only time we care about the distinction is when\n+ * remove_empty_trees is in effect, in which case we care only about\n+ * whether the whole change is REV_TREE_NEW, or if there's another type\n+ * of change. Which means we can stop the diff early in either of these\n+ * cases:\n+ *\n+ *   1. We're not using remove_empty_trees at all.\n+ *\n+ *   2. We saw anything except REV_TREE_NEW.\n  */\n static int tree_difference = REV_TREE_SAME;\n \n@@ -411,9 +419,10 @@ static void file_add_remove(struct diff_options *options,\n \t\t    const char *fullpath, unsigned dirty_submodule)\n {\n \tint diff = addremove == '+' ? REV_TREE_NEW : REV_TREE_OLD;\n+\tstruct rev_info *revs = options->change_fn_data;\n \n \ttree_difference |= diff;\n-\tif (tree_difference == REV_TREE_DIFFERENT)\n+\tif (!revs->remove_empty_trees || tree_difference != REV_TREE_NEW)\n \t\tDIFF_OPT_SET(options, HAS_CHANGES);\n }\n \n@@ -1351,6 +1360,7 @@ void init_revisions(struct rev_info *revs, const char *prefix)\n \tDIFF_OPT_SET(&revs->pruning, QUICK);\n \trevs->pruning.add_remove = file_add_remove;\n \trevs->pruning.change = file_change;\n+\trevs->pruning.change_fn_data = revs;\n \trevs->sort_order = REV_SORT_IN_GRAPH_ORDER;\n \trevs->dense = 1;\n \trevs->prefix = prefix;\n-- \n2.15.0.rc1.395.ga4290b5804\n\n"},{"id":"330366","messageId":"e2db8086-29af-a8bc-1e12-8642e430fcb3@gmail.com","threadId":"46964","inReplyTo":"20171013152745.cgqt3qgvcngyr5ew@sigill.intra.peff.net","subject":"Re: [PATCH] revision: quit pruning diff more quickly when possible","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2017-10-13T15:37:50Z","receivedAt":"2017-10-13T15:37:59Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 10/13/2017 11:27 AM, Jeff King wrote:\n> On Fri, Oct 13, 2017 at 10:26:46AM -0400, Jeff King wrote:\n>\n>> On Fri, Oct 13, 2017 at 10:25:10AM -0400, Derrick Stolee wrote:\n>>\n>>> This does appear to be the problem. The missing DIFF_OPT_HAS_CHANGES is\n>>> causing diff_can_quit_early() to return false. Due to the corner-case of the\n>>> bug it seems it will not be a huge performance improvement in most cases.\n>>> Still worth fixing and I'm looking at your suggestions to try and learn this\n>>> area better.\n>> Yeah, I just timed some pathspec limits on linux.git, and it makes at\n>> best a fraction of a percent improvement (but any improvement is well\n>> within run-to-run noise). Which is not surprising.\n>>\n>> I agree it's worth fixing, though.\n> Here it is cleaned up and with a commit message. There's another case\n> that can be optimized, too: --remove-empty with an all-deletions commit.\n> That's probably even more obscure and pathological, but it was easy to\n> cover in the same breath.\n>\n> I didn't bother making a perf script, since this really isn't indicative\n> of real-world performance. If we wanted to do perf regression tests\n> here, I think the best path forward would be:\n>\n>    1. Make sure there the perf tests cover pathspecs (maybe in p0001?).\n>\n>    2. Make it easy to run the whole perf suite against a \"bomb\" repo.\n>       This surely isn't the only slow thing of interest.\n>\n> -- >8 --\n> Subject: revision: quit pruning diff more quickly when possible\n>\n> When the revision traversal machinery is given a pathspec,\n> we must compute the parent-diff for each commit to determine\n> which ones are TREESAME. We set the QUICK diff flag to avoid\n> looking at more entries than we need; we really just care\n> whether there are any changes at all.\n>\n> But there is one case where we want to know a bit more: if\n> --remove-empty is set, we care about finding cases where the\n> change consists only of added entries (in which case we may\n> prune the parent in try_to_simplify_commit()). To cover that\n> case, our file_add_remove() callback does not quit the diff\n> upon seeing an added entry; it keeps looking for other types\n> of entries.\n>\n> But this means when --remove-empty is not set (and it is not\n> by default), we compute more of the diff than is necessary.\n> You can see this in a pathological case where a commit adds\n> a very large number of entries, and we limit based on a\n> broad pathspec. E.g.:\n>\n>    perl -e '\n>      chomp(my $blob = `git hash-object -w --stdin </dev/null`);\n>      for my $a (1..1000) {\n>        for my $b (1..1000) {\n>          print \"100644 $blob\\t$a/$b\\n\";\n>        }\n>      }\n>    ' | git update-index --index-info\n>    git commit -qm add\n>\n>    git rev-list HEAD -- .\n>\n> This case takes about 100ms now, but after this patch only\n> needs 6ms. That's not a huge improvement, but it's easy to\n> get and it protects us against even more pathological cases\n> (e.g., going from 1 million to 10 million files would take\n> ten times as long with the current code, but not increase at\n> all after this patch).\n>\n> This is reported to minorly speed-up pathspec limiting in\n> real world repositories (like the 100-million-file Windows\n> repository), but probably won't make a noticeable difference\n> outside of pathological setups.\n>\n> This patch actually covers the case without --remove-empty,\n> and the case where we see only deletions. See the in-code\n> comment for details.\n>\n> Note that we have to add a new member to the diff_options\n> struct so that our callback can see the value of\n> revs->remove_empty_trees. This callback parameter could be\n> passed to the \"add_remove\" and \"change\" callbacks, but\n> there's not much point. They already receive the\n> diff_options struct, and doing it this way avoids having to\n> update the function signature of the other callbacks\n> (arguably the format_callback and output_prefix functions\n> could benefit from the same simplification).\n>\n> Signed-off-by: Jeff King <peff@peff.net>\n> ---\n>   diff.h     |  1 +\n>   revision.c | 16 +++++++++++++---\n>   2 files changed, 14 insertions(+), 3 deletions(-)\n>\n> diff --git a/diff.h b/diff.h\n> index 7dcfcfbef7..4a34d256f1 100644\n> --- a/diff.h\n> +++ b/diff.h\n> @@ -180,6 +180,7 @@ struct diff_options {\n>   \tpathchange_fn_t pathchange;\n>   \tchange_fn_t change;\n>   \tadd_remove_fn_t add_remove;\n> +\tvoid *change_fn_data;\n>   \tdiff_format_fn_t format_callback;\n>   \tvoid *format_callback_data;\n>   \tdiff_prefix_fn_t output_prefix;\n> diff --git a/revision.c b/revision.c\n> index 8fd222f3bf..a3f245e2cc 100644\n> --- a/revision.c\n> +++ b/revision.c\n> @@ -399,8 +399,16 @@ static struct commit *one_relevant_parent(const struct rev_info *revs,\n>    * if the whole diff is removal of old data, and otherwise\n>    * REV_TREE_DIFFERENT (of course if the trees are the same we\n>    * want REV_TREE_SAME).\n> - * That means that once we get to REV_TREE_DIFFERENT, we do not\n> - * have to look any further.\n> + *\n> + * The only time we care about the distinction is when\n> + * remove_empty_trees is in effect, in which case we care only about\n> + * whether the whole change is REV_TREE_NEW, or if there's another type\n> + * of change. Which means we can stop the diff early in either of these\n> + * cases:\n> + *\n> + *   1. We're not using remove_empty_trees at all.\n> + *\n> + *   2. We saw anything except REV_TREE_NEW.\n>    */\n>   static int tree_difference = REV_TREE_SAME;\n>   \n> @@ -411,9 +419,10 @@ static void file_add_remove(struct diff_options *options,\n>   \t\t    const char *fullpath, unsigned dirty_submodule)\n>   {\n>   \tint diff = addremove == '+' ? REV_TREE_NEW : REV_TREE_OLD;\n> +\tstruct rev_info *revs = options->change_fn_data;\n>   \n>   \ttree_difference |= diff;\n> -\tif (tree_difference == REV_TREE_DIFFERENT)\n> +\tif (!revs->remove_empty_trees || tree_difference != REV_TREE_NEW)\n>   \t\tDIFF_OPT_SET(options, HAS_CHANGES);\n>   }\n>   \n> @@ -1351,6 +1360,7 @@ void init_revisions(struct rev_info *revs, const char *prefix)\n>   \tDIFF_OPT_SET(&revs->pruning, QUICK);\n>   \trevs->pruning.add_remove = file_add_remove;\n>   \trevs->pruning.change = file_change;\n> +\trevs->pruning.change_fn_data = revs;\n>   \trevs->sort_order = REV_SORT_IN_GRAPH_ORDER;\n>   \trevs->dense = 1;\n>   \trevs->prefix = prefix;\n\nThanks, Peff. This patch looks good to me.\n\nI tried a few other things like adding a flag DIFF_OPT_HAS_ANY_CHANGE \nnext to DIFF_OPT_HAS_CHANGES that we could check in \ndiff_can_quit_early() but it had side-effects that broke existing tests. \nFrom this exploration, it does seem necessary to be aware of \n'remove_empty_trees'.\n\nThanks,\n-Stolee\n"},{"id":"330367","messageId":"20171013154401.hwwvl2xi5quv2sg3@sigill.intra.peff.net","threadId":"46964","inReplyTo":"e2db8086-29af-a8bc-1e12-8642e430fcb3@gmail.com","subject":"Re: [PATCH] revision: quit pruning diff more quickly when possible","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-10-13T15:44:01Z","receivedAt":"2017-10-13T15:44:08Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Oct 13, 2017 at 11:37:50AM -0400, Derrick Stolee wrote:\n\n> Thanks, Peff. This patch looks good to me.\n> \n> I tried a few other things like adding a flag DIFF_OPT_HAS_ANY_CHANGE next\n> to DIFF_OPT_HAS_CHANGES that we could check in diff_can_quit_early() but it\n> had side-effects that broke existing tests. From this exploration, it does\n> seem necessary to be aware of 'remove_empty_trees'.\n\nKeep in mind that the regular diff_change callbacks already handle this\ncase[1].\n\nThe file_change callbacks are specific to the revision machinery's\npruning diff, and intentionally hold back the HAS_CHANGES flag.\n\n-Peff\n\n[1] I tried \"git diff-tree --root -r --quiet 45546f17e\" on the bomb\n    repo, and it went quickly. Dropping --quiet makes it take a really\n    long time.\n"},{"id":"330390","messageId":"xmqqy3oetpnx.fsf@gitster.mtv.corp.google.com","threadId":"46964","inReplyTo":"20171013152745.cgqt3qgvcngyr5ew@sigill.intra.peff.net","subject":"Re: [PATCH] revision: quit pruning diff more quickly when possible","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-10-14T02:43:30Z","receivedAt":"2017-10-14T02:43:38Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> Here it is cleaned up and with a commit message. There's another case\n> that can be optimized, too: --remove-empty with an all-deletions commit.\n> That's probably even more obscure and pathological, but it was easy to\n> cover in the same breath.\n\nThis one looks good.  It appears that again you guys had all the fun\nwhile I was offline ;-).  And I am happy to see that we didn't veer\nin the direction to optimize for a wrong case by keeping track of\nwhat trees we already saw and things like that, of course.\n\nThanks.\n\n> Subject: revision: quit pruning diff more quickly when possible\n>\n> When the revision traversal machinery is given a pathspec,\n> we must compute the parent-diff for each commit to determine\n> which ones are TREESAME. We set the QUICK diff flag to avoid\n> looking at more entries than we need; we really just care\n> whether there are any changes at all.\n>\n> But there is one case where we want to know a bit more: if\n> --remove-empty is set, we care about finding cases where the\n> change consists only of added entries (in which case we may\n> prune the parent in try_to_simplify_commit()). To cover that\n> case, our file_add_remove() callback does not quit the diff\n> upon seeing an added entry; it keeps looking for other types\n> of entries.\n>\n> But this means when --remove-empty is not set (and it is not\n> by default), we compute more of the diff than is necessary.\n> You can see this in a pathological case where a commit adds\n> a very large number of entries, and we limit based on a\n> broad pathspec. E.g.:\n>\n>   perl -e '\n>     chomp(my $blob = `git hash-object -w --stdin </dev/null`);\n>     for my $a (1..1000) {\n>       for my $b (1..1000) {\n>         print \"100644 $blob\\t$a/$b\\n\";\n>       }\n>     }\n>   ' | git update-index --index-info\n>   git commit -qm add\n>\n>   git rev-list HEAD -- .\n>\n> This case takes about 100ms now, but after this patch only\n> needs 6ms. That's not a huge improvement, but it's easy to\n> get and it protects us against even more pathological cases\n> (e.g., going from 1 million to 10 million files would take\n> ten times as long with the current code, but not increase at\n> all after this patch).\n>\n> This is reported to minorly speed-up pathspec limiting in\n> real world repositories (like the 100-million-file Windows\n> repository), but probably won't make a noticeable difference\n> outside of pathological setups.\n>\n> This patch actually covers the case without --remove-empty,\n> and the case where we see only deletions. See the in-code\n> comment for details.\n>\n> Note that we have to add a new member to the diff_options\n> struct so that our callback can see the value of\n> revs->remove_empty_trees. This callback parameter could be\n> passed to the \"add_remove\" and \"change\" callbacks, but\n> there's not much point. They already receive the\n> diff_options struct, and doing it this way avoids having to\n> update the function signature of the other callbacks\n> (arguably the format_callback and output_prefix functions\n> could benefit from the same simplification).\n>\n> Signed-off-by: Jeff King <peff@peff.net>\n> ---\n>  diff.h     |  1 +\n>  revision.c | 16 +++++++++++++---\n>  2 files changed, 14 insertions(+), 3 deletions(-)\n>\n> diff --git a/diff.h b/diff.h\n> index 7dcfcfbef7..4a34d256f1 100644\n> --- a/diff.h\n> +++ b/diff.h\n> @@ -180,6 +180,7 @@ struct diff_options {\n>  \tpathchange_fn_t pathchange;\n>  \tchange_fn_t change;\n>  \tadd_remove_fn_t add_remove;\n> +\tvoid *change_fn_data;\n>  \tdiff_format_fn_t format_callback;\n>  \tvoid *format_callback_data;\n>  \tdiff_prefix_fn_t output_prefix;\n> diff --git a/revision.c b/revision.c\n> index 8fd222f3bf..a3f245e2cc 100644\n> --- a/revision.c\n> +++ b/revision.c\n> @@ -399,8 +399,16 @@ static struct commit *one_relevant_parent(const struct rev_info *revs,\n>   * if the whole diff is removal of old data, and otherwise\n>   * REV_TREE_DIFFERENT (of course if the trees are the same we\n>   * want REV_TREE_SAME).\n> - * That means that once we get to REV_TREE_DIFFERENT, we do not\n> - * have to look any further.\n> + *\n> + * The only time we care about the distinction is when\n> + * remove_empty_trees is in effect, in which case we care only about\n> + * whether the whole change is REV_TREE_NEW, or if there's another type\n> + * of change. Which means we can stop the diff early in either of these\n> + * cases:\n> + *\n> + *   1. We're not using remove_empty_trees at all.\n> + *\n> + *   2. We saw anything except REV_TREE_NEW.\n>   */\n>  static int tree_difference = REV_TREE_SAME;\n>  \n> @@ -411,9 +419,10 @@ static void file_add_remove(struct diff_options *options,\n>  \t\t    const char *fullpath, unsigned dirty_submodule)\n>  {\n>  \tint diff = addremove == '+' ? REV_TREE_NEW : REV_TREE_OLD;\n> +\tstruct rev_info *revs = options->change_fn_data;\n>  \n>  \ttree_difference |= diff;\n> -\tif (tree_difference == REV_TREE_DIFFERENT)\n> +\tif (!revs->remove_empty_trees || tree_difference != REV_TREE_NEW)\n>  \t\tDIFF_OPT_SET(options, HAS_CHANGES);\n>  }\n>  \n> @@ -1351,6 +1360,7 @@ void init_revisions(struct rev_info *revs, const char *prefix)\n>  \tDIFF_OPT_SET(&revs->pruning, QUICK);\n>  \trevs->pruning.add_remove = file_add_remove;\n>  \trevs->pruning.change = file_change;\n> +\trevs->pruning.change_fn_data = revs;\n>  \trevs->sort_order = REV_SORT_IN_GRAPH_ORDER;\n>  \trevs->dense = 1;\n>  \trevs->prefix = prefix;\n"}]}