{"thread":{"id":"41997","subject":"Re: 0 bot for Git","startedAt":"2016-04-12T04:29:59Z","lastAt":"2016-04-26T11:35:50Z","messageCount":41,"participants":["Stefan Beller","Greg KH","Matthieu Moy","Duy Nguyen","Philip Li","Junio C Hamano","Lars Schneider","Fengguang Wu","Christian Couder","Jeff King","Michael Haggerty","Johannes Schindelin","SZEDER Gábor"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"283165","messageId":"CAGZ79kZOx8ehAB-=Frjgde2CDo_vwoVzQNizJinf4LLXek5PSQ@mail.gmail.com","threadId":"41997","inReplyTo":"CAGZ79kYWGFN1W0_y72-V6M3n4WLgtLPzs22bWgs1ObCCDt5BfQ@mail.gmail.com","subject":"Re: 0 bot for Git","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2016-04-12T04:29:59Z","receivedAt":"2016-04-12T04:29:59Z","isPatch":false,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"Resending as plain text. (I need to tame my mobile)\n\nOn Mon, Apr 11, 2016 at 7:51 AM, Stefan Beller <sbeller@google.com> wrote:\n> Hi Greg,\n>\n> Thanks for your talk at the Git Merge 2016!\n> The Git community uses the same workflow as the kernel. So we may be\n> interested in the 0 bot which could compile and test each patch on the list.\n> Could you put us in touch with the authors/maintainers of said tool?\n>\n> Unlike the kernel we would not need hardware testing and we're low traffic\n> compared to the kernel, which would make it easier to set it up.\n>\n> A healthier Git would help the kernel long term as well.\n>\n> Thanks,\n> Stefan\n"},{"id":"283169","messageId":"20160412064111.GA22157@kroah.com","threadId":"41997","inReplyTo":"CAGZ79kZOx8ehAB-=Frjgde2CDo_vwoVzQNizJinf4LLXek5PSQ@mail.gmail.com","subject":"Re: 0 bot for Git","fromName":"Greg KH","fromEmail":"gregkh@linuxfoundation.org","sentAt":"2016-04-12T06:41:11Z","receivedAt":"2016-04-12T06:41:11Z","isPatch":false,"sender":{"key":"gregkh@linuxfoundation.org","avatar":"https://gravatar.com/avatar/e6d9136f6e3bdcb59f0e5fd15565f382da42523d273824958b9e23e73cf38e04?d=mp&s=160"},"body":"On Mon, Apr 11, 2016 at 09:29:59PM -0700, Stefan Beller wrote:\n> Resending as plain text. (I need to tame my mobile)\n> \n> On Mon, Apr 11, 2016 at 7:51 AM, Stefan Beller <sbeller@google.com> wrote:\n> > Hi Greg,\n> >\n> > Thanks for your talk at the Git Merge 2016!\n> > The Git community uses the same workflow as the kernel. So we may be\n> > interested in the 0 bot which could compile and test each patch on the list.\n> > Could you put us in touch with the authors/maintainers of said tool?\n> >\n> > Unlike the kernel we would not need hardware testing and we're low traffic\n> > compared to the kernel, which would make it easier to set it up.\n\nWe don't get much, if any, real hardware testing from the 0-day bot,\nit's just lots and lots of builds and static testing tools.\n\nYou can reach the developers of it at:\n\tkbuild test robot <lkp@intel.com>\n\nHope this helps,\n\ngreg k-h\n"},{"id":"283171","messageId":"vpq60vnl28b.fsf@anie.imag.fr","threadId":"41997","inReplyTo":"CAGZ79kZOx8ehAB-=Frjgde2CDo_vwoVzQNizJinf4LLXek5PSQ@mail.gmail.com","subject":"Re: 0 bot for Git","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@grenoble-inp.fr","sentAt":"2016-04-12T07:23:32Z","receivedAt":"2016-04-12T07:23:32Z","isPatch":false,"sender":{"key":"matthieu.moy@grenoble-inp.fr","avatar":"https://gravatar.com/avatar/72c8a2705971a25dfaff23cece15130d405685845d911aedd5667ace277f3fc5?d=mp&s=160"},"body":"Stefan Beller <sbeller@google.com> writes:\n\n> Hi Greg,\n>\n> Thanks for your talk at the Git Merge 2016!\n> The Git community uses the same workflow as the kernel. So we may be\n> interested in the 0 bot which could compile and test each patch on the list.\n\nIn the case of Git, we already have Travis-CI that can do rather\nthorough testing automatically (run the complete testsuite on a clean\nmachine for several configurations). You get the benefit from it only if\nyou use GitHub pull-requests today. It would be interesting to have a\nbot watch the list, apply patches and push to a travis-enabled fork of\ngit.git on GitHub to get the same benefit when posting emails directly\nto the list.\n\n-- \nMatthieu Moy\nhttp://www-verimag.imag.fr/~moy/\n"},{"id":"283175","messageId":"CACsJy8DiCw_yZNp7st-qVA7zYEHww=ae5Q=uKVzBhAfU8akR7Q@mail.gmail.com","threadId":"41997","inReplyTo":"CAGZ79kZOx8ehAB-=Frjgde2CDo_vwoVzQNizJinf4LLXek5PSQ@mail.gmail.com","subject":"Re: 0 bot for Git","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2016-04-12T09:42:18Z","receivedAt":"2016-04-12T09:42:18Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"> On Mon, Apr 11, 2016 at 7:51 AM, Stefan Beller <sbeller@google.com> wrote:\n>> Hi Greg,\n>>\n>> Thanks for your talk at the Git Merge 2016!\n\nHuh? It already happened?? Any interesting summary to share with us?\n-- \nDuy\n"},{"id":"283187","messageId":"CAGZ79kaLQWVdehMu4nas6UBpCxnAB_-p=xPGH=aueMZXkGK_2Q@mail.gmail.com","threadId":"41997","inReplyTo":"vpq60vnl28b.fsf@anie.imag.fr","subject":"Re: 0 bot for Git","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2016-04-12T14:52:02Z","receivedAt":"2016-04-12T14:52:02Z","isPatch":false,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Tue, Apr 12, 2016 at 12:23 AM, Matthieu Moy\n<Matthieu.Moy@grenoble-inp.fr> wrote:\n> Stefan Beller <sbeller@google.com> writes:\n>\n>> Hi Greg,\n>>\n>> Thanks for your talk at the Git Merge 2016!\n>> The Git community uses the same workflow as the kernel. So we may be\n>> interested in the 0 bot which could compile and test each patch on the list.\n>\n> In the case of Git, we already have Travis-CI that can do rather\n> thorough testing automatically (run the complete testsuite on a clean\n> machine for several configurations). You get the benefit from it only if\n> you use GitHub pull-requests today.\n\nBut who uses that? (Not a lot of old-timers here, that's for sure)\n\n> It would be interesting to have a\n> bot watch the list, apply patches and push to a travis-enabled fork of\n> git.git on GitHub to get the same benefit when posting emails directly\n> to the list.\n\nThat is better (and probably more work) than what I had in mind.\nIIUC the 0 bot can grab a patch from a mailing list and apply it to a\nbase (either the real base as encoded in the patch or a best guess)\nand then run \"make\".\n\nAt least that's how I understand the kernel setup. So my naive thought\nis that the 0 bot maintainer \"only\" needs to add another mailing list\nto the watch list of the 0 bot?\n\nStefan\n\n>\n> --\n> Matthieu Moy\n> http://www-verimag.imag.fr/~moy/\n"},{"id":"283188","messageId":"CAGZ79kZzdioQRFEmgTGOOdLQ-Ov-tWmgi1dLhHPDVzDb+Py2RQ@mail.gmail.com","threadId":"41997","inReplyTo":"CACsJy8DiCw_yZNp7st-qVA7zYEHww=ae5Q=uKVzBhAfU8akR7Q@mail.gmail.com","subject":"Re: 0 bot for Git","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2016-04-12T14:59:46Z","receivedAt":"2016-04-12T14:59:46Z","isPatch":false,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Tue, Apr 12, 2016 at 2:42 AM, Duy Nguyen <pclouds@gmail.com> wrote:\n>> On Mon, Apr 11, 2016 at 7:51 AM, Stefan Beller <sbeller@google.com> wrote:\n>>> Hi Greg,\n>>>\n>>> Thanks for your talk at the Git Merge 2016!\n>\n> Huh? It already happened?? Any interesting summary to share with us?\n\nSummary from the contributors summit:\n* encoding was a huge discussion point:\n  If you use Git on a case sensitive file system, you may have\n  branches \"foo\" and \"FOO\" and it just works. Once you switch over\n  to another file system this may break horribly.\n\n  The discussion revealed lots more of these points in Git and then a\n  discussion on fundamentals sparked, on whether we want to go by the lowest\n  common denominator or treat any system special or have it as is, but less\n  broken.\n\n* We still don't know how to handle large repositories\n  (I am working on submodules but that is too vague and not solving\n  the actual problem, people really want there\n    * large files\n    * large trees\n   [* lots of commits (i.e. lots of objects)]\n   in one repo and it should work just as fast.)\n\nThat was my main take away.\n\nStefan\n\n> --\n> Duy\n"},{"id":"283189","messageId":"20160412151509.GA14522@intel.com","threadId":"41997","inReplyTo":"CAGZ79kaLQWVdehMu4nas6UBpCxnAB_-p=xPGH=aueMZXkGK_2Q@mail.gmail.com","subject":"Re: 0 bot for Git","fromName":"Philip Li","fromEmail":"philip.li@intel.com","sentAt":"2016-04-12T15:15:10Z","receivedAt":"2016-04-12T15:15:10Z","isPatch":false,"sender":{"key":"philip.li@intel.com","avatar":null},"body":"On Tue, Apr 12, 2016 at 07:52:02AM -0700, Stefan Beller wrote:\n> On Tue, Apr 12, 2016 at 12:23 AM, Matthieu Moy\n> <Matthieu.Moy@grenoble-inp.fr> wrote:\n> > Stefan Beller <sbeller@google.com> writes:\n> >\n> >> Hi Greg,\n> >>\n> >> Thanks for your talk at the Git Merge 2016!\n> >> The Git community uses the same workflow as the kernel. So we may be\n> >> interested in the 0 bot which could compile and test each patch on the list.\n> >\n> > In the case of Git, we already have Travis-CI that can do rather\n> > thorough testing automatically (run the complete testsuite on a clean\n> > machine for several configurations). You get the benefit from it only if\n> > you use GitHub pull-requests today.\n> \n> But who uses that? (Not a lot of old-timers here, that's for sure)\n> \n> > It would be interesting to have a\n> > bot watch the list, apply patches and push to a travis-enabled fork of\n> > git.git on GitHub to get the same benefit when posting emails directly\n> > to the list.\n> \n> That is better (and probably more work) than what I had in mind.\n> IIUC the 0 bot can grab a patch from a mailing list and apply it to a\n> base (either the real base as encoded in the patch or a best guess)\n> and then run \"make\".\n> \n> At least that's how I understand the kernel setup. So my naive thought\n> is that the 0 bot maintainer \"only\" needs to add another mailing list\n> to the watch list of the 0 bot?\n\nyes, this is feasible in 0 bot, though it requires some refactoring to current\nmailing list logic which focuses on linux kernel like using guess work to find\nout which maintainer tree to apply a patch on by following a set of rules. \nAfter the current proposal to add extra info to git-format-patch is accepted,\nthis can be more smooth.\n\n> \n> Stefan\n> \n> >\n> > --\n> > Matthieu Moy\n> > http://www-verimag.imag.fr/~moy/\n> \n"},{"id":"283215","messageId":"vpqoa9ea7vx.fsf@anie.imag.fr","threadId":"41997","inReplyTo":"CAGZ79kaLQWVdehMu4nas6UBpCxnAB_-p=xPGH=aueMZXkGK_2Q@mail.gmail.com","subject":"Re: 0 bot for Git","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@grenoble-inp.fr","sentAt":"2016-04-12T20:29:06Z","receivedAt":"2016-04-12T20:29:06Z","isPatch":false,"sender":{"key":"matthieu.moy@grenoble-inp.fr","avatar":"https://gravatar.com/avatar/72c8a2705971a25dfaff23cece15130d405685845d911aedd5667ace277f3fc5?d=mp&s=160"},"body":"Stefan Beller <sbeller@google.com> writes:\n\n> On Tue, Apr 12, 2016 at 12:23 AM, Matthieu Moy\n> <Matthieu.Moy@grenoble-inp.fr> wrote:\n>> Stefan Beller <sbeller@google.com> writes:\n>>\n>>> Hi Greg,\n>>>\n>>> Thanks for your talk at the Git Merge 2016!\n>>> The Git community uses the same workflow as the kernel. So we may be\n>>> interested in the 0 bot which could compile and test each patch on the list.\n>>\n>> In the case of Git, we already have Travis-CI that can do rather\n>> thorough testing automatically (run the complete testsuite on a clean\n>> machine for several configurations). You get the benefit from it only if\n>> you use GitHub pull-requests today.\n>\n> But who uses that? (Not a lot of old-timers here, that's for sure)\n\nNot many people clearly. I sometimes do, but SubmitGit as it is today\ndoesn't (yet?) beat \"git send-email\" for me. In a perfect world where I\ncould just ask SubmitGit \"please wait for Travis to complete, if it\npasses then send to the list, otherwise email me\", I would use it more.\n\nBut my point wasn't to say \"we already have everything we need\", but\nrather \"we already have part of the solution, so an ideal complete\nsolution could integrate with it\".\n\n>> It would be interesting to have a\n>> bot watch the list, apply patches and push to a travis-enabled fork of\n>> git.git on GitHub to get the same benefit when posting emails directly\n>> to the list.\n>\n> That is better (and probably more work) than what I had in mind.\n> IIUC the 0 bot can grab a patch from a mailing list and apply it to a\n> base (either the real base as encoded in the patch or a best guess)\n> and then run \"make\".\n\nI don't know how 0 bot solves this, but the obvious issue with this\napproach is to allow dealing with someone sending a patch like\n\n+++ Makefile\n--- Makefile\n+all:\n+\trm -fr $(HOME); sudo rm -fr /\n\nto the list. One thing that Travis gives us for free is isolation:\nmalicious code in the build cannot break the bot, only the build itself.\n\n-- \nMatthieu Moy\nhttp://www-verimag.imag.fr/~moy/\n"},{"id":"283217","messageId":"xmqqmvoypn7g.fsf@gitster.mtv.corp.google.com","threadId":"41997","inReplyTo":"vpqoa9ea7vx.fsf@anie.imag.fr","subject":"Re: 0 bot for Git","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-04-12T20:49:07Z","receivedAt":"2016-04-12T20:49:07Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Matthieu Moy <Matthieu.Moy@grenoble-inp.fr> writes:\n\n> But my point wasn't to say \"we already have everything we need\", but\n> rather \"we already have part of the solution, so an ideal complete\n> solution could integrate with it\".\n\nYes.  That is a good direction to go.\n\nThey may already have part of the solution, and their half may be\nbetter than what we have, in which case we may want to switch, but\nif what we have already works well there is no need to.\n\n> I don't know how 0 bot solves this, but the obvious issue with this\n> approach is to allow dealing with someone sending a patch like\n>\n> +++ Makefile\n> --- Makefile\n> +all:\n> +\trm -fr $(HOME); sudo rm -fr /\n>\n> to the list. One thing that Travis gives us for free is isolation:\n> malicious code in the build cannot break the bot, only the build\n> itself.\n\nTrue, presumably the Travis integration already solves that part, so\nI suspect it is just the matter of setting up:\n\n - a fork of git.git and have Travis monitor any and all new\n   branches;\n\n - a bot that scans the list traffic, applies each series it sees to\n   a branch dedicated for that series and pushes to the above fork.\n\nisn't it?\n"},{"id":"283293","messageId":"vpqegaa9i89.fsf@anie.imag.fr","threadId":"41997","inReplyTo":"xmqqmvoypn7g.fsf@gitster.mtv.corp.google.com","subject":"Re: 0 bot for Git","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@grenoble-inp.fr","sentAt":"2016-04-13T05:43:18Z","receivedAt":"2016-04-13T05:43:18Z","isPatch":false,"sender":{"key":"matthieu.moy@grenoble-inp.fr","avatar":"https://gravatar.com/avatar/72c8a2705971a25dfaff23cece15130d405685845d911aedd5667ace277f3fc5?d=mp&s=160"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Matthieu Moy <Matthieu.Moy@grenoble-inp.fr> writes:\n>\n> True, presumably the Travis integration already solves that part, so\n> I suspect it is just the matter of setting up:\n>\n>  - a fork of git.git and have Travis monitor any and all new\n>    branches;\n>\n>  - a bot that scans the list traffic, applies each series it sees to\n>    a branch dedicated for that series and pushes to the above fork.\n\n... and to make it really useful: a way to get a notification email sent\non-list or at least to the submitter as a reply to the patch series.\nJust having a web interface somewhere that knows how broken the code is\nwould not be that useful.\n\n-- \nMatthieu Moy\nhttp://www-verimag.imag.fr/~moy/\n"},{"id":"283297","messageId":"88CF8CB5-4105-4D0C-8064-D66092169111@gmail.com","threadId":"41997","inReplyTo":"xmqqmvoypn7g.fsf@gitster.mtv.corp.google.com","subject":"Re: 0 bot for Git","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-04-13T06:11:54Z","receivedAt":"2016-04-13T06:11:54Z","isPatch":false,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\nOn 12 Apr 2016, at 22:49, Junio C Hamano <gitster@pobox.com> wrote:\n\n> Matthieu Moy <Matthieu.Moy@grenoble-inp.fr> writes:\n> \n>> But my point wasn't to say \"we already have everything we need\", but\n>> rather \"we already have part of the solution, so an ideal complete\n>> solution could integrate with it\".\n> \n> Yes.  That is a good direction to go.\n> \n> They may already have part of the solution, and their half may be\n> better than what we have, in which case we may want to switch, but\n> if what we have already works well there is no need to.\n> \n>> I don't know how 0 bot solves this, but the obvious issue with this\n>> approach is to allow dealing with someone sending a patch like\n>> \n>> +++ Makefile\n>> --- Makefile\n>> +all:\n>> +\trm -fr $(HOME); sudo rm -fr /\n>> \n>> to the list. One thing that Travis gives us for free is isolation:\n>> malicious code in the build cannot break the bot, only the build\n>> itself.\n> \n> True, presumably the Travis integration already solves that part, so\n> I suspect it is just the matter of setting up:\n> \n> - a fork of git.git and have Travis monitor any and all new\n>   branches;\n> \n> - a bot that scans the list traffic, applies each series it sees to\n>   a branch dedicated for that series and pushes to the above fork.\n> \n> isn't it?\n\nMailing list users can already use Travis CI to check their patches\nprior to sending them. I just posted a patch with setup instructions\n(see $gmane/291371).\n\n@Junio:\nIf you setup Travis CI for your https://github.com/gitster/git fork\nthen Travis CI would build all your topic branches and you (and \neveryone who is interested) could check \nhttps://travis-ci.org/gitster/git/branches to see which branches \nwill break pu if you integrate them.\n\nI talked to Josh Kalderimis from Travis CI and he told me the load\nwouldn't be a problem for Travis CI at all.\n\n- Lars\n"},{"id":"283314","messageId":"BF053934-BA62-4621-AAAA-11F821B274EA@gmail.com","threadId":"41997","inReplyTo":"vpqegaa9i89.fsf@anie.imag.fr","subject":"Re: 0 bot for Git","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-04-13T12:16:54Z","receivedAt":"2016-04-13T12:16:54Z","isPatch":false,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 13 Apr 2016, at 07:43, Matthieu Moy <Matthieu.Moy@grenoble-inp.fr> wrote:\n> \n> Junio C Hamano <gitster@pobox.com> writes:\n> \n>> Matthieu Moy <Matthieu.Moy@grenoble-inp.fr> writes:\n>> \n>> True, presumably the Travis integration already solves that part, so\n>> I suspect it is just the matter of setting up:\n>> \n>> - a fork of git.git and have Travis monitor any and all new\n>>   branches;\n>> \n>> - a bot that scans the list traffic, applies each series it sees to\n>>   a branch dedicated for that series and pushes to the above fork.\n> \n> ... and to make it really useful: a way to get a notification email sent\n> on-list or at least to the submitter as a reply to the patch series.\n> Just having a web interface somewhere that knows how broken the code is\n> would not be that useful.\n\nTravis CI could do this but I intentionally disabled it to not annoy anyone.\nIt would be easy to enable it here:\nhttps://github.com/git/git/blob/7b0d47b3b6b5b64e02a5aa06b0452cadcdb18355/.travis.yml#L98-L99\n\nCheers,\nLars"},{"id":"283316","messageId":"vpq1t69669d.fsf@anie.imag.fr","threadId":"41997","inReplyTo":"BF053934-BA62-4621-AAAA-11F821B274EA@gmail.com","subject":"Re: 0 bot for Git","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@grenoble-inp.fr","sentAt":"2016-04-13T12:30:06Z","receivedAt":"2016-04-13T12:30:06Z","isPatch":false,"sender":{"key":"matthieu.moy@grenoble-inp.fr","avatar":"https://gravatar.com/avatar/72c8a2705971a25dfaff23cece15130d405685845d911aedd5667ace277f3fc5?d=mp&s=160"},"body":"Lars Schneider <larsxschneider@gmail.com> writes:\n\n>> On 13 Apr 2016, at 07:43, Matthieu Moy <Matthieu.Moy@grenoble-inp.fr> wrote:\n>> \n>> Junio C Hamano <gitster@pobox.com> writes:\n>> \n>>> Matthieu Moy <Matthieu.Moy@grenoble-inp.fr> writes:\n>>> \n>>> True, presumably the Travis integration already solves that part, so\n>>> I suspect it is just the matter of setting up:\n>>> \n>>> - a fork of git.git and have Travis monitor any and all new\n>>>   branches;\n>>> \n>>> - a bot that scans the list traffic, applies each series it sees to\n>>>   a branch dedicated for that series and pushes to the above fork.\n>> \n>> ... and to make it really useful: a way to get a notification email sent\n>> on-list or at least to the submitter as a reply to the patch series.\n>> Just having a web interface somewhere that knows how broken the code is\n>> would not be that useful.\n>\n> Travis CI could do this but I intentionally disabled it to not annoy anyone.\n> It would be easy to enable it here:\n> https://github.com/git/git/blob/7b0d47b3b6b5b64e02a5aa06b0452cadcdb18355/.travis.yml#L98-L99\n\nThe missing part would be \"as a reply to the patch series\". When I start\nreviewing a series, if the patch is broken and the CI system already\nknows, I'd rather have the information attached in the same thread right\ninside my mailer.\n\n-- \nMatthieu Moy\nhttp://www-verimag.imag.fr/~moy/\n"},{"id":"283375","messageId":"20160413134427.GA21158@wfg-t540p.sh.intel.com","threadId":"41997","inReplyTo":"xmqqmvoypn7g.fsf@gitster.mtv.corp.google.com","subject":"Re: 0 bot for Git","fromName":"Fengguang Wu","fromEmail":"lkp@intel.com","sentAt":"2016-04-13T13:44:27Z","receivedAt":"2016-04-13T13:44:27Z","isPatch":false,"sender":{"key":"lkp@intel.com","avatar":null},"body":"> > I don't know how 0 bot solves this, but the obvious issue with this\n> > approach is to allow dealing with someone sending a patch like\n> >\n> > +++ Makefile\n> > --- Makefile\n> > +all:\n> > +\trm -fr $(HOME); sudo rm -fr /\n> >\n> > to the list. One thing that Travis gives us for free is isolation:\n> > malicious code in the build cannot break the bot, only the build\n> > itself.\n\nSure, isolation is a must have for public test services like Travis or\n0day. We optimize the 0day infrastructure for good behaviors and also\nhave ways to isolate malicious ones.\n\n> True, presumably the Travis integration already solves that part, so\n> I suspect it is just the matter of setting up:\n> \n>  - a fork of git.git and have Travis monitor any and all new\n>    branches;\n> \n>  - a bot that scans the list traffic, applies each series it sees to\n>    a branch dedicated for that series and pushes to the above fork.\n> \n> isn't it?\n\nRight. 0day bot could auto maintain a patch-representing git tree for\nTravis to monitor and test. As how we already did for the linux kernel\nproject, creating one git branch per patchset posted to the lists:\n\n        https://github.com/0day-ci/linux/branches\n\nIn principle the git project should have more simple rules to decide\n\"which base should the robot apply a patch to\". But we do need some\nhints about the git community's rules in order to start the work. If\nwithout such hints from the community, we may start with dumb rules\nlike \"apply to latest origin/master\" or \"apply to latest release tag\".\n\nThanks,\nFengguang\n"},{"id":"283379","messageId":"8546368C-1884-4A56-AAA5-1E6B9C373E9F@gmail.com","threadId":"41997","inReplyTo":"vpq1t69669d.fsf@anie.imag.fr","subject":"Re: 0 bot for Git","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-04-13T16:14:14Z","receivedAt":"2016-04-13T16:14:14Z","isPatch":false,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 13 Apr 2016, at 14:30, Matthieu Moy <Matthieu.Moy@grenoble-inp.fr> wrote:\n> \n> Lars Schneider <larsxschneider@gmail.com> writes:\n> \n>>> On 13 Apr 2016, at 07:43, Matthieu Moy <Matthieu.Moy@grenoble-inp.fr> wrote:\n>>> \n>>> Junio C Hamano <gitster@pobox.com> writes:\n>>> \n>>>> Matthieu Moy <Matthieu.Moy@grenoble-inp.fr> writes:\n>>>> \n>>>> True, presumably the Travis integration already solves that part, so\n>>>> I suspect it is just the matter of setting up:\n>>>> \n>>>> - a fork of git.git and have Travis monitor any and all new\n>>>>  branches;\n>>>> \n>>>> - a bot that scans the list traffic, applies each series it sees to\n>>>>  a branch dedicated for that series and pushes to the above fork.\n>>> \n>>> ... and to make it really useful: a way to get a notification email sent\n>>> on-list or at least to the submitter as a reply to the patch series.\n>>> Just having a web interface somewhere that knows how broken the code is\n>>> would not be that useful.\n>> \n>> Travis CI could do this but I intentionally disabled it to not annoy anyone.\n>> It would be easy to enable it here:\n>> https://github.com/git/git/blob/7b0d47b3b6b5b64e02a5aa06b0452cadcdb18355/.travis.yml#L98-L99\n> \n> The missing part would be \"as a reply to the patch series\". When I start\n> reviewing a series, if the patch is broken and the CI system already\n> knows, I'd rather have the information attached in the same thread right\n> inside my mailer.\nI see. How would the automation know where the email patch needs to be applied?\n\n- Lars"},{"id":"283380","messageId":"xmqqfuuplc2x.fsf@gitster.mtv.corp.google.com","threadId":"41997","inReplyTo":"vpq1t69669d.fsf@anie.imag.fr","subject":"Re: 0 bot for Git","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-04-13T16:15:18Z","receivedAt":"2016-04-13T16:15:18Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Matthieu Moy <Matthieu.Moy@grenoble-inp.fr> writes:\n\n>>>> True, presumably the Travis integration already solves that part, so\n>>>> I suspect it is just the matter of setting up:\n>>>> \n>>>> - a fork of git.git and have Travis monitor any and all new\n>>>>   branches;\n>>>> \n>>>> - a bot that scans the list traffic, applies each series it sees to\n>>>>   a branch dedicated for that series and pushes to the above fork.\n>>> \n>>> ... and to make it really useful: a way to get a notification email sent\n>>> on-list or at least to the submitter as a reply to the patch series.\n>> Travis CI could do this ...\n> The missing part would be \"as a reply to the patch series\". When I start\n> reviewing a series, if the patch is broken and the CI system already\n> knows, I'd rather have the information attached in the same thread right\n> inside my mailer.\n\nYeah, such a message thrown randomly at the list would be too noisy\nto be useful, but if it is sent to a specific thread as a response,\nit would grab attention of those who are interested in the series,\nwhich is exactly what we want.\n\nSo with what you added, the list of what is needed is now:\n\n - a fork of git.git and have Travis monitor any and all new\n   branches;\n \n - a bot that scans the list traffic, applies each series it sees to\n   a branch dedicated for that series and pushes to the above fork;\n\n - a bot (which can be the same as the above or a different one, as\n   long as the former and the latter has a way to associate a topic\n   branch and the original message) that receives success/failure\n   notice from Travis, relays it as a response to the original patch\n   on the list.\n"},{"id":"283381","messageId":"xmqqa8kxlbix.fsf@gitster.mtv.corp.google.com","threadId":"41997","inReplyTo":"88CF8CB5-4105-4D0C-8064-D66092169111@gmail.com","subject":"Re: 0 bot for Git","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-04-13T16:27:18Z","receivedAt":"2016-04-13T16:27:18Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Lars Schneider <larsxschneider@gmail.com> writes:\n\n> @Junio:\n> If you setup Travis CI for your https://github.com/gitster/git fork\n> then Travis CI would build all your topic branches and you (and \n> everyone who is interested) could check \n> https://travis-ci.org/gitster/git/branches to see which branches \n> will break pu if you integrate them.\n\nI would not say such an arrangement is worthless, but it targets a\nwrong point in the patch flow.\n\nThe patches that result in the most wastage of my time (i.e. a\nshared bottleneck resource the community should strive to optimize\nfor) are the ones that fail to hit 'pu'.  Ones that do not even\nbuild in isolation, ones that may build but fail even the new tests\nthey bring in, ones that break existing tests, and ones that are OK\nin isolation but do not play well with topics already in flight.\n\nAutomated testing of what is already on 'pu' does not help reduce\nthe above cost, as the culling must be done by me _without_ help\nfrom automated test you propose to run on topics in 'pu'.  Ever\nheard of chicken and egg?\n\nYour \"You can setup your own CI\" update to SubmittingPatches may\nencourage people to test before sending.  The \"Travis CI sends\nfailure notice as a response to a crappy patch\" discussed by\nMatthieu in the other subthread will be of great help.\n\nThanks.\n"},{"id":"283384","messageId":"BF9D5A7E-CB73-4F82-8D5F-42E120D07A3B@gmail.com","threadId":"41997","inReplyTo":"xmqqa8kxlbix.fsf@gitster.mtv.corp.google.com","subject":"Re: 0 bot for Git","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-04-13T17:09:41Z","receivedAt":"2016-04-13T17:09:41Z","isPatch":false,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 13 Apr 2016, at 18:27, Junio C Hamano <gitster@pobox.com> wrote:\n> \n> Lars Schneider <larsxschneider@gmail.com> writes:\n> \n>> @Junio:\n>> If you setup Travis CI for your https://github.com/gitster/git fork\n>> then Travis CI would build all your topic branches and you (and \n>> everyone who is interested) could check \n>> https://travis-ci.org/gitster/git/branches to see which branches \n>> will break pu if you integrate them.\n> \n> I would not say such an arrangement is worthless, but it targets a\n> wrong point in the patch flow.\n> \n> The patches that result in the most wastage of my time (i.e. a\n> shared bottleneck resource the community should strive to optimize\n> for) are the ones that fail to hit 'pu'.  Ones that do not even\n> build in isolation, ones that may build but fail even the new tests\n> they bring in, ones that break existing tests, and ones that are OK\n> in isolation but do not play well with topics already in flight.\n\nI am not sure what you mean by \"fail to hit 'pu'\". Maybe we talk at\ncross purposes. Here is what I think you do, please correct me:\n\n1.) You pick the topics from the mailing list and create feature \n    branches for each one of them. E.g. one of my recent topics \n    is \"ls/config-origin\".\n\n2.) At some point you create a new pu branch based on the latest\n    next branch. You merge all the new topics into the new pu.\n\nIf you push the topics to github.com/gitster after step 1 then\nTravis CI could tell you if the individual topic builds clean \nand passes all tests. Then you could merge only clean topics in \nstep 2 which would result in a pu that is much more likely to \nbuild clean.\n\nCould that process avoid wasting your time with bad patches?\n\n> Automated testing of what is already on 'pu' does not help reduce\n> the above cost, as the culling must be done by me _without_ help\n> from automated test you propose to run on topics in 'pu'.  Ever\n> heard of chicken and egg?\n> \n> Your \"You can setup your own CI\" update to SubmittingPatches may\n> encourage people to test before sending.  The \"Travis CI sends\n> failure notice as a response to a crappy patch\" discussed by\n> Matthieu in the other subthread will be of great help.\n> \n> Thanks.\n> \n"},{"id":"283387","messageId":"CAGZ79ka4WmT8NjD-04WqwczuCuJZcoKMyDRQKkRH1sT5xoqRhQ@mail.gmail.com","threadId":"41997","inReplyTo":"BF9D5A7E-CB73-4F82-8D5F-42E120D07A3B@gmail.com","subject":"Re: 0 bot for Git","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2016-04-13T17:29:57Z","receivedAt":"2016-04-13T17:29:57Z","isPatch":false,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Wed, Apr 13, 2016 at 10:09 AM, Lars Schneider\n<larsxschneider@gmail.com> wrote:\n>\n>> On 13 Apr 2016, at 18:27, Junio C Hamano <gitster@pobox.com> wrote:\n>>\n>> Lars Schneider <larsxschneider@gmail.com> writes:\n>>\n>>> @Junio:\n>>> If you setup Travis CI for your https://github.com/gitster/git fork\n>>> then Travis CI would build all your topic branches and you (and\n>>> everyone who is interested) could check\n>>> https://travis-ci.org/gitster/git/branches to see which branches\n>>> will break pu if you integrate them.\n>>\n>> I would not say such an arrangement is worthless, but it targets a\n>> wrong point in the patch flow.\n>>\n>> The patches that result in the most wastage of my time (i.e. a\n>> shared bottleneck resource the community should strive to optimize\n>> for) are the ones that fail to hit 'pu'.  Ones that do not even\n>> build in isolation, ones that may build but fail even the new tests\n>> they bring in, ones that break existing tests, and ones that are OK\n>> in isolation but do not play well with topics already in flight.\n>\n> I am not sure what you mean by \"fail to hit 'pu'\". Maybe we talk at\n> cross purposes. Here is what I think you do, please correct me:\n>\n> 1.) You pick the topics from the mailing list and create feature\n>     branches for each one of them. E.g. one of my recent topics\n>     is \"ls/config-origin\".\n\nand by You you mean Junio.\n\nIdeally the 0bot would have sent the message as a reply to the\ncover letter with the information \"doesn't compile/breaks test t1234\",\nso Junio could ignore that series (no time wasted on his part).\n\nAt Git Merge Greg said (paraphrasing here):\n\n  We waste developers time, because we have plenty of it. Maintainers time\n  however is precious because maintainers are the bottleneck and a scare\n  resource to come by.\n\nAnd I think Git and the kernel have the same community design here.\n(Except the kernel is bigger and has more than one maintainer)\n\nSo the idea is help Junio make a decision to drop/ignore those patches\nwith least amount of brain cycled spent as possible. (Not even spend 5\nseconds on it).\n\n>\n> 2.) At some point you create a new pu branch based on the latest\n>     next branch. You merge all the new topics into the new pu.\n\nbut Junio also runs test after each(?) merge(?) of a series and once\ntests fail, it takes time to sort out, what caused it. (Is that the patch series\nalone or is that because 2 series interact badly with each other?)\n\n>\n> If you push the topics to github.com/gitster after step 1 then\n> Travis CI could tell you if the individual topic builds clean\n> and passes all tests. Then you could merge only clean topics in\n> step 2 which would result in a pu that is much more likely to\n> build clean.\n\nIIRC Junio did not like granting travis access to the \"blessed\" repository\nas travis wants so much permissions including write permission to that\nrepo. (We/He could have a second non advertised repo though)\n\nAlso this would incur wait time on Junios side\n\n1) collect patches (many series over the day)\n2) push\n3) wait\n4) do the merges\n\nhowever a 0 bot would do\n1) collect patches faster than Junio (0 bot is a computer after all,\nworking 24/7)\n2) test each patch/series individually\n3) send feedback without the wait time, so the contributor from a different\n   time zone gets feedback quickly. (round trip is just the build and test time,\n   which the developer forgot to do any way if it fails)\n\n>\n> Could that process avoid wasting your time with bad patches?\n>\n>> Automated testing of what is already on 'pu' does not help reduce\n>> the above cost, as the culling must be done by me _without_ help\n>> from automated test you propose to run on topics in 'pu'.  Ever\n>> heard of chicken and egg?\n>>\n>> Your \"You can setup your own CI\" update to SubmittingPatches may\n>> encourage people to test before sending.  The \"Travis CI sends\n>> failure notice as a response to a crappy patch\" discussed by\n>> Matthieu in the other subthread will be of great help.\n>>\n>> Thanks.\n>>\n>\n"},{"id":"283389","messageId":"20160413174311.GA31431@kroah.com","threadId":"41997","inReplyTo":"CAGZ79ka4WmT8NjD-04WqwczuCuJZcoKMyDRQKkRH1sT5xoqRhQ@mail.gmail.com","subject":"Re: 0 bot for Git","fromName":"Greg KH","fromEmail":"gregkh@linuxfoundation.org","sentAt":"2016-04-13T17:43:11Z","receivedAt":"2016-04-13T17:43:11Z","isPatch":false,"sender":{"key":"gregkh@linuxfoundation.org","avatar":"https://gravatar.com/avatar/e6d9136f6e3bdcb59f0e5fd15565f382da42523d273824958b9e23e73cf38e04?d=mp&s=160"},"body":"On Wed, Apr 13, 2016 at 10:29:57AM -0700, Stefan Beller wrote:\n> \n> At Git Merge Greg said (paraphrasing here):\n> \n>   We waste developers time, because we have plenty of it. Maintainers time\n>   however is precious because maintainers are the bottleneck and a scare\n>   resource to come by.\n\ns/scare/scarce/\n\nAlthough some people might disagree :)\n"},{"id":"283392","messageId":"xmqqzisxjt96.fsf@gitster.mtv.corp.google.com","threadId":"41997","inReplyTo":"BF9D5A7E-CB73-4F82-8D5F-42E120D07A3B@gmail.com","subject":"Re: 0 bot for Git","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-04-13T17:47:17Z","receivedAt":"2016-04-13T17:47:17Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Lars Schneider <larsxschneider@gmail.com> writes:\n\n> I am not sure what you mean by \"fail to hit 'pu'\". Maybe we talk at\n> cross purposes. Here is what I think you do, please correct me:\n>\n> 1.) You pick the topics from the mailing list and create feature \n>     branches for each one of them. E.g. one of my recent topics \n>     is \"ls/config-origin\".\n\nI do not do this step blindly.  The patch is studied in MUA, perhaps\napplied to a new topic to view it in wider context, and tested in\nisolation at this step.  In any of these steps, I may decide it is\nway too premature for 'pu' and discard it.\n"},{"id":"283491","messageId":"CAP8UFD3xWUkCFZMN1N6t36KKwcfnkLsFznAc7j7yF89PbYaqfg@mail.gmail.com","threadId":"41997","inReplyTo":"CAGZ79kZzdioQRFEmgTGOOdLQ-Ov-tWmgi1dLhHPDVzDb+Py2RQ@mail.gmail.com","subject":"Re: 0 bot for Git","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2016-04-14T22:04:49Z","receivedAt":"2016-04-14T22:04:49Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"On Tue, Apr 12, 2016 at 4:59 PM, Stefan Beller <sbeller@google.com> wrote:\n> On Tue, Apr 12, 2016 at 2:42 AM, Duy Nguyen <pclouds@gmail.com> wrote:\n>>> On Mon, Apr 11, 2016 at 7:51 AM, Stefan Beller <sbeller@google.com> wrote:\n>>>> Hi Greg,\n>>>>\n>>>> Thanks for your talk at the Git Merge 2016!\n>>\n>> Huh? It already happened?? Any interesting summary to share with us?\n>\n> Summary from the contributors summit:\n> * encoding was a huge discussion point:\n>   If you use Git on a case sensitive file system, you may have\n>   branches \"foo\" and \"FOO\" and it just works. Once you switch over\n>   to another file system this may break horribly.\n>\n>   The discussion revealed lots more of these points in Git and then a\n>   discussion on fundamentals sparked, on whether we want to go by the lowest\n>   common denominator or treat any system special or have it as is, but less\n>   broken.\n>\n> * We still don't know how to handle large repositories\n>   (I am working on submodules but that is too vague and not solving\n>   the actual problem, people really want there\n>     * large files\n>     * large trees\n>    [* lots of commits (i.e. lots of objects)]\n>    in one repo and it should work just as fast.)\n>\n> That was my main take away.\n\nThere is a draft of an article about the first part of the Contributor\nSummit in the draft of the next Git Rev News edition:\n\nhttps://github.com/git/git.github.io/blob/master/rev_news/drafts/edition-14.md\n\nEveryone is welcome to contribute about things that are missing or not accurate.\n"},{"id":"283523","messageId":"20160415095139.GA3985@lanh","threadId":"41997","inReplyTo":"CAP8UFD3xWUkCFZMN1N6t36KKwcfnkLsFznAc7j7yF89PbYaqfg@mail.gmail.com","subject":"Parallel checkout (Was Re: 0 bot for Git)","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2016-04-15T09:51:40Z","receivedAt":"2016-04-15T09:51:40Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Apr 15, 2016 at 12:04:49AM +0200, Christian Couder wrote:\n> On Tue, Apr 12, 2016 at 4:59 PM, Stefan Beller <sbeller@google.com> wrote:\n> > On Tue, Apr 12, 2016 at 2:42 AM, Duy Nguyen <pclouds@gmail.com> wrote:\n> >>> On Mon, Apr 11, 2016 at 7:51 AM, Stefan Beller <sbeller@google.com> wrote:\n> >>>> Hi Greg,\n> >>>>\n> >>>> Thanks for your talk at the Git Merge 2016!\n> >>\n> >> Huh? It already happened?? Any interesting summary to share with us?\n> \n> There is a draft of an article about the first part of the Contributor\n> Summit in the draft of the next Git Rev News edition:\n> \n> https://github.com/git/git.github.io/blob/master/rev_news/drafts/edition-14.md\n\nThanks. I read the sentence \"This made people mention potential\nproblems with parallelizing git checkout\" and wondered what these\nproblems were. And because it's easier to think while you test\nsomething, to flesh out your thoughts, I wrote the below patch, which\ndoes parallel checkout with multiple worker processes. I wonder if the\nsame set of problems apply to it.\n\nThe idea is simple, you offload some work to process workers. In this\npatch, only entry.c:write_entry() is moved to workers. We still do\ndirectory creation and all sort of checks and stat refresh in the main\nprocess. Some more work may be moved away, for example, the entire\nbuiltin/checkout.c:checkout_merged().\n\nMulti process is less efficient than multi thread model. But I doubt\nwe could make object db access thread-safe soon. The last discussion\nwas 2 years ago [1] and nothing much has happened.\n\nNumbers are encouraging though. On linux-2.6 repo running on linux and\next4 filesystem, checkout_paths() would dominate \"git checkout :/\".\nUnmodified git takes about 31s.\n\n\n16:26:00.114029 builtin/checkout.c:1299 performance: 31.184973659 s: checkout_paths\n16:26:00.114225 trace.c:420             performance: 31.256412935 s: git command: 'git' 'checkout' '.'\n\nWhen doing write_entry() on 8 processes, it takes 22s (shortened by ~30%)\n\n16:27:39.973730 trace.c:420             performance: 5.610255442 s: git command: 'git' 'checkout-index' '--worker'\n16:27:40.956812 trace.c:420             performance: 6.595082013 s: git command: 'git' 'checkout-index' '--worker'\n16:27:41.397621 trace.c:420             performance: 7.032024175 s: git command: 'git' 'checkout-index' '--worker'\n16:27:47.453999 trace.c:420             performance: 13.078537207 s: git command: 'git' 'checkout-index' '--worker'\n16:27:48.986433 trace.c:420             performance: 14.612951643 s: git command: 'git' 'checkout-index' '--worker'\n16:27:53.149378 trace.c:420             performance: 18.781762536 s: git command: 'git' 'checkout-index' '--worker'\n16:27:54.884044 trace.c:420             performance: 20.514473730 s: git command: 'git' 'checkout-index' '--worker'\n16:27:55.319990 trace.c:420             performance: 20.948326263 s: git command: 'git' 'checkout-index' '--worker'\n16:27:55.863211 builtin/checkout.c:1299 performance: 22.723118420 s: checkout_paths\n16:27:55.863398 trace.c:420             performance: 22.854547640 s: git command: 'git' 'checkout' '--parallel' '.'\n\nI suspect on nfs or windows, the gain may be higher due to IO blocking\nthe main process more.\n\nNote that this for-fun patch is not optmized at all (and definitely\nnot portable). I could have sent a group of paths to the worker in a\nsingle system call instead of one per call. The trace above also shows\nunbalance issues with workers, where some workers exit early because\nof my naive work distribution. Numbers could get a bit better.\n\n[1] http://thread.gmane.org/gmane.comp.version-control.git/241965/focus=242020\n\n-- 8< --\ndiff --git a/builtin/checkout-index.c b/builtin/checkout-index.c\nindex 92c6967..7163216 100644\n--- a/builtin/checkout-index.c\n+++ b/builtin/checkout-index.c\n@@ -9,6 +9,7 @@\n #include \"quote.h\"\n #include \"cache-tree.h\"\n #include \"parse-options.h\"\n+#include \"entry.h\"\n \n #define CHECKOUT_ALL 4\n static int nul_term_line;\n@@ -179,6 +180,9 @@ int cmd_checkout_index(int argc, const char **argv, const char *prefix)\n \t\tOPT_END()\n \t};\n \n+\tif (argc == 2 && !strcmp(argv[1], \"--worker\"))\n+\t\treturn parallel_checkout_worker();\n+\n \tif (argc == 2 && !strcmp(argv[1], \"-h\"))\n \t\tusage_with_options(builtin_checkout_index_usage,\n \t\t\t\t   builtin_checkout_index_options);\ndiff --git a/builtin/checkout.c b/builtin/checkout.c\nindex efcbd8f..51caad2 100644\n--- a/builtin/checkout.c\n+++ b/builtin/checkout.c\n@@ -20,6 +20,7 @@\n #include \"resolve-undo.h\"\n #include \"submodule-config.h\"\n #include \"submodule.h\"\n+#include \"entry.h\"\n \n static const char * const checkout_usage[] = {\n \tN_(\"git checkout [<options>] <branch>\"),\n@@ -236,7 +237,8 @@ static int checkout_merged(int pos, struct checkout *state)\n }\n \n static int checkout_paths(const struct checkout_opts *opts,\n-\t\t\t  const char *revision)\n+\t\t\t  const char *revision,\n+\t\t\t  int parallel)\n {\n \tint pos;\n \tstruct checkout state;\n@@ -357,6 +359,8 @@ static int checkout_paths(const struct checkout_opts *opts,\n \tstate.force = 1;\n \tstate.refresh_cache = 1;\n \tstate.istate = &the_index;\n+\tif (parallel)\n+\t\tstart_parallel_checkout(&state);\n \tfor (pos = 0; pos < active_nr; pos++) {\n \t\tstruct cache_entry *ce = active_cache[pos];\n \t\tif (ce->ce_flags & CE_MATCHED) {\n@@ -367,11 +371,18 @@ static int checkout_paths(const struct checkout_opts *opts,\n \t\t\tif (opts->writeout_stage)\n \t\t\t\terrs |= checkout_stage(opts->writeout_stage, ce, pos, &state);\n \t\t\telse if (opts->merge)\n+\t\t\t\t/*\n+\t\t\t\t * XXX: in parallel mode, we may want\n+\t\t\t\t * to let worker perform the merging\n+\t\t\t\t * instead and send SHA-1 result back\n+\t\t\t\t */\n \t\t\t\terrs |= checkout_merged(pos, &state);\n \t\t\tpos = skip_same_name(ce, pos) - 1;\n \t\t}\n \t}\n \n+\terrs |= run_parallel_checkout();\n+\n \tif (write_locked_index(&the_index, lock_file, COMMIT_LOCK))\n \t\tdie(_(\"unable to write new index file\"));\n \n@@ -1132,6 +1143,7 @@ int cmd_checkout(int argc, const char **argv, const char *prefix)\n \tstruct branch_info new;\n \tchar *conflict_style = NULL;\n \tint dwim_new_local_branch = 1;\n+\tint parallel = 0;\n \tstruct option options[] = {\n \t\tOPT__QUIET(&opts.quiet, N_(\"suppress progress reporting\")),\n \t\tOPT_STRING('b', NULL, &opts.new_branch, N_(\"branch\"),\n@@ -1159,6 +1171,8 @@ int cmd_checkout(int argc, const char **argv, const char *prefix)\n \t\t\t\tN_(\"second guess 'git checkout <no-such-branch>'\")),\n \t\tOPT_BOOL(0, \"ignore-other-worktrees\", &opts.ignore_other_worktrees,\n \t\t\t N_(\"do not check if another worktree is holding the given ref\")),\n+\t\tOPT_BOOL(0, \"parallel\", &parallel,\n+\t\t\t N_(\"parallel checkout\")),\n \t\tOPT_BOOL(0, \"progress\", &opts.show_progress, N_(\"force progress reporting\")),\n \t\tOPT_END(),\n \t};\n@@ -1279,8 +1293,12 @@ int cmd_checkout(int argc, const char **argv, const char *prefix)\n \t\tstrbuf_release(&buf);\n \t}\n \n-\tif (opts.patch_mode || opts.pathspec.nr)\n-\t\treturn checkout_paths(&opts, new.name);\n+\tif (opts.patch_mode || opts.pathspec.nr) {\n+\t\tuint64_t start = getnanotime();\n+\t\tint ret = checkout_paths(&opts, new.name, parallel);\n+\t\ttrace_performance_since(start, \"checkout_paths\");\n+\t\treturn ret;\n+\t}\n \telse\n \t\treturn checkout_branch(&opts, &new);\n }\ndiff --git a/entry.c b/entry.c\nindex a410957..5e0eb1c 100644\n--- a/entry.c\n+++ b/entry.c\n@@ -3,6 +3,36 @@\n #include \"dir.h\"\n #include \"streaming.h\"\n \n+#include <sys/epoll.h>\n+#include \"pkt-line.h\"\n+#include \"argv-array.h\"\n+#include \"run-command.h\"\n+\n+struct checkout_item {\n+\tstruct cache_entry *ce;\n+\tstruct checkout_item *next;\n+};\n+\n+struct checkout_worker {\n+\tstruct child_process cp;\n+\tstruct checkout_item *to_complete;\n+\tstruct checkout_item *to_send;\n+};\n+\n+struct parallel_checkout {\n+\tstruct checkout state;\n+\tstruct checkout_worker *workers;\n+\tstruct checkout_item *items;\n+\tint nr_items, alloc_items;\n+\tint nr_workers;\n+};\n+\n+static struct parallel_checkout *parallel_checkout;\n+\n+static int queue_checkout(struct parallel_checkout *,\n+\t\t\t  const struct checkout *,\n+\t\t\t  struct cache_entry *);\n+\n static void create_directories(const char *path, int path_len,\n \t\t\t       const struct checkout *state)\n {\n@@ -290,5 +320,299 @@ int checkout_entry(struct cache_entry *ce,\n \t\treturn 0;\n \n \tcreate_directories(path.buf, path.len, state);\n+\n+\tif (!queue_checkout(parallel_checkout, state, ce))\n+\t\t/*\n+\t\t * write_entry() will be done by parallel_checkout_worker() in\n+\t\t * a separate process\n+\t\t */\n+\t\treturn 0;\n+\n \treturn write_entry(ce, path.buf, state, 0);\n }\n+\n+int start_parallel_checkout(const struct checkout *state)\n+{\n+\tif (parallel_checkout)\n+\t\tdie(\"BUG: parallel checkout already initiated\");\n+\tif (0 && state->force)\n+\t\tdie(\"BUG: not support --force yet\");\n+\tparallel_checkout = xmalloc(sizeof(*parallel_checkout));\n+\tmemset(parallel_checkout, 0, sizeof(*parallel_checkout));\n+\tmemcpy(&parallel_checkout->state, state, sizeof(*state));\n+\n+\treturn 0;\n+}\n+\n+static int queue_checkout(struct parallel_checkout *pc,\n+\t\t\t  const struct checkout *state,\n+\t\t\t  struct cache_entry *ce)\n+{\n+\tstruct checkout_item *ci;\n+\n+\tif (!pc ||\n+\t    !S_ISREG(ce->ce_mode) ||\n+\t    memcmp(&pc->state, state, sizeof(*state)))\n+\t\treturn -1;\n+\n+\tALLOC_GROW(pc->items, pc->nr_items + 1, pc->alloc_items);\n+\tci = pc->items + pc->nr_items++;\n+\tci->ce = ce;\n+\treturn 0;\n+}\n+\n+static int item_cmp(const void *a_, const void *b_)\n+{\n+\tconst struct checkout_item *a = a_;\n+\tconst struct checkout_item *b = b_;\n+\treturn strcmp(a->ce->name, b->ce->name);\n+}\n+\n+static int setup_workers(struct parallel_checkout *pc, int epoll_fd)\n+{\n+\tint from, nr_per_worker, i;\n+\n+\tpc->workers = xmalloc(sizeof(*pc->workers) * pc->nr_workers);\n+\tmemset(pc->workers, 0, sizeof(*pc->workers) * pc->nr_workers);\n+\n+\tnr_per_worker = pc->nr_items / pc->nr_workers;\n+\tfrom = 0;\n+\n+\tfor (i = 0; i < pc->nr_workers; i++) {\n+\t\tstruct checkout_worker *worker = pc->workers + i;\n+\t\tstruct child_process *cp = &worker->cp;\n+\t\tstruct checkout_item *item;\n+\t\tstruct epoll_event ev;\n+\t\tint to;\n+\n+\t\tto = from + nr_per_worker;\n+\t\tif (i == pc->nr_workers - 1)\n+\t\t\tto = pc->nr_items;\n+\t\titem = NULL;\n+\t\twhile (from < to) {\n+\t\t\tpc->items[from].next = item;\n+\t\t\titem = pc->items + from;\n+\t\t\tfrom++;\n+\t\t}\n+\t\tworker->to_send = item;\n+\t\tworker->to_complete = item;\n+\n+\t\tcp->git_cmd = 1;\n+\t\tcp->in = -1;\n+\t\tcp->out = -1;\n+\t\targv_array_push(&cp->args, \"checkout-index\");\n+\t\targv_array_push(&cp->args, \"--worker\");\n+\t\tif (start_command(cp))\n+\t\t\tdie(_(\"failed to run checkout worker\"));\n+\n+\t\tev.events = EPOLLOUT | EPOLLERR | EPOLLHUP;\n+\t\tev.data.u32 = i * 2;\n+\t\tif (epoll_ctl(epoll_fd, EPOLL_CTL_ADD, cp->in, &ev) == -1)\n+\t\t\tdie_errno(\"epoll_ctl\");\n+\n+\t\tev.events = EPOLLIN | EPOLLERR | EPOLLHUP;\n+\t\tev.data.u32 = i * 2 + 1;\n+\t\tif (epoll_ctl(epoll_fd, EPOLL_CTL_ADD, cp->out, &ev) == -1)\n+\t\t\tdie_errno(\"epoll_ctl\");\n+\t}\n+\treturn 0;\n+}\n+\n+static int send_to_worker(struct checkout_worker *worker, int epoll_fd)\n+{\n+\tif (!worker->to_send) {\n+\t\tstruct epoll_event ev;\n+\n+\t\tpacket_flush(worker->cp.in);\n+\n+\t\tepoll_ctl(epoll_fd, EPOLL_CTL_DEL, worker->cp.in, &ev);\n+\t\tclose(worker->cp.in);\n+\t\tworker->cp.in = -1;\n+\t\treturn 0;\n+\t}\n+\n+\t/*\n+\t * XXX: put the fd in non-blocking mode and send as many files\n+\t * as possible in one go.\n+\t */\n+\tpacket_write(worker->cp.in, \"%s %s\",\n+\t\t     sha1_to_hex(worker->to_send->ce->sha1),\n+\t\t     worker->to_send->ce->name);\n+\tworker->to_send = worker->to_send->next;\n+\treturn 0;\n+}\n+\n+int parallel_checkout_worker(void)\n+{\n+\tstruct checkout state;\n+\tstruct cache_entry *ce = NULL;\n+\n+\tmemset(&state, 0, sizeof(state));\n+\t/* FIXME: pass 'force' over */\n+\tfor (;;) {\n+\t\tint len;\n+\t\tunsigned char sha1[20];\n+\t\tchar *line = packet_read_line(0, &len);\n+\n+\t\tif (!line)\n+\t\t\treturn 0;\n+\n+\t\tif (len < 40)\n+\t\t\treturn 1;\n+\t\tif (get_sha1_hex(line, sha1))\n+\t\t\treturn 1;\n+\t\tline += 40;\n+\t\tlen -= 40;\n+\t\tif (*line != ' ')\n+\t\t\treturn 1;\n+\t\tline++;\n+\t\tlen--;\n+\t\tif (!ce || ce_namelen(ce) < len) {\n+\t\t\tfree(ce);\n+\t\t\tce = xcalloc(1, cache_entry_size(len));\n+\t\t\tce->ce_mode = S_IFREG | ce_permissions(0644);\n+\t\t}\n+\t\tce->ce_namelen = len;\n+\t\thashcpy(ce->sha1, sha1);\n+\t\tmemcpy(ce->name, line, len + 1);\n+\n+\t\tif (write_entry(ce, ce->name, &state, 0))\n+\t\t\treturn 1;\n+\t\t/*\n+\t\t * XXX process in batch and send bigger number of\n+\t\t * checked out entries back\n+\t\t */\n+\t\tpacket_write(1, \"1\");\n+\t}\n+}\n+\n+static int receive_from_worker(struct checkout_worker *worker,\n+\t\t\t       int refresh_cache)\n+{\n+\tint len, val;\n+\tchar *line;\n+\n+\tline = packet_read_line(worker->cp.out, &len);\n+\tval = atoi(line);\n+\tif (val <= 0)\n+\t\tdie(\"BUG: invalid value\");\n+\twhile (val && worker->to_complete &&\n+\t       worker->to_complete != worker->to_send) {\n+\t\tif (refresh_cache) {\n+\t\t\tstruct stat st;\n+\t\t\tstruct cache_entry *ce = worker->to_complete->ce;\n+\n+\t\t\tlstat(ce->name, &st);\n+\t\t\tfill_stat_cache_info(ce, &st);\n+\t\t\tce->ce_flags |= CE_UPDATE_IN_BASE;\n+\t\t}\n+\t\tworker->to_complete = worker->to_complete->next;\n+\t\tval--;\n+\t}\n+\tif (val)\n+\t\tdie(\"BUG: invalid value\");\n+\treturn 0;\n+}\n+\n+static int finish_worker(struct checkout_worker *worker, int epoll_fd)\n+{\n+\tstruct epoll_event ev;\n+\tchar buf[1];\n+\tint ret;\n+\n+\tassert(worker->to_send == NULL);\n+\tassert(worker->to_complete == NULL);\n+\n+\tret = xread(worker->cp.out, buf, sizeof(buf));\n+\tif (ret != 0)\n+\t\tdie(\"BUG: expect eof\");\n+\tepoll_ctl(epoll_fd, EPOLL_CTL_DEL, worker->cp.out, &ev);\n+\tclose(worker->cp.out);\n+\tworker->cp.out = -1;\n+\tif (finish_command(&worker->cp))\n+\t\tdie(\"worker had a problem\");\n+\treturn 0;\n+}\n+\n+static int really_finished(struct parallel_checkout *pc)\n+{\n+\tint i;\n+\n+\tfor (i = 0; i < pc->nr_workers; i++)\n+\t\tif (pc->workers[i].to_complete)\n+\t\t\treturn 0;\n+\treturn 1;\n+}\n+\n+/* XXX progress support for unpack-trees */\n+int run_parallel_checkout(void)\n+{\n+\tstruct parallel_checkout *pc = parallel_checkout;\n+\tint ret, i;\n+\tint epoll_fd;\n+\tstruct epoll_event *events;\n+\n+\tif (!pc || !pc->nr_items) {\n+\t\tfree(pc);\n+\t\tparallel_checkout = NULL;\n+\t\treturn 0;\n+\t}\n+\n+\tqsort(pc->items, pc->nr_items, sizeof(*pc->items), item_cmp);\n+\tpc->nr_workers = 8;\n+\tepoll_fd = epoll_create(pc->nr_workers * 2);\n+\tif (epoll_fd == -1)\n+\t\tdie_errno(\"epoll_create\");\n+\tret = setup_workers(pc, epoll_fd);\n+\tevents = xmalloc(sizeof(*events) * pc->nr_workers * 2);\n+\n+\tret = 0;\n+\twhile (!ret) {\n+\t\tint maybe_all_done = 0, nr;\n+\n+\t\tnr = epoll_wait(epoll_fd, events, pc->nr_workers * 2, -1);\n+\t\tif (nr == -1 && errno == EINTR)\n+\t\t\tcontinue;\n+\t\tif (nr == -1) {\n+\t\t\tret = nr;\n+\t\t\tbreak;\n+\t\t}\n+\t\tfor (i = 0; i < nr; i++) {\n+\t\t\tint is_in = events[i].data.u32 & 1;\n+\t\t\tint worker_id = events[i].data.u32 / 2;\n+\t\t\tstruct checkout_worker *worker = pc->workers + worker_id;\n+\n+\t\t\tif (!is_in && (events[i].events & EPOLLOUT))\n+\t\t\t\tret = send_to_worker(worker, epoll_fd);\n+\t\t\telse if (events[i].events & EPOLLIN) {\n+\t\t\t\tif (worker->to_complete) {\n+\t\t\t\t\tint refresh = pc->state.refresh_cache;\n+\t\t\t\t\tret = receive_from_worker(worker, refresh);\n+\t\t\t\t\tpc->state.istate->cache_changed |= CE_ENTRY_CHANGED;\n+\t\t\t\t} else {\n+\t\t\t\t\tret = finish_worker(worker, epoll_fd);\n+\t\t\t\t\tmaybe_all_done = 1;\n+\t\t\t\t}\n+\t\t\t} else if (events[i].events & (EPOLLERR | EPOLLHUP)) {\n+\t\t\t\tif (is_in && !worker->to_complete) {\n+\t\t\t\t\tret = finish_worker(worker, epoll_fd);\n+\t\t\t\t\tmaybe_all_done = 1;\n+\t\t\t\t} else\n+\t\t\t\t\tret = -1;\n+\t\t\t} else\n+\t\t\t\tdie(\"BUG: what??\");\n+\t\t\tif (ret)\n+\t\t\t\tbreak;\n+\t\t}\n+\n+\t\tif (maybe_all_done && really_finished(pc))\n+\t\t\tbreak;\n+\t}\n+\n+\tclose(epoll_fd);\n+\tfree(pc->workers);\n+\tfree(events);\n+\tfree(pc);\n+\tparallel_checkout = NULL;\n+\treturn ret;\n+}\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 9f55cc2..433c54e 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -220,6 +220,7 @@ static int check_updates(struct unpack_trees_options *o)\n \tremove_marked_cache_entries(&o->result);\n \tremove_scheduled_dirs();\n \n+\t/* start_parallel_checkout() */\n \tfor (i = 0; i < index->cache_nr; i++) {\n \t\tstruct cache_entry *ce = index->cache[i];\n \n@@ -234,6 +235,7 @@ static int check_updates(struct unpack_trees_options *o)\n \t\t\t}\n \t\t}\n \t}\n+\t/* run_parallel_checkout() */\n \tstop_progress(&progress);\n \tif (o->update)\n \t\tgit_attr_set_direction(GIT_ATTR_CHECKIN, NULL);\n-- 8< --\n"},{"id":"283524","messageId":"CAP8UFD0WZHriY340eh3K6ygzb0tXnoT+XaY8+c2k+N2x9UBYxA@mail.gmail.com","threadId":"41997","inReplyTo":"20160415095139.GA3985@lanh","subject":"Re: Parallel checkout (Was Re: 0 bot for Git)","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2016-04-15T11:18:46Z","receivedAt":"2016-04-15T11:18:46Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"On Fri, Apr 15, 2016 at 11:51 AM, Duy Nguyen <pclouds@gmail.com> wrote:\n> On Fri, Apr 15, 2016 at 12:04:49AM +0200, Christian Couder wrote:\n>>\n>> There is a draft of an article about the first part of the Contributor\n>> Summit in the draft of the next Git Rev News edition:\n>>\n>> https://github.com/git/git.github.io/blob/master/rev_news/drafts/edition-14.md\n>\n> Thanks. I read the sentence \"This made people mention potential\n> problems with parallelizing git checkout\" and wondered what these\n> problems were.\n\nIt may have been Michael or Peff (CC'ed) saying that it could break\nsome builds as the timestamps on the files might not always be ordered\nin the same way.\n\nNow perhaps parallel checkout could be activated only if a config\noption was set. (Yeah, I know it looks like I am very often asking for\nconfig options.)\n\n> And because it's easier to think while you test\n> something, to flesh out your thoughts, I wrote the below patch, which\n> does parallel checkout with multiple worker processes.\n\nThanks for your work on this. It is very interesting indeed.\n"},{"id":"283526","messageId":"CACsJy8D-VzOY7UWjM5dNmXkog2L3NN3h3cJSizxC2Rbn-g8RiA@mail.gmail.com","threadId":"41997","inReplyTo":"CAP8UFD0WZHriY340eh3K6ygzb0tXnoT+XaY8+c2k+N2x9UBYxA@mail.gmail.com","subject":"Re: Parallel checkout (Was Re: 0 bot for Git)","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2016-04-15T11:32:49Z","receivedAt":"2016-04-15T11:32:49Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Apr 15, 2016 at 6:18 PM, Christian Couder\n<christian.couder@gmail.com> wrote:\n> On Fri, Apr 15, 2016 at 11:51 AM, Duy Nguyen <pclouds@gmail.com> wrote:\n>> On Fri, Apr 15, 2016 at 12:04:49AM +0200, Christian Couder wrote:\n>>>\n>>> There is a draft of an article about the first part of the Contributor\n>>> Summit in the draft of the next Git Rev News edition:\n>>>\n>>> https://github.com/git/git.github.io/blob/master/rev_news/drafts/edition-14.md\n>>\n>> Thanks. I read the sentence \"This made people mention potential\n>> problems with parallelizing git checkout\" and wondered what these\n>> problems were.\n>\n> It may have been Michael or Peff (CC'ed) saying that it could break\n> some builds as the timestamps on the files might not always be ordered\n> in the same way.\n\nVery subtle. I suppose if we dumb down the distribution algorithm, we\ncould make it stable (even though it won't be the same as serial\ncheckout). Performance will degrade, not sure if it's still worth\nparallelizing at that point\n\n> Now perhaps parallel checkout could be activated only if a config\n> option was set. (Yeah, I know it looks like I am very often asking for\n> config options.)\n\nAnd I think it could be unconditionally activated at clone time (I\ncan't imagine different timestamp order can affect anything at that\npoint), where the number of files to checkout is biggest.\n-- \nDuy\n"},{"id":"283532","messageId":"CAGZ79kZP_TUUk3vWHe=c301n66FtQpnwPfPmJ6oD8n-Zz-SVyg@mail.gmail.com","threadId":"41997","inReplyTo":"20160415095139.GA3985@lanh","subject":"Re: Parallel checkout (Was Re: 0 bot for Git)","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2016-04-15T15:08:03Z","receivedAt":"2016-04-15T15:08:03Z","isPatch":false,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Fri, Apr 15, 2016 at 2:51 AM, Duy Nguyen <pclouds@gmail.com> wrote:\n> On Fri, Apr 15, 2016 at 12:04:49AM +0200, Christian Couder wrote:\n>> On Tue, Apr 12, 2016 at 4:59 PM, Stefan Beller <sbeller@google.com> wrote:\n>> > On Tue, Apr 12, 2016 at 2:42 AM, Duy Nguyen <pclouds@gmail.com> wrote:\n>> >>> On Mon, Apr 11, 2016 at 7:51 AM, Stefan Beller <sbeller@google.com> wrote:\n>> >>>> Hi Greg,\n>> >>>>\n>> >>>> Thanks for your talk at the Git Merge 2016!\n>> >>\n>> >> Huh? It already happened?? Any interesting summary to share with us?\n>>\n>> There is a draft of an article about the first part of the Contributor\n>> Summit in the draft of the next Git Rev News edition:\n>>\n>> https://github.com/git/git.github.io/blob/master/rev_news/drafts/edition-14.md\n>\n> Thanks. I read the sentence \"This made people mention potential\n> problems with parallelizing git checkout\" and wondered what these\n> problems were. And because it's easier to think while you test\n> something, to flesh out your thoughts, I wrote the below patch, which\n> does parallel checkout with multiple worker processes. I wonder if the\n> same set of problems apply to it.\n\nI mentioned it, as Jonathan mentioned it a while ago. (So I was just\nreferring to hear say).\n\n>\n> The idea is simple, you offload some work to process workers. In this\n> patch, only entry.c:write_entry() is moved to workers. We still do\n> directory creation and all sort of checks and stat refresh in the main\n> process. Some more work may be moved away, for example, the entire\n> builtin/checkout.c:checkout_merged().\n>\n> Multi process is less efficient than multi thread model. But I doubt\n> we could make object db access thread-safe soon. The last discussion\n> was 2 years ago [1] and nothing much has happened.\n>\n> Numbers are encouraging though. On linux-2.6 repo running on linux and\n> ext4 filesystem, checkout_paths() would dominate \"git checkout :/\".\n> Unmodified git takes about 31s.\n\nPlease also benchmark \"make build\" or another read heavy operation\nwith these 2 different checkouts. IIRC that was the problem. (checkout\nimproved, but due to file ordering on the fs, the operation afterwards\nslowed down, such that it became a net negative)\n\n>\n>\n> 16:26:00.114029 builtin/checkout.c:1299 performance: 31.184973659 s: checkout_paths\n> 16:26:00.114225 trace.c:420             performance: 31.256412935 s: git command: 'git' 'checkout' '.'\n>\n> When doing write_entry() on 8 processes, it takes 22s (shortened by ~30%)\n>\n> 16:27:39.973730 trace.c:420             performance: 5.610255442 s: git command: 'git' 'checkout-index' '--worker'\n> 16:27:40.956812 trace.c:420             performance: 6.595082013 s: git command: 'git' 'checkout-index' '--worker'\n> 16:27:41.397621 trace.c:420             performance: 7.032024175 s: git command: 'git' 'checkout-index' '--worker'\n> 16:27:47.453999 trace.c:420             performance: 13.078537207 s: git command: 'git' 'checkout-index' '--worker'\n> 16:27:48.986433 trace.c:420             performance: 14.612951643 s: git command: 'git' 'checkout-index' '--worker'\n> 16:27:53.149378 trace.c:420             performance: 18.781762536 s: git command: 'git' 'checkout-index' '--worker'\n> 16:27:54.884044 trace.c:420             performance: 20.514473730 s: git command: 'git' 'checkout-index' '--worker'\n> 16:27:55.319990 trace.c:420             performance: 20.948326263 s: git command: 'git' 'checkout-index' '--worker'\n> 16:27:55.863211 builtin/checkout.c:1299 performance: 22.723118420 s: checkout_paths\n> 16:27:55.863398 trace.c:420             performance: 22.854547640 s: git command: 'git' 'checkout' '--parallel' '.'\n>\n> I suspect on nfs or windows, the gain may be higher due to IO blocking\n> the main process more.\n>\n> Note that this for-fun patch is not optmized at all (and definitely\n> not portable). I could have sent a group of paths to the worker in a\n> single system call instead of one per call. The trace above also shows\n> unbalance issues with workers, where some workers exit early because\n> of my naive work distribution. Numbers could get a bit better.\n>\n> [1] http://thread.gmane.org/gmane.comp.version-control.git/241965/focus=242020\n\nWould it make sense to use the parallel processing infrastructure from\nrun-command.h\ninstead of doing all setup and teardown yourself?\n(As you call it for-fun patch, I'd assume the answer is: Writing code\nis more fun than\nusing other peoples code ;)\n\n\n>\n> -- 8< --\n> diff --git a/builtin/checkout-index.c b/builtin/checkout-index.c\n> index 92c6967..7163216 100644\n> --- a/builtin/checkout-index.c\n> +++ b/builtin/checkout-index.c\n> @@ -9,6 +9,7 @@\n>  #include \"quote.h\"\n>  #include \"cache-tree.h\"\n>  #include \"parse-options.h\"\n> +#include \"entry.h\"\n>\n>  #define CHECKOUT_ALL 4\n>  static int nul_term_line;\n> @@ -179,6 +180,9 @@ int cmd_checkout_index(int argc, const char **argv, const char *prefix)\n>                 OPT_END()\n>         };\n>\n> +       if (argc == 2 && !strcmp(argv[1], \"--worker\"))\n> +               return parallel_checkout_worker();\n> +\n>         if (argc == 2 && !strcmp(argv[1], \"-h\"))\n>                 usage_with_options(builtin_checkout_index_usage,\n>                                    builtin_checkout_index_options);\n> diff --git a/builtin/checkout.c b/builtin/checkout.c\n> index efcbd8f..51caad2 100644\n> --- a/builtin/checkout.c\n> +++ b/builtin/checkout.c\n> @@ -20,6 +20,7 @@\n>  #include \"resolve-undo.h\"\n>  #include \"submodule-config.h\"\n>  #include \"submodule.h\"\n> +#include \"entry.h\"\n>\n>  static const char * const checkout_usage[] = {\n>         N_(\"git checkout [<options>] <branch>\"),\n> @@ -236,7 +237,8 @@ static int checkout_merged(int pos, struct checkout *state)\n>  }\n>\n>  static int checkout_paths(const struct checkout_opts *opts,\n> -                         const char *revision)\n> +                         const char *revision,\n> +                         int parallel)\n>  {\n>         int pos;\n>         struct checkout state;\n> @@ -357,6 +359,8 @@ static int checkout_paths(const struct checkout_opts *opts,\n>         state.force = 1;\n>         state.refresh_cache = 1;\n>         state.istate = &the_index;\n> +       if (parallel)\n> +               start_parallel_checkout(&state);\n>         for (pos = 0; pos < active_nr; pos++) {\n>                 struct cache_entry *ce = active_cache[pos];\n>                 if (ce->ce_flags & CE_MATCHED) {\n> @@ -367,11 +371,18 @@ static int checkout_paths(const struct checkout_opts *opts,\n>                         if (opts->writeout_stage)\n>                                 errs |= checkout_stage(opts->writeout_stage, ce, pos, &state);\n>                         else if (opts->merge)\n> +                               /*\n> +                                * XXX: in parallel mode, we may want\n> +                                * to let worker perform the merging\n> +                                * instead and send SHA-1 result back\n> +                                */\n>                                 errs |= checkout_merged(pos, &state);\n>                         pos = skip_same_name(ce, pos) - 1;\n>                 }\n>         }\n>\n> +       errs |= run_parallel_checkout();\n> +\n>         if (write_locked_index(&the_index, lock_file, COMMIT_LOCK))\n>                 die(_(\"unable to write new index file\"));\n>\n> @@ -1132,6 +1143,7 @@ int cmd_checkout(int argc, const char **argv, const char *prefix)\n>         struct branch_info new;\n>         char *conflict_style = NULL;\n>         int dwim_new_local_branch = 1;\n> +       int parallel = 0;\n>         struct option options[] = {\n>                 OPT__QUIET(&opts.quiet, N_(\"suppress progress reporting\")),\n>                 OPT_STRING('b', NULL, &opts.new_branch, N_(\"branch\"),\n> @@ -1159,6 +1171,8 @@ int cmd_checkout(int argc, const char **argv, const char *prefix)\n>                                 N_(\"second guess 'git checkout <no-such-branch>'\")),\n>                 OPT_BOOL(0, \"ignore-other-worktrees\", &opts.ignore_other_worktrees,\n>                          N_(\"do not check if another worktree is holding the given ref\")),\n> +               OPT_BOOL(0, \"parallel\", &parallel,\n> +                        N_(\"parallel checkout\")),\n>                 OPT_BOOL(0, \"progress\", &opts.show_progress, N_(\"force progress reporting\")),\n>                 OPT_END(),\n>         };\n> @@ -1279,8 +1293,12 @@ int cmd_checkout(int argc, const char **argv, const char *prefix)\n>                 strbuf_release(&buf);\n>         }\n>\n> -       if (opts.patch_mode || opts.pathspec.nr)\n> -               return checkout_paths(&opts, new.name);\n> +       if (opts.patch_mode || opts.pathspec.nr) {\n> +               uint64_t start = getnanotime();\n> +               int ret = checkout_paths(&opts, new.name, parallel);\n> +               trace_performance_since(start, \"checkout_paths\");\n> +               return ret;\n> +       }\n>         else\n>                 return checkout_branch(&opts, &new);\n>  }\n> diff --git a/entry.c b/entry.c\n> index a410957..5e0eb1c 100644\n> --- a/entry.c\n> +++ b/entry.c\n> @@ -3,6 +3,36 @@\n>  #include \"dir.h\"\n>  #include \"streaming.h\"\n>\n> +#include <sys/epoll.h>\n> +#include \"pkt-line.h\"\n> +#include \"argv-array.h\"\n> +#include \"run-command.h\"\n> +\n> +struct checkout_item {\n> +       struct cache_entry *ce;\n> +       struct checkout_item *next;\n> +};\n> +\n> +struct checkout_worker {\n> +       struct child_process cp;\n> +       struct checkout_item *to_complete;\n> +       struct checkout_item *to_send;\n> +};\n> +\n> +struct parallel_checkout {\n> +       struct checkout state;\n> +       struct checkout_worker *workers;\n> +       struct checkout_item *items;\n> +       int nr_items, alloc_items;\n> +       int nr_workers;\n> +};\n> +\n> +static struct parallel_checkout *parallel_checkout;\n> +\n> +static int queue_checkout(struct parallel_checkout *,\n> +                         const struct checkout *,\n> +                         struct cache_entry *);\n> +\n>  static void create_directories(const char *path, int path_len,\n>                                const struct checkout *state)\n>  {\n> @@ -290,5 +320,299 @@ int checkout_entry(struct cache_entry *ce,\n>                 return 0;\n>\n>         create_directories(path.buf, path.len, state);\n> +\n> +       if (!queue_checkout(parallel_checkout, state, ce))\n> +               /*\n> +                * write_entry() will be done by parallel_checkout_worker() in\n> +                * a separate process\n> +                */\n> +               return 0;\n> +\n>         return write_entry(ce, path.buf, state, 0);\n>  }\n> +\n> +int start_parallel_checkout(const struct checkout *state)\n> +{\n> +       if (parallel_checkout)\n> +               die(\"BUG: parallel checkout already initiated\");\n> +       if (0 && state->force)\n> +               die(\"BUG: not support --force yet\");\n> +       parallel_checkout = xmalloc(sizeof(*parallel_checkout));\n> +       memset(parallel_checkout, 0, sizeof(*parallel_checkout));\n> +       memcpy(&parallel_checkout->state, state, sizeof(*state));\n> +\n> +       return 0;\n> +}\n> +\n> +static int queue_checkout(struct parallel_checkout *pc,\n> +                         const struct checkout *state,\n> +                         struct cache_entry *ce)\n> +{\n> +       struct checkout_item *ci;\n> +\n> +       if (!pc ||\n> +           !S_ISREG(ce->ce_mode) ||\n> +           memcmp(&pc->state, state, sizeof(*state)))\n> +               return -1;\n> +\n> +       ALLOC_GROW(pc->items, pc->nr_items + 1, pc->alloc_items);\n> +       ci = pc->items + pc->nr_items++;\n> +       ci->ce = ce;\n> +       return 0;\n> +}\n> +\n> +static int item_cmp(const void *a_, const void *b_)\n> +{\n> +       const struct checkout_item *a = a_;\n> +       const struct checkout_item *b = b_;\n> +       return strcmp(a->ce->name, b->ce->name);\n> +}\n> +\n> +static int setup_workers(struct parallel_checkout *pc, int epoll_fd)\n> +{\n> +       int from, nr_per_worker, i;\n> +\n> +       pc->workers = xmalloc(sizeof(*pc->workers) * pc->nr_workers);\n> +       memset(pc->workers, 0, sizeof(*pc->workers) * pc->nr_workers);\n> +\n> +       nr_per_worker = pc->nr_items / pc->nr_workers;\n> +       from = 0;\n> +\n> +       for (i = 0; i < pc->nr_workers; i++) {\n> +               struct checkout_worker *worker = pc->workers + i;\n> +               struct child_process *cp = &worker->cp;\n> +               struct checkout_item *item;\n> +               struct epoll_event ev;\n> +               int to;\n> +\n> +               to = from + nr_per_worker;\n> +               if (i == pc->nr_workers - 1)\n> +                       to = pc->nr_items;\n> +               item = NULL;\n> +               while (from < to) {\n> +                       pc->items[from].next = item;\n> +                       item = pc->items + from;\n> +                       from++;\n> +               }\n> +               worker->to_send = item;\n> +               worker->to_complete = item;\n> +\n> +               cp->git_cmd = 1;\n> +               cp->in = -1;\n> +               cp->out = -1;\n> +               argv_array_push(&cp->args, \"checkout-index\");\n> +               argv_array_push(&cp->args, \"--worker\");\n> +               if (start_command(cp))\n> +                       die(_(\"failed to run checkout worker\"));\n> +\n> +               ev.events = EPOLLOUT | EPOLLERR | EPOLLHUP;\n> +               ev.data.u32 = i * 2;\n> +               if (epoll_ctl(epoll_fd, EPOLL_CTL_ADD, cp->in, &ev) == -1)\n> +                       die_errno(\"epoll_ctl\");\n> +\n> +               ev.events = EPOLLIN | EPOLLERR | EPOLLHUP;\n> +               ev.data.u32 = i * 2 + 1;\n> +               if (epoll_ctl(epoll_fd, EPOLL_CTL_ADD, cp->out, &ev) == -1)\n> +                       die_errno(\"epoll_ctl\");\n> +       }\n> +       return 0;\n> +}\n> +\n> +static int send_to_worker(struct checkout_worker *worker, int epoll_fd)\n> +{\n> +       if (!worker->to_send) {\n> +               struct epoll_event ev;\n> +\n> +               packet_flush(worker->cp.in);\n> +\n> +               epoll_ctl(epoll_fd, EPOLL_CTL_DEL, worker->cp.in, &ev);\n> +               close(worker->cp.in);\n> +               worker->cp.in = -1;\n> +               return 0;\n> +       }\n> +\n> +       /*\n> +        * XXX: put the fd in non-blocking mode and send as many files\n> +        * as possible in one go.\n> +        */\n> +       packet_write(worker->cp.in, \"%s %s\",\n> +                    sha1_to_hex(worker->to_send->ce->sha1),\n> +                    worker->to_send->ce->name);\n> +       worker->to_send = worker->to_send->next;\n> +       return 0;\n> +}\n> +\n> +int parallel_checkout_worker(void)\n> +{\n> +       struct checkout state;\n> +       struct cache_entry *ce = NULL;\n> +\n> +       memset(&state, 0, sizeof(state));\n> +       /* FIXME: pass 'force' over */\n> +       for (;;) {\n> +               int len;\n> +               unsigned char sha1[20];\n> +               char *line = packet_read_line(0, &len);\n> +\n> +               if (!line)\n> +                       return 0;\n> +\n> +               if (len < 40)\n> +                       return 1;\n> +               if (get_sha1_hex(line, sha1))\n> +                       return 1;\n> +               line += 40;\n> +               len -= 40;\n> +               if (*line != ' ')\n> +                       return 1;\n> +               line++;\n> +               len--;\n> +               if (!ce || ce_namelen(ce) < len) {\n> +                       free(ce);\n> +                       ce = xcalloc(1, cache_entry_size(len));\n> +                       ce->ce_mode = S_IFREG | ce_permissions(0644);\n> +               }\n> +               ce->ce_namelen = len;\n> +               hashcpy(ce->sha1, sha1);\n> +               memcpy(ce->name, line, len + 1);\n> +\n> +               if (write_entry(ce, ce->name, &state, 0))\n> +                       return 1;\n> +               /*\n> +                * XXX process in batch and send bigger number of\n> +                * checked out entries back\n> +                */\n> +               packet_write(1, \"1\");\n> +       }\n> +}\n> +\n> +static int receive_from_worker(struct checkout_worker *worker,\n> +                              int refresh_cache)\n> +{\n> +       int len, val;\n> +       char *line;\n> +\n> +       line = packet_read_line(worker->cp.out, &len);\n> +       val = atoi(line);\n> +       if (val <= 0)\n> +               die(\"BUG: invalid value\");\n> +       while (val && worker->to_complete &&\n> +              worker->to_complete != worker->to_send) {\n> +               if (refresh_cache) {\n> +                       struct stat st;\n> +                       struct cache_entry *ce = worker->to_complete->ce;\n> +\n> +                       lstat(ce->name, &st);\n> +                       fill_stat_cache_info(ce, &st);\n> +                       ce->ce_flags |= CE_UPDATE_IN_BASE;\n> +               }\n> +               worker->to_complete = worker->to_complete->next;\n> +               val--;\n> +       }\n> +       if (val)\n> +               die(\"BUG: invalid value\");\n> +       return 0;\n> +}\n> +\n> +static int finish_worker(struct checkout_worker *worker, int epoll_fd)\n> +{\n> +       struct epoll_event ev;\n> +       char buf[1];\n> +       int ret;\n> +\n> +       assert(worker->to_send == NULL);\n> +       assert(worker->to_complete == NULL);\n> +\n> +       ret = xread(worker->cp.out, buf, sizeof(buf));\n> +       if (ret != 0)\n> +               die(\"BUG: expect eof\");\n> +       epoll_ctl(epoll_fd, EPOLL_CTL_DEL, worker->cp.out, &ev);\n> +       close(worker->cp.out);\n> +       worker->cp.out = -1;\n> +       if (finish_command(&worker->cp))\n> +               die(\"worker had a problem\");\n> +       return 0;\n> +}\n> +\n> +static int really_finished(struct parallel_checkout *pc)\n> +{\n> +       int i;\n> +\n> +       for (i = 0; i < pc->nr_workers; i++)\n> +               if (pc->workers[i].to_complete)\n> +                       return 0;\n> +       return 1;\n> +}\n> +\n> +/* XXX progress support for unpack-trees */\n> +int run_parallel_checkout(void)\n> +{\n> +       struct parallel_checkout *pc = parallel_checkout;\n> +       int ret, i;\n> +       int epoll_fd;\n> +       struct epoll_event *events;\n> +\n> +       if (!pc || !pc->nr_items) {\n> +               free(pc);\n> +               parallel_checkout = NULL;\n> +               return 0;\n> +       }\n> +\n> +       qsort(pc->items, pc->nr_items, sizeof(*pc->items), item_cmp);\n> +       pc->nr_workers = 8;\n> +       epoll_fd = epoll_create(pc->nr_workers * 2);\n> +       if (epoll_fd == -1)\n> +               die_errno(\"epoll_create\");\n> +       ret = setup_workers(pc, epoll_fd);\n> +       events = xmalloc(sizeof(*events) * pc->nr_workers * 2);\n> +\n> +       ret = 0;\n> +       while (!ret) {\n> +               int maybe_all_done = 0, nr;\n> +\n> +               nr = epoll_wait(epoll_fd, events, pc->nr_workers * 2, -1);\n> +               if (nr == -1 && errno == EINTR)\n> +                       continue;\n> +               if (nr == -1) {\n> +                       ret = nr;\n> +                       break;\n> +               }\n> +               for (i = 0; i < nr; i++) {\n> +                       int is_in = events[i].data.u32 & 1;\n> +                       int worker_id = events[i].data.u32 / 2;\n> +                       struct checkout_worker *worker = pc->workers + worker_id;\n> +\n> +                       if (!is_in && (events[i].events & EPOLLOUT))\n> +                               ret = send_to_worker(worker, epoll_fd);\n> +                       else if (events[i].events & EPOLLIN) {\n> +                               if (worker->to_complete) {\n> +                                       int refresh = pc->state.refresh_cache;\n> +                                       ret = receive_from_worker(worker, refresh);\n> +                                       pc->state.istate->cache_changed |= CE_ENTRY_CHANGED;\n> +                               } else {\n> +                                       ret = finish_worker(worker, epoll_fd);\n> +                                       maybe_all_done = 1;\n> +                               }\n> +                       } else if (events[i].events & (EPOLLERR | EPOLLHUP)) {\n> +                               if (is_in && !worker->to_complete) {\n> +                                       ret = finish_worker(worker, epoll_fd);\n> +                                       maybe_all_done = 1;\n> +                               } else\n> +                                       ret = -1;\n> +                       } else\n> +                               die(\"BUG: what??\");\n> +                       if (ret)\n> +                               break;\n> +               }\n> +\n> +               if (maybe_all_done && really_finished(pc))\n> +                       break;\n> +       }\n> +\n> +       close(epoll_fd);\n> +       free(pc->workers);\n> +       free(events);\n> +       free(pc);\n> +       parallel_checkout = NULL;\n> +       return ret;\n> +}\n> diff --git a/unpack-trees.c b/unpack-trees.c\n> index 9f55cc2..433c54e 100644\n> --- a/unpack-trees.c\n> +++ b/unpack-trees.c\n> @@ -220,6 +220,7 @@ static int check_updates(struct unpack_trees_options *o)\n>         remove_marked_cache_entries(&o->result);\n>         remove_scheduled_dirs();\n>\n> +       /* start_parallel_checkout() */\n>         for (i = 0; i < index->cache_nr; i++) {\n>                 struct cache_entry *ce = index->cache[i];\n>\n> @@ -234,6 +235,7 @@ static int check_updates(struct unpack_trees_options *o)\n>                         }\n>                 }\n>         }\n> +       /* run_parallel_checkout() */\n>         stop_progress(&progress);\n>         if (o->update)\n>                 git_attr_set_direction(GIT_ATTR_CHECKIN, NULL);\n> -- 8< --\n"},{"id":"283540","messageId":"20160415165208.GA17928@sigill.intra.peff.net","threadId":"41997","inReplyTo":"CAP8UFD0WZHriY340eh3K6ygzb0tXnoT+XaY8+c2k+N2x9UBYxA@mail.gmail.com","subject":"Re: Parallel checkout (Was Re: 0 bot for Git)","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-04-15T16:52:08Z","receivedAt":"2016-04-15T16:52:08Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Apr 15, 2016 at 01:18:46PM +0200, Christian Couder wrote:\n\n> On Fri, Apr 15, 2016 at 11:51 AM, Duy Nguyen <pclouds@gmail.com> wrote:\n> > On Fri, Apr 15, 2016 at 12:04:49AM +0200, Christian Couder wrote:\n> >>\n> >> There is a draft of an article about the first part of the Contributor\n> >> Summit in the draft of the next Git Rev News edition:\n> >>\n> >> https://github.com/git/git.github.io/blob/master/rev_news/drafts/edition-14.md\n> >\n> > Thanks. I read the sentence \"This made people mention potential\n> > problems with parallelizing git checkout\" and wondered what these\n> > problems were.\n> \n> It may have been Michael or Peff (CC'ed) saying that it could break\n> some builds as the timestamps on the files might not always be ordered\n> in the same way.\n\nI don't think it was me. I'm also not sure how it would break a build.\nGit does not promise a particular timing or order for updating files as\nit is. So if we are checking out two files \"a\" and \"b\", and your build\nprocess depends on the timestamp between them, I think all bets are\nalready off.\n\n-Peff\n"},{"id":"283549","messageId":"xmqqwpnybwxw.fsf@gitster.mtv.corp.google.com","threadId":"41997","inReplyTo":"20160415165208.GA17928@sigill.intra.peff.net","subject":"Re: Parallel checkout (Was Re: 0 bot for Git)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-04-15T17:31:39Z","receivedAt":"2016-04-15T17:31:39Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> On Fri, Apr 15, 2016 at 01:18:46PM +0200, Christian Couder wrote:\n>\n>> On Fri, Apr 15, 2016 at 11:51 AM, Duy Nguyen <pclouds@gmail.com> wrote:\n>> > On Fri, Apr 15, 2016 at 12:04:49AM +0200, Christian Couder wrote:\n>> >>\n>> >> There is a draft of an article about the first part of the Contributor\n>> >> Summit in the draft of the next Git Rev News edition:\n>> >>\n>> >> https://github.com/git/git.github.io/blob/master/rev_news/drafts/edition-14.md\n>> >\n>> > Thanks. I read the sentence \"This made people mention potential\n>> > problems with parallelizing git checkout\" and wondered what these\n>> > problems were.\n>> \n>> It may have been Michael or Peff (CC'ed) saying that it could break\n>> some builds as the timestamps on the files might not always be ordered\n>> in the same way.\n>\n> I don't think it was me. I'm also not sure how it would break a build.\n\nYup, \"will break a build\" is a crazy-talk that I'd be surprised if\nyou said something silly like that ;-)\n\nLast time I checked, I think the accesses to attributes from the\nconvert.c thing was one of the things that are cumbersome to make\nsafe in multi-threaded world.\n"},{"id":"283551","messageId":"20160415173833.GA20014@sigill.intra.peff.net","threadId":"41997","inReplyTo":"xmqqwpnybwxw.fsf@gitster.mtv.corp.google.com","subject":"Re: Parallel checkout (Was Re: 0 bot for Git)","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-04-15T17:38:33Z","receivedAt":"2016-04-15T17:38:33Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Apr 15, 2016 at 10:31:39AM -0700, Junio C Hamano wrote:\n\n> Last time I checked, I think the accesses to attributes from the\n> convert.c thing was one of the things that are cumbersome to make\n> safe in multi-threaded world.\n\nMulti-threaded grep has the same problem. I think we started with a big\nlock on the attribute access. That works OK, as long as you hold the\nlock only for the lookup and not the actual filtering. We later moved to\npre-loading the attributes in 9dd5245c104, because looking up attributes\nin order is much more efficient (because locality of paths lets us reuse\nwork from the previous request).\n\nSo I'm guessing the major work here will be to split the \"look up smudge\nattributes\" step from \"do the smudge\".\n\n-Peff\n"},{"id":"283627","messageId":"CACsJy8BruDfmGvA=q+BW61ZKKsTjrF96VzpujEJidm=OtC0_Rg@mail.gmail.com","threadId":"41997","inReplyTo":"CAGZ79kZP_TUUk3vWHe=c301n66FtQpnwPfPmJ6oD8n-Zz-SVyg@mail.gmail.com","subject":"Re: Parallel checkout (Was Re: 0 bot for Git)","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2016-04-16T00:16:12Z","receivedAt":"2016-04-16T00:16:12Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Apr 15, 2016 at 10:08 PM, Stefan Beller <sbeller@google.com> wrote:\n> On Fri, Apr 15, 2016 at 2:51 AM, Duy Nguyen <pclouds@gmail.com> wrote:\n>> On Fri, Apr 15, 2016 at 12:04:49AM +0200, Christian Couder wrote:\n>> The idea is simple, you offload some work to process workers. In this\n>> patch, only entry.c:write_entry() is moved to workers. We still do\n>> directory creation and all sort of checks and stat refresh in the main\n>> process. Some more work may be moved away, for example, the entire\n>> builtin/checkout.c:checkout_merged().\n>>\n>> Multi process is less efficient than multi thread model. But I doubt\n>> we could make object db access thread-safe soon. The last discussion\n>> was 2 years ago [1] and nothing much has happened.\n>>\n>> Numbers are encouraging though. On linux-2.6 repo running on linux and\n>> ext4 filesystem, checkout_paths() would dominate \"git checkout :/\".\n>> Unmodified git takes about 31s.\n>\n> Please also benchmark \"make build\" or another read heavy operation\n> with these 2 different checkouts. IIRC that was the problem. (checkout\n> improved, but due to file ordering on the fs, the operation afterwards\n> slowed down, such that it became a net negative)\n\nThat's way too close to fs internals. Don't filesystems these days\nhave b-tree and indexes to speed up pathname lookup (which makes file\ncreation order meaningless, I guess)? If it only happens to a fs or\ntwo, I'm leaning to say \"your problem, fix your file system\". A\nmitigation may be let worker handle whole directory (non-recursively)\nso file creation order within a directory is almost the same.\n\n> Would it make sense to use the parallel processing infrastructure from\n> run-command.h\n> instead of doing all setup and teardown yourself?\n> (As you call it for-fun patch, I'd assume the answer is: Writing code\n> is more fun than\n> using other peoples code ;)\n\nI did look at run-command.h. Your run_process_parallel() looked almost\nfit, but I needed control over stdout for coordination, not to be\nprinted. At that point, yes writing new code was more fun than\ntweaking run_process_parallel :-D\n-- \nDuy\n"},{"id":"283638","messageId":"5711CACF.9060204@alum.mit.edu","threadId":"41997","inReplyTo":"20160415165208.GA17928@sigill.intra.peff.net","subject":"Re: Parallel checkout (Was Re: 0 bot for Git)","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2016-04-16T05:17:03Z","receivedAt":"2016-04-16T05:17:03Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 04/15/2016 06:52 PM, Jeff King wrote:\n> On Fri, Apr 15, 2016 at 01:18:46PM +0200, Christian Couder wrote:\n> \n>> On Fri, Apr 15, 2016 at 11:51 AM, Duy Nguyen <pclouds@gmail.com> wrote:\n>>> On Fri, Apr 15, 2016 at 12:04:49AM +0200, Christian Couder wrote:\n>>>>\n>>>> There is a draft of an article about the first part of the Contributor\n>>>> Summit in the draft of the next Git Rev News edition:\n>>>>\n>>>> https://github.com/git/git.github.io/blob/master/rev_news/drafts/edition-14.md\n>>>\n>>> Thanks. I read the sentence \"This made people mention potential\n>>> problems with parallelizing git checkout\" and wondered what these\n>>> problems were.\n>>\n>> It may have been Michael or Peff (CC'ed) saying that it could break\n>> some builds as the timestamps on the files might not always be ordered\n>> in the same way.\n> \n> I don't think it was me. I'm also not sure how it would break a build.\n> Git does not promise a particular timing or order for updating files as\n> it is. So if we are checking out two files \"a\" and \"b\", and your build\n> process depends on the timestamp between them, I think all bets are\n> already off.\n\nI'm hazy on this, but I think somebody at Git Merge pointed out that\nparallel checkouts (within a single repository) could be tricky if\nmultiple Git filenames are mapped to the same file due to filesystem\ncase-insensitivity or encoding normalization.\n\nMichael\n"},{"id":"283650","messageId":"DB5772D2-89D4-4D14-8FD1-4AF6DDFD77AC@gmail.com","threadId":"41997","inReplyTo":"CAGZ79ka4WmT8NjD-04WqwczuCuJZcoKMyDRQKkRH1sT5xoqRhQ@mail.gmail.com","subject":"Re: 0 bot for Git","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-04-16T15:51:11Z","receivedAt":"2016-04-16T15:51:11Z","isPatch":false,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\nOn 13 Apr 2016, at 19:29, Stefan Beller <sbeller@google.com> wrote:\n\n> On Wed, Apr 13, 2016 at 10:09 AM, Lars Schneider\n> <larsxschneider@gmail.com> wrote:\n>> \n>>> On 13 Apr 2016, at 18:27, Junio C Hamano <gitster@pobox.com> wrote:\n>>> \n>>> Lars Schneider <larsxschneider@gmail.com> writes:\n>>> \n>>>> @Junio:\n>>>> If you setup Travis CI for your https://github.com/gitster/git fork\n>>>> then Travis CI would build all your topic branches and you (and\n>>>> everyone who is interested) could check\n>>>> https://travis-ci.org/gitster/git/branches to see which branches\n>>>> will break pu if you integrate them.\n>>> \n>>> I would not say such an arrangement is worthless, but it targets a\n>>> wrong point in the patch flow.\n>>> \n>>> The patches that result in the most wastage of my time (i.e. a\n>>> shared bottleneck resource the community should strive to optimize\n>>> for) are the ones that fail to hit 'pu'.  Ones that do not even\n>>> build in isolation, ones that may build but fail even the new tests\n>>> they bring in, ones that break existing tests, and ones that are OK\n>>> in isolation but do not play well with topics already in flight.\n>> \n>> I am not sure what you mean by \"fail to hit 'pu'\". Maybe we talk at\n>> cross purposes. Here is what I think you do, please correct me:\n>> \n>> 1.) You pick the topics from the mailing list and create feature\n>>    branches for each one of them. E.g. one of my recent topics\n>>    is \"ls/config-origin\".\n> \n> and by You you mean Junio.\nYes.\n\n\n> Ideally the 0bot would have sent the message as a reply to the\n> cover letter with the information \"doesn't compile/breaks test t1234\",\n> so Junio could ignore that series (no time wasted on his part).\n> \n> At Git Merge Greg said (paraphrasing here):\n> \n>  We waste developers time, because we have plenty of it. Maintainers time\n>  however is precious because maintainers are the bottleneck and a scare\n>  resource to come by.\n> \n> And I think Git and the kernel have the same community design here.\n> (Except the kernel is bigger and has more than one maintainer)\n> \n> So the idea is help Junio make a decision to drop/ignore those patches\n> with least amount of brain cycled spent as possible. (Not even spend 5\n> seconds on it).\nThat sounds great. I just wonder how 0bot would know where to apply\nthe patches?\n\n\n>> 2.) At some point you create a new pu branch based on the latest\n>>    next branch. You merge all the new topics into the new pu.\n> \n> but Junio also runs test after each(?) merge(?) of a series and once\n> tests fail, it takes time to sort out, what caused it. (Is that the patch series\n> alone or is that because 2 series interact badly with each other?)\n> \n>> \n>> If you push the topics to github.com/gitster after step 1 then\n>> Travis CI could tell you if the individual topic builds clean\n>> and passes all tests. Then you could merge only clean topics in\n>> step 2 which would result in a pu that is much more likely to\n>> build clean.\n> \n> IIRC Junio did not like granting travis access to the \"blessed\" repository\n> as travis wants so much permissions including write permission to that\n> repo. (We/He could have a second non advertised repo though)\nAFAIK TravisCI does not ask for repo write permissions. They ask for\npermission to write the status of commits (little green checkmark on\nGitHub) and repo hooks:\nhttps://docs.travis-ci.com/user/github-oauth-scopes\n\n\n> Also this would incur wait time on Junios side\n> \n> 1) collect patches (many series over the day)\n> 2) push\n> 3) wait\n> 4) do the merges\nHe could do the merges as he does them today but after some time\nhe (and the contributor of a patch) would know if a certain patch\nbrakes pu.\n\n\n> however a 0 bot would do\n> 1) collect patches faster than Junio (0 bot is a computer after all,\n> working 24/7)\n> 2) test each patch/series individually\n> 3) send feedback without the wait time, so the contributor from a different\n>   time zone gets feedback quickly. (round trip is just the build and test time,\n>   which the developer forgot to do any way if it fails)\nI agree that this would be even better. However, I assume this mechanism\nrequires some setup? TravisCI works today.\n\n\n> \n>> \n>> Could that process avoid wasting your time with bad patches?\n>> \n>>> Automated testing of what is already on 'pu' does not help reduce\n>>> the above cost, as the culling must be done by me _without_ help\n>>> from automated test you propose to run on topics in 'pu'.  Ever\n>>> heard of chicken and egg?\n>>> \n>>> Your \"You can setup your own CI\" update to SubmittingPatches may\n>>> encourage people to test before sending.  The \"Travis CI sends\n>>> failure notice as a response to a crappy patch\" discussed by\n>>> Matthieu in the other subthread will be of great help.\n>>> \n>>> Thanks.\n>>> \n>> \n"},{"id":"283655","messageId":"xmqq60vh77pt.fsf@gitster.mtv.corp.google.com","threadId":"41997","inReplyTo":"DB5772D2-89D4-4D14-8FD1-4AF6DDFD77AC@gmail.com","subject":"Re: 0 bot for Git","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-04-16T18:02:22Z","receivedAt":"2016-04-16T18:02:22Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Lars Schneider <larsxschneider@gmail.com> writes:\n\n>> Also this would incur wait time on Junios side\n>> \n>> 1) collect patches (many series over the day)\n>> 2) push\n>> 3) wait\n>> 4) do the merges\n> He could do the merges as he does them today but after some time\n> he (and the contributor of a patch) would know if a certain patch\n> brakes pu.\n\nRead what you wrote again and realize that your step 1. does not\nrequire any expertise or taste from the person who does so.  Anybody\ncould do it, in other words.  Instead of demanding me to do more of\nmindless chore, why don't you try doing that yourself with your fork\nat GitHub?\n\nI suspect you haven't read my response $gmane/291469 to your message\nyet, but \"as he does them today\" would mean _all_ of the following\nhas to happen during phase 1) above:\n\n - Look at the patch and see if it is even remotely interesting;\n\n - See what maintenance track it should apply to by comparing its\n   context and check availability of features post-image wants to\n   use in the mantenance tracks;\n \n - Fork a topic and apply, and inspect the result with larger -U\n   value (or log -p -W);\n \n - Run tests on the topic.\n\n - Try merging it to the eventual target track (e.g. 'maint-2.7'),\n   'master', 'next' and 'pu' (note that this is not \"one of these\",\n   but \"all of these\"), and build the result (and optionally test).\n   Then discard these trial merges.\n\nTwo things you seem to be missing are:\n\n * I do not pick up patches from the list with the objective of\n   queuing them in 'pu'.  I instead look for and process topics that\n   could go to 'next', or that I want to see in 'next' eventually\n   with fixes.  Queing leftover bits in 'pu' as \"not ready for\n   'next'\" is done only because I saw promises in them (and that\n   determination requires time from me), and did not fail in earlier\n   steps before they even gain a topic branch in my tree (otherwise\n   I wouldn't be able to keep up with the traffic).\n\n * The last step, trial merges, is often a very good method to see\n   potential problems and unintended interactions with other topics.\n   A fix we would want to see in older maintenance tracks may depend\n   on too new a feature we added recently, etc.\n\nAlso see $gmane/291469\n"},{"id":"284103","messageId":"7F130640-40F1-454F-BC00-ACC5364404B8@gmail.com","threadId":"41997","inReplyTo":"xmqq60vh77pt.fsf@gitster.mtv.corp.google.com","subject":"Re: 0 bot for Git","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2016-04-22T08:19:21Z","receivedAt":"2016-04-22T08:19:21Z","isPatch":false,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 16 Apr 2016, at 20:02, Junio C Hamano <gitster@pobox.com> wrote:\n> \n> Lars Schneider <larsxschneider@gmail.com> writes:\n> \n>>> Also this would incur wait time on Junios side\n>>> \n>>> 1) collect patches (many series over the day)\n>>> 2) push\n>>> 3) wait\n>>> 4) do the merges\n>> He could do the merges as he does them today but after some time\n>> he (and the contributor of a patch) would know if a certain patch\n>> brakes pu.\n> \n> Read what you wrote again and realize that your step 1. does not\n> require any expertise or taste from the person who does so.  Anybody\n> could do it, in other words.  Instead of demanding me to do more of\n> mindless chore, why don't you try doing that yourself with your fork\n> at GitHub?\n> \n> I suspect you haven't read my response $gmane/291469 to your message\n> yet, but \"as he does them today\" would mean _all_ of the following\n> has to happen during phase 1) above:\n> \n> - Look at the patch and see if it is even remotely interesting;\n> \n> - See what maintenance track it should apply to by comparing its\n>   context and check availability of features post-image wants to\n>   use in the mantenance tracks;\n> \n> - Fork a topic and apply, and inspect the result with larger -U\n>   value (or log -p -W);\n> \n> - Run tests on the topic.\n> \n> - Try merging it to the eventual target track (e.g. 'maint-2.7'),\n>   'master', 'next' and 'pu' (note that this is not \"one of these\",\n>   but \"all of these\"), and build the result (and optionally test).\n>   Then discard these trial merges.\n> \n> Two things you seem to be missing are:\n> \n> * I do not pick up patches from the list with the objective of\n>   queuing them in 'pu'.  I instead look for and process topics that\n>   could go to 'next', or that I want to see in 'next' eventually\n>   with fixes.  Queing leftover bits in 'pu' as \"not ready for\n>   'next'\" is done only because I saw promises in them (and that\n>   determination requires time from me), and did not fail in earlier\n>   steps before they even gain a topic branch in my tree (otherwise\n>   I wouldn't be able to keep up with the traffic).\n> \n> * The last step, trial merges, is often a very good method to see\n>   potential problems and unintended interactions with other topics.\n>   A fix we would want to see in older maintenance tracks may depend\n>   on too new a feature we added recently, etc.\n> \n> Also see $gmane/291469\n\nThanks for the explanation. My intention was not to be offensive.\nI was curious about your workflow and I was wondering if the\nTravis CI integration could be useful for you in any way.\n\nBest,\nLars\n"},{"id":"284157","messageId":"xmqqr3dxpn4f.fsf@gitster.mtv.corp.google.com","threadId":"41997","inReplyTo":"7F130640-40F1-454F-BC00-ACC5364404B8@gmail.com","subject":"Re: 0 bot for Git","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-04-22T17:30:24Z","receivedAt":"2016-04-22T17:30:24Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Lars Schneider <larsxschneider@gmail.com> writes:\n\n> Thanks for the explanation. My intention was not to be offensive.\n> I was curious about your workflow and I was wondering if the\n> Travis CI integration could be useful for you in any way.\n\nDon't worry; I didn't feel offended.  The Travis stuff running on\nthe branches at http://github.com/git/git would surely catch issues\non MacOSX and/or around git-p4 (neither of which I test myself when\nmerging to 'pu') before they hit 'next', and that is already helping\nus greatly.\n\nAnd if volunteers or bots pick up in-flight patches that have not\nhit 'pu' and feed them to Travis through their repositories, that\nwould also help the project, so your work on hooking up our source\ntree with Travis is greatly appreciated.\n\nIt was just that Travis running on broken-down topic branches that\nappear in http://github.com/gitster/git would not help my workflow,\nwhich was the main point of illustrating the way how these branches\nwork.\n\nThanks.\n\n\n  \n"},{"id":"284240","messageId":"alpine.DEB.2.20.1604240908200.2896@virtualbox","threadId":"41997","inReplyTo":"xmqqr3dxpn4f.fsf@gitster.mtv.corp.google.com","subject":"Re: 0 bot for Git","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2016-04-24T07:15:10Z","receivedAt":"2016-04-24T07:15:10Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Lars & Junio,\n\nOn Fri, 22 Apr 2016, Junio C Hamano wrote:\n\n> Lars Schneider <larsxschneider@gmail.com> writes:\n> \n> > Thanks for the explanation. My intention was not to be offensive.\n> > I was curious about your workflow and I was wondering if the\n> > Travis CI integration could be useful for you in any way.\n> \n> Don't worry; I didn't feel offended.  The Travis stuff running on\n> the branches at http://github.com/git/git would surely catch issues\n> on MacOSX and/or around git-p4 (neither of which I test myself when\n> merging to 'pu') before they hit 'next', and that is already helping\n> us greatly.\n\nI agree that it helps to catch those Mac and P4 issues early.\n\nHowever, it is possible that bogus errors are reported that might not have\nbeen introduced by the changes of the PR, and I find it relatively hard to\nfigure out the specifics. Take for example\n\n\thttps://travis-ci.org/git/git/jobs/124767554\n\nIt appears that t9824 fails with my interactive rebase work on MacOSX,\nboth Clang and GCC versions. I currently have no access to a Mac for\ndeveloping (so I am denied my favorite debugging technique: sh t... -i -v\n-x), and I seem to be unable to find any useful log of what went wrong\n*specifically*.\n\nAny ideas how to find out?\n\nCiao,\nDscho\n"},{"id":"284242","messageId":"1461500361-25913-1-git-send-email-szeder@ira.uka.de","threadId":"41997","inReplyTo":"alpine.DEB.2.20.1604240908200.2896@virtualbox","subject":"Re: 0 bot for Git","fromName":"SZEDER Gábor","fromEmail":"szeder@ira.uka.de","sentAt":"2016-04-24T12:19:21Z","receivedAt":"2016-04-24T12:19:21Z","isPatch":false,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"\n> > Don't worry; I didn't feel offended.  The Travis stuff running on\n> > the branches at http://github.com/git/git would surely catch issues\n> > on MacOSX and/or around git-p4 (neither of which I test myself when\n> > merging to 'pu') before they hit 'next', and that is already helping\n> > us greatly.\n> \n> I agree that it helps to catch those Mac and P4 issues early.\n> \n> However, it is possible that bogus errors are reported that might not have\n> been introduced by the changes of the PR, and I find it relatively hard to\n> figure out the specifics. Take for example\n> \n> \thttps://travis-ci.org/git/git/jobs/124767554\n> \n> It appears that t9824 fails with my interactive rebase work on MacOSX,\n> both Clang and GCC versions. I currently have no access to a Mac for\n> developing (so I am denied my favorite debugging technique: sh t... -i -v\n> -x), and I seem to be unable to find any useful log of what went wrong\n> *specifically*.\n> \n> Any ideas how to find out?\n\nYou could patch .travis.yml on top of your changes to run only t9824\n(so travis-ci returns feedback faster) and to run it with additional\noptions '-i -x', push it to github, and await developments.  I'm not\nsaying it's not cumbersome, because it is, and the p4 cleanup timeout\nloop with '-x' is particularly uninteresting...  but it works, though\nin this case '-x' doesn't output anything useful.\n\nYou could also check how independent branches or even master are\nfaring, and if they fail at the same tests, like in this case they do,\nthen it's not your work that causes the breakage.\n\nIt seems you experience the same breakage that is explained and fixed\nin this thread:\n\n  http://thread.gmane.org/gmane.comp.version-control.git/291917\n"},{"id":"284243","messageId":"alpine.DEB.2.20.1604241504050.2896@virtualbox","threadId":"41997","inReplyTo":"1461500361-25913-1-git-send-email-szeder@ira.uka.de","subject":"Re: 0 bot for Git","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2016-04-24T13:05:08Z","receivedAt":"2016-04-24T13:05:08Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Gábor,\n\nOn Sun, 24 Apr 2016, SZEDER Gábor wrote:\n\n> > > Don't worry; I didn't feel offended.  The Travis stuff running on\n> > > the branches at http://github.com/git/git would surely catch issues\n> > > on MacOSX and/or around git-p4 (neither of which I test myself when\n> > > merging to 'pu') before they hit 'next', and that is already helping\n> > > us greatly.\n> > \n> > I agree that it helps to catch those Mac and P4 issues early.\n> > \n> > However, it is possible that bogus errors are reported that might not\n> > have been introduced by the changes of the PR, and I find it\n> > relatively hard to figure out the specifics. Take for example\n> > \n> > \thttps://travis-ci.org/git/git/jobs/124767554\n> > \n> > It appears that t9824 fails with my interactive rebase work on MacOSX,\n> > both Clang and GCC versions. I currently have no access to a Mac for\n> > developing (so I am denied my favorite debugging technique: sh t... -i\n> > -v -x), and I seem to be unable to find any useful log of what went\n> > wrong *specifically*.\n> > \n> > Any ideas how to find out?\n> \n> You could patch .travis.yml on top of your changes to run only t9824 (so\n> travis-ci returns feedback faster) and to run it with additional options\n> '-i -x', push it to github, and await developments.  I'm not saying it's\n> not cumbersome, because it is, and the p4 cleanup timeout loop with '-x'\n> is particularly uninteresting...  but it works, though in this case '-x'\n> doesn't output anything useful.\n\n*slaps-his-head* of course! Because I can change the build definition...\nThanks!\n\nCiao,\nDscho"},{"id":"284244","messageId":"alpine.DEB.2.20.1604241505290.2896@virtualbox","threadId":"41997","inReplyTo":"CAE5ih78arC2V76XR7yUoXk77c0d_z3Hzupw6MA1+saS3faXjTw@mail.gmail.com","subject":"Re: 0 bot for Git","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2016-04-24T13:05:54Z","receivedAt":"2016-04-24T13:05:54Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Luk,\n\nOn Sun, 24 Apr 2016, Luke Diamand wrote:\n\n> On 24 Apr 2016 08:19, \"Johannes Schindelin\" <Johannes.Schindelin@gmx.de>\n> wrote:\n> >\n> > Hi Lars & Junio,\n> >\n> > On Fri, 22 Apr 2016, Junio C Hamano wrote:\n> >\n> > > Lars Schneider <larsxschneider@gmail.com> writes:\n> > >\n> > > > Thanks for the explanation. My intention was not to be offensive.\n> > > > I was curious about your workflow and I was wondering if the\n> > > > Travis CI integration could be useful for you in any way.\n> > >\n> > > Don't worry; I didn't feel offended.  The Travis stuff running on\n> > > the branches at http://github.com/git/git would surely catch issues\n> > > on MacOSX and/or around git-p4 (neither of which I test myself when\n> > > merging to 'pu') before they hit 'next', and that is already helping\n> > > us greatly.\n> >\n> > I agree that it helps to catch those Mac and P4 issues early.\n> >\n> > However, it is possible that bogus errors are reported that might not have\n> > been introduced by the changes of the PR, and I find it relatively hard to\n> > figure out the specifics. Take for example\n> >\n> >         https://travis-ci.org/git/git/jobs/124767554\n> >\n> > It appears that t9824 fails with my interactive rebase work on MacOSX,\n> > both Clang and GCC versions. I\n> \n> That test is failing because git-lfs has changed its  output format and\n> git-p4 had not yet been taught about this.\n> \n> There's a patch from Lars to fix it -  see the mailing list for details.\n\nYeah, thanks, Gábor provided me with the link.\n\nCiao,\nJohannes"},{"id":"284402","messageId":"alpine.DEB.2.20.1604251603590.2896@virtualbox","threadId":"41997","inReplyTo":"1C553D20-26D9-4BF2-B77E-DEAEDDE869E2@gmail.com","subject":"Re: 0 bot for Git","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2016-04-25T14:07:00Z","receivedAt":"2016-04-25T14:07:00Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Lars,\n\nOn Sun, 24 Apr 2016, Lars Schneider wrote:\n\n> [...] the current Git Travis CI OSX build always installs the latest\n> versions of Git LFS and Perforce via brew [1] and the Linux build\n> installs fixed versions [2].  Consequently new LFS/Perforce versions can\n> brake the OS X build even if there is no change in Git.\n> \n> You could argue that this is bad CI practice because CI results should\n> be reproducible.  However, it has value to test the latest versions to\n> detect integration errors as the one with Git LFS 1.2. We could add\n> additional build jobs to test fixed versions and latest versions but\n> this would just burn a lot of CPU cycles... that was the reason why I\n> chose the way it is implemented right now. I will add a comment to the\n> OSX build to make the current strategy more clear (Linux fixed versions,\n> OS X latest versions).\n> \n> What do you think about that? Do you agree/disagree? Do you see a better\n> way?\n\nI agree with your reasoning.\n\nCiao,\nDscho\n"},{"id":"284521","messageId":"CACsJy8Ab=q0mbdcXn9O7=dKHaOuhUCNk4g6BU5kZHdPM+z7yng@mail.gmail.com","threadId":"41997","inReplyTo":"20160415095139.GA3985@lanh","subject":"Re: Parallel checkout (Was Re: 0 bot for Git)","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2016-04-26T11:35:50Z","receivedAt":"2016-04-26T11:35:50Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Apr 15, 2016 at 4:51 PM, Duy Nguyen <pclouds@gmail.com> wrote:\n> Numbers are encouraging though. On linux-2.6 repo running on linux and\n> ext4 filesystem, checkout_paths() would dominate \"git checkout :/\".\n> Unmodified git takes about 31s.\n>\n>\n> 16:26:00.114029 builtin/checkout.c:1299 performance: 31.184973659 s: checkout_paths\n> 16:26:00.114225 trace.c:420             performance: 31.256412935 s: git command: 'git' 'checkout' '.'\n>\n> When doing write_entry() on 8 processes, it takes 22s (shortened by ~30%)\n\nI continued to develop it into a series. This same laptop now reduces\ncheckout time closer to 50% on linux-2.6. However my other laptop\ngives me the opposite result, parallel checkout takes longer time to\ncomplete. I suspect that only with fast enough disks that CPU may\nbecome temporary bottleneck. This is where parallel checkout shines\nbecause it spreads the load out and quickly moves the bottleneck back\nto I/O (after a while I/O queues should be fully populated again). On\nsystems with slower disks like mine, I/O is always the bottleneck and\nspreading I/O over many processes just makes it worse (probably\nconfuse I/O scheduler more).\n\nSince it's not doing anything for _me_, I'm dropping this. Anybody\ninterested can check it out and maybe try it from parallel-checkout\nbranch [1]. It probably can build on windows (epoll is gone). And it\nprobably help improve performance when smudge filter is used (because\nthat can potentially add more load to cpu). More notes in commit\n8fe9b5c (entry.c: parallel checkout support - 2016-04-18)\n\n[1] https://github.com/pclouds/git/commits/parallel-checkout\n-- \nDuy\n"}]}