{"thread":{"id":"51754","subject":"Git in Outreachy December 2019?","startedAt":"2019-08-27T05:17:59Z","lastAt":"2019-10-22T21:16:48Z","messageCount":63,"participants":["Jeff King","Christian Couder","Olga Telezhnaya","Emily Shaffer","Carlo Arenas","Pratyush Yadav","Jonathan Tan","Eric Wong","SZEDER Gábor","Jonathan Nieder","Johannes Schindelin","Philip Oakley","Junio C Hamano"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"381307","messageId":"20190827051756.GA12795@sigill.intra.peff.net","threadId":"51754","inReplyTo":null,"subject":"Git in Outreachy December 2019?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-08-27T05:17:57Z","receivedAt":"2019-08-27T05:17:59Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"Do we have interested mentors for the next round of Outreachy?\n\nThe deadline for Git to apply to the program is September 5th. The\ndeadline for mentors to have submitted project descriptions is September\n24th. Intern applications would start on October 1st.\n\nIf there are mentors who want to participate, I can handle the project\napplication and can start asking around for funding.\n\n-Peff\n"},{"id":"381621","messageId":"CAP8UFD31Pp9XMDpaNfYP9ph_W0LV43sXvSvppXDjrTSp89S7ZQ@mail.gmail.com","threadId":"51754","inReplyTo":"20190827051756.GA12795@sigill.intra.peff.net","subject":"Re: Git in Outreachy December 2019?","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2019-08-31T07:58:02Z","receivedAt":"2019-08-31T08:01:50Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"On Tue, Aug 27, 2019 at 7:17 AM Jeff King <peff@peff.net> wrote:\n>\n> Do we have interested mentors for the next round of Outreachy?\n\nI am interested to co-mentor.\n\n> The deadline for Git to apply to the program is September 5th. The\n> deadline for mentors to have submitted project descriptions is September\n> 24th. Intern applications would start on October 1st.\n>\n> If there are mentors who want to participate, I can handle the project\n> application and can start asking around for funding.\n\nThat would be really nice, thank you!\n"},{"id":"381627","messageId":"CAL21BmnZr6ubcZOmJebST8e4UccWnPHwDGqLimwrHVCmf61Jmg@mail.gmail.com","threadId":"51754","inReplyTo":"CAP8UFD31Pp9XMDpaNfYP9ph_W0LV43sXvSvppXDjrTSp89S7ZQ@mail.gmail.com","subject":"Re: Git in Outreachy December 2019?","fromName":"Olga Telezhnaya","fromEmail":"olyatelezhnaya@gmail.com","sentAt":"2019-08-31T19:44:53Z","receivedAt":"2019-08-31T19:46:06Z","isPatch":false,"sender":{"key":"olyatelezhnaya@gmail.com","avatar":"https://avatars.githubusercontent.com/u/11246099?v=4"},"body":"сб, 31 авг. 2019 г. в 10:58, Christian Couder <christian.couder@gmail.com>:\n>\n> On Tue, Aug 27, 2019 at 7:17 AM Jeff King <peff@peff.net> wrote:\n> >\n> > Do we have interested mentors for the next round of Outreachy?\n>\n> I am interested to co-mentor.\n\nI am not ready to give the answer right now, but I will definitely\nthink about the project proposals, I want to help with that part. By\nthe way, we can take some projects (with description) from this summer\nround of GSoC. 3 of 4 projects were not selected.\n\n>\n> > The deadline for Git to apply to the program is September 5th. The\n> > deadline for mentors to have submitted project descriptions is September\n> > 24th. Intern applications would start on October 1st.\n> >\n> > If there are mentors who want to participate, I can handle the project\n> > application and can start asking around for funding.\n>\n> That would be really nice, thank you!\n\nThank you!\n"},{"id":"381817","messageId":"20190904194114.GA31398@sigill.intra.peff.net","threadId":"51754","inReplyTo":"20190827051756.GA12795@sigill.intra.peff.net","subject":"Re: Git in Outreachy December 2019?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-09-04T19:41:15Z","receivedAt":"2019-09-04T19:41:17Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Aug 27, 2019 at 01:17:57AM -0400, Jeff King wrote:\n\n> Do we have interested mentors for the next round of Outreachy?\n> \n> The deadline for Git to apply to the program is September 5th. The\n> deadline for mentors to have submitted project descriptions is September\n> 24th. Intern applications would start on October 1st.\n> \n> If there are mentors who want to participate, I can handle the project\n> application and can start asking around for funding.\n\nFunding is still up in the air, but in the meantime I've tentatively\nsigned us up (we have until the 24th to have the funding committed).\nNext we need mentors to submit projects, as well as first-time\ncontribution micro-projects.\n\nProject proposals can be made here:\n\n  https://www.outreachy.org/communities/cfp/git/\n\nIf you want to know more about the program, there's a mentor FAQ here:\n\n  https://www.outreachy.org/mentor/mentor-faq/\n\nor just ask in this thread.\n\nThe project page has a section to point people in the right direction\nfor first-time contributions. I've left it blank for now, but I think it\nmakes sense to point one (or both) of:\n\n  - https://git-scm.com/docs/MyFirstContribution\n\n  - https://matheustavares.gitlab.io/posts/first-steps-contributing-to-git\n\nas well as a list of micro-projects (or at least instructions on how to\nfind #leftoverbits, though we'd definitely have to step up our labeling,\nas I do not recall having seen one for a while).\n\n-Peff\n"},{"id":"381861","messageId":"CAP8UFD19Lop7jLsBcN6dEARCNX19asQKTszFpJ51=TEAMzAp2g@mail.gmail.com","threadId":"51754","inReplyTo":"20190904194114.GA31398@sigill.intra.peff.net","subject":"Re: Git in Outreachy December 2019?","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2019-09-05T07:24:28Z","receivedAt":"2019-09-05T07:24:43Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"On Wed, Sep 4, 2019 at 9:41 PM Jeff King <peff@peff.net> wrote:\n>\n> Funding is still up in the air, but in the meantime I've tentatively\n> signed us up (we have until the 24th to have the funding committed).\n> Next we need mentors to submit projects, as well as first-time\n> contribution micro-projects.\n\nGreat! Thanks for signing up and for all the information!\n"},{"id":"381906","messageId":"20190905193959.GA17913@google.com","threadId":"51754","inReplyTo":"20190904194114.GA31398@sigill.intra.peff.net","subject":"Re: Git in Outreachy December 2019?","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2019-09-05T19:39:59Z","receivedAt":"2019-09-05T19:40:08Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Wed, Sep 04, 2019 at 03:41:15PM -0400, Jeff King wrote:\n> On Tue, Aug 27, 2019 at 01:17:57AM -0400, Jeff King wrote:\n> \n> > Do we have interested mentors for the next round of Outreachy?\n> > \n> > The deadline for Git to apply to the program is September 5th. The\n> > deadline for mentors to have submitted project descriptions is September\n> > 24th. Intern applications would start on October 1st.\n> > \n> > If there are mentors who want to participate, I can handle the project\n> > application and can start asking around for funding.\n> \n> Funding is still up in the air, but in the meantime I've tentatively\n> signed us up (we have until the 24th to have the funding committed).\n> Next we need mentors to submit projects, as well as first-time\n> contribution micro-projects.\n\nI'm interested to mentor too, but I haven't done anything like this -\nofficial mentoring, intern hosting, anything - so I will need to learn\n:)\n\n - Emily\n"},{"id":"381951","messageId":"CAPUEspgyLHSwLBn2EkFyfxuU9KTx+CURTvjmenz2edw-htRxBA@mail.gmail.com","threadId":"51754","inReplyTo":"20190905193959.GA17913@google.com","subject":"Re: Git in Outreachy December 2019?","fromName":"Carlo Arenas","fromEmail":"carenas@gmail.com","sentAt":"2019-09-06T11:55:46Z","receivedAt":"2019-09-06T11:56:00Z","isPatch":false,"sender":{"key":"carenas@gmail.com","avatar":"https://avatars.githubusercontent.com/u/76036?v=4"},"body":"I'm interested to mentor/help too, but I am definitely not a (some\npeople would even argue against \"reliable\") contributor but I might be\nbetter than nothing and could pass my \"lessons learned\" along, so\nhopefully next contributors are less of a pain to deal with than I am\n\nCarlo\n"},{"id":"381999","messageId":"20190907063616.GA28860@sigill.intra.peff.net","threadId":"51754","inReplyTo":"20190905193959.GA17913@google.com","subject":"Re: Git in Outreachy December 2019?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-09-07T06:36:16Z","receivedAt":"2019-09-07T06:36:19Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Sep 05, 2019 at 12:39:59PM -0700, Emily Shaffer wrote:\n\n> > Funding is still up in the air, but in the meantime I've tentatively\n> > signed us up (we have until the 24th to have the funding committed).\n> > Next we need mentors to submit projects, as well as first-time\n> > contribution micro-projects.\n> \n> I'm interested to mentor too, but I haven't done anything like this -\n> official mentoring, intern hosting, anything - so I will need to learn\n> :)\n\nGreat! I think a big challenge is coming up with an appropriately scoped\nproject for the intern. You can take a look at previous Outreachy and\nGSoC proposals[1] to get the general idea, and then think about some\narea you feel comfortable working in. You're welcome to send ideas or\ndrafts to the list to get feedback.\n\nIf you're feeling overwhelmed (or even if you're not), it might make\nsense to try to partner with somebody else as a co-mentor, especially if\nthey've participated in the past. (It might also be that you can offer\nto co-mentor with somebody else's project proposal).\n\n-Peff\n\n[1] Older materials can be found at https://git.github.io/, though\n    these days the proposed projects go directly into the Outreachy\n    system.\n"},{"id":"382000","messageId":"20190907063958.GB28860@sigill.intra.peff.net","threadId":"51754","inReplyTo":"CAPUEspgyLHSwLBn2EkFyfxuU9KTx+CURTvjmenz2edw-htRxBA@mail.gmail.com","subject":"Re: Git in Outreachy December 2019?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-09-07T06:39:59Z","receivedAt":"2019-09-07T06:40:01Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Sep 06, 2019 at 04:55:46AM -0700, Carlo Arenas wrote:\n\n> I'm interested to mentor/help too, but I am definitely not a (some\n> people would even argue against \"reliable\") contributor but I might be\n> better than nothing and could pass my \"lessons learned\" along, so\n> hopefully next contributors are less of a pain to deal with than I am\n\nI just wrote a response to Emily, but I think a lot of it applies to\nyou, as well.\n\nIn particular, I think both of you are a bit newer to the project than\nmost of the other people who have mentored in the past. In some ways\nthat may be a good thing, as it likely makes it easier to see things\nfrom the intern's perspective. :) But it may also introduce some\ncomplications if you're working in an area of the code you're not\nfamiliar with. Co-mentoring may help with that.\n\n-Peff\n"},{"id":"382002","messageId":"CAPUEspgY23L-bjojL1yEftW6WWddZf1ORY+rowxSTYx2c+c=xQ@mail.gmail.com","threadId":"51754","inReplyTo":"20190907063958.GB28860@sigill.intra.peff.net","subject":"Re: Git in Outreachy December 2019?","fromName":"Carlo Arenas","fromEmail":"carenas@gmail.com","sentAt":"2019-09-07T10:13:55Z","receivedAt":"2019-09-07T10:14:10Z","isPatch":false,"sender":{"key":"carenas@gmail.com","avatar":"https://avatars.githubusercontent.com/u/76036?v=4"},"body":"On Fri, Sep 6, 2019 at 11:40 PM Jeff King <peff@peff.net> wrote:\n>\n> I just wrote a response to Emily, but I think a lot of it applies to\n> you, as well.\n\nWith the exception of course that Emily can definitely write better\ncode than my attempted hacks\n\n> In particular, I think both of you are a bit newer to the project than\n> most of the other people who have mentored in the past. In some ways\n> that may be a good thing, as it likely makes it easier to see things\n> from the intern's perspective. :) But it may also introduce some\n> complications if you're working in an area of the code you're not\n> familiar with. Co-mentoring may help with that.\n\nagree, and glad to help any way I can, specially if I can help offset\nsome of the boring work required that would free resources somewhere\nelse; then again by replying to Emily, I didn't meant we were both in\nthe same level both technically and resource wise (after all she wrote\nthe documentation on how to contribute and I only found a silly bug on\nher code that will prevent really obsolete systems to compile and\ndidn't even wrote the patch)\n\nCarlo\n"},{"id":"382046","messageId":"20190908145610.3ho2wo5qqiw3u4lz@yadavpratyush.com","threadId":"51754","inReplyTo":"20190904194114.GA31398@sigill.intra.peff.net","subject":"Re: Git in Outreachy December 2019?","fromName":"Pratyush Yadav","fromEmail":"me@yadavpratyush.com","sentAt":"2019-09-08T14:56:10Z","receivedAt":"2019-09-08T14:56:17Z","isPatch":false,"sender":{"key":"me@yadavpratyush.com","avatar":"https://avatars.githubusercontent.com/u/8817931?v=4"},"body":"Hi Jeff,\n\nOn 04/09/19 03:41PM, Jeff King wrote:\n[snip]\n> The project page has a section to point people in the right direction\n> for first-time contributions. I've left it blank for now, but I think it\n> makes sense to point one (or both) of:\n> \n>   - https://git-scm.com/docs/MyFirstContribution\n> \n>   - https://matheustavares.gitlab.io/posts/first-steps-contributing-to-git\n> \n> as well as a list of micro-projects (or at least instructions on how to\n> find #leftoverbits, though we'd definitely have to step up our labeling,\n> as I do not recall having seen one for a while).\n\nI'd like to put out a proposal regarding first contributions and micro \nprojects.\n\nI have a small list of small isolated features and bug fixes that\n_I think_ git-gui would benefit with. And other people using it can \nprobably add their pet peeves and issues as well. My question is, are \nthese something new contributors should try to work on as an \nintroduction to the community? Since most of these features and fixes \nare small and isolated, they should be pretty easy to work on. And I \nthink people generally find UI apps a little easier to work on.\n\nBut I'll play the devil's advocate on my proposal and point out some \nproblems/flaws:\n- Git-gui is written in Tcl, and git in C (and other languages too, but \n  not Tcl). That means while people do get a feel of the community and \n  general workflow, they don't necessarily get a feel of the actual git \n  internal codebase.\n- Since I don't see a git-gui related project worth being into the \n  Outreachy program, it essentially means they will likely not work on \n  anything related to their project.\n- Git-gui is essentially a wrapper on top of git, so people won't get \n  exposure to the git internals.\n\nI'd like to hear your and the rest of the community's thoughts about \nthis proposal, and whether it will be a good idea or not.\n\nIf people do like this idea, I can do a write up on \"things to fix in \ngit-gui\" that people can add to (and they get a chance to call me stupid \nfor even thinking feature X is a good idea ;)).\n\n-- \nRegards,\nPratyush Yadav\n"},{"id":"382071","messageId":"20190909170002.GA30399@sigill.intra.peff.net","threadId":"51754","inReplyTo":"20190908145610.3ho2wo5qqiw3u4lz@yadavpratyush.com","subject":"Re: Git in Outreachy December 2019?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-09-09T17:00:03Z","receivedAt":"2019-09-09T17:00:05Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Sep 08, 2019 at 08:26:10PM +0530, Pratyush Yadav wrote:\n\n> I'd like to put out a proposal regarding first contributions and micro \n> projects.\n> \n> I have a small list of small isolated features and bug fixes that\n> _I think_ git-gui would benefit with. And other people using it can \n> probably add their pet peeves and issues as well. My question is, are \n> these something new contributors should try to work on as an \n> introduction to the community? Since most of these features and fixes \n> are small and isolated, they should be pretty easy to work on. And I \n> think people generally find UI apps a little easier to work on.\n> \n> But I'll play the devil's advocate on my proposal and point out some \n> problems/flaws:\n> - Git-gui is written in Tcl, and git in C (and other languages too, but \n>   not Tcl). That means while people do get a feel of the community and \n>   general workflow, they don't necessarily get a feel of the actual git \n>   internal codebase.\n> - Since I don't see a git-gui related project worth being into the \n>   Outreachy program, it essentially means they will likely not work on \n>   anything related to their project.\n> - Git-gui is essentially a wrapper on top of git, so people won't get \n>   exposure to the git internals.\n> \n> I'd like to hear your and the rest of the community's thoughts about \n> this proposal, and whether it will be a good idea or not.\n\nRight, I came up with similar devil's advocate arguments. :) I'm not\ntotally opposed, because part of the point of these microprojects just\ngetting people familiar with interacting with the community and\nsubmitting a patch. They're not always in the same area the intern\nintends to work, just because there's not always a trivial problem to be\nsolved there.\n\nSo we do look at it mostly as a \"can you do this basic test\" test, and\nnot necessarily as a prelude to the project.\n\nBut it would be nice if it were at least in the same _language_ that the\nultimate project will be done in. Because we're evaluating the\napplicant's ability to write code in that language, too.\n\nSo I dunno. I am on the fence.\n\n-Peff\n"},{"id":"382325","messageId":"20190913200317.68440-1-jonathantanmy@google.com","threadId":"51754","inReplyTo":"20190827051756.GA12795@sigill.intra.peff.net","subject":"Re: Git in Outreachy December 2019?","fromName":"Jonathan Tan","fromEmail":"jonathantanmy@google.com","sentAt":"2019-09-13T20:03:17Z","receivedAt":"2019-09-13T20:03:24Z","isPatch":false,"sender":{"key":"jonathantanmy@fastmail.com","avatar":null},"body":"> Do we have interested mentors for the next round of Outreachy?\n> \n> The deadline for Git to apply to the program is September 5th. The\n> deadline for mentors to have submitted project descriptions is September\n> 24th. Intern applications would start on October 1st.\n> \n> If there are mentors who want to participate, I can handle the project\n> application and can start asking around for funding.\n\nI probably should have replied earlier, but if Git has applied to the\nprogram, feel free to include me as a mentor.\n\nThere was a discussion about mentors/co-mentors possibly working in a\npart of a codebase that they are not familiar with [1] - firstly, I\nthink that's possible and even likely for most of us. :-) If any\nquestion arises, maybe it would be sufficient for the mentors to just\nhelp formulate the question (or pose the question themselves) to the\nmailing list. If \"[Outreachy]\" appears in the subject, I'll make it a\nhigher priority for myself to answer those.\n\n[1] https://public-inbox.org/git/20190907063958.GB28860@sigill.intra.peff.net/\n"},{"id":"382336","messageId":"20190913205148.GA8799@sigill.intra.peff.net","threadId":"51754","inReplyTo":"20190913200317.68440-1-jonathantanmy@google.com","subject":"Re: Git in Outreachy December 2019?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-09-13T20:51:49Z","receivedAt":"2019-09-13T20:51:51Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Sep 13, 2019 at 01:03:17PM -0700, Jonathan Tan wrote:\n\n> > Do we have interested mentors for the next round of Outreachy?\n> > \n> > The deadline for Git to apply to the program is September 5th. The\n> > deadline for mentors to have submitted project descriptions is September\n> > 24th. Intern applications would start on October 1st.\n> > \n> > If there are mentors who want to participate, I can handle the project\n> > application and can start asking around for funding.\n> \n> I probably should have replied earlier, but if Git has applied to the\n> program, feel free to include me as a mentor.\n\nGreat!  See my followup here:\n\n  https://public-inbox.org/git/20190904194114.GA31398@sigill.intra.peff.net/\n\nProspective mentors need to sign up on that site, and should propose a\nproject they'd be willing to mentor.\n\n> There was a discussion about mentors/co-mentors possibly working in a\n> part of a codebase that they are not familiar with [1] - firstly, I\n> think that's possible and even likely for most of us. :-) If any\n> question arises, maybe it would be sufficient for the mentors to just\n> help formulate the question (or pose the question themselves) to the\n> mailing list. If \"[Outreachy]\" appears in the subject, I'll make it a\n> higher priority for myself to answer those.\n\nI do think it's OK for mentors to not be intimately familiar with the\npart of the code that is being touched, as long as the project is simple\nenough that they can pick up the technical details easily as-needed. A\nlot of what mentors will help mentees with is the overall process (both\nGit-specific parts, but also more general development issues). But I\nthink the proposed projects do need to be feasible.\n\nI'm happy to discuss possible projects if anybody has an idea but isn't\nsure how to develop it into a proposal.\n\n-Peff\n"},{"id":"382438","messageId":"20190916184208.GB17913@google.com","threadId":"51754","inReplyTo":"20190913205148.GA8799@sigill.intra.peff.net","subject":"Re: Git in Outreachy December 2019?","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2019-09-16T18:42:08Z","receivedAt":"2019-09-16T18:42:16Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Fri, Sep 13, 2019 at 04:51:49PM -0400, Jeff King wrote:\n> On Fri, Sep 13, 2019 at 01:03:17PM -0700, Jonathan Tan wrote:\n> \n> > > Do we have interested mentors for the next round of Outreachy?\n> > > \n> > > The deadline for Git to apply to the program is September 5th. The\n> > > deadline for mentors to have submitted project descriptions is September\n> > > 24th. Intern applications would start on October 1st.\n> > > \n> > > If there are mentors who want to participate, I can handle the project\n> > > application and can start asking around for funding.\n> > \n> > I probably should have replied earlier, but if Git has applied to the\n> > program, feel free to include me as a mentor.\n> \n> Great!  See my followup here:\n> \n>   https://public-inbox.org/git/20190904194114.GA31398@sigill.intra.peff.net/\n> \n> Prospective mentors need to sign up on that site, and should propose a\n> project they'd be willing to mentor.\n> \n> > There was a discussion about mentors/co-mentors possibly working in a\n> > part of a codebase that they are not familiar with [1] - firstly, I\n> > think that's possible and even likely for most of us. :-) If any\n> > question arises, maybe it would be sufficient for the mentors to just\n> > help formulate the question (or pose the question themselves) to the\n> > mailing list. If \"[Outreachy]\" appears in the subject, I'll make it a\n> > higher priority for myself to answer those.\n> \n> I do think it's OK for mentors to not be intimately familiar with the\n> part of the code that is being touched, as long as the project is simple\n> enough that they can pick up the technical details easily as-needed. A\n> lot of what mentors will help mentees with is the overall process (both\n> Git-specific parts, but also more general development issues). But I\n> think the proposed projects do need to be feasible.\n> \n> I'm happy to discuss possible projects if anybody has an idea but isn't\n> sure how to develop it into a proposal.\n\nHi Peff,\n\nJonathan Tan, Jonathan Nieder, Josh Steadmon and I met on Friday to talk\nabout projects and we came up with a trimmed list; not sure what more\nneeds to be done to make them into fully-fledged proposals.\n\nFor starter microprojects, we came up with:\n\n - cleanup a test script (although we need to identify particularly\n   which ones and what counts as \"clean\")\n - moving doc from documentation/technical/api-* to comments in the\n   appropriate header instead\n - teach a command which currently handles its own argv how to use\n   parse-options instead\n - add a user.timezone option which Git can use if present rather than\n   checking system local time\n\nFor the longer projects, we came up with a few more:\n\n - find places where we can pass in the_repository as arg instead of\n   using global the_repository\n - convert sh/pl commands to C, including:\n   - git-submodules.sh\n   - git-bisect.sh\n   - rebase --preserve-merges\n   - add -i\n   (We were afraid this might be too boring, though.)\n - reduce/eliminate use of fetch_if_missing global\n - create a better difftool/mergetool for format of choice (this one\n   ends up existing outside of the Git codebase, but still may be pretty\n   adjacent and big impact)\n - training wheels/intro/tutorial mode? (We thought it may be useful to\n   make available a very basic \"I just want to make a single PR and not\n   learn graph theory\" mode, toggled by config switch)\n - \"did you mean?\" for common use cases, e.g. commit with a dirty\n   working tree and no staged files - either offer a hint or offer a\n   prompt to continue (\"Stage changed files and commit? [Y/n]\")\n - new `git partial-clone` command to interactively set a filter,\n   configure other partial clone settings\n - add progress bars in various situations\n - add a TUI to deal more easily with the mailing list. Jonathan Tan has\n   a strong idea of what this TUI would do... This one would also end up\n   external but adjacent to the Git codebase.\n - try and make progress towards running many tests from a single test\n   file in parallel - maybe this is too big, I'm not sure if we know how\n   many of our tests are order-dependent within a file for now...\n\nIt might make sense to only focus on scoping the ones we feel most\ninterested in. We came up with a pretty big list because we had some\nother programs in mind, so I suppose it's not necessary to develop all\nof them for this program.\n\n - Emily\n"},{"id":"382463","messageId":"20190916213301.mybxocvdhdhd7xlg@whir","threadId":"51754","inReplyTo":"20190916184208.GB17913@google.com","subject":"Re: Git in Outreachy December 2019?","fromName":"Eric Wong","fromEmail":"e@80x24.org","sentAt":"2019-09-16T21:33:01Z","receivedAt":"2019-09-16T21:33:04Z","isPatch":false,"sender":{"key":"e@80x24.org","avatar":null},"body":"Emily Shaffer <emilyshaffer@google.com> wrote:\n> Jonathan Tan, Jonathan Nieder, Josh Steadmon and I met on Friday to talk\n> about projects and we came up with a trimmed list; not sure what more\n> needs to be done to make them into fully-fledged proposals.\n\n<snip>\n\n> For the longer projects, we came up with a few more:\n\n<snip>\n\n>  - add a TUI to deal more easily with the mailing list. Jonathan Tan has\n>    a strong idea of what this TUI would do... This one would also end up\n>    external but adjacent to the Git codebase.\n\nAFAIK, Konstantin is/was interested in exploring some of these\nideas with Linux Foundation, too (but he's on vacation atm)\n\n<snip>\n\n> It might make sense to only focus on scoping the ones we feel most\n> interested in. We came up with a pretty big list because we had some\n> other programs in mind, so I suppose it's not necessary to develop all\n> of them for this program.\n"},{"id":"382464","messageId":"20190916214452.GC6190@szeder.dev","threadId":"51754","inReplyTo":"20190916184208.GB17913@google.com","subject":"Re: Git in Outreachy December 2019?","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2019-09-16T21:44:52Z","receivedAt":"2019-09-16T21:44:58Z","isPatch":false,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Mon, Sep 16, 2019 at 11:42:08AM -0700, Emily Shaffer wrote:\n>  - try and make progress towards running many tests from a single test\n>    file in parallel - maybe this is too big, I'm not sure if we know how\n>    many of our tests are order-dependent within a file for now...\n\nForget it, too many (most?) of them are order-dependent.\n\n"},{"id":"382470","messageId":"20190916231348.GB67467@google.com","threadId":"51754","inReplyTo":"20190916214452.GC6190@szeder.dev","subject":"Re: Git in Outreachy December 2019?","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2019-09-16T23:13:49Z","receivedAt":"2019-09-16T23:13:53Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"SZEDER Gábor wrote:\n> On Mon, Sep 16, 2019 at 11:42:08AM -0700, Emily Shaffer wrote:\n\n>>  - try and make progress towards running many tests from a single test\n>>    file in parallel - maybe this is too big, I'm not sure if we know how\n>>    many of our tests are order-dependent within a file for now...\n>\n> Forget it, too many (most?) of them are order-dependent.\n\nHm, I remember a conversation about this with Thomas Rast a while ago.\nIt seemed possible at the time.\n\nMost tests use \"setup\" or \"set up\" in the names of test assertions\nthat are required by later tests.  It's very helpful for debugging and\nmaintenance to be able to skip or reorder some tests, so I've been\nable to rely on this a bit.  Of course there's no automated checking\nin place for that, so there are plenty of test scripts that are\nexceptions to it.\n\nIf we introduce a test_setup helper, then we would not have to rely on\nconvention any more.  A test_setup test assertion would represent a\n\"barrier\" that all later tests in the file can rely on.  We could\nintroduce some automated checking that these semantics are respected,\nand then we get a maintainability improvement in every test script\nthat uses test_setup.  (In scripts without any test_setup, treat all\ntest assertions as barriers since they haven't been vetted.)\n\nWith such automated tests in place, we can then try updating all tests\nthat say \"setup\" or \"set up\" to use test_setup and see what fails.\n\nSome other tests cannot run in parallel for other reasons (e.g. HTTP\ntests).  These can be declared as such, and then we have the ability\nto run arbitrary individual tests in parallel.\n\nMost of the time in a test run involves multiple test scripts running\nin parallel already, so this isn't a huge win for the time to complete\na normal test run.  It helps more with expensive runs like --valgrind.\n\nThanks,\nJonathan\n"},{"id":"382473","messageId":"20190917005928.GA27926@sigill.intra.peff.net","threadId":"51754","inReplyTo":"20190916231348.GB67467@google.com","subject":"Re: Git in Outreachy December 2019?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-09-17T00:59:28Z","receivedAt":"2019-09-17T00:59:30Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Sep 16, 2019 at 04:13:49PM -0700, Jonathan Nieder wrote:\n\n> Most tests use \"setup\" or \"set up\" in the names of test assertions\n> that are required by later tests.  It's very helpful for debugging and\n> maintenance to be able to skip or reorder some tests, so I've been\n> able to rely on this a bit.  Of course there's no automated checking\n> in place for that, so there are plenty of test scripts that are\n> exceptions to it.\n> \n> If we introduce a test_setup helper, then we would not have to rely on\n> convention any more.  A test_setup test assertion would represent a\n> \"barrier\" that all later tests in the file can rely on.  We could\n> introduce some automated checking that these semantics are respected,\n> and then we get a maintainability improvement in every test script\n> that uses test_setup.  (In scripts without any test_setup, treat all\n> test assertions as barriers since they haven't been vetted.)\n> \n> With such automated tests in place, we can then try updating all tests\n> that say \"setup\" or \"set up\" to use test_setup and see what fails.\n> \n> Some other tests cannot run in parallel for other reasons (e.g. HTTP\n> tests).  These can be declared as such, and then we have the ability\n> to run arbitrary individual tests in parallel.\n\nThis isn't quite the same, but couldn't we get most of the gain just by\nsplitting the tests into more scripts? As you note, we already run those\nin parallel, so it increases the granularity of our parallelism. And you\ndon't have to worry about skipping tests 1 through 18 if they're in\nanother file; you just don't consider them at all. It also Just Works\nwith things like HTTP, which choose ports under the assumption that the\nother tests are running simultaneously.\n\nIt doesn't help with the case that test 1 does setup, and then tests 2,\n3, and 4 are logically independent (and some could be skipped or not).\n\nIf anybody is interested in splitting up scripts, the obvious ones to\nlook at are the ones that take the longest (t9001 takes 55s on my\nsystem, though the whole suite runs in only 95s). Of course you can get\nmost of the parallelism benefit by using \"prove --state=slow,save\",\nwhich ends up with lots of short scripts at the end (rather than one\nslow one chewing one CPU while the rest sit idle).\n\n> Most of the time in a test run involves multiple test scripts running\n> in parallel already, so this isn't a huge win for the time to complete\n> a normal test run.  It helps more with expensive runs like --valgrind.\n\nTwo easier suggestions than trying to make --valgrind faster:\n\n  - use SANITIZE=address, which is way cheaper than valgrind (and\n    catches more things!)\n\n  - use --valgrind-only=17 to run everything else in \"fast\" mode, but\n    check the test you care about\n\n-Peff\n"},{"id":"382482","messageId":"nycvar.QRO.7.76.6.1909171158090.15067@tvgsbejvaqbjf.bet","threadId":"51754","inReplyTo":"20190916184208.GB17913@google.com","subject":"Re: Git in Outreachy December 2019?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-09-17T11:23:18Z","receivedAt":"2019-09-17T11:23:47Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Emily,\n\nOn Mon, 16 Sep 2019, Emily Shaffer wrote:\n\n> Jonathan Tan, Jonathan Nieder, Josh Steadmon and I met on Friday to\n> talk about projects and we came up with a trimmed list; not sure what\n> more needs to be done to make them into fully-fledged proposals.\n\nThank you for doing this!\n\n> For starter microprojects, we came up with:\n>\n>  - cleanup a test script (although we need to identify particularly\n>    which ones and what counts as \"clean\")\n>  - moving doc from documentation/technical/api-* to comments in the\n>    appropriate header instead\n>  - teach a command which currently handles its own argv how to use\n>    parse-options instead\n>  - add a user.timezone option which Git can use if present rather than\n>    checking system local time\n\nNice projects, all. There are a couple more ideas on\nhttps://github.com/gitgitgadget/git/issues, they could probably use some\ntagging.\n\n> For the longer projects, we came up with a few more:\n>\n>  - find places where we can pass in the_repository as arg instead of\n>    using global the_repository\n\nGood project, if a bit boring ;-) Also, `the_index` is used a lot in\n`builtin/*.c`, still.\n\n>  - convert sh/pl commands to C, including:\n>    - git-submodules.sh\n\nI am of two minds there. But mostly, I am of the \"friends don't let\nfriends use submodules\" camp, so I would not even want to mentor for\nthis project: it just makes me shudder too much every time I have to\nwork with/on submodules.\n\nOf course, if others are interested, I'd hardly object to turn this into\na built-in.\n\n>    - git-bisect.sh\n\nThat would be my top recommendation, especially given how much effort\nTanushree put in last winter to make this conversion to C so much more\nachievable than before.\n\n>    - rebase --preserve-merges\n\nNo. `rebase -p` is already deprecated in favor of `rebase -r` (which\n_is_ already built-in).\n\nI already have patches lined up to drop that rebase backend. Let's not\nwaste effort on converting this script to C.\n\n>    - add -i\n\nPlease see PRs #170-#175 on https://github.com/gitgitgadget/git/pulls,\nand please do help by adding your review of #170 (which was already\nsubmitted as v4:\nhttps://public-inbox.org/git/pull.170.v4.git.gitgitgadget@gmail.com/\n\nIn other words: this project is well under way. In fact, Git for Windows\nusers enjoy this as an opt-in already.\n\n>    (We were afraid this might be too boring, though.)\n\nConverting shell/Perl scripts into built-in C never looks as much fun as\nopen-ended projects with lots of playing around, but the advantage of\nthe former is that they can be easily structured, offer a lot of\nopportunity for learning, and they are ultimately more rewarding because\nthe goals are much better defined than many other projects'.\n\nAnother script that would _really_ benefit from being converted to C:\n`mergetool`. Especially on Windows, where the over-use of spawned\nprocesses really hurts, it is awfully slow a command.\n\nTo complete the list of sh/pl commands:\n\n- git-merge-octopus.sh\n- git-merge-one-file.sh\n- git-merge-resolve.sh\n\nThese seem to be good candidates for conversion to built-ins. Their\nfunctionality is well-exercised in the test suite, their complexity is\nquite manageable, and there is no good reason that these should be\nscripted.\n\nThe only slightly challenging aspect might be that `merge-one-file` is\nactually not a merge strategy, but it is used as helper to be passed to\n`git merge-index` via the `-o <helper>` option, which makes it slightly\nawkward to be implemented as a built-in. A better approach would\ntherefore be to special-case this value in `merge-index` and execute the\nC code directly, without the detour of spawning a built-in.\n\n- git-difftool--helper.sh\n- git-mergetool--lib.sh\n\nThese would be converted as part of making `mergetool` a built-in, I\nbelieve.\n\n- git-filter-branch.sh\n\nThis one is in the process of being deprecated in favor of `git\nfilter-repo` (which is an external tool), so I don't think there would\nbe much use in wasting energy on trying to convert it to C. Especially\ngiven that it wants to call shell script snippets all over the place,\nand those shell script snippets are supposed to run in the same context,\nwhich might actually make it completely impossible to convert this to C\nat all.\n\n- git-legacy-stash.sh\n\nThis will go away once the built-in stash is considered good enough.\n\n- git-instaweb.sh\n- git-request-pull.sh\n- git-send-email.perl\n- git-web--browse.sh\n\nI don't think that any of these should be converted. They are just too\nunimportant from a performance point of view, and obscure enough that\neven their portability issues don't matter too much.\n\nAs to `send-email` in particular: I would not want anybody to drag in\nall the dependencies required to convert `send-email` to a built-in to\nbegin with.\n\n- git-archimport.perl\n- git-cvsexportcommit.perl\n- git-cvsimport.perl\n- git-cvsserver.perl\n- git-quiltimport.sh\n- git-svn.perl\n\nThese are all connectors of some sort to other version control software.\nIt also feels like they become less and less important, as Git really\ntakes over the world.\n\nAt some stage, I think, it would make sense to push those scripts out\ninto their own repositories, looking for new maintainers (and if none\ncan be found, then there really is not enough need for them to begin\nwith, and they can be archived).\n\n>  - reduce/eliminate use of fetch_if_missing global\n>  - create a better difftool/mergetool for format of choice (this one\n>    ends up existing outside of the Git codebase, but still may be pretty\n>    adjacent and big impact)\n>  - training wheels/intro/tutorial mode? (We thought it may be useful to\n>    make available a very basic \"I just want to make a single PR and not\n>    learn graph theory\" mode, toggled by config switch)\n>  - \"did you mean?\" for common use cases, e.g. commit with a dirty\n>    working tree and no staged files - either offer a hint or offer a\n>    prompt to continue (\"Stage changed files and commit? [Y/n]\")\n>  - new `git partial-clone` command to interactively set a filter,\n>    configure other partial clone settings\n>  - add progress bars in various situations\n>  - add a TUI to deal more easily with the mailing list. Jonathan Tan has\n>    a strong idea of what this TUI would do... This one would also end up\n>    external but adjacent to the Git codebase.\n\nI don't think that this would be a good project for anybody except\npeople who are already really, really familiar with our mailing\nlist-centric workflow.\n\n>  - try and make progress towards running many tests from a single test\n>    file in parallel - maybe this is too big, I'm not sure if we know how\n>    many of our tests are order-dependent within a file for now...\n\nAnother, potentially more rewarding, project would be to modernize our\ntest suite framework, so that it is not based on Unix shell scripting,\nbut on C instead.\n\nThe fact that it is based on Unix shell scripting not only costs a lot\nof speed, especially on Windows, it also limits us quite a bit, and I am\ntalking about a lot more than just the awkwardness of having to think\nabout options of BSD vs GNU variants of common command-line tools.\n\nFor example, many, many, if not all, test cases, spend the majority of\ntheir code on setting up specific scenarios. I don't know about you,\nbut personally I have to dive into many of them when things fail (and I\n_dread_ the numbers 0021, 0025 and 3070, let me tell you) and I really\nhave to say that most of that code is hard to follow and does not make\nit easy to form a mental model of what the code tries to accomplish.\n\nTo address this, a while ago Thomas Rast started to use `fast-export`ed\ncommit histories in test scripts (see e.g. `t/t3206/history.export`). I\nstill find that this fails to make it easier for occasional readers to\nunderstand the ideas underlying the test cases.\n\nAnother approach is to document heavily the ideas first, then use code\nto implement them. For example, t3430 starts with this:\n\n\t[...]\n\n\tInitial setup:\n\n\t    -- B --                   (first)\n\t   /       \\\n\t A - C - D - E - H            (master)\n\t   \\    \\       /\n\t    \\    F - G                (second)\n\t     \\\n\t      Conflicting-G\n\n\t[...]\n\n\ttest_commit A &&\n\tgit checkout -b first &&\n\ttest_commit B &&\n\tgit checkout master &&\n\ttest_commit C &&\n\ttest_commit D &&\n\tgit merge --no-commit B &&\n\ttest_tick &&\n\tgit commit -m E &&\n\tgit tag -m E E &&\n\tgit checkout -b second C &&\n\ttest_commit F &&\n\ttest_commit G &&\n\tgit checkout master &&\n\tgit merge --no-commit G &&\n\ttest_tick &&\n\tgit commit -m H &&\n\tgit tag -m H H &&\n\tgit checkout A &&\n\ttest_commit conflicting-G G.t\n\n\t[...]\n\nWhile this is _somewhat_ better than having only the code, I am still\nunhappy about it: this wall of `test_commit` lines interspersed with\nother commands is very hard to follow.\n\nIf we were to (slowly) convert our test suite framework to C, we could\nchange that.\n\nOne idea would be to allow recreating commit history from something that\nlooks like the output of `git log`, or even `git log --graph --oneline`,\nmuch like `git mktree` (which really should have been a test helper\ninstead of a Git command, but I digress) takes something that looks like\nthe output of `git ls-tree` and creates a tree object from it.\n\nAnother thing that would be much easier if we moved more and more parts\nof the test suite framework to C: we could implement more powerful\nassertions, a lot more easily. For example, the trace output of a failed\n`test_i18ngrep` (or `mingw_test_cmp`!!!) could be made a lot more\nfocused on what is going wrong than on cluttering the terminal window\nwith almost useless lines which are tedious to sift through.\n\nLikewise, having a framework in C would make it a lot easier to improve\ndebugging, e.g. by making test scripts \"resumable\" (guarded by an\noption, it could store a complete state, including a copy of the trash\ndirectory, before executing commands, which would allow \"going back in\ntime\" and calling a failing command with a debugger, or with valgrind, or\njust seeing whether the command would still fail, i.e. whether the test\ncase is flaky).\n\nAlso, things like the code tracing via `-x` (which relies on Bash\nfunctionality in order to work properly, and which _still_ does not work\nas intended if your test case evaluates a lazy prereq that has not been\nevaluated before) could be \"done right\".\n\nIn many ways, our current test suite seems to test Git's functionality\nas much as (core) contributors' abilities to implement test cases in\nUnix shell script, _correctly_, and maybe also contributors' patience.\nYou could say that it tests for the wrong thing at least half of the\ntime, by design.\n\nIt might look like a somewhat less important project, but given that we\nexercise almost 150,000 test cases with every CI build, I think it does\nmake sense to grind our axe for a while, so to say.\n\nTherefore, it might be a really good project to modernize our test\nsuite. To take ideas from modern test frameworks such as Jest and try to\nbring them to C. Which means that new contributors would probably be\nbetter suited to work on this project than Git old-timers!\n\nAnd the really neat thing about this project is that it could be done\nincrementally.\n\n> It might make sense to only focus on scoping the ones we feel most\n> interested in. We came up with a pretty big list because we had some\n> other programs in mind, so I suppose it's not necessary to develop all\n> of them for this program.\n\nI don't find that list particularly big, to be honest ;-)\n\nCiao,\nDscho\n"},{"id":"382484","messageId":"20190917120230.GA27531@szeder.dev","threadId":"51754","inReplyTo":"nycvar.QRO.7.76.6.1909171158090.15067@tvgsbejvaqbjf.bet","subject":"Re: Git in Outreachy December 2019?","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2019-09-17T12:02:44Z","receivedAt":"2019-09-17T12:02:56Z","isPatch":false,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Tue, Sep 17, 2019 at 01:23:18PM +0200, Johannes Schindelin wrote:\n> Also, things like the code tracing via `-x` (which relies on Bash\n> functionality in order to work properly,\n\nNot really.\n\n> and which _still_ does not work\n> as intended if your test case evaluates a lazy prereq that has not been\n> evaluated before\n\nI don't see any striking differences between the trace output of a test\ninvolving a lazy prereq from Bash or dash:\n\n  $ cat t9999-test.sh \n  #!/bin/sh\n  \n  test_description='test'\n  \n  . ./test-lib.sh\n  \n  test_lazy_prereq DUMMY_PREREQ '\n          : lazily evaluating a dummy prereq\n  '\n  \n  test_expect_success DUMMY_PREREQ 'test' '\n          true\n  '\n  \n  test_done\n  $ ./t9999-test.sh -x\n  Initialized empty Git repository in /home/szeder/src/git/t/trash directory.t9999-test/.git/\n  checking prerequisite: DUMMY_PREREQ\n  \n  mkdir -p \"$TRASH_DIRECTORY/prereq-test-dir\" &&\n  (\n          cd \"$TRASH_DIRECTORY/prereq-test-dir\" &&\n          : lazily evaluating a dummy prereq\n  \n  )\n  + mkdir -p /home/szeder/src/git/t/trash directory.t9999-test/prereq-test-dir\n  + cd /home/szeder/src/git/t/trash directory.t9999-test/prereq-test-dir\n  + : lazily evaluating a dummy prereq\n  prerequisite DUMMY_PREREQ ok\n  expecting success of 9999.1 'test': \n          true\n  \n  + true\n  ok 1 - test\n  \n  # passed all 1 test(s)\n  1..1\n  $ bash ./t9999-test.sh -x\n  Initialized empty Git repository in /home/szeder/src/git/t/trash directory.t9999-test/.git/\n  checking prerequisite: DUMMY_PREREQ\n  \n  mkdir -p \"$TRASH_DIRECTORY/prereq-test-dir\" &&\n  (\n          cd \"$TRASH_DIRECTORY/prereq-test-dir\" &&\n          : lazily evaluating a dummy prereq\n  \n  )\n  ++ mkdir -p '/home/szeder/src/git/t/trash directory.t9999-test/prereq-test-dir'\n  ++ cd '/home/szeder/src/git/t/trash directory.t9999-test/prereq-test-dir'\n  ++ : lazily evaluating a dummy prereq\n  prerequisite DUMMY_PREREQ ok\n  expecting success of 9999.1 'test': \n          true\n  \n  ++ true\n  ok 1 - test\n  \n  # passed all 1 test(s)\n  1..1\n\n"},{"id":"382490","messageId":"CAP8UFD38S_nV2NmjeadZ0J5ftJgBwghOZ+BNHZaNQ72nZmLtNA@mail.gmail.com","threadId":"51754","inReplyTo":"nycvar.QRO.7.76.6.1909171158090.15067@tvgsbejvaqbjf.bet","subject":"Re: Git in Outreachy December 2019?","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2019-09-17T15:10:43Z","receivedAt":"2019-09-17T15:10:58Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"Hi Emily and Dscho,\n\nOn Tue, Sep 17, 2019 at 1:28 PM Johannes Schindelin\n<Johannes.Schindelin@gmx.de> wrote:\n> On Mon, 16 Sep 2019, Emily Shaffer wrote:\n>\n> > Jonathan Tan, Jonathan Nieder, Josh Steadmon and I met on Friday to\n> > talk about projects and we came up with a trimmed list; not sure what\n> > more needs to be done to make them into fully-fledged proposals.\n>\n> Thank you for doing this!\n\nYeah, great!\n\n> > For starter microprojects, we came up with:\n> >\n> >  - cleanup a test script (although we need to identify particularly\n> >    which ones and what counts as \"clean\")\n> >  - moving doc from documentation/technical/api-* to comments in the\n> >    appropriate header instead\n> >  - teach a command which currently handles its own argv how to use\n> >    parse-options instead\n> >  - add a user.timezone option which Git can use if present rather than\n> >    checking system local time\n>\n> Nice projects, all. There are a couple more ideas on\n> https://github.com/gitgitgadget/git/issues, they could probably use some\n> tagging.\n\nThanks! Maybe we should have a page with Outreachy microprojects on\nhttps://git.github.io/\n\nI will see if I find the time to create one soon with the above information.\n\n> > For the longer projects, we came up with a few more:\n\n[...]\n\n> >    - git-bisect.sh\n>\n> That would be my top recommendation, especially given how much effort\n> Tanushree put in last winter to make this conversion to C so much more\n> achievable than before.\n\nI just added a project in the Outreachy system about it. I would have\nadded the link but Outreachy asks to not share the link publicly. I am\nwilling to co-mentor (or maybe mentor alone if no one else wants to\nco-mentor) it. Anyone willing to co-mentor can register on the\noutreachy website.\n\nThanks for the other suggestions by the way.\n\n> Converting shell/Perl scripts into built-in C never looks as much fun as\n> open-ended projects with lots of playing around, but the advantage of\n> the former is that they can be easily structured, offer a lot of\n> opportunity for learning, and they are ultimately more rewarding because\n> the goals are much better defined than many other projects'.\n\nI agree. Outreachy also suggest avoiding projects that have to be\ndiscussed a lot or are not clear enough.\n\n> >  - reduce/eliminate use of fetch_if_missing global\n\nI like this one!\n\n> > It might make sense to only focus on scoping the ones we feel most\n> > interested in. We came up with a pretty big list because we had some\n> > other programs in mind, so I suppose it's not necessary to develop all\n> > of them for this program.\n\nI agree as I don't think we will have enough mentors or co-mentors for\na big number of projects.\n\nBest,\nChristian.\n"},{"id":"382681","messageId":"20190920170448.226942-1-jonathantanmy@google.com","threadId":"51754","inReplyTo":"20190913205148.GA8799@sigill.intra.peff.net","subject":"Re: Git in Outreachy December 2019?","fromName":"Jonathan Tan","fromEmail":"jonathantanmy@google.com","sentAt":"2019-09-20T17:04:48Z","receivedAt":"2019-09-20T17:04:55Z","isPatch":false,"sender":{"key":"jonathantanmy@fastmail.com","avatar":null},"body":"> Prospective mentors need to sign up on that site, and should propose a\n> project they'd be willing to mentor.\n\n[snip]\n\n> I'm happy to discuss possible projects if anybody has an idea but isn't\n> sure how to develop it into a proposal.\n\nI'm new to Outreachy and programs like this, so does anyone have an\nopinion on my draft proposal below? It does not have any immediate\nuser-facing benefit, but it does have a definite end point.\n\nAlso let me know if an Outreachy proposal should have more detail, etc.\n\n    Refactor \"git index-pack\" logic into library code\n\n    Currently, whenever any Git code needs a pack to be indexed, it\n    needs to spawn a new \"git index-pack\" process, passing command-line\n    arguments and communicating with it using file descriptors (standard\n    input and output), much like an end-user would if invoking \"git\n    index-pack\" directly. Refactor the pack indexing logic into library\n    code callable from other Git code, make \"git index-pack\" a thin\n    wrapper around that library code, and (to demonstrate that the\n    refactoring works) change fetch-pack.c to use the library code\n    instead of spawning the \"git index-pack\" process.\n\n    This allows the pack indexing code to communicate with its callers\n    with the full power of C (structs, callbacks, etc.) instead of being\n    restricted to command-line arguments and file descriptors. It also\n    simplifies debugging in that there will no longer be 2\n    inter-communicating processes to deal with, only 1.\n"},{"id":"382717","messageId":"20190921014701.GA191795@google.com","threadId":"51754","inReplyTo":"20190920170448.226942-1-jonathantanmy@google.com","subject":"Re: Git in Outreachy December 2019?","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2019-09-21T01:47:01Z","receivedAt":"2019-09-21T01:48:37Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Fri, Sep 20, 2019 at 10:04:48AM -0700, Jonathan Tan wrote:\n> > Prospective mentors need to sign up on that site, and should propose a\n> > project they'd be willing to mentor.\n> \n> [snip]\n> \n> > I'm happy to discuss possible projects if anybody has an idea but isn't\n> > sure how to develop it into a proposal.\n> \n> I'm new to Outreachy and programs like this, so does anyone have an\n> opinion on my draft proposal below? It does not have any immediate\n> user-facing benefit, but it does have a definite end point.\n\nI'd appreciate similar opinion if anybody has it - and I'd also really\nfeel more comfortable with a co-mentor.\n\n\"\"\"\n\"Did You Mean..?\"\n\nThere are some situations where it's fairly clear what a user meant to\ndo, even though they did not do that thing correctly. For example, if a\nuser runs `git commit` with tracked, modified, unstaged files in their\nworktree, but no staged files at all, it's fairly likely that they\nsimply forgot to add the files they wanted. In this case, the error\nmessage is slightly obtuse:\n\n$ git commit\nOn branch master\nChanges not staged for commit:\n\tmodified:   foo.txt\n\nno changes added to commit\n\n\nSince we have an idea of what the user _meant_ to do, we can offer\nsomething more like:\n\n$ git commit\nOn branch master\nChanges not staged for commit:\n\tmodified:   foo.txt\n\nStage listed changes and continue? [Y/n]\n\nWhile the above case is a good starting place, other similar cases can\nbe added afterwards if time permits. These helper prompts should be\nenabled/disabled via a config option so that people who are used to\ntheir current workflow won't be impacted.\n\"\"\"\n\nThanks in advance for feedback.\n\n - Emily\n"},{"id":"382766","messageId":"CAP8UFD3NPYJr5PXLDyRD=qbEPft8E-HwtGUo_FxoG=q5jfY5Ng@mail.gmail.com","threadId":"51754","inReplyTo":"20190920170448.226942-1-jonathantanmy@google.com","subject":"Re: Git in Outreachy December 2019?","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2019-09-23T11:49:05Z","receivedAt":"2019-09-23T11:49:20Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"On Mon, Sep 23, 2019 at 10:53 AM Jonathan Tan <jonathantanmy@google.com> wrote:\n>\n> > Prospective mentors need to sign up on that site, and should propose a\n> > project they'd be willing to mentor.\n\nYeah, you are very welcome to sign up soon if you haven't already done\nso as I think the deadline is really soon.\n\n> > I'm happy to discuss possible projects if anybody has an idea but isn't\n> > sure how to develop it into a proposal.\n>\n> I'm new to Outreachy and programs like this, so does anyone have an\n> opinion on my draft proposal below? It does not have any immediate\n> user-facing benefit, but it does have a definite end point.\n\nNo need for user-facing benefits. Refactoring or improving the code in\nother useful ways are very good subjects (as I already said in my\nreply to Emily and Dscho).\n\n> Also let me know if an Outreachy proposal should have more detail, etc.\n>\n>     Refactor \"git index-pack\" logic into library code\n>\n>     Currently, whenever any Git code needs a pack to be indexed, it\n>     needs to spawn a new \"git index-pack\" process, passing command-line\n>     arguments and communicating with it using file descriptors (standard\n>     input and output), much like an end-user would if invoking \"git\n>     index-pack\" directly. Refactor the pack indexing logic into library\n>     code callable from other Git code, make \"git index-pack\" a thin\n>     wrapper around that library code, and (to demonstrate that the\n>     refactoring works) change fetch-pack.c to use the library code\n>     instead of spawning the \"git index-pack\" process.\n>\n>     This allows the pack indexing code to communicate with its callers\n>     with the full power of C (structs, callbacks, etc.) instead of being\n>     restricted to command-line arguments and file descriptors. It also\n>     simplifies debugging in that there will no longer be 2\n>     inter-communicating processes to deal with, only 1.\n\nI think this is really great, both the idea and the description! No\nneed for more details.\n\nThanks,\nChristian.\n"},{"id":"382768","messageId":"nycvar.QRO.7.76.6.1909231444590.15067@tvgsbejvaqbjf.bet","threadId":"51754","inReplyTo":"20190917120230.GA27531@szeder.dev","subject":"Re: Git in Outreachy December 2019?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-09-23T12:47:23Z","receivedAt":"2019-09-23T12:47:50Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 17 Sep 2019, SZEDER Gábor wrote:\n\n> On Tue, Sep 17, 2019 at 01:23:18PM +0200, Johannes Schindelin wrote:\n> > Also, things like the code tracing via `-x` (which relies on Bash\n> > functionality in order to work properly,\n>\n> Not really.\n\nTo work properly. What I meant was the trick we need to play with\n`BASH_XTRACEFD`.\n\n> > and which _still_ does not work as intended if your test case\n> > evaluates a lazy prereq that has not been evaluated before\n>\n> I don't see any striking differences between the trace output of a test\n> involving a lazy prereq from Bash or dash:\n>\n> [...]\n\nThe evaluation of the lazy prereq is indeed not different between Bash\nor dash. It is nevertheless quite disruptive in the trace of a test\nscript, especially when it is evaluated for a test case that is skipped\nexplicitly via the `--run` option.\n\nCiao,\nDscho\n"},{"id":"382769","messageId":"nycvar.QRO.7.76.6.1909231448340.15067@tvgsbejvaqbjf.bet","threadId":"51754","inReplyTo":"CAP8UFD38S_nV2NmjeadZ0J5ftJgBwghOZ+BNHZaNQ72nZmLtNA@mail.gmail.com","subject":"Re: Git in Outreachy December 2019?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-09-23T12:50:20Z","receivedAt":"2019-09-23T12:50:41Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 17 Sep 2019, Christian Couder wrote:\n\n> On Tue, Sep 17, 2019 at 1:28 PM Johannes Schindelin\n> <Johannes.Schindelin@gmx.de> wrote:\n> > On Mon, 16 Sep 2019, Emily Shaffer wrote:\n> >\n> > >  - reduce/eliminate use of fetch_if_missing global\n>\n> I like this one!\n\nIt looks as if a (non-Outreachy) contributor also does like this one\nalready: https://github.com/git/git/pull/650\n\nCiao,\nDscho\n"},{"id":"382771","messageId":"CAP8UFD3zw1dYUZ8Sei+kzcYmcsgQsRLoy1uHU+ZQp6CBDbCVkQ@mail.gmail.com","threadId":"51754","inReplyTo":"20190921014701.GA191795@google.com","subject":"Re: Git in Outreachy December 2019?","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2019-09-23T14:23:25Z","receivedAt":"2019-09-23T14:23:40Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"On Mon, Sep 23, 2019 at 3:35 PM Emily Shaffer <emilyshaffer@google.com> wrote:\n>\n> On Fri, Sep 20, 2019 at 10:04:48AM -0700, Jonathan Tan wrote:\n>\n> > I'm new to Outreachy and programs like this, so does anyone have an\n> > opinion on my draft proposal below? It does not have any immediate\n> > user-facing benefit, but it does have a definite end point.\n>\n> I'd appreciate similar opinion if anybody has it - and I'd also really\n> feel more comfortable with a co-mentor.\n\nFirst as the deadline is tomorrow, I think it is important to submit\nprojects to Outreachy even if they are not perfect and even if we\nwould like a co-mentor (which is also my case by the way).\n\nWe can hopefully improve the projects after the deadline or perhaps\ndrop them if we cannot find enough co-mentors or if we don't agree\nwith the goal or find it too difficult.\n\n> \"\"\"\n> \"Did You Mean..?\"\n>\n> There are some situations where it's fairly clear what a user meant to\n> do, even though they did not do that thing correctly. For example, if a\n> user runs `git commit` with tracked, modified, unstaged files in their\n> worktree, but no staged files at all, it's fairly likely that they\n> simply forgot to add the files they wanted. In this case, the error\n> message is slightly obtuse:\n>\n> $ git commit\n> On branch master\n> Changes not staged for commit:\n>         modified:   foo.txt\n>\n> no changes added to commit\n>\n>\n> Since we have an idea of what the user _meant_ to do, we can offer\n> something more like:\n>\n> $ git commit\n> On branch master\n> Changes not staged for commit:\n>         modified:   foo.txt\n>\n> Stage listed changes and continue? [Y/n]\n>\n> While the above case is a good starting place, other similar cases can\n> be added afterwards if time permits. These helper prompts should be\n> enabled/disabled via a config option so that people who are used to\n> their current workflow won't be impacted.\n> \"\"\"\n\nI agree that it might help. There could be significant discussion\nabout what the UI should be though. For example maybe we could just\npromote `git commit -p` in the tutorials instead of doing the above.\nOr have a commit.patch config option if we haven't one already.\n\nThanks,\nChristian.\n"},{"id":"382774","messageId":"20190923165828.GA27068@szeder.dev","threadId":"51754","inReplyTo":"nycvar.QRO.7.76.6.1909231444590.15067@tvgsbejvaqbjf.bet","subject":"Re: Git in Outreachy December 2019?","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2019-09-23T16:58:28Z","receivedAt":"2019-09-23T16:58:38Z","isPatch":false,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Mon, Sep 23, 2019 at 02:47:23PM +0200, Johannes Schindelin wrote:\n> Hi,\n> \n> On Tue, 17 Sep 2019, SZEDER Gábor wrote:\n> \n> > On Tue, Sep 17, 2019 at 01:23:18PM +0200, Johannes Schindelin wrote:\n> > > Also, things like the code tracing via `-x` (which relies on Bash\n> > > functionality in order to work properly,\n> >\n> > Not really.\n> \n> To work properly. What I meant was the trick we need to play with\n> `BASH_XTRACEFD`.\n\nI'm still unsure what BASH_XTRACEFD trick you mean.  AFAICT we don't\nplay any tricks with it to make '-x' work properly, and indeed '-x'\ntracing works properly even without BASH_XTRACEFD (and to achive that\nwe did have to play some tricks, but not any with BASH_XTRACEFD;\nperhaps these tricks are what you meant?).\n\n> > > and which _still_ does not work as intended if your test case\n> > > evaluates a lazy prereq that has not been evaluated before\n> >\n> > I don't see any striking differences between the trace output of a test\n> > involving a lazy prereq from Bash or dash:\n> >\n> > [...]\n> \n> The evaluation of the lazy prereq is indeed not different between Bash\n> or dash. It is nevertheless quite disruptive in the trace of a test\n> script, especially when it is evaluated for a test case that is skipped\n> explicitly via the `--run` option.\n\nBut then the actual issue is the unnecessary evaluation of the prereq\neven when the test framework could know in advance that the test case\nshould be skipped anyway, and the trace from it is a mere side effect,\nno?\n\n"},{"id":"382776","messageId":"20190923175805.58457-1-jonathantanmy@google.com","threadId":"51754","inReplyTo":"CAP8UFD3NPYJr5PXLDyRD=qbEPft8E-HwtGUo_FxoG=q5jfY5Ng@mail.gmail.com","subject":"Re: Git in Outreachy December 2019?","fromName":"Jonathan Tan","fromEmail":"jonathantanmy@google.com","sentAt":"2019-09-23T17:58:05Z","receivedAt":"2019-09-23T17:58:11Z","isPatch":false,"sender":{"key":"jonathantanmy@fastmail.com","avatar":null},"body":"> No need for user-facing benefits. Refactoring or improving the code in\n> other useful ways are very good subjects (as I already said in my\n> reply to Emily and Dscho).\n\nThanks!\n\n> I think this is really great, both the idea and the description! No\n> need for more details.\n\nThanks! I've just submitted the project proposal - hopefully it will be\napproved soon. In any case, it seems that project information can be\nedited after submission.\n\nThere was a \"How can applicants make a contribution to your project?\"\nquestion and a few questions about communication channels. I answered\nthem as best I could but if anyone has already answered them, it would\nbe great to just use the same answer everywhere. (I can't see all\nproject information of other projects since I haven't filled out a\n\"short initial application\", but I don't think that applies to me.)\n"},{"id":"382777","messageId":"20190923180649.GA2886@szeder.dev","threadId":"51754","inReplyTo":"20190904194114.GA31398@sigill.intra.peff.net","subject":"Re: Git in Outreachy December 2019?","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2019-09-23T18:07:09Z","receivedAt":"2019-09-23T18:07:19Z","isPatch":false,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Wed, Sep 04, 2019 at 03:41:15PM -0400, Jeff King wrote:\n> The project page has a section to point people in the right direction\n> for first-time contributions. I've left it blank for now, but I think it\n> makes sense to point one (or both) of:\n> \n>   - https://git-scm.com/docs/MyFirstContribution\n> \n>   - https://matheustavares.gitlab.io/posts/first-steps-contributing-to-git\n> \n> as well as a list of micro-projects (or at least instructions on how to\n> find #leftoverbits, though we'd definitely have to step up our labeling,\n> as I do not recall having seen one for a while).\n\nAnd we should make sure that all microprojects are indeed micro in\nsize.  Matheus sent v8 of a 10 patch series in July that started out\nas a microproject back in February...\n\nHere is one more idea for microprojects:\n\n  Find a group of related preprocessor constants and turn them into an\n  enum.  Also find where those constants are stored in variables and\n  in structs and passed around as function parameters, and change the\n  type of those variables, fields and parameters to the new enum.\n\n\n\n"},{"id":"382778","messageId":"20190923180751.GA21344@sigill.intra.peff.net","threadId":"51754","inReplyTo":"nycvar.QRO.7.76.6.1909171158090.15067@tvgsbejvaqbjf.bet","subject":"Re: Git in Outreachy December 2019?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-09-23T18:07:51Z","receivedAt":"2019-09-23T18:07:53Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Sep 17, 2019 at 01:23:18PM +0200, Johannes Schindelin wrote:\n\n> The only slightly challenging aspect might be that `merge-one-file` is\n> actually not a merge strategy, but it is used as helper to be passed to\n> `git merge-index` via the `-o <helper>` option, which makes it slightly\n> awkward to be implemented as a built-in. A better approach would\n> therefore be to special-case this value in `merge-index` and execute the\n> C code directly, without the detour of spawning a built-in.\n\nI think it could make sense for merge-index to be able to directly run\nthe merge-one-file code[1]. But I think we'd want to keep its ability to\nrun an arbitrary script, and for people to call merge-one-file\nseparately, since right now you can do:\n\n  git merge-index my-script\n\nand have \"my-script\" do some processing of its own, then hand off more\nwork to merge-one-file.\n\nSo the weird calling convention is actually a user-visible and\npotentially useful interface. So it would need a deprecation period\nbefore being removed.\n\n-Peff\n\n[1] Certainly doing it in-process would be faster for the common case of\n    \"git merge-index git-merge-one-file\", but I wonder if anybody really\n    does that. These days most people would just merge-recursive anyway.\n"},{"id":"382779","messageId":"20190923181950.GB21344@sigill.intra.peff.net","threadId":"51754","inReplyTo":"nycvar.QRO.7.76.6.1909231444590.15067@tvgsbejvaqbjf.bet","subject":"Re: Git in Outreachy December 2019?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-09-23T18:19:51Z","receivedAt":"2019-09-23T18:19:53Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Sep 23, 2019 at 02:47:23PM +0200, Johannes Schindelin wrote:\n\n> The evaluation of the lazy prereq is indeed not different between Bash\n> or dash. It is nevertheless quite disruptive in the trace of a test\n> script, especially when it is evaluated for a test case that is skipped\n> explicitly via the `--run` option.\n\nThat sounds like a bug: if we know we are not going to run the test\nanyway due to --run or GIT_TEST_SKIP, we should probably avoid checking\nthe prereq at all.\n\n-Peff\n"},{"id":"382787","messageId":"20190923191509.GC21344@sigill.intra.peff.net","threadId":"51754","inReplyTo":"20190920170448.226942-1-jonathantanmy@google.com","subject":"Re: Git in Outreachy December 2019?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-09-23T19:15:10Z","receivedAt":"2019-09-23T19:15:12Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Sep 20, 2019 at 10:04:48AM -0700, Jonathan Tan wrote:\n\n> > I'm happy to discuss possible projects if anybody has an idea but isn't\n> > sure how to develop it into a proposal.\n> \n> I'm new to Outreachy and programs like this, so does anyone have an\n> opinion on my draft proposal below? It does not have any immediate\n> user-facing benefit, but it does have a definite end point.\n> \n> Also let me know if an Outreachy proposal should have more detail, etc.\n> \n>     Refactor \"git index-pack\" logic into library code\n> \n>     Currently, whenever any Git code needs a pack to be indexed, it\n>     needs to spawn a new \"git index-pack\" process, passing command-line\n>     arguments and communicating with it using file descriptors (standard\n>     input and output), much like an end-user would if invoking \"git\n>     index-pack\" directly. Refactor the pack indexing logic into library\n>     code callable from other Git code, make \"git index-pack\" a thin\n>     wrapper around that library code, and (to demonstrate that the\n>     refactoring works) change fetch-pack.c to use the library code\n>     instead of spawning the \"git index-pack\" process.\n> \n>     This allows the pack indexing code to communicate with its callers\n>     with the full power of C (structs, callbacks, etc.) instead of being\n>     restricted to command-line arguments and file descriptors. It also\n>     simplifies debugging in that there will no longer be 2\n>     inter-communicating processes to deal with, only 1.\n\nI think this is an OK level of detail. I'm not sure quite sure about the\ngoal of the project, though. In particular:\n\n  - I'm not clear what we'd hope to gain. I.e., what richer information\n    would we want to pass back and forth between index-pack and the\n    other processes? It might also be more efficient, but I'm not sure\n    it's measurably so (we save a single process, and we save some pipe\n    traffic, but the sideband demuxer would probably end up passing it\n    over a self-pipe anyway).\n\n  - index-pack is prone to dying on bad input, and we wouldn't want it\n    to take down the outer fetch-pack or receive-pack, which are what\n    produce useful messages to the user. That's something that could be\n    fixed as part of the libification, but I suspect the control flow\n    might be a little tricky.\n\n  - we don't always call index-pack, but sometimes call unpack-objects.\n    I suppose we could continue to call an external unpack-objects in\n    that path, but that eliminates the utility of having richer\n    communication if we sometimes have to take the \"dumb\" path. A while\n    ago I took a stab at teaching index-pack to unpack. It works, but\n    there are a few ugly bits, as discussed in:\n\n      https://github.com/peff/git/commit/7df82454a855281e9c147f3023225f8a6f72e303\n\n    Maybe that would be worth making part of the project?\n\n-Peff\n"},{"id":"382789","messageId":"20190923192704.GD21344@sigill.intra.peff.net","threadId":"51754","inReplyTo":"20190923175805.58457-1-jonathantanmy@google.com","subject":"Re: Git in Outreachy December 2019?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-09-23T19:27:04Z","receivedAt":"2019-09-23T19:27:06Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Sep 23, 2019 at 10:58:05AM -0700, Jonathan Tan wrote:\n\n> > I think this is really great, both the idea and the description! No\n> > need for more details.\n> \n> Thanks! I've just submitted the project proposal - hopefully it will be\n> approved soon. In any case, it seems that project information can be\n> edited after submission.\n\nI approved this. I did leave some comments elsewhere in the thread, but\nI think we can continue to iterate on the idea.\n\n> There was a \"How can applicants make a contribution to your project?\"\n> question and a few questions about communication channels. I answered\n> them as best I could but if anyone has already answered them, it would\n> be great to just use the same answer everywhere. (I can't see all\n> project information of other projects since I haven't filled out a\n> \"short initial application\", but I don't think that applies to me.)\n\nWhat you wrote there looks pretty good to me. Copying it here for\npurposes of discussion:\n\n> Please introduce yourself on the public project chat:\n>\n> IRC - Follow this link to join this project's public chat. If you are\n> asked for username, pick any username! You will not need a password to\n> join this chat.\n>\n> Git mailing list - Follow this link to join this project's public\n> chat.\n\nwhere the IRC link goes to Freenode's webchat, and the mailing list link\ngoes to https://git-scm.com/community.\n\nThe other proposal is from Christian, who wrote:\n\n> Please introduce yourself on the public project chat:\n>\n>     a mailing list - Once you join the project's communication\n>     channel, the mentors have some additional instructions for you to\n>     follow:\n> \n>     Start the subject of your emails to the mailing list with \"[Outreachy]\" so that mentors and people interested in Outreachy can easily notice your emails.\n>     Follow this link to join this project's public chat.\n> \n>     Send an email to majordomo@vger.kernel.org with \"subscribe git\" in\n>     the body of the email.\n\nI think the \"[Outreachy]\" advice is good, and worth adding to yours. I\nthink linking to the \"community\" page is a good idea for Christian's, as\nit has more tips on using the mailing list.\n\nI'd leave it up to individual mentors whether they want to mention IRC\n(based on whether or not they actually use IRC themselves).\n\n-Peff\n"},{"id":"382790","messageId":"20190923193009.GE21344@sigill.intra.peff.net","threadId":"51754","inReplyTo":"CAP8UFD38S_nV2NmjeadZ0J5ftJgBwghOZ+BNHZaNQ72nZmLtNA@mail.gmail.com","subject":"Re: Git in Outreachy December 2019?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-09-23T19:30:09Z","receivedAt":"2019-09-23T19:30:12Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Sep 17, 2019 at 05:10:43PM +0200, Christian Couder wrote:\n\n> > Nice projects, all. There are a couple more ideas on\n> > https://github.com/gitgitgadget/git/issues, they could probably use some\n> > tagging.\n> \n> Thanks! Maybe we should have a page with Outreachy microprojects on\n> https://git.github.io/\n> \n> I will see if I find the time to create one soon with the above information.\n\nPlease let me know if you do; there's a spot in the project page for:\n\n  Description of your first time contribution tutorial:\n\n  If your applicants need to complete a tutorial before working on\n  contributions with mentors, please provide a description and the URL\n  for the tutorial. For example, the Linux kernel asks applicants to\n  complete a tutorial for compiling and installing a custom kernel, and\n  sending in a simple whitespace change patch. Once applicants complete\n  this tutorial, they can start to work with mentors on more complex\n  contributions.\n\nwhich could link to that list (and the list should probably discuss how\nwe consider micro-projects, etc, and link to MyFirstContribution or\nsimilar).\n\nI think this suggestion is a good one to add, as well:\n\n  https://public-inbox.org/git/20190923180649.GA2886@szeder.dev/\n\n-Peff\n"},{"id":"382792","messageId":"20190923194004.GF21344@sigill.intra.peff.net","threadId":"51754","inReplyTo":"20190921014701.GA191795@google.com","subject":"Re: Git in Outreachy December 2019?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-09-23T19:40:05Z","receivedAt":"2019-09-23T19:40:07Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Sep 20, 2019 at 06:47:01PM -0700, Emily Shaffer wrote:\n\n> \"\"\"\n> \"Did You Mean..?\"\n> \n> There are some situations where it's fairly clear what a user meant to\n> do, even though they did not do that thing correctly. For example, if a\n> user runs `git commit` with tracked, modified, unstaged files in their\n> worktree, but no staged files at all, it's fairly likely that they\n> simply forgot to add the files they wanted. In this case, the error\n> message is slightly obtuse:\n> \n> $ git commit\n> On branch master\n> Changes not staged for commit:\n> \tmodified:   foo.txt\n> \n> no changes added to commit\n> \n> \n> Since we have an idea of what the user _meant_ to do, we can offer\n> something more like:\n> \n> $ git commit\n> On branch master\n> Changes not staged for commit:\n> \tmodified:   foo.txt\n> \n> Stage listed changes and continue? [Y/n]\n> \n> While the above case is a good starting place, other similar cases can\n> be added afterwards if time permits. These helper prompts should be\n> enabled/disabled via a config option so that people who are used to\n> their current workflow won't be impacted.\n> \"\"\"\n\nThis is an interesting idea. At first I thought it might be too small\nfor a project, but I think it could be expanded or contracted as much as\nthe time allows by just looking for more \"did you mean\" spots.\n\nI have mixed feelings on making things interactive. For one, it gets\nawkward when Git commands are called as part of a script or other\nprogram (and a lot of programs like git-commit ride the line of plumbing\nand porcelain). I know this would kick in only when a config option is\nset, but I think that might things even _more_ confusing, as something\nthat works for one user (without the config) would start behaving\nweirdly for another.\n\nI also think it might be an opportunity to educate. Instead of giving a\nyes/no prompt, we can actually recommend one (or more!) sets of commands\nto get the desired effect. I _thought_ we already did for this case by\ndefault (triggered by advice.statusHints, which is true by default). But\nit looks like those don't get printed for git-commit?\n\n-Peff\n"},{"id":"382794","messageId":"20190923203854.171170-1-jonathantanmy@google.com","threadId":"51754","inReplyTo":"20190923191509.GC21344@sigill.intra.peff.net","subject":"Re: Git in Outreachy December 2019?","fromName":"Jonathan Tan","fromEmail":"jonathantanmy@google.com","sentAt":"2019-09-23T20:38:54Z","receivedAt":"2019-09-23T20:39:00Z","isPatch":false,"sender":{"key":"jonathantanmy@fastmail.com","avatar":null},"body":"> I think this is an OK level of detail. I'm not sure quite sure about the\n> goal of the project, though. In particular:\n> \n>   - I'm not clear what we'd hope to gain. I.e., what richer information\n>     would we want to pass back and forth between index-pack and the\n>     other processes? It might also be more efficient, but I'm not sure\n>     it's measurably so (we save a single process, and we save some pipe\n>     traffic, but the sideband demuxer would probably end up passing it\n>     over a self-pipe anyway).\n\nI didn't have any concrete ideas so I didn't include those, but some\nunrefined ideas:\n\n - index-pack has the CLI option to specify a message to be written into\n   the .promisor file, but in my patch to write fetched refs to\n   .promisor [1], I ended up making fetch-pack.c write the information\n   because I didn't know how many refs were going to be written (and I\n   didn't want to bump into CLI argument length limits). If we had this\n   feature, I might have been able to pass a callback to index-pack that\n   writes the list of refs once we have the fd into .promisor,\n   eliminating some code duplication (but I haven't verified this).\n\n - In your reply [2] to the above [1], you mentioned the possibility of\n   keeping a list of cutoff points. One way of doing this, as I state in\n   [3], is my original suggestion back in 2017 of one such\n   repository-wide list. If we do this, it would be better for\n   fetch-pack to handle this instead of index-pack, and it seems more\n   efficient to me to have index-pack be able to pass objects to\n   fetch-pack as they are inflated instead of fetch-pack rereading the\n   compressed forms on disk (but again, I haven't verified this).\n\n[1] https://public-inbox.org/git/20190826214737.164132-1-jonathantanmy@google.com/\n[2] https://public-inbox.org/git/20190905070153.GE21450@sigill.intra.peff.net/\n[3] https://public-inbox.org/git/20190905183926.137490-1-jonathantanmy@google.com/\n\nThere are also the debuggability improvements of not having to deal with\n2 processes.\n\n>   - index-pack is prone to dying on bad input, and we wouldn't want it\n>     to take down the outer fetch-pack or receive-pack, which are what\n>     produce useful messages to the user. That's something that could be\n>     fixed as part of the libification, but I suspect the control flow\n>     might be a little tricky.\n\nGood point.\n\n>   - we don't always call index-pack, but sometimes call unpack-objects.\n>     I suppose we could continue to call an external unpack-objects in\n>     that path, but that eliminates the utility of having richer\n>     communication if we sometimes have to take the \"dumb\" path. A while\n>     ago I took a stab at teaching index-pack to unpack. It works, but\n>     there are a few ugly bits, as discussed in:\n> \n>       https://github.com/peff/git/commit/7df82454a855281e9c147f3023225f8a6f72e303\n> \n>     Maybe that would be worth making part of the project?\n\nI'm reluctant to do so because I don't want to increase the scope too\nmuch - although if my project has relatively narrow scope for an\nOutreachy project, we can do so. As for eliminating the utility of\nhaving richer communication, I don't think so, because in the situations\nwhere we require richer communication (right now, situations to do with\npartial clone), we specifically run index-pack anyway.\n"},{"id":"382796","messageId":"20190923204807.173287-1-jonathantanmy@google.com","threadId":"51754","inReplyTo":"20190923192704.GD21344@sigill.intra.peff.net","subject":"Re: Git in Outreachy December 2019?","fromName":"Jonathan Tan","fromEmail":"jonathantanmy@google.com","sentAt":"2019-09-23T20:48:07Z","receivedAt":"2019-09-23T20:48:13Z","isPatch":false,"sender":{"key":"jonathantanmy@fastmail.com","avatar":null},"body":"> I approved this. I did leave some comments elsewhere in the thread, but\n> I think we can continue to iterate on the idea.\n\nThanks.\n\n> > There was a \"How can applicants make a contribution to your project?\"\n> > question and a few questions about communication channels. I answered\n> > them as best I could but if anyone has already answered them, it would\n> > be great to just use the same answer everywhere. (I can't see all\n> > project information of other projects since I haven't filled out a\n> > \"short initial application\", but I don't think that applies to me.)\n> \n> What you wrote there looks pretty good to me. Copying it here for\n> purposes of discussion:\n> \n> > Please introduce yourself on the public project chat:\n> >\n> > IRC - Follow this link to join this project's public chat. If you are\n> > asked for username, pick any username! You will not need a password to\n> > join this chat.\n> >\n> > Git mailing list - Follow this link to join this project's public\n> > chat.\n> \n> where the IRC link goes to Freenode's webchat, and the mailing list link\n> goes to https://git-scm.com/community.\n\nThanks. For the record, most of the text is from Outreachy - I just\nsupplied the links.\n\n> The other proposal is from Christian, who wrote:\n> \n> > Please introduce yourself on the public project chat:\n> >\n> >     a mailing list - Once you join the project's communication\n> >     channel, the mentors have some additional instructions for you to\n> >     follow:\n> > \n> >     Start the subject of your emails to the mailing list with \"[Outreachy]\" so that mentors and people interested in Outreachy can easily notice your emails.\n> >     Follow this link to join this project's public chat.\n> > \n> >     Send an email to majordomo@vger.kernel.org with \"subscribe git\" in\n> >     the body of the email.\n> \n> I think the \"[Outreachy]\" advice is good, and worth adding to yours. I\n> think linking to the \"community\" page is a good idea for Christian's, as\n> it has more tips on using the mailing list.\n\nGood idea on the \"[Outreachy]\" - I've added it to my project as well.\n"},{"id":"382798","messageId":"20190923212834.GA19504@sigill.intra.peff.net","threadId":"51754","inReplyTo":"20190923203854.171170-1-jonathantanmy@google.com","subject":"Re: Git in Outreachy December 2019?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-09-23T21:28:34Z","receivedAt":"2019-09-23T21:28:37Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Sep 23, 2019 at 01:38:54PM -0700, Jonathan Tan wrote:\n\n> I didn't have any concrete ideas so I didn't include those, but some\n> unrefined ideas:\n\nOne risk to a mentoring project like this is that the intern does a good\njob of steps 1-5, and then in step 6 we realize that the whole thing is\nnot useful, and upstream doesn't want it. Which isn't to say the intern\ndidn't learn something, and the project didn't benefit. Negative results\ncan be useful; but it can also be demoralizing.\n\nI'm not arguing that's going to be the case here. But I do think it's\nworth talking through these things a bit as part of thinking about\nproposals.\n\n>  - index-pack has the CLI option to specify a message to be written into\n>    the .promisor file, but in my patch to write fetched refs to\n>    .promisor [1], I ended up making fetch-pack.c write the information\n>    because I didn't know how many refs were going to be written (and I\n>    didn't want to bump into CLI argument length limits). If we had this\n>    feature, I might have been able to pass a callback to index-pack that\n>    writes the list of refs once we have the fd into .promisor,\n>    eliminating some code duplication (but I haven't verified this).\n\nThat makes some sense. We could pass the data over a pipe, but obviously\nstdin is already in use to receive the pack here. Ideally we'd be able\nto pass multiple streams between the programs, but I think due to\nWindows support, we can't assume that arbitrary pipe descriptors will\nmake it across the run-command boundary. So I think we'd be left with\ncommunicating via temporary files (which really isn't the worst thing in\nthe world, but has its own complications).\n\n>  - In your reply [2] to the above [1], you mentioned the possibility of\n>    keeping a list of cutoff points. One way of doing this, as I state in\n>    [3], is my original suggestion back in 2017 of one such\n>    repository-wide list. If we do this, it would be better for\n>    fetch-pack to handle this instead of index-pack, and it seems more\n>    efficient to me to have index-pack be able to pass objects to\n>    fetch-pack as they are inflated instead of fetch-pack rereading the\n>    compressed forms on disk (but again, I haven't verified this).\n\nAnd this is the flip-side problem: we need to get data back, but we have\nonly stdout, which is already in use (so we need some kind of protocol).\nThat leads to things like the horrible NUL-byte added by 83558686ce\n(receive-pack: send keepalives during quiet periods, 2016-07-15).\n\n> There are also the debuggability improvements of not having to deal with\n> 2 processes.\n\nI think it can sometimes be easier to debug with two separate processes,\nbecause the input to index-pack is well-defined and can be repeated\nwithout hitting the network (though you do have to figure out how to\nrecord the network response, which can be non-trivial). I've also done\nsimilar things for running performance simulations.\n\nWe'll still have the stand-alone index-pack command, so it can be used\nfor those cases. But as we add more features that utilize the in-process\ninterface, that may eventually stop being feasible.\n\n> > [dropping unpack-objects]\n> >     Maybe that would be worth making part of the project?\n> \n> I'm reluctant to do so because I don't want to increase the scope too\n> much - although if my project has relatively narrow scope for an\n> Outreachy project, we can do so. As for eliminating the utility of\n> having richer communication, I don't think so, because in the situations\n> where we require richer communication (right now, situations to do with\n> partial clone), we specifically run index-pack anyway.\n\nYeah, we're in kind of a weird situation there, where unpack-objects is\nused less and less. I wonder how many surprises are lurking where\nsomebody reasoned about index-pack behavior, but unpack-objects may do\nsomething slightly differently (I know this came up when we looked at\nfsck-ing incoming objects for submodule vulnerabilities).\n\nI kind of wonder if it would be reasonable to just always use index-pack\nfor the sake of simplicity, even if it never learns to actually unpack\nobjects. We've been doing that for years on the server side at GitHub\nwithout ill effects (I think the unpack route is slightly more efficient\nfor a thin pack, but since it only kicks in when there are few objects\nanyway, I wonder how big an advantage it is in general).\n\n-Peff\n"},{"id":"382805","messageId":"dba9791a-ea7a-2e79-0eb3-27fd8504853e@iee.email","threadId":"51754","inReplyTo":"20190923194004.GF21344@sigill.intra.peff.net","subject":"Re: Git in Outreachy December 2019?","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2019-09-23T22:29:36Z","receivedAt":"2019-09-23T22:29:44Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"On 23/09/2019 20:40, Jeff King wrote:\n> On Fri, Sep 20, 2019 at 06:47:01PM -0700, Emily Shaffer wrote:\n>\n>> \"\"\"\n>> \"Did You Mean..?\"\n>>\n>> There are some situations where it's fairly clear what a user meant to\n>> do, even though they did not do that thing correctly. For example, if a\n>> user runs `git commit` with tracked, modified, unstaged files in their\n>> worktree, but no staged files at all, it's fairly likely that they\n>> simply forgot to add the files they wanted. In this case, the error\n>> message is slightly obtuse:\n>>\n>> $ git commit\n>> On branch master\n>> Changes not staged for commit:\n>> \tmodified:   foo.txt\n>>\n>> no changes added to commit\n>>\n>>\n>> Since we have an idea of what the user _meant_ to do, we can offer\n>> something more like:\n>>\n>> $ git commit\n>> On branch master\n>> Changes not staged for commit:\n>> \tmodified:   foo.txt\n>>\n>> Stage listed changes and continue? [Y/n]\n>>\n>> While the above case is a good starting place, other similar cases can\n>> be added afterwards if time permits. These helper prompts should be\n>> enabled/disabled via a config option so that people who are used to\n>> their current workflow won't be impacted.\n>> \"\"\"\n> This is an interesting idea. At first I thought it might be too small\n> for a project, but I think it could be expanded or contracted as much as\n> the time allows by just looking for more \"did you mean\" spots.\n>\n> I have mixed feelings on making things interactive. For one, it gets\n> awkward when Git commands are called as part of a script or other\n> program (and a lot of programs like git-commit ride the line of plumbing\n> and porcelain). I know this would kick in only when a config option is\n> set, but I think that might things even _more_ confusing, as something\n> that works for one user (without the config) would start behaving\n> weirdly for another.\n>\n> I also think it might be an opportunity to educate. Instead of giving a\n> yes/no prompt, we can actually recommend one (or more!) sets of commands\n> to get the desired effect. I _thought_ we already did for this case by\n> default (triggered by advice.statusHints, which is true by default). But\n> it looks like those don't get printed for git-commit?\n>\n> -Peff\nAlso there is a lot of common problems and issues that can be mined from \nStackOverflow for similar \"Did You Mean..?\"user problems.\n\nPhilip\n\n  \n\n"},{"id":"382807","messageId":"20190924005529.GA8354@dcvr","threadId":"51754","inReplyTo":"nycvar.QRO.7.76.6.1909171158090.15067@tvgsbejvaqbjf.bet","subject":"Re: Git in Outreachy December 2019?","fromName":"Eric Wong","fromEmail":"e@80x24.org","sentAt":"2019-09-24T00:55:29Z","receivedAt":"2019-09-24T00:55:31Z","isPatch":false,"sender":{"key":"e@80x24.org","avatar":null},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> On Mon, 16 Sep 2019, Emily Shaffer wrote:\n> >  - try and make progress towards running many tests from a single test\n> >    file in parallel - maybe this is too big, I'm not sure if we know how\n> >    many of our tests are order-dependent within a file for now...\n> \n> Another, potentially more rewarding, project would be to modernize our\n> test suite framework, so that it is not based on Unix shell scripting,\n> but on C instead.\n\nI worry more C would reduce the amount of contributors (some of\nthe C rewrites already scared me off hacking years ago).  I\nfigure more users are familiar with sh than C.\n\nIt would also increase the disparity between tests and use of\nactual users from the command-line.\n\n> The fact that it is based on Unix shell scripting not only costs a lot\n> of speed, especially on Windows, it also limits us quite a bit, and I am\n> talking about a lot more than just the awkwardness of having to think\n> about options of BSD vs GNU variants of common command-line tools.\n\nI agree that it costs a lot of time, and I'm even on Linux using\ndash as /bin/sh + eatmydata (but ancient laptop)\n\n> For example, many, many, if not all, test cases, spend the majority of\n> their code on setting up specific scenarios. I don't know about you,\n> but personally I have to dive into many of them when things fail (and I\n> _dread_ the numbers 0021, 0025 and 3070, let me tell you) and I really\n> have to say that most of that code is hard to follow and does not make\n> it easy to form a mental model of what the code tries to accomplish.\n> \n> To address this, a while ago Thomas Rast started to use `fast-export`ed\n> commit histories in test scripts (see e.g. `t/t3206/history.export`). I\n> still find that this fails to make it easier for occasional readers to\n> understand the ideas underlying the test cases.\n> \n> Another approach is to document heavily the ideas first, then use code\n> to implement them. For example, t3430 starts with this:\n> \n> \t[...]\n> \n> \tInitial setup:\n> \n> \t    -- B --                   (first)\n> \t   /       \\\n> \t A - C - D - E - H            (master)\n> \t   \\    \\       /\n> \t    \\    F - G                (second)\n> \t     \\\n> \t      Conflicting-G\n> \n> \t[...]\n> \n> \ttest_commit A &&\n> \tgit checkout -b first &&\n> \ttest_commit B &&\n> \tgit checkout master &&\n> \ttest_commit C &&\n> \ttest_commit D &&\n> \tgit merge --no-commit B &&\n> \ttest_tick &&\n> \tgit commit -m E &&\n> \tgit tag -m E E &&\n> \tgit checkout -b second C &&\n> \ttest_commit F &&\n> \ttest_commit G &&\n> \tgit checkout master &&\n> \tgit merge --no-commit G &&\n> \ttest_tick &&\n> \tgit commit -m H &&\n> \tgit tag -m H H &&\n> \tgit checkout A &&\n> \ttest_commit conflicting-G G.t\n> \n> \t[...]\n> \n> While this is _somewhat_ better than having only the code, I am still\n> unhappy about it: this wall of `test_commit` lines interspersed with\n> other commands is very hard to follow.\n\nAgreed.  More on the readability part below...\n\nAs far as speeding that up, I think moving some parts\nof test setup to Makefiles + fast-import/fast-export would give\nus a nice balance of speed + maintainability:\n\n1. initial setup is done using normal commands (or graph drawing tool)\n2. the result of setup is \"built\" with fast-export\n3. test uses fast-import\n\nMakefile rules would prevent subsequent test runs from repeating\n1. and 2.\n\n> If we were to (slowly) convert our test suite framework to C, we could\n> change that.\n> \n> One idea would be to allow recreating commit history from something that\n> looks like the output of `git log`, or even `git log --graph --oneline`,\n> much like `git mktree` (which really should have been a test helper\n> instead of a Git command, but I digress) takes something that looks like\n> the output of `git ls-tree` and creates a tree object from it.\n\nI've been playing with Graph::Easy (Perl5 module) in other\nprojects, and I also think the setup could be more easily\nexpressed with a declarative language (e.g. GNU make)\n\n> Another thing that would be much easier if we moved more and more parts\n> of the test suite framework to C: we could implement more powerful\n> assertions, a lot more easily. For example, the trace output of a failed\n> `test_i18ngrep` (or `mingw_test_cmp`!!!) could be made a lot more\n> focused on what is going wrong than on cluttering the terminal window\n> with almost useless lines which are tedious to sift through.\n\nI fail to see how language choice here matters.  But then again,\nI have plenty of experience writing bad code in ALL languages I\nknow :>\n\n> Likewise, having a framework in C would make it a lot easier to improve\n> debugging, e.g. by making test scripts \"resumable\" (guarded by an\n> option, it could store a complete state, including a copy of the trash\n> directory, before executing commands, which would allow \"going back in\n> time\" and calling a failing command with a debugger, or with valgrind, or\n> just seeing whether the command would still fail, i.e. whether the test\n> case is flaky).\n\nResumability sounds like a perfect job for GNU make.\n(that said, I don't know if you use make or something else to build gfw)\n\n> In many ways, our current test suite seems to test Git's functionality\n> as much as (core) contributors' abilities to implement test cases in\n> Unix shell script, _correctly_, and maybe also contributors' patience.\n> You could say that it tests for the wrong thing at least half of the\n> time, by design.\n\nBasic (not advanced) sh is already a prerequisite for using git.\n\nWriting correct code and tests in ANY language is still a\nchallenge for me; but I'm least convinced a low-level language\nsuch as C is the right language for writing integration tests in.\n\nC is fine for unit tests, and maybe we can use more unit tests\nand less integration tests.\n\n> It might look like a somewhat less important project, but given that we\n> exercise almost 150,000 test cases with every CI build, I think it does\n> make sense to grind our axe for a while, so to say.\n\nSomething that would benefit both users and regular contributors\nis the use and adoption of more batch and eval-friendly interfaces.\ne.g. fast-import/export, cat-file --batch, for-each-ref --perl...\n\nI haven't used hg since 2005, but I know \"hg server\" exists\nnowadays to get rid of a lot of startup overhead in Mercurial,\nand maybe git could steal that idea, too...\n\n> Therefore, it might be a really good project to modernize our test\n> suite. To take ideas from modern test frameworks such as Jest and try to\n> bring them to C. Which means that new contributors would probably be\n> better suited to work on this project than Git old-timers!\n> \n> And the really neat thing about this project is that it could be done\n> incrementally.\n\nI hope to find time to hack some more batch/eval-friendly stuff\nthat can make scripting git more performant; but no idea on my\navailability :<\n"},{"id":"382853","messageId":"nycvar.QRO.7.76.6.1909241624300.15067@tvgsbejvaqbjf.bet","threadId":"51754","inReplyTo":"20190923180751.GA21344@sigill.intra.peff.net","subject":"Re: Git in Outreachy December 2019?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-09-24T14:25:45Z","receivedAt":"2019-09-24T14:26:13Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Peff,\n\nOn Mon, 23 Sep 2019, Jeff King wrote:\n\n> On Tue, Sep 17, 2019 at 01:23:18PM +0200, Johannes Schindelin wrote:\n>\n> > The only slightly challenging aspect might be that `merge-one-file` is\n> > actually not a merge strategy, but it is used as helper to be passed to\n> > `git merge-index` via the `-o <helper>` option, which makes it slightly\n> > awkward to be implemented as a built-in. A better approach would\n> > therefore be to special-case this value in `merge-index` and execute the\n> > C code directly, without the detour of spawning a built-in.\n>\n> I think it could make sense for merge-index to be able to directly run\n> the merge-one-file code[1]. But I think we'd want to keep its ability to\n> run an arbitrary script, and for people to call merge-one-file\n> separately, since right now you can do:\n>\n>   git merge-index my-script\n>\n> and have \"my-script\" do some processing of its own, then hand off more\n> work to merge-one-file.\n\nOh, sorry, I did not mean to say that we should do away with this at\nall! Rather, I meant to say that `merge-index` could detect when it was\nasked to run `git-merge-one-file` and re-route to internal code instead\nof spawning a process. If any other script/program was specified, it\nshould be spawned off, just like it is done today.\n\nThanks,\nDscho\n\n> So the weird calling convention is actually a user-visible and\n> potentially useful interface. So it would need a deprecation period\n> before being removed.\n>\n> -Peff\n>\n> [1] Certainly doing it in-process would be faster for the common case of\n>     \"git merge-index git-merge-one-file\", but I wonder if anybody really\n>     does that. These days most people would just merge-recursive anyway.\n>\n"},{"id":"382854","messageId":"nycvar.QRO.7.76.6.1909241629530.15067@tvgsbejvaqbjf.bet","threadId":"51754","inReplyTo":"20190923181950.GB21344@sigill.intra.peff.net","subject":"Re: Git in Outreachy December 2019?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-09-24T14:30:20Z","receivedAt":"2019-09-24T14:30:43Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Peff,\n\nOn Mon, 23 Sep 2019, Jeff King wrote:\n\n> On Mon, Sep 23, 2019 at 02:47:23PM +0200, Johannes Schindelin wrote:\n>\n> > The evaluation of the lazy prereq is indeed not different between Bash\n> > or dash. It is nevertheless quite disruptive in the trace of a test\n> > script, especially when it is evaluated for a test case that is skipped\n> > explicitly via the `--run` option.\n>\n> That sounds like a bug: if we know we are not going to run the test\n> anyway due to --run or GIT_TEST_SKIP, we should probably avoid checking\n> the prereq at all.\n\nYep. Likewise, the `&&` chain checker seems to run even if the test case\nis skipped.\n\nCiao,\nDscho\n"},{"id":"382856","messageId":"20190924153316.GA1801@sigill.intra.peff.net","threadId":"51754","inReplyTo":"nycvar.QRO.7.76.6.1909241624300.15067@tvgsbejvaqbjf.bet","subject":"Re: Git in Outreachy December 2019?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-09-24T15:33:17Z","receivedAt":"2019-09-24T15:33:19Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Sep 24, 2019 at 04:25:45PM +0200, Johannes Schindelin wrote:\n\n> > I think it could make sense for merge-index to be able to directly run\n> > the merge-one-file code[1]. But I think we'd want to keep its ability to\n> > run an arbitrary script, and for people to call merge-one-file\n> > separately, since right now you can do:\n> >\n> >   git merge-index my-script\n> >\n> > and have \"my-script\" do some processing of its own, then hand off more\n> > work to merge-one-file.\n> \n> Oh, sorry, I did not mean to say that we should do away with this at\n> all! Rather, I meant to say that `merge-index` could detect when it was\n> asked to run `git-merge-one-file` and re-route to internal code instead\n> of spawning a process. If any other script/program was specified, it\n> should be spawned off, just like it is done today.\n\nOK, great, then we are completely on the same page. :)\n\n-Peff\n"},{"id":"382862","messageId":"20190924170746.100302-1-jonathantanmy@google.com","threadId":"51754","inReplyTo":"20190923212834.GA19504@sigill.intra.peff.net","subject":"Re: Git in Outreachy December 2019?","fromName":"Jonathan Tan","fromEmail":"jonathantanmy@google.com","sentAt":"2019-09-24T17:07:46Z","receivedAt":"2019-09-24T17:07:53Z","isPatch":false,"sender":{"key":"jonathantanmy@fastmail.com","avatar":null},"body":"> On Mon, Sep 23, 2019 at 01:38:54PM -0700, Jonathan Tan wrote:\n> \n> > I didn't have any concrete ideas so I didn't include those, but some\n> > unrefined ideas:\n> \n> One risk to a mentoring project like this is that the intern does a good\n> job of steps 1-5, and then in step 6 we realize that the whole thing is\n> not useful, and upstream doesn't want it. Which isn't to say the intern\n> didn't learn something, and the project didn't benefit. Negative results\n> can be useful; but it can also be demoralizing.\n\nThat's true. I think that libification is in itself a useful and\nnon-controversial goal.\n\n> I'm not arguing that's going to be the case here. But I do think it's\n> worth talking through these things a bit as part of thinking about\n> proposals.\n\n[snip]\n\n> >  - index-pack has the CLI option to specify a message to be written into\n> >    the .promisor file, but in my patch to write fetched refs to\n> >    .promisor [1], I ended up making fetch-pack.c write the information\n> >    because I didn't know how many refs were going to be written (and I\n> >    didn't want to bump into CLI argument length limits). If we had this\n> >    feature, I might have been able to pass a callback to index-pack that\n> >    writes the list of refs once we have the fd into .promisor,\n> >    eliminating some code duplication (but I haven't verified this).\n> \n> That makes some sense. We could pass the data over a pipe, but obviously\n> stdin is already in use to receive the pack here. Ideally we'd be able\n> to pass multiple streams between the programs, but I think due to\n> Windows support, we can't assume that arbitrary pipe descriptors will\n> make it across the run-command boundary. So I think we'd be left with\n> communicating via temporary files (which really isn't the worst thing in\n> the world, but has its own complications).\n> \n> >  - In your reply [2] to the above [1], you mentioned the possibility of\n> >    keeping a list of cutoff points. One way of doing this, as I state in\n> >    [3], is my original suggestion back in 2017 of one such\n> >    repository-wide list. If we do this, it would be better for\n> >    fetch-pack to handle this instead of index-pack, and it seems more\n> >    efficient to me to have index-pack be able to pass objects to\n> >    fetch-pack as they are inflated instead of fetch-pack rereading the\n> >    compressed forms on disk (but again, I haven't verified this).\n> \n> And this is the flip-side problem: we need to get data back, but we have\n> only stdout, which is already in use (so we need some kind of protocol).\n> That leads to things like the horrible NUL-byte added by 83558686ce\n> (receive-pack: send keepalives during quiet periods, 2016-07-15).\n\nSounds good. With this, do you think that there is enough likelihood of\nacceptance that we can move ahead with my proposed project?\n\nBesides discussing the likelihood of patches being accepted/rejected,\nshould we record the result of discussion somewhere (or, if only the\nmentor should give their ideas, for me to write in more detail)? I don't\nrecall a place in the Outreachy form to write this, so I just mentioned\nthe benefits in outline, but maybe I can just include it somewhere\nanyway.\n\n> > There are also the debuggability improvements of not having to deal with\n> > 2 processes.\n> \n> I think it can sometimes be easier to debug with two separate processes,\n> because the input to index-pack is well-defined and can be repeated\n> without hitting the network (though you do have to figure out how to\n> record the network response, which can be non-trivial). I've also done\n> similar things for running performance simulations.\n\nHmm...that's true, but I think this is a matter of degree. The input to\na lib function for index-pack can be similarly simple and well-defined\n(a C interface that we can exercise using a throwaway patch to\ntest-tool, for example), but I agree that it usually won't be as simple\nas input to CLI (but this is because of limitations that the CLI\nimposes).\n\n> > > [dropping unpack-objects]\n> > >     Maybe that would be worth making part of the project?\n> > \n> > I'm reluctant to do so because I don't want to increase the scope too\n> > much - although if my project has relatively narrow scope for an\n> > Outreachy project, we can do so. As for eliminating the utility of\n> > having richer communication, I don't think so, because in the situations\n> > where we require richer communication (right now, situations to do with\n> > partial clone), we specifically run index-pack anyway.\n> \n> Yeah, we're in kind of a weird situation there, where unpack-objects is\n> used less and less. I wonder how many surprises are lurking where\n> somebody reasoned about index-pack behavior, but unpack-objects may do\n> something slightly differently (I know this came up when we looked at\n> fsck-ing incoming objects for submodule vulnerabilities).\n> \n> I kind of wonder if it would be reasonable to just always use index-pack\n> for the sake of simplicity, even if it never learns to actually unpack\n> objects. We've been doing that for years on the server side at GitHub\n> without ill effects (I think the unpack route is slightly more efficient\n> for a thin pack, but since it only kicks in when there are few objects\n> anyway, I wonder how big an advantage it is in general).\n\nThis sounds reasonable to me.\n"},{"id":"382951","messageId":"20190926070900.GA20653@sigill.intra.peff.net","threadId":"51754","inReplyTo":"20190924170746.100302-1-jonathantanmy@google.com","subject":"Re: Git in Outreachy December 2019?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-09-26T07:09:01Z","receivedAt":"2019-09-26T07:09:14Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Sep 24, 2019 at 10:07:46AM -0700, Jonathan Tan wrote:\n\n> > >  - In your reply [2] to the above [1], you mentioned the possibility of\n> > >    keeping a list of cutoff points. One way of doing this, as I state in\n> > >    [3], is my original suggestion back in 2017 of one such\n> > >    repository-wide list. If we do this, it would be better for\n> > >    fetch-pack to handle this instead of index-pack, and it seems more\n> > >    efficient to me to have index-pack be able to pass objects to\n> > >    fetch-pack as they are inflated instead of fetch-pack rereading the\n> > >    compressed forms on disk (but again, I haven't verified this).\n> > \n> > And this is the flip-side problem: we need to get data back, but we have\n> > only stdout, which is already in use (so we need some kind of protocol).\n> > That leads to things like the horrible NUL-byte added by 83558686ce\n> > (receive-pack: send keepalives during quiet periods, 2016-07-15).\n> \n> Sounds good. With this, do you think that there is enough likelihood of\n> acceptance that we can move ahead with my proposed project?\n> \n> Besides discussing the likelihood of patches being accepted/rejected,\n> should we record the result of discussion somewhere (or, if only the\n> mentor should give their ideas, for me to write in more detail)? I don't\n> recall a place in the Outreachy form to write this, so I just mentioned\n> the benefits in outline, but maybe I can just include it somewhere\n> anyway.\n\nYeah, I think it's OK to go ahead. I think an intern who is interested\nin the project would get in touch with you either directly, or via the\nlist. So that gives some opportunity to discuss the ideas with them, and\nto go into more detail on the proposal in an interactive way (it would\nalso be fine to point at this thread, too, of course).\n\n-Peff\n"},{"id":"382973","messageId":"20190926094723.GE2637@szeder.dev","threadId":"51754","inReplyTo":"20190923180649.GA2886@szeder.dev","subject":"Re: Git in Outreachy December 2019?","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2019-09-26T09:47:23Z","receivedAt":"2019-09-26T09:47:30Z","isPatch":false,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Mon, Sep 23, 2019 at 08:07:09PM +0200, SZEDER Gábor wrote:\n> Here is one more idea for microprojects:\n> \n>   Find a group of related preprocessor constants and turn them into an\n>   enum.  Also find where those constants are stored in variables and\n>   in structs and passed around as function parameters, and change the\n>   type of those variables, fields and parameters to the new enum.\n\nPeff thought elsewhere in the thread that this is a good idea, so I\nwanted to try out how this microproject would work in practice, and to \nadd a commit that we can show as a good example, and therefore set out \nto convert 'cache_entry->ce_flags' to an enum...  and will soon send\nout a RFH patch, because I hit a snag, and am not sure what to do\nabout it :)  Anyway:\n\n  - Finding a group of related preprocessor constants is trivial: the\n    common prefixes and vertically aligned values of related constants\n    stand out in output of 'git grep #define'.  Converting them to an\n    enum is fairly trivial as well.\n\n  - Converting various integer types of variables, struct fields, and\n    function parameters to the new enum is... well, I wouldn't say\n    that it's hard, but it's tedious (but 'ce_flags' with about 20\n    related constants is perhaps the biggest we have).  OTOH, it's all \n    fairly mechanical, and doesn't require any understanding of Git\n    internals.  Overall I think that this is indeed a micro-sized\n    microproject, but...\n\n  - The bad news is that I expect that reviewing the variable, etc.\n    type conversions will be just as tedious, and it's quite easy to\n    miss a conversion or three, so I'm afraid that several rerolls\n    will be necessary.\n\n"},{"id":"382975","messageId":"nycvar.QRO.7.76.6.1909261257160.15067@tvgsbejvaqbjf.bet","threadId":"51754","inReplyTo":"20190923165828.GA27068@szeder.dev","subject":"Re: Git in Outreachy December 2019?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-09-26T11:04:48Z","receivedAt":"2019-09-26T11:05:16Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 23 Sep 2019, SZEDER Gábor wrote:\n\n> On Mon, Sep 23, 2019 at 02:47:23PM +0200, Johannes Schindelin wrote:\n> >\n> > On Tue, 17 Sep 2019, SZEDER Gábor wrote:\n> >\n> > > On Tue, Sep 17, 2019 at 01:23:18PM +0200, Johannes Schindelin wrote:\n> > > > Also, things like the code tracing via `-x` (which relies on Bash\n> > > > functionality in order to work properly,\n> > >\n> > > Not really.\n> >\n> > To work properly. What I meant was the trick we need to play with\n> > `BASH_XTRACEFD`.\n>\n> I'm still unsure what BASH_XTRACEFD trick you mean.  AFAICT we don't\n> play any tricks with it to make '-x' work properly, and indeed '-x'\n> tracing works properly even without BASH_XTRACEFD (and to achive that\n> we did have to play some tricks, but not any with BASH_XTRACEFD;\n> perhaps these tricks are what you meant?).\n\nIt works okay some of the time. But IIRC `-x -V` requires the\n`BASH_XTRACEFD` trick.\n\nHowever, I start to feel like I am distracted deliberately from my main\nargument: that shell scripting is simply an awful language to implement\na highly reliable test framework. That we need to rely on Bash, at least\nsome of the time, is just _one_ of the many shortcomings.\n\n> > > > and which _still_ does not work as intended if your test case\n> > > > evaluates a lazy prereq that has not been evaluated before\n> > >\n> > > I don't see any striking differences between the trace output of a test\n> > > involving a lazy prereq from Bash or dash:\n> > >\n> > > [...]\n> >\n> > The evaluation of the lazy prereq is indeed not different between Bash\n> > or dash. It is nevertheless quite disruptive in the trace of a test\n> > script, especially when it is evaluated for a test case that is skipped\n> > explicitly via the `--run` option.\n>\n> But then the actual issue is the unnecessary evaluation of the prereq\n> even when the test framework could know in advance that the test case\n> should be skipped anyway, and the trace from it is a mere side effect,\n> no?\n\nI forgot a crucial tidbit: if you run with `-x` and a lazy prereq is\nevaluated, not only is the output disruptive, the trace is also turned\noff after the lazy prereq, _before_ the actual test case is run. So you\ndon't see any trace of the actual test case.\n\nIn any case, I really do not want to see this thread derailed into\nspecifics of Bashisms and bugs in our test framework.\n\nMy main point should not be diluted: a test framework should be\nimplemented in a language that offers speedy execution of even\ncomplicated logic, proper error checking, and higher data types (i.e.\nother than \"everything is a string\"). Unix shell script is not it.\n\nCiao,\nDscho\n"},{"id":"382976","messageId":"nycvar.QRO.7.76.6.1909261341300.15067@tvgsbejvaqbjf.bet","threadId":"51754","inReplyTo":"20190923180649.GA2886@szeder.dev","subject":"Re: Git in Outreachy December 2019?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-09-26T11:42:12Z","receivedAt":"2019-09-26T11:42:38Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 23 Sep 2019, SZEDER Gábor wrote:\n\n> On Wed, Sep 04, 2019 at 03:41:15PM -0400, Jeff King wrote:\n> > The project page has a section to point people in the right direction\n> > for first-time contributions. I've left it blank for now, but I think it\n> > makes sense to point one (or both) of:\n> >\n> >   - https://git-scm.com/docs/MyFirstContribution\n> >\n> >   - https://matheustavares.gitlab.io/posts/first-steps-contributing-to-git\n> >\n> > as well as a list of micro-projects (or at least instructions on how to\n> > find #leftoverbits, though we'd definitely have to step up our labeling,\n> > as I do not recall having seen one for a while).\n>\n> And we should make sure that all microprojects are indeed micro in\n> size.  Matheus sent v8 of a 10 patch series in July that started out\n> as a microproject back in February...\n\nIndeed.\n\n> Here is one more idea for microprojects:\n>\n>   Find a group of related preprocessor constants and turn them into an\n>   enum.  Also find where those constants are stored in variables and\n>   in structs and passed around as function parameters, and change the\n>   type of those variables, fields and parameters to the new enum.\n\nI agree that this is a good suggestion, and turned this #leftoverbits\ninto https://github.com/gitgitgadget/git/issues/357.\n\nCiao,\nDscho\n"},{"id":"382978","messageId":"nycvar.QRO.7.76.6.1909261343590.15067@tvgsbejvaqbjf.bet","threadId":"51754","inReplyTo":"20190924005529.GA8354@dcvr","subject":"Re: Git in Outreachy December 2019?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-09-26T12:45:09Z","receivedAt":"2019-09-26T12:45:39Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Eric,\n\nOn Tue, 24 Sep 2019, Eric Wong wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> > On Mon, 16 Sep 2019, Emily Shaffer wrote:\n> > >  - try and make progress towards running many tests from a single test\n> > >    file in parallel - maybe this is too big, I'm not sure if we know how\n> > >    many of our tests are order-dependent within a file for now...\n> >\n> > Another, potentially more rewarding, project would be to modernize our\n> > test suite framework, so that it is not based on Unix shell scripting,\n> > but on C instead.\n>\n> I worry more C would reduce the amount of contributors (some of\n> the C rewrites already scared me off hacking years ago).  I\n> figure more users are familiar with sh than C.\n\nSeeing as most of the patches/patch series require contributors not only\nto write test cases in Unix shell script, but also to patch or implement\ncode in C, I fail to be concerned about that.\n\n> It would also increase the disparity between tests and use of\n> actual users from the command-line.\n\nI find it really endearing whenever I hear anybody talking about Git\nusers as if every single one of them crafted extensive shell scripts\naround Git. In my experience, most of the people who do that gather on\nthis here mailing list, or on a nearby one, and that is but a tiny\nfraction of Git users.\n\nBesides, it was my understanding that Git's test suite tries to prevent\nregressions in Git's code, whether it is called in scripts or not. As\nsuch, it would not matter _how_ that functionality is tested. Does that\nnot match your understanding?\n\n> > The fact that it is based on Unix shell scripting not only costs a lot\n> > of speed, especially on Windows, it also limits us quite a bit, and I am\n> > talking about a lot more than just the awkwardness of having to think\n> > about options of BSD vs GNU variants of common command-line tools.\n>\n> I agree that it costs a lot of time, and I'm even on Linux using\n> dash as /bin/sh + eatmydata (but ancient laptop)\n\nOne thing that I meant to play with (but which is hampered by too many\nparts of Git being implemented in Unix shell/Perl, still, therefore\nmaking code coverage analysis hard) is Test Impact Analysis\n(https://docs.microsoft.com/en-us/azure/devops/pipelines/test/test-impact-analysis).\nIn short, a way to avoid running tests the code touched by them was\nalready tested before. Example: if I change the `README.md`, no\nregression test needs to be run at all. If I change `git-p4.py`, the\nmajority of test scripts can be skipped, only t98*.sh need to be run\n(and maybe not even all of them).\n\nThere is a lot of work to be done on the built-in'ification, still,\nbefore that becomes feasible, of course.\n\n> > For example, many, many, if not all, test cases, spend the majority of\n> > their code on setting up specific scenarios. I don't know about you,\n> > but personally I have to dive into many of them when things fail (and I\n> > _dread_ the numbers 0021, 0025 and 3070, let me tell you) and I really\n> > have to say that most of that code is hard to follow and does not make\n> > it easy to form a mental model of what the code tries to accomplish.\n> >\n> > To address this, a while ago Thomas Rast started to use `fast-export`ed\n> > commit histories in test scripts (see e.g. `t/t3206/history.export`). I\n> > still find that this fails to make it easier for occasional readers to\n> > understand the ideas underlying the test cases.\n> >\n> > Another approach is to document heavily the ideas first, then use code\n> > to implement them. For example, t3430 starts with this:\n> >\n> > \t[...]\n> >\n> > \tInitial setup:\n> >\n> > \t    -- B --                   (first)\n> > \t   /       \\\n> > \t A - C - D - E - H            (master)\n> > \t   \\    \\       /\n> > \t    \\    F - G                (second)\n> > \t     \\\n> > \t      Conflicting-G\n> >\n> > \t[...]\n> >\n> > \ttest_commit A &&\n> > \tgit checkout -b first &&\n> > \ttest_commit B &&\n> > \tgit checkout master &&\n> > \ttest_commit C &&\n> > \ttest_commit D &&\n> > \tgit merge --no-commit B &&\n> > \ttest_tick &&\n> > \tgit commit -m E &&\n> > \tgit tag -m E E &&\n> > \tgit checkout -b second C &&\n> > \ttest_commit F &&\n> > \ttest_commit G &&\n> > \tgit checkout master &&\n> > \tgit merge --no-commit G &&\n> > \ttest_tick &&\n> > \tgit commit -m H &&\n> > \tgit tag -m H H &&\n> > \tgit checkout A &&\n> > \ttest_commit conflicting-G G.t\n> >\n> > \t[...]\n> >\n> > While this is _somewhat_ better than having only the code, I am still\n> > unhappy about it: this wall of `test_commit` lines interspersed with\n> > other commands is very hard to follow.\n>\n> Agreed.  More on the readability part below...\n>\n> As far as speeding that up, I think moving some parts\n> of test setup to Makefiles + fast-import/fast-export would give\n> us a nice balance of speed + maintainability:\n>\n> 1. initial setup is done using normal commands (or graph drawing tool)\n> 2. the result of setup is \"built\" with fast-export\n> 3. test uses fast-import\n\nI actually talked about this in my mail. If you find it easy to deduce\nthe intent behind a commit history that was exported via fast-export,\nmore power to you. (Was the committer name crucial? The file name? Or\nthe temporal order of the commits?)\n\nIn contrast, I find it very challenging, myself. And please keep in mind\nthat the first thing any contributor needs to do who sees a failing\nregression test (where the failure is most likely caused by the patch\nthey plan on contributing): understand what the heck the regression test\ncase is trying to ensure. The harder the code makes that, the worse it\ndoes its primary job: to (help) prevent regressions.\n\nSo no, I am not at all on board with moving to fast-imported commit\nhistories in Git's test suite. They provide some convenience to the\nauthors of those regression tests, which is not the audience you need to\ncater for in this case: instead, regression tests should make it not\nonly easy to catch, but _especially_ easy to fix, regressions. And that\naudience would pay dearly for that erstwhile convenience.\n\n> Makefile rules would prevent subsequent test runs from repeating\n> 1. and 2.\n\nThat is a cute idea, until you realize that the number of developers\nfluent in `make` is even smaller than the number of developers fluent in\n`C`. In other words, you would again _increase_ the the number of\nprerequisites instead of reducing it.\n\n> > If we were to (slowly) convert our test suite framework to C, we could\n> > change that.\n> >\n> > One idea would be to allow recreating commit history from something that\n> > looks like the output of `git log`, or even `git log --graph --oneline`,\n> > much like `git mktree` (which really should have been a test helper\n> > instead of a Git command, but I digress) takes something that looks like\n> > the output of `git ls-tree` and creates a tree object from it.\n>\n> I've been playing with Graph::Easy (Perl5 module) in other\n> projects, and I also think the setup could be more easily\n> expressed with a declarative language (e.g. GNU make)\n\nI am dubious. But hey, if you show me something that looks _dead_ easy\nto understand, and even easier to write for new contributors, who am I\nto object?\n\nBut mind, I am not fluent in Perl. I can probably hack my way through\n`git-svn`, but that's a far cry from knowing what I am doing there.\n\nAnd wouldn't that _again_ increase the number of prerequisites on\ncontributors? I mean, you sounded genuinely concerned about that at the\nbeginning of the mail. And I share that concern.\n\n> > Another thing that would be much easier if we moved more and more parts\n> > of the test suite framework to C: we could implement more powerful\n> > assertions, a lot more easily. For example, the trace output of a failed\n> > `test_i18ngrep` (or `mingw_test_cmp`!!!) could be made a lot more\n> > focused on what is going wrong than on cluttering the terminal window\n> > with almost useless lines which are tedious to sift through.\n>\n> I fail to see how language choice here matters.\n\nIn Unix shell script, we either trace (everything) or we don't. In C,\nyou can be a lot more fine-grained with log messages. A *lot*.\n\nWe even already have such a fine-grained log machinery in place, see\n`trace.h`.\n\nAlso, because of the ability to perform more sophisticated locking, lazy\nprerequisites can easily be cached in C, whereas that is not so easy in\nUnix shell (and hence it is not done).\n\n> > Likewise, having a framework in C would make it a lot easier to improve\n> > debugging, e.g. by making test scripts \"resumable\" (guarded by an\n> > option, it could store a complete state, including a copy of the trash\n> > directory, before executing commands, which would allow \"going back in\n> > time\" and calling a failing command with a debugger, or with valgrind, or\n> > just seeing whether the command would still fail, i.e. whether the test\n> > case is flaky).\n>\n> Resumability sounds like a perfect job for GNU make.\n\nUmm.\n\nSo let me give you an example of something I had to debug recently. A\n`git stash apply` marked files outside the sparse checkout as deleted,\nwhen they actually had staged changes during the `git stash` call.\n\nIf this had been a regression test case, it would have looked like this:\n\ntest_expect_success 'stash handles skip-worktree entries nicely' '\n        test_commit A &&\n\techo changed >A.t &&\n\tgit add A.t &&\n\tgit update-index --skip-worktree A.t &&\n\trm A.t &&\n\tgit stash &&\n\n\t: this should not mark A.t as deleted &&\n\tgit stash apply &&\n\ttest -n \"$(git ls-files A.t)\"\n'\n\nNow, the problem would have occurred in the very last command: `A.t`\nwould have been missing from the index.\n\nIn order to debug this, you would have had to put in \"breakpoints\"\n(inserting `false &&` at strategic places), or prefix commands with\n`debug` to start them in GDB, then re-run the test case.\n\nLather, rinse and repeat, until you figured out that `git stash` was\nfailing to record `A.t` properly.\n\nThen dig into that, recreating the same worktree situation every time\nyou run `git stash`, until you find out that there is a call to `git\nupdate-index -z --add --remove --stdin` that removes that file.\n\nFurther investigation would show you that this command pretty much does\nwhat it is told to, because it is fed `A.t` in its `stdin`.\n\nThe next step would probably be to go back even one more step and see\nthat `diff_index()` reported this file as modified between the index and\n`HEAD`.\n\nSlowly, you would form an understanding of what is going wrong, and you\nwould have to go back and forth between blaming `diff_index()` and\n`update-index`, or the options `git stash` passes to them.\n\nYou would have to recreate the worktree many, many times, in order to\ndig in deep, and of course you would need to understand the intention\nnot only of the regression test, but also of the code it calls.\n\nIn this instance, it is merely tedious, but possible. I know, because I\ndid it. For flaky tests, not so much.\n\n*That* is the scenario I tried to get at.\n\nWriting the test cases in `make` would not help that. Not one bit. It\nwould actually make things a lot worse.\n\n> (that said, I don't know if you use make or something else to build\n> gfw)\n\nWe are talking about the test suite, yes?\n\nNot about building Git? Because it does not matter whether we use `make`\nor not (we do by default, although we also have an option to build in\nVisual Studio and/or via MSBuild).\n\nOn Windows, just like on Linux, we use a Unix shell interpreter (Bash).\nSure, to run the entire test suite, we use `make` (sometimes in\nconjunction with `prove`), but the tests themselves are written in Unix\nshell script, so I have a hard time imagining a different method to run\nthem -- whether on Windows or not -- than to use a Unix shell\ninterpreter such as Bash.\n\n> > In many ways, our current test suite seems to test Git's\n> > functionality as much as (core) contributors' abilities to implement\n> > test cases in Unix shell script, _correctly_, and maybe also\n> > contributors' patience.  You could say that it tests for the wrong\n> > thing at least half of the time, by design.\n>\n> Basic (not advanced) sh is already a prerequisite for using git.\n\nWell, if you are happy with that prerequisite to be set in stone, I am\nnot. Why should any Git user *need* to know sh?\n\n> Writing correct code and tests in ANY language is still a\n> challenge for me; but I'm least convinced a low-level language\n> such as C is the right language for writing integration tests in.\n\nI would be delighted to go for a proper, well-maintained test framework\nsuch as Jest. But of course, that would not be accepted in this project.\nSo I won't even think about it.\n\n> C is fine for unit tests, and maybe we can use more unit tests and\n> less integration tests.\n\nWe don't have integration tests.\n\nUnless your concept of what constitutes an \"integration test\" is very\ndifferent from mine. For me, an integration test would set up an\nenvironment that involves multiple points of failure, e.g. setting up an\nSSH server on Linux and accessing that via Git from Ubuntu. Or setting\nup a web server with a self-signed certificate, import the public key\ninto the Windows Certificate Store and then accessing the server both\nusing OpenSSL and Secure Channel (the native Windows way to communicate\nvia TLS).\n\nThere is nothing even close to that in Git's test suite.\n\n> > It might look like a somewhat less important project, but given that we\n> > exercise almost 150,000 test cases with every CI build, I think it does\n> > make sense to grind our axe for a while, so to say.\n>\n> Something that would benefit both users and regular contributors\n> is the use and adoption of more batch and eval-friendly interfaces.\n> e.g. fast-import/export, cat-file --batch, for-each-ref --perl...\n\nGiven how hard it is to deduce the intention behind such invocations, I\nam rather doubtful that this would improve our test suite.\n\n> I haven't used hg since 2005, but I know \"hg server\" exists\n> nowadays to get rid of a lot of startup overhead in Mercurial,\n> and maybe git could steal that idea, too...\n\nI have no idea what `hg server` does. Care to enlighten me?\n\n> > Therefore, it might be a really good project to modernize our test\n> > suite. To take ideas from modern test frameworks such as Jest and try to\n> > bring them to C. Which means that new contributors would probably be\n> > better suited to work on this project than Git old-timers!\n> >\n> > And the really neat thing about this project is that it could be done\n> > incrementally.\n>\n> I hope to find time to hack some more batch/eval-friendly stuff\n> that can make scripting git more performant; but no idea on my\n> availability :<\n\nKnock yourself out, if you enjoy that type of project. And who knows,\nmaybe you will convince me yet that it benefits the tests...\n\nCiao,\nDscho\n"},{"id":"382986","messageId":"20190926132852.GF2637@szeder.dev","threadId":"51754","inReplyTo":"nycvar.QRO.7.76.6.1909261257160.15067@tvgsbejvaqbjf.bet","subject":"Re: Git in Outreachy December 2019?","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2019-09-26T13:28:52Z","receivedAt":"2019-09-26T13:28:59Z","isPatch":false,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Thu, Sep 26, 2019 at 01:04:48PM +0200, Johannes Schindelin wrote:\n> > > > > Also, things like the code tracing via `-x` (which relies on Bash\n> > > > > functionality in order to work properly,\n> > > >\n> > > > Not really.\n> > >\n> > > To work properly. What I meant was the trick we need to play with\n> > > `BASH_XTRACEFD`.\n> >\n> > I'm still unsure what BASH_XTRACEFD trick you mean.  AFAICT we don't\n> > play any tricks with it to make '-x' work properly, and indeed '-x'\n> > tracing works properly even without BASH_XTRACEFD (and to achive that\n> > we did have to play some tricks, but not any with BASH_XTRACEFD;\n> > perhaps these tricks are what you meant?).\n> \n> It works okay some of the time.\n\nAs far as I can tell it works all the time.\n\n(Well, Ok, with the exception of t1510, but only because back then I\ncouldn't be bothered to figure out how that test script works.  But\neven that script handles '-x' without BASH_XTRACEFD gracefully, and\nit's safe to run the whole test suite with '-x'.)\n\n>  But IIRC `-x -V` requires the `BASH_XTRACEFD` trick.\n\nNo, it doesn't; '-V' should have no effect on the '-x' trace\nwhatsoever.\n\nAs soon as I fixed running the test suite with '-x' and /bin/sh I\nadded GIT_TEST_OPTS=\"--verbose-log -x\" to my 'config.mak' and to our\nCI scripts.  The default shell running the test suite in our Linux CI\njobs is dash, and in our macOS jobs it's an ancient Bash version that\ndoesn't yet have BASH_XTRACEFD.  As far as I know they all work as\nthey should.\n\n> However, I start to feel like I am distracted deliberately from my main\n> argument\n\nThat was definitely not my intention.  However, if there are any open\nissues with '-x', then I do want to know about it and fix it sooner\nrather than later.  Alas, I still don't have the slightest clue about\nwhat your issue actually is.\n\n> I forgot a crucial tidbit: if you run with `-x` and a lazy prereq is\n> evaluated, not only is the output disruptive, the trace is also turned\n> off after the lazy prereq, _before_ the actual test case is run. So you\n> don't see any trace of the actual test case.\n\nTracing is always turned on before running the test case,\nindependently from whether a lazy prereq was evaluated or not, so we\ndo always see the trace of the actual test case.  Notice the '+ true'\nand '++ true' lines in my earlier reply including the test traces:\nthose lines are the trace of the actual test case.\n\n  https://public-inbox.org/git/20190917120230.GA27531@szeder.dev/\n\n"},{"id":"383019","messageId":"nycvar.QRO.7.76.6.1909262132090.15067@tvgsbejvaqbjf.bet","threadId":"51754","inReplyTo":"20190926094723.GE2637@szeder.dev","subject":"Re: Git in Outreachy December 2019?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-09-26T19:32:35Z","receivedAt":"2019-09-26T19:32:57Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 26 Sep 2019, SZEDER Gábor wrote:\n\n> On Mon, Sep 23, 2019 at 08:07:09PM +0200, SZEDER Gábor wrote:\n> > Here is one more idea for microprojects:\n> >\n> >   Find a group of related preprocessor constants and turn them into an\n> >   enum.  Also find where those constants are stored in variables and\n> >   in structs and passed around as function parameters, and change the\n> >   type of those variables, fields and parameters to the new enum.\n>\n> Peff thought elsewhere in the thread that this is a good idea, so I\n> wanted to try out how this microproject would work in practice, and to\n> add a commit that we can show as a good example, and therefore set out\n> to convert 'cache_entry->ce_flags' to an enum...  and will soon send\n> out a RFH patch, because I hit a snag, and am not sure what to do\n> about it :)  Anyway:\n>\n>   - Finding a group of related preprocessor constants is trivial: the\n>     common prefixes and vertically aligned values of related constants\n>     stand out in output of 'git grep #define'.  Converting them to an\n>     enum is fairly trivial as well.\n>\n>   - Converting various integer types of variables, struct fields, and\n>     function parameters to the new enum is... well, I wouldn't say\n>     that it's hard, but it's tedious (but 'ce_flags' with about 20\n>     related constants is perhaps the biggest we have).  OTOH, it's all\n>     fairly mechanical, and doesn't require any understanding of Git\n>     internals.  Overall I think that this is indeed a micro-sized\n>     microproject, but...\n>\n>   - The bad news is that I expect that reviewing the variable, etc.\n>     type conversions will be just as tedious, and it's quite easy to\n>     miss a conversion or three, so I'm afraid that several rerolls\n>     will be necessary.\n\nI thought Coccinelle could help with that?\n\nCiao,\nDscho\n"},{"id":"383021","messageId":"nycvar.QRO.7.76.6.1909262138450.15067@tvgsbejvaqbjf.bet","threadId":"51754","inReplyTo":"20190926132852.GF2637@szeder.dev","subject":"Re: Git in Outreachy December 2019?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-09-26T19:39:58Z","receivedAt":"2019-09-26T19:40:21Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 26 Sep 2019, SZEDER Gábor wrote:\n\n> On Thu, Sep 26, 2019 at 01:04:48PM +0200, Johannes Schindelin wrote:\n> > > > > > Also, things like the code tracing via `-x` (which relies on Bash\n> > > > > > functionality in order to work properly,\n> > > > >\n> > > > > Not really.\n> > > >\n> > > > To work properly. What I meant was the trick we need to play with\n> > > > `BASH_XTRACEFD`.\n> > >\n> > > I'm still unsure what BASH_XTRACEFD trick you mean.  AFAICT we don't\n> > > play any tricks with it to make '-x' work properly, and indeed '-x'\n> > > tracing works properly even without BASH_XTRACEFD (and to achive that\n> > > we did have to play some tricks, but not any with BASH_XTRACEFD;\n> > > perhaps these tricks are what you meant?).\n> >\n> > It works okay some of the time.\n>\n> As far as I can tell it works all the time.\n\nI must be misinterpreting this part of `t/test-lib.sh`, then:\n\n-- snipsnap --\nif test -n \"$trace\" && test -n \"$test_untraceable\"\nthen\n\t# '-x' tracing requested, but this test script can't be reliably\n\t# traced, unless it is run with a Bash version supporting\n\t# BASH_XTRACEFD (introduced in Bash v4.1).\n\t#\n\t# Perform this version check _after_ the test script was\n\t# potentially re-executed with $TEST_SHELL_PATH for '--tee' or\n\t# '--verbose-log', so the right shell is checked and the\n\t# warning is issued only once.\n\tif test -n \"$BASH_VERSION\" && eval '\n\t     test ${BASH_VERSINFO[0]} -gt 4 || {\n\t       test ${BASH_VERSINFO[0]} -eq 4 &&\n\t       test ${BASH_VERSINFO[1]} -ge 1\n\t     }\n\t   '\n\tthen\n\t\t: Executed by a Bash version supporting BASH_XTRACEFD.  Good.\n\telse\n\t\techo >&2 \"warning: ignoring -x; '$0' is untraceable without BASH_XTRACEFD\"\n\t\ttrace=\n\tfi\nfi\n"},{"id":"383039","messageId":"20190926214448.GI2637@szeder.dev","threadId":"51754","inReplyTo":"nycvar.QRO.7.76.6.1909262138450.15067@tvgsbejvaqbjf.bet","subject":"Re: Git in Outreachy December 2019?","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2019-09-26T21:44:48Z","receivedAt":"2019-09-26T21:44:55Z","isPatch":false,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Thu, Sep 26, 2019 at 09:39:58PM +0200, Johannes Schindelin wrote:\n> Hi,\n> \n> On Thu, 26 Sep 2019, SZEDER Gábor wrote:\n> \n> > On Thu, Sep 26, 2019 at 01:04:48PM +0200, Johannes Schindelin wrote:\n> > > > > > > Also, things like the code tracing via `-x` (which relies on Bash\n> > > > > > > functionality in order to work properly,\n> > > > > >\n> > > > > > Not really.\n> > > > >\n> > > > > To work properly. What I meant was the trick we need to play with\n> > > > > `BASH_XTRACEFD`.\n> > > >\n> > > > I'm still unsure what BASH_XTRACEFD trick you mean.  AFAICT we don't\n> > > > play any tricks with it to make '-x' work properly, and indeed '-x'\n> > > > tracing works properly even without BASH_XTRACEFD (and to achive that\n> > > > we did have to play some tricks, but not any with BASH_XTRACEFD;\n> > > > perhaps these tricks are what you meant?).\n> > >\n> > > It works okay some of the time.\n> >\n> > As far as I can tell it works all the time.\n> \n> I must be misinterpreting this part of `t/test-lib.sh`, then:\n\nOk, let me try to clarify.\n\nThere are a couple of things that we can't do in our tests without\nBASH_XTRACEFD, e.g. redirecting the standard error of a subshell or a\nloop to a file and then check that file with 'test_cmp' or\n'test_must_be_empty'.  With tracing enabled but without BASH_XTRACEFD,\nthe trace of the commands executed within the subshell or loop end up\nin that file as well, and cause failure (grepping through that file is\nmostly ok, though).  Back then we had 23 test cases failing because\nthey were doing things like this and needed to be fixed, so\nconsidering the total number of test cases we only rarely used such\nproblematic constructs.\n\nStill, as I recall, Peff was concerned that these limitations might\nlead to maintenance burden on the long run, so I decided to add an\nescape hatch, just in case someone constructs such an elaborate test\nscript, where redirecting the stderr of a compound command could\nconsiderably simplify the tests. \n\nThat snippet of code that you copied is this escape hatch: if \n$test_untraceable is set to a non-empty value before sourcing\n'test-lib.sh', then tracing will only be enabled if BASH_XTRACEFD is\navailable.\n\nAll that was over a year and a half ago, and these limitations weren't\na maintenance burden at all so far, and nobody needed that escape\nhatch.\n\nWell, nobody except me, that is :)  When I saw back then that t1510\nsaves the stderr of nested function calls with 7 parameters, I\nshrugged in disgust, admitted defeat, and simply reached for that\nescape hatch: partly because I couldn't be bothered to figure out how\nthat test script works, but more importantly because I didn't want to\nrisk that any cleanup inadvertently hides a bug in the future.\n\nSo that's the only user that piece of code ever had, and I certainly\nhope that no other test script will ever grow so complicated that it\nwill need this escape hatch.  I would actually prefer to remove it,\nbut t1510 must be cleaned up first...  so I'm afraid it will be with\nus for a while.\n\n\n> -- snipsnap --\n> if test -n \"$trace\" && test -n \"$test_untraceable\"\n> then\n> \t# '-x' tracing requested, but this test script can't be reliably\n> \t# traced, unless it is run with a Bash version supporting\n> \t# BASH_XTRACEFD (introduced in Bash v4.1).\n> \t#\n> \t# Perform this version check _after_ the test script was\n> \t# potentially re-executed with $TEST_SHELL_PATH for '--tee' or\n> \t# '--verbose-log', so the right shell is checked and the\n> \t# warning is issued only once.\n> \tif test -n \"$BASH_VERSION\" && eval '\n> \t     test ${BASH_VERSINFO[0]} -gt 4 || {\n> \t       test ${BASH_VERSINFO[0]} -eq 4 &&\n> \t       test ${BASH_VERSINFO[1]} -ge 1\n> \t     }\n> \t   '\n> \tthen\n> \t\t: Executed by a Bash version supporting BASH_XTRACEFD.  Good.\n> \telse\n> \t\techo >&2 \"warning: ignoring -x; '$0' is untraceable without BASH_XTRACEFD\"\n> \t\ttrace=\n> \tfi\n> fi\n\n"},{"id":"383040","messageId":"20190926215448.GJ2637@szeder.dev","threadId":"51754","inReplyTo":"nycvar.QRO.7.76.6.1909262132090.15067@tvgsbejvaqbjf.bet","subject":"Re: Git in Outreachy December 2019?","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2019-09-26T21:54:48Z","receivedAt":"2019-09-26T21:54:54Z","isPatch":false,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Thu, Sep 26, 2019 at 09:32:35PM +0200, Johannes Schindelin wrote:\n> Hi,\n> \n> On Thu, 26 Sep 2019, SZEDER Gábor wrote:\n> \n> > On Mon, Sep 23, 2019 at 08:07:09PM +0200, SZEDER Gábor wrote:\n> > > Here is one more idea for microprojects:\n> > >\n> > >   Find a group of related preprocessor constants and turn them into an\n> > >   enum.  Also find where those constants are stored in variables and\n> > >   in structs and passed around as function parameters, and change the\n> > >   type of those variables, fields and parameters to the new enum.\n> >\n> > Peff thought elsewhere in the thread that this is a good idea, so I\n> > wanted to try out how this microproject would work in practice, and to\n> > add a commit that we can show as a good example, and therefore set out\n> > to convert 'cache_entry->ce_flags' to an enum...  and will soon send\n> > out a RFH patch, because I hit a snag, and am not sure what to do\n> > about it :)  Anyway:\n> >\n> >   - Finding a group of related preprocessor constants is trivial: the\n> >     common prefixes and vertically aligned values of related constants\n> >     stand out in output of 'git grep #define'.  Converting them to an\n> >     enum is fairly trivial as well.\n> >\n> >   - Converting various integer types of variables, struct fields, and\n> >     function parameters to the new enum is... well, I wouldn't say\n> >     that it's hard, but it's tedious (but 'ce_flags' with about 20\n> >     related constants is perhaps the biggest we have).  OTOH, it's all\n> >     fairly mechanical, and doesn't require any understanding of Git\n> >     internals.  Overall I think that this is indeed a micro-sized\n> >     microproject, but...\n> >\n> >   - The bad news is that I expect that reviewing the variable, etc.\n> >     type conversions will be just as tedious, and it's quite easy to\n> >     miss a conversion or three, so I'm afraid that several rerolls\n> >     will be necessary.\n> \n> I thought Coccinelle could help with that?\n\nMaybe it could, I don't know.  I mean, it should be able to e.g.\nchange the data type of the function parameter 'param' if the function\nbody contains a\n\n  param & <any named value of the new enum or their combinarion>\n\nstatement, or similar statements with operators '|=', '&=', or '='.\nI have no idea how to tell Coccinelle to do that.\n\n"},{"id":"383074","messageId":"20190927221857.GB31237@sigill.intra.peff.net","threadId":"51754","inReplyTo":"20190926214448.GI2637@szeder.dev","subject":"Re: Git in Outreachy December 2019?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-09-27T22:18:58Z","receivedAt":"2019-09-27T22:19:00Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Sep 26, 2019 at 11:44:48PM +0200, SZEDER Gábor wrote:\n\n> All that was over a year and a half ago, and these limitations weren't\n> a maintenance burden at all so far, and nobody needed that escape\n> hatch.\n> \n> Well, nobody except me, that is :)  When I saw back then that t1510\n> saves the stderr of nested function calls with 7 parameters, I\n> shrugged in disgust, admitted defeat, and simply reached for that\n> escape hatch: partly because I couldn't be bothered to figure out how\n> that test script works, but more importantly because I didn't want to\n> risk that any cleanup inadvertently hides a bug in the future.\n> \n> So that's the only user that piece of code ever had, and I certainly\n> hope that no other test script will ever grow so complicated that it\n> will need this escape hatch.  I would actually prefer to remove it,\n> but t1510 must be cleaned up first...  so I'm afraid it will be with\n> us for a while.\n\nI'm actually surprised we haven't run into it more. We have some custom\ntest scripts in our fork of Git at GitHub. We usually just use\nTEST_SHELL_PATH=bash, but curious, I tried running with dash and \"-x\",\nand three of them failed.\n\nProbably they'd be easy enough to fix (and they're out of tree anyway),\nso I'm not really arguing against the escape hatch exactly. Mostly I'm\njust surprised that if I introduced 3 cases (out of probably a dozen\nscripts), I'm surprised that more contributors aren't accidentally doing\nso upstream.\n\n-Peff\n"},{"id":"383079","messageId":"xmqqblv5kr9u.fsf@gitster-ct.c.googlers.com","threadId":"51754","inReplyTo":"20190924153316.GA1801@sigill.intra.peff.net","subject":"Re: Git in Outreachy December 2019?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-09-28T03:56:29Z","receivedAt":"2019-09-28T03:56:35Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> On Tue, Sep 24, 2019 at 04:25:45PM +0200, Johannes Schindelin wrote:\n>\n>> > I think it could make sense for merge-index to be able to directly run\n>> > the merge-one-file code[1]. But I think we'd want to keep its ability to\n>> > run an arbitrary script, and for people to call merge-one-file\n>> > separately, since right now you can do:\n>> >\n>> >   git merge-index my-script\n>> >\n>> > and have \"my-script\" do some processing of its own, then hand off more\n>> > work to merge-one-file.\n>> \n>> Oh, sorry, I did not mean to say that we should do away with this at\n>> all! Rather, I meant to say that `merge-index` could detect when it was\n>> asked to run `git-merge-one-file` and re-route to internal code instead\n>> of spawning a process. If any other script/program was specified, it\n>> should be spawned off, just like it is done today.\n>\n> OK, great, then we are completely on the same page. :)\n\nI wondered briefly if we want tospecial case the string\n'git-merge-one-file' (iow, teaching merge-index a new option\n\"--use-builtin-merge-one-file\" and update our own use in\ngit-merge-octopus.sh, git-merge-resolve.sh, etc. would be safer).\n\nBut \"git merge-index git-merge-one-file\" called in the context of\nscrpted Porcelain has the directory that has *OUR*\ngit-merge-one-file as the first element on $PATH, I think, so it may\nbe perfectly safe to use the \"ah, the command name we are asked to\nrun happens to be git-merge-one-file, so let's use the internal\nversion instead without spawning\" short-cut.\n\n"},{"id":"383080","messageId":"xmqq7e5tkr28.fsf@gitster-ct.c.googlers.com","threadId":"51754","inReplyTo":"20190924005529.GA8354@dcvr","subject":"Re: Git in Outreachy December 2019?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-09-28T04:01:03Z","receivedAt":"2019-09-28T04:01:10Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Eric Wong <e@80x24.org> writes:\n\n> C is fine for unit tests, and maybe we can use more unit tests\n> and less integration tests.\n\nNicely put.  I often find it somewhat disturbing that what some of\nthe t/helper/ tests are trying to exercise is at too low a level\nthat the distance from the real-world observable effect is too many\nhops detached.  For unit tests (of an API, for example), that is\nexactly what we want.  For a test of an entire command, it feels\nlike scratching foot from outside while still wearing a shoe.\n\n> I hope to find time to hack some more batch/eval-friendly stuff\n> that can make scripting git more performant; but no idea on my\n> availability :<\n\n"},{"id":"383130","messageId":"20190930085512.GA21522@dcvr","threadId":"51754","inReplyTo":"nycvar.QRO.7.76.6.1909261343590.15067@tvgsbejvaqbjf.bet","subject":"Re: Git in Outreachy December 2019?","fromName":"Eric Wong","fromEmail":"e@80x24.org","sentAt":"2019-09-30T08:55:12Z","receivedAt":"2019-09-30T08:55:16Z","isPatch":false,"sender":{"key":"e@80x24.org","avatar":null},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> On Tue, 24 Sep 2019, Eric Wong wrote:\n> > Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> > > On Mon, 16 Sep 2019, Emily Shaffer wrote:\n> > > >  - try and make progress towards running many tests from a single test\n> > > >    file in parallel - maybe this is too big, I'm not sure if we know how\n> > > >    many of our tests are order-dependent within a file for now...\n> > >\n> > > Another, potentially more rewarding, project would be to modernize our\n> > > test suite framework, so that it is not based on Unix shell scripting,\n> > > but on C instead.\n> >\n> > I worry more C would reduce the amount of contributors (some of\n> > the C rewrites already scared me off hacking years ago).  I\n> > figure more users are familiar with sh than C.\n> \n> Seeing as most of the patches/patch series require contributors not only\n> to write test cases in Unix shell script, but also to patch or implement\n> code in C, I fail to be concerned about that.\n\nMy point was that it was easier to experiment using a less\ntedious and verbose language than C.  git already has a lot of\ncommands and options, but they're mostly documented.  Having to\nlearn our internal C APIs on top of that is a lot of cognitive\noverhead (and I'm still learning our APIs, myself).\n\n> > It would also increase the disparity between tests and use of\n> > actual users from the command-line.\n> \n> I find it really endearing whenever I hear anybody talking about Git\n> users as if every single one of them crafted extensive shell scripts\n> around Git. In my experience, most of the people who do that gather on\n> this here mailing list, or on a nearby one, and that is but a tiny\n> fraction of Git users.\n\nI only mean stuff like: \"git add foo && git commit -m ...\"\nin tests, which is standard command-line usage and using\nour documented UI.\n\nFwiw, I don't have extensive scripts or customizations\naround git or any of the tools I use, either.\n\n> Besides, it was my understanding that Git's test suite tries to prevent\n> regressions in Git's code, whether it is called in scripts or not. As\n> such, it would not matter _how_ that functionality is tested. Does that\n> not match your understanding?\n\nThat matches my understanding.  My concern that most of our\nfunctionality is exposed to users as commands; so new developers\nwould feel more comfortable recreating tests off failures\nthey encounter under normal use, instead of having to learn\nC API internals.\n\n> > > The fact that it is based on Unix shell scripting not only costs a lot\n> > > of speed, especially on Windows, it also limits us quite a bit, and I am\n> > > talking about a lot more than just the awkwardness of having to think\n> > > about options of BSD vs GNU variants of common command-line tools.\n> >\n> > I agree that it costs a lot of time, and I'm even on Linux using\n> > dash as /bin/sh + eatmydata (but ancient laptop)\n> \n> One thing that I meant to play with (but which is hampered by too many\n> parts of Git being implemented in Unix shell/Perl, still, therefore\n> making code coverage analysis hard) is Test Impact Analysis\n> (https://docs.microsoft.com/en-us/azure/devops/pipelines/test/test-impact-analysis).\n> In short, a way to avoid running tests the code touched by them was\n> already tested before. Example: if I change the `README.md`, no\n> regression test needs to be run at all. If I change `git-p4.py`, the\n> majority of test scripts can be skipped, only t98*.sh need to be run\n> (and maybe not even all of them).\n\nCool.  I wonder how much effort it would take to do with gcov +\nDevel::Cover.\n\n> There is a lot of work to be done on the built-in'ification, still,\n> before that becomes feasible, of course.\n> \n> > > For example, many, many, if not all, test cases, spend the majority of\n> > > their code on setting up specific scenarios. I don't know about you,\n> > > but personally I have to dive into many of them when things fail (and I\n> > > _dread_ the numbers 0021, 0025 and 3070, let me tell you) and I really\n> > > have to say that most of that code is hard to follow and does not make\n> > > it easy to form a mental model of what the code tries to accomplish.\n> > >\n> > > To address this, a while ago Thomas Rast started to use `fast-export`ed\n> > > commit histories in test scripts (see e.g. `t/t3206/history.export`). I\n> > > still find that this fails to make it easier for occasional readers to\n> > > understand the ideas underlying the test cases.\n> > >\n> > > Another approach is to document heavily the ideas first, then use code\n> > > to implement them. For example, t3430 starts with this:\n> > >\n> > > \t[...]\n> > >\n> > > \tInitial setup:\n> > >\n> > > \t    -- B --                   (first)\n> > > \t   /       \\\n> > > \t A - C - D - E - H            (master)\n> > > \t   \\    \\       /\n> > > \t    \\    F - G                (second)\n> > > \t     \\\n> > > \t      Conflicting-G\n> > >\n> > > \t[...]\n> > >\n> > > \ttest_commit A &&\n> > > \tgit checkout -b first &&\n> > > \ttest_commit B &&\n> > > \tgit checkout master &&\n> > > \ttest_commit C &&\n> > > \ttest_commit D &&\n> > > \tgit merge --no-commit B &&\n> > > \ttest_tick &&\n> > > \tgit commit -m E &&\n> > > \tgit tag -m E E &&\n> > > \tgit checkout -b second C &&\n> > > \ttest_commit F &&\n> > > \ttest_commit G &&\n> > > \tgit checkout master &&\n> > > \tgit merge --no-commit G &&\n> > > \ttest_tick &&\n> > > \tgit commit -m H &&\n> > > \tgit tag -m H H &&\n> > > \tgit checkout A &&\n> > > \ttest_commit conflicting-G G.t\n> > >\n> > > \t[...]\n> > >\n> > > While this is _somewhat_ better than having only the code, I am still\n> > > unhappy about it: this wall of `test_commit` lines interspersed with\n> > > other commands is very hard to follow.\n> >\n> > Agreed.  More on the readability part below...\n> >\n> > As far as speeding that up, I think moving some parts\n> > of test setup to Makefiles + fast-import/fast-export would give\n> > us a nice balance of speed + maintainability:\n> >\n> > 1. initial setup is done using normal commands (or graph drawing tool)\n> > 2. the result of setup is \"built\" with fast-export\n> > 3. test uses fast-import\n> \n> I actually talked about this in my mail. If you find it easy to deduce\n> the intent behind a commit history that was exported via fast-export,\n> more power to you. (Was the committer name crucial? The file name? Or\n> the temporal order of the commits?)\n> \n> In contrast, I find it very challenging, myself. And please keep in mind\n> that the first thing any contributor needs to do who sees a failing\n> regression test (where the failure is most likely caused by the patch\n> they plan on contributing): understand what the heck the regression test\n> case is trying to ensure. The harder the code makes that, the worse it\n> does its primary job: to (help) prevent regressions.\n> \n> So no, I am not at all on board with moving to fast-imported commit\n> histories in Git's test suite. They provide some convenience to the\n> authors of those regression tests, which is not the audience you need to\n> cater for in this case: instead, regression tests should make it not\n> only easy to catch, but _especially_ easy to fix, regressions. And that\n> audience would pay dearly for that erstwhile convenience.\n\nSorry I wasn't clear.  I only want fast-export/import to be used\nas a local cache mechanism for tests.  I don't want to be\nwriting or maintaining fast-import histories by hand, either :>\n\n> > Makefile rules would prevent subsequent test runs from repeating\n> > 1. and 2.\n> \n> That is a cute idea, until you realize that the number of developers\n> fluent in `make` is even smaller than the number of developers fluent in\n> `C`. In other words, you would again _increase_ the the number of\n> prerequisites instead of reducing it.\n\nWouldn't any C developer need to know SOME build system?\n\nWe already use make, and make uses sh.  I'm not especially\nfluent in make or sh, either; but I can hack my way through\nthem.\n\n> > > If we were to (slowly) convert our test suite framework to C, we could\n> > > change that.\n> > >\n> > > One idea would be to allow recreating commit history from something that\n> > > looks like the output of `git log`, or even `git log --graph --oneline`,\n> > > much like `git mktree` (which really should have been a test helper\n> > > instead of a Git command, but I digress) takes something that looks like\n> > > the output of `git ls-tree` and creates a tree object from it.\n> >\n> > I've been playing with Graph::Easy (Perl5 module) in other\n> > projects, and I also think the setup could be more easily\n> > expressed with a declarative language (e.g. GNU make)\n> \n> I am dubious. But hey, if you show me something that looks _dead_ easy\n> to understand, and even easier to write for new contributors, who am I\n> to object?\n> \n> But mind, I am not fluent in Perl. I can probably hack my way through\n> `git-svn`, but that's a far cry from knowing what I am doing there.\n> \n> And wouldn't that _again_ increase the number of prerequisites on\n> contributors? I mean, you sounded genuinely concerned about that at the\n> beginning of the mail. And I share that concern.\n\nI'm not saying we use Graph::Easy, but we could probably take\ninspiration from it for creating something that makes creating\nhistories easy and maintainable.\n\nBut it could also be writing better comments, using \"git commit\"\nas normal, and letting fast-export cache to speed up subsequent\nruns.\n\n> > > Another thing that would be much easier if we moved more and more parts\n> > > of the test suite framework to C: we could implement more powerful\n> > > assertions, a lot more easily. For example, the trace output of a failed\n> > > `test_i18ngrep` (or `mingw_test_cmp`!!!) could be made a lot more\n> > > focused on what is going wrong than on cluttering the terminal window\n> > > with almost useless lines which are tedious to sift through.\n> >\n> > I fail to see how language choice here matters.\n> \n> In Unix shell script, we either trace (everything) or we don't. In C,\n> you can be a lot more fine-grained with log messages. A *lot*.\n> \n> We even already have such a fine-grained log machinery in place, see\n> `trace.h`.\n> \n> Also, because of the ability to perform more sophisticated locking, lazy\n> prerequisites can easily be cached in C, whereas that is not so easy in\n> Unix shell (and hence it is not done).\n\n`make` seems good at driving stuff to cache.\n\n> > > Likewise, having a framework in C would make it a lot easier to improve\n> > > debugging, e.g. by making test scripts \"resumable\" (guarded by an\n> > > option, it could store a complete state, including a copy of the trash\n> > > directory, before executing commands, which would allow \"going back in\n> > > time\" and calling a failing command with a debugger, or with valgrind, or\n> > > just seeing whether the command would still fail, i.e. whether the test\n> > > case is flaky).\n> >\n> > Resumability sounds like a perfect job for GNU make.\n> \n> Umm.\n> \n> So let me give you an example of something I had to debug recently. A\n> `git stash apply` marked files outside the sparse checkout as deleted,\n> when they actually had staged changes during the `git stash` call.\n> \n> If this had been a regression test case, it would have looked like this:\n> \n> test_expect_success 'stash handles skip-worktree entries nicely' '\n>         test_commit A &&\n> \techo changed >A.t &&\n> \tgit add A.t &&\n> \tgit update-index --skip-worktree A.t &&\n> \trm A.t &&\n> \tgit stash &&\n> \n> \t: this should not mark A.t as deleted &&\n> \tgit stash apply &&\n> \ttest -n \"$(git ls-files A.t)\"\n> '\n> \n> Now, the problem would have occurred in the very last command: `A.t`\n> would have been missing from the index.\n> \n> In order to debug this, you would have had to put in \"breakpoints\"\n> (inserting `false &&` at strategic places), or prefix commands with\n> `debug` to start them in GDB, then re-run the test case.\n> \n> Lather, rinse and repeat, until you figured out that `git stash` was\n> failing to record `A.t` properly.\n> \n> Then dig into that, recreating the same worktree situation every time\n> you run `git stash`, until you find out that there is a call to `git\n> update-index -z --add --remove --stdin` that removes that file.\n> \n> Further investigation would show you that this command pretty much does\n> what it is told to, because it is fed `A.t` in its `stdin`.\n> \n> The next step would probably be to go back even one more step and see\n> that `diff_index()` reported this file as modified between the index and\n> `HEAD`.\n> \n> Slowly, you would form an understanding of what is going wrong, and you\n> would have to go back and forth between blaming `diff_index()` and\n> `update-index`, or the options `git stash` passes to them.\n> \n> You would have to recreate the worktree many, many times, in order to\n> dig in deep, and of course you would need to understand the intention\n> not only of the regression test, but also of the code it calls.\n> \n> In this instance, it is merely tedious, but possible. I know, because I\n> did it. For flaky tests, not so much.\n> \n> *That* is the scenario I tried to get at.\n> \n> Writing the test cases in `make` would not help that. Not one bit. It\n> would actually make things a lot worse.\n\nI agree debugging can be tedious, but I'm not sure how to get\naround that besides making test cases smaller, more isolated\n(which can also lead to missing coverage), and better\ndocumented.\n\nHow would your ideal test framework approach it?\n\nWas using git tracing in the above scenarios not enough?\n\nI've never cared much about project-specific tracers; but\ninstead prefer to rely on project-agnostic (but OS-specific)\nfacilities such as truss/strace since I can reuse that knowledge\nacross languages/projects.\n\n> > (that said, I don't know if you use make or something else to build\n> > gfw)\n> \n> We are talking about the test suite, yes?\n> \n> Not about building Git? Because it does not matter whether we use `make`\n> or not (we do by default, although we also have an option to build in\n> Visual Studio and/or via MSBuild).\n> \n> On Windows, just like on Linux, we use a Unix shell interpreter (Bash).\n> Sure, to run the entire test suite, we use `make` (sometimes in\n> conjunction with `prove`), but the tests themselves are written in Unix\n> shell script, so I have a hard time imagining a different method to run\n> them -- whether on Windows or not -- than to use a Unix shell\n> interpreter such as Bash.\n> \n> > > In many ways, our current test suite seems to test Git's\n> > > functionality as much as (core) contributors' abilities to implement\n> > > test cases in Unix shell script, _correctly_, and maybe also\n> > > contributors' patience.  You could say that it tests for the wrong\n> > > thing at least half of the time, by design.\n> >\n> > Basic (not advanced) sh is already a prerequisite for using git.\n> \n> Well, if you are happy with that prerequisite to be set in stone, I am\n> not. Why should any Git user *need* to know sh?\n\nFor chaining commands together, why not?\n\nI expect anybody who's learned C to also be able to figure out\nsome basic scripting to tie a series of commands together in the\nsame way you can poke around in Perl without being fluent.\n\nAnd there's plenty of overlap between C and sh when it comes to\nwith control flow (if/else/while/&&/for).\n\n> > Writing correct code and tests in ANY language is still a\n> > challenge for me; but I'm least convinced a low-level language\n> > such as C is the right language for writing integration tests in.\n> \n> I would be delighted to go for a proper, well-maintained test framework\n> such as Jest. But of course, that would not be accepted in this project.\n> So I won't even think about it.\n> \n> > C is fine for unit tests, and maybe we can use more unit tests and\n> > less integration tests.\n> \n> We don't have integration tests.\n> \n> Unless your concept of what constitutes an \"integration test\" is very\n> different from mine. For me, an integration test would set up an\n> environment that involves multiple points of failure, e.g. setting up an\n> SSH server on Linux and accessing that via Git from Ubuntu. Or setting\n> up a web server with a self-signed certificate, import the public key\n> into the Windows Certificate Store and then accessing the server both\n> using OpenSSL and Secure Channel (the native Windows way to communicate\n> via TLS).\n\nI think of integration tests as anything which covers\nreal-world use cases (e.g. running documented commands as\na normal user would).\n\nUnit tests would be t/helper/*.c in my mind...\n\nBut maybe my terminology is all wrong *shrug*\n\n> There is nothing even close to that in Git's test suite.\n> \n> > > It might look like a somewhat less important project, but given that we\n> > > exercise almost 150,000 test cases with every CI build, I think it does\n> > > make sense to grind our axe for a while, so to say.\n> >\n> > Something that would benefit both users and regular contributors\n> > is the use and adoption of more batch and eval-friendly interfaces.\n> > e.g. fast-import/export, cat-file --batch, for-each-ref --perl...\n> \n> Given how hard it is to deduce the intention behind such invocations, I\n> am rather doubtful that this would improve our test suite.\n> \n> > I haven't used hg since 2005, but I know \"hg server\" exists\n> > nowadays to get rid of a lot of startup overhead in Mercurial,\n> > and maybe git could steal that idea, too...\n> \n> I have no idea what `hg server` does. Care to enlighten me?\n\nSorry, \"hg serve\" (no 'r').  It starts a long-lived server to\nexpose the API over a pipe or socket:\nhttps://www.mercurial-scm.org/wiki/CommandServer\nClosest we have is fast-import and \"cat-file --batch*\"\n"},{"id":"383756","messageId":"20191009172551.GI29845@szeder.dev","threadId":"51754","inReplyTo":"20190927221857.GB31237@sigill.intra.peff.net","subject":"Re: Git in Outreachy December 2019?","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2019-10-09T17:25:51Z","receivedAt":"2019-10-09T17:25:59Z","isPatch":false,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Fri, Sep 27, 2019 at 06:18:58PM -0400, Jeff King wrote:\n> On Thu, Sep 26, 2019 at 11:44:48PM +0200, SZEDER Gábor wrote:\n> \n> > All that was over a year and a half ago, and these limitations weren't\n> > a maintenance burden at all so far, and nobody needed that escape\n> > hatch.\n\n> I'm actually surprised we haven't run into it more. We have some custom\n> test scripts in our fork of Git at GitHub. We usually just use\n> TEST_SHELL_PATH=bash, but curious, I tried running with dash and \"-x\",\n> and three of them failed.\n\nI try to avoid using TEST_SHELL_PATH at all costs, it is far too keen\nto not do what I naively expect.\n\n> Probably they'd be easy enough to fix (and they're out of tree anyway),\n> so I'm not really arguing against the escape hatch exactly. Mostly I'm\n> just surprised that if I introduced 3 cases (out of probably a dozen\n> scripts), I'm surprised that more contributors aren't accidentally doing\n> so upstream.\n\nI see it a bit differently.  Over a decade we gathered about\ntwenty-something such tests cases: that's about two cases per year.\nYou added three such cases in about a year and a half: that's two\ncases per year.  The numbers add up perfectly, you singlehandedly took\ncare of everything ;)\n\n\nAnyway, I did some more digging, and, unfortunately, it turned out\nthat Dscho is somewhat right.  While the situation is not as bad as he\nmade it look like (\"We need Bash!\"), it's not as good as I thought it\nis (\"But it Just Works!!\") either.\n\n  - Some shells do include file descriptor redirections in the trace\n    output, and it varies between implementations to which fd the\n    trace of the redirection goes.\n    \n      - 'ksh/ksh93' and NetBSD's /bin/sh send the trace of\n        redirections to the \"wrong\" fd, in the sense that e.g. the\n        trace of commands invoked in 'test_must_fail' goes to the\n        function's standard error, and checking its stderr with\n        'test_cmp' would then fail.\n \n        (But 'ksh/ksh93' doesn't really matter, because they don't\n        support the 'local' keyword, so they fail a bunch of tests\n        even without '-x' anyway.)\n\n        I don't think we can do anything about these shells.\n\n      - 'mksh/lksh' send the trace of redirections to the \"right\" fd,\n        so they won't pollute the stderr of test helper functions.\n        And indeed the test suite passes when run with 'mksh' (well,\n        at least the subset of the test suite that I usually run).\n\n  - We do call 'test_have_prereq' from within test cases as well,\n    notably from the 'test_i18ngrep', 'test_i18ncmp' and\n    'test_ln_s_add' helper functions.  In those cases all trace output\n    from 'test_have_prereq' is included in the test case's trace\n    output, which means that during the first invocation:\n\n      - there is lots of distracting and confusing trace output, as\n        the script evaluating the prereq is passed around to a bunch\n        of functions.\n\n      - after running the script evaluating the prereq 'test_eval_'\n        does indeed turn off tracing, so there will be no trace from\n        the remainder of that test case (except with 'mksh': while it\n        does run 'set +x' in 'test_eval_', that somehow doesn't turn\n        off tracing...  I have no idea whether that's a bug or a\n        feature).\n\n    As far as 'test_i18ngrep' is concerned, which accounts for the\n    majority of 'test_have_prereq' invocations within test cases, I\n    don't understand why it uses 'test_have_prereq' in the first place\n    instead of checking the GIT_TEST_GETTEXT_POISON environment\n    variable; and 6cdccfce1e (i18n: make GETTEXT_POISON a runtime\n    option, 2018-11-08) doesn't give me any insight.\n\n    I recall that some months ago we discussed the idea of how to\n    disable trace output from within test helper functions; that would\n    help with this 'test_have_prereq' issue as well, at least in case\n    of the more \"common\" shells.\n\n"},{"id":"383879","messageId":"20191011063443.GB25741@sigill.intra.peff.net","threadId":"51754","inReplyTo":"20191009172551.GI29845@szeder.dev","subject":"Re: Git in Outreachy December 2019?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-10-11T06:34:43Z","receivedAt":"2019-10-11T06:34:46Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Oct 09, 2019 at 07:25:51PM +0200, SZEDER Gábor wrote:\n\n> > Probably they'd be easy enough to fix (and they're out of tree anyway),\n> > so I'm not really arguing against the escape hatch exactly. Mostly I'm\n> > just surprised that if I introduced 3 cases (out of probably a dozen\n> > scripts), I'm surprised that more contributors aren't accidentally doing\n> > so upstream.\n> \n> I see it a bit differently.  Over a decade we gathered about\n> twenty-something such tests cases: that's about two cases per year.\n> You added three such cases in about a year and a half: that's two\n> cases per year.  The numbers add up perfectly, you singlehandedly took\n> care of everything ;)\n\nThose cases are actually much older than that. I just didn't bother to\nclean them up until recently. So my rate is even lower :)\n\n>   - Some shells do include file descriptor redirections in the trace\n>     output, and it varies between implementations to which fd the\n>     trace of the redirection goes.\n>     \n>       - 'ksh/ksh93' and NetBSD's /bin/sh send the trace of\n>         redirections to the \"wrong\" fd, in the sense that e.g. the\n>         trace of commands invoked in 'test_must_fail' goes to the\n>         function's standard error, and checking its stderr with\n>         'test_cmp' would then fail.\n>  \n>         (But 'ksh/ksh93' doesn't really matter, because they don't\n>         support the 'local' keyword, so they fail a bunch of tests\n>         even without '-x' anyway.)\n> \n>         I don't think we can do anything about these shells.\n\nYeah, unless somebody is complaining, I don't know that it's worth\nworrying about too much. The test suite is certainly useful without\nbeing able to use \"-x\" on every single test run (you can still run it\nwithout \"-x\" obviously, or selectively use \"-x\" to debug a single test\nor script). So if it is only unreliable on a few tests on a few obscure\nshells, we can probably live with it until somebody demonstrates a\nreal-world problem (e.g., that they're running automated CI on an\nobscure platform that is stuck with an old shell, and really want \"-x\n--verbose-log\" to get more verbose failures).\n\n>   - We do call 'test_have_prereq' from within test cases as well,\n>     notably from the 'test_i18ngrep', 'test_i18ncmp' and\n>     'test_ln_s_add' helper functions.  In those cases all trace output\n>     from 'test_have_prereq' is included in the test case's trace\n>     output, which means that during the first invocation:\n> \n>       - there is lots of distracting and confusing trace output, as\n>         the script evaluating the prereq is passed around to a bunch\n>         of functions.\n\nYeah, I think this is probably an issue even with bash.\n\n>     As far as 'test_i18ngrep' is concerned, which accounts for the\n>     majority of 'test_have_prereq' invocations within test cases, I\n>     don't understand why it uses 'test_have_prereq' in the first place\n>     instead of checking the GIT_TEST_GETTEXT_POISON environment\n>     variable; and 6cdccfce1e (i18n: make GETTEXT_POISON a runtime\n>     option, 2018-11-08) doesn't give me any insight.\n\nI think it's just that checking the environment variable is non-trivial:\nwe invoke env--helper to handle bool interpretation. So we'd prefer to\ncache the result (and not to run it at all if a test script doesn't use\ni18ngrep, though it's perhaps ubiquitous enough that we should just run\nit up front for every script).\n\n>     I recall that some months ago we discussed the idea of how to\n>     disable trace output from within test helper functions; that would\n>     help with this 'test_have_prereq' issue as well, at least in case\n>     of the more \"common\" shells.\n\nThat might be worth doing, though IIRC it got kind of hairy. :)\n\n-Peff\n"},{"id":"384631","messageId":"20191022211639.GF9323@google.com","threadId":"51754","inReplyTo":"20190921014701.GA191795@google.com","subject":"Re: Git in Outreachy December 2019?","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2019-10-22T21:16:39Z","receivedAt":"2019-10-22T21:16:48Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Fri, Sep 20, 2019 at 06:47:01PM -0700, Emily Shaffer wrote:\n> On Fri, Sep 20, 2019 at 10:04:48AM -0700, Jonathan Tan wrote:\n> > > Prospective mentors need to sign up on that site, and should propose a\n> > > project they'd be willing to mentor.\n> > \n> > [snip]\n> > \n> > > I'm happy to discuss possible projects if anybody has an idea but isn't\n> > > sure how to develop it into a proposal.\n> > \n> > I'm new to Outreachy and programs like this, so does anyone have an\n> > opinion on my draft proposal below? It does not have any immediate\n> > user-facing benefit, but it does have a definite end point.\n> \n> I'd appreciate similar opinion if anybody has it - and I'd also really\n> feel more comfortable with a co-mentor.\n\nI know early on in this thread about Outreachy projects some folks\nexpressed interest in comentoring. Is anybody still interested in doing\nso?\n\nFor context, I've been in contact with 3 applicants who have either sent\ntheir first patch already or are getting ready to (and have needed some\ninvolved discussion or help) plus another few applicants who have\ninquired and may or may not send patches in the future. I've also\nreceived quite a few mails outside of my timezone working hours (I\nusually am awake/working 18:00GMT-02:00GMT), which I feel badly about\nnot being able to respond to in a timely fashion. If anybody wants to\ncomentor I would be so excited to have the help :)\n\n - Emily\n"}]}