Re: [RFC PATCH 1/1] SubmittingPatches: allow responsible AI assistance
- From
Scott Chacon <schacon@gmail.com>
- Date
- Oct 8, 2026, 13:53 UTC
- Message-ID
- <CAP2yMa+o7zv=8bzHUa8FRkTaU46BQ3w_ZAr6zQ=rFC_a4cg2yA@mail.gmail.com>
- In-Reply-To
- <xmqqcxtl3zda.fsf@gitster.g>
On Thu, Oct 8, 2026 at 12:44 AM Junio C Hamano <gitster@pobox.com> wrote:
Show 12 quoted lines
> Scott Chacon <scott@gitbutler.net> writes: > > As an example, an OpenAI model was used to help me research, compare and > > craft the appropriate legal language for this policy change to help us > > match the modern, legally reviewed approaches now taken by peer GPL > > projects such as the Linux kernel [1]. > > > > [1] https://docs.kernel.org/process/coding-assistants.html > > That makes it sound as if this is just as legally sound as what the > kernel project uses. However, the only assurance we get (unless you > are willing to act as our lawyer, and I do not know if you are one) > is that an OpenAI model produced plausible-sounding utterances.
Well, honestly, most lawyers I know only barely produce plausible sounding utterances.
I read the change and edited it where I thought clarification was needed. But this is different from, say, a blog post or emails, which I personally never write via LLM because I don't like the voice. SubmittingPatches is supposed to be dry and factual. Legal guidance is supposed to be neutral. I've found LLMs quite good at producing correct legal documents. Again, unjokingly this time, better than most human lawyers (and I deal with a lot of them).
It's not unlike a code-based LLM contribution. I had it generated with specific guidance and context of what I wanted the change to be, because it's faster, and I spent my time reviewing and editing the result to get the text I was looking for.
I would love to have you pass this by the SFC, because most of the guidance I gave my agent was _their_ guidelines.
Show 14 quoted lines
> It looks, at least to me, that there is not much that can be > meaningfully enforced by reviewers and followed by contributors in > the above text. It seems to be little more than "the world would be > a wonderful place if everybody behaved this way." > > A violation of "concise and relevant" seems to be the recent trend > of much AI-generated slop, so it may be a good suggestion to give > today. But would we need to update it once the trend of text > generated by AI tools becomes "concise and relevant" nonsense that > merely sounds plausible? What if an "AI-assisted" contributor lacks > common sense to tell between plausible-sounding nonsense and a > well-written description? What if reviewers get too many such > "contributions" and cannot allocate enough review bandwidth to sift > good contributions from plausible-sounding nonsense?
This is a fair point, but you'll get unreviewed crap either way. I'm sure you already are. However, I don't think people who submit complete bullshit are reading the SubmittingPatches file in the first place, so I'm not sure that opening this wording up a little is going to make much of a difference here.
My recent patch series converting the sha1dc is a possible example. I don't understand all of the code it wrote. I read through it, but there are some crazy tables and complex math in there. The first pass did a weird Rust to C machine translation rather than reimplement it in more idiomatic C, so I had it rewrite that - so there was some approach guidance, but again, I wasn't hand crafting the code. I did, however, spend a lot of time and resources testing and benchmarking it on multiple architectures so that I was reasonably confident that it was fast and correct.
But I hesitated to submit it at all because I knew the policy. I only sent it so that if someone at GitHub or OpenAI or whatever wanted to use it in an internal fork so they could save a ton of CPU, this would be a way to get the implementation. I was aware that, although I believe the patch is quite reasonable and valuable, due to the conservative AI policies of this project, it would not seriously be considered no matter what.
My point with this change is to open the possibility for AI assisted change that is reasonable, similar to the Linux kernel's approach.
Show 16 quoted lines
> > +The <<dco,Developer's Certificate of Origin>> applies unchanged. Only a > > +human can make that certification; an AI tool cannot sign off on your > > +behalf. Consider the origin and licensing of generated material, > > +including any third-party material it reproduces, and comply with > > +applicable license and attribution requirements. A tool's assurance > > +that its output is original or compatible with our license is not a > > +substitute for checking those requirements. If you cannot certify the > > +DCO for a contribution, do not submit it. > > Again, this is a good aspiration to have, but I doubt that anyone > can practically certify that the output of an LLM is devoid of > content borrowed from problematic sources under the rule the text > above gives. Would it not be more useful to help contributors by > defining what not to do more clearly? Our current text says as much > more directly: you cannot practically certify, so do not send in > AI-generated slop, period.
There is a good section on this DCO issue in the Red Hat article on navigating legal issues around AI [1] where they state that "the DCO has never been interpreted to require that every line of a contribution must be the personal creative expression of the contributor or another human developer". I think that this suggested paragraph in my patch is a fairly clear interpretation of this stance, but I can give it another pass if there are more specifics you would like covered.
[1] https://www.redhat.com/en/blog/ai-assisted-development-and-open-source-navigating-legal-issues
Show 15 quoted lines
> > +Disclose substantial AI assistance in each affected commit with an > > +`Assisted-by:` trailer naming the tool and, when available, its model > > +or version. For example: > > + > > +.... > > + Assisted-by: ExampleTool version 1.2 > > +.... > > I thought the kernel guidelines instructed us to say only "LLM" > these days, to avoid giving free advertising. On the other hand, > they ask contributors to also list non-LLM tools, like coccinelle > and clang-tidy, that were used in their machine-assisted > contributions. I am undecided on the merit of specifying the > exact model and version, but listing non-LLM tools alongside > materials for independent reproduction looks like a good idea.
The kernel guidelines do say "LLM" only, but the SFC guidelines (part 5) says:
"Part of the contribution process should (at least) include a disclosure of what LLM-gen-AI system was used, its version (as these system change over time), and a brief description of how the system assisted the contributor. This information should be included in a machine-readable format in commit logs."
I'm happy to go back to the kernel guidelines, but since Git is an SFC project, I figured this is what they would be more comfortable with.
Show 12 quoted lines
> > +Maintainers may request more explanation, testing, or information about > > +provenance, and may decline contributions they cannot confidently > > +assess. > > The text of the kernel guidelines appears to give maintainers more > latitude (cf. https://docs.kernel.org/process/generated-content.html). > They can treat it just like any other contribution, reject it > outright, or choose any approach in between. The proposed text above > does not account for cases where reviewers simply lack the bandwidth > to even think about what explanation and proof to request, and it > makes it sound as if declining a submission in such a case an unfair > rejection.
I mean, no matter what this says, you can pretty much do whatever you want. I'm happy to change it to any amount of latitude the maintainer has that you want to convey. Or remove it entirely.
Scott