Re: [RFC PATCH 1/1] SubmittingPatches: allow responsible AI assistance
- From
Scott Chacon <schacon@gmail.com>
- Date
- Oct 8, 2026, 04:49 UTC
- Message-ID
- <CAP2yMa+kgphMe-cpcZSvPSqwm-npUDVp=HaNRW+MPmPzZ_aOXw@mail.gmail.com>
- In-Reply-To
- <asa8ymCv4hoRJcZM@fruit.crustytoothpaste.net>
On Wed, Oct 7, 2026 at 11:42 PM brian m. carlson <sandals@crustytoothpaste.net> wrote:
Show 10 quoted lines
> > On 2026-10-07 at 14:29:54, Scott Chacon wrote: > I don't think I'm in favour of this policy. All the major models have > been trained on a large variety of code from a large variety of sources, > including sources such as news reports or personal websites that do not > allow copying, modification, or distribution. Given that LLMs are known > to reproduce portions of their training set or craft code or text which > is very similar to items in the training set, how can anyone honestly > assert the DCO without knowing all of the sources that were used to > create it?
I agree that generated output can reproduce material we don't have permission to distribute (though I think this is incredibly rare for anything complex). What I question is whether that possibility means knowing every source in the training set is necessary to make any DCO certification.
Human contributors have also read code under many different licenses (and news articles and blogs) . We don't ask them to account for everything they've ever read before signing off on a patch. We do expect them to have the right to submit the actual contribution and to respect the licenses of material they incorporate. Again, Red Hat, the Linux Foundation and the SFC all now state that LLM generated code is acceptable and compatible with DCO requirements.
Show 8 quoted lines
> I'm a distributor of Git and I don't want to be sued or arrested because > I end up distributing code that I don't have the right to distribute. > Large companies may have lawyers and lots of money to fight those > claims, but I do not (nor does the Git project) and I don't want to > spend my resources fighting allegations of copyright infringement or > have my reputation besmirched for that reason. Just because other > projects think it's okay to do legally and ethically questionable things > doesn't mean we should as well.
Again, a lot of my argumentation here was directly taken from the SFC's recommendations [1], which Git is a member project of.
https://sfconservancy.org/llm-gen-ai/llm-backed-generative-ai-recommendations.html
Show 10 quoted lines
> I'll add that if Git were to include a portion of my MIT- or > BSD-licensed code without including a copyright or permission notice > because it was laundered through an LLM, I would absolutely file a > copyright complaint, and rightfully so. > > I refer you to policies from other major open source projects that cover > this exact provenance issue: > > * Gentoo: https://wiki.gentoo.org/wiki/Project:Council/AI_policy > * NetBSD: https://www.netbsd.org/developers/commit-guidelines.html
I'm not saying that other projects haven't taken positions as conservative as Git's current policy. I'm saying that much larger projects with much larger legal surface area such as Linux have adopted more progressive ones. Linus is fine with it on a project with the same DCO, the same license, and honestly, a lot more legal scrutiny.
> I also will point out the notes from the Contributor Summit where we > discussed this issue in some depth and proposed an approach for further > discussion.
It was unclear from the notes what a "vote" meant exactly. It looked like Taylor was roughly tasked with providing a Linux-style position, so I thought I would help out by providing a first version of that position.
This was written to follow the current SFC guidelines. I would encourage Junio to have them read it, but I tried my best to follow it's guidelines and advice.
Scott