Re: [RFC PATCH 1/1] SubmittingPatches: allow responsible AI assistance
- From
Luca Milanesio <luca.milanesio@gmail.com>
- Date
- Oct 8, 2026, 05:29 UTC
- Message-ID
- <6CCE2DB2-E2E2-48A0-B443-95B0AABFE83B@gmail.com>
- In-Reply-To
- <CAP2yMa+kgphMe-cpcZSvPSqwm-npUDVp=HaNRW+MPmPzZ_aOXw@mail.gmail.com>
Show 24 quoted lines
> On 8 Oct 2026, at 05:49, Scott Chacon <schacon@gmail.com> wrote: > > On Wed, Oct 7, 2026 at 11:42 PM brian m. carlson > <sandals@crustytoothpaste.net> wrote: >> >> On 2026-10-07 at 14:29:54, Scott Chacon wrote: >> I don't think I'm in favour of this policy. All the major models have >> been trained on a large variety of code from a large variety of sources, >> including sources such as news reports or personal websites that do not >> allow copying, modification, or distribution. Given that LLMs are known >> to reproduce portions of their training set or craft code or text which >> is very similar to items in the training set, how can anyone honestly >> assert the DCO without knowing all of the sources that were used to >> create it? > > I agree that generated output can reproduce material we don't have > permission to distribute (though I think this is incredibly rare for > anything complex). What I question is whether that possibility means > knowing every source in the training set is necessary to make any DCO > certification. > > Human contributors have also read code under many different licenses > (and news articles and blogs) . We don't ask them to account for > everything they've ever read before signing off on a patch.
There is a difference between “learning from existing code” which builds up experience and professional capability and “copy & pasting” code from different sources and putting it together.
The first (learning from existing code) is a product of whoever writes the code, based on its mental model and experience developed, the second is a simple violation of the contribution guidelines.
Where AI stays? LLMs are a simple processing of a large amount of code for statistically detecting which part of the copy need to be copy&pasted and blended together. Even though they had a “learning” path in terms of developing an understanding and building a mental model (they don’t, at least at the moment), *THEY* would be the author of the code, not the person that just “pressed the button” on the AI model.
Should we use fully-generated AI contributions? When “fully-generated” means code that has been totally written and reviewed by agents autonomously, then I believe the answer should be a sound no, even just for the lack of a legal framework around it on who is taking the responsibility of what has been contributed.
Show 5 quoted lines
> We do > expect them to have the right to submit the actual contribution and to > respect the licenses of material they incorporate. Again, Red Hat, the > Linux Foundation and the SFC all now state that LLM generated code is > acceptable and compatible with DCO requirements.
Do you have a link to their exact statement? Do they really allow *fully AI generated* code contributions where the human has no understanding of the lines of code generated?
If that was the case, how can you manage a review of code that wasn’t fully guided and understood by the author? Are you foreseeing an agent answering the mailing list to the comments of the reviewers?
Even though we would allow *fully* AI-generated code, who is really the contribution from? The “assumed human author”? The LLM? The data that LLM is trained on? The company that developed the LLM? Nobody?
Also, imagine that the code generated *did contain* malicious backdoors, who is liable for it?
There are currently lawsuit in progress against companies that have trained LLMs on allegedly copyrighted material, we don’t know yet where they’ll end up to.
Show 14 quoted lines
> >> I'm a distributor of Git and I don't want to be sued or arrested because >> I end up distributing code that I don't have the right to distribute. >> Large companies may have lawyers and lots of money to fight those >> claims, but I do not (nor does the Git project) and I don't want to >> spend my resources fighting allegations of copyright infringement or >> have my reputation besmirched for that reason. Just because other >> projects think it's okay to do legally and ethically questionable things >> doesn't mean we should as well. > > Again, a lot of my argumentation here was directly taken from the > SFC's recommendations [1], which Git is a member project of. > > https://sfconservancy.org/llm-gen-ai/llm-backed-generative-ai-recommendations.html
I am fully aligned and in agreement with SFC’s recommendations, which is very much aligned with my thinking. If you are the author of the code, even if was created with the assistance of of LLMs, then you are fully responsible for each and every line of it and you must understand it and approve it yourself.
I personally use LLMs *a lot* for giving me ideas and reviewing my code or learning about something new that I never explored before. The end goal of LLMs for me is to improve my capabilities as a developer, not to generate contributions on my behalf.
My success criteria after having used LLMs is: I am a better developer now? Am I more capable of writing better software and having a deeper understanding of existing code?
P.S. My comment was 100% human-generated, including spelling mistakes and awkward expressions (I’m not mother tongue).
Luca.
Show 34 quoted lines
> >> I'll add that if Git were to include a portion of my MIT- or >> BSD-licensed code without including a copyright or permission notice >> because it was laundered through an LLM, I would absolutely file a >> copyright complaint, and rightfully so. >> >> I refer you to policies from other major open source projects that cover >> this exact provenance issue: >> >> * Gentoo: https://wiki.gentoo.org/wiki/Project:Council/AI_policy >> * NetBSD: https://www.netbsd.org/developers/commit-guidelines.html > > I'm not saying that other projects haven't taken positions as > conservative as Git's current policy. I'm saying that much larger > projects with much larger legal surface area such as Linux have > adopted more progressive ones. Linus is fine with it on a project with > the same DCO, the same license, and honestly, a lot more legal > scrutiny. > >> I also will point out the notes from the Contributor Summit where we >> discussed this issue in some depth and proposed an approach for further >> discussion. > > It was unclear from the notes what a "vote" meant exactly. It looked > like Taylor was roughly tasked with providing a Linux-style position, > so I thought I would help out by providing a first version of that > position. > > This was written to follow the current SFC guidelines. I would > encourage Junio to have them read it, but I tried my best to follow > it's guidelines and advice. > > Scott >