From: Collin Funk Date: Thu, 09 Oct 2025 01:13:43 GMT Subject: Re: [PATCH v2] SubmittingPatches: add section about AI Message-ID: <87cy6wdghk.fsf@gmail.com> In-Reply-To: Christian Couder writes: > On Sat, Oct 4, 2025 at 12:20 AM brian m. carlson > wrote: >> >> On 2025-10-03 at 20:48:40, Elijah Newren wrote: >> > Would this mean that you wanted to ban contributions like d12166d3c8bb >> > (Merge branch 'en/docfixes', 2023-10-23), available on the list over >> > at https://lore.kernel.org/git/pull.1595.git.1696747527.gitgitgadget@gmail.com/ >> > ? We don't need to go theoretical, I've already contributed such a >> > patch series before -- 2 years ago -- and it was merged. Granted, >> > that was entirely documentation, and I called out the usage of AI in >> > the cover letter, and I manually checked every change (discarding many >> > of them) and split it into commits on my own, could easily explain any >> > change and why it was good, etc. And I was upfront about all of it. >> >> I think the main problem here is that we don't know the copyright >> status of LLM outputs. > > It's very unlikely that whatever is decided about the copyright status > of LLM outputs will fundamentally change copyright law. So for example > small changes, or changes where a human has been involved a lot, or > changes that are very specific, and so on, are very likely acceptable. The issue is lack of law, from my understanding. There has been zero political will in the US for copyright legislation with respect to the output of AI. Therefore, we are left with case law that is still ongoing, that is, no precedent. >> I remember the SCO situation with Linux and how it really created a lot >> of uncertainty with Linux because SCO created FUD around Linux licensing >> and how that led to the DCO being created. I am aware of the fact that >> many open source contributors are very unhappy that their code has been >> used to train LLMs without retaining credits and copyright notices or >> honouring the license terms[2]. > > I don't think it's very relevant for your position on this. On the > contrary, if LLMs have been trained mostly with open source code, then > if they produce copyrighted output, that output is more likely to be > compatible with the GPL. It has even been suggested (and discussed in > this thread) that some AIs should be trained only with open source > material (for example MIT licensed material?) so that we could stop > worrying about including it. If that happens, there would be no reason > to outright ban AI generated content, right? Not all open source code is compatible with other open source code. If you use the output of a model trained on GPLv3+ code in a GPLv2-only project, then the creator of the GPLv3+ code could claim that you violated the license since they are not compatible. Whether they would win in court or not, I have no clue, but it is probably best to avoid that situation. Collin