git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH 0/2] Fix misaligned output of git repo structure

From
Junio C Hamano <gitster@pobox.com>
Date
Nov 14, 2025, 19:22 UTC
Message-ID
<xmqqecq0ifld.fsf@gitster.g>
In-Reply-To
<CANYiYbGyGKy=S6a3NJFyrv-bOZos+BXdR=nPXDT3W_dGxeiNPA@mail.gmail.com>
Jiang Xin <worldhello.net@gmail.com> writes:
Show 14 quoted lines
>> Is `Co-developed-by` supposed to have a different meaning than the more
>> common `Co-authored-by`?
>
> This is a very good question.
>
> **Background**
>
> At Alibaba Cloud, our development team uses a variety of AI coding tools,
> including Cursor, Claude Code, Gemini-CLI, Lingma, and Qoder, etc. To
> measure adoption—specifically, how many developers are using AI coding
> tools and how much code is AI-generated—we needed a unified tracking
> mechanism compatible with all these tools. I chose to implement a git
> commit-msg hook that automatically detects the AI coding tool responsible
> for a commit based on environment variables at commit time.

In other words, addition of this is solely to help corporations like Alibaba to measure which AI tools are used (and what correlation there are between success rate of the patches and the tools that generated them, etc..

What is in it for us? What benefit are we getting in exchange for tolerating these additional trailer lines in our log messages?

A few random thoughts about generated contents:
 * Disclosing the tools that were used during the development of a
   patch is a good practice in principle, but this is not limited to
   use of AI tools.  We have fixes for issues found with existing
   Coccinelle checks, sanitizers, static checkers, and it is the
   usual practice for the patches that fix them to disclose how the
   author discovered the issue.  When making mechanical replacement
   changes en masse, it is the usual practice for the patches to
   describe what scripts were used to make the changes in them.  But
   we do not dedicate a trailer line for such a disclosure, and
   there is no reason why AI tools has to be treated specially here.
   Instead of "Co-developed-by" that only tells what tool was used,
   why not disclose what prompts (again, somehow AI tools are
   treated specially here, too---we call the input to these tools
   "scripts" when the changes were made with sed or perl or
   coccinelle) were used?
 * Whether some or all contents in a submitted patch were generated
   by tools, it does not change the obligation of the person who
   submits the patch.  They need to make sure that the changes are
   reviewable, its goal and implementation are described in the
   proposed log message appropriately, the updated code does what
   the proposed log message claims to do.  They need to make sure
   that they have the right to contribute the patch under DCO, and
   sign off their patch accordingly.
 * What is made more difficult for a submitter with AI tools is that
   it is often not obvious to the human developer how much of the
   tools' generated output is parroting what the tools saw during
   their training session, and what the licensing terms of these
   training materials are.  Even if a hypothetical AI tool were
   trained only with BSD licensed material, the output from such a
   tool is likely to hold you under certain obligations like
   including the original copyright notice, but without the tool
   disclosing to you the human developer, you do not even know whose
   copyright notice to include.
 * Worse yet, the above difficulty is only for the submitter of such
   a patch, not the project that, trusting what the sign-off of the
   submitter certifies, reviews and accepts such a patch.  It does
   not make any difference if the original submitter copied and
   pasted proprietary code of their employer in the patch, or
   included code that AI tools "borrowed" from elsewhere without
   following proper procedure to honor the licensing terms.  In
   either case, the project may have accepted what was stolen
   without knowing, and it is very likely that the submitter but not
   the project is primarily held liable.  In a sense, the project
   would be better off if the patch does not say it was generated
   with AI tools---if the project does not know, it cannot possibly
   held liable for it, even though the project will have to waste
   engineering resources to rewrite or remove the remnant from such
   a faulty contribution.
Previous: Jiang XinNext: Jiang Xin
Message 15 of 22 in “Fix misaligned output of git repo structure”
  1. 0/2 Fix misaligned output of git repo structureJiang Xin, Nov 14, 2025
  2. 1/2 t/unit-tests: add UTF-8 width tests for CJK charsJiang Xin, Nov 14, 2025
  3. Junio C HamanoNov 14, 2025
  4. Jiang XinNov 15, 2025
  5. 2/2 builtin/repo: fix table alignment for UTF-8 charactersJiang Xin, Nov 14, 2025
  6. Justin ToblerNov 14, 2025
  7. Jiang XinNov 15, 2025
  8. Junio C HamanoNov 14, 2025
  9. Jiang XinNov 15, 2025
  10. Junio C HamanoNov 15, 2025
  11. Jiang XinNov 16, 2025
  12. Junio C HamanoNov 16, 2025
  13. Kristoffer HaugsbakkNov 14, 2025
  14. Jiang XinNov 14, 2025
  15. Junio C HamanoNov 14, 2025
  16. Jiang XinNov 15, 2025
  17. Junio C HamanoNov 14, 2025
  18. 0/2 Fix misaligned output of git repo structureJiang Xin, Nov 15, 2025
  19. 1/2 t/unit-tests: add UTF-8 width tests for CJK charsJiang Xin, Nov 15, 2025
  20. 2/2 builtin/repo: fix table alignment for UTF-8 charactersJiang Xin, Nov 15, 2025
  21. Phillip WoodNov 15, 2025
  22. Junio C HamanoNov 15, 2025

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.