{"thread":{"id":"63245","subject":"Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","startedAt":"2025-04-02T18:48:14Z","lastAt":"2025-08-19T16:45:11Z","messageCount":118,"participants":["Martin von Zweigbergk","Remo Senekowitsch","Konstantin Ryabitsev","Patrick Steinhardt","Elijah Newren","Theodore Ts'o","Nico Williams","Kane York","Junio C Hamano","Phillip Wood","Eric Sunshine","D. Ben Knoble","Jacob Keller","Toon Claes","brian m. carlson","Kristoffer Haugsbakk","Oswald Buddenhagen","Askar Safin","Ben Knoble"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"515527","messageId":"CAESOdVAspxUJKGAA58i0tvks4ZOfoGf1Aa5gPr0FXzdcywqUUw@mail.gmail.com","threadId":"63245","inReplyTo":null,"subject":"Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Martin von Zweigbergk","fromEmail":"martinvonz@google.com","sentAt":"2025-04-02T18:48:01Z","receivedAt":"2025-04-02T18:48:14Z","isPatch":false,"sender":{"key":"martinvonz@google.com","avatar":"https://avatars.githubusercontent.com/u/891642?v=4"},"body":"Hi,\n\nThe Gerrit, GitButler, and Jujutsu projects all have a concept of\na \"change id\", and it behaves in a similar way between the three\ntools. The change id is conceptually associated with a commit.\nIt follows a commit as its rewritten (e.g. by amending and\nrebasing). The three projects currently store and format the\nchange id differently. We would like to unify that so we can\ninteroperate better. We hope the Git project is also interested\nin preserving and using this header.\n\nThere are many benefits to having a change id even if it's just\nlocal. I mentioned some in my email to this mailing list in [1].\nFor example, it enables\n`git rebase main <change ID>; git switch <change ID>` without\nrequiring the user to look up the hash of the rewritten commit.\nIf the change id also transferred between repos and preserved by\na forge (such as Gerrit), it enables the change id to be used to\nidentify a code review.\n\nHere's how the change ids are currently stored and formatted:\n\n * Gerrit currently stores change ids in a commit trailer called\n   `Change-Id`. It always starts with the letter 'I' and is\n   followed by 40 hex digits. For example:\n   `Change-Id: Ib563e78c3fedcff262255fa025441daa3202311b`.\n\n * GitButler currently stores change ids in a commit footer\n   called `gitbutler-change-id` (older versions used\n   `change-id`). It's written as 32 hex digits separated by\n   dashes as in the UUID  format. For example:\n   `gitbutler-change-id  7d0fbc63-032d-413c-8ae8-610fbeb713c0`.\n\n * Jujutsu currently stores change ids in a local storage outside\n   of the Git repo and is therefore not part of the Git commit\n   id. It is stored as 16 bytes. It is rendered to the user as\n  \"reverse hex\" using 'z' through 'k' as hex digits ('z' = 0,\n  'k' = 15). This allows even short prefixes to be distinguished\n   from commit  ids, which is a very useful property when used in\n   the CLI.\n\nAs mentioned, the three projects would like to use the same\nstorage and format. I think we have a consensus to store it in a\nGit commit header called `change-id` as a 32 reverse-hex digis.\nFor example: `change-id ywlktllmukprnxnmzzprukpuwyztylwt`.\n\nThere is a design doc [2] about the impact on Gerrit and how to\nhandle various cases where the client doesn't understand the\n`change-id` header. That also includes some discussion about\nwhether cherry-picking should preserve the change id or create a\nnew one. I think there is a lot of value in having a\nstandardized header regardless of what we decide about\ncherry-picks.\n\nSo, to be clear, this is mostly a heads up at this point; we don't\ndepend on any immediate changes from the Git project.\n\nThanks,\nMartin\n\n\n[1] https://lore.kernel.org/git/CANiSa6gwup5vXU235mG+Ybbc+P=SbwoNFEmuhg=iYu0yGvSXVA@mail.gmail.com/\n[2] https://gerrit-review.googlesource.com/c/homepage/+/464287\n"},{"id":"515531","messageId":"D8WEKR5QQD3W.23CD3CXEPONGB@buenzli.dev","threadId":"63245","inReplyTo":"CAESOdVAspxUJKGAA58i0tvks4ZOfoGf1Aa5gPr0FXzdcywqUUw@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Remo Senekowitsch","fromEmail":"remo@buenzli.dev","sentAt":"2025-04-02T19:34:18Z","receivedAt":"2025-04-02T19:34:24Z","isPatch":false,"sender":{"key":"remo@buenzli.dev","avatar":"https://gravatar.com/avatar/7df680b096206886db5a2dc983926f314985bbee662eb406cf23c65322cf98b7?d=mp&s=160"},"body":"Hi,\n\nI would like to add one benefit to the list that I think is very important.\nA change-id header would allow standalone code review tooling and Git forges\n(e.g. GitHub, GitLab, Forgejo) to reliably track different versions of a patch.\nThis would enable much improved code review experiences, not unlike what Gerrit\nusers are enjoying already.\n\nMailing list oriented projects like Git itself and the Linux kernel actually\ndo code review more similar to the Gerrit model. There is a focus on the\nindividual commits -- what will eventually end up on the master branch.\n\nA \"force-push\" to a PR/MR on a Git forge is conceptually similar to sending a\nnew patchset to a mailing list. However, some people don't like that, because\nit makes it hard for reviewers to track what they have and haven't reviewed\nalready. Their answer to this problem is to sequentially add meaningless \"fixup\ncommits\" on top of a branch, making incremental code review easier. Often the\nbranch is later squashed, to clean up the history a little bit. But that also\nloses important information in the process.\n\nMailing lists don't suffer from this as much, because mail clients don't go\nout of their way to hide old patchset versions from users. However, they also\ndon't provide any tools to help code reviewers associate old and new versions\nof patchsets. A change-id header could be essential in developing tooling\nfor mailing lists that track patchsets and even individual patches within\nthem across versions. Easily being able to view the interdiff between the\nlast-reviewed version of a patchset and its most recent version is tremendously\nuseful for any code review workflow.\n\nSo I think that almost all members of the broader Git ecosystem would benefit\nin many ways from Git supporting this header directly.\n\nRemo\n"},{"id":"515533","messageId":"20250402-adventurous-mustard-raccoon-ed93e3@meerkat","threadId":"63245","inReplyTo":"CAESOdVAspxUJKGAA58i0tvks4ZOfoGf1Aa5gPr0FXzdcywqUUw@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2025-04-02T19:45:15Z","receivedAt":"2025-04-02T19:45:17Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On Wed, Apr 02, 2025 at 11:48:01AM -0700, Martin von Zweigbergk wrote:\n> Hi,\n> \n> The Gerrit, GitButler, and Jujutsu projects all have a concept of\n> a \"change id\", and it behaves in a similar way between the three\n> tools. The change id is conceptually associated with a commit.\n> It follows a commit as its rewritten (e.g. by amending and\n> rebasing). The three projects currently store and format the\n> change id differently. We would like to unify that so we can\n> interoperate better. We hope the Git project is also interested\n> in preserving and using this header.\n\nNotably, b4 also uses change-id, but as a series identifier, not as a commit\nidentifier. It is not intended to ever make its way into git commits and is\nalways passed in the patch footer.\n\nThe format is an arbitrary unique string. :)\n\n-K\n"},{"id":"515534","messageId":"20250402-classic-hilarious-barnacle-7d0d0f@meerkat","threadId":"63245","inReplyTo":"D8WEKR5QQD3W.23CD3CXEPONGB@buenzli.dev","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2025-04-02T19:49:32Z","receivedAt":"2025-04-02T19:49:34Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On Wed, Apr 02, 2025 at 09:34:18PM +0200, Remo Senekowitsch wrote:\n> Mailing lists don't suffer from this as much, because mail clients don't go\n> out of their way to hide old patchset versions from users. However, they also\n> don't provide any tools to help code reviewers associate old and new versions\n> of patchsets. A change-id header could be essential in developing tooling\n> for mailing lists that track patchsets and even individual patches within\n> them across versions. Easily being able to view the interdiff between the\n> last-reviewed version of a patchset and its most recent version is tremendously\n> useful for any code review workflow.\n\nYes, this already exists, using change-id footers.\nE.g. you can run this inside your git checkout:\n\n\t$ b4 diff 20250331-b4-pks-collect-build-fixes-v2-0-6b06136808f3@pks.im\n\tGrabbing thread from lore.kernel.org/all/20250331-b4-pks-collect-build-fixes-v2-0-6b06136808f3@pks.im/t.mbox.gz\n\tChecking for older revisions\n\tGrabbing search results from lore.kernel.org\n\t---\n\tAnalyzing 19 messages in the thread\n\tPreparing fake-am for v1: meson: fix handling of '-Dcurl=auto'\n\t  range: 98f2b1fb3587..163b98dec916\n\tPreparing fake-am for v2: meson: fix handling of '-Dcurl=auto'\n\t  range: 1406ab7e183b..df3c15fbd1ce\n\t---\n\tDiffing v1 and v2\n\t\tRunning: git range-diff 98f2b1fb3587..163b98dec916 1406ab7e183b..df3c15fbd1ce\n\t---\n\t1:  6587b42aec = 1:  f41c06addd meson: fix handling of '-Dcurl=auto'\n\t2:  d8d124b84d = 2:  8c2301bcd5 gitweb: fix generation of \"gitweb.js\"\n\t3:  1f27a035f9 < -:  ---------- meson: require Perl when building docs\n\t4:  163b98dec9 = 3:  e1962003a5 meson: respect 'tests' build option in contrib\n\t-:  ---------- > 4:  6f07c417a8 meson: distinguish build and target host binaries\n\t-:  ---------- > 5:  df3c15fbd1 ci: use Visual Studio for win+meson job on GitHub Workflows\n\n-K\n"},{"id":"515535","messageId":"CAESOdVBBeQDtRmRSQeHomuxQubTP5ggKZWGG88n88qKYBHR=+w@mail.gmail.com","threadId":"63245","inReplyTo":"CAESOdVAspxUJKGAA58i0tvks4ZOfoGf1Aa5gPr0FXzdcywqUUw@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Martin von Zweigbergk","fromEmail":"martinvonz@google.com","sentAt":"2025-04-02T19:52:35Z","receivedAt":"2025-04-02T19:52:49Z","isPatch":false,"sender":{"key":"martinvonz@google.com","avatar":"https://avatars.githubusercontent.com/u/891642?v=4"},"body":"Sorry about the typo in the subject line. I meant \"commit *header*\".\n"},{"id":"515571","messageId":"Z-5QR57zgSsm6jNP@pks.im","threadId":"63245","inReplyTo":"CAESOdVAspxUJKGAA58i0tvks4ZOfoGf1Aa5gPr0FXzdcywqUUw@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-04-03T09:09:27Z","receivedAt":"2025-04-03T09:09:39Z","isPatch":false,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Hi Martin,\n\nOn Wed, Apr 02, 2025 at 11:48:01AM -0700, Martin von Zweigbergk wrote:\n> Hi,\n> \n> The Gerrit, GitButler, and Jujutsu projects all have a concept of\n> a \"change id\", and it behaves in a similar way between the three\n> tools. The change id is conceptually associated with a commit.\n> It follows a commit as its rewritten (e.g. by amending and\n> rebasing). The three projects currently store and format the\n> change id differently. We would like to unify that so we can\n> interoperate better. We hope the Git project is also interested\n> in preserving and using this header.\n> \n> There are many benefits to having a change id even if it's just\n> local. I mentioned some in my email to this mailing list in [1].\n> For example, it enables\n> `git rebase main <change ID>; git switch <change ID>` without\n> requiring the user to look up the hash of the rewritten commit.\n> If the change id also transferred between repos and preserved by\n> a forge (such as Gerrit), it enables the change id to be used to\n> identify a code review.\n\nAgreed, change IDs solve a couple of issues that many users face:\n\n  - You can reliably track how a patch evolves over time. This helps\n    various different tools to track identity of commits, like for\n    example forges, but also tools like git-range-diff(1).\n\n  - It becomes trivial to see whether a commit has been cherry-picked\n    into another branch. We do have git-cherry(1) to do that right now,\n    but that command is based on heuristics and fails as soon as the\n    patch itself needed to be adapted.\n\n  - Working with history rewrites becomes easier in the general case as\n    you don't have to adapt to constantly changing commit IDs.\n\nThe mere fact that different tools eventually ended up with similar\ndesigns around change IDs is a good indicator that there is a real need\nfor them out there.\n\n> Here's how the change ids are currently stored and formatted:\n> \n>  * Gerrit currently stores change ids in a commit trailer called\n>    `Change-Id`. It always starts with the letter 'I' and is\n>    followed by 40 hex digits. For example:\n>    `Change-Id: Ib563e78c3fedcff262255fa025441daa3202311b`.\n> \n>  * GitButler currently stores change ids in a commit footer\n>    called `gitbutler-change-id` (older versions used\n>    `change-id`). It's written as 32 hex digits separated by\n>    dashes as in the UUID  format. For example:\n>    `gitbutler-change-id  7d0fbc63-032d-413c-8ae8-610fbeb713c0`.\n> \n>  * Jujutsu currently stores change ids in a local storage outside\n>    of the Git repo and is therefore not part of the Git commit\n>    id. It is stored as 16 bytes. It is rendered to the user as\n>   \"reverse hex\" using 'z' through 'k' as hex digits ('z' = 0,\n>   'k' = 15). This allows even short prefixes to be distinguished\n>    from commit  ids, which is a very useful property when used in\n>    the CLI.\n> \n> As mentioned, the three projects would like to use the same\n> storage and format. I think we have a consensus to store it in a\n> Git commit header called `change-id` as a 32 reverse-hex digis.\n> For example: `change-id ywlktllmukprnxnmzzprukpuwyztylwt`.\n\nI don't mind the actual format too much at this point, so I won't\ncomment on this part.\n\n> There is a design doc [2] about the impact on Gerrit and how to\n> handle various cases where the client doesn't understand the\n> `change-id` header. That also includes some discussion about\n> whether cherry-picking should preserve the change id or create a\n> new one. I think there is a lot of value in having a\n> standardized header regardless of what we decide about\n> cherry-picks.\n> \n> So, to be clear, this is mostly a heads up at this point; we don't\n> depend on any immediate changes from the Git project.\n\nScott has already been reaching out to me before your mail, and I also\nmentioned to him that I have been thinking about the problem of change\nIDs for quite a while already. This has mostly been triggered by Jujutsu\nand how it uses change IDs, which is one of the good improvements over\nGit from my perspective.\n\nWhile there may not be a need to do anything in Git itself I would think\nthat supporting change IDs natively in Git would still be sensible.\nSure, you can emulate them via commit trailers. But I don't consider\ntrailers to be particularly great as a storage format for this metadata.\nAfter all, you will want to filter the commit graph by change ID for\nsome of the usecases, and doing that based on a loosely-defined format\nprobably isn't great.\n\nSo what would it take to get change IDs into Git? I think the most\nimportant items would be:\n\n  - Generating and writing change IDs in commands that support them.\n    This includes e.g. git-commit(1), git-commit-tree(1), git-merge(1),\n    git-merge-tree(1). This should of course be completely optional and\n    probably be disabled by default.\n\n  - Making tools that rewrite commits aware of change IDs so that they\n    know to retain change IDs. This involves e.g. git-cherry-pick(1),\n    git-rebase(1), git-replay(1).\n\n  - Extending revisions to allow specifying commits by change ID.\n\n  - Allowing us to filter commit graphs by change ID.\n\nI don't think any of these should be particularly hard to do. Sure,\naddressing and filtering commits by change IDs would be slowish at first\nbecause we have to basically read all commits, but this is something\nthat can be sped up via indices.\n\nThe biggest question is of course backwards compatibility -- can we\nintroduce a change ID into the commit metadata without breaking existing\nusers? I guess you'll already have a lot of experience with this given\nthat you essentially already inject change IDs into metadata, and tools\ngenerally handle this just fine?\n\nI'd certainly be happy to help out with an effort to introduce change\nIDs into Git if the community is amenable to such a proposal.\n\nPatrick\n\nNB: I'm also quite happy that Jujutsu brings a bit of a new contender\n    to Git into the picture. It has a lot of nice ideas, and in the best\n    case Git might be able to learn a few nice tricks from JJ. After\n    all, I think we can all benefit from some friendly competition.\n"},{"id":"515577","messageId":"D8WXTCOESY86.3RRJOR5GPUL47@buenzli.dev","threadId":"63245","inReplyTo":"Z-5QR57zgSsm6jNP@pks.im","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Remo Senekowitsch","fromEmail":"remo@buenzli.dev","sentAt":"2025-04-03T10:38:52Z","receivedAt":"2025-04-03T10:39:04Z","isPatch":false,"sender":{"key":"remo@buenzli.dev","avatar":"https://gravatar.com/avatar/7df680b096206886db5a2dc983926f314985bbee662eb406cf23c65322cf98b7?d=mp&s=160"},"body":"Hi Patrick,\n\nOn Thu Apr 3, 2025 at 11:09 AM CEST, Patrick Steinhardt wrote:\n> On Wed, Apr 02, 2025 at 11:48:01AM -0700, Martin von Zweigbergk wrote:\n>>\n>> As mentioned, the three projects would like to use the same\n>> storage and format. I think we have a consensus to store it in a\n>> Git commit header called `change-id` as a 32 reverse-hex digis.\n>> For example: `change-id ywlktllmukprnxnmzzprukpuwyztylwt`.\n>\n> I don't mind the actual format too much at this point, so I won't\n> comment on this part.\n\nGerrit and GitButler also did not mind the format, which is why they\nagreed to adopt the one of Jujutsu. There is also no technical reason\nwhy Jujutsu wouldn't be able to support a free-form id. However,\ndiscussing a standard for the ecosystem gives us the opportunity to\npick something that everybody can rely on and benefit from.\n\nSome benefits of the proposed format include:\n- known memory requirement\n- change-id as part of a URL never has to be escaped\n- it being a hash means the smallest unambiguous prefix is minimized\n\nSo, these are mostly practical considerations. If there are notable\nbenefits to free-form IDs, Jujutsu can hash that again to get an ID in\nits internal format if necessary. But it's always easier to go from a\nstrict format to a loose one later, as opposed to the other way around.\n\n> While there may not be a need to do anything in Git itself I would think\n> that supporting change IDs natively in Git would still be sensible.\n> Sure, you can emulate them via commit trailers. But I don't consider\n> trailers to be particularly great as a storage format for this metadata.\n> After all, you will want to filter the commit graph by change ID for\n> some of the usecases, and doing that based on a loosely-defined format\n> probably isn't great.\n>\n> So what would it take to get change IDs into Git? I think the most\n> important items would be:\n>\n>   - Generating and writing change IDs in commands that support them.\n>     This includes e.g. git-commit(1), git-commit-tree(1), git-merge(1),\n>     git-merge-tree(1). This should of course be completely optional and\n>     probably be disabled by default.\n>\n>   - Making tools that rewrite commits aware of change IDs so that they\n>     know to retain change IDs. This involves e.g. git-cherry-pick(1),\n>     git-rebase(1), git-replay(1).\n>\n>   - Extending revisions to allow specifying commits by change ID.\n>\n>   - Allowing us to filter commit graphs by change ID.\n\nI agree with all of that. The first two points are the ones that would\nactually allow the ecosystem to start relying on this new header as a\nstandard and develop related features while staying interoperable with\nthe rest of the ecosystem. E.g. if the header is preserved by\ngit-rebase, Git & Jujutsu users will enjoy stable change-ids when a\nbranch is rebase-merged on a forge. And if git generated the header\nitself with git-commit, Gerrit could drop its requirement for clients\nto generate a change-id footer via their commit-msg hook.\n\n> The biggest question is of course backwards compatibility -- can we\n> introduce a change ID into the commit metadata without breaking existing\n> users? I guess you'll already have a lot of experience with this given\n> that you essentially already inject change IDs into metadata, and tools\n> generally handle this just fine?\n\nJujutsu has been injecting a 'jj:trees' header into commits to track\nmore metadata around merge conflicts. There weren't any problems with\nthat, unless one uses git to rewrite these commits with e.g. git-rebase,\nin which case that header is simply lost. But commits with conflicts are\nusually not pushed to a remote anyway, so the risk there was minimal.\nScott Chacon with GitButler has more experience in this regard, since\nthey actually push commits with a change-id in its header to remotes.\nHe told the Jujutsu community that they didn't encounter any problems,\nno misbehaving tools that are fussy about unknown headers. The only\nproblem is unknown commit headers being dropped by Git itself, depending\non how it is invoked by the remote. (GitHub seems to preserve the header\nduring a rebase-merge, because they use git-replay. GitLab and Forgejo\ndrop the header.) With these insights from Scott, Jujutsu is moving\nforward to put the change-id in the commit header.\n\n> NB: I'm also quite happy that Jujutsu brings a bit of a new contender\n>     to Git into the picture. It has a lot of nice ideas, and in the best\n>     case Git might be able to learn a few nice tricks from JJ. After\n>     all, I think we can all benefit from some friendly competition.\n\nOne of the best features of Jujutsu is that it plays nice with Git.\nMost users work in \"colocated\" repos that have both a .git and a .jj\ndirectory. Git commands continue to work as usual (mostly). If Git\nadopted the change-id header, interleaving of Git and Jujutsu commands\nwould work even better.\n\nSo, +1 from me for friendly competition / collaboration. :-)\n\nRemo\n"},{"id":"515579","messageId":"Z-5rpWKAVPmz32jC@pks.im","threadId":"63245","inReplyTo":"D8WXTCOESY86.3RRJOR5GPUL47@buenzli.dev","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-04-03T11:06:13Z","receivedAt":"2025-04-03T11:06:18Z","isPatch":false,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Thu, Apr 03, 2025 at 12:38:52PM +0200, Remo Senekowitsch wrote:\n> On Thu Apr 3, 2025 at 11:09 AM CEST, Patrick Steinhardt wrote:\n> > On Wed, Apr 02, 2025 at 11:48:01AM -0700, Martin von Zweigbergk wrote:\n> > The biggest question is of course backwards compatibility -- can we\n> > introduce a change ID into the commit metadata without breaking existing\n> > users? I guess you'll already have a lot of experience with this given\n> > that you essentially already inject change IDs into metadata, and tools\n> > generally handle this just fine?\n> \n> Jujutsu has been injecting a 'jj:trees' header into commits to track\n> more metadata around merge conflicts. There weren't any problems with\n> that, unless one uses git to rewrite these commits with e.g. git-rebase,\n> in which case that header is simply lost. But commits with conflicts are\n> usually not pushed to a remote anyway, so the risk there was minimal.\n> Scott Chacon with GitButler has more experience in this regard, since\n> they actually push commits with a change-id in its header to remotes.\n> He told the Jujutsu community that they didn't encounter any problems,\n> no misbehaving tools that are fussy about unknown headers. The only\n> problem is unknown commit headers being dropped by Git itself, depending\n> on how it is invoked by the remote. (GitHub seems to preserve the header\n> during a rebase-merge, because they use git-replay. GitLab and Forgejo\n> drop the header.) With these insights from Scott, Jujutsu is moving\n> forward to put the change-id in the commit header.\n\nYeah, Scott made me aware of the limitations in GitLab already. We\nwanted to migrate to git-replay(1) for a long time already, but never\ngot around to actually doing this. Coincidentally I have recently been\ntalking with Chris, who proposed to finally go through with this change.\nI guess this here is another factor that will make us schedule this\nchange sooner rather than later.\n\nSo: we'll soon start working on it, but I won't promise any timeline.\n\nPatrick\n"},{"id":"515592","messageId":"CABPp-BFRz-yjnti4W17AEBozb0v52kmNsgTLUZW6-MF34R-xdw@mail.gmail.com","threadId":"63245","inReplyTo":"CAESOdVAspxUJKGAA58i0tvks4ZOfoGf1Aa5gPr0FXzdcywqUUw@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2025-04-03T15:39:31Z","receivedAt":"2025-04-03T15:39:43Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Wed, Apr 2, 2025 at 11:48 AM Martin von Zweigbergk\n<martinvonz@google.com> wrote:\n>\n> Hi,\n>\n> The Gerrit, GitButler, and Jujutsu projects all have a concept of\n> a \"change id\", and it behaves in a similar way between the three\n> tools. The change id is conceptually associated with a commit.\n> It follows a commit as its rewritten (e.g. by amending and\n> rebasing). The three projects currently store and format the\n> change id differently. We would like to unify that so we can\n> interoperate better. We hope the Git project is also interested\n> in preserving and using this header.\n>\n> There are many benefits to having a change id even if it's just\n> local. I mentioned some in my email to this mailing list in [1].\n> For example, it enables\n> `git rebase main <change ID>; git switch <change ID>` without\n> requiring the user to look up the hash of the rewritten commit.\n\nBut <change ID> isn't unique, right?  The whole point of having the\nchange ID is to preserve it despite edits (e.g. rebase, commit\n--amend, cherry-pick), meaning that you end up with multiple commits\nwith the same <change ID>.\n\nWhy would this work?\n\nAnd if it does work, isn't it expensive since you'd need to walk\nhistory to find it?  Or do you keep an extra lookup table on the side\nsomewhere?\n\n> If the change id also transferred between repos and preserved by\n> a forge (such as Gerrit), it enables the change id to be used to\n> identify a code review.\n>\n> Here's how the change ids are currently stored and formatted:\n>\n>  * Gerrit currently stores change ids in a commit trailer called\n>    `Change-Id`. It always starts with the letter 'I' and is\n>    followed by 40 hex digits. For example:\n>    `Change-Id: Ib563e78c3fedcff262255fa025441daa3202311b`.\n>\n>  * GitButler currently stores change ids in a commit footer\n>    called `gitbutler-change-id` (older versions used\n>    `change-id`). It's written as 32 hex digits separated by\n>    dashes as in the UUID  format. For example:\n>    `gitbutler-change-id  7d0fbc63-032d-413c-8ae8-610fbeb713c0`.\n>\n>  * Jujutsu currently stores change ids in a local storage outside\n>    of the Git repo and is therefore not part of the Git commit\n>    id. It is stored as 16 bytes. It is rendered to the user as\n>   \"reverse hex\" using 'z' through 'k' as hex digits ('z' = 0,\n>   'k' = 15). This allows even short prefixes to be distinguished\n>    from commit  ids, which is a very useful property when used in\n>    the CLI.\n>\n> As mentioned, the three projects would like to use the same\n> storage and format. I think we have a consensus to store it in a\n> Git commit header called `change-id` as a 32 reverse-hex digis.\n> For example: `change-id ywlktllmukprnxnmzzprukpuwyztylwt`.\n\nYaay, I always hated it as a trailer.\n\n> There is a design doc [2] about the impact on Gerrit and how to\n> handle various cases where the client doesn't understand the\n> `change-id` header. That also includes some discussion about\n> whether cherry-picking should preserve the change id or create a\n> new one. I think there is a lot of value in having a\n> standardized header regardless of what we decide about\n> cherry-picks.\n\ncherry-pick & rebase preserve author name, email & time, while\ncreating a new committer name, email, & time.  To me, the change-id is\nabout the authorship, and since these commands already preserve\nauthorship, it'd seem weird to me to have cherry-pick not preserve the\nchange-id by default.\n\n> So, to be clear, this is mostly a heads up at this point; we don't\n> depend on any immediate changes from the Git project.\n\nI appreciate the heads up, and agree based on what I've seen so far\nthat you can at least get started without Git changes.\n\nHowever, I think you'd want git to preserve the change-id headers upon\ngit commit --amend, rebase, or cherry-pick, which would require some\ngit changes.  And you may want git to preserve them when doing a\nfast-export, and be able to read them in with fast-import.\n\nAnyway, I was a voice in the past that was kind of against these,\nthough that was mostly as a commit footer.  Plus, the number of\nprojects using them, hearing about their experience at Git Merge, and\nrealizing that part of my objection was due to misuse of Gerrit by\nsome folks in the past have all lead me to change my opinion.\n"},{"id":"515594","messageId":"CABPp-BGwXaiohvfSdr96hzKNPYXQqz+_okxLNj7P9KSjX2PW6g@mail.gmail.com","threadId":"63245","inReplyTo":"Z-5QR57zgSsm6jNP@pks.im","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2025-04-03T15:56:01Z","receivedAt":"2025-04-03T15:56:14Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Thu, Apr 3, 2025 at 2:13 AM Patrick Steinhardt <ps@pks.im> wrote:\n>\n[...]\n> Agreed, change IDs solve a couple of issues that many users face:\n>\n>   - You can reliably track how a patch evolves over time. This helps\n>     various different tools to track identity of commits, like for\n>     example forges, but also tools like git-range-diff(1).\n>\n>   - It becomes trivial to see whether a commit has been cherry-picked\n>     into another branch. We do have git-cherry(1) to do that right now,\n>     but that command is based on heuristics and fails as soon as the\n>     patch itself needed to be adapted.\n>\n>   - Working with history rewrites becomes easier in the general case as\n>     you don't have to adapt to constantly changing commit IDs.\n\nCould you elaborate?  I agree with the other points you raise, but I'm\nunsure how this helps with a history rewrite.  Do you mean the\nrewriting of history, or someone trying to consume the history\nrewrite?  If the former, I don't see it, and if the latter, didn't you\nalready cover that in the two bullets above?  Or is there something\nelse you are also getting at?\n\n> So what would it take to get change IDs into Git? I think the most\n> important items would be:\n>\n>   - Generating and writing change IDs in commands that support them.\n>     This includes e.g. git-commit(1), git-commit-tree(1), git-merge(1),\n>     git-merge-tree(1). This should of course be completely optional and\n>     probably be disabled by default.\n>\n>   - Making tools that rewrite commits aware of change IDs so that they\n>     know to retain change IDs. This involves e.g. git-cherry-pick(1),\n>     git-rebase(1), git-replay(1).\n\nAnd also git-commit(1) [when passing --amend], and git-fast-export(1)\nand git-fast-import(1) -- though possibly with options for the last\ntwo to expunge them instead of preserving them, but probably\ndefaulting to preserving them.\n\nHowever, I think some of these might already handle this.  Commands\nwhich call read_commit_extra_headers() and pass those along to\ncommit_tree_extended() may already preserve these.  It appears commit\n--amend and replay both do this.  sequencer has some code that looks\nrelevant, but it appears to only be reading the headers from HEAD (at\nthe time the nth commit is being replayed), which seems like it'd be\nlooking at the wrong commit.  That might actually be a bug...\n\n>   - Extending revisions to allow specifying commits by change ID.\n\nWould this essentially be similar to <rev>^{/<text>} except searching\nspecifically change-id headers rather than commit message?\n"},{"id":"515596","messageId":"D8X571K4M77Y.2PVKK2KQCRBOM@buenzli.dev","threadId":"63245","inReplyTo":"CABPp-BGwXaiohvfSdr96hzKNPYXQqz+_okxLNj7P9KSjX2PW6g@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Remo Senekowitsch","fromEmail":"remo@buenzli.dev","sentAt":"2025-04-03T16:25:53Z","receivedAt":"2025-04-03T16:26:00Z","isPatch":false,"sender":{"key":"remo@buenzli.dev","avatar":"https://gravatar.com/avatar/7df680b096206886db5a2dc983926f314985bbee662eb406cf23c65322cf98b7?d=mp&s=160"},"body":"On Thu Apr 3, 2025 at 5:56 PM CEST, Elijah Newren wrote:\n> On Thu, Apr 3, 2025 at 2:13 AM Patrick Steinhardt <ps@pks.im> wrote:\n>>\n>>   - Extending revisions to allow specifying commits by change ID.\n>\n> Would this essentially be similar to <rev>^{/<text>} except searching\n> specifically change-id headers rather than commit message?\n\nOne benefit of using the \"reverse-hex\" format (hex with a different\nalphabet: z(0) through k(25)) we're proposing is that it allows a\nchange-id or its prefix to be used in the same place as a commit hash,\nwithout ambiguity.\n\nRemo\n"},{"id":"515597","messageId":"CABPp-BFr+4iy7awessWSY8NzzswY1-=30L4VvOZMpFDoOxJUgg@mail.gmail.com","threadId":"63245","inReplyTo":"D8X571K4M77Y.2PVKK2KQCRBOM@buenzli.dev","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2025-04-03T16:38:55Z","receivedAt":"2025-04-03T16:39:07Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Thu, Apr 3, 2025 at 9:25 AM Remo Senekowitsch <remo@buenzli.dev> wrote:\n>\n> On Thu Apr 3, 2025 at 5:56 PM CEST, Elijah Newren wrote:\n> > On Thu, Apr 3, 2025 at 2:13 AM Patrick Steinhardt <ps@pks.im> wrote:\n> >>\n> >>   - Extending revisions to allow specifying commits by change ID.\n> >\n> > Would this essentially be similar to <rev>^{/<text>} except searching\n> > specifically change-id headers rather than commit message?\n>\n> One benefit of using the \"reverse-hex\" format (hex with a different\n> alphabet: z(0) through k(25)) we're proposing is that it allows a\n> change-id or its prefix to be used in the same place as a commit hash,\n> without ambiguity.\n\nI already saw that, but this doesn't address my question.\n"},{"id":"515598","messageId":"D8X5I3W7K1DI.2JYHGNY9L7ZD3@buenzli.dev","threadId":"63245","inReplyTo":"CABPp-BFRz-yjnti4W17AEBozb0v52kmNsgTLUZW6-MF34R-xdw@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Remo Senekowitsch","fromEmail":"remo@buenzli.dev","sentAt":"2025-04-03T16:40:20Z","receivedAt":"2025-04-03T16:40:31Z","isPatch":false,"sender":{"key":"remo@buenzli.dev","avatar":"https://gravatar.com/avatar/7df680b096206886db5a2dc983926f314985bbee662eb406cf23c65322cf98b7?d=mp&s=160"},"body":"On Thu Apr 3, 2025 at 5:39 PM CEST, Elijah Newren wrote:\n> On Wed, Apr 2, 2025 at 11:48 AM Martin von Zweigbergk\n> <martinvonz@google.com> wrote:\n>>\n>> There are many benefits to having a change id even if it's just\n>> local. I mentioned some in my email to this mailing list in [1].\n>> For example, it enables\n>> `git rebase main <change ID>; git switch <change ID>` without\n>> requiring the user to look up the hash of the rewritten commit.\n>\n> But <change ID> isn't unique, right?  The whole point of having the\n> change ID is to preserve it despite edits (e.g. rebase, commit\n> --amend, cherry-pick), meaning that you end up with multiple commits\n> with the same <change ID>.\n>\n> Why would this work?\n>\n> And if it does work, isn't it expensive since you'd need to walk\n> history to find it?  Or do you keep an extra lookup table on the side\n> somewhere?\n\nFor rebase and commit --amend, the way Jujutsu deals with those is that\nall descendants are immediately rebased on top of the new commit, and\nrefs to those descendants are updated as well. That means, the old\nversion of the patch with the same change-id becomes unreachable. So,\nat least most of the time, the change-id is indeed unique.\n\nThis doesn't work for cherry-pick, more on that below.\n\nSome of these features are not in Git yet, at least not to my knowledge.\nThat means getting the full benefit of change-ids with Git itself\nwould indeed require some more work. I know of rebase.updateRefs\nand rebase.rebaseMerges, which move the Git experience closer to\nJujutsu, but don't go all the way. AFAIK it's not possible with Git to\nautomatically rebase --update-refs all descendants of a commit that is\namended or rebased.\n\nJujutsu does keep a separate index of change-ids, yes.\n\n>> There is a design doc [2] about the impact on Gerrit and how to\n>> handle various cases where the client doesn't understand the\n>> `change-id` header. That also includes some discussion about\n>> whether cherry-picking should preserve the change id or create a\n>> new one. I think there is a lot of value in having a\n>> standardized header regardless of what we decide about\n>> cherry-picks.\n>\n> cherry-pick & rebase preserve author name, email & time, while\n> creating a new committer name, email, & time.  To me, the change-id is\n> about the authorship, and since these commands already preserve\n> authorship, it'd seem weird to me to have cherry-pick not preserve the\n> change-id by default.\n\nI'd say Jujutsu, Gerrit and GitButler think of a change-id as associated\nwith a unit of review. (Although it will naturally support reviewing\nsets of patches as well.) Usually only one person will push commits with\nthe same change-id, just like people don't usually force-push over each\nothers branches. But that's mostly about avoiding logistical problems.\nWhen an employee leaves a company or is on vacation, it can be perfectly\nreasonable for someone else to take over their work. In that case, it\nwould be appropriate to preserve the change-id, even though authorship\nhas changed, because the history of code review on that patch should\nstay associated with the new version.\n\nCherry-picking on the other hand often represents a separate unit of\nreview. That review may revolve around whether it makes sense to\nbackport a bugfix at all or any additional changes that may have been\nnecessary to make the bugfix work in the different, older codebase.\n\nAs mentioned above, there's also the issue that preserving the change-id\non cherry-pick likely results in duplicates. For Jujutsu, it would be\nnice it this was avoided. But it's not infeasible to deal with that\neither.\n\nFor Gerrit, it would be important to be able to track a change across\ncherry-picks somehow, since that is a feature they already have. If Git\ndecides to preserve the change-id on cherry-pick, there's no problem\nfor Gerrit. Alternatives include storing a separate cherry-picked-from\nheader or enabling the -x flag on cherry-pick by default.\n\nRemo\n"},{"id":"515600","messageId":"20250403174847.GB3051250@mit.edu","threadId":"63245","inReplyTo":"CABPp-BFRz-yjnti4W17AEBozb0v52kmNsgTLUZW6-MF34R-xdw@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Theodore Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2025-04-03T17:48:47Z","receivedAt":"2025-04-03T17:49:11Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Thu, Apr 03, 2025 at 08:39:31AM -0700, Elijah Newren wrote:\n> \n> But <change ID> isn't unique, right?  The whole point of having the\n> change ID is to preserve it despite edits (e.g. rebase, commit\n> --amend, cherry-pick), meaning that you end up with multiple commits\n> with the same <change ID>.\n\nIt's supposed to be unique, but it isn't always.  I've certainly seen\ncases where it might not be, but that's arguably a bug.  I suspect in\nsome cases it's because users are cutting and pasting commit\ndescriptions, and sometimes when they rebase a patch series, patches\nwill get collapsed or split apart --- especially when backporting to\nan older LTS release.\n\nPerhaps because of this, in some communities, their tooling in front\nof Gerrit will always regenerate the Commit-ID when doing a\ncherry-pick (For example, when cherry-picking from the development\nHEAD branch back to a release branch).\n\nSo as a cauaionary note, as people use Change ID's in Gerrit today,\nsometimes the Change ID changes between rebases, and I've certainly\nseen cases where the sematic meaning of the commit has changed\nsignificantly without changing the Change ID.  So it's great as a\nhint, but in practice, at least today, it might not be completely safe\nto assume the semantics are as advertised....\n\n\t\t\t\t- Ted\n"},{"id":"515602","messageId":"Z+7PDi5y4wXJBK4r@ubby","threadId":"63245","inReplyTo":"CABPp-BFRz-yjnti4W17AEBozb0v52kmNsgTLUZW6-MF34R-xdw@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-03T18:10:22Z","receivedAt":"2025-04-03T18:10:32Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Thu, Apr 03, 2025 at 08:39:31AM -0700, Elijah Newren wrote:\n> On Wed, Apr 2, 2025 at 11:48 AM Martin von Zweigbergk\n> > There are many benefits to having a change id even if it's just\n> > local. I mentioned some in my email to this mailing list in [1].\n> > For example, it enables\n> > `git rebase main <change ID>; git switch <change ID>` without\n> > requiring the user to look up the hash of the rewritten commit.\n> \n> But <change ID> isn't unique, right?  The whole point of having the\n> change ID is to preserve it despite edits (e.g. rebase, commit\n> --amend, cherry-pick), meaning that you end up with multiple commits\n> with the same <change ID>.\n> \n> Why would this work?\n\nI agree that `git rebase main <change ID>; git switch <change ID>` is\nnot a good UI, and I wouldn't want it even though I want change IDs.\n\nChange IDs are documentary, and they help people understand that various\ncommits are just edits/rewrites of the same original commit.\n\n> And if it does work, isn't it expensive since you'd need to walk\n> history to find it?  Or do you keep an extra lookup table on the side\n> somewhere?\n\nWorse: since there can be many commits with the same change ID they\ncan't be used as refs because Git can't possibly be expected to find\n_the one_ you really intend -- how could it?  I suppose Git could let\nyou pick from a list, but that's not likely going to have enough\ncontext.  Maybe Git could give you a list of named branches in which it\nfound some change ID's commits to pick one branch from, or maybe one\ncould `git cherry-pick --from $some_branch $cid` and have Git find the\ncommit(s) on `$some_branch` that match change ID `$cid`.\n\nAlso, maybe commits need to support multiple change IDs.  For example,\none might want one change ID for a set of commits (eg implementing a\nfeature), and each commit in the set having a second and more unique\nchange ID for just that commit (and all edits/rewrites of it).\n\n> Yaay, I always hated it as a trailer.\n\nMore headers is better.\n\n> cherry-pick & rebase preserve author name, email & time, while\n> creating a new committer name, email, & time.  To me, the change-id is\n> about the authorship, and since these commands already preserve\n> authorship, it'd seem weird to me to have cherry-pick not preserve the\n> change-id by default.\n\n+1.  Besides, rebase is just a a bunch of cherry-picks, so cherry-pick\nis the fundamental operation here.  If rebase is to preserver change-id\nthen so is cherry-pick.\n\nNico\n-- \n"},{"id":"515606","messageId":"D8XAETM3PJ4A.1RBPVHJWE9DFT@buenzli.dev","threadId":"63245","inReplyTo":"20250403174847.GB3051250@mit.edu","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Remo Senekowitsch","fromEmail":"remo@buenzli.dev","sentAt":"2025-04-03T20:31:08Z","receivedAt":"2025-04-03T20:31:15Z","isPatch":false,"sender":{"key":"remo@buenzli.dev","avatar":"https://gravatar.com/avatar/7df680b096206886db5a2dc983926f314985bbee662eb406cf23c65322cf98b7?d=mp&s=160"},"body":"On Thu Apr 3, 2025 at 7:48 PM CEST, Theodore Ts'o wrote:\n> On Thu, Apr 03, 2025 at 08:39:31AM -0700, Elijah Newren wrote:\n>> \n>> But <change ID> isn't unique, right?  The whole point of having the\n>> change ID is to preserve it despite edits (e.g. rebase, commit\n>> --amend, cherry-pick), meaning that you end up with multiple commits\n>> with the same <change ID>.\n>\n> It's supposed to be unique, but it isn't always.  I've certainly seen\n> cases where it might not be, but that's arguably a bug.  I suspect in\n> some cases it's because users are cutting and pasting commit\n> descriptions, and sometimes when they rebase a patch series, patches\n> will get collapsed or split apart --- especially when backporting to\n> an older LTS release.\n>\n> Perhaps because of this, in some communities, their tooling in front\n> of Gerrit will always regenerate the Commit-ID when doing a\n> cherry-pick (For example, when cherry-picking from the development\n> HEAD branch back to a release branch).\n>\n> So as a cauaionary note, as people use Change ID's in Gerrit today,\n> sometimes the Change ID changes between rebases, and I've certainly\n> seen cases where the sematic meaning of the commit has changed\n> significantly without changing the Change ID.  So it's great as a\n> hint, but in practice, at least today, it might not be completely safe\n> to assume the semantics are as advertised....\n\nThat is all true. I would just say that some change-ids not being unique\ndoesn't systematically take away from the benefits. If a given change-id\nis unique, you get all the benefits for that patch, independent of\nthe uniqueness of the other change-ids. If some patch changes a lot\nsemantically while keeping its change-id, that will degrade its review\nhistory, but without affecting the review history of any other patch.\n\nRemo\n"},{"id":"515609","messageId":"D8XC02N85O7W.2LQY7B8Q2Z98G@buenzli.dev","threadId":"63245","inReplyTo":"Z+7PDi5y4wXJBK4r@ubby","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Remo Senekowitsch","fromEmail":"remo@buenzli.dev","sentAt":"2025-04-03T21:45:55Z","receivedAt":"2025-04-03T21:46:01Z","isPatch":false,"sender":{"key":"remo@buenzli.dev","avatar":"https://gravatar.com/avatar/7df680b096206886db5a2dc983926f314985bbee662eb406cf23c65322cf98b7?d=mp&s=160"},"body":"On Thu Apr 3, 2025 at 8:10 PM CEST, Nico Williams wrote:\n> On Thu, Apr 03, 2025 at 08:39:31AM -0700, Elijah Newren wrote:\n\n> Also, maybe commits need to support multiple change IDs.  For example,\n> one might want one change ID for a set of commits (eg implementing a\n> feature), and each commit in the set having a second and more unique\n> change ID for just that commit (and all edits/rewrites of it).\n\nSemantically grouping a set of commits is an intersting idea, but not\nreally related to the proposed change-id? Mercurial has \"topics\", which\nI don't know too much about. Even if such a \"topic-identifier\" were\nto be stored in the commit header, it probably shouldn't be called the\nsame thing as the change-id header. Otherwise it becomes impossible to\ndetermine if a change-id header is supposed to identify an individual\npatch across its versions or associate it semantically with others.\n\n>> cherry-pick & rebase preserve author name, email & time, while\n>> creating a new committer name, email, & time.  To me, the change-id is\n>> about the authorship, and since these commands already preserve\n>> authorship, it'd seem weird to me to have cherry-pick not preserve the\n>> change-id by default.\n>\n> +1.  Besides, rebase is just a a bunch of cherry-picks, so cherry-pick\n> is the fundamental operation here.  If rebase is to preserver change-id\n> then so is cherry-pick.\n\nWell, if rebase is just a bunch of cherry-picks, that sounds like an\nimplementation detail. The ways rebase and cherry-pick are most often\nused are semantically very different from each other. (interactive)\nrebase is often used to amend commits that already have descendants. In\nthat case, it makes sense for the change-id to be preserved. cherry-pick\non the other hand is often used to create a dublicate of a patch at a\ndifferent location in the commit tree, e.g. for backporting purposes.\nThe equivalent of cherry-pick in Jujutsu is called \"duplicate\" and\ndoesn't preserve the change-id for that reason. So if cherry-pick\nretains the change-id, that would lead to more duplicate change-ids in\nthe tree, degrading their usefulness in those cases for little benefit.\nBut again, if Git decides to go that route it's still better for the\nother projects than not supporting the header at all.\n\nRemo\n"},{"id":"515610","messageId":"CAESOdVB7VhEhBJJYVY8ZdbShQPRKcoWu=7YQPBwRp93iH2yvWA@mail.gmail.com","threadId":"63245","inReplyTo":"CABPp-BFr+4iy7awessWSY8NzzswY1-=30L4VvOZMpFDoOxJUgg@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Martin von Zweigbergk","fromEmail":"martinvonz@google.com","sentAt":"2025-04-03T21:46:22Z","receivedAt":"2025-04-03T21:46:35Z","isPatch":false,"sender":{"key":"martinvonz@google.com","avatar":"https://avatars.githubusercontent.com/u/891642?v=4"},"body":"On Thu, 3 Apr 2025 at 09:39, Elijah Newren <newren@gmail.com> wrote:\n>\n> On Thu, Apr 3, 2025 at 9:25 AM Remo Senekowitsch <remo@buenzli.dev> wrote:\n> >\n> > On Thu Apr 3, 2025 at 5:56 PM CEST, Elijah Newren wrote:\n> > > On Thu, Apr 3, 2025 at 2:13 AM Patrick Steinhardt <ps@pks.im> wrote:\n> > >>\n> > >>   - Extending revisions to allow specifying commits by change ID.\n> > >\n> > > Would this essentially be similar to <rev>^{/<text>} except searching\n> > > specifically change-id headers rather than commit message?\n> >\n> > One benefit of using the \"reverse-hex\" format (hex with a different\n> > alphabet: z(0) through k(25)) we're proposing is that it allows a\n> > change-id or its prefix to be used in the same place as a commit hash,\n> > without ambiguity.\n>\n> I already saw that, but this doesn't address my question.\n\nJujutsu has a persistent index of change ids to commit ids. It's\nsimilar to Git's commit graph. As you might expect, this index often\nhas multiple commits for a single change id. When needed, we then\nbuild an in-memory index of only the subset of commits currently\nreachable from the visible head commits (you can think of it as all\ncurrently reachable commits from Git branches, but we don't require a\nbranch to keep a commit visible). When you do something like `jj show\nxyz`, we use that in-memory index to find the relevant commits. That's\nusually just one commit.\n\nBy the way, we actually have another level of lookup that happens\nbefore that. You can configure an expression for the set of commits\nyou want to prioritize short change id prefixes and commit id prefixes\nfor. If you say that you want all commits that are not ancestors of\nany remote-tracking branch to be shorter, for example, then we build\nan additional in-memory index of only those commits. That means that\nboth change id prefixes and commit id prefixes are often just a few\ncharacters long, even in huge repos like the monorepo at Google.\n"},{"id":"515611","messageId":"CAESOdVAd+X=6nEULHtKKotH_W5yNaJAcUajRU79EuG+0SF3m1A@mail.gmail.com","threadId":"63245","inReplyTo":"Z+7PDi5y4wXJBK4r@ubby","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Martin von Zweigbergk","fromEmail":"martinvonz@google.com","sentAt":"2025-04-03T22:05:41Z","receivedAt":"2025-04-03T22:05:54Z","isPatch":false,"sender":{"key":"martinvonz@google.com","avatar":"https://avatars.githubusercontent.com/u/891642?v=4"},"body":"I think I may have answered some of your questions here in my other\nreply to Elijah, so consider reading that too.\n\nOn Thu, 3 Apr 2025 at 11:10, Nico Williams <nico@cryptonector.com> wrote:\n>\n> I agree that `git rebase main <change ID>; git switch <change ID>` is\n> not a good UI, and I wouldn't want it even though I want change IDs.\n\nWhy do you think it's not a good UI? Is it because the change ID isn't\nmeaningful? That's correct, but they are also very convenient. The\nunique prefix is usually two letters or so, depending on how many\n\"local\" commits you have in your repo. That makes them easy to type. I\nbasically never refer to a commit by a branch name anymore.\n\n> > And if it does work, isn't it expensive since you'd need to walk\n> > history to find it?  Or do you keep an extra lookup table on the side\n> > somewhere?\n>\n> Worse: since there can be many commits with the same change ID they\n> can't be used as refs because Git can't possibly be expected to find\n> _the one_ you really intend -- how could it?  I suppose Git could let\n> you pick from a list, but that's not likely going to have enough\n> context.  Maybe Git could give you a list of named branches in which it\n> found some change ID's commits to pick one branch from, or maybe one\n> could `git cherry-pick --from $some_branch $cid` and have Git find the\n> commit(s) on `$some_branch` that match change ID `$cid`.\n\nSee my reply to Elijah. There's usually just one visible commit with a\ngiven change id a repo.\n"},{"id":"515613","messageId":"CABeNrKX3fY8qmASgyKaSv99LkGsrcExKFwNtgaKqjfJdQn8vrQ@mail.gmail.com","threadId":"63245","inReplyTo":"D8X5I3W7K1DI.2JYHGNY9L7ZD3@buenzli.dev","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Kane York","fromEmail":"kanepyork@gmail.com","sentAt":"2025-04-03T22:11:08Z","receivedAt":"2025-04-03T22:11:10Z","isPatch":false,"sender":{"key":"kanepyork@gmail.com","avatar":null},"body":"On Thu, 03 Apr 2025 18:40:20 +0200, Remo Senekowitsch wrote:\n> On Thu Apr 3, 2025 at 5:39 PM CEST, Elijah Newren wrote:\n>> cherry-pick & rebase preserve author name, email & time, while creating a\n>> new committer name, email, & time.  To me, the change-id is about the\n>> authorship, and since these commands already preserve authorship, it'd seem\n>> weird to me to have cherry-pick not preserve the change-id by default.\n\n> I'd say Jujutsu, Gerrit and GitButler think of a change-id as associated with\n> a unit of review.\n>\n> [...]\n>\n> Cherry-picking on the other hand often represents a separate unit of review.\n> That review may revolve around whether it makes sense to backport a bugfix at\n> all or any additional changes that may have been necessary to make the bugfix\n> work in the different, older codebase.\n>\n> As mentioned above, there's also the issue that preserving the change-id on\n> cherry-pick likely results in duplicates. For Jujutsu, it would be nice it\n> this was avoided. But it's not infeasible to deal with that either.\n>\n> For Gerrit, it would be important to be able to track a change across\n> cherry-picks somehow, since that is a feature they already have. If Git\n> decides to preserve the change-id on cherry-pick, there's no problem for\n> Gerrit. Alternatives include storing a separate cherry-picked-from header or\n> enabling the -x flag on cherry-pick by default.\n\nI agree, and propose this concrete behavior:\n\n- git cherry-pick generates a fresh `change-id`, and places the old change-id\n  in a `cherry-picked-change` header\n- git cherry-pick preserves the old `change-id` if passed the new\n  `--preserve-change-id` flag\n- git rebase passes the `--preserve-change-id` flag on 'pick' actions, unless\n  passed the new `--no-preserve-change-id` flag\n- git rebase uses the earlier commit's change-id on 'fixup' actions\n- git rebase prints the change-ids into the COMMIT_MSG on 'squash' actions, and\n  tries to read the user's choice of which to use, defaulting to the earlier\n  commit if both or none are present\n- git rebase creates a new change-id on 'merge' actions\n- git rebase needs no special behavior specified for 'edit', 'exec', 'break',\n  'drop', 'label', 'reset', 'update-ref'\n"},{"id":"515615","messageId":"Z+8IF67AC8gSouYc@ubby","threadId":"63245","inReplyTo":"CAESOdVAd+X=6nEULHtKKotH_W5yNaJAcUajRU79EuG+0SF3m1A@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-03T22:13:43Z","receivedAt":"2025-04-03T22:20:58Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Thu, Apr 03, 2025 at 03:05:41PM -0700, Martin von Zweigbergk wrote:\n> On Thu, 3 Apr 2025 at 11:10, Nico Williams <nico@cryptonector.com> wrote:\n> >\n> > I agree that `git rebase main <change ID>; git switch <change ID>` is\n> > not a good UI, and I wouldn't want it even though I want change IDs.\n> \n> Why do you think it's not a good UI? Is it because the change ID isn't\n> meaningful? That's correct, but they are also very convenient. The\n> unique prefix is usually two letters or so, depending on how many\n> \"local\" commits you have in your repo. That makes them easy to type. I\n> basically never refer to a commit by a branch name anymore.\n\nWhat would `git rebase main <change ID>` do?  I assumed that it would\nfind a commit `<change ID>` in the current branch and rebase it onto the\nmain branch.  That seems workable, I suppose.  It would be akin to\n`git cherry-pick --from $branch --change-id $change_id`.\n\nWhat would `git switch <change ID>` do?  `git switch` switches between\nbranches, but a change ID can't possibly identify a branch since many\ncommits could exist with the same change ID all in different branches.\n\n> > > And if it does work, isn't it expensive since you'd need to walk\n> > > history to find it?  Or do you keep an extra lookup table on the side\n> > > somewhere?\n> >\n> > Worse: since there can be many commits with the same change ID they\n> > can't be used as refs because Git can't possibly be expected to find\n> > _the one_ you really intend -- how could it?  I suppose Git could let\n> > you pick from a list, but that's not likely going to have enough\n> > context.  Maybe Git could give you a list of named branches in which it\n> > found some change ID's commits to pick one branch from, or maybe one\n> > could `git cherry-pick --from $some_branch $cid` and have Git find the\n> > commit(s) on `$some_branch` that match change ID `$cid`.\n> \n> See my reply to Elijah. There's usually just one visible commit with a\n> given change id a repo.\n\ns/a repo/in a repo/ ?\ns/a repo/in a branch/ ?\n\nI'd expect many commits with the same change ID in a _repo_, but at most\none in a _branch_.  Except see my other comments about having a second\nkind of change ID to identify a set of related commits (e.g., if you\nhave a rule that features, bug fixes, and tests all must go in separate\ncommits, which some do, so you might need to have a main commit for some\nfeature, several commits with fixes to earlier bugs required for that\nfeature, and several test commits, all related to that feature).\n\nNico\n-- \n"},{"id":"515619","messageId":"CAESOdVAWWP=Rte4bx3zUZc6p0XiZaJS2OZr8ezRPkfq8K1TYfw@mail.gmail.com","threadId":"63245","inReplyTo":"Z+8IF67AC8gSouYc@ubby","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Martin von Zweigbergk","fromEmail":"martinvonz@google.com","sentAt":"2025-04-03T22:47:30Z","receivedAt":"2025-04-03T22:47:43Z","isPatch":false,"sender":{"key":"martinvonz@google.com","avatar":"https://avatars.githubusercontent.com/u/891642?v=4"},"body":"On Thu, 3 Apr 2025 at 15:13, Nico Williams <nico@cryptonector.com> wrote:\n>\n> On Thu, Apr 03, 2025 at 03:05:41PM -0700, Martin von Zweigbergk wrote:\n> > On Thu, 3 Apr 2025 at 11:10, Nico Williams <nico@cryptonector.com> wrote:\n> > >\n> > > I agree that `git rebase main <change ID>; git switch <change ID>` is\n> > > not a good UI, and I wouldn't want it even though I want change IDs.\n> >\n> > Why do you think it's not a good UI? Is it because the change ID isn't\n> > meaningful? That's correct, but they are also very convenient. The\n> > unique prefix is usually two letters or so, depending on how many\n> > \"local\" commits you have in your repo. That makes them easy to type. I\n> > basically never refer to a commit by a branch name anymore.\n>\n> What would `git rebase main <change ID>` do?  I assumed that it would\n> find a commit `<change ID>` in the current branch and rebase it onto the\n> main branch.  That seems workable, I suppose.  It would be akin to\n> `git cherry-pick --from $branch --change-id $change_id`.\n\nI think part of the problem is that I didn't consider that Git doesn't\nreally like to work in detached HEAD mode and doesn't automatically\nupdate refs pointing to rewritten commits. This does take away a lot\nof the usefulness, unfortunately. It would still be a bit useful as an\nargument to readonly commands.\n\n> What would `git switch <change ID>` do?  `git switch` switches between\n> branches, but a change ID can't possibly identify a branch since many\n> commits could exist with the same change ID all in different branches.\n\nYes, the same change id *can* exist on many branches, but it's pretty\nuncommon. It might happen after cherry-picking, depending on what we\ndecide there, but it should very rarely happen in other cases. When\nyou rewrite a commit, the old commit usually becomes unreachable, so\nif your change id resolved to one commit before the rewrite, then it\nwill resolve to one commit after the rewrite. I know Git often leaves\ndescendant branches until you manually rebase them, but at least\nthat's probably typically a pretty short-lived state. I had a script\nfor this back when I used Git.\n\n> > > > And if it does work, isn't it expensive since you'd need to walk\n> > > > history to find it?  Or do you keep an extra lookup table on the side\n> > > > somewhere?\n> > >\n> > > Worse: since there can be many commits with the same change ID they\n> > > can't be used as refs because Git can't possibly be expected to find\n> > > _the one_ you really intend -- how could it?  I suppose Git could let\n> > > you pick from a list, but that's not likely going to have enough\n> > > context.  Maybe Git could give you a list of named branches in which it\n> > > found some change ID's commits to pick one branch from, or maybe one\n> > > could `git cherry-pick --from $some_branch $cid` and have Git find the\n> > > commit(s) on `$some_branch` that match change ID `$cid`.\n> >\n> > See my reply to Elijah. There's usually just one visible commit with a\n> > given change id a repo.\n>\n> s/a repo/in a repo/ ?\n> s/a repo/in a branch/ ?\n\nThe former.\n\n> I'd expect many commits with the same change ID in a _repo_, but at most\n> one in a _branch_.\n\nThere may be multiple commits with the same change ID in a repo if you\nhad cherry-picked the commit, depending on what we decide to do with\nthe change ID on cherry-pick. But cherry-picks are not very common\nanyway. Or maybe it's common in some workflow? Oh, are you thinking of\na scenario where you cherry-pick your own commit to see if an\nalternative approach is better? Sure, if cherry-pick preserve the\nchange ID, then you would have multiple commits with the same change\nID in that case.\n"},{"id":"515621","messageId":"D8XEYQ9TRB10.L5S89IAC2LZ9@buenzli.dev","threadId":"63245","inReplyTo":"Z+8GoNrdaJlmNpGm@ubby","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Remo Senekowitsch","fromEmail":"remo@buenzli.dev","sentAt":"2025-04-04T00:05:13Z","receivedAt":"2025-04-04T00:05:18Z","isPatch":false,"sender":{"key":"remo@buenzli.dev","avatar":"https://gravatar.com/avatar/7df680b096206886db5a2dc983926f314985bbee662eb406cf23c65322cf98b7?d=mp&s=160"},"body":"On Fri Apr 4, 2025 at 12:07 AM CEST, Nico Williams wrote:\n> On Thu, Apr 03, 2025 at 11:45:55PM +0200, Remo Senekowitsch wrote:\n>\n> Regardless, all operations that \"alter\" a commit, such as by cherry-\n> picking or rebasing it onto some other commit, should have the same\n> defaults and options for preserving/dropping metadata such as \"change\n> ID\".\n>\n> If I cherry-pick a commit then I absolutely want its \"change ID\" to be\n> preserved by default.  If I want to drop that I can always ask for that\n> or amend the commit to remove it.  I will want the same behavior for\n> rebase and cherry-pick.  Having to remember different defaults and\n> options for the two would be a cognitive load I do not need.\n\nYeah, that's a very valid argument. In Jujutsus CLI, there is a very\nclear separation between \"rebase\" and \"duplicate\", so there's no risk\nof confusion if one preserves the change-id and the other doesn't. In\nGit, the saparation between rebase and cherry-pick is less clear-cut.\nMaking them behave the same way can be seen as simpler.\n\n>>                 [...]. The ways rebase and cherry-pick are most often\n>> used are semantically very different from each other. (interactive)\n>\n> How do you know this is \"most often\" so? [...]\n\nI haven't conducted a study, this is my impression from talking to peers\nand reading chatter from other Git users online. Maybe the impression is\nwrong.\n\n>> rebase is often used to amend commits that already have descendants. In\n>> that case, it makes sense for the change-id to be preserved. cherry-pick\n>> on the other hand is often used to create a dublicate of a patch at a\n>> different location in the commit tree, e.g. for backporting purposes.\n>\n> That's not how I use rebase.  I rebase to:\n>\n>  - catch up with upstream changes\n>  - reorder commits\n>  - combine commits\n>  - split commits\n>  - edit commit metadata\n>  - edit commit contents\n\nYeah, I was (over-)simplifying. rebase is the swiss-army knife of git\ncommands. But for all of these operations, it holds that the previous\nversion of the patch(es) won't be reachable in the commit tree anymore\nafter the rebase is complete. (assuming potential descendant branches\nare also rebased, which is usually the case) So rebase doesn't generally\ncause duplicate change-ids, which is what I wanted to get at.\n\n> [...] The whole point of a \"change ID\" is to let the user notice that\n> some set of commits all share the same origin, such as all being a fix\n> to the same bug each in a different release (backports).  If you're\n> backporting bug fixes you'll really want the change IDs to be preserved.\n>\n> When would you not want to preserve a change ID on cherry-pick?  I can't\n> say I would ever have wanted to do that had Git had change IDs from day\n> 1, and I've been using Git for more than twenty years.\n\nThat's not exactly how Jujutsu thinks about the change-id, but it's a\nuseful piece of information. Gerrit does indeed use its change-id to\ntrack cherry-picks. I am in favor of measures to track that metadata\n(although duplicating change-ids is not my preferred option for that).\n\nLet's assume the change-id represents the origin of a patch. What should\nhappen if a patch is split in two? Should they have the same change-id,\nbecause they ultimately have the same origin? Maybe.\n\nI don't attach too much semantic meaning to the change-id. It's a\nnormally unique identifier for a change that persists as the change\nevolves. That's useful. The more commits with the same change-id as\nothers there are, the less useful the concept becomes.\n\n>> doesn't preserve the change-id for that reason. So if cherry-pick\n>\n> I have _never_ used cherry-pick to cause there to be duplicate commits\n> in the same branch.  Therefore calling it \"duplicate\" seems terribly\n> wrong to me.\n\nWell, obviously not in the same branch. I meant duplicate among all\nvisible commits (reachable from any branch). That's the issue we're\ndiscussing w.r.t. change-ids not always being unique identifiers for\na single commit. What would you like me to call that siuation instead\nof duplicate?\n\nCan you maybe give some examples of how you use cherry-pick? I'd be\ninterested in your use cases to maybe better understand where you're\ncoming from. I myself almost never use cherry-pick, simply because I'm\nnot involved in any backporting. I've seen cherry-pick used to get a\nbugfix from another branch onto your own, in order to avoid having to\nwait for the other branch to be merged. But that practice has always\nrubbed me the wrong way. I feel like the correct thing to do in that\nsituation is to extract the bugfix to a separate dependency-free branch\nand make the two feature branches depend on it. That way, both feature\nbranches can more easily track changes in the bugfix by rebasing. If the\nbugfix was cherry-picked, it's much harder to keep the two versions in\nsync. (And finally, the latter approach probably makes the bugfix land\nfaster.) So yeah, interested to hear your use-cases for cherry-pick.\n\nRemo\n"},{"id":"515623","messageId":"CABPp-BHHz3zASVEmePxEW4qsu85Q_kfO8JHeb67CB6DMRHmUKg@mail.gmail.com","threadId":"63245","inReplyTo":"CAESOdVAWWP=Rte4bx3zUZc6p0XiZaJS2OZr8ezRPkfq8K1TYfw@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2025-04-04T02:06:33Z","receivedAt":"2025-04-04T02:06:46Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Thu, Apr 3, 2025 at 3:47 PM Martin von Zweigbergk\n<martinvonz@google.com> wrote:\n>\n[...]\n> > What would `git switch <change ID>` do?  `git switch` switches between\n> > branches, but a change ID can't possibly identify a branch since many\n> > commits could exist with the same change ID all in different branches.\n>\n> Yes, the same change id *can* exist on many branches, but it's pretty\n> uncommon.\n\nIn my experience with Gerrit-based projects, it was actually pretty common...\n\n> > I'd expect many commits with the same change ID in a _repo_, but at most\n> > one in a _branch_.\n>\n> There may be multiple commits with the same change ID in a repo if you\n> had cherry-picked the commit, depending on what we decide to do with\n> the change ID on cherry-pick. But cherry-picks are not very common\n> anyway. Or maybe it's common in some workflow?\n\nIt's quite common for projects with multiple supported LTS releases.\nImportant fixes are backported to each LTS branch, and people\ntypically check that the fix has been backported by looking for the\nsame change id on other branches.\n\nIn such projects, 'git switch <change-id>' feels non-sensical to me;\nwhich particular copy of that change-id from which branch would that\nswitch you to?\n\nOr were the dozens of projects I was using that were hosted in Gerrit\nabusing change-ids in your view?  If so, how would you get them to\nstop?\n\nIf you don't have multiple LTS releases to support, or avoid\nbackporting fixes, then I could see how you'd arrive at the conclusion\nthat duplicate change-ids were rare, and then you could start using\nchange-ids interchangably with commit identifiers in commands so long\nas you only tried it on such projects.  But I'm a bit uncomfortable\nsuggesting we just assume all projects fit that mold.  Am I still\nmissing something here?\n"},{"id":"515624","messageId":"CABPp-BECTrVp9X6bVmzU8LEeYsC3KbzeJvAaDPN+FgZz_uEhmA@mail.gmail.com","threadId":"63245","inReplyTo":"D8X5I3W7K1DI.2JYHGNY9L7ZD3@buenzli.dev","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2025-04-04T02:28:39Z","receivedAt":"2025-04-04T02:28:51Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Thu, Apr 3, 2025 at 9:40 AM Remo Senekowitsch <remo@buenzli.dev> wrote:\n>\n> On Thu Apr 3, 2025 at 5:39 PM CEST, Elijah Newren wrote:\n> > On Wed, Apr 2, 2025 at 11:48 AM Martin von Zweigbergk\n> > <martinvonz@google.com> wrote:\n> >>\n> >> There are many benefits to having a change id even if it's just\n> >> local. I mentioned some in my email to this mailing list in [1].\n> >> For example, it enables\n> >> `git rebase main <change ID>; git switch <change ID>` without\n> >> requiring the user to look up the hash of the rewritten commit.\n> >\n> > But <change ID> isn't unique, right?  The whole point of having the\n> > change ID is to preserve it despite edits (e.g. rebase, commit\n> > --amend, cherry-pick), meaning that you end up with multiple commits\n> > with the same <change ID>.\n> >\n> > Why would this work?\n> >\n> > And if it does work, isn't it expensive since you'd need to walk\n> > history to find it?  Or do you keep an extra lookup table on the side\n> > somewhere?\n>\n> For rebase and commit --amend, the way Jujutsu deals with those is that\n> all descendants are immediately rebased on top of the new commit, and\n> refs to those descendants are updated as well. That means, the old\n> version of the patch with the same change-id becomes unreachable. So,\n> at least most of the time, the change-id is indeed unique.\n>\n> This doesn't work for cherry-pick, more on that below.\n>\n> Some of these features are not in Git yet, at least not to my knowledge.\n> That means getting the full benefit of change-ids with Git itself\n> would indeed require some more work. I know of rebase.updateRefs\n> and rebase.rebaseMerges, which move the Git experience closer to\n> Jujutsu, but don't go all the way. AFAIK it's not possible with Git to\n> automatically rebase --update-refs all descendants of a commit that is\n> amended or rebased.\n\nCorrect; that doesn't exist currently.\n\n> Jujutsu does keep a separate index of change-ids, yes.\n\nThanks.\n\n> >> There is a design doc [2] about the impact on Gerrit and how to\n> >> handle various cases where the client doesn't understand the\n> >> `change-id` header. That also includes some discussion about\n> >> whether cherry-picking should preserve the change id or create a\n> >> new one. I think there is a lot of value in having a\n> >> standardized header regardless of what we decide about\n> >> cherry-picks.\n> >\n> > cherry-pick & rebase preserve author name, email & time, while\n> > creating a new committer name, email, & time.  To me, the change-id is\n> > about the authorship, and since these commands already preserve\n> > authorship, it'd seem weird to me to have cherry-pick not preserve the\n> > change-id by default.\n>\n> I'd say Jujutsu, Gerrit and GitButler think of a change-id as associated\n> with a unit of review. (Although it will naturally support reviewing\n> sets of patches as well.) Usually only one person will push commits with\n> the same change-id, just like people don't usually force-push over each\n> others branches. But that's mostly about avoiding logistical problems.\n> When an employee leaves a company or is on vacation, it can be perfectly\n> reasonable for someone else to take over their work. In that case, it\n> would be appropriate to preserve the change-id, even though authorship\n> has changed, because the history of code review on that patch should\n> stay associated with the new version.\n>\n> Cherry-picking on the other hand often represents a separate unit of\n> review. That review may revolve around whether it makes sense to\n> backport a bugfix at all or any additional changes that may have been\n> necessary to make the bugfix work in the different, older codebase.\n\nI've worked with many projects hosted in Gerrit, and they all had a\nvery different view of change-ids than what you've espoused here.\nThey cherry-picked changes to other branches, fully expecting the\nchange-id to be kept the same.  They often checked to verify that\nimportant fixes had been backported to all the relevant LTS branches\nby looking for the change-id.  So, we'd typically have N+1 commits\nsharing the same change-id, all reachable from existing branches,\nwhere N is the number of LTS versions still supported at the time (and\nthe +1 comes from the main branch development).\n\n> As mentioned above, there's also the issue that preserving the change-id\n> on cherry-pick likely results in duplicates. For Jujutsu, it would be\n> nice it this was avoided. But it's not infeasible to deal with that\n> either.\n>\n> For Gerrit, it would be important to be able to track a change across\n> cherry-picks somehow, since that is a feature they already have. If Git\n> decides to preserve the change-id on cherry-pick, there's no problem\n> for Gerrit. Alternatives include storing a separate cherry-picked-from\n> header or enabling the -x flag on cherry-pick by default.\n\nCherry-picked-from trailers can be nice when it exists, but much more\nfrequently than one would want it provides a dead-end.  People will\ncherry-pick a commit that was local-only, or only found in some\nsecurity-embargoed repository, and you'd end up with dead ends.  You\nalso occasionally get chains: E cherry-picked from D, which was\ncherry-picked from C, which was cherry-picked from B, etc.  And more\ncomplex structures are possible.  And maybe part of that chain was a\nlocal-only commit or some commit from a security-embargoed repository\nthat you don't have access to.  Then folks get to write scripts and\ntry to deduce relationships from those trailers (e.g. hey, these two\ncommits both claim they were cherry-picked from the same non-existent\ncommit, and this other commit was a cherry-pick of one of these two,\nso they're a representation of the same logical change on these\ndifferent LTS branches).  It makes it a hassle to try to determine\nwhich LTS branches have the appropriate fixes backported and applied.\nI've done it, but I thought this problem was logically the point of\nchange-ids as found in Gerrit, honestly (well, that and its byzantine\npush to refs/for/$BRANCH stuff so it could automagically determine\nwhich CR that your push was supposed to be correlated with instead of\njust letting you specify via a real refname in your push command).\nWhile I understand that having nearly-unique change-ids let you use\nchange-ids interchangably with commits, that seems like a questionable\nbenefit over being able to actually track which logical changes are\nthe same and have been applied to which LTS branches.  I fully realize\nfolks may disagree...but if we're suggesting commands like `git switch\n<change-id>` which can only possibly be meaningful if <change-id> is\nunique across all branches, then what are we supposed to do for the\nmany projects which use change-ids for LTS backport tracking?  What\ndoes `git switch <change-id>` (and any other command where you attempt\nto use a non-unique change-id in place of a unique commit identifier)\ndo for them?\n"},{"id":"515625","messageId":"CABPp-BFYoZ1cuUMJPhWhtgntS0D-E=ZF+8_KS7gC+ShXjTrEDg@mail.gmail.com","threadId":"63245","inReplyTo":"CABPp-BECTrVp9X6bVmzU8LEeYsC3KbzeJvAaDPN+FgZz_uEhmA@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2025-04-04T02:40:36Z","receivedAt":"2025-04-04T02:40:48Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Thu, Apr 3, 2025 at 7:28 PM Elijah Newren <newren@gmail.com> wrote:\n>\n> On Thu, Apr 3, 2025 at 9:40 AM Remo Senekowitsch <remo@buenzli.dev> wrote:\n> >\n> > On Thu Apr 3, 2025 at 5:39 PM CEST, Elijah Newren wrote:\n> > > On Wed, Apr 2, 2025 at 11:48 AM Martin von Zweigbergk\n> > > <martinvonz@google.com> wrote:\n> > >>\n> > >> There are many benefits to having a change id even if it's just\n> > >> local. I mentioned some in my email to this mailing list in [1].\n> > >> For example, it enables\n> > >> `git rebase main <change ID>; git switch <change ID>` without\n> > >> requiring the user to look up the hash of the rewritten commit.\n> > >\n> > > But <change ID> isn't unique, right?  The whole point of having the\n> > > change ID is to preserve it despite edits (e.g. rebase, commit\n> > > --amend, cherry-pick), meaning that you end up with multiple commits\n> > > with the same <change ID>.\n> > >\n> > > Why would this work?\n> > >\n> > > And if it does work, isn't it expensive since you'd need to walk\n> > > history to find it?  Or do you keep an extra lookup table on the side\n> > > somewhere?\n> >\n> > For rebase and commit --amend, the way Jujutsu deals with those is that\n> > all descendants are immediately rebased on top of the new commit, and\n> > refs to those descendants are updated as well. That means, the old\n> > version of the patch with the same change-id becomes unreachable. So,\n> > at least most of the time, the change-id is indeed unique.\n> >\n> > This doesn't work for cherry-pick, more on that below.\n> >\n> > Some of these features are not in Git yet, at least not to my knowledge.\n> > That means getting the full benefit of change-ids with Git itself\n> > would indeed require some more work. I know of rebase.updateRefs\n> > and rebase.rebaseMerges, which move the Git experience closer to\n> > Jujutsu, but don't go all the way. AFAIK it's not possible with Git to\n> > automatically rebase --update-refs all descendants of a commit that is\n> > amended or rebased.\n>\n> Correct; that doesn't exist currently.\n>\n> > Jujutsu does keep a separate index of change-ids, yes.\n>\n> Thanks.\n>\n> > >> There is a design doc [2] about the impact on Gerrit and how to\n> > >> handle various cases where the client doesn't understand the\n> > >> `change-id` header. That also includes some discussion about\n> > >> whether cherry-picking should preserve the change id or create a\n> > >> new one. I think there is a lot of value in having a\n> > >> standardized header regardless of what we decide about\n> > >> cherry-picks.\n> > >\n> > > cherry-pick & rebase preserve author name, email & time, while\n> > > creating a new committer name, email, & time.  To me, the change-id is\n> > > about the authorship, and since these commands already preserve\n> > > authorship, it'd seem weird to me to have cherry-pick not preserve the\n> > > change-id by default.\n> >\n> > I'd say Jujutsu, Gerrit and GitButler think of a change-id as associated\n> > with a unit of review. (Although it will naturally support reviewing\n> > sets of patches as well.) Usually only one person will push commits with\n> > the same change-id, just like people don't usually force-push over each\n> > others branches. But that's mostly about avoiding logistical problems.\n> > When an employee leaves a company or is on vacation, it can be perfectly\n> > reasonable for someone else to take over their work. In that case, it\n> > would be appropriate to preserve the change-id, even though authorship\n> > has changed, because the history of code review on that patch should\n> > stay associated with the new version.\n> >\n> > Cherry-picking on the other hand often represents a separate unit of\n> > review. That review may revolve around whether it makes sense to\n> > backport a bugfix at all or any additional changes that may have been\n> > necessary to make the bugfix work in the different, older codebase.\n>\n> I've worked with many projects hosted in Gerrit, and they all had a\n> very different view of change-ids than what you've espoused here.\n> They cherry-picked changes to other branches, fully expecting the\n> change-id to be kept the same.  They often checked to verify that\n> important fixes had been backported to all the relevant LTS branches\n> by looking for the change-id.  So, we'd typically have N+1 commits\n> sharing the same change-id, all reachable from existing branches,\n> where N is the number of LTS versions still supported at the time (and\n> the +1 comes from the main branch development).\n>\n> > As mentioned above, there's also the issue that preserving the change-id\n> > on cherry-pick likely results in duplicates. For Jujutsu, it would be\n> > nice it this was avoided. But it's not infeasible to deal with that\n> > either.\n> >\n> > For Gerrit, it would be important to be able to track a change across\n> > cherry-picks somehow, since that is a feature they already have. If Git\n> > decides to preserve the change-id on cherry-pick, there's no problem\n> > for Gerrit. Alternatives include storing a separate cherry-picked-from\n> > header or enabling the -x flag on cherry-pick by default.\n>\n> Cherry-picked-from trailers can be nice when it exists, but much more\n> frequently than one would want it provides a dead-end.  People will\n> cherry-pick a commit that was local-only, or only found in some\n> security-embargoed repository, and you'd end up with dead ends.  You\n> also occasionally get chains: E cherry-picked from D, which was\n> cherry-picked from C, which was cherry-picked from B, etc.  And more\n> complex structures are possible.  And maybe part of that chain was a\n> local-only commit or some commit from a security-embargoed repository\n> that you don't have access to.  Then folks get to write scripts and\n> try to deduce relationships from those trailers (e.g. hey, these two\n> commits both claim they were cherry-picked from the same non-existent\n> commit, and this other commit was a cherry-pick of one of these two,\n> so they're a representation of the same logical change on these\n> different LTS branches).  It makes it a hassle to try to determine\n> which LTS branches have the appropriate fixes backported and applied.\n> I've done it, but I thought this problem was logically the point of\n> change-ids as found in Gerrit, honestly (well, that and its byzantine\n> push to refs/for/$BRANCH stuff so it could automagically determine\n> which CR that your push was supposed to be correlated with instead of\n> just letting you specify via a real refname in your push command).\n> While I understand that having nearly-unique change-ids let you use\n> change-ids interchangably with commits, that seems like a questionable\n> benefit over being able to actually track which logical changes are\n> the same and have been applied to which LTS branches.  I fully realize\n> folks may disagree...but if we're suggesting commands like `git switch\n> <change-id>` which can only possibly be meaningful if <change-id> is\n> unique across all branches, then what are we supposed to do for the\n> many projects which use change-ids for LTS backport tracking?  What\n> does `git switch <change-id>` (and any other command where you attempt\n> to use a non-unique change-id in place of a unique commit identifier)\n> do for them?\n\nOne possible simple solution here is just to treat change-ids (or\nthere abbreviations) kind of like abbreviated hashes -- they aren't\nguaranteed to be unique.  If the user specifies a change-id and there\nare multiple branches with such a change-id, we provide the user an\nerror much like we do for abbreviated hashes.\n\nIs that what folks have in mind?  If so, I'll be happy to drop my\nreservations about this aspect.\n"},{"id":"515626","messageId":"Z+9N2REkYZhrbkzb@ubby","threadId":"63245","inReplyTo":"CAESOdVAWWP=Rte4bx3zUZc6p0XiZaJS2OZr8ezRPkfq8K1TYfw@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-04T03:11:21Z","receivedAt":"2025-04-04T03:28:08Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Thu, Apr 03, 2025 at 03:47:30PM -0700, Martin von Zweigbergk wrote:\n> I think part of the problem is that I didn't consider that Git doesn't\n> really like to work in detached HEAD mode and doesn't automatically\n> update refs pointing to rewritten commits. This does take away a lot\n> of the usefulness, unfortunately. It would still be a bit useful as an\n> argument to readonly commands.\n\nI work in detached HEAD mode almost all the time.  And yeah, Git won't\nupdate refs when in detached HEAD mode because... what ref should it\nupdate?  The whole point of detached HEAD mode is that it does not.\n\n> > What would `git switch <change ID>` do?  `git switch` switches between\n> > branches, but a change ID can't possibly identify a branch since many\n> > commits could exist with the same change ID all in different branches.\n> \n> Yes, the same change id *can* exist on many branches, but it's pretty\n> uncommon. It might happen after cherry-picking, depending on what we\n\nThe whole point of change IDs for me is that if I need to backport bug\nfixes [0] then I can identify the bug fixes by change ID and then\ncherry-pick them onto support branches, which means that yes, there will\nbe many commits with the same change ID, each on different branches.\n\nBesides backports another use case that leaves multiple commits with the\nsame change ID is when I'm working on multiple different approaches to\nimplementing some feature on a complex codebase.  I might have two or\nthree branches exploring different ways to implement some feature, and\nof course I would want to use the same change IDs for similar commits\neven if I didn't use cherry-pick to create all but the first.\n\nHere's a third use-case that's quite real for me at $WORK where I have\nrepos for Debian-style packaging and use different branches for\ndifferent OS releases.  Some such pkgs are ones that we patch locally,\nso each release-specific branch targets a different OS release, but they\nall mostly have the same contents, in which case this is similar to\nbackports.  Other pkgs are locally developed and rarely differ on\ndifferent OS releases but occasionally they have to for <reasons> --\nthis is not similar to backports.\n\n[0] Granted, the industry has moved away from backports.  But many open\nsource projects still backport security fixes.  E.g., OpenSSL,\nPostgreSQL, etc.  So this is relevant.\n\n> decide there, but it should very rarely happen in other cases. When\n> you rewrite a commit, the old commit usually becomes unreachable, so\n> if your change id resolved to one commit before the rewrite, then it\n> will resolve to one commit after the rewrite. [...]\n\nNo matter what you will not have a guarantee of one commit for any given\nchange ID.  Therefore `git switch $change_id` is not workable, not\nwithout interaction, and even then still not workable because Git might\nhave to search _many_ branches to find commits matching the given change\nID.  (Fossil could have an index on change ID and trivially make that\nsearch possible, but for Git adding an index is more complicated.)\n\n> > I'd expect many commits with the same change ID in a _repo_, but at most\n> > one in a _branch_.\n> \n> There may be multiple commits with the same change ID in a repo if you\n> had cherry-picked the commit, depending on what we decide to do with\n> the change ID on cherry-pick. But cherry-picks are not very common\n> anyway. Or maybe it's common in some workflow? Oh, are you thinking of\n> a scenario where you cherry-pick your own commit to see if an\n> alternative approach is better? Sure, if cherry-pick preserve the\n> change ID, then you would have multiple commits with the same change\n> ID in that case.\n\nEven if cherry-picking were not common it's supported, therefore you\ncan't have change IDs be unique like this.  That doesn't make change IDs\nuseless -- on the contrary, their utility comes from the fact that they\nare not repo-wide unique!\n\nNico\n-- \n"},{"id":"515627","messageId":"CAESOdVB4yrDQ1v1BZtPiHDJwbaRVN6tixWg9eWNmBitXyqAh6w@mail.gmail.com","threadId":"63245","inReplyTo":"CABPp-BFYoZ1cuUMJPhWhtgntS0D-E=ZF+8_KS7gC+ShXjTrEDg@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Martin von Zweigbergk","fromEmail":"martinvonz@google.com","sentAt":"2025-04-04T03:47:37Z","receivedAt":"2025-04-04T03:47:50Z","isPatch":false,"sender":{"key":"martinvonz@google.com","avatar":"https://avatars.githubusercontent.com/u/891642?v=4"},"body":"On Thu, 3 Apr 2025 at 19:40, Elijah Newren <newren@gmail.com> wrote:\n>\n> One possible simple solution here is just to treat change-ids (or\n> there abbreviations) kind of like abbreviated hashes -- they aren't\n> guaranteed to be unique.  If the user specifies a change-id and there\n> are multiple branches with such a change-id, we provide the user an\n> error much like we do for abbreviated hashes.\n>\n> Is that what folks have in mind?  If so, I'll be happy to drop my\n> reservations about this aspect.\n\nYes, that's close to what we have in mind. I think I just didn't\nexplain clearly that it's mostly harmless in at least Jujutsu if there\nare multiple commits with the same change id. If there are multiple\nvisible commits with the same change id, then you'll just have to\ndecide what should happen when the user tries to refer to commits by\nchange id. We currently let it resolve to all the visible commits with\nthe given change id. We may change that to be an error instead [1].\nThe user can always fall back to using the commit id in such cases. We\ncall change ids with multiple visible commits \"divergent\". They\ncurrently show up in red in `jj log`, which I think we all agree makes\nthem seem unnecessarily scary. We'll probably change that soon [2]\n[3].\n\nSo when I said that I think it's quite uncommon to have multiple\ncommits with the same change id, I didn't mean that as an excuse to\nnot consider the other cases at all. I just mean that I think the vast\nmajority of commits are not cherry-picked, so we don't need to\noptimize the user experience for that case - it's fine if it's a bit\nmore complicated to refer to such commits.\n\nI hope that clarifies. Let me know if there are still unanswered\nquestions that I have missed.\n\n[1] https://github.com/jj-vcs/jj/issues/5632\n[2] https://github.com/jj-vcs/jj/pull/5800\n[3] https://github.com/jj-vcs/jj/pull/5850\n"},{"id":"515628","messageId":"Z+9aFffZJ1pP9QQK@ubby","threadId":"63245","inReplyTo":"CAESOdVB4yrDQ1v1BZtPiHDJwbaRVN6tixWg9eWNmBitXyqAh6w@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-04T04:03:33Z","receivedAt":"2025-04-04T04:03:43Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Thu, Apr 03, 2025 at 08:47:37PM -0700, Martin von Zweigbergk wrote:\n> Yes, that's close to what we have in mind. I think I just didn't\n> explain clearly that it's mostly harmless in at least Jujutsu if there\n> are multiple commits with the same change id. If there are multiple\n> visible commits with the same change id, then you'll just have to\n> decide what should happen when the user tries to refer to commits by\n> change id. We currently let it resolve to all the visible commits with\n> the given change id. We may change that to be an error instead [1].\n> The user can always fall back to using the commit id in such cases. We\n> call change ids with multiple visible commits \"divergent\". They\n> currently show up in red in `jj log`, which I think we all agree makes\n> them seem unnecessarily scary. We'll probably change that soon [2]\n> [3].\n\nTo support search by change ID you can use refs named something like\nrefs/change-IDs/<change-ID>, and the ref can either point to a singular\ncommit if the change ID is intended to be unique, or it can point to a\nroot commit which lists the commits with that change ID.\n\n> So when I said that I think it's quite uncommon to have multiple\n> commits with the same change id, I didn't mean that as an excuse to\n> not consider the other cases at all. I just mean that I think the vast\n> majority of commits are not cherry-picked, so we don't need to\n> optimize the user experience for that case - it's fine if it's a bit\n> more complicated to refer to such commits.\n\nFair.\n\nNico\n-- \n"},{"id":"515629","messageId":"CAESOdVCekFDxOWTTF71dpH1id_H2t9SaNo6buJ1MbvTnaENY7g@mail.gmail.com","threadId":"63245","inReplyTo":"Z+9N2REkYZhrbkzb@ubby","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Martin von Zweigbergk","fromEmail":"martinvonz@google.com","sentAt":"2025-04-04T04:08:59Z","receivedAt":"2025-04-04T04:09:12Z","isPatch":false,"sender":{"key":"martinvonz@google.com","avatar":"https://avatars.githubusercontent.com/u/891642?v=4"},"body":"On Thu, 3 Apr 2025 at 20:11, Nico Williams <nico@cryptonector.com> wrote:\n>\n> On Thu, Apr 03, 2025 at 03:47:30PM -0700, Martin von Zweigbergk wrote:\n> > I think part of the problem is that I didn't consider that Git doesn't\n> > really like to work in detached HEAD mode and doesn't automatically\n> > update refs pointing to rewritten commits. This does take away a lot\n> > of the usefulness, unfortunately. It would still be a bit useful as an\n> > argument to readonly commands.\n>\n> I work in detached HEAD mode almost all the time.  And yeah, Git won't\n> update refs when in detached HEAD mode because... what ref should it\n> update?  The whole point of detached HEAD mode is that it does not.\n\nJujutsu (and Mercurial) keep track of the set of visible heads. There\ncan be branches but they are not necessary. When you rewrite a commit,\nJutjutsu always rewrites all descendants. It also updates all branches\npointing to those commits automatically. For example, if you update\nthe description of some commit with `jj describe -m 'new description\n--revision xyz', then commit xyz and all its descendants will be\nupdated, and any branches pointing to any of the rewritten commits\nwill be updated.\n\nI think this is getting off topic. I can provide more detail to you in\na private message if necessary, or maybe we should start a new thread.\nI don't know what the convention on this list is. Or join the Discord\nchannel [1], or the #jujutsu channel on Libera.Chat.\n\n> > > What would `git switch <change ID>` do?  `git switch` switches between\n> > > branches, but a change ID can't possibly identify a branch since many\n> > > commits could exist with the same change ID all in different branches.\n> >\n> > Yes, the same change id *can* exist on many branches, but it's pretty\n> > uncommon. It might happen after cherry-picking, depending on what we\n>\n> The whole point of change IDs for me is that if I need to backport bug\n> fixes [0] then I can identify the bug fixes by change ID and then\n> cherry-pick them onto support branches, which means that yes, there will\n> be many commits with the same change ID, each on different branches.\n\nI agree that that can be one reason to use change IDs but I disagree\nthat it's the only reason. Having a stable way to refer to an evolving\ncommit is also important.\n\n> Besides backports another use case that leaves multiple commits with the\n> same change ID is when I'm working on multiple different approaches to\n> implementing some feature on a complex codebase.  I might have two or\n> three branches exploring different ways to implement some feature, and\n> of course I would want to use the same change IDs for similar commits\n> even if I didn't use cherry-pick to create all but the first.\n\nThis sounds more like the idea of \"topics\". It's an interesting\ndiscussion, but it seems off topic for this thread.\n\n> and even then still not workable because Git might\n> have to search _many_ branches to find commits matching the given change\n> ID.  (Fossil could have an index on change ID and trivially make that\n> search possible, but for Git adding an index is more complicated.)\n\nYes, I understand that it would be significant work to add support in\nGit. I hope that Git can gain the feature eventually, but we have no\nexpectation that it will be implemented soon, especially not the UX\npart (the preservation-on-rewrite part should be simpler, I think).\n\n[1] https://discord.gg/dkmfj3aGQN\n"},{"id":"515630","messageId":"Z+9XfLtodUOvqnse@ubby","threadId":"63245","inReplyTo":"D8XEYQ9TRB10.L5S89IAC2LZ9@buenzli.dev","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-04T03:52:28Z","receivedAt":"2025-04-04T04:09:32Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Fri, Apr 04, 2025 at 02:05:13AM +0200, Remo Senekowitsch wrote:\n> On Fri Apr 4, 2025 at 12:07 AM CEST, Nico Williams wrote:\n> > If I cherry-pick a commit then I absolutely want its \"change ID\" to be\n> > preserved by default.  If I want to drop that I can always ask for that\n> > or amend the commit to remove it.  I will want the same behavior for\n> > rebase and cherry-pick.  Having to remember different defaults and\n> > options for the two would be a cognitive load I do not need.\n> \n> Yeah, that's a very valid argument. In Jujutsus CLI, there is a very\n> clear separation between \"rebase\" and \"duplicate\", so there's no risk\n> of confusion if one preserves the change-id and the other doesn't.\n> [...]\n\n\"no risk of confusion\"?  It's higher cognitive load for the user, and\nhigher cognitive load does lead to confusion.\n\nI don't understand whence the need to introduce a new name for an old\nfeature.  Even Fossil -which lacks rebase- did not introduce a new name\nfor cherry-pick.  Given cherry-pick one can always implement rebase, but\nFossil's devs hate rebase workflows and love merge workflows, so they\nwon't implement or accept rebase.  But still, they did not think up a\nnew name for cherry-pick just to be different.\n\nMercurial also went through a phase of \"we hate rebase\", then later they\nadded: light-weight branches (i.e., Git-style branching), rebase, and\nhistedit (because mixing history editing with rebase is scary please\nno!!), and the end result was just more unnecessary cognitive load.\n\nI hope the next better-VCS project has more empathy for its users.\n\n> Making them behave the same way can be seen as simpler.\n\nYes: because it reduces users' cognitive load!\n\nHere's an example: Git's rebase, cherry-pick, and am commands have the\nsame --abort, --continue, and --skip options because they all do similar\nthings.  This means I only ever needed to learn that once for one\nsubcommand and then that knowledge carried over to the others.  For all\nthe UI hate Git gets it gets things like that right, and I love that.\n\n> >>                 [...]. The ways rebase and cherry-pick are most often\n> >> used are semantically very different from each other. (interactive)\n> >\n> > How do you know this is \"most often\" so? [...]\n> \n> I haven't conducted a study, this is my impression from talking to peers\n> and reading chatter from other Git users online. Maybe the impression is\n> wrong.\n\nRebase is notionally built out of cherry-pick, therefore they are\nsemantically similar even if users don't notice it.\n\nBe careful with \"chatter\".  You might only be hearing from merge-happy\nusers who don't use rebase and rarely use cherry-pick.  Many users\nprefer merge-based workflows because that elides the rebase/cherry-pick\ncognitive load.  I won't deny that rebasing requires more thinking than\nmerging, but leaving behind useful history always requires more thinking\nthan merging.  I recommend watching huge projects like OpenSSL,\nPostgreSQL, or Illumos to get a good idea of how power users use Git.\n\nFor example, Illumos uses the same policy that Sun used to: strictly\nlinear history only in the upstream master branch, with no merge commits\never.  You can have topic branches, naturally, and release branches too,\nbut in each branch linear history only, which means you have to rebase\nbefore pushing to upstream branches.  If you've never run into users who\nuse or have to use such rebase-heavy workflows then you might reach the\nwrong conclusions about what is common.  But at any rate in this case\nthe fundamentals make it clear that you cannot have just one commit with\na given change ID and you should not have different defaults for copying\nchange IDs for rebase vs. cherry-pick vs. am.\n\n> Yeah, I was (over-)simplifying. rebase is the swiss-army knife of git\n> commands. But for all of these operations, it holds that the previous\n> version of the patch(es) won't be reachable in the commit tree anymore\n> after the rebase is complete. (assuming potential descendant branches\n> are also rebased, which is usually the case) So rebase doesn't generally\n> cause duplicate change-ids, which is what I wanted to get at.\n\n> it holds\n\nNot so.  For example, when forward-porting our local patches to\n$external_open_source_project from 1.2.3 to 1.3.4 I do the following:\n\n: ; git checkout 1.2.3-patched\n: ; git checkout -b 1.3.4-patched\n: ; git rebase --onto 1.3.4 1.2.3\n: ; <address merge conflicts...>\n: ;\n: ; # Now branch 1.3.4-patched has the local patches from 1.2.3-patched\n: ; # but is based on 1.3.4.\n\nand now I'll have multiple commits with the same change IDs.\n\n> > When would you not want to preserve a change ID on cherry-pick?  I can't\n> > say I would ever have wanted to do that had Git had change IDs from day\n> > 1, and I've been using Git for more than twenty years.\n> \n> That's not exactly how Jujutsu thinks about the change-id, but it's a\n> useful piece of information. Gerrit does indeed use its change-id to\n> track cherry-picks. I am in favor of measures to track that metadata\n> (although duplicating change-ids is not my preferred option for that).\n\nSay you insisted on adding a prefix or suffix to the change ID when\n\"duplicating\" commits, how would you have Git enforce repo-wide\nuniqueness of change IDs?  The only tool Git has for this is refs, so\nyou'd have to create a ref for each change ID that points to the commit\nwith that change ID.  But now you've defeated the whole point of change\nIDs beyond code review, so I would insist on a multitude of types of\nchange ID so that I could have one that lets me have more than one\ncommit with the same change ID.\n\n> Let's assume the change-id represents the origin of a patch. What should\n> happen if a patch is split in two? Should they have the same change-id,\n> because they ultimately have the same origin? Maybe.\n\nI, the author, get to decide whether a) they both keep the same change\nID, or b) one of them gets a new change ID, or c) both get new change\nIDs, and/or maybe even d) with support for multiple change IDs I can\ntrack that both came from the same original commit but also they now\nhave different additional change IDs.  If it's just a header I can do\nthis by convention.  If the VCS was going to implement global uniqueness\nfor change IDs then my life would get more complicated in this case and\nI would not appreciate it.\n\n> I don't attach too much semantic meaning to the change-id. It's a\n> normally unique identifier for a change that persists as the change\n> evolves. That's useful. The more commits with the same change-id as\n> others there are, the less useful the concept becomes.\n\nBut it's perfect for all the other use-cases I mentioned, such as\nbackporting and forward-porting.  Those are the use-cases that most\nwould benefit from change IDs.  But even for the code-review-only case\nthe fact that you could look at a commit and use it to find\ncorresponding code review(s) is nice, even if you've had to cherry-pick\na commit for backports or for forward-porting.\n\n> >> doesn't preserve the change-id for that reason. So if cherry-pick\n> >\n> > I have _never_ used cherry-pick to cause there to be duplicate commits\n> > in the same branch.  Therefore calling it \"duplicate\" seems terribly\n> > wrong to me.\n> \n> Well, obviously not in the same branch. I meant duplicate among all\n> visible commits (reachable from any branch). That's the issue we're\n> discussing w.r.t. change-ids not always being unique identifiers for\n> a single commit. What would you like me to call that siuation instead\n> of duplicate?\n\nCherry-pick.  Because that's the name we already have.\n\n> Can you maybe give some examples of how you use cherry-pick? I'd be\n> interested in your use cases to maybe better understand where you're\n> coming from. [...]\n\n - [take over someone else's work and] decide to pick some of their\n   commits and drop others, but maybe do it in a new branch because the\n   old branch is still useful due to dropped commits still being useful\n   history (e.g., in case they are ever needed in the future)\n\n - fetch a branch from one upstream and pick selected commits onto\n   another branch that normally tracks a different upstream (I keep\n   local commits always \"on top\" of the upstream, rebasing as needed)\n\n - backports\n\n - forward-porting (this is arguably symmetric with backporting)\n\n - maintaining multiple related but different branches when researching\n   different ways to implement some feature\n\n>       [...]. I myself almost never use cherry-pick, simply because I'm\n> not involved in any backporting. I've seen cherry-pick used to get a\n> bugfix from another branch onto your own, in order to avoid having to\n> wait for the other branch to be merged. But that practice has always\n> rubbed me the wrong way. [...]\n\nYou may not have worked in sufficiently complex environments/projects.\n\nIn my world cherry-pick is an essential tool I cannot do without.  Where\na VCS was forced upon me that did not implement cherry-pick I've simply\nused patch(1) to apply diffs from the commit I wanted to cherry-pick --\nit's precisely because this is possible _and_ necessary that the VCS\nmight as well provide it.  Given cherry-pick then rebase follows, ergo\nthe VCS might as well also provide rebase.\n\n>                   [...]. I feel like the correct thing to do in that\n> situation is to extract the bugfix to a separate dependency-free branch\n> and make the two feature branches depend on it. That way, both feature\n\n\"extract the bugfix to a separate [...] branch\" -- that's exactly what\ncherry-pick does!\n\n> branches can more easily track changes in the bugfix by rebasing. If the\n> bugfix was cherry-picked, it's much harder to keep the two versions in\n> sync. (And finally, the latter approach probably makes the bugfix land\n> faster.) So yeah, interested to hear your use-cases for cherry-pick.\n\nI don't see how using a built-in cherry-pick feature vs manually\n\"extracting\" a commit makes it easier to \"keep the two versions in\nsync\".  On the contrary, manual operations always involve more cogntive\nload that automated ones, and at any rate the thing that woul dhelp you\n\"keep the two versions in sync\" is.. change IDs!\n\nNico\n-- \n"},{"id":"515631","messageId":"Z+9ez7kbh/L0Iq4k@ubby","threadId":"63245","inReplyTo":"CAESOdVCekFDxOWTTF71dpH1id_H2t9SaNo6buJ1MbvTnaENY7g@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-04T04:23:43Z","receivedAt":"2025-04-04T04:39:35Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Thu, Apr 03, 2025 at 09:08:59PM -0700, Martin von Zweigbergk wrote:\n> Jujutsu (and Mercurial) keep track of the set of visible heads. There\n> can be branches but they are not necessary. When you rewrite a commit,\n> Jutjutsu always rewrites all descendants. It also updates all branches\n> pointing to those commits automatically. For example, if you update\n> the description of some commit with `jj describe -m 'new description\n> --revision xyz', then commit xyz and all its descendants will be\n> updated, and any branches pointing to any of the rewritten commits\n> will be updated.\n\nI think I would find that too patronising, probably unbearable.\n\n> > and even then still not workable because Git might\n> > have to search _many_ branches to find commits matching the given change\n> > ID.  (Fossil could have an index on change ID and trivially make that\n> > search possible, but for Git adding an index is more complicated.)\n> \n> Yes, I understand that it would be significant work to add support in\n> Git. I hope that Git can gain the feature eventually, but we have no\n> expectation that it will be implemented soon, especially not the UX\n> part (the preservation-on-rewrite part should be simpler, I think).\n\nAh, well, Git does have an index: refs.  You could use\nrefs/change-IDs/<change-ID> to index by change ID.\n"},{"id":"515632","messageId":"CABPp-BHWFaUHAXwuddNpD1w=Fe7BK=9-Bc=-b9yXbqqWsQ8_pw@mail.gmail.com","threadId":"63245","inReplyTo":"CAESOdVB4yrDQ1v1BZtPiHDJwbaRVN6tixWg9eWNmBitXyqAh6w@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2025-04-04T04:59:53Z","receivedAt":"2025-04-04T05:00:05Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Thu, Apr 3, 2025 at 8:47 PM Martin von Zweigbergk\n<martinvonz@google.com> wrote:\n>\n> On Thu, 3 Apr 2025 at 19:40, Elijah Newren <newren@gmail.com> wrote:\n> >\n> > One possible simple solution here is just to treat change-ids (or\n> > there abbreviations) kind of like abbreviated hashes -- they aren't\n> > guaranteed to be unique.  If the user specifies a change-id and there\n> > are multiple branches with such a change-id, we provide the user an\n> > error much like we do for abbreviated hashes.\n> >\n> > Is that what folks have in mind?  If so, I'll be happy to drop my\n> > reservations about this aspect.\n>\n> Yes, that's close to what we have in mind. I think I just didn't\n> explain clearly that it's mostly harmless in at least Jujutsu if there\n> are multiple commits with the same change id. If there are multiple\n> visible commits with the same change id, then you'll just have to\n> decide what should happen when the user tries to refer to commits by\n> change id. We currently let it resolve to all the visible commits with\n> the given change id.\n\nresolve to all visible commits?  So the Jujutsu equivalent of 'git\nswitch <change-id>' would simultaneously check out N different\nbranches?  Or do commands which cannot accept multiple commits just\nthrow an error in such a case?\n\nDoing a \"git log --no-walk <change-id>\" and have it resolve to several\ncommits would be kinda cool...\n\n> We may change that to be an error instead [1].\n> The user can always fall back to using the commit id in such cases. We\n> call change ids with multiple visible commits \"divergent\". They\n> currently show up in red in `jj log`, which I think we all agree makes\n> them seem unnecessarily scary. We'll probably change that soon [2]\n> [3].\n>\n> So when I said that I think it's quite uncommon to have multiple\n> commits with the same change id, I didn't mean that as an excuse to\n> not consider the other cases at all. I just mean that I think the vast\n> majority of commits are not cherry-picked, so we don't need to\n> optimize the user experience for that case - it's fine if it's a bit\n> more complicated to refer to such commits.\n>\n> I hope that clarifies. Let me know if there are still unanswered\n> questions that I have missed.\n\nYep, thanks, that answered all my previous questions...though you\nraised one new one that I mentioned above.\n\n> [1] https://github.com/jj-vcs/jj/issues/5632\n> [2] https://github.com/jj-vcs/jj/pull/5800\n> [3] https://github.com/jj-vcs/jj/pull/5850\n"},{"id":"515634","messageId":"CAESOdVArh6Vksd9bktBz4DBqOzvoydfh6_DZcm2t9kJ5F-s1EQ@mail.gmail.com","threadId":"63245","inReplyTo":"CABPp-BHWFaUHAXwuddNpD1w=Fe7BK=9-Bc=-b9yXbqqWsQ8_pw@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Martin von Zweigbergk","fromEmail":"martinvonz@google.com","sentAt":"2025-04-04T05:21:57Z","receivedAt":"2025-04-04T05:22:11Z","isPatch":false,"sender":{"key":"martinvonz@google.com","avatar":"https://avatars.githubusercontent.com/u/891642?v=4"},"body":"On Thu, 3 Apr 2025 at 22:00, Elijah Newren <newren@gmail.com> wrote:\n>\n> On Thu, Apr 3, 2025 at 8:47 PM Martin von Zweigbergk\n> <martinvonz@google.com> wrote:\n> >\n> > On Thu, 3 Apr 2025 at 19:40, Elijah Newren <newren@gmail.com> wrote:\n> > >\n> > > One possible simple solution here is just to treat change-ids (or\n> > > there abbreviations) kind of like abbreviated hashes -- they aren't\n> > > guaranteed to be unique.  If the user specifies a change-id and there\n> > > are multiple branches with such a change-id, we provide the user an\n> > > error much like we do for abbreviated hashes.\n> > >\n> > > Is that what folks have in mind?  If so, I'll be happy to drop my\n> > > reservations about this aspect.\n> >\n> > Yes, that's close to what we have in mind. I think I just didn't\n> > explain clearly that it's mostly harmless in at least Jujutsu if there\n> > are multiple commits with the same change id. If there are multiple\n> > visible commits with the same change id, then you'll just have to\n> > decide what should happen when the user tries to refer to commits by\n> > change id. We currently let it resolve to all the visible commits with\n> > the given change id.\n>\n> resolve to all visible commits?  So the Jujutsu equivalent of 'git\n> switch <change-id>' would simultaneously check out N different\n> branches?  Or do commands which cannot accept multiple commits just\n> throw an error in such a case?\n\nYes, the latter.\n\nThe closest equivalent of `git switch` is `jj new` and that command\nactually does support multiple commits - that's how you create a merge\ncommit. But we also have a little safeguard in the form of requiring\nyou to say `jj new all:xyz` if you really want that. For commands\nwhere it's more expected and harmless to take multiple commits as\ninput, we don't require the `all:` prefix.\n\n> Doing a \"git log --no-walk <change-id>\" and have it resolve to several\n> commits would be kinda cool...\n\nYes, this is one of those commands where it's harmless so `jj log -r\nxyz` simply shows all (visible) commits with the given change id.\nOther examples include `jj abandon`, `jj describe`, `jj diff -r`, `jj\nrebase -r`, `jj squash --from`, while e.g. `jj rebase -d` (the\ndestination) requires the `all:` prefix to allow multiple parents\n(making the roots of the rebased set into merge commits).\n\n> Yep, thanks, that answered all my previous questions...though you\n> raised one new one that I mentioned above.\n\nYay!\n"},{"id":"515638","messageId":"D8XOO0H2GSK0.JQSNCLB1QUQR@buenzli.dev","threadId":"63245","inReplyTo":"Z+9XfLtodUOvqnse@ubby","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Remo Senekowitsch","fromEmail":"remo@buenzli.dev","sentAt":"2025-04-04T07:41:24Z","receivedAt":"2025-04-04T07:41:30Z","isPatch":false,"sender":{"key":"remo@buenzli.dev","avatar":"https://gravatar.com/avatar/7df680b096206886db5a2dc983926f314985bbee662eb406cf23c65322cf98b7?d=mp&s=160"},"body":"I'll try to join some of our threads and summarize... please correct me\nif you disagree with that summary (or I've left something out you think\nis important).\n\nOne of my points was that rebase usually doesn't lead to multiple\nvisible commits with the same change-id, whereas cherry-pick usually\ndoes. You pointed out that rebase can in fact lead to multiple visible\ncommits with the same change-id and cherry-pick sometimes doesn't (and\nsome of these examples represent valid use cases). So making these two\ncommands behave in a different way that only makes sense in the \"common\"\ncase makes them more complicated - and less consistent - in the general\ncase. I find this convincing and I now agree that both cherry-pick and\nrebase should preserve the change-id by default.\n\nThe Gerrit developers will be happy about this anyway, and it's worth\nrepeating that Jujutsu can deal with it perfectly fine while preserving\nalmost all benefits of change-ids.\n\nWhen discussing the uniqueness of change-ids or lack thereof, I'd\nlike to introduce one more factor: At which point in time during the\ndevelopment cycle a change-id is unique. Jujutsu users derive most\nof the benefits of change-ids during active development. The use case\nyou find most important - tracking (forward- and) back-ports - happens\nat a different time during development, when a change has been merged\nto a public branch already. So I think there is no conflict at all.\nChange-ids will naturally tend to be unique during active development\nand once they are merged, whether the change-id stays unique or not\ndoesn't matter anymore for those active development use cases. We can\nhave our cake and eat it too.\n\nRemo\n"},{"id":"515648","messageId":"Z--mbrqpUGOcGMVi@pks.im","threadId":"63245","inReplyTo":"CAESOdVArh6Vksd9bktBz4DBqOzvoydfh6_DZcm2t9kJ5F-s1EQ@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-04-04T09:29:18Z","receivedAt":"2025-04-04T09:29:23Z","isPatch":false,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Thu, Apr 03, 2025 at 10:21:57PM -0700, Martin von Zweigbergk wrote:\n> On Thu, 3 Apr 2025 at 22:00, Elijah Newren <newren@gmail.com> wrote:\n> >\n> > On Thu, Apr 3, 2025 at 8:47 PM Martin von Zweigbergk\n> > <martinvonz@google.com> wrote:\n> > >\n> > > On Thu, 3 Apr 2025 at 19:40, Elijah Newren <newren@gmail.com> wrote:\n> > > >\n> > > > One possible simple solution here is just to treat change-ids (or\n> > > > there abbreviations) kind of like abbreviated hashes -- they aren't\n> > > > guaranteed to be unique.  If the user specifies a change-id and there\n> > > > are multiple branches with such a change-id, we provide the user an\n> > > > error much like we do for abbreviated hashes.\n> > > >\n> > > > Is that what folks have in mind?  If so, I'll be happy to drop my\n> > > > reservations about this aspect.\n> > >\n> > > Yes, that's close to what we have in mind. I think I just didn't\n> > > explain clearly that it's mostly harmless in at least Jujutsu if there\n> > > are multiple commits with the same change id. If there are multiple\n> > > visible commits with the same change id, then you'll just have to\n> > > decide what should happen when the user tries to refer to commits by\n> > > change id. We currently let it resolve to all the visible commits with\n> > > the given change id.\n> >\n> > resolve to all visible commits?  So the Jujutsu equivalent of 'git\n> > switch <change-id>' would simultaneously check out N different\n> > branches?  Or do commands which cannot accept multiple commits just\n> > throw an error in such a case?\n> \n> Yes, the latter.\n\nThat sounds like sensible behaviour to me. We already know to print\nambiguous hashes, so we could probably do the same with change IDs:\n\n    $ git rev-parse 1234\n    error: short object ID 1234 is ambiguous\n    hint: The candidates are:\n    hint:   1234d8d9179 commit 2018-06-08 - Merge 'add-p-many-files'\n    hint:   1234e8297f3 commit 2020-10-26 - Merge branch 'en/sequencer-rollback-lock-cleanup' into next\n    hint:   123456ff3f0 tree\n    hint:   12349764e65 tree\n    hint:   1234b76f424 tree\n    hint:   1234c687269 tree\n    hint:   1234ebb5d8f blob\n    fatal: ambiguous argument '1234': unknown revision or path not in the working tree.\n    Use '--' to separate paths from revisions, like this:\n    'git <command> [<revision>...] -- [<file>...]'\n    1234\n\nIt should be rather easy to adapt this mechanism to also handle the case\nwhere the same change IDs (or abbreviated change IDs) exist across\nmultiple commits.\n\nPatrick\n"},{"id":"515649","messageId":"Z--nnOhsUCaqo45z@pks.im","threadId":"63245","inReplyTo":"Z+9ez7kbh/L0Iq4k@ubby","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-04-04T09:34:20Z","receivedAt":"2025-04-04T09:34:25Z","isPatch":false,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Thu, Apr 03, 2025 at 11:23:43PM -0500, Nico Williams wrote:\n> On Thu, Apr 03, 2025 at 09:08:59PM -0700, Martin von Zweigbergk wrote:\n> > > and even then still not workable because Git might\n> > > have to search _many_ branches to find commits matching the given change\n> > > ID.  (Fossil could have an index on change ID and trivially make that\n> > > search possible, but for Git adding an index is more complicated.)\n> > \n> > Yes, I understand that it would be significant work to add support in\n> > Git. I hope that Git can gain the feature eventually, but we have no\n> > expectation that it will be implemented soon, especially not the UX\n> > part (the preservation-on-rewrite part should be simpler, I think).\n> \n> Ah, well, Git does have an index: refs.  You could use\n> refs/change-IDs/<change-ID> to index by change ID.\n\nI don't think references are a good mechanism to track change IDs. The\nexpectation around refs is that users can change them basically at will,\nbut it certainly does not make any sense to let them update change IDs.\nFurthermore, as we have already discussed, change IDs are not unique,\nbut refs can only point to a single commit ID. So that's another\nmismatch that we cannot address.\n\nI think caching the information in an auxiliary data structure would\nthus be a lot more reasonable. This could for example be part of commit\ngraphs, but could also be a separate index specific to change IDs\nthemselves.\n\nPatirck\n"},{"id":"515650","messageId":"Z--pQ2ge47J_489a@pks.im","threadId":"63245","inReplyTo":"CABPp-BGwXaiohvfSdr96hzKNPYXQqz+_okxLNj7P9KSjX2PW6g@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-04-04T09:41:23Z","receivedAt":"2025-04-04T09:41:29Z","isPatch":false,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Thu, Apr 03, 2025 at 08:56:01AM -0700, Elijah Newren wrote:\n> On Thu, Apr 3, 2025 at 2:13 AM Patrick Steinhardt <ps@pks.im> wrote:\n> >\n> [...]\n> > Agreed, change IDs solve a couple of issues that many users face:\n> >\n> >   - You can reliably track how a patch evolves over time. This helps\n> >     various different tools to track identity of commits, like for\n> >     example forges, but also tools like git-range-diff(1).\n> >\n> >   - It becomes trivial to see whether a commit has been cherry-picked\n> >     into another branch. We do have git-cherry(1) to do that right now,\n> >     but that command is based on heuristics and fails as soon as the\n> >     patch itself needed to be adapted.\n> >\n> >   - Working with history rewrites becomes easier in the general case as\n> >     you don't have to adapt to constantly changing commit IDs.\n> \n> Could you elaborate?  I agree with the other points you raise, but I'm\n> unsure how this helps with a history rewrite.  Do you mean the\n> rewriting of history, or someone trying to consume the history\n> rewrite?  If the former, I don't see it, and if the latter, didn't you\n> already cover that in the two bullets above?  Or is there something\n> else you are also getting at?\n\nYeah, this point wasn't quite clear. It's mostly based around my own\nfindings that I always end up copying a lot of object IDs around while\nworking on rebases. And the most annoying part to me is that those OIDs\nalso change on every rewrite, and as a consequence I always have to look\nup the rewritten object IDs.\n\nBy using change IDs this issue would become easier as I only need to\nremember one set of constant IDs that don't change on ever rewrite.\n\n> > So what would it take to get change IDs into Git? I think the most\n> > important items would be:\n> >\n> >   - Generating and writing change IDs in commands that support them.\n> >     This includes e.g. git-commit(1), git-commit-tree(1), git-merge(1),\n> >     git-merge-tree(1). This should of course be completely optional and\n> >     probably be disabled by default.\n> >\n> >   - Making tools that rewrite commits aware of change IDs so that they\n> >     know to retain change IDs. This involves e.g. git-cherry-pick(1),\n> >     git-rebase(1), git-replay(1).\n> \n> And also git-commit(1) [when passing --amend], and git-fast-export(1)\n> and git-fast-import(1) -- though possibly with options for the last\n> two to expunge them instead of preserving them, but probably\n> defaulting to preserving them.\n> \n> However, I think some of these might already handle this.  Commands\n> which call read_commit_extra_headers() and pass those along to\n> commit_tree_extended() may already preserve these.  It appears commit\n> --amend and replay both do this.  sequencer has some code that looks\n> relevant, but it appears to only be reading the headers from HEAD (at\n> the time the nth commit is being replayed), which seems like it'd be\n> looking at the wrong commit.  That might actually be a bug...\n\nYes, some tools already handle this correctly indeed.\n\n> >   - Extending revisions to allow specifying commits by change ID.\n> \n> Would this essentially be similar to <rev>^{/<text>} except searching\n> specifically change-id headers rather than commit message?\n\nYeah, something like that. The exact format for such a new revision\nwould be up for debate, but I quite like the reverse hex format that JJ\nitself uses as it is unambiguous compared to object IDs. Only problem of\ncourse is that it's not unambiguous compared to refnames, so accepting\nreverse object IDS as-is without any kind of prefix is probably a no-go.\n\nPatrick\n"},{"id":"515679","messageId":"Z/AD4CsUZTPQyKeT@ubby","threadId":"63245","inReplyTo":"D8XOO0H2GSK0.JQSNCLB1QUQR@buenzli.dev","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-04T16:08:00Z","receivedAt":"2025-04-04T16:08:10Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Fri, Apr 04, 2025 at 09:41:24AM +0200, Remo Senekowitsch wrote:\n> I'll try to join some of our threads and summarize... please correct me\n> if you disagree with that summary (or I've left something out you think\n> is important).\n>\n> [...]\n\n+1.\n\n> When discussing the uniqueness of change-ids or lack thereof, I'd\n> like to introduce one more factor: At which point in time during the\n> development cycle a change-id is unique. Jujutsu users derive most\n> of the benefits of change-ids during active development. The use case\n> you find most important - tracking (forward- and) back-ports - happens\n> at a different time during development, when a change has been merged\n> to a public branch already. So I think there is no conflict at all.\n> Change-ids will naturally tend to be unique during active development\n> and once they are merged, whether the change-id stays unique or not\n> doesn't matter anymore for those active development use cases. We can\n> have our cake and eat it too.\n\nI think that's right.  One more benefit of change IDs is that you can\nfind a commit in a back-/forward-port and track it all the way to its\nvery initial integration and its code review history, perhaps all the\nway to its very first version.\n"},{"id":"515684","messageId":"Z/ADDp/mIzuOofYl@ubby","threadId":"63245","inReplyTo":"Z--nnOhsUCaqo45z@pks.im","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-04T16:04:30Z","receivedAt":"2025-04-04T18:28:08Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Fri, Apr 04, 2025 at 11:34:20AM +0200, Patrick Steinhardt wrote:\n> > Ah, well, Git does have an index: refs.  You could use\n> > refs/change-IDs/<change-ID> to index by change ID.\n> \n> I don't think references are a good mechanism to track change IDs. The\n> expectation around refs is that users can change them basically at will,\n> but it certainly does not make any sense to let them update change IDs.\n\ngit meta (a separate project) uses refs named refs/commits/$full_hash.\nUsers can always break their repos in a many many ways.  Using refs is\nfine.\n\n> Furthermore, as we have already discussed, change IDs are not unique,\n> but refs can only point to a single commit ID. So that's another\n> mismatch that we cannot address.\n\nOh I agree with that!  But I think I've made that pretty clear by now :]\n\nHowever one could have refs/change-IDs/<change-ID> ultimately point to\nan object that lists the commits with that change ID, and then that\nwould function as an index for non-unique change IDs.\n\n> I think caching the information in an auxiliary data structure would\n> thus be a lot more reasonable. This could for example be part of commit\n> graphs, but could also be a separate index specific to change IDs\n> themselves.\n\nSure.\n"},{"id":"515689","messageId":"20250405020918.GC3051250@mit.edu","threadId":"63245","inReplyTo":"D8XAETM3PJ4A.1RBPVHJWE9DFT@buenzli.dev","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Theodore Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2025-04-05T02:09:18Z","receivedAt":"2025-04-05T02:09:49Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Thu, Apr 03, 2025 at 10:31:08PM +0200, Remo Senekowitsch wrote:\n> > So as a cauaionary note, as people use Change ID's in Gerrit today,\n> > sometimes the Change ID changes between rebases, and I've certainly\n> > seen cases where the sematic meaning of the commit has changed\n> > significantly without changing the Change ID.  So it's great as a\n> > hint, but in practice, at least today, it might not be completely safe\n> > to assume the semantics are as advertised....\n> \n> That is all true. I would just say that some change-ids not being unique\n> doesn't systematically take away from the benefits. If a given change-id\n> is unique, you get all the benefits for that patch, independent of\n> the uniqueness of the other change-ids. If some patch changes a lot\n> semantically while keeping its change-id, that will degrade its review\n> history, but without affecting the review history of any other patch.\n\nOh, agreed.  I've certainly used a Change-ID by pasting it into the\nGerrit Search Box, and when I saw two or three unrelated commits where\nthe one-line subject were for entirely different kernel subsystem, as\na human I could do the disambiguation myself.  And if I did't find a\nparticular bug fix for the various release brances, I would cut and\npaste the one-line commit summary into the Gerrit Search Box, and find\nthe right commit, again aided by human intelligence.\n\nSo I'm super-glad that it's there, and it *is* useful.  But I just\nwanted to inject a note of caution if people were thinking about using\nit in some kind of automated tooling --- especilly if there is a\nproposal to inject this into git as a first-class object.\n\nCertainly from a git perspective, I will admit that more often than\nnot, I generally don't do something like \"git log --grep\nI23e218cd964f16c0b2b26127d4a5ca6529867673\", but rather, \"git log\n--grep 'mediatek: Don't modify spi_transfer'\".  If I have to cut and\npaste a Change-ID, most of the time I aso have access to the one-line\ncommit description, and grepping for something human readable is just\neasier.\n\n\t\t\t\t\t- Ted\n"},{"id":"515744","messageId":"Z_OGMb-1oV0Ex05e@pks.im","threadId":"63245","inReplyTo":"Z/ADDp/mIzuOofYl@ubby","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-04-07T08:00:49Z","receivedAt":"2025-04-07T08:00:54Z","isPatch":false,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Fri, Apr 04, 2025 at 11:04:30AM -0500, Nico Williams wrote:\n> On Fri, Apr 04, 2025 at 11:34:20AM +0200, Patrick Steinhardt wrote:\n> > > Ah, well, Git does have an index: refs.  You could use\n> > > refs/change-IDs/<change-ID> to index by change ID.\n> > \n> > I don't think references are a good mechanism to track change IDs. The\n> > expectation around refs is that users can change them basically at will,\n> > but it certainly does not make any sense to let them update change IDs.\n> \n> git meta (a separate project) uses refs named refs/commits/$full_hash.\n> Users can always break their repos in a many many ways.  Using refs is\n> fine.\n\nTrue, they can. But if we use refs as a storage mechanism this means\nthat a remote user can also mess with _your_ repository by introducing a\nbroken `refs/change-IDs/<change-ID>` ref that points to a commit with a\ndifferent change ID.\n\nYou can of course work around this by treating such references\nspecially. But that would introduce more complexity into a central\nsubsystem of Git, and the result would probably be another mechanism\nwith leaky abstractions and confusing behaviour.\n\n> > Furthermore, as we have already discussed, change IDs are not unique,\n> > but refs can only point to a single commit ID. So that's another\n> > mismatch that we cannot address.\n> \n> Oh I agree with that!  But I think I've made that pretty clear by now :]\n> \n> However one could have refs/change-IDs/<change-ID> ultimately point to\n> an object that lists the commits with that change ID, and then that\n> would function as an index for non-unique change IDs.\n\nThat would be possible indeed, but now it's getting even more awkward.\n\nSo I'm not yet convinced that refs are a good fit for storing cached\nchange IDs.\n\nPatrick\n"},{"id":"515802","messageId":"xmqq4iyzn0vn.fsf@gitster.g","threadId":"63245","inReplyTo":"CAESOdVAspxUJKGAA58i0tvks4ZOfoGf1Aa5gPr0FXzdcywqUUw@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-04-07T20:59:56Z","receivedAt":"2025-04-07T20:59:59Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Martin von Zweigbergk <martinvonz@google.com> writes:\n\n> For example, it enables\n> `git rebase main <change ID>; git switch <change ID>` without\n> requiring the user to look up the hash of the rewritten commit.\n\nI do not quite see why this can be listed even as an advantage,\nunless you are going to allow end users to name the changes, instead\nof using auto-generated impossible-to-remember hexadecimal string\n(perhaps prefixed with a single \"I\" or something).\n\n> If the change id also transferred between repos and preserved by\n> a forge (such as Gerrit), it enables the change id to be used to\n> identify a code review.\n\nPeople often talk about rebasing and rewriting in the context of\ndiscussing \"change IDs\", and for 80% of the use cases where a simple\nsingle-commit topic is involved, it would perfectly work fine.\nAfter making a new commit C0 on top of 'main', updating 'main' with\nothers' changes, and then rebasing that C0 on top of updated 'main'\nto produce C1, you would expect that C0 and C1 are moral equivalents\nso it is natural that you wish there is a name to give to these\nmoral equivalents.\n\nBut stepping back a bit, if they are not just moral equivalents but\nrecord identical changes that are so same that an earlier review of\nC0 makes it unnecessary to review C1, why are you even rebasing in\nthe first place?  Just merging C0 to the updated 'main' would retain\nthe earlier review made on C0 and things should merge just fine.\n\nI have more problems with the remaining 20% use case, where you need\nto deal with multiple commits.\n\nPerhaps your initial changeset is a single commit C0 that is so\nlarge and does too many things at once, and reviewers would\nnaturally advise you to split things up.  You'll come up with a\nseries of commits, C1_0 and C1_1.  The net effect of applying these\ntwo patches may be the same as applying the original C0, but each of\nthem is more cleanly separated to address one issue at a time, and\nthe explanation given in the proposed log message more clearly\ndescribes the issue each of them addresses.  Now you gave a change ID\nto C0, and want to somehow relate C1_0 and C1_1 to the original C0.\nWhich one gets the same change ID?  Earlier one?  The last one?\nBoth gets the same change ID?\n\nOr your initial changeset is a two-commit series, C0_0 and C0_1, but\nreviewers find that each one of them alone is not complete, and\nbecause the issue addressed by these two is small and isolated\nenough, you are advised to make them into a single commit C1.  Did\nyou start with two change IDs for these two original commits?  If\nso, whose change ID the updated commit C1 inherit?  Or does C1 have\ntwo change IDs now?  Or did you start with a single change ID\nassigned to both of these two original commits?\n\nQuite frankly, I think the concept of \"change ID\" is nice but it is\nnot mechanically trustable.  Recording them in the trailers is fine,\nbut I somehow feel that they have a clear-cut semantics everybody\ncan agree on to deserve to be in the header part of commit objects.\n\n"},{"id":"515804","messageId":"Z/RFQY433muaCW44@ubby","threadId":"63245","inReplyTo":"xmqq4iyzn0vn.fsf@gitster.g","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-07T21:36:01Z","receivedAt":"2025-04-07T21:36:11Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Mon, Apr 07, 2025 at 08:59:56PM +0000, Junio C Hamano wrote:\n> I have more problems with the remaining 20% use case, where you need\n> to deal with multiple commits.\n> \n> Perhaps your initial changeset is a single commit C0 that is so\n> large and does too many things at once, and reviewers would\n> naturally advise you to split things up.  You'll come up with a\n> series of commits, C1_0 and C1_1.  The net effect of applying these\n> two patches may be the same as applying the original C0, but each of\n> them is more cleanly separated to address one issue at a time, and\n> the explanation given in the proposed log message more clearly\n> describes the issue each of them addresses.  Now you gave a change ID\n> to C0, and want to somehow relate C1_0 and C1_1 to the original C0.\n> Which one gets the same change ID?  Earlier one?  The last one?\n> Both gets the same change ID?\n> \n> Or your initial changeset is a two-commit series, C0_0 and C0_1, but\n> reviewers find that each one of them alone is not complete, and\n> because the issue addressed by these two is small and isolated\n> enough, you are advised to make them into a single commit C1.  Did\n> you start with two change IDs for these two original commits?  If\n> so, whose change ID the updated commit C1 inherit?  Or does C1 have\n> two change IDs now?  Or did you start with a single change ID\n> assigned to both of these two original commits?\n\nThis is why I suggested earlier that there need to be multiple change\nIDs, not just one.  Perhaps one is a \"code review ID\" and another is\na \"commit change ID\".  The code review ID would let you link together\nall commits that were reviewed together, so if you have to split or\nsquash commits they would all still have that one code review ID.  The\ncommit change ID would be shared by all sufficiently-similar versions of\na commit.  If a commit is dropped or split or squashed then its commit\nID might get dropped too, but the code review ID would stay the same.\n\nI think that a code review ID is a lot saner to use as a ref, as in OP's\n`git switch ...` example.  Though one might still want symbolic tag-like\nnames for individual commits in a set of related commits (e.g., for\nforward- and back-ports), though I'm inclined to say that's not needed.\n\n> Quite frankly, I think the concept of \"change ID\" is nice but it is\n> not mechanically trustable.  Recording them in the trailers is fine,\n> but I somehow feel that they have a clear-cut semantics everybody\n> can agree on to deserve to be in the header part of commit objects.\n\nI don't think they need to have such extremely detailed semantics in\norder to be able to get a header.  The semantics will ultimately be\nsomewhat project-defined, typically something like \"during code review\nyou can use these to related newer updates to an MR/PR/CR to older\nversions\" and \"once integrated you can use these to find the approved\ncode review as follows [details]\".  The [details] (probably a URI\ntemplate) for finding concluded CRs might vary.  The CR tool might vary.\nThe construction of the change IDs might vary.  The intent might not\nvary at all.\n\nNico\n-- \n"},{"id":"515808","messageId":"D90RWL4FEBQA.1UNOR59T3U98R@buenzli.dev","threadId":"63245","inReplyTo":"xmqq4iyzn0vn.fsf@gitster.g","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Remo Senekowitsch","fromEmail":"remo@buenzli.dev","sentAt":"2025-04-07T22:51:39Z","receivedAt":"2025-04-07T22:51:45Z","isPatch":false,"sender":{"key":"remo@buenzli.dev","avatar":"https://gravatar.com/avatar/7df680b096206886db5a2dc983926f314985bbee662eb406cf23c65322cf98b7?d=mp&s=160"},"body":"On Mon Apr 7, 2025 at 10:59 PM CEST, Junio C Hamano wrote:\n> Martin von Zweigbergk <martinvonz@google.com> writes:\n>\n>> For example, it enables\n>> `git rebase main <change ID>; git switch <change ID>` without\n>> requiring the user to look up the hash of the rewritten commit.\n>\n> I do not quite see why this can be listed even as an advantage,\n> unless you are going to allow end users to name the changes, instead\n> of using auto-generated impossible-to-remember hexadecimal string\n> (perhaps prefixed with a single \"I\" or something).\n\nSince the change-id will use a \"reverse-hex\" alphabet (z-k instead of\n0-f), prefixing an \"I\" won't be necessary for disambiguation.\n\nIn the case of Jujutsu, there is a configurable \"immutable revset\",\nwhich represents the idea that people usually don't want to force-push\nover the master branch. If the user specifies a change-id prefix that's\nambiguous, but only one of the commits it could refer to is \"mutable\",\nthat one takes precedence. So in practice, the commits one is currently\nworking on can be identified with 1-3 characters, which fit comfortably\ninto short-term memory.\n\nWhile that feature isn't implemented in Git (yet), it shows how such a\nusage pattern can be very ergonomic.\n\n>> If the change id also transferred between repos and preserved by\n>> a forge (such as Gerrit), it enables the change id to be used to\n>> identify a code review.\n>\n> People often talk about rebasing and rewriting in the context of\n> discussing \"change IDs\", and for 80% of the use cases where a simple\n> single-commit topic is involved, it would perfectly work fine.\n> After making a new commit C0 on top of 'main', updating 'main' with\n> others' changes, and then rebasing that C0 on top of updated 'main'\n> to produce C1, you would expect that C0 and C1 are moral equivalents\n> so it is natural that you wish there is a name to give to these\n> moral equivalents.\n>\n> But stepping back a bit, if they are not just moral equivalents but\n> record identical changes that are so same that an earlier review of\n> C0 makes it unnecessary to review C1, why are you even rebasing in\n> the first place?  Just merging C0 to the updated 'main' would retain\n> the earlier review made on C0 and things should merge just fine.\n\nSome people (including me) like to / are used to rebasing often and\nkeeping a linear history on master as possible. I'm not aware that\nthere's anything wrong with that.\n\nIn practice, I don't think this will cause any problems. The change-id\nwill help the code review tool to identify the commits as morally\nequivalent. After checking that the other review-relevant data (patch,\nmessage, author...) is identical, the review tool may choose to mark\nthe commit C1 as \"already reviewed\" without any user intervention.\n\nAnd if the data has somehow changed, only the interdiff of that can be\nshown for review, which is precisely the kind of benefit we're aiming\nfor with this header. It has been pointed out that git-range-diff can do\nthis in a limited fashion already, which we can hopefully expand upon.\n\n> I have more problems with the remaining 20% use case, where you need\n> to deal with multiple commits.\n>\n> Perhaps your initial changeset is a single commit C0 that is so\n> large and does too many things at once, and reviewers would\n> naturally advise you to split things up.  You'll come up with a\n> series of commits, C1_0 and C1_1.  The net effect of applying these\n> two patches may be the same as applying the original C0, but each of\n> them is more cleanly separated to address one issue at a time, and\n> the explanation given in the proposed log message more clearly\n> describes the issue each of them addresses.  Now you gave a change ID\n> to C0, and want to somehow relate C1_0 and C1_1 to the original C0.\n> Which one gets the same change ID?  Earlier one?  The last one?\n> Both gets the same change ID?\n>\n> Or your initial changeset is a two-commit series, C0_0 and C0_1, but\n> reviewers find that each one of them alone is not complete, and\n> because the issue addressed by these two is small and isolated\n> enough, you are advised to make them into a single commit C1.  Did\n> you start with two change IDs for these two original commits?  If\n> so, whose change ID the updated commit C1 inherit?  Or does C1 have\n> two change IDs now?  Or did you start with a single change ID\n> assigned to both of these two original commits?\n\nThese are good descriptions of realistic scenarios. At a high level, I'd\nsay change-ids don't help much with these problems. But that's OK, they\ndon't need to solve all problems to be worthwhile.\n\nA commit only ever has one change-id. (Other headers may be useful, but\nthey should have another name and should be disussed separately IMO.)\n\nIn both of the described scenarios, it doesn't matter which one\ninherits the previous one. When a commit is split into two, the one that\ninherits the change-id will show up in review tools with half of its\npatch deleted, the other commit will show up as new. When two commits\nare combined into one, any one of the change-ids is inherited. Review\ntools will show the patch from the commit the ID was inherited from as\npreserved (maybe even hidden) while the patches from the other commits\nwill show up as added. If they were previously reviewed, those reviews\nwill be \"lost\".\n\nSo, while these scenarios are not ideal and somehow ambiguous, the\nambiguity doesn't present any real problems in practice.\n\nMore importantly, I disagree with the 80%-20% split between these\nscenarios. There is another one that comes up a lot, and that scenario\nis the one that benefits the most from change-ids. Namely, when multiple\ncommits are carried together as a patchset. These commits can split\nout into little subtrees that are constantly rebased to keep up with\nupstream, while the commits are amended, reworded and so on. They can\nsometimes be reviewed as a whole, where people are generally looking at\nall commits at once. But they can also represent dependencies between\nentirely different features. For example, bugfixes will often be at the\nbase of this tree / patchset, while features that depend on it fan out\nfrom there. Here is where the change-id shines: As commits are amended,\nreordered, rebased and reworded, their evolution can be trivially\ntracked. There are discussions around the topic of \"stacked PRs\", which\nis basically this exact scenario, just managed in a specific way.\n\nI find myself working this way a lot.\n\n> Quite frankly, I think the concept of \"change ID\" is nice but it is\n> not mechanically trustable.  Recording them in the trailers is fine,\n> but I somehow feel that they have a clear-cut semantics everybody\n> can agree on to deserve to be in the header part of commit objects.\n\nI think the concrete examples above show that while change-ids don't\nsolve all version control problems, these remaining problems don't\ndiminish the value of change-ids either. I also don't see any potential\nfor different tools in the ecosystem to disagree about the semantics so\nprofoundly that interoperability would be hampered.\n\nRemo\n"},{"id":"515810","messageId":"xmqqzfgrjyws.fsf@gitster.g","threadId":"63245","inReplyTo":"xmqq4iyzn0vn.fsf@gitster.g","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-04-08T00:10:43Z","receivedAt":"2025-04-08T00:10:46Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> I have more problems with the remaining 20% use case, where you need\n> to deal with multiple commits.\n>\n> Perhaps your initial changeset is a single commit C0 that is so\n> large and does too many things at once, and reviewers would\n> naturally advise you to split things up.  You'll come up with a\n> series of commits, C1_0 and C1_1.  The net effect of applying these\n> two patches may be the same as applying the original C0, but each of\n> them is more cleanly separated to address one issue at a time, and\n> the explanation given in the proposed log message more clearly\n> describes the issue each of them addresses.  Now you gave a change ID\n> to C0, and want to somehow relate C1_0 and C1_1 to the original C0.\n> Which one gets the same change ID?  Earlier one?  The last one?\n> Both gets the same change ID?\n>\n> Or your initial changeset is a two-commit series, C0_0 and C0_1, but\n> reviewers find that each one of them alone is not complete, and\n> because the issue addressed by these two is small and isolated\n> enough, you are advised to make them into a single commit C1.  Did\n> you start with two change IDs for these two original commits?  If\n> so, whose change ID the updated commit C1 inherit?  Or does C1 have\n> two change IDs now?  Or did you start with a single change ID\n> assigned to both of these two original commits?\n\nAnother thing I forgot to mention.\n\nA well-kept reflog on a topic branch can keep track of how the set\nof patches that makes the topic evolved.  E.g. with something like\nthis:\n\n    $ git switch topic\n    $ git range-diff @{1}...\n    $ git range-diff @{2}...@{1}\n    $ git range-diff @{3}...@{2}\n\nyou can view how the topic as a whole has evolved.  A downside of\nusing reflog entries this way is that it is not well-suited for\ndistributed workflow.\n\nA set of individual commits that share the same \"change ID\" is,\nunlike reflog entries which is an ordered set of tip of topics, not\ninherently ordered.  This is inevitable in the distributed world\nwhere many people can simultaneously work on improving a single\n\"change\" in many different ways, but making it difficult if not\nimpossible to see how things evolved, simply because you first need\nto figure out the order of these commits that share the same \"change\nID\".  Some may be independently evolved from the same ancestor\niteration.  Some may be repeatedly worked on on a single strand of\npearls (much like how development recorded in reflog entries of a\nsingle branch in a single user set-up goes).  I guess you would need\na way to record the predecessor vs successor relationship of various\ncommits that share the same \"change ID\", much like commits form DAG\nto represent ancestor vs descendant relationship.\n\n"},{"id":"515821","messageId":"CAESOdVC8m6VjQtyVi8O8bLWyJFaq7wnQ8U2kxW6SHnoXpCd14w@mail.gmail.com","threadId":"63245","inReplyTo":"xmqqzfgrjyws.fsf@gitster.g","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Martin von Zweigbergk","fromEmail":"martinvonz@google.com","sentAt":"2025-04-08T05:35:03Z","receivedAt":"2025-04-08T05:35:16Z","isPatch":false,"sender":{"key":"martinvonz@google.com","avatar":"https://avatars.githubusercontent.com/u/891642?v=4"},"body":"On Mon, 7 Apr 2025 at 17:10, Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n> A well-kept reflog on a topic branch can keep track of how the set\n> of patches that makes the topic evolved.  E.g. with something like\n> this:\n>\n>     $ git switch topic\n>     $ git range-diff @{1}...\n>     $ git range-diff @{2}...@{1}\n>     $ git range-diff @{3}...@{2}\n>\n> you can view how the topic as a whole has evolved.  A downside of\n> using reflog entries this way is that it is not well-suited for\n> distributed workflow.\n>\n> A set of individual commits that share the same \"change ID\" is,\n> unlike reflog entries which is an ordered set of tip of topics, not\n> inherently ordered.  This is inevitable in the distributed world\n> where many people can simultaneously work on improving a single\n> \"change\" in many different ways, but making it difficult if not\n> impossible to see how things evolved, simply because you first need\n> to figure out the order of these commits that share the same \"change\n> ID\".  Some may be independently evolved from the same ancestor\n> iteration.  Some may be repeatedly worked on on a single strand of\n> pearls (much like how development recorded in reflog entries of a\n> single branch in a single user set-up goes).  I guess you would need\n> a way to record the predecessor vs successor relationship of various\n> commits that share the same \"change ID\", much like commits form DAG\n> to represent ancestor vs descendant relationship.\n\nThat is correct. The change ID should be sufficient for handling\nsimple distributed cases involving a single remote but it's not a full\nreplacement for something like Mercurial's Changeset Evolution [1].\nFor example:\n\n1. I push commit C1 with change ID X to remote R\n2. You fetch commit C1 from remote R\n3. You rewrite commit C1 as commit C2 (still change ID X)\n4. You push C2 to the remote\n5. I fetch from the remote\n\nMy client will now be able to tell that the remote rewrote C1 to C2\n(because you told it to do so). Any local commits I had on top of C1\ncan then be rebased on top of C2.\n\nHowever, let's say you had instead pushed C2 remote R2. If my client\nhad never fetched from that remote before, then when I fetch from R2,\nit would not know whether the new C2 was a successor of C1 or actually\na predecessor of it (or even a sibling in this \"obsolescence graph\" as\nMercurial calls it).\n\n\n[1] https://wiki.mercurial-scm.org/ChangesetEvolution\n"},{"id":"515864","messageId":"20250408125521.GA17892@mit.edu","threadId":"63245","inReplyTo":"Z/RFQY433muaCW44@ubby","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Theodore Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2025-04-08T12:55:21Z","receivedAt":"2025-04-08T12:57:59Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Mon, Apr 07, 2025 at 04:36:01PM -0500, Nico Williams wrote:\n> This is why I suggested earlier that there need to be multiple change\n> IDs, not just one.  Perhaps one is a \"code review ID\" and another is\n> a \"commit change ID\".  The code review ID would let you link together\n> all commits that were reviewed together, so if you have to split or\n> squash commits they would all still have that one code review ID.  The\n> commit change ID would be shared by all sufficiently-similar versions of\n> a commit.  If a commit is dropped or split or squashed then its commit\n> ID might get dropped too, but the code review ID would stay the same.\n\nI think \"code review ID\" makes a lot of sense, although what I would\ncall it is \"patch series ID\".  This has very clear semantic: it ties\ncommits which should be grouped together as a single higher-level set\nof changes.  It could be used by \"git format-patch\" / \"git send-email\"\nto automatically send a group of patches as a logical unit.\n\nI'd include the \"patch series ID\" in the e-mail that gets sent out, so\nthat \"git apply-message\" would be able to retain the patch series ID.\nPatchwork could use the Patch series ID to automatically mark a v2\nversion of patch series as obsoleting the v1 version of the patch\nseries.  So it would be a lot more useful for than just for\nGerrit-style workflows, and that's a good sign that feature makes\nsense from a design perspective.\n\nI'll note that even without the \"commit change ID\", just simply\nknowing that one patch series is a newer version of a pre-existing\npatch series is enough to allow Gerrit to intuit which commit is a\nnewer version of another commit.  For singleton commits, nothing else\nis necessary.  For multi-commit patch series, gerrit could use the\none-line commit description to associate commits; it could use\nordering of the patches; it could just see which commit contents are\nsimilar to previous commits, much like how git detects renames.\n\nIn my experience looking at how kernel developers use gerrit versus\ne-mail workflows, in general, gerrit patch series tend to involve a\nsmaller number of commits, because looking at how various files change\nbetween commtis is awkward; and with e-mail workflows, the patch\nseries tend involve a larger number of commits, because reviewing\nsmaller commits is easier with e-mail.\n\nSo if this true for other communities using web-based review\nworkflows, using an hueristics instead of a \"commit change ID\" might\nbe sufficient --- and for those communities that run into problems,\nthey could continue to use a gerrit-style \"Change-ID: \" in the footer,\nwith the hueristics being used if for some reason commits that don't\nhave the Change-ID make it into Gerrit.\n\n> > Quite frankly, I think the concept of \"change ID\" is nice but it is\n> > not mechanically trustable.  Recording them in the trailers is fine,\n> > but I somehow feel that they have a clear-cut semantics everybody\n> > can agree on to deserve to be in the header part of commit objects.\n> \n> I don't think they need to have such extremely detailed semantics in\n> order to be able to get a header.  The semantics will ultimately be\n> somewhat project-defined, typically something like \"during code review\n> you can use these to related newer updates to an MR/PR/CR to older\n> versions\" and \"once integrated you can use these to find the approved\n> code review as follows [details]\".  The [details] (probably a URI\n> template) for finding concluded CRs might vary.  The CR tool might vary.\n> The construction of the change IDs might vary.  The intent might not\n> vary at all.\n\nI disagree.  From long experience, allowing something into an\ninterface that doesn't have strongly defined semantics has lead to\n*huge* problems.  This has certainly been the case for\nKernel<->Userspace interfaces; so my bias is that if we can't define\nstrong semantics, then we should probably avoid adding that interface\nuntil we can.  Otherwise, this can lead to a huge number of headaches,\nboth for developers and users.\n\nPeople *will* develop automation tools suing an official \"commit\nchange ID\", assuming that how their project (or their forge site) uses\nthe ill-defined Change ID is the One True Way that the badly defined\nfield should be used.  And other people will developer *other* tools\nassuming some other interpreation for that field.  And then the git\ndevelopers and users will be left trying to pick up the pieces.\n\nCheers,\n\n\t\t\t\t\t\t- Ted\n"},{"id":"515869","messageId":"xmqqwmbuybhg.fsf@gitster.g","threadId":"63245","inReplyTo":"CAESOdVC8m6VjQtyVi8O8bLWyJFaq7wnQ8U2kxW6SHnoXpCd14w@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-04-08T14:27:39Z","receivedAt":"2025-04-08T14:27:42Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Martin von Zweigbergk <martinvonz@google.com> writes:\n\n>> A set of individual commits that share the same \"change ID\" is,\n>> unlike reflog entries which is an ordered set of tip of topics, not\n>> inherently ordered.  This is inevitable in the distributed world\n>> where many people can simultaneously work on improving a single\n>> \"change\" in many different ways, but making it difficult if not\n>> impossible to see how things evolved, simply because you first need\n>> to figure out the order of these commits that share the same \"change\n>> ID\".  Some may be independently evolved from the same ancestor\n>> iteration.  Some may be repeatedly worked on on a single strand of\n>> pearls (much like how development recorded in reflog entries of a\n>> single branch in a single user set-up goes).  I guess you would need\n>> a way to record the predecessor vs successor relationship of various\n>> commits that share the same \"change ID\", much like commits form DAG\n>> to represent ancestor vs descendant relationship.\n>\n> That is correct. The change ID should be sufficient for handling\n> simple distributed cases involving a single remote but it's not a full\n> replacement for something like Mercurial's Changeset Evolution [1].\n\nJust a random thought.  We could very easily replace \"change ID\"\nwith a concept of predecessor-successor commits.\n\nJust like we can represent parents-children NxM transitive relation\nonly with 0 or more \"parent\" commit object headers, we can record\nzero or more \"predecessor\" trailer in the commit log.\n\n (1) a commit with no \"predecessor\" is like \"root commit\" in the\n     commit history topology.  It is a brand new change that took\n     inspiration from nobody else and that is not a polished form of\n     any other existing commit.\n\n (2) a commit created as a refinement for one or more existing\n     commits record each of them as \"predecessor\" to it.  Having\n     more than one of them is like a \"merge commit\" in the commit\n     history topology and represents that two patches were squashed\n     into one.\n\n (3) Splitting an originally large change into multiple changes can\n     be represented the same way.  They share the same commit as\n     their \"predecessor\".  Perhaps you have originally two-commit\n     series, A and B, and split them differently in such a way that\n     C has half of a and D has the rest of A plus B.  In which case,\n     C has A as its predecessor while D has both A and B as its\n     predecessor.\n\n (4) Just like we can use auxiliary data structures like bitmaps to\n     figure out reachability without following all the links in the\n     commit history topology, we should be able to learn how a new\n     change was born, and trace how it evolved into newer iteration\n     of the moral equivalent of the change, possibly as a series\n     with mutiple commits, using auxiliary data structure, which\n     would represent predecessor-successor NxM transitive relation\n     in a similar way in a form that is efficient to access.\n\nSomething like this should allow us avoid relying on \"change ID\"s\nthat can collide elsewhere in the world without having a central\nauthority to assign them.\n"},{"id":"515870","messageId":"xmqqv7reybh5.fsf@gitster.g","threadId":"63245","inReplyTo":"CAESOdVC8m6VjQtyVi8O8bLWyJFaq7wnQ8U2kxW6SHnoXpCd14w@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-04-08T14:27:50Z","receivedAt":"2025-04-08T14:27:52Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Martin von Zweigbergk <martinvonz@google.com> writes:\n\n>> A set of individual commits that share the same \"change ID\" is,\n>> unlike reflog entries which is an ordered set of tip of topics, not\n>> inherently ordered.  This is inevitable in the distributed world\n>> where many people can simultaneously work on improving a single\n>> \"change\" in many different ways, but making it difficult if not\n>> impossible to see how things evolved, simply because you first need\n>> to figure out the order of these commits that share the same \"change\n>> ID\".  Some may be independently evolved from the same ancestor\n>> iteration.  Some may be repeatedly worked on on a single strand of\n>> pearls (much like how development recorded in reflog entries of a\n>> single branch in a single user set-up goes).  I guess you would need\n>> a way to record the predecessor vs successor relationship of various\n>> commits that share the same \"change ID\", much like commits form DAG\n>> to represent ancestor vs descendant relationship.\n>\n> That is correct. The change ID should be sufficient for handling\n> simple distributed cases involving a single remote but it's not a full\n> replacement for something like Mercurial's Changeset Evolution [1].\n\nJust a random thought.  We could very easily replace \"change ID\"\nwith a concept of predecessor-successor commits.\n\nJust like we can represent parents-children NxM transitive relation\nonly with 0 or more \"parent\" commit object headers, we can record\nzero or more \"predecessor\" trailer in the commit log.\n\n (1) a commit with no \"predecessor\" is like \"root commit\" in the\n     commit history topology.  It is a brand new change that took\n     inspiration from nobody else and that is not a polished form of\n     any other existing commit.\n\n (2) a commit created as a refinement for one or more existing\n     commits record each of them as \"predecessor\" to it.  Having\n     more than one of them is like a \"merge commit\" in the commit\n     history topology and represents that two patches were squashed\n     into one.\n\n (3) Splitting an originally large change into multiple changes can\n     be represented the same way.  They share the same commit as\n     their \"predecessor\".  Perhaps you have originally two-commit\n     series, A and B, and split them differently in such a way that\n     C has half of a and D has the rest of A plus B.  In which case,\n     C has A as its predecessor while D has both A and B as its\n     predecessor.\n\n (4) Just like we can use auxiliary data structures like bitmaps to\n     figure out reachability without following all the links in the\n     commit history topology, we should be able to learn how a new\n     change was born, and trace how it evolved into newer iteration\n     of the moral equivalent of the change, possibly as a series\n     with mutiple commits, using auxiliary data structure, which\n     would represent predecessor-successor NxM transitive relation\n     in a similar way in a form that is efficient to access.\n\nSomething like this should allow us avoid relying on \"change ID\"s\nthat can collide elsewhere in the world without having a central\nauthority to assign them.\n"},{"id":"515889","messageId":"3a5eeaef-05a1-4e04-8bc5-0d023e63f27c@gmail.com","threadId":"63245","inReplyTo":"xmqqwmbuybhg.fsf@gitster.g","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2025-04-08T15:58:58Z","receivedAt":"2025-04-08T15:59:09Z","isPatch":false,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"On 08/04/2025 15:27, Junio C Hamano wrote:\n> Martin von Zweigbergk <martinvonz@google.com> writes:\n> \n>>> A set of individual commits that share the same \"change ID\" is,\n>>> unlike reflog entries which is an ordered set of tip of topics, not\n>>> inherently ordered.  This is inevitable in the distributed world\n>>> where many people can simultaneously work on improving a single\n>>> \"change\" in many different ways, but making it difficult if not\n>>> impossible to see how things evolved, simply because you first need\n>>> to figure out the order of these commits that share the same \"change\n>>> ID\".  Some may be independently evolved from the same ancestor\n>>> iteration.  Some may be repeatedly worked on on a single strand of\n>>> pearls (much like how development recorded in reflog entries of a\n>>> single branch in a single user set-up goes).  I guess you would need\n>>> a way to record the predecessor vs successor relationship of various\n>>> commits that share the same \"change ID\", much like commits form DAG\n>>> to represent ancestor vs descendant relationship.\n>>\n>> That is correct. The change ID should be sufficient for handling\n>> simple distributed cases involving a single remote but it's not a full\n>> replacement for something like Mercurial's Changeset Evolution [1].\n> \n> Just a random thought.  We could very easily replace \"change ID\"\n> with a concept of predecessor-successor commits.\n> \n> Just like we can represent parents-children NxM transitive relation\n> only with 0 or more \"parent\" commit object headers, we can record\n> zero or more \"predecessor\" trailer in the commit log.\n> \n>   (1) a commit with no \"predecessor\" is like \"root commit\" in the\n>       commit history topology.  It is a brand new change that took\n>       inspiration from nobody else and that is not a polished form of\n>       any other existing commit.\n> \n>   (2) a commit created as a refinement for one or more existing\n>       commits record each of them as \"predecessor\" to it.  Having\n>       more than one of them is like a \"merge commit\" in the commit\n>       history topology and represents that two patches were squashed\n>       into one.\n> \n>   (3) Splitting an originally large change into multiple changes can\n>       be represented the same way.  They share the same commit as\n>       their \"predecessor\".  Perhaps you have originally two-commit\n>       series, A and B, and split them differently in such a way that\n>       C has half of a and D has the rest of A plus B.  In which case,\n>       C has A as its predecessor while D has both A and B as its\n>       predecessor.\n> \n>   (4) Just like we can use auxiliary data structures like bitmaps to\n>       figure out reachability without following all the links in the\n>       commit history topology, we should be able to learn how a new\n>       change was born, and trace how it evolved into newer iteration\n>       of the moral equivalent of the change, possibly as a series\n>       with mutiple commits, using auxiliary data structure, which\n>       would represent predecessor-successor NxM transitive relation\n>       in a similar way in a form that is efficient to access.\n> \n> Something like this should allow us avoid relying on \"change ID\"s\n> that can collide elsewhere in the world without having a central\n> authority to assign them.\n\nThis is similar in spirit to the \"git evolve\" proposal [1]. One of the \nobjections to that was that it required all of the rewritten commits to \nbe pushed back to the remote, rather than just the current version. So \nif I rewrite a branch three times and push the result for review all of \nthe intermediate state gets pushed as well. That is because the \nintermediate commits were needed to track the chain of rewritten commits \n  to avoid the problem Elijah described [2] when trying to follow \ncherry-picked-from trailers. If the predecessor information was stored \nseparately to the commit it refers to (in a notes ref for example) then \nwe could in principle simplify the chain of rewrites when pushing so \nthat we only need to push the final version of the commit and a mapping \nfrom the version that we fetched from the remote.\n\nTracking predecessors as you describe is certainly a more complete \nsolution to tracking the evolution of commits and it addresses the \nshortcomings of change-ids you outlined in your previous mail. It is a \nlot more work to implement though.\n\nBest Wishes\n\nPhillip\n\n[1] \nhttps://lore.kernel.org/git/pull.1356.git.1663959324.gitgitgadget@gmail.com/\n[2] \nhttps://lore.kernel.org/git/CABPp-BECTrVp9X6bVmzU8LEeYsC3KbzeJvAaDPN+FgZz_uEhmA@mail.gmail.com/\n\n"},{"id":"515890","messageId":"Z/VGYrrVZYQ13TLj@ubby","threadId":"63245","inReplyTo":"20250408125521.GA17892@mit.edu","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-08T15:53:06Z","receivedAt":"2025-04-08T16:01:05Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Tue, Apr 08, 2025 at 08:55:21AM -0400, Theodore Ts'o wrote:\n> On Mon, Apr 07, 2025 at 04:36:01PM -0500, Nico Williams wrote:\n> > This is why I suggested earlier that there need to be multiple change\n> > IDs, not just one.  Perhaps one is a \"code review ID\" and another is\n> > a \"commit change ID\".  [...]\n> \n> I think \"code review ID\" makes a lot of sense, although what I would\n> call it is \"patch series ID\".  This has very clear semantic: it ties\n> commits which should be grouped together as a single higher-level set\n> of changes.  It could be used by \"git format-patch\" / \"git send-email\"\n> to automatically send a group of patches as a logical unit.\n> \n> [...]\n\nYes.\n\n> I'll note that even without the \"commit change ID\", just simply\n> knowing that one patch series is a newer version of a pre-existing\n> patch series is enough to allow Gerrit to intuit which commit is a\n> newer version of another commit.  For singleton commits, nothing else\n> is necessary.  For multi-commit patch series, gerrit could use the\n> one-line commit description to associate commits; it could use\n> ordering of the patches; it could just see which commit contents are\n> similar to previous commits, much like how git detects renames.\n\nI'm not keen on CR tools \"intuiting\" from.. similarity checks.  I don't\nlove Git's similarity checks for file renames.  I get that for a\ndistributed VCS assigning something like \"inode numbers\" is tricky, but\nas long as devs don't race to create the same files it was always\npossible to have UUIDs as \"inode numbers\" and avoid the similarity\nchecks.  Strictly speaking we don't even need any of these change IDs to\nmake it possible for tools to use similarity checks to find all versions\nof a commit or patch series or whatever, but it's very nice to have\nsomething less heuristic and more exact.\n\n> In my experience looking at how kernel developers use gerrit versus\n> e-mail workflows, in general, gerrit patch series tend to involve a\n> smaller number of commits, because looking at how various files change\n> between commtis is awkward; and with e-mail workflows, the patch\n> series tend involve a larger number of commits, because reviewing\n> smaller commits is easier with e-mail.\n\nYes.\n\n> So if this true for other communities using web-based review\n> workflows, using an hueristics instead of a [...]\n\nI'm not keen :)\n\n> > I don't think they need to have such extremely detailed semantics in\n> > order to be able to get a header.  The semantics will ultimately be\n> > somewhat project-defined, typically something like \"during code review\n> > you can use these to related newer updates to an MR/PR/CR to older\n> > versions\" and \"once integrated you can use these to find the approved\n> > code review as follows [details]\".  The [details] (probably a URI\n> > template) for finding concluded CRs might vary.  The CR tool might vary.\n> > The construction of the change IDs might vary.  The intent might not\n> > vary at all.\n> \n> I disagree.  From long experience, allowing something into an\n> interface that doesn't have strongly defined semantics has lead to\n> *huge* problems.  This has certainly been the case for\n> Kernel<->Userspace interfaces; so my bias is that if we can't define\n> strong semantics, then we should probably avoid adding that interface\n> until we can.  Otherwise, this can lead to a huge number of headaches,\n> both for developers and users.\n\nSo how much of the [details] do you want specified?  If you want to be\nable to go from \"change ID\" to CR generically for all CR tools then the\nthe best -and perhaps only reasonable- way is to make the change ID a\nURI.  Or if you think the [details] can be elided and still have\nsemantics that are well-defined enough then I think you agree with me\nmore than you disagree :)\n\nIf we want to leave some details to be site-/project-local then perhaps\nchange IDs should have some type and domain/project identifier.  Users\nwho cannot make use of that metadata (e.g., because the CR tool is not\nreachable) can still use the change IDs to link commits and patch\nseries.  I think that linking is the only thing we absolutely must\ndefine semantics for, and the rest can be site-/project-local.  IMO.\n\nNico\n-- \n"},{"id":"515891","messageId":"Z/VOekAaq+n45ex1@ubby","threadId":"63245","inReplyTo":"3a5eeaef-05a1-4e04-8bc5-0d023e63f27c@gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-08T16:27:38Z","receivedAt":"2025-04-08T16:27:48Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Tue, Apr 08, 2025 at 04:58:58PM +0100, Phillip Wood wrote:\n> On 08/04/2025 15:27, Junio C Hamano wrote:\n> > Something like this should allow us avoid relying on \"change ID\"s\n> > that can collide elsewhere in the world without having a central\n> > authority to assign them.\n> \n> This is similar in spirit to the \"git evolve\" proposal [1]. One of the\n> objections to that was that it required all of the rewritten commits to be\n> pushed back to the remote, rather than just the current version. So if I\n> rewrite a branch three times and push the result for review all of the\n> intermediate state gets pushed as well. That is because the intermediate\n> commits were needed to track the chain of rewritten commits  to avoid the\n> problem Elijah described [2] when trying to follow cherry-picked-from\n> trailers. If the predecessor information was stored separately to the commit\n> it refers to (in a notes ref for example) then we could in principle\n> simplify the chain of rewrites when pushing so that we only need to push the\n> final version of the commit and a mapping from the version that we fetched\n> from the remote.\n\nOne might as well have a way to push reflogs.  But so much internal\nhistory is just distracting.  Half the point of linear history upstream\nis to drop all the internal history that no one should care about.\n\nEven during code review this is too much and too distracting.  Some\nusers commit very often, and for them this would make their work rather\ndifficult to review.\n\n> Tracking predecessors as you describe is certainly a more complete solution\n> to tracking the evolution of commits and it addresses the shortcomings of\n> change-ids you outlined in your previous mail. It is a lot more work to\n> implement though.\n\nThere's also that.\n\nIMO that's ETOOMUCH.  Just change IDs / series IDs should suffice.\n"},{"id":"515910","messageId":"20250409121924.GA148735@mit.edu","threadId":"63245","inReplyTo":"Z/VGYrrVZYQ13TLj@ubby","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Theodore Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2025-04-09T12:19:24Z","receivedAt":"2025-04-09T12:20:11Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Tue, Apr 08, 2025 at 10:53:06AM -0500, Nico Williams wrote:\n> I'm not keen on CR tools \"intuiting\" from.. similarity checks.  I don't\n> love Git's similarity checks for file renames.  I get that for a\n> distributed VCS assigning something like \"inode numbers\" is tricky, but\n> as long as devs don't race to create the same files it was always\n> possible to have UUIDs as \"inode numbers\" and avoid the similarity\n> checks.\n\nI'm not keen on fields that can have essentially random semantics.\nPart of this is because today Change-ID is in the footer, and so\nhumans can randomly set it to any value they like.  Sometimes they cut\nand paste footers, and so completely unrelated commits have the same\nChange-Id which show up when you do a Gerrit lookup by Chnage-Id.\nAdmittedly, this aspect gets better if we shove it into the git commit\nheader.\n\nPart of it is because some tools will edit the Change-Id when doing a\ncherry-pick.  (For example, one tool that I'm familiar which is a CLI\nfront-end to Gerrit, when you run the command \"kdt cherry-pick\", will\nunconditionally edit the Change-Id to a completely new value) --- and\nsome will not, because they are just do a \"git cherry-pick\" without\ndoing anything else.  And if you live in an ecosystem where some\npoeple use \"git cherry-pick\", and other people do \"kdt cherry-pick\",\nyou basically have *no* guarantees about how Change-Id might behave\nfor different commits.  This *might* get better if we shove it into a\ngit commit header, although if you give people tools to edit the\nChange-Id as part of a \"git commit --amend\", some tools might end up\nchanging the Change-Id in random ways again.\n\nBut then we have the problem where if patches get merged or split,\nwhat Change-Id is really undefined today.  I could imagine that if a\ncommit gets split, both descedent commits should retain the same\nChange-Id.  Or maybe if a patch stack gets collapsed, all of the\npredecessor Change-Id should be included in that collapsed commit,\nmuch like how an \"Octopus Merge\" might have a half-dozen or more\nparent commits.  Defining the semantics here is part of the battle;\nthe other part of the battle would be how would the tools make sure\nthese semantics get obeyed.\n\nPerhaps one approach might be that the hueristics that you hate being\nused as an automated way to sort it out, might get used to set the\nsemantics at commit time, with perhaps a way for the user to override\nthe hueristics, or where the user has to explicitly acknowledge that\nthe hueristics correctly noticed that the patch has changed radically\nand maybe the Change-Id shouldn't be retained any more?\n\nFinally, perhaps there should be some discussion about whether we\nthink git should be maintaining indexes based on the Commit-Id.\nPersonally, cutting and pasting a random 17 character ID is painful\nand annoying, and when I see it in my shell history, I have no idea\nwhat might have been going on.  So if I need to cut and paste a\nCommit-Id, I might as well cut and paste the one-line commit summary,\nand do a \"git log --grep\" search based on that.  But if the Commit-Id\nis indexed, then maybe it might be more useful?  I dunno....\n\n> So how much of the [details] do you want specified?  If you want to be\n> able to go from \"change ID\" to CR generically for all CR tools then the\n> the best -and perhaps only reasonable- way is to make the change ID a\n> URI.  Or if you think the [details] can be elided and still have\n> semantics that are well-defined enough then I think you agree with me\n> more than you disagree :)\n\nWell, see above about some possible semantics.  I'm *still* not\nconvinced even with the better-defined semantics it's worth storing\nthe extra baggage in the commit header.  But that's more of a\nvalue/philosophical question, much like how we \"could\" store explicit\nfile rename information in the git commit, but in the very early days\nof the git design history, although BitKeeper did track file names,\nLinus consciously decided to go down a much simpler path.  So that's\nreally more of a SMTP vs X.400 preference of simplicity versus\ncomplexity in the protocol versus implementation, which is something\nwhere people of good will might disagree --- and there Junio's\nopinions matter far more then mine.  :-)\n\nCheers,\n\n\t\t\t\t\t- Ted\n\t\t\t\t\t\n"},{"id":"515912","messageId":"xmqqlds9trwv.fsf@gitster.g","threadId":"63245","inReplyTo":"20250409121924.GA148735@mit.edu","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-04-09T12:56:16Z","receivedAt":"2025-04-09T12:56:20Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Theodore Ts'o\" <tytso@mit.edu> writes:\n\n> ...  This *might* get better if we shove it into a\n> git commit header, although if you give people tools to edit the\n> Change-Id as part of a \"git commit --amend\", some tools might end up\n> changing the Change-Id in random ways again.\n\nThanks for pointing these out.  I agree with the above, and it is\none of the reasons why I doubt that it would be a win to have this\ninformation in the header part.\n\n> ... So if I need to cut and paste a\n> Commit-Id, I might as well cut and paste the one-line commit summary,\n> and do a \"git log --grep\" search based on that.  But if the Commit-Id\n> is indexed, then maybe it might be more useful?  I dunno....\n\nIf the information becomes useful enough, we will definitely start\nadding index for it, just like we only have \"parent\" header field in\ncommit objects to represent transitive NxM parent-child relationship\nto start with but have in the form of reachability bitmaps an index\nof which commit can or cannot reach which other commits.  With squash\nand split you outlined above (omitted from my quote), it is likely\nthat you'd want similar transitive NxM predecessor-successor relationship\namong a family of commits that represents patchset evolution, and an\nindex constructed with a similar principle should work well.\n\n> Well, see above about some possible semantics.  I'm *still* not\n> convinced even with the better-defined semantics it's worth storing\n> the extra baggage in the commit header.  But that's more of a\n> value/philosophical question, much like how we \"could\" store explicit\n> file rename information in the git commit, but in the very early days\n> of the git design history, although BitKeeper did track file names,\n> Linus consciously decided to go down a much simpler path.  So that's\n> really more of a SMTP vs X.400 preference of simplicity versus\n> complexity in the protocol versus implementation...\n\nIt is not \"simpler is more manageable\".\n\nThe early days' design decision, which still lives to this day, was\na bit stronger than that.  As can be read from [*1*] (which by the\nway I consider one of the most important message regarding the\ndesign in early days of Git), the design started from \"recording\nrenames is pointless\".\n\nThanks.\n\n\n[Reference]\n\n*1* https://lore.kernel.org/git/Pine.LNX.4.58.0504150753440.7211@ppc970.osdl.org/\n"},{"id":"515925","messageId":"Z/amMj/eg0RbXdkS@ubby","threadId":"63245","inReplyTo":"20250409121924.GA148735@mit.edu","subject":"Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-09T16:54:10Z","receivedAt":"2025-04-09T16:54:21Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Wed, Apr 09, 2025 at 08:19:24AM -0400, Theodore Ts'o wrote:\n> On Tue, Apr 08, 2025 at 10:53:06AM -0500, Nico Williams wrote:\n> > I'm not keen on CR tools \"intuiting\" from.. similarity checks.\n> > [...]\n> \n> I'm not keen on fields that can have essentially random semantics.\n> Part of this is because today Change-ID is in the footer, and so\n> humans can randomly set it to any value they like.  Sometimes they cut\n> and paste footers, and so completely unrelated commits have the same\n> Change-Id which show up when you do a Gerrit lookup by Chnage-Id.\n> Admittedly, this aspect gets better if we shove it into the git commit\n> header.\n>\n> Part of it is because some tools will edit the Change-Id when doing a\n> cherry-pick.  [...]\n\nI was only proposing to leave some details out, not to have completely\nundefined semantics.  The particular details we might want to leave out\nare about resolving change IDs to URIs.  In particular this editing of\nchange IDs on cherry-pick you mention has to not be permitted, or\nperhaps a new change ID could be added -- i.e., are these headers\nsingle-valued or multi-valued?\n\nLet's nail down the semantics of these change ID headers.  Here is a\nproposal to bang on:\n\n - change IDs get preserved on cherry-pick and on `pick`s in rebases\n\n - users can manually remove or change these change IDs, naturally,\n   though generall they would not\n\n - the actual change IDs are either free-form or they are URIs -- pick\n   one, but if they are URIs they should be URIs to CRs, and approved\n   CRs should perhaps have links to integration reports etc.\n\n - there should be one header for a change ID for the patch series (the\n   MR/PR/whateverR); patch series IDs can be shared by many commits in\n   one branch, so they are not in any way unique\n\n - there may be one header for a change ID for each commit, which should\n   be unique in any _branch_, but not unique in any repo (due to back-\n   and forward-ports for example)\n\n - there should be another header to list change IDs from which a commit\n   was derived that nonetheless has a different commit change ID\n\n - these headers should be multi-valued to handle squashes and merges\n\n - if a commit change ID is missing but a path series change ID is\n   present then similarity checks could be used to link multiple\n   versions of any one such commit\n\nOptional:\n\n - a commit change ID could be used as a ref to an object that lists the\n   commits that have that change ID\n\n - a patch series change ID could be used as a ref to an object that lists\n   the head commit of of that patch series in every branch that contains\n   it\n\n> Perhaps one approach might be that the hueristics that you hate being\n> used as an automated way to sort it out, might get used to set the\n> semantics at commit time, with perhaps a way for the user to override\n> the hueristics, or where the user has to explicitly acknowledge that\n> the hueristics correctly noticed that the patch has changed radically\n> and maybe the Change-Id shouldn't be retained any more?\n\nYes, heuristics can be used to help the user make such decisions.  I've\nno issue with that.\n\n> Finally, perhaps there should be some discussion about whether we\n> think git should be maintaining indexes based on the Commit-Id.\n\nIf they can be refs, then they should be.  Since they can't be unique\nthe ref should be to an object listing the actual commits (see above).\n\nThere could also be a non-ref index for these.\n\n> Personally, cutting and pasting a random 17 character ID is painful\n> and annoying, and when I see it in my shell history, I have no idea\n> what might have been going on.  So if I need to cut and paste a\n> Commit-Id, I might as well cut and paste the one-line commit summary,\n> and do a \"git log --grep\" search based on that.  But if the Commit-Id\n> is indexed, then maybe it might be more useful?  I dunno....\n\n+1\n\n> Well, see above about some possible semantics.  I'm *still* not\n> convinced even with the better-defined semantics it's worth storing\n> the extra baggage in the commit header.  But that's more of a\n> value/philosophical question, much like how we \"could\" store explicit\n> file rename information in the git commit, but in the very early days\n> of the git design history, although BitKeeper did track file names,\n> Linus consciously decided to go down a much simpler path.  So that's\n> really more of a SMTP vs X.400 preference of simplicity versus\n> complexity in the protocol versus implementation, which is something\n> where people of good will might disagree --- and there Junio's\n> opinions matter far more then mine.  :-)\n\nI don't find file rename heuristics to be \"simple\", and they're often\nwrong, though I've fully internalized that copies and renames have to be\ndone alone in separate commits with no contents changes so as to make\nincorrect rename determinations much less likely.\n\nNico\n-- \n"},{"id":"515928","messageId":"xmqqv7rdqkla.fsf@gitster.g","threadId":"63245","inReplyTo":"Z/amMj/eg0RbXdkS@ubby","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-04-09T18:02:41Z","receivedAt":"2025-04-09T18:02:45Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nico Williams <nico@cryptonector.com> writes:\n\n> I don't find file rename heuristics to be \"simple\", and they're often\n\nI think you need to re-read what Tytso wrote before reacting.  He\nnever said rename and copy heuristics is simple.  His statement was\nthat choosing *NOT* to record the \"file A was renamed to B and C was\ncopied to D in this commit\" in commit objects can be thought as a\nchoice to keep the system simple.\n\nBut the decision not to record renames and copies in commits is not\nabout simplicity but it is more about correctness.  People record\nwrong renames and not all renaming changes are made with \"git mv\".\n\nThrough your GUI IDE, you may choose \"Copy file\" menu and make a\ncopy of an existing file F to file G, and edit file G extensively\nand record the result in a commit.  No matter how heavy the edit\nafter copying is, your IDE would remember that G came originally by\ncopying from F so it should be able to record the fact that G was\ncopied from F in the resulting commit.\n\nBut recording that as a copy may be a *wrong* thing to do in the\nfirst place.  The IDE cannot guess *why* you copied F to G in the\nfirst place.  Perhaps it was because your project requires you to\nreproduce boilerplate license notice in the comment at the beginning\nof each and every file, and copying an existing file as a whole,\nremoving everything after that initial boilderplate comment, and\nwrite everything else afresh was the easiest way for you to work.\nIf you inspect the resulting change brought in by the commit, with\nintelligence, a human may say \"G is a completely new file, even\nthough it shares the leading boilerplate comment with F and every\nother file in the project\".  Your IDE wouldn't say that.\n\nWe designed not to etch such wrong renames/copoies in stone by\nrecording them at the commit time.  Instead we compare the before\nand after image to intuit the _intention_ of what the user wanted to\ndo _when_ you _ask_ (i.e. when you run \"git diff\" or \"git log\").\nThis has an additional benefit that the heuristics to find out about\nrenames and copies *can* be improved long after commits are made, so\na commit that used to be misidentified to have renamed a file (when\nthere is no such renaming) may later say that a file is removed and\na new file is made independently with newer version of Git.\n"},{"id":"515930","messageId":"Z/a+AVopz+HLa1eL@ubby","threadId":"63245","inReplyTo":"xmqqv7rdqkla.fsf@gitster.g","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-09T18:35:45Z","receivedAt":"2025-04-09T18:51:23Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Wed, Apr 09, 2025 at 11:02:41AM -0700, Junio C Hamano wrote:\n> Nico Williams <nico@cryptonector.com> writes:\n> \n> > I don't find file rename heuristics to be \"simple\", and they're often\n> \n> I think you need to re-read what Tytso wrote before reacting.  He\n> never said rename and copy heuristics is simple.  His statement was\n> that choosing *NOT* to record the \"file A was renamed to B and C was\n> copied to D in this commit\" in commit objects can be thought as a\n> choice to keep the system simple.\n\nI was using file rename heuristics to explain that I wouldn't like more\nof the same for other things; I was not trying to litigate renames.\n\nI'm trying to litigate the _addition_ of more similarity-based\nheuristics for _other_ things.\n\nIf similarity heuristics were enough for CR tools then none would have\nintroduced anything like change IDs.  Or perhaps CR tools authors have\nbeen flat out wrong to not try or use similarity heuristics exclusively\nover change IDs.  That's a topic worth discussing.  I've stated my\npreference for not relying solely on similarity heuristics.\n\n> [...]\n> \n> But recording that as a copy may be a *wrong* thing to do in the\n> first place.  The IDE cannot guess *why* you copied F to G in the\n> first place.  [...]\n\nExactly, software can't read human minds, which is why recording\nuser-expressed intent (rename, copy) would be a nice option to have.\n\n> If you inspect the resulting change brought in by the commit, with\n> intelligence, a human may say \"G is a completely new file, even\n> though it shares the leading boilerplate comment with F and every\n> other file in the project\".  Your IDE wouldn't say that.\n\nRight, the IDE should let one rename/copy the file _and_ choose to ether\nrecord that as a rename/copy in the version control system _or not_.\n\nOnly the user can truly know their own intent.\n\nSometimes users may slip up and not record their intent correctly,\nperhaps because going back to fix an earlier choice is ETOOHARD to\nfigure out how to do in their IDE's UI.  Fair enough.  So similarity\nchecks can help one understand history, but where one can have intent\nrecorded, that's a better indicator.  IMO.\n\n> We designed not to etch such wrong renames/copoies in stone by\n> recording them at the commit time.  Instead we compare the before\n> and after image to intuit the _intention_ of what the user wanted to\n> do _when_ you _ask_ (i.e. when you run \"git diff\" or \"git log\").\n\nWell, I suspect more likely that Linus didn't want to have some sort of\ninode number nor some sort of explicit rename/copy indication as a\nsignificant simplification that allowed Git to get shipped sooner.  I'm\nnot questioning that nor trying to litigate rename/copy.\n\nNico\n-- \n"},{"id":"515933","messageId":"Z/bGzvDfsIcclBW+@ubby","threadId":"63245","inReplyTo":"xmqqlds9trwv.fsf@gitster.g","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-09T19:13:18Z","receivedAt":"2025-04-09T19:13:22Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Wed, Apr 09, 2025 at 05:56:16AM -0700, Junio C Hamano wrote:\n> It is not \"simpler is more manageable\".\n> \n> The early days' design decision, which still lives to this day, was\n> a bit stronger than that.  As can be read from [*1*] (which by the\n> way I consider one of the most important message regarding the\n> design in early days of Git), the design started from \"recording\n> renames is pointless\".\n> \n> *1* https://lore.kernel.org/git/Pine.LNX.4.58.0504150753440.7211@ppc970.osdl.org/\n\nAllow me to withdraw my use of similarity heuristics for renames as an\nargument against similarity heuristics over change IDs.  I still think\nthat explicit change IDs would be better than using only commit\nsimilarity heuristics.\n\nNico\n-- \n"},{"id":"515934","messageId":"CAPig+cSN97oyYbF=mRijbgxUtED2q=u2PFAV+gPP3qM6Vm0OPg@mail.gmail.com","threadId":"63245","inReplyTo":"Z/a+AVopz+HLa1eL@ubby","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2025-04-09T19:14:46Z","receivedAt":"2025-04-09T19:15:04Z","isPatch":false,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Wed, Apr 9, 2025 at 2:51 PM Nico Williams <nico@cryptonector.com> wrote:\n> On Wed, Apr 09, 2025 at 11:02:41AM -0700, Junio C Hamano wrote:\n> > We designed not to etch such wrong renames/copoies in stone by\n> > recording them at the commit time.  Instead we compare the before\n> > and after image to intuit the _intention_ of what the user wanted to\n> > do _when_ you _ask_ (i.e. when you run \"git diff\" or \"git log\").\n>\n> Well, I suspect more likely that Linus didn't want to have some sort of\n> inode number nor some sort of explicit rename/copy indication as a\n> significant simplification that allowed Git to get shipped sooner.  I'm\n> not questioning that nor trying to litigate rename/copy.\n\nContrary to your suspicion, what Junio describes above was a conscious\nand deliberate design decision by Linus[*].\n\n[*]: https://lore.kernel.org/git/Pine.LNX.4.58.0504150753440.7211@ppc970.osdl.org/\n"},{"id":"515936","messageId":"Z/bLEtQRUYEIzSne@ubby","threadId":"63245","inReplyTo":"CAPig+cSN97oyYbF=mRijbgxUtED2q=u2PFAV+gPP3qM6Vm0OPg@mail.gmail.com","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-09T19:31:30Z","receivedAt":"2025-04-09T19:31:35Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Wed, Apr 09, 2025 at 03:14:46PM -0400, Eric Sunshine wrote:\n> Contrary to your suspicion, what Junio describes above was a conscious\n> and deliberate design decision by Linus[*].\n\nFair enough.  Please look past that.\n\nNico\n-- \n"},{"id":"515941","messageId":"xmqqsemgpggd.fsf@gitster.g","threadId":"63245","inReplyTo":"Z/bGzvDfsIcclBW+@ubby","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-04-10T08:29:38Z","receivedAt":"2025-04-10T08:29:42Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nico Williams <nico@cryptonector.com> writes:\n\n> argument against similarity heuristics over change IDs.  I still think\n> that explicit change IDs would be better than using only commit\n> similarity heuristics.\n\nI do think it makes sense to explicitly record that this commit was\n(or \"these commits were\") created to refine and replace that commit\n(or \"these other commits\"), if we want to keep track of how a set of\npatches evolved, if such a determination can be reliably done.  And\n\nI suspect that IDEs can do a much better job keeping track of such\ncorrespondence than they can keep track of renames and copies, which\nI mentioned in an earlier message.\n\nIt is insufficient to just record a single \"change ID\" to each\ncommit, in order to handle anything other than \"a single commit gets\nupdated by another single commit\" case.  It is insufficient to even\nkeep track of \"a single commit gets updated by another single\ncommit, which in turn gets updated by yet another single commit\"\ncase, without assuming globally synchronised clock in a distributed\nenvironment, simply because you only have three commit objects that\nshare the same \"change ID\" string among themselves, and you cannot\ntell between A becoming B becoming C (in which case people would\nconsider C is the latest in the iterations), or two developers\nstarted from A to produce B and C indenendently (in which case it is\nnot yet decided which one between B and C should be considered the\nlatest).\n\nSince we are all human, it is possible that we think things through\nand make a design as complete as humanly possible but it later turns\nout to be insufficient.  If we make such a mistake, we'd then need\nto deal with it and that is just simply a part of developers' life.\n\nBut something that is _known_ to be structurally insufficient before\nit is added to the system?  We should refuse to make such a thing a\npart of very core part of the data structure, like the header fields\nin commit objects.\n\nThanks.\n"},{"id":"515959","messageId":"20250410134426.GB13132@mit.edu","threadId":"63245","inReplyTo":"Z/a+AVopz+HLa1eL@ubby","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Theodore Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2025-04-10T13:44:26Z","receivedAt":"2025-04-10T13:45:02Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Wed, Apr 09, 2025 at 01:35:45PM -0500, Nico Williams wrote:\n> I was using file rename heuristics to explain that I wouldn't like more\n> of the same for other things; I was not trying to litigate renames.\n> \n> I'm trying to litigate the _addition_ of more similarity-based\n> heuristics for _other_ things.\n> \n> If similarity heuristics were enough for CR tools then none would have\n> introduced anything like change IDs.  Or perhaps CR tools authors have\n> been flat out wrong to not try or use similarity heuristics exclusively\n> over change IDs.  That's a topic worth discussing.  I've stated my\n> preference for not relying solely on similarity heuristics.\n\nThere is quite a lot of similarity between trying to record file names\n(and having the concept of \"inode numbers\" for files tracked by git),\nand the discussion we've had about how to track user intent when a\ncommit gets split or merged, and having a \"Change-ID\" which exactly\nfunctions like an \"inode number\", except for an individual commit\nintead of a file.\n\nThe arguments about why we don't have an \"inode number\" for files,\nbecause it *is* complicated and hard to getr right, are *precisely*\nthe same argument for why I remmain unconvinced that having an \"inode\nnumber\" of the semantic idea of a commit (read: Change-Id).\n\nIf you are someone who very much believes in the importance of doing\nper-commit Code Review using something like Gerrit, then you might\nthink that a Change-ID is more *important* than an \"inode number\", and\nso it is therefore worth the greater amount of complexity and/or\nambiguity when the Change-Id gets subject to the same levels of\nincorrectness that having the IDE track the user intent behind a file\ncopy or rename might have.  That's a value judgement, and there's no\nreal right answer here.\n\nAfter all, there are still people, for example as seen on a thread on\nthe The Unix Heritage Sociey mailing list, who have argued that git\nis a hot mess because we don't track file renames and copies the way\n\"real\" source code management systems like BitKeeper and Perforce does\nthings.  I happen to disagree, but that's a value judgement about\nwhat's important in a SCM design.\n\nRegardless how we come out on whethe having an \"inode number\" for the\nhigh-level semantic value of a commit is worth it, I do think having a\n\"patch set ID\" which ties related commits together does make sense,\nthough.  That would solve some interesting problems both for the\nweb/forge review workflow as wel as the mailing list review workflow.\nI'd be curious what people might think about that.\n\nCheers,\n\n\t\t\t\t\t\t- Ted\n"},{"id":"515963","messageId":"xmqqy0w8ng5r.fsf@gitster.g","threadId":"63245","inReplyTo":"20250410134426.GB13132@mit.edu","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-04-10T16:18:56Z","receivedAt":"2025-04-10T16:18:59Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Theodore Ts'o\" <tytso@mit.edu> writes:\n\n> Regardless how we come out on whethe having an \"inode number\" for the\n> high-level semantic value of a commit is worth it, I do think having a\n> \"patch set ID\" which ties related commits together does make sense,\n> though.  That would solve some interesting problems both for the\n> web/forge review workflow as wel as the mailing list review workflow.\n> I'd be curious what people might think about that.\n\nAs a concept, I agree that a mechanism to identify these iterations\nof the same topic collectively is a very valuable thing to have.\n\nFWIW, I use the Message-ID of the cover letter e-mail as a rough\napproximation for \"patch set ID\", and it is quite usable once you\ntrain your contributors to always make the cover letter for\niteration N a direct reply to the cover letter for iteration N-1,\nand also make the individual patches a direct reply to the cover\nletter for the same iteration.\n\nThen visiting lore.kernel.org/$mid/ will give me at a glance some\nessential information about the series, like\n\n - how hotly the topic is being discussed?\n\n - does the iteration $mid I happened to have picked the latest, a\n   bit older, or irrelevantly older?\n\n - has the topic been extending its scope?\n\nThanks to the \"cover for iteration N is a direct response for\niteration N-1\" and \"cover is marked as [PATCH 0/$n]\" conventions,\n\"b4 am\" grabs, by default, the patches from the latest iteration\nwhen given the message-id of the cover letter of any iteration of a\npatch set.  Because most of the time a consumer of an evolving\npatchset is interested in the latest iteration (unless the\ncontributor screws up, in which case we may need to go back and\nexplicitly grab an older iteration), I never felt a need for an\nofficial \"patch set ID\", though.\n"},{"id":"515970","messageId":"CAESOdVA7O2sg3t2YjrWAbHv5edREOh94ckf4oqdi70eH4Q+QNA@mail.gmail.com","threadId":"63245","inReplyTo":"xmqqsemgpggd.fsf@gitster.g","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Martin von Zweigbergk","fromEmail":"martinvonz@google.com","sentAt":"2025-04-10T21:40:34Z","receivedAt":"2025-04-10T21:40:48Z","isPatch":false,"sender":{"key":"martinvonz@google.com","avatar":"https://avatars.githubusercontent.com/u/891642?v=4"},"body":"On Thu, 10 Apr 2025 at 01:29, Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Nico Williams <nico@cryptonector.com> writes:\n>\n> > argument against similarity heuristics over change IDs.  I still think\n> > that explicit change IDs would be better than using only commit\n> > similarity heuristics.\n>\n> I do think it makes sense to explicitly record that this commit was\n> (or \"these commits were\") created to refine and replace that commit\n> (or \"these other commits\"), if we want to keep track of how a set of\n> patches evolved, if such a determination can be reliably done.  And\n>\n> I suspect that IDEs can do a much better job keeping track of such\n> correspondence than they can keep track of renames and copies, which\n> I mentioned in an earlier message.\n>\n> It is insufficient to just record a single \"change ID\" to each\n> commit, in order to handle anything other than \"a single commit gets\n> updated by another single commit\" case.  It is insufficient to even\n> keep track of \"a single commit gets updated by another single\n> commit, which in turn gets updated by yet another single commit\"\n> case, without assuming globally synchronised clock in a distributed\n> environment, simply because you only have three commit objects that\n> share the same \"change ID\" string among themselves, and you cannot\n> tell between A becoming B becoming C (in which case people would\n> consider C is the latest in the iterations), or two developers\n> started from A to produce B and C indenendently (in which case it is\n> not yet decided which one between B and C should be considered the\n> latest).\n>\n> Since we are all human, it is possible that we think things through\n> and make a design as complete as humanly possible but it later turns\n> out to be insufficient.  If we make such a mistake, we'd then need\n> to deal with it and that is just simply a part of developers' life.\n>\n> But something that is _known_ to be structurally insufficient before\n> it is added to the system?  We should refuse to make such a thing a\n> part of very core part of the data structure, like the header fields\n> in commit objects.\n\nI think we are talking about slightly different things. The change ID\nproposal is about providing a stable way of referring to an evolving\ncommit. You can think of it almost like an automatically generated Git\nbranch name that follows the commit as it's rewritten (as if you had\npassed `--update-refs` to every command). I think what you're\ndescribing is more like the \"Git Evolve\" proposal [1] (also linked to\nfrom elsewhere in this thread). I think that's also an interesting\nfeature but I see it as mostly a separate feature. There is certainly\noverlap between the two features, such as how the simple centralized\nflow I mentioned can work pretty well by relying only on the change\nID.\n\nChange IDs in Jujutsu are very useful without support for changeset\nevolution. With the indexing I mentioned and the prefix lookup\nprioritizing \"mutable commits\" (roughly those that are not on a\nremote), it's quite convenient to run `jj log` and see a highlighted\nprefix of, say, \"xu\", and then you can do e.g. `jj show xu` instead of\nhaving to paste a longer ID or type a branch name. A further advantage\nof preferring change IDs over commit IDs in commands is that you don't\nrisk creating \"divergent\" commits (similar copies) by rewriting the\nsame commit twice. For example, `jj describe xu -m foo && jj describe\nxu -m bar` will rewrite the original commit to have message \"foo\" and\nthen rewrite the rewritten commit to have message \"bar\". I understand\nthat this use case is less useful to Git users because Git doesn't\nlike to work in detached HEAD mode and doesn't rewrite descendants\nautomatically. Consider experimenting with jj to get a better sense of\nhow it works :)\n\n\n\n[1] https://lore.kernel.org/git/pull.1356.git.1663959324.gitgitgadget@gmail.com/\n"},{"id":"516000","messageId":"20250411154839.GC648081@mit.edu","threadId":"63245","inReplyTo":"xmqqy0w8ng5r.fsf@gitster.g","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Theodore Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2025-04-11T15:48:39Z","receivedAt":"2025-04-11T15:49:09Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Thu, Apr 10, 2025 at 09:18:56AM -0700, Junio C Hamano wrote:\n> Thanks to the \"cover for iteration N is a direct response for\n> iteration N-1\" and \"cover is marked as [PATCH 0/$n]\" conventions,\n\nEven if the cover for iteration N isn't a reply-to the cover for\ninteration N-1, b4 will search based on the subject line for a cover\nletter with higher version number, and this mostly works.\n\nMy one (admittedly minor) pain point is where someone replies to a\npatch series with something like \"you should really also fix FOO\", and\nthen someone replies with a single patch (without a cover letter,\npossibly created with git; possibly not) that addresses issue FOO.\n\nThis can confuse \"b4 am -c\" into thinking that the patch to address\nFOO was in fact a newer version of the patch being reviewed.  It's not\na big deal; I can deal with this manually.  But having a patch set ID\nwould help with this.\n\nThe other things that would help with having an official patch set ID\nwould be to allow patchwork to automatically supercede an older\nversion of the patch series (possibly with a link to the older version\nof the patch series in the Web UI).\n\nI'd also love if lore.kernel.org and maybe b4 also had an automatic\nway to get at the older versions of the patch series, and the patch\nset ID would help with the automation.  Admittedly it's not strictly\nspeaking necessary, since b4 is already using the cover letter subject\nline to search newer versions of the patch series.  The number of\nmessages it would need to search to find older versions would be\ngreater, though.\n\nCheers,\n\n\t\t\t\t\t- Ted\n"},{"id":"516001","messageId":"20250411-arboreal-ultra-dachshund-f34a54@lemur","threadId":"63245","inReplyTo":"20250411154839.GC648081@mit.edu","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2025-04-11T16:38:56Z","receivedAt":"2025-04-11T16:39:00Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On Fri, Apr 11, 2025 at 11:48:39AM -0400, Theodore Ts'o wrote:\n> On Thu, Apr 10, 2025 at 09:18:56AM -0700, Junio C Hamano wrote:\n> > Thanks to the \"cover for iteration N is a direct response for\n> > iteration N-1\" and \"cover is marked as [PATCH 0/$n]\" conventions,\n> \n> Even if the cover for iteration N isn't a reply-to the cover for\n> interation N-1, b4 will search based on the subject line for a cover\n> letter with higher version number, and this mostly works.\n\nNote, that we try to use the series change-id for this, if we find it. We only\nfall back to matching by subject (+author) if that's not present.\n\n> My one (admittedly minor) pain point is where someone replies to a\n> patch series with something like \"you should really also fix FOO\", and\n> then someone replies with a single patch (without a cover letter,\n> possibly created with git; possibly not) that addresses issue FOO.\n> \n> This can confuse \"b4 am -c\" into thinking that the patch to address\n> FOO was in fact a newer version of the patch being reviewed.  It's not\n> a big deal; I can deal with this manually.  But having a patch set ID\n> would help with this.\n\nI'm not sure there's ever going to be a clear \"do what I mean\" solution here.\nWe try to pick the most common course of action in such case, which is to\nassume that it's a quick followup bugfix for the patch.\n\n> I'd also love if lore.kernel.org and maybe b4 also had an automatic\n> way to get at the older versions of the patch series, and the patch\n> set ID would help with the automation.\n\nYou can, for patches sent with b4 that contain change-id. E.g.:\nhttps://lore.kernel.org/all/?q=changeid%3A20250313-a4-a5-reset-6696e5b18e10\nhttps://lore.kernel.org/lkml/?q=changeid%3A20250313-try_with-cc9f91dd3b60\n\n-K\n"},{"id":"516003","messageId":"xmqqfriemw38.fsf@gitster.g","threadId":"63245","inReplyTo":"20250411154839.GC648081@mit.edu","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-04-11T17:44:43Z","receivedAt":"2025-04-11T17:44:46Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Theodore Ts'o\" <tytso@mit.edu> writes:\n\n> On Thu, Apr 10, 2025 at 09:18:56AM -0700, Junio C Hamano wrote:\n>> Thanks to the \"cover for iteration N is a direct response for\n>> iteration N-1\" and \"cover is marked as [PATCH 0/$n]\" conventions,\n>\n> Even if the cover for iteration N isn't a reply-to the cover for\n> interation N-1, b4 will search based on the subject line for a cover\n> letter with higher version number, and this mostly works.\n\nThat is nice.\n\n> My one (admittedly minor) pain point is where someone replies to a\n> patch series with something like \"you should really also fix FOO\", and\n> then someone replies with a single patch (without a cover letter,\n> possibly created with git; possibly not) that addresses issue FOO.\n>\n> This can confuse \"b4 am -c\" into thinking that the patch to address\n> FOO was in fact a newer version of the patch being reviewed.  It's not\n> a big deal; I can deal with this manually.  But having a patch set ID\n> would help with this.\n\nExcellent.  I've seen this happen often; even though I usually pick\nthese small things up directly from within my newsreader, it would\nbe unpleasant when it happens when you are trying to grab a large\nseries with \"b4 am\".\n\nThe submitting contributor must make a conscious arrangement to give\na \"patch set ID\" shared among the messages in a single iteration,\nand everybody who are responding must make sure they do not add the\nsame ID to the messages they throw at the thread in response.  Those\nwho use format-patch and send-email can do that with convention and\nautomation and there is no reason to rely on In-Reply-To: header\n(which may confuse the automated recipient of manually created\nfollow-up messages).\n\n> The other things that would help with having an official patch set ID\n> would be to allow patchwork to automatically supercede an older\n> version of the patch series (possibly with a link to the older version\n> of the patch series in the Web UI).\n\nLovely.\n\n> I'd also love if lore.kernel.org and maybe b4 also had an automatic\n> way to get at the older versions of the patch series, and the patch\n> set ID would help with the automation.  Admittedly it's not strictly\n> speaking necessary, since b4 is already using the cover letter subject\n> line to search newer versions of the patch series.  The number of\n> messages it would need to search to find older versions would be\n> greater, though.\n\nTrue.\n\nThanks.\n"},{"id":"516052","messageId":"xmqqh62tm5fo.fsf@gitster.g","threadId":"63245","inReplyTo":"3a5eeaef-05a1-4e04-8bc5-0d023e63f27c@gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-04-12T21:32:43Z","receivedAt":"2025-04-12T21:32:46Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Phillip Wood <phillip.wood123@gmail.com> writes:\n\n> This is similar in spirit to the \"git evolve\" proposal [1]. One of the\n> objections to that was that it required all of the rewritten commits\n> to be pushed back to the remote, rather than just the current\n> version.\n\nI am not sure if that particular \"objection\" is valid.\n\nWe can make the predecessor link not participate in the reachability\ndag, and even if we made them contribute to commit reachability,\nthey can be filtered out with --filter= facility, can't they?\n"},{"id":"516053","messageId":"20250412231318.GG13132@mit.edu","threadId":"63245","inReplyTo":"xmqqfriemw38.fsf@gitster.g","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Theodore Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2025-04-12T23:13:18Z","receivedAt":"2025-04-12T23:13:46Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Fri, Apr 11, 2025 at 10:44:43AM -0700, Junio C Hamano wrote:\n> \n> The submitting contributor must make a conscious arrangement to give\n> a \"patch set ID\" shared among the messages in a single iteration,\n> and everybody who are responding must make sure they do not add the\n> same ID to the messages they throw at the thread in response.  Those\n> who use format-patch and send-email can do that with convention and\n> automation and there is no reason to rely on In-Reply-To: header\n> (which may confuse the automated recipient of manually created\n> follow-up messages).\n\nSo it all depends on how the patch set ID is implemented.  Here's one\nway that I had in mind.  The reason why I like like this over the\nChange-ID approach is that the semantics can be very clearly defined,\nand the only thing we rely on is the user saying \"this new commit is\npart of patch series which I'm putting together\". \n\nBy default when creating a new commit, the field is empty (in which\ncase the patch set ID is presumed to be the same as the commit ID), or\nif the user gives a command-line flag say, \"git commit --series\"\nwhich indicates that it is part of a patch series in which case the\npatch set ID of the commit is set to the patch set ID of the current\ncommit (i.e., eventully, its parent commit).\n\nWhenever the commit is amended or rebased or cherry picked, if the\npatch series ID is NULL, then it is set to the original commit ID.\nOtherwise, the existing patch set ID is preserved.\n\nThe patch set ID will be output by git format-patch (perhaps as \"Patch\nSeries ID: sha has\" immediately after the --- line.  And if it is\npresent, \"git am\" will import that patch series ID into git commit\nwhich creates when it sucks in the e-mail.\n\nThe net affect of this is that for new versions of git which implement\nthe Patch Set ID, all new commits are treated as patch series of\nlength 1, unless a subsequent commit is created using \"git commit\n--series\".  And the Patch Set ID will be preserved across\ncherry-picks, rebase operations, and git send-email/git apply-message\noperations.\n\nSo if someone replies to an existing e-mail thread with a new commit,\ngit format-patch will give it a different patch set ID, so we can\ndistinguish it from an amended  copy of a patch in the patch series.\n\nIt also means that singleton commits, the patch ID effectively acts\nmuch like the tranditonal Change-ID.  For multi-commit patch series,\nall of the commits will have the same patch set ID.\n\n       \t   \t   \t     \t      - Ted\n"},{"id":"516104","messageId":"xmqq8qo2srn5.fsf@gitster.g","threadId":"63245","inReplyTo":"20250412231318.GG13132@mit.edu","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-04-14T15:13:18Z","receivedAt":"2025-04-14T15:13:21Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Theodore Ts'o\" <tytso@mit.edu> writes:\n\n> On Fri, Apr 11, 2025 at 10:44:43AM -0700, Junio C Hamano wrote:\n>> \n>> The submitting contributor must make a conscious arrangement to give\n>> a \"patch set ID\" shared among the messages in a single iteration,\n>> and everybody who are responding must make sure they do not add the\n>> same ID to the messages they throw at the thread in response.  Those\n>> who use format-patch and send-email can do that with convention and\n>> automation and there is no reason to rely on In-Reply-To: header\n>> (which may confuse the automated recipient of manually created\n>> follow-up messages).\n>\n> So it all depends on how the patch set ID is implemented.  Here's one\n> way that I had in mind.  The reason why I like like this over the\n> Change-ID approach is that the semantics can be very clearly defined,\n> and the only thing we rely on is the user saying \"this new commit is\n> part of patch series which I'm putting together\". \n>\n> By default when creating a new commit, the field is empty (in which\n> case the patch set ID is presumed to be the same as the commit ID), or\n> if the user gives a command-line flag say, \"git commit --series\"\n> which indicates that it is part of a patch series in which case the\n> patch set ID of the commit is set to the patch set ID of the current\n> commit (i.e., eventully, its parent commit).\n>\n> Whenever the commit is amended or rebased or cherry picked, if the\n> patch series ID is NULL, then it is set to the original commit ID.\n> Otherwise, the existing patch set ID is preserved.\n>\n> The patch set ID will be output by git format-patch (perhaps as \"Patch\n> Series ID: sha has\" immediately after the --- line.  And if it is\n> present, \"git am\" will import that patch series ID into git commit\n> which creates when it sucks in the e-mail.\n>\n> The net affect of this is that for new versions of git which implement\n> the Patch Set ID, all new commits are treated as patch series of\n> length 1, unless a subsequent commit is created using \"git commit\n> --series\".  And the Patch Set ID will be preserved across\n> cherry-picks, rebase operations, and git send-email/git apply-message\n> operations.\n>\n> So if someone replies to an existing e-mail thread with a new commit,\n> git format-patch will give it a different patch set ID, so we can\n> distinguish it from an amended  copy of a patch in the patch series.\n>\n> It also means that singleton commits, the patch ID effectively acts\n> much like the tranditonal Change-ID.  For multi-commit patch series,\n> all of the commits will have the same patch set ID.\n\nYeah, I like that aspect the best---the case for single commit\nseries falling out as a natural degenerate case of the more general\ncase to support multi-commit series is a good sign that the design\ngot something right ;-)\n\nI am still not sure what to think about the lack of explicit the\nevolution history of one patch set that share the same patch set ID.\n\nWhen we have 10 commits that share the same patch set ID, I can\nimagine that we can easily tell 3 are from one iteration, and 3 and\n4 among the rest are from another two iterations by noticing that\nthere are three strand of pearls, having 3, 3, and 4 commits on it.\nAnd we can identify the initial round by noticing that one of the\ncommits have its name as the patch set ID, but I am not sure if we\nshould be OK by not having anything but the committter timestamp to\ntell which one among the other two iterations are earlier, and we\ncannot tell anything about these two other iterations if they are\nindependent rewrites of the original round.\n\nBut other than that, I like something with clearly defined semantics\n(and the definition coming naturally out of the structure, not out\nof some arbitrary convention that forces to bring in some\nsemantics), and what you outlined above looks reasonably clean and\neasy to use.\n\nThanks.\n"},{"id":"516126","messageId":"CALnO6CC_Gvqhcxp4AknwM+YSsngv_0zngKb2XHXN4u0AvKEMMg@mail.gmail.com","threadId":"63245","inReplyTo":"Z/amMj/eg0RbXdkS@ubby","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"D. Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2025-04-14T19:54:23Z","receivedAt":"2025-04-14T19:54:35Z","isPatch":false,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"On Wed, Apr 9, 2025 at 12:56 PM Nico Williams <nico@cryptonector.com> wrote:\n>\n> On Wed, Apr 09, 2025 at 08:19:24AM -0400, Theodore Ts'o wrote:\n> > On Tue, Apr 08, 2025 at 10:53:06AM -0500, Nico Williams wrote:\n> > > I'm not keen on CR tools \"intuiting\" from.. similarity checks.\n> > > [...]\n> >\n> > I'm not keen on fields that can have essentially random semantics.\n> > Part of this is because today Change-ID is in the footer, and so\n> > humans can randomly set it to any value they like.  Sometimes they cut\n> > and paste footers, and so completely unrelated commits have the same\n> > Change-Id which show up when you do a Gerrit lookup by Chnage-Id.\n> > Admittedly, this aspect gets better if we shove it into the git commit\n> > header.\n> >\n> > Part of it is because some tools will edit the Change-Id when doing a\n> > cherry-pick.  [...]\n>\n> I was only proposing to leave some details out, not to have completely\n> undefined semantics.  The particular details we might want to leave out\n> are about resolving change IDs to URIs.  In particular this editing of\n> change IDs on cherry-pick you mention has to not be permitted, or\n> perhaps a new change ID could be added -- i.e., are these headers\n> single-valued or multi-valued?\n>\n> Let's nail down the semantics of these change ID headers.  Here is a\n> proposal to bang on:\n>\n>  - change IDs get preserved on cherry-pick and on `pick`s in rebases\n>\n>  - users can manually remove or change these change IDs, naturally,\n>    though generall they would not\n>\n>  - the actual change IDs are either free-form or they are URIs -- pick\n>    one, but if they are URIs they should be URIs to CRs, and approved\n>    CRs should perhaps have links to integration reports etc.\n\nUsing URIs [to code reviews] looks to me like it makes some\nassumptions about what creates or consumes these headers, right?\nEspecially since the URI should point to a code review… Is there a way\nto do that which is downstream-agnostic?\n\nFurther, and maybe this is my ignorance of Gerrit showing: how would\nyou attach a URI to a local commit when authoring it? You don't have\nthe review URI when running `git commit`, do you? (Maybe I\nmisunderstood; I'm seeing an odd chicken-egg problem here.)\n\nWhich begs another question: what/who applies the initial change ID to\na commit and when?\n\n[…]\n\nI've skimmed most of the discussion, and I think a unique ID for an\nin-flight series could be useful for ergonomics and to support more\ntools that link between versions of the series.\n\nRe-reading the original post [1] (which didn't mention this kind of\nID?), I'm having a hard time seeing the problem statement. There's a\nlot said here about the specifics of the solution, and some other neat\nthings it might unlock… meanwhile, I'm wondering if all the\nconsternation about change IDs is because the problem being solved is\nunderspecified for a core Git feature? (That might tie to Ted's\ninitial concerns about semantic meaning, on which I think I concur:\nthe parent and committer/author headers have unambiguous meaning to\nGit, independent of anything else.)\n\nIt looks to me, an outsider, like the problem is some combination of\n\"I want to track a commit's evolution\" and \"I want to see related\ncommits in review, esp. when it's an identical and already-approved\ncommit.\" But I might be misreading, and clarifying the problem\nstatement might help bring us to a better core solution?\n\n[1]: https://lore.kernel.org/git/xmqqh62tm5fo.fsf@gitster.g/T/#m038be849b9b4020c16c562d810cf77bad91a2c87\n\nCheers,\nD. Ben Knoble\n\nPS This discussion feels somewhat related to the classic GitHub\nproblem of not presenting interdiffs/range-diffs: GitHub shows a\ntoo-flat source diff on force-pushes. Perhaps better web UI tooling\nabout interdiff review (which I think is one of the things Gerrit\ndoes/wants to do?) makes change IDs less necessary, since interdiffs\nhelp connect evolutions of commits?\n"},{"id":"516153","messageId":"Z/1/bDQnWc8Lj29S@ubby","threadId":"63245","inReplyTo":"CALnO6CC_Gvqhcxp4AknwM+YSsngv_0zngKb2XHXN4u0AvKEMMg@mail.gmail.com","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-14T21:34:36Z","receivedAt":"2025-04-15T02:05:46Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Mon, Apr 14, 2025 at 03:54:23PM -0400, D. Ben Knoble wrote:\n> Using URIs [to code reviews] looks to me like it makes some\n> assumptions about what creates or consumes these headers, right?\n> Especially since the URI should point to a code review… Is there a way\n> to do that which is downstream-agnostic?\n\nYou could tag the URI(s) with purposes, but URIs are already pretty\nagnostic as to what is being referenced by them as they are merely the\nreference.  That said, if you need to decompose the URI into specific\nsubitems then you need to understand the underlying application's\nlocal-part (and q-param, if any) scheme.\n\n> Further, and maybe this is my ignorance of Gerrit showing: how would\n> you attach a URI to a local commit when authoring it? You don't have\n> the review URI when running `git commit`, do you? (Maybe I\n> misunderstood; I'm seeing an odd chicken-egg problem here.)\n\nExcellent outlook-changing question.  Local tooling would be needed,\nwhich would be annoying if that tooling were not Git itself, but if it's\nGit then how would it interface with Gerrit or any other such tools?\nWe'd have to define APIs for that, and that too would be annoying.  So\nit has to be `git commit` (or `git rebase -i` and then use a new verb to\nstop and set a commit's change ID metadata, like `reword`, but for\nmetadata), which means the user has to acquire a CR before creating the\nCR, so the CR tools would have to support that.\n\nOn the other hand if it's not CR URIs but more like ticket URIs (as in\nJIRA, bugzilla, etc.) then it's much easier.\n\nPeople already use ticket IDs all the time, typically in the commit\nsubject, else in the commit commentary, typically using some specific\nform.  For example Illumos has devs put one ticket ID in the subject and\nif there are more ticket IDs then the body of the commit message must\nstart with each additional ticket ID on a line by itself with the\nticket's synopsis following the ID.  E.g.,\nhttps://src.illumos.org/source/history/illumos-gate/ (I think they\ndon't allow any actual commentary in the commit message body, with all\ncommentary having to be in the tickets).\n\nTypically tickets have to exist before the commits get created, and in\ncases like Illumos' tickets have to exist before the code review is\ncreated, and the commits have to reference the relevant ticket(s).\n\nOTOH in the Illumos case you see that in a CR one might have multiple\ncommits for different tickets each, and still all be related.  A change\nID/URI could be used to link those together without having to go\nspelunking in the ticket system.  Also ticket IDs (and URIs) could be\nhandy as a header in the commits because otherwise one has to know the\ncommit naming conventions of the project.  Illumos, for example, could\nlink tickets by ID in the subject and commit message body and by URI in\ncommit headers (or footers).\n\n> Which begs another question: what/who applies the initial change ID to\n> a commit and when?\n\nSee above.\n\n> I've skimmed most of the discussion, and I think a unique ID for an\n> in-flight series could be useful for ergonomics and to support more\n> tools that link between versions of the series.\n\nAlso to ease back- and forward-ports.  Though to be fair that's a mostly\nsolved probalm as when users do those typically they start with a\nticket, find a CR linked from the ticket, find the corresponding commits\nin the main branch, then go from there.  A change ID might not be more\nhelpful than that.  Then again, if you're doing a second or third\nbackport then one could find earlier backports that might be easier to\nport from than the main branch commits, and here then a change ID might\nhelp.\n\n> Re-reading the original post [1] (which didn't mention this kind of\n> ID?), I'm having a hard time seeing the problem statement. [...]\n\nIt mentions a \"change ID\".  I'm supposing it could be a URI, but I don't\ncare if it's not, and if it's easier then ignore the URI thing.\n\n> It looks to me, an outsider, like the problem is some combination of\n> \"I want to track a commit's evolution\" and \"I want to see related\n> commits in review, esp. when it's an identical and already-approved\n> commit.\" But I might be misreading, and clarifying the problem\n> statement might help bring us to a better core solution?\n\nI'm the one introducing the second of those, and perhaps I should butt\noff.\n\n> [1]: https://lore.kernel.org/git/xmqqh62tm5fo.fsf@gitster.g/T/#m038be849b9b4020c16c562d810cf77bad91a2c87\n> \n> Cheers,\n> D. Ben Knoble\n> \n> PS This discussion feels somewhat related to the classic GitHub\n> problem of not presenting interdiffs/range-diffs: GitHub shows a\n> too-flat source diff on force-pushes. Perhaps better web UI tooling\n> about interdiff review (which I think is one of the things Gerrit\n> does/wants to do?) makes change IDs less necessary, since interdiffs\n> help connect evolutions of commits?\n\nNothing here could force GH to make their UI nicer and more featureful.\n\nNico\n-- \n"},{"id":"516223","messageId":"acadf677-502b-4c55-8c7c-4d0929d55603@intel.com","threadId":"63245","inReplyTo":"20250412231318.GG13132@mit.edu","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Jacob Keller","fromEmail":"jacob.e.keller@intel.com","sentAt":"2025-04-15T21:38:54Z","receivedAt":"2025-04-15T21:39:28Z","isPatch":false,"sender":{"key":"jacob.e.keller@intel.com","avatar":"https://avatars.githubusercontent.com/u/874719?v=4"},"body":"\n\nOn 4/12/2025 4:13 PM, Theodore Ts'o wrote:\n> On Fri, Apr 11, 2025 at 10:44:43AM -0700, Junio C Hamano wrote:\n>>\n>> The submitting contributor must make a conscious arrangement to give\n>> a \"patch set ID\" shared among the messages in a single iteration,\n>> and everybody who are responding must make sure they do not add the\n>> same ID to the messages they throw at the thread in response.  Those\n>> who use format-patch and send-email can do that with convention and\n>> automation and there is no reason to rely on In-Reply-To: header\n>> (which may confuse the automated recipient of manually created\n>> follow-up messages).\n> \n> So it all depends on how the patch set ID is implemented.  Here's one\n> way that I had in mind.  The reason why I like like this over the\n> Change-ID approach is that the semantics can be very clearly defined,\n> and the only thing we rely on is the user saying \"this new commit is\n> part of patch series which I'm putting together\". \n> \n\n\nI've been catching up on this thread, trying to get a sense of the\ndiscussion. I like the notion of this patch-id. I think dealing with\npatch series as a single entity with one patch-id is nice.\n\nIn my experiences with gerrit, a patch series being treated as\nindividual reviews with their own  change-ids usually discouraged doing\nthings in series and especially discouraged splitting a patch into two\nafter a review started. I prefer being able to collate the series\ntogether, so a patch-id is useful.\n\nHaving the singleton change-id semantics naturally emerge is nice.\n\nOne thing I really liked from previous in the thread was the \"reverse\nhex\" where they suggested encoding a change-id uses letters from the end\nof the alphabet. I really like that it was immediately unambiguous when\nyou see a change-id value vs seeing a commit-id. Obviously there are\nlots of other ways to encode this and I think the thread has discussed\nnumerous options.\n\n> By default when creating a new commit, the field is empty (in which\n> case the patch set ID is presumed to be the same as the commit ID), or\n> if the user gives a command-line flag say, \"git commit --series\"\n> which indicates that it is part of a patch series in which case the\n> patch set ID of the commit is set to the patch set ID of the current\n> commit (i.e., eventully, its parent commit).\n> \n\n> Whenever the commit is amended or rebased or cherry picked, if the\n> patch series ID is NULL, then it is set to the original commit ID.\n> Otherwise, the existing patch set ID is preserved.\n\nOk, so it starts null, but as soon as you rebase/amend/etc the commit\nthen its set. This is how the change-id semantics fall out for singleton\ncommits.\n\n> \n> The patch set ID will be output by git format-patch (perhaps as \"Patch\n> Series ID: sha has\" immediately after the --- line.  And if it is\n> present, \"git am\" will import that patch series ID into git commit\n> which creates when it sucks in the e-mail.\n> \n> The net affect of this is that for new versions of git which implement\n> the Patch Set ID, all new commits are treated as patch series of\n> length 1, unless a subsequent commit is created using \"git commit\n> --series\".  And the Patch Set ID will be preserved across\n> cherry-picks, rebase operations, and git send-email/git apply-message\n> operations.\n\nI think its likely for tooling to emerge to retroactively convert a\nbranch into a \"series\" with the same patch-id after the fact once a\nseries is formed.\n\n> \n> So if someone replies to an existing e-mail thread with a new commit,\n> git format-patch will give it a different patch set ID, so we can\n> distinguish it from an amended  copy of a patch in the patch series.\n> \n> It also means that singleton commits, the patch ID effectively acts\n> much like the tranditonal Change-ID.  For multi-commit patch series,\n> all of the commits will have the same patch set ID.\n> \n>        \t   \t   \t     \t      - Ted\n> \n\n"},{"id":"516224","messageId":"f5f58fef-16d4-4e98-8429-1e10fd9ce07a@intel.com","threadId":"63245","inReplyTo":"CALnO6CC_Gvqhcxp4AknwM+YSsngv_0zngKb2XHXN4u0AvKEMMg@mail.gmail.com","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Jacob Keller","fromEmail":"jacob.e.keller@intel.com","sentAt":"2025-04-15T21:44:53Z","receivedAt":"2025-04-15T21:45:31Z","isPatch":false,"sender":{"key":"jacob.e.keller@intel.com","avatar":"https://avatars.githubusercontent.com/u/874719?v=4"},"body":"\n\nOn 4/14/2025 12:54 PM, D. Ben Knoble wrote:\n> \n> It looks to me, an outsider, like the problem is some combination of\n> \"I want to track a commit's evolution\" and \"I want to see related\n> commits in review, esp. when it's an identical and already-approved\n> commit.\" But I might be misreading, and clarifying the problem\n> statement might help bring us to a better core solution?\n> \n> [1]: https://lore.kernel.org/git/xmqqh62tm5fo.fsf@gitster.g/T/#m038be849b9b4020c16c562d810cf77bad91a2c87\n> \n\nTo me, it seems like multiple different and independent problems are\nbeing solved with something that is almost but not quite the same in\neach of the major projects shown as examples. All of these projects\nwould benefit from having something built into git... but its a\nchallenge when they don't have the same semantics and don't quite solve\nthe same use cases.\n\nIt is hard to come up with something that is general enough to cover all\nof the uses cases.\n\n> Cheers,\n> D. Ben Knoble\n> \n> PS This discussion feels somewhat related to the classic GitHub\n> problem of not presenting interdiffs/range-diffs: GitHub shows a\n> too-flat source diff on force-pushes. Perhaps better web UI tooling\n> about interdiff review (which I think is one of the things Gerrit\n> does/wants to do?) makes change IDs less necessary, since interdiffs\n> help connect evolutions of commits?\n> \n\nI think interdiffs and range-diffs are very helpful. More exposure of\nthese in the various forges would be good. I suspect that creating an\neasy to use web UI for these is a hard problem, especially as there are\na number of corner cases to get right.\n"},{"id":"516227","messageId":"D97KGN6TV8F7.1KKO8GYI65W59@buenzli.dev","threadId":"63245","inReplyTo":"xmqq8qo2srn5.fsf@gitster.g","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Remo Senekowitsch","fromEmail":"remo@buenzli.dev","sentAt":"2025-04-15T22:30:22Z","receivedAt":"2025-04-15T22:30:33Z","isPatch":false,"sender":{"key":"remo@buenzli.dev","avatar":"https://gravatar.com/avatar/7df680b096206886db5a2dc983926f314985bbee662eb406cf23c65322cf98b7?d=mp&s=160"},"body":"On Mon Apr 14, 2025 at 5:13 PM CEST, Junio C Hamano wrote:\n> \"Theodore Ts'o\" <tytso@mit.edu> writes:\n>\n>> On Fri, Apr 11, 2025 at 10:44:43AM -0700, Junio C Hamano wrote:\n>>> \n>>> The submitting contributor must make a conscious arrangement to give\n>>> a \"patch set ID\" shared among the messages in a single iteration,\n>>> and everybody who are responding must make sure they do not add the\n>>> same ID to the messages they throw at the thread in response.  Those\n>>> who use format-patch and send-email can do that with convention and\n>>> automation and there is no reason to rely on In-Reply-To: header\n>>> (which may confuse the automated recipient of manually created\n>>> follow-up messages).\n>>\n>> So it all depends on how the patch set ID is implemented.  Here's one\n>> way that I had in mind.  The reason why I like like this over the\n>> Change-ID approach is that the semantics can be very clearly defined,\n>> and the only thing we rely on is the user saying \"this new commit is\n>> part of patch series which I'm putting together\". \n>>\n>> By default when creating a new commit, the field is empty (in which\n>> case the patch set ID is presumed to be the same as the commit ID), or\n>> if the user gives a command-line flag say, \"git commit --series\"\n>> which indicates that it is part of a patch series in which case the\n>> patch set ID of the commit is set to the patch set ID of the current\n>> commit (i.e., eventully, its parent commit).\n>>\n>> Whenever the commit is amended or rebased or cherry picked, if the\n>> patch series ID is NULL, then it is set to the original commit ID.\n>> Otherwise, the existing patch set ID is preserved.\n>>\n>> The patch set ID will be output by git format-patch (perhaps as \"Patch\n>> Series ID: sha has\" immediately after the --- line.  And if it is\n>> present, \"git am\" will import that patch series ID into git commit\n>> which creates when it sucks in the e-mail.\n>>\n>> The net affect of this is that for new versions of git which implement\n>> the Patch Set ID, all new commits are treated as patch series of\n>> length 1, unless a subsequent commit is created using \"git commit\n>> --series\".  And the Patch Set ID will be preserved across\n>> cherry-picks, rebase operations, and git send-email/git apply-message\n>> operations.\n>>\n>> So if someone replies to an existing e-mail thread with a new commit,\n>> git format-patch will give it a different patch set ID, so we can\n>> distinguish it from an amended  copy of a patch in the patch series.\n>>\n>> It also means that singleton commits, the patch ID effectively acts\n>> much like the tranditonal Change-ID.  For multi-commit patch series,\n>> all of the commits will have the same patch set ID.\n>\n> Yeah, I like that aspect the best---the case for single commit\n> series falling out as a natural degenerate case of the more general\n> case to support multi-commit series is a good sign that the design\n> got something right ;-)\n>\n> I am still not sure what to think about the lack of explicit the\n> evolution history of one patch set that share the same patch set ID.\n>\n> When we have 10 commits that share the same patch set ID, I can\n> imagine that we can easily tell 3 are from one iteration, and 3 and\n> 4 among the rest are from another two iterations by noticing that\n> there are three strand of pearls, having 3, 3, and 4 commits on it.\n> And we can identify the initial round by noticing that one of the\n> commits have its name as the patch set ID, but I am not sure if we\n> should be OK by not having anything but the committter timestamp to\n> tell which one among the other two iterations are earlier, and we\n> cannot tell anything about these two other iterations if they are\n> independent rewrites of the original round.\n>\n> But other than that, I like something with clearly defined semantics\n> (and the definition coming naturally out of the structure, not out\n> of some arbitrary convention that forces to bring in some\n> semantics), and what you outlined above looks reasonably clean and\n> easy to use.\n\nDoesn't a patch set ID suffer from the same kind of ambiguity the\nchange-id supposedly does? Patch sets can be split and merged, a commit\nfrom one patch set can be cherry-picked into another. What patch set ID\nshould such a cherry-picked commit have?\n\nAnd I think the argument that a change-id for a singleton patch\nset naturally falls out of the patch set ID can easily be reversed.\nAdmittedly, I don't have the most experience with the mailing list\nworkflow, but a multi-commit patch set usually comes with a cover\nletter, right? And people like to track their cover letter in a commit?\nIIUC, b4 is designed around that too.\n\nIn that case, the cover letter has its own change-id as any other\ncommit, which will naturally remain stable across every version of the\npatch set. It would be non-sensical to squash, split or cherry-pick the\ncover letter commit. Sounds like a great candidate for the patch set ID.\n\nSo the patch set ID can just as naturally flow out from the change-id.\n\nI can see two concrete disadvantages of the patch set ID:\n\n* It's strictly less powerful. As explained, the change-id can do\n  everything the patch set ID can via the cover letter. But the patch\n  set ID cannot help you track how individual commits within the patch\n  set evolved.\n\n* It's more complicated. While many Git users work with patch sets every\n  day, it's not a concept in Git iself. Git only knows about commits.\n  The patch set ID would introduce a new concept into Git unnecessarily,\n  while the change-id naturally extends the language Git already speaks,\n  that of commits.\n\nRemo\n"},{"id":"516241","messageId":"xmqqy0w1j7an.fsf@gitster.g","threadId":"63245","inReplyTo":"D97KGN6TV8F7.1KKO8GYI65W59@buenzli.dev","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-04-16T00:09:52Z","receivedAt":"2025-04-16T00:09:55Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Remo Senekowitsch\" <remo@buenzli.dev> writes:\n\n> Doesn't a patch set ID suffer from the same kind of ambiguity the\n> change-id supposedly does? Patch sets can be split and merged, a commit\n> from one patch set can be cherry-picked into another. What patch set ID\n\nCorrect.\n\nI still prefer it over the change IDs between these two incomplete\nmechanisms.  To resolve the ambiguity, you'd probably go all the way\nto an approach like \"in addition to the usual parent-child\nrelationship, we record the change evolution relationship so that\nthere is a record on each commit what 'predecessor commit(s)' it was\nderived from\".\n"},{"id":"516242","messageId":"12a56142-d786-40d7-8ae9-2ebdc3467e4c@intel.com","threadId":"63245","inReplyTo":"D97KGN6TV8F7.1KKO8GYI65W59@buenzli.dev","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Jacob Keller","fromEmail":"jacob.e.keller@intel.com","sentAt":"2025-04-16T00:21:20Z","receivedAt":"2025-04-16T00:21:39Z","isPatch":false,"sender":{"key":"jacob.e.keller@intel.com","avatar":"https://avatars.githubusercontent.com/u/874719?v=4"},"body":"\n\nOn 4/15/2025 3:30 PM, Remo Senekowitsch wrote:\n> Doesn't a patch set ID suffer from the same kind of ambiguity the\n> change-id supposedly does? Patch sets can be split and merged, a commit\n> from one patch set can be cherry-picked into another. What patch set ID\n> should such a cherry-picked commit have?\n> \n> And I think the argument that a change-id for a singleton patch\n> set naturally falls out of the patch set ID can easily be reversed.\n> Admittedly, I don't have the most experience with the mailing list\n> workflow, but a multi-commit patch set usually comes with a cover\n> letter, right? And people like to track their cover letter in a commit?\n> IIUC, b4 is designed around that too.\n> \n> In that case, the cover letter has its own change-id as any other\n> commit, which will naturally remain stable across every version of the\n> patch set. It would be non-sensical to squash, split or cherry-pick the\n> cover letter commit. Sounds like a great candidate for the patch set ID.\n> \n\nIf you commit your cover letter, that would work fine. If you don't\ncommit your cover letter, you could probably also generate this.\n\nHowever, the commit itself wouldn't necessary have the patch-id as part\nof its metadata so you may not be able to easily look this up without\nreferring to the external place where you published. Generally cover\nletters are commits only for the submitter. Once you submit and it is\napplied/merged, the patch-id would vanish unless\n\n> So the patch set ID can just as naturally flow out from the change-id.\n> \n> I can see two concrete disadvantages of the patch set ID:\n> \n> * It's strictly less powerful. As explained, the change-id can do\n>   everything the patch set ID can via the cover letter. But the patch\n>   set ID cannot help you track how individual commits within the patch\n>   set evolved.\n\nFair.\n\n> \n> * It's more complicated. While many Git users work with patch sets every\n>   day, it's not a concept in Git iself. Git only knows about commits.\n>   The patch set ID would introduce a new concept into Git unnecessarily,\n>   while the change-id naturally extends the language Git already speaks,\n>   that of commits.\n> \n\nSure. I guess it depends somewhat on what you want to get out of the\nchange-ids.\n\nIf you only care about \"how do commits fit together in series\", I think\na set of change-ids per commit is not that helpful on its own. You could\nlikely build the equivalent of patch-id out of it with proper tooling.\n\nHowever, if patch-id is not sufficient for what most folks interested in\nthis topic want, then I would agree with you its not the right direction.\n\n> Remo\n> \n\n"},{"id":"516243","messageId":"fac75248-0a32-4c84-b37c-0c72b0b189e6@intel.com","threadId":"63245","inReplyTo":"xmqqwmbuybhg.fsf@gitster.g","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Jacob Keller","fromEmail":"jacob.e.keller@intel.com","sentAt":"2025-04-16T00:24:58Z","receivedAt":"2025-04-16T00:25:03Z","isPatch":false,"sender":{"key":"jacob.e.keller@intel.com","avatar":"https://avatars.githubusercontent.com/u/874719?v=4"},"body":"\n\nOn 4/8/2025 7:27 AM, Junio C Hamano wrote:\n> Martin von Zweigbergk <martinvonz@google.com> writes:\n> \n>>> A set of individual commits that share the same \"change ID\" is,\n>>> unlike reflog entries which is an ordered set of tip of topics, not\n>>> inherently ordered.  This is inevitable in the distributed world\n>>> where many people can simultaneously work on improving a single\n>>> \"change\" in many different ways, but making it difficult if not\n>>> impossible to see how things evolved, simply because you first need\n>>> to figure out the order of these commits that share the same \"change\n>>> ID\".  Some may be independently evolved from the same ancestor\n>>> iteration.  Some may be repeatedly worked on on a single strand of\n>>> pearls (much like how development recorded in reflog entries of a\n>>> single branch in a single user set-up goes).  I guess you would need\n>>> a way to record the predecessor vs successor relationship of various\n>>> commits that share the same \"change ID\", much like commits form DAG\n>>> to represent ancestor vs descendant relationship.\n>>\n>> That is correct. The change ID should be sufficient for handling\n>> simple distributed cases involving a single remote but it's not a full\n>> replacement for something like Mercurial's Changeset Evolution [1].\n> \n> Just a random thought.  We could very easily replace \"change ID\"\n> with a concept of predecessor-successor commits.\n> \n> Just like we can represent parents-children NxM transitive relation\n> only with 0 or more \"parent\" commit object headers, we can record\n> zero or more \"predecessor\" trailer in the commit log.\n> \n>  (1) a commit with no \"predecessor\" is like \"root commit\" in the\n>      commit history topology.  It is a brand new change that took\n>      inspiration from nobody else and that is not a polished form of\n>      any other existing commit.\n> \n>  (2) a commit created as a refinement for one or more existing\n>      commits record each of them as \"predecessor\" to it.  Having\n>      more than one of them is like a \"merge commit\" in the commit\n>      history topology and represents that two patches were squashed\n>      into one.\n> \n>  (3) Splitting an originally large change into multiple changes can\n>      be represented the same way.  They share the same commit as\n>      their \"predecessor\".  Perhaps you have originally two-commit\n>      series, A and B, and split them differently in such a way that\n>      C has half of a and D has the rest of A plus B.  In which case,\n>      C has A as its predecessor while D has both A and B as its\n>      predecessor.\n> \n>  (4) Just like we can use auxiliary data structures like bitmaps to\n>      figure out reachability without following all the links in the\n>      commit history topology, we should be able to learn how a new\n>      change was born, and trace how it evolved into newer iteration\n>      of the moral equivalent of the change, possibly as a series\n>      with mutiple commits, using auxiliary data structure, which\n>      would represent predecessor-successor NxM transitive relation\n>      in a similar way in a form that is efficient to access.\n> \n> Something like this should allow us avoid relying on \"change ID\"s\n> that can collide elsewhere in the world without having a central\n> authority to assign them.\n> \n\nThis does seem like the most \"powerful\" form of this, but does lose one\nof the \"simplicity\"-based advantages of change ids.\n\nOf course, you could simply use the root commit ID in most cases and\nthat would be sufficient, and in cases where its not unique you could\nhave tooling show more data and allow users to disambiguate.\n\nThis approach also likely requires the most \"work\" to implement on the\ngit side, vs storing a simpler single-value header.\n"},{"id":"516278","messageId":"D9816I5AX1RG.AA4A7H2D8SJ7@buenzli.dev","threadId":"63245","inReplyTo":"CALnO6CC_Gvqhcxp4AknwM+YSsngv_0zngKb2XHXN4u0AvKEMMg@mail.gmail.com","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Remo Senekowitsch","fromEmail":"remo@buenzli.dev","sentAt":"2025-04-16T11:36:26Z","receivedAt":"2025-04-16T11:36:33Z","isPatch":false,"sender":{"key":"remo@buenzli.dev","avatar":"https://gravatar.com/avatar/7df680b096206886db5a2dc983926f314985bbee662eb406cf23c65322cf98b7?d=mp&s=160"},"body":"On Mon Apr 14, 2025 at 9:54 PM CEST, D. Ben Knoble wrote:\n> On Wed, Apr 9, 2025 at 12:56 PM Nico Williams <nico@cryptonector.com> wrote:\n>> Let's nail down the semantics of these change ID headers.  Here is a\n>> proposal to bang on:\n>>\n>>  - change IDs get preserved on cherry-pick and on `pick`s in rebases\n>>\n>>  - users can manually remove or change these change IDs, naturally,\n>>    though generall they would not\n>>\n>>  - the actual change IDs are either free-form or they are URIs -- pick\n>>    one, but if they are URIs they should be URIs to CRs, and approved\n>>    CRs should perhaps have links to integration reports etc.\n>\n> Using URIs [to code reviews] looks to me like it makes some\n> assumptions about what creates or consumes these headers, right?\n> Especially since the URI should point to a code review… Is there a way\n> to do that which is downstream-agnostic?\n>\n> Further, and maybe this is my ignorance of Gerrit showing: how would\n> you attach a URI to a local commit when authoring it? You don't have\n> the review URI when running `git commit`, do you? (Maybe I\n> misunderstood; I'm seeing an odd chicken-egg problem here.)\n>\n> Which begs another question: what/who applies the initial change ID to\n> a commit and when?\n\nThese are all great questions, which the originally proposed format\n(fixed-width reverse-hex) has answers to. I think a URI would be\nstrictly worse.\n\n* Using a reverse-hex ID makes no assumptions about what consumes these\n  headers. There can be multiple different consumers which treat the ID\n  differently with different URI schemes.\n\n* Attaching a reverse-hex ID to a local commit when authoring it is\n  trivial: you generate it randomly.\n\nThis is one of those cases where being maximally restrictive about the\nformat will enable maximal flexibility downstream.\n\nOne example: GitHub has a URL scheme that looks like this:\ngithub.com/org/repo/compare/<ref1>..<ref2>\n\nThis doesn't work if the refs contain slashes, as branches sometimes do\n(e.g. feat/foo, username/bar). If the change-id is a URI, this type of\nURL scheme doesn't work reliably.\n\nThat is not to say we should design the change-id around GitHub, it's\njust an example how making the format more free-form (URI is more\nfree-form than fixed-width reverse-hex) makes it more difficult to get\nstuff working downstream.\n\nAnd lastly, laser-etching the URI scheme of one particular tool into\nyour commit history means the history is at great risk of degrading\nover time. URI schemes change, domains change, tools become outdated\nand are replaced.\n\nAdding some ephemeral configuration to a tool that constructs a URI out\nof a reverse-hex ID on the other hand is trivial.\n\n> PS This discussion feels somewhat related to the classic GitHub\n> problem of not presenting interdiffs/range-diffs: GitHub shows a\n> too-flat source diff on force-pushes. Perhaps better web UI tooling\n> about interdiff review (which I think is one of the things Gerrit\n> does/wants to do?) makes change IDs less necessary, since interdiffs\n> help connect evolutions of commits?\n\nI think it's the other way around: Building a code review UI built on\ngit and centered around interdiffs today is _hard_, that's why we don't\nhave it yet. Adding change-ids to commits will make it much easier,\npaving the way for these tools to be implemented.\n\nRemo\n"},{"id":"516529","messageId":"CALnO6CCjkxv40+5wZ_vwZTKv7Te8Xh--M1fY2wbuOfgJm5LZxw@mail.gmail.com","threadId":"63245","inReplyTo":"D9816I5AX1RG.AA4A7H2D8SJ7@buenzli.dev","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"D. Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2025-04-22T20:17:03Z","receivedAt":"2025-04-22T20:17:15Z","isPatch":false,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"On Wed, Apr 16, 2025 at 7:36 AM Remo Senekowitsch <remo@buenzli.dev> wrote:\n>\n> On Mon Apr 14, 2025 at 9:54 PM CEST, D. Ben Knoble wrote:\n> > On Wed, Apr 9, 2025 at 12:56 PM Nico Williams <nico@cryptonector.com> wrote:\n> >> Let's nail down the semantics of these change ID headers.  Here is a\n> >> proposal to bang on:\n> >>\n> >>  - change IDs get preserved on cherry-pick and on `pick`s in rebases\n> >>\n> >>  - users can manually remove or change these change IDs, naturally,\n> >>    though generall they would not\n> >>\n> >>  - the actual change IDs are either free-form or they are URIs -- pick\n> >>    one, but if they are URIs they should be URIs to CRs, and approved\n> >>    CRs should perhaps have links to integration reports etc.\n> >\n> > Using URIs [to code reviews] looks to me like it makes some\n> > assumptions about what creates or consumes these headers, right?\n> > Especially since the URI should point to a code review… Is there a way\n> > to do that which is downstream-agnostic?\n> >\n> > Further, and maybe this is my ignorance of Gerrit showing: how would\n> > you attach a URI to a local commit when authoring it? You don't have\n> > the review URI when running `git commit`, do you? (Maybe I\n> > misunderstood; I'm seeing an odd chicken-egg problem here.)\n> >\n> > Which begs another question: what/who applies the initial change ID to\n> > a commit and when?\n>\n> These are all great questions, which the originally proposed format\n> (fixed-width reverse-hex) has answers to. I think a URI would be\n> strictly worse.\n\nWell, I think we still missed \"what/who applies the initial change ID\nto a commit and when.\"\n\nBut the treatment below is something I agree with and failed to\nconvey, I think: namely, URIs seem to encode too much\n\"unportable\"/\"specific\" information in Git. I feel like the current\ndesign is not really \"tool-agnostic\" as much as \"built on a universal\ncore.\" That seems valuable and prone to more longevity.\n\n>\n> * Using a reverse-hex ID makes no assumptions about what consumes these\n>   headers. There can be multiple different consumers which treat the ID\n>   differently with different URI schemes.\n>\n> * Attaching a reverse-hex ID to a local commit when authoring it is\n>   trivial: you generate it randomly.\n>\n> This is one of those cases where being maximally restrictive about the\n> format will enable maximal flexibility downstream.\n>\n> One example: GitHub has a URL scheme that looks like this:\n> github.com/org/repo/compare/<ref1>..<ref2>\n>\n> This doesn't work if the refs contain slashes, as branches sometimes do\n> (e.g. feat/foo, username/bar). If the change-id is a URI, this type of\n> URL scheme doesn't work reliably.\n>\n> That is not to say we should design the change-id around GitHub, it's\n> just an example how making the format more free-form (URI is more\n> free-form than fixed-width reverse-hex) makes it more difficult to get\n> stuff working downstream.\n>\n> And lastly, laser-etching the URI scheme of one particular tool into\n> your commit history means the history is at great risk of degrading\n> over time. URI schemes change, domains change, tools become outdated\n> and are replaced.\n>\n> Adding some ephemeral configuration to a tool that constructs a URI out\n> of a reverse-hex ID on the other hand is trivial.\n\nYep.\n\n> > PS This discussion feels somewhat related to the classic GitHub\n> > problem of not presenting interdiffs/range-diffs: GitHub shows a\n> > too-flat source diff on force-pushes. Perhaps better web UI tooling\n> > about interdiff review (which I think is one of the things Gerrit\n> > does/wants to do?) makes change IDs less necessary, since interdiffs\n> > help connect evolutions of commits?\n>\n> I think it's the other way around: Building a code review UI built on\n> git and centered around interdiffs today is _hard_, that's why we don't\n> have it yet. Adding change-ids to commits will make it much easier,\n> paving the way for these tools to be implemented.\n>\n> Remo\n\nFair point, although GitHub's detection of force-pushes makes me think\nit could split a PR into versions at that point, cross-link backwards\nand forwards by one version (from the force-push detection), show\nrange-diffs between versions based on the target branch of the merge,\nand even follow the cross-links to show an overall sequence of\nversions.\n\nBut I don't work there, so presumably it's harder than that :)\n\nI sincerely hope to make it easier if it's really that hard! And,\nthough my opinion matters little, I'm having a hard time piercing the\nconversation to see a \"universal core\" that solves the desired\nproblem. Maybe I'm not reading carefully enough, and maybe a summary\nwould help. I greatly appreciated the work of previous folks to\nsummarize the current thread status.\n\n-- \nD. Ben Knoble\n"},{"id":"516538","messageId":"D9DIPNY431IJ.23DG6UL5CIQJ@buenzli.dev","threadId":"63245","inReplyTo":"CALnO6CCjkxv40+5wZ_vwZTKv7Te8Xh--M1fY2wbuOfgJm5LZxw@mail.gmail.com","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Remo Senekowitsch","fromEmail":"remo@buenzli.dev","sentAt":"2025-04-22T22:24:06Z","receivedAt":"2025-04-22T22:24:12Z","isPatch":false,"sender":{"key":"remo@buenzli.dev","avatar":"https://gravatar.com/avatar/7df680b096206886db5a2dc983926f314985bbee662eb406cf23c65322cf98b7?d=mp&s=160"},"body":"On Tue Apr 22, 2025 at 10:17 PM CEST, D. Ben Knoble wrote:\n> On Wed, Apr 16, 2025 at 7:36 AM Remo Senekowitsch <remo@buenzli.dev> wrote:\n>>\n>> On Mon Apr 14, 2025 at 9:54 PM CEST, D. Ben Knoble wrote:\n>> > On Wed, Apr 9, 2025 at 12:56 PM Nico Williams <nico@cryptonector.com> wrote:\n>> >> Let's nail down the semantics of these change ID headers.  Here is a\n>> >> proposal to bang on:\n>> >>\n>> >>  - change IDs get preserved on cherry-pick and on `pick`s in rebases\n>> >>\n>> >>  - users can manually remove or change these change IDs, naturally,\n>> >>    though generall they would not\n>> >>\n>> >>  - the actual change IDs are either free-form or they are URIs -- pick\n>> >>    one, but if they are URIs they should be URIs to CRs, and approved\n>> >>    CRs should perhaps have links to integration reports etc.\n>> >\n>> > Using URIs [to code reviews] looks to me like it makes some\n>> > assumptions about what creates or consumes these headers, right?\n>> > Especially since the URI should point to a code review… Is there a way\n>> > to do that which is downstream-agnostic?\n>> >\n>> > Further, and maybe this is my ignorance of Gerrit showing: how would\n>> > you attach a URI to a local commit when authoring it? You don't have\n>> > the review URI when running `git commit`, do you? (Maybe I\n>> > misunderstood; I'm seeing an odd chicken-egg problem here.)\n>> >\n>> > Which begs another question: what/who applies the initial change ID to\n>> > a commit and when?\n>>\n>> These are all great questions, which the originally proposed format\n>> (fixed-width reverse-hex) has answers to. I think a URI would be\n>> strictly worse.\n>\n> Well, I think we still missed \"what/who applies the initial change ID\n> to a commit and when.\"\n\nThe tool that creates the commit, when it creates the commit. In the\ncontext of this discussion, that's Git or Jujutsu. If the commit is\nbrand new, generate the change-id randomly. If it's the \"spiritual\nsuccessor\" of another commit, use that commit's change-id.\n\nBtw. since the thread was started, the implementation in Jujutsu has\nbeen completed and I've been pushing commits with the change-id header\nto various remotes for a while now. It works well. Forges can start\ntaking advantage of it. (I hope I find time to help work on that.)\n\n> But the treatment below is something I agree with and failed to\n> convey, I think: namely, URIs seem to encode too much\n> \"unportable\"/\"specific\" information in Git. I feel like the current\n> design is not really \"tool-agnostic\" as much as \"built on a universal\n> core.\" That seems valuable and prone to more longevity.\n>\n>>\n>> * Using a reverse-hex ID makes no assumptions about what consumes these\n>>   headers. There can be multiple different consumers which treat the ID\n>>   differently with different URI schemes.\n>>\n>> * Attaching a reverse-hex ID to a local commit when authoring it is\n>>   trivial: you generate it randomly.\n>>\n>> This is one of those cases where being maximally restrictive about the\n>> format will enable maximal flexibility downstream.\n>>\n>> One example: GitHub has a URL scheme that looks like this:\n>> github.com/org/repo/compare/<ref1>..<ref2>\n>>\n>> This doesn't work if the refs contain slashes, as branches sometimes do\n>> (e.g. feat/foo, username/bar). If the change-id is a URI, this type of\n>> URL scheme doesn't work reliably.\n>>\n>> That is not to say we should design the change-id around GitHub, it's\n>> just an example how making the format more free-form (URI is more\n>> free-form than fixed-width reverse-hex) makes it more difficult to get\n>> stuff working downstream.\n>>\n>> And lastly, laser-etching the URI scheme of one particular tool into\n>> your commit history means the history is at great risk of degrading\n>> over time. URI schemes change, domains change, tools become outdated\n>> and are replaced.\n>>\n>> Adding some ephemeral configuration to a tool that constructs a URI out\n>> of a reverse-hex ID on the other hand is trivial.\n>\n> Yep.\n>\n>> > PS This discussion feels somewhat related to the classic GitHub\n>> > problem of not presenting interdiffs/range-diffs: GitHub shows a\n>> > too-flat source diff on force-pushes. Perhaps better web UI tooling\n>> > about interdiff review (which I think is one of the things Gerrit\n>> > does/wants to do?) makes change IDs less necessary, since interdiffs\n>> > help connect evolutions of commits?\n>>\n>> I think it's the other way around: Building a code review UI built on\n>> git and centered around interdiffs today is _hard_, that's why we don't\n>> have it yet. Adding change-ids to commits will make it much easier,\n>> paving the way for these tools to be implemented.\n>\n> Fair point, although GitHub's detection of force-pushes makes me think\n> it could split a PR into versions at that point, cross-link backwards\n> and forwards by one version (from the force-push detection), show\n> range-diffs between versions based on the target branch of the merge,\n> and even follow the cross-links to show an overall sequence of\n> versions.\n>\n> But I don't work there, so presumably it's harder than that :)\n\nI agree that forges still have a lot of potential to improve their code\nreview UIs even without change-ids. But tracking individual patches\nacross force-pushes is not as easy as calling git range-diff. Its\nmanpage says the output is not stable for machines to read and the\nalgorithm it uses has cubic runtime complexity. Not something I'd want\nto use on my backend if the user controls the input. So the next best\nthing is to reimplement it using cheaper (but worse) heuristics. The\nauthor timestamp would probably be a good start. But at that point,\nyou're looking at a lot of work for what users will perceive as an\nunreliable and inconsistent experience.\n\n> I sincerely hope to make it easier if it's really that hard! And,\n> though my opinion matters little, I'm having a hard time piercing the\n> conversation to see a \"universal core\" that solves the desired\n> problem. Maybe I'm not reading carefully enough, and maybe a summary\n> would help. I greatly appreciated the work of previous folks to\n> summarize the current thread status.\n\n--\nBest regards,\nRemo\n"},{"id":"516540","messageId":"xmqq8qnr3jji.fsf@gitster.g","threadId":"63245","inReplyTo":"D9DIPNY431IJ.23DG6UL5CIQJ@buenzli.dev","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-04-22T22:42:25Z","receivedAt":"2025-04-22T22:42:28Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Remo Senekowitsch\" <remo@buenzli.dev> writes:\n\n> Btw. since the thread was started, the implementation in Jujutsu has\n> been completed and I've been pushing commits with the change-id header\n> to various remotes for a while now. It works well. Forges can start\n> taking advantage of it. (I hope I find time to help work on that.)\n\nIt should work well, until somebody finds your random is not random\nenough, right?  Unlike our object name that depends on the contents\n(hence a duplicate unless the cryptographic hash function collides\nmeans they are truly the same commit), there is no grabally unique\nID assigner involved in your implementation, right?  Until\nsufficiently large number of people start using and large number of\nchanges gets assigned IDs, it won't become an issue, but then how\nwould it be different from what was raised in the earlier discussion\nto use the commit object name itself as the change ID for a commit\nthat is not derived from anybody else, and copy that ID to commits\nthat are derived from the original commit as the change ID shared\namong them?  At least that would give us a much better uniqueness\nguarantee, wouldn't it?  If you want to be able to tell between\ncommit object name and change ID, I wouldn't object if you encode\nthem using whatever mechanism.\n\n"},{"id":"516541","messageId":"D9DJXDL6JJJ0.3BM1YMOSV5P0F@buenzli.dev","threadId":"63245","inReplyTo":"xmqq8qnr3jji.fsf@gitster.g","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Remo Senekowitsch","fromEmail":"remo@buenzli.dev","sentAt":"2025-04-22T23:21:12Z","receivedAt":"2025-04-22T23:21:23Z","isPatch":false,"sender":{"key":"remo@buenzli.dev","avatar":"https://gravatar.com/avatar/7df680b096206886db5a2dc983926f314985bbee662eb406cf23c65322cf98b7?d=mp&s=160"},"body":"On Wed Apr 23, 2025 at 12:42 AM CEST, Junio C Hamano wrote:\n> \"Remo Senekowitsch\" <remo@buenzli.dev> writes:\n>\n>> Btw. since the thread was started, the implementation in Jujutsu has\n>> been completed and I've been pushing commits with the change-id header\n>> to various remotes for a while now. It works well. Forges can start\n>> taking advantage of it. (I hope I find time to help work on that.)\n>\n> It should work well, until somebody finds your random is not random\n> enough, right?  Unlike our object name that depends on the contents\n> (hence a duplicate unless the cryptographic hash function collides\n> means they are truly the same commit), there is no grabally unique\n> ID assigner involved in your implementation, right?  Until\n> sufficiently large number of people start using and large number of\n> changes gets assigned IDs, it won't become an issue, but then how\n> would it be different from what was raised in the earlier discussion\n> to use the commit object name itself as the change ID for a commit\n> that is not derived from anybody else, and copy that ID to commits\n> that are derived from the original commit as the change ID shared\n> among them?  At least that would give us a much better uniqueness\n> guarantee, wouldn't it?  If you want to be able to tell between\n> commit object name and change ID, I wouldn't object if you encode\n> them using whatever mechanism.\n\nRandom number generators are well suited for this purpose, the only\nrequirement is that the distribution of generated numbers does not\nmassively deviate from a uniform probability distribution.\n\nCommit hashes have 160 bits of information, the change-ids used by\nJujutsu today have 128 bits. That's still plenty to make random\ncollisions practically irrelevant.\n\nAnd anyway, the situation of colliding change-ids will be somewhat of\na normal event due to cherry-picking etc. Jujutsu can deal with this\nsituation just fine today. Colliding change-ids are simply not as much\nof a problem as colliding commit hashes.\n"},{"id":"516542","messageId":"aAgdauFt/mdCY+GZ@ubby","threadId":"63245","inReplyTo":"xmqq8qnr3jji.fsf@gitster.g","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-22T22:51:22Z","receivedAt":"2025-04-22T23:30:12Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Tue, Apr 22, 2025 at 03:42:25PM -0700, Junio C Hamano wrote:\n> \"Remo Senekowitsch\" <remo@buenzli.dev> writes:\n> > Btw. since the thread was started, the implementation in Jujutsu has\n> > been completed and I've been pushing commits with the change-id header\n> > to various remotes for a while now. It works well. Forges can start\n> > taking advantage of it. (I hope I find time to help work on that.)\n> \n> It should work well, until somebody finds your random is not random\n> enough, right?  Unlike our object name that depends on the contents\n> (hence a duplicate unless the cryptographic hash function collides\n> means they are truly the same commit), there is no grabally unique\n\nEh, if the hash function is weak then collisions might not be so rare,\nespecially when intentional.\n\nIn a content addressed storage system hash collisions are (can be) bad,\nbut typically the hash functions used are good enough that collisions\nare tolerable, and one just assumes that if the hash matches then the\ndata is right, and there is no way to verify that that does not require\ntrusting the hash function.\n\n> ID assigner involved in your implementation, right?  [...]\n\nChange IDs for this purpose need not be unique, IMO.  If they get\nduplicated either you can notice the problem before pushing, or you can\naccept the dups and use context to resolve them correctly as needed.\n\nOne could make every commit have the same change ID, as a joke or spam,\nbut presumably most upstreams wouldn't do that.\n\nUsing ticket IDs as change IDs implies a globally unique ID assigner,\nand should work well enough where things like bugzilla are used.\n\nNico\n-- \n"},{"id":"516543","messageId":"D9DKHI316ER9.PNEG774QLFL8@buenzli.dev","threadId":"63245","inReplyTo":"aAgdauFt/mdCY+GZ@ubby","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Remo Senekowitsch","fromEmail":"remo@buenzli.dev","sentAt":"2025-04-22T23:47:29Z","receivedAt":"2025-04-22T23:47:35Z","isPatch":false,"sender":{"key":"remo@buenzli.dev","avatar":"https://gravatar.com/avatar/7df680b096206886db5a2dc983926f314985bbee662eb406cf23c65322cf98b7?d=mp&s=160"},"body":"On Wed Apr 23, 2025 at 12:51 AM CEST, Nico Williams wrote:\n>\n> Using ticket IDs as change IDs implies a globally unique ID assigner,\n> and should work well enough where things like bugzilla are used.\n\nThis email thread contains recurring ideas of stuffing unrelated\nmetadata into the change-id header (patch-id, ticket-id). I think we\nshould be careful not to do that.\n\nThe purpose of the proposed change-id is to identify and track how a\nchange evolves over time. We have talked about how those semantics may\nor may not be clear enough, and that's a good thing to discuss.\n\nIt's also a good thing to discuss potential alternatives. E.g. \"If we\nhave a patch-id, we don't need a change-id.\" I don't agree with that,\nbut it's a good thing to discuss.\n\nBut _deriving a change-id from_ a patch-id or ticket-id or whatever\ncompletely destroys its purpose of tracking how a change evolves over\ntime. The change-id can only do that job if it _doesn't_ have other\nunrelated semantics attached to it.\n\nA patch can contain multiple changes. A ticket can be associated with\nmultiple changes. Deriving a change-id from either of them makes it\nimpossible to identify and track a single change.\n\nSo let's avoid mixing these concepts and talk about them as distinct\nalternatives.\n"},{"id":"516544","messageId":"xmqqy0vr21vq.fsf@gitster.g","threadId":"63245","inReplyTo":"aAgdauFt/mdCY+GZ@ubby","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-04-22T23:49:13Z","receivedAt":"2025-04-22T23:49:17Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nico Williams <nico@cryptonector.com> writes:\n\n> Change IDs for this purpose need not be unique, IMO.  If they get\n> duplicated either you can notice the problem before pushing, or you can\n> accept the dups and use context to resolve them correctly as needed.\n\nNot having to worry about duplicates (as long as you trust Git well\nenough not to worry about object name collisions) is a powerful\nthing, though.  Unless there is a strong reason to stick to slightly\nshorter (like 160-bit vs 128-bit) random string, being able to\ndirectly compute the object name of the very first iteration is also\na plus, if you choose to use its commit object name as its change ID\n(hence you do not have to worry about how you come up with a new ID)\npossibly in some encoded form (if you want to be able to use both\nobject name and change ID in your UI), and have later iterations\nrecord the same change ID in the trailer would be quite simple,\nrobust, and would be just as useful, if not more, as a scheme that\nuses randomly generated string as change ID.\n\nIn any case, I am not all that interested in how change ID is\nassigned than what its semantics would be and how it would help\nkeeping track of the changes made to changes.  Personally, I am\nuninterested in a scheme that does not let me even tell which one\namong many commit objects that share the same change ID is the\noriginal and how they evolved (in other words, how they relate to\neach other).\n\nThanks.\n"},{"id":"516545","messageId":"D9DLAQTMJYU6.RJLLVMQZOICK@buenzli.dev","threadId":"63245","inReplyTo":"aAgWytQNqtLzg2TU@ubby","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Remo Senekowitsch","fromEmail":"remo@buenzli.dev","sentAt":"2025-04-23T00:25:40Z","receivedAt":"2025-04-23T00:25:47Z","isPatch":false,"sender":{"key":"remo@buenzli.dev","avatar":"https://gravatar.com/avatar/7df680b096206886db5a2dc983926f314985bbee662eb406cf23c65322cf98b7?d=mp&s=160"},"body":"On Wed Apr 23, 2025 at 12:23 AM CEST, Nico Williams wrote:\n> On Tue, Apr 22, 2025 at 04:17:03PM -0400, D. Ben Knoble wrote:\n>> On Wed, Apr 16, 2025 at 7:36 AM Remo Senekowitsch <remo@buenzli.dev> wrote:\n>> > On Mon Apr 14, 2025 at 9:54 PM CEST, D. Ben Knoble wrote:\n>\n> [Responding out of order.]\n>\n>> > I think it's the other way around: Building a code review UI built on\n>> > git and centered around interdiffs today is _hard_, that's why we don't\n>> > have it yet. Adding change-ids to commits will make it much easier,\n>> > paving the way for these tools to be implemented.\n>> \n>> Fair point, although GitHub's detection of force-pushes makes me think\n>> it could split a PR into versions at that point, cross-link backwards\n>> and forwards by one version (from the force-push detection), show\n>> range-diffs between versions based on the target branch of the merge,\n>> and even follow the cross-links to show an overall sequence of\n>> versions.\n>\n> GitLab does something like this for review comments where the author\n> updates the commented you get a link you can click on that shows you the\n> differences between version n-1 and version n.  GitLab does this without\n> change IDs.\n>\n> GitLab seems to figure it out -- an existence proof that it can be done.\n> So maybe Junio and Theodore are quite right that similarity checks\n> should be enough.\n\nI haven't used GitLab in a while so I had to test it. I found the\nfeature you describe and it's nice, but it falls short in important and\nsomewhat predicable ways.\n\nFirstly, it only happens when you make a comment on the whole diff.\nThen it shows you an interdiff between the version you commented on and\nthe immediate next version. (I didn't find a way to make it show the\ninterdiff between the commented-on and the _latest_ version, but that\nseems doable implementation wise.)\n\nBut that's not really what we're talking about here. That's taking the\ninterdiff between two versions _of the entire branch_ and selecting a\nrelevant _hunk_ from it. Neat, but completely unrelated.\n\nIf you make a comment on a _specific commit_, the feature doesn't kick\nin at all. GitLab just tells you that this comment was made \"on an\noutdated change in commit xyz\".\n\nA truly patch-based review UI would show the interdiff between the\nprevious version of the specific commit that was commented on and the\nnext or latest version of _that specific patch_.\n\nThat requires tracking how the individual patches evolve and as far as\nI can tell, GitLab makes no attempt to do so. I didn't even give it a\nhard time, I kept the number of commits, the commit messages and author\ntimestamps the same.\n\nSo, while I must say GitLab is doing pretty well here, claiming that\nthey \"figured out\" patch-based review on top of git without change-ids\nis not accurate.\n"},{"id":"516546","messageId":"aAg1ALkWaRQswZtK@ubby","threadId":"63245","inReplyTo":"D9DKHI316ER9.PNEG774QLFL8@buenzli.dev","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-23T00:32:00Z","receivedAt":"2025-04-23T00:47:17Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Wed, Apr 23, 2025 at 01:47:29AM +0200, Remo Senekowitsch wrote:\n> On Wed Apr 23, 2025 at 12:51 AM CEST, Nico Williams wrote:\n> > Using ticket IDs as change IDs implies a globally unique ID assigner,\n> > and should work well enough where things like bugzilla are used.\n> \n> This email thread contains recurring ideas of stuffing unrelated\n> metadata into the change-id header (patch-id, ticket-id). I think we\n> should be careful not to do that.\n\nWe, the users, have been doing this for decades by convention (i.e.,\nstarting commit subject lines with ticket IDs and putting all other\nrelated ticket IDs in the rest of the commit comment.  Why would that be\nwrong _now_?\n\n> The purpose of the proposed change-id is to identify and track how a\n> change evolves over time. We have talked about how those semantics may\n> or may not be clear enough, and that's a good thing to discuss.\n\nAnd ticket IDs happen to be very good at being this sort of change ID.\n\nSun did this starting in the early 90s.  They used a rebase workflow at\nleast since 1992.  Folks still do this in Illumos.  I'm certain that the\nrump Solaris engineering at Oracle still does this.  That's more than\nthree decades of that.\n\nNico\n-- \n"},{"id":"516547","messageId":"D9DMCVD6EG00.317YDVDW95P45@buenzli.dev","threadId":"63245","inReplyTo":"aAg1ALkWaRQswZtK@ubby","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Remo Senekowitsch","fromEmail":"remo@buenzli.dev","sentAt":"2025-04-23T01:15:28Z","receivedAt":"2025-04-23T01:15:40Z","isPatch":false,"sender":{"key":"remo@buenzli.dev","avatar":"https://gravatar.com/avatar/7df680b096206886db5a2dc983926f314985bbee662eb406cf23c65322cf98b7?d=mp&s=160"},"body":"On Wed Apr 23, 2025 at 2:32 AM CEST, Nico Williams wrote:\n> On Wed, Apr 23, 2025 at 01:47:29AM +0200, Remo Senekowitsch wrote:\n>> On Wed Apr 23, 2025 at 12:51 AM CEST, Nico Williams wrote:\n>> > Using ticket IDs as change IDs implies a globally unique ID assigner,\n>> > and should work well enough where things like bugzilla are used.\n>> \n>> This email thread contains recurring ideas of stuffing unrelated\n>> metadata into the change-id header (patch-id, ticket-id). I think we\n>> should be careful not to do that.\n>\n> We, the users, have been doing this for decades by convention (i.e.,\n> starting commit subject lines with ticket IDs and putting all other\n> related ticket IDs in the rest of the commit comment.  Why would that be\n> wrong _now_?\n\nYou can stuff as much free-form metadata into the commit message as you\nwant, because git itself doesn't care much about what's in there. The\nbetter analogy would be to put the names of your mom and dad in the\n\"parent\" header as a free-form piece of metadata about the heritage of\nthe commit author. That's gonna break stuff.\n\nPutting any sort of unrelated metadata into a change-id breaks how it\nworks. There is no reason to do it, you can put that metadata in its\nown, dedicated place. If not the commit message, then at least in its\nown header.\n"},{"id":"516552","messageId":"aAg8HN6sgFu4mj1/@ubby","threadId":"63245","inReplyTo":"xmqqy0vr21vq.fsf@gitster.g","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-23T01:02:20Z","receivedAt":"2025-04-23T04:27:43Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Tue, Apr 22, 2025 at 04:49:13PM -0700, Junio C Hamano wrote:\n> Not having to worry about duplicates (as long as you trust Git well\n> enough not to worry about object name collisions) is a powerful\n> thing, though.  Unless there is a strong reason to stick to slightly\n> [...]\n\nIs that really true?  Yes, because if two different objects have the\nsame hash then the first one pushed to a repository will \"win\" and the\nsecond's content will never appear, right?  I suppose that the server\ncould detect the collision at push time and reject the push to make sure\nthat there's no surprises for the pusher.\n\nBut Git requires objects and refs to be unique _and_ it indexes by them.\nThis is great.  It might be fine for change IDs too, or not.\n\nBut remember that proponents want change IDs to not quite be unique,\nsince the point is that they tie multiple different versions of a commit\nseries together for the purpose of code review.  In the end, when the\ncode review is completed and approved then the change ID might be unique\nagain, but then cherry-picking onto other branches for forward- or\nback-porting might render them non-unique again.\n\nIf you accept change IDs, I highly recommend not requiring them to be\nunique.\n\nNico\n-- \n"},{"id":"516553","messageId":"CAESOdVDG_tfrWMvV6V_Ad76EqXU3Be+EpJDLvtgPcfCRHoJoYQ@mail.gmail.com","threadId":"63245","inReplyTo":"xmqq8qnr3jji.fsf@gitster.g","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Martin von Zweigbergk","fromEmail":"martinvonz@google.com","sentAt":"2025-04-23T05:07:01Z","receivedAt":"2025-04-23T05:07:14Z","isPatch":false,"sender":{"key":"martinvonz@google.com","avatar":"https://avatars.githubusercontent.com/u/891642?v=4"},"body":"On Tue, 22 Apr 2025 at 15:42, Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Remo Senekowitsch\" <remo@buenzli.dev> writes:\n>\n> > Btw. since the thread was started, the implementation in Jujutsu has\n> > been completed and I've been pushing commits with the change-id header\n> > to various remotes for a while now. It works well. Forges can start\n> > taking advantage of it. (I hope I find time to help work on that.)\n>\n> It should work well, until somebody finds your random is not random\n> enough, right?  Unlike our object name that depends on the contents\n> (hence a duplicate unless the cryptographic hash function collides\n> means they are truly the same commit), there is no grabally unique\n> ID assigner involved in your implementation, right?\n\nA forge can decide to enforce that no two commits on the main branch\nhave the same change ID, for example.\n"},{"id":"516554","messageId":"aAg4JR+rCDqO5ljV@ubby","threadId":"63245","inReplyTo":"D9DLAQTMJYU6.RJLLVMQZOICK@buenzli.dev","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-23T00:45:25Z","receivedAt":"2025-04-23T05:18:14Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Wed, Apr 23, 2025 at 02:25:40AM +0200, Remo Senekowitsch wrote:\n> On Wed Apr 23, 2025 at 12:23 AM CEST, Nico Williams wrote:\n> > GitLab seems to figure it out -- an existence proof that it can be done.\n> > So maybe Junio and Theodore are quite right that similarity checks\n> > should be enough.\n> \n> I haven't used GitLab in a while so I had to test it. I found the\n> feature you describe and it's nice, but it falls short in important and\n> somewhat predicable ways.\n\nTrue.\n\n> Firstly, it only happens when you make a comment on the whole diff.\n> Then it shows you an interdiff between the version you commented on and\n> the immediate next version. (I didn't find a way to make it show the\n> interdiff between the commented-on and the _latest_ version, but that\n> seems doable implementation wise.)\n\nI suspect that's a UI design issue: how cluttered do you want the UI to\nbe?  GitLab's UI is already quite cluttered.  Often I struggle to find\nthe thing I want, and I use it often.\n\n> But that's not really what we're talking about here. That's taking the\n> interdiff between two versions _of the entire branch_ and selecting a\n> relevant _hunk_ from it. Neat, but completely unrelated.\n\nI think GitHub and Gitlab both can handle ... commit ranges, can they\nnot?  The problem is that then you have to find refs (or commit hashes)\nfor each version you want to compare, and once again UI complexity gets\nugly.\n\nThis is why I often do code reviews in the terminal, using `git diff`\nand `git log --patch` as needed.  The problem then is that actually\nleaving a comment on the CR requires going back to the browser,\nnavigating to the CR/MR/PR/WhateverR, navigating to the file and diff,\nfinding the relevant hunk, and finally authoring my comment -- this is\nvery painful.  In principle I want to be able to do everything in the\nbrowser because I don't see how to design a TUI that lets me do all of\nthis, but I'm really a shell and TUI type of user, so if we could design\na TUI I'd rather use that than the browser.\n\n> If you make a comment on a _specific commit_, the feature doesn't kick\n> in at all. GitLab just tells you that this comment was made \"on an\n> outdated change in commit xyz\".\n\nI suppose GL could give you a link to a diff view of the same commit in\ndifferent versions.\n\nThe point is that GL demonstrates that these things can be done.  And I\ndon't see how a change ID would have helped GL much except in cases\nwhere one re-does all the commits with different subject lines etc, but\nleaves the actual patches mostly the same.  Now it does happen that I\nsplit and squash commits, but it's rare that I completely redo them.\n\nSo I think I'm coming back to my original view, that change IDs are\nuseful for forward- and back-porting, and not much else, and that\nexisting conventions are probably good enough for forward- and\nback-porting anyways.  Though I don't object to a change ID header, mind\nyou.\n\n> A truly patch-based review UI would show the interdiff between the\n> previous version of the specific commit that was commented on and the\n> next or latest version of _that specific patch_.\n\nI suspect most users are not pedantic like me and don't really go to\nthat sort of length, know they could, would want to, or would if they\nknew they can.\n\n> So, while I must say GitLab is doing pretty well here, claiming that\n> they \"figured out\" patch-based review on top of git without change-ids\n> is not accurate.\n\nThey've demonstrated that this is possible now and without change IDs --\nthat is very relevant here since the assertion has been made that change\nIDs would help do better than most CR tools do w/o them.  GL probably\nhave done what they think their users want them to do.  GL might yet add\nmore of this functionality.\n\nNico\n-- \n"},{"id":"516555","messageId":"aAhwcgzSIaE6l1O9@ubby","threadId":"63245","inReplyTo":"D9DMCVD6EG00.317YDVDW95P45@buenzli.dev","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-23T04:45:38Z","receivedAt":"2025-04-23T05:23:49Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Wed, Apr 23, 2025 at 03:15:28AM +0200, Remo Senekowitsch wrote:\n> You can stuff as much free-form metadata into the commit message as you\n> want, because git itself doesn't care much about what's in there. The\n> better analogy would be to put the names of your mom and dad in the\n> \"parent\" header as a free-form piece of metadata about the heritage of\n> the commit author. That's gonna break stuff.\n\nThe point was that lacking a change ID header people have been resorting\nto conventions, and to point out that people have been using ticket IDs\nas change IDs.  It's just an observation.\n"},{"id":"516556","messageId":"aAhwzs62VPZrWr7+@ubby","threadId":"63245","inReplyTo":"aAg8HN6sgFu4mj1/@ubby","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-23T04:47:10Z","receivedAt":"2025-04-23T07:08:52Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Tue, Apr 22, 2025 at 08:02:20PM -0500, Nico Williams wrote:\n> But remember that proponents want change IDs to not quite be unique,\n> since the point is that they tie multiple different versions of a commit\n> series together for the purpose of code review.  In the end, when the\n> code review is completed and approved then the change ID might be unique\n> again, but then cherry-picking onto other branches for forward- or\n> back-porting might render them non-unique again.\n\nAh, right, the thing about change IDs getting copied to new commits when\ncherry-picking was my suggestion.  I stand by it, but it's not what\nothers are proposing.\n"},{"id":"516604","messageId":"87cyd3f306.fsf@iotcl.com","threadId":"63245","inReplyTo":"aAg4JR+rCDqO5ljV@ubby","subject":"How GitLab does/doesn't need change IDs (was Re: Semantics of change IDs)","fromName":"Toon Claes","fromEmail":"toon@iotcl.com","sentAt":"2025-04-23T12:58:49Z","receivedAt":"2025-04-23T12:59:01Z","isPatch":false,"sender":{"key":"toon@iotcl.com","avatar":"https://avatars.githubusercontent.com/u/121621?v=4"},"body":"Nico Williams <nico@cryptonector.com> writes:\n\n> On Wed, Apr 23, 2025 at 02:25:40AM +0200, Remo Senekowitsch wrote:\n\n>> Firstly, it only happens when you make a comment on the whole diff.\n>> Then it shows you an interdiff between the version you commented on and\n>> the immediate next version. (I didn't find a way to make it show the\n>> interdiff between the commented-on and the _latest_ version, but that\n>> seems doable implementation wise.)\n>\n> I suspect that's a UI design issue: how cluttered do you want the UI to\n> be?  GitLab's UI is already quite cluttered.  Often I struggle to find\n> the thing I want, and I use it often.\n\nI'm not sure if you've been using Merge Requests here, but if you do you\ncan compare each version against another. There are two dropdowns near\nthe top of the \"Changes\" tab: \"Compare <master> and <latest version>\".\nHere you can compare versions against the target branch and other\nversions.\n\n> I think GitHub and Gitlab both can handle ... commit ranges, can they\n> not?  The problem is that then you have to find refs (or commit hashes)\n> for each version you want to compare, and once again UI complexity gets\n> ugly.\n\nAt GitLab we keep track of the commit IDs a branch has been (maybe only\nif there is a Merge Request for that branch, I'm not sure). We can use\nthis information to compare versions. I think, to some people in this\nthread, the use-case for Change-IDs is pretty similar: you want to tie\ninformation about two commits together. If there is a branch, you have\nthis information, but in a mail-based workflow (or in Jujutsu) there is\nno branch. In that case a Change-Id can fill that gap.\n\n>> If you make a comment on a _specific commit_, the feature doesn't kick\n>> in at all. GitLab just tells you that this comment was made \"on an\n>> outdated change in commit xyz\".\n\nAs I mentioned, you might need to create a MR to retain/see this\ninformation, I'm not entirely sure.\n\n> The point is that GL demonstrates that these things can be done.  And I\n> don't see how a change ID would have helped GL much except in cases\n> where one re-does all the commits with different subject lines etc, but\n> leaves the actual patches mostly the same.  Now it does happen that I\n> split and squash commits, but it's rare that I completely redo them.\n\nThat's because GL stores history about a branch ref (outside the Git\nobject/ref database). If you don't do that, you can't. Having a\nChange-Id embedded in the commit, retains that information in Git's DB.\n\n--\nToon\n"},{"id":"516609","messageId":"xmqqjz7a27ww.fsf@gitster.g","threadId":"63245","inReplyTo":"CAESOdVDG_tfrWMvV6V_Ad76EqXU3Be+EpJDLvtgPcfCRHoJoYQ@mail.gmail.com","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-04-23T15:51:11Z","receivedAt":"2025-04-23T15:51:15Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Martin von Zweigbergk <martinvonz@google.com> writes:\n\n> On Tue, 22 Apr 2025 at 15:42, Junio C Hamano <gitster@pobox.com> wrote:\n>>\n>> \"Remo Senekowitsch\" <remo@buenzli.dev> writes:\n>>\n>> > Btw. since the thread was started, the implementation in Jujutsu has\n>> > been completed and I've been pushing commits with the change-id header\n>> > to various remotes for a while now. It works well. Forges can start\n>> > taking advantage of it. (I hope I find time to help work on that.)\n>>\n>> It should work well, until somebody finds your random is not random\n>> enough, right?  Unlike our object name that depends on the contents\n>> (hence a duplicate unless the cryptographic hash function collides\n>> means they are truly the same commit), there is no grabally unique\n>> ID assigner involved in your implementation, right?\n>\n> A forge can decide to enforce that no two commits on the main branch\n> have the same change ID, for example.\n\nWould it make sense, though?  Imagine that a contributor in your\nproject did not refactor code properly and instead made a\ncopy-and-paste duplicates of a very similar code.  I find a bug in\none of them, without realizing that the old mistake of duplicating\ncode (instead of making it a shared helper that is called from the\ntwo places) and create a fix for it.  Later somebody else realizes\nthe same fix is needed for the other copy---attempting to cherry\npick the original fix may find that remaining copy of a buggy code\nas the logic to perform a three-way merge across renames that is\nsufficiently clever kicks in.  Shouldn't these two commits to fix\nthe same bug in two places share the same change ID so that it is\nclear to the later developers that the latter fix was derived from\nthe former one?\n\n\n"},{"id":"516610","messageId":"CAESOdVCjc1kvQSKnxGfNNSTvFhLRjH_vzwMauP8ZWQ5hhfBnEw@mail.gmail.com","threadId":"63245","inReplyTo":"xmqqjz7a27ww.fsf@gitster.g","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Martin von Zweigbergk","fromEmail":"martinvonz@google.com","sentAt":"2025-04-23T16:19:17Z","receivedAt":"2025-04-23T16:19:30Z","isPatch":false,"sender":{"key":"martinvonz@google.com","avatar":"https://avatars.githubusercontent.com/u/891642?v=4"},"body":"On Wed, 23 Apr 2025 at 08:51, Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Martin von Zweigbergk <martinvonz@google.com> writes:\n>\n> > On Tue, 22 Apr 2025 at 15:42, Junio C Hamano <gitster@pobox.com> wrote:\n> >>\n> >> \"Remo Senekowitsch\" <remo@buenzli.dev> writes:\n> >>\n> >> > Btw. since the thread was started, the implementation in Jujutsu has\n> >> > been completed and I've been pushing commits with the change-id header\n> >> > to various remotes for a while now. It works well. Forges can start\n> >> > taking advantage of it. (I hope I find time to help work on that.)\n> >>\n> >> It should work well, until somebody finds your random is not random\n> >> enough, right?  Unlike our object name that depends on the contents\n> >> (hence a duplicate unless the cryptographic hash function collides\n> >> means they are truly the same commit), there is no grabally unique\n> >> ID assigner involved in your implementation, right?\n> >\n> > A forge can decide to enforce that no two commits on the main branch\n> > have the same change ID, for example.\n>\n> Would it make sense, though?  Imagine that a contributor in your\n> project did not refactor code properly and instead made a\n> copy-and-paste duplicates of a very similar code.  I find a bug in\n> one of them, without realizing that the old mistake of duplicating\n> code (instead of making it a shared helper that is called from the\n> two places) and create a fix for it.  Later somebody else realizes\n> the same fix is needed for the other copy---attempting to cherry\n> pick the original fix may find that remaining copy of a buggy code\n> as the logic to perform a three-way merge across renames that is\n> sufficiently clever kicks in.  Shouldn't these two commits to fix\n> the same bug in two places share the same change ID so that it is\n> clear to the later developers that the latter fix was derived from\n> the former one?\n\nMaybe it depends on how the forge uses the change id. If the forge is\nGerrit, it will use the (change id, target branch) to identify a\nreview (IIUC), so then it will require a new change ID because it\nrequires a new review. If it doesn't have that requirement (maybe it's\nPR-style forge), then it could at least highlight to the reviewer that\nthere was an old version of the change that has already been merged.\n\nAs Remo said, we've implemented this feature in Jujutsu and we'll\nprobably enable it by default soon. I think GitButler has had it\nenabled for quite some time. So maybe we'll see in a year or two what\nproblems it causes :)\n"},{"id":"516628","messageId":"aAk4kf2EM9pXaHZG@ubby","threadId":"63245","inReplyTo":"87cyd3f306.fsf@iotcl.com","subject":"Re: How GitLab does/doesn't need change IDs (was Re: Semantics of change IDs)","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-04-23T18:59:29Z","receivedAt":"2025-04-23T18:59:33Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Wed, Apr 23, 2025 at 02:58:49PM +0200, Toon Claes wrote:\n> Nico Williams <nico@cryptonector.com> writes:\n> \n> At GitLab we keep track of the commit IDs a branch has been (maybe only\n> if there is a Merge Request for that branch, I'm not sure). [...]\n\nDo you mean \"we keep track of the commit _hashes_ a branch has _seen_\"?\nBut it can't be commit hashes, and there's no commit IDs, so GL could be\nassigning synthetic, internal commit IDs based on commit similarity,\nwhich proves Junio's and Theodore's point that similarity checking can\nbe enough.\n\n> > The point is that GL demonstrates that these things can be done.  And I\n> > don't see how a change ID would have helped GL much except in cases\n> > where one re-does all the commits with different subject lines etc, but\n> > leaves the actual patches mostly the same.  Now it does happen that I\n> > split and squash commits, but it's rare that I completely redo them.\n> \n> That's because GL stores history about a branch ref (outside the Git\n> object/ref database). If you don't do that, you can't. Having a\n> Change-Id embedded in the commit, retains that information in Git's DB.\n\nI.e., GL has an internal reflog on the server side.  I've sometimes\nwished that I could push and fetch reflogs (or subsets thereof anyways).\n\nWhen doing code reviews I use [local, obv.] reflogs to see the diffs\nbetween an earlier version of a branch that I fetched and reviewed\nearlier and the latest that I just fetched and am reviewing, and\ngenerally I don't need to see any other versions I never fetched, but\noccasionally I've wished I could fetch those other versions, but since\nthere are no server-side refs for them, I can't.  [Or maybe I'm about to\nlearn of some feature I didn't know about :)]\n\nI agree that change IDs / commit IDs in commit headers can help one keep\ntrack of versions of a branch w/o a server-side reflog, but how would\nyou keep track of their chnronology?  I.e., how do you know which is\nversion 1, which is version 2, .., and which is version N-1?  (Version N\nbeing the head of the branch.)  If you don't index these then finding\nthem is a full table scan, and if you index them then you've implemented\na server-side reflog.\n\nWhich makes me think that all that's needed for a good CR tool here is\na) a server-side reflog, b) similarity checking for commits.  (a)\ndoesn't seem like a radical idea (that can be implemented with server\nside hooks), and (b) is also not radical given that file rename / copy\noperations are detected by Git using similarity checking already.\n\nFrom a UI/UX perspective not having to take extra steps to get those\nchange IDs into the commits is nice and user-friendly.\n\nSo I've come around to not wanting a change ID header :)  For the back-\nand forward-port use-case having commit subject conventions that make\nuse of \"ticket IDs\" or similar is a 90% solution that many users have\nbeen living with for decades.\n\nNico\n-- \n"},{"id":"517758","messageId":"CALnO6CBq2cqBAhzMh8rnXzc8cPTsB4hz98YVn3B4+PGdiyn9_A@mail.gmail.com","threadId":"63245","inReplyTo":"aAgWytQNqtLzg2TU@ubby","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"D. Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2025-05-10T19:32:17Z","receivedAt":"2025-05-10T19:32:30Z","isPatch":false,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"On Tue, Apr 22, 2025 at 6:23 PM Nico Williams <nico@cryptonector.com> wrote:\n>\n> On Tue, Apr 22, 2025 at 04:17:03PM -0400, D. Ben Knoble wrote:\n> > On Wed, Apr 16, 2025 at 7:36 AM Remo Senekowitsch <remo@buenzli.dev> wrote:\n> > > On Mon Apr 14, 2025 at 9:54 PM CEST, D. Ben Knoble wrote:\n>\n> [Responding out of order.]\n>\n> > > I think it's the other way around: Building a code review UI built on\n> > > git and centered around interdiffs today is _hard_, that's why we don't\n> > > have it yet. Adding change-ids to commits will make it much easier,\n> > > paving the way for these tools to be implemented.\n> >\n> > Fair point, although GitHub's detection of force-pushes makes me think\n> > it could split a PR into versions at that point, cross-link backwards\n> > and forwards by one version (from the force-push detection), show\n> > range-diffs between versions based on the target branch of the merge,\n> > and even follow the cross-links to show an overall sequence of\n> > versions.\n>\n> GitLab does something like this for review comments where the author\n> updates the commented you get a link you can click on that shows you the\n> differences between version n-1 and version n.  GitLab does this without\n> change IDs.\n\nDoes GitLab's output resemble that of git-range-diff? If not, it's not\nquite what I had in mind ;)\n\nGitHub's \"compare versions\" button (after a force-push) is more like\n\"git diff\" on the 2 trees, which is only a fraction of the relevant\ninformation.\n\n[…]\n\n> > But the treatment below is something I agree with and failed to\n> > convey, I think: namely, URIs seem to encode too much\n> > \"unportable\"/\"specific\" information in Git. I feel like the current\n> > design is not really \"tool-agnostic\" as much as \"built on a universal\n> > core.\" That seems valuable and prone to more longevity.\n>\n> On the other hand URIs can easily be dereferenced, as long as they are\n> not rotted.\n\nRot happens, though… your point is well-taken! Most of Git seems to be\nvaluable in the presence of linkrot, though, and I don't see how\nadding URIs keeps that principle alive.\n\n> > > > Which begs another question: what/who applies the initial change ID to\n> > > > a commit and when?\n> > >\n> > > These are all great questions, which the originally proposed format\n> > > (fixed-width reverse-hex) has answers to. I think a URI would be\n> > > strictly worse.\n> >\n> > Well, I think we still missed \"what/who applies the initial change ID\n> > to a commit and when.\"\n>\n> Some possible answers:\n>\n>  - The user constructs the change ID from other things like\n>    bugzilla/jira/... ticket IDs.\n>\n>    This is OK, but probably not what OP had in mind.\n>\n>    If the user needs additional sub-IDs just add a -{N} suffix.\n>\n>    It's on the user and utilities to keep usage consistent.  CR tooling\n>    could refuse to allow reuse of change IDs after a CR is merged.\n\nYep; I think trailers are the common version of this today. And ofc,\nsee linkrot.\n\n>  - CR tooling edits your history to add that change ID when you create a\n>    CR\n>\n>    This feels wrong.\n\nYep.\n\n>  - CR tooling creates an 'empty' CR just to acquire a change ID that the\n>    user can then add to their commits, either manually or with utility\n>    that does it for them (but which they invoke).\n>\n>    This is alright, though a bit annoying.  If the user fails to set the\n>    change ID, then it's no worse than today.\n>\n>  - The user adds it when they update an existing CR by manually editing\n>    their history or invoking a utility that does it for them.  The\n>    change ID comes from the CR.\n>\n>    This is alright, though a bit annoying.  If the user fails to set the\n>    change ID, then it's no worse than today.\n>\n>  - The CR tooling adds an empty commit at the head just to hold the\n>    change ID and any additional metadata for each of the commits in the\n>    series.\n>\n>    The user has to remember to fetch this commit, unless it's generated\n>    locally, but they're going to fetch anyways so this is ok.\n>\n>    IMO empty commits are annoying.  (That includes merge commits, but\n>    then I'm a rebase workflow / linear history zealot, so merge commits\n>    are also annoying because merge workflows are annoying.)  But I would\n>    think that empty commits that include metadata about code review are\n>    reasonable and tolerable, especially if their subject lines are\n>    descriptive, like \"CR: {cr title here}\" or \"CR {ID}: {title}\", say.\n>\n> I think I'd only really like the first and last of the above.\n>\n> Nico\n> --\n\nThanks for thinking out some concrete options on where IDs come from!\nI think I've gotten even more confused on how they are supposed to be\nused (which should probably inform the implementation), but hopefully\nwe're getting somewhere :)\n\n-- \nD. Ben Knoble\n"},{"id":"517759","messageId":"CALnO6CD8JTnNGfuCtb1QKFhx+Vv1txUZ+wCL1nZCDGAvHx6A6g@mail.gmail.com","threadId":"63245","inReplyTo":"CALnO6CBq2cqBAhzMh8rnXzc8cPTsB4hz98YVn3B4+PGdiyn9_A@mail.gmail.com","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"D. Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2025-05-10T19:46:10Z","receivedAt":"2025-05-10T19:46:22Z","isPatch":false,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"On Sat, May 10, 2025 at 3:32 PM D. Ben Knoble <ben.knoble@gmail.com> wrote:\n[…]\n>\n> Thanks for thinking out some concrete options on where IDs come from!\n> I think I've gotten even more confused on how they are supposed to be\n> used (which should probably inform the implementation), but hopefully\n> we're getting somewhere :)\n>\n> --\n> D. Ben Knoble\n\n[Trying to summarize?]\n\nOn a re-read of\nhttps://lore.kernel.org/git/CANiSa6gwup5vXU235mG+Ybbc+P=SbwoNFEmuhg=iYu0yGvSXVA@mail.gmail.com/,\nI see that change IDs were motivated partly by identifying (related?)\ncommits after rewrites. I can certainly see how it would be nice to\ntrack down how a commit I'm working on evolved; I can even imagine\nmost of the problems brought up in this thread wrt splitting or\ncombining commits (not to mention, say, cherry-picks where the\ncommitter makes non-trivial changes to the patch).\n\nThere was also a note about using a change ID to identify a code\nreview in supporting tools. Neat!\n\nI'll leave it to someone else to summarize the open questions? (I now\nhave a few of my own about how tools in Gits ecosystem respond to…\nunexpected… headers.)\n\nIn the meantime, I think I'll repost this, since I'm not sure I ever\ngot clarity:\n\nRe-reading the original post [1] (which didn't mention this kind of\nID?), I'm having a hard time seeing the problem statement. There's a\nlot said here about the specifics of the solution, and some other neat\nthings it might unlock… meanwhile, I'm wondering if all the\nconsternation about change IDs is because the problem being solved is\nunderspecified for a core Git feature? (That might tie to Ted's\ninitial concerns about semantic meaning, on which I think I concur:\nthe parent and committer/author headers have unambiguous meaning to\nGit, independent of anything else.)\n\nIt looks to me, an outsider, like the problem is some combination of\n\"I want to track a commit's evolution\" and \"I want to see related\ncommits in review, esp. when it's an identical and already-approved\ncommit.\" But I might be misreading, and clarifying the problem\nstatement might help bring us to a better core solution?\n\n[1]: https://lore.kernel.org/git/xmqqh62tm5fo.fsf@gitster.g/T/#m038be849b9b4020c16c562d810cf77bad91a2c87\n\n-- \nD. Ben Knoble\n"},{"id":"517761","messageId":"CAESOdVCKTnUbVuXq-=F3df4i2T-GcDpJMENr8wwm-ZXR95+59w@mail.gmail.com","threadId":"63245","inReplyTo":"CALnO6CD8JTnNGfuCtb1QKFhx+Vv1txUZ+wCL1nZCDGAvHx6A6g@mail.gmail.com","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Martin von Zweigbergk","fromEmail":"martinvonz@google.com","sentAt":"2025-05-10T20:31:32Z","receivedAt":"2025-05-10T20:31:45Z","isPatch":false,"sender":{"key":"martinvonz@google.com","avatar":"https://avatars.githubusercontent.com/u/891642?v=4"},"body":"Hi,\n\nOn Sat, 10 May 2025 at 12:46, D. Ben Knoble <ben.knoble@gmail.com> wrote:\n>\n> On a re-read of\n> https://lore.kernel.org/git/CANiSa6gwup5vXU235mG+Ybbc+P=SbwoNFEmuhg=iYu0yGvSXVA@mail.gmail.com/,\n> I see that change IDs were motivated partly by identifying (related?)\n> commits after rewrites. I can certainly see how it would be nice to\n> track down how a commit I'm working on evolved; I can even imagine\n> most of the problems brought up in this thread wrt splitting or\n> combining commits (not to mention, say, cherry-picks where the\n> committer makes non-trivial changes to the patch).\n>\n> There was also a note about using a change ID to identify a code\n> review in supporting tools. Neat!\n>\n> I'll leave it to someone else to summarize the open questions? (I now\n> have a few of my own about how tools in Gits ecosystem respond to…\n> unexpected… headers.)\n>\n> In the meantime, I think I'll repost this, since I'm not sure I ever\n> got clarity:\n>\n> Re-reading the original post [1] (which didn't mention this kind of\n> ID?), I'm having a hard time seeing the problem statement. There's a\n> lot said here about the specifics of the solution, and some other neat\n> things it might unlock… meanwhile, I'm wondering if all the\n> consternation about change IDs is because the problem being solved is\n> underspecified for a core Git feature? (That might tie to Ted's\n> initial concerns about semantic meaning, on which I think I concur:\n> the parent and committer/author headers have unambiguous meaning to\n> Git, independent of anything else.)\n>\n> It looks to me, an outsider, like the problem is some combination of\n> \"I want to track a commit's evolution\" and \"I want to see related\n> commits in review, esp. when it's an identical and already-approved\n> commit.\" But I might be misreading, and clarifying the problem\n> statement might help bring us to a better core solution?\n\nTo me, the main benefit is being able to refer to an evolving change\nby a stable ID. That enables things like `jj describe qx -m 'new\ndescription'; jj new qx` (update commit message, then switch to it)\nwithout having to look up the new commit ID after setting the\ndescription. That's sufficient benefit for me, and I think most\nJujutsu users would agree. That's basically the only benefit we've\ngotten from it so far since we have not started transferring it to\nremotes. (There are other minor benefits like being able to highlight\nto the user if they have two related commits so they may want to\ndelete one or somehow combine them.)\n\nGiven that we already have this stable ID, it would be nice to also\ntransfer it to remotes and have it be preserved by the remote,\nincluding when the remote rewrites the commit. If we can use it for\nthings like identifying a code review so we don't need to link it\nusing a `Change-Id:` commit footer, then that's even better.\n\nIf we instead had something like Mercurial's Changeset Evolution\n(explicitly recording how commits have evolved), then we could have a\nsimilar identifier that was based on the original version of a commit.\nTo make lookup by this kind of change ID faster, we could have an\nindex from commit ID to change ID (i.e. original commit ID). This\nseems to imply a commit can have 0 or 1 predecessors (0 for brand new\ncommits, 1 for rewrites), which is different from Mercurial's\nChangeset Evolution, but not necessarily bad. For this kind of change\nID to be the same across repos, and assuming the predecessor pointer\nis stored in the commit, we need to make sure to transfer all commits\nback to the original commit when we push to a remote. As I think we've\ntalked about before here, that can be problematic because the user has\nto be careful to check that the intermediate commits did not have\nanything sensitive in them. It's also often wasteful to share all the\nintermediate commits with other developers. Another option is to\ntransfer the predecessor pointer outside of the commit object. That\nhas its own problems, like being able to create cycles in the\npredecessor graph.\n\n>\n> [1]: https://lore.kernel.org/git/xmqqh62tm5fo.fsf@gitster.g/T/#m038be849b9b4020c16c562d810cf77bad91a2c87\n>\n> --\n> D. Ben Knoble\n"},{"id":"517873","messageId":"xmqqtt5pu5g8.fsf@gitster.g","threadId":"63245","inReplyTo":"CAESOdVCKTnUbVuXq-=F3df4i2T-GcDpJMENr8wwm-ZXR95+59w@mail.gmail.com","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-05-12T17:03:35Z","receivedAt":"2025-05-12T17:03:38Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Martin von Zweigbergk <martinvonz@google.com> writes:\n\n> If we instead had something like Mercurial's Changeset Evolution\n> (explicitly recording how commits have evolved), then we could have a\n> similar identifier that was based on the original version of a commit.\n> To make lookup by this kind of change ID faster, we could have an\n> index from commit ID to change ID (i.e. original commit ID). This\n> seems to imply a commit can have 0 or 1 predecessors (0 for brand new\n> commits, 1 for rewrites), which is different from Mercurial's\n> Changeset Evolution, but not necessarily bad. For this kind of change\n> ID to be the same across repos, and assuming the predecessor pointer\n> is stored in the commit, we need to make sure to transfer all commits\n> back to the original commit when we push to a remote. As I think we've\n> talked about before here, that can be problematic because the user has\n> to be careful to check that the intermediate commits did not have\n> anything sensitive in them. It's also often wasteful to share all the\n> intermediate commits with other developers. Another option is to\n> transfer the predecessor pointer outside of the commit object. That\n> has its own problems, like being able to create cycles in the\n> predecessor graph.\n\nA few comments (not necessarily strong suggestions).\n\n - I do not think you need to limit the predecessor pointers to 0 or\n   1; when you started from N commits and worked to produce the\n   final single commit, the result would naturally have N predecessors.\n\n - The predecessor pointers do not necessarily have to participate\n   in the object transfer, just like filtered/lazy clones can ignore\n   the tree pointer in a commit object when making a commit-only\n   clone and the contained trees are fetched from the promisor\n   on-demand.  It can even be set to be filtered out by default,\n   since it would make unnecessary transfer cost people would not\n   care most of the time, and only made available when the user\n   expresses that they want to know how the change resulted in the\n   current shape.\n\n"},{"id":"517880","messageId":"CAESOdVD-8j9k2Dq9WgiR9WWO09mpfR9Xxe3pMUWg-KoTfELG8w@mail.gmail.com","threadId":"63245","inReplyTo":"xmqqtt5pu5g8.fsf@gitster.g","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Martin von Zweigbergk","fromEmail":"martinvonz@google.com","sentAt":"2025-05-12T17:19:16Z","receivedAt":"2025-05-12T17:19:29Z","isPatch":false,"sender":{"key":"martinvonz@google.com","avatar":"https://avatars.githubusercontent.com/u/891642?v=4"},"body":"On Mon, 12 May 2025 at 10:03, Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Martin von Zweigbergk <martinvonz@google.com> writes:\n>\n> > If we instead had something like Mercurial's Changeset Evolution\n> > (explicitly recording how commits have evolved), then we could have a\n> > similar identifier that was based on the original version of a commit.\n> > To make lookup by this kind of change ID faster, we could have an\n> > index from commit ID to change ID (i.e. original commit ID). This\n> > seems to imply a commit can have 0 or 1 predecessors (0 for brand new\n> > commits, 1 for rewrites), which is different from Mercurial's\n> > Changeset Evolution, but not necessarily bad. For this kind of change\n> > ID to be the same across repos, and assuming the predecessor pointer\n> > is stored in the commit, we need to make sure to transfer all commits\n> > back to the original commit when we push to a remote. As I think we've\n> > talked about before here, that can be problematic because the user has\n> > to be careful to check that the intermediate commits did not have\n> > anything sensitive in them. It's also often wasteful to share all the\n> > intermediate commits with other developers. Another option is to\n> > transfer the predecessor pointer outside of the commit object. That\n> > has its own problems, like being able to create cycles in the\n> > predecessor graph.\n>\n> A few comments (not necessarily strong suggestions).\n>\n>  - I do not think you need to limit the predecessor pointers to 0 or\n>    1; when you started from N commits and worked to produce the\n>    final single commit, the result would naturally have N predecessors.\n>\n>  - The predecessor pointers do not necessarily have to participate\n>    in the object transfer, just like filtered/lazy clones can ignore\n>    the tree pointer in a commit object when making a commit-only\n>    clone and the contained trees are fetched from the promisor\n>    on-demand.  It can even be set to be filtered out by default,\n>    since it would make unnecessary transfer cost people would not\n>    care most of the time, and only made available when the user\n>    expresses that they want to know how the change resulted in the\n>    current shape.\n\nSorry, what I meant that those two things (limiting it to 0 or 1\npredecessor and transferring predecessors to the remote) would be\nneeded if you want to be able to use the predecessor pointers to infer\na \"change ID\" that allows you to do things like `jj describe qx -m\n'new\ndescription'; jj new qx`. Oh, I suppose the former can be relaxed by\nsimply saying that the change ID is defined as (or derived from) the\nlast commit ID you reach by walking predecessors backwards by\nfollowing the first predecessor pointer if there are several. You\nstill need to transfer at least that whole chain to the remote if you\nwant to make sure that others agree about the change ID (though you\nonly need to transfer the chain of first predecessors).\n"},{"id":"517913","messageId":"CAESOdVD_Cse6AjwLb-4QKjdo4ESWwF3FzSS5JaHbE6ZrMjFeZw@mail.gmail.com","threadId":"63245","inReplyTo":"aCJi+4q6DZhnfdy+@ubby","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Martin von Zweigbergk","fromEmail":"martinvonz@google.com","sentAt":"2025-05-12T21:43:46Z","receivedAt":"2025-05-12T21:43:59Z","isPatch":false,"sender":{"key":"martinvonz@google.com","avatar":"https://avatars.githubusercontent.com/u/891642?v=4"},"body":"On Mon, 12 May 2025 at 14:07, Nico Williams <nico@cryptonector.com> wrote:\n>\n> On Sat, May 10, 2025 at 01:31:32PM -0700, Martin von Zweigbergk wrote:\n> > To me, the main benefit is being able to refer to an evolving change\n> > by a stable ID. That enables things like `jj describe qx -m 'new\n> > description'; jj new qx` (update commit message, then switch to it)\n> > without having to look up the new commit ID after setting the\n> > description.\n>\n> Notionally this is not different from renaming a file.  You have a name\n> (file name, commit message subject) and you have the thing it refers to\n> (file contents, tree object).\n>\n>   <insert sub-thread about why Git does not have inode numbers for\n>    files, does not record rename/copy intent, and depends on file\n>    content similarity checks to detect renames>\n>\n> If Git can do file content similarity checking to discover renames, then\n> surely so can jj and other CR tools do commit similarity checking to\n> discover commit message changes.  Is there anything that makes the\n> preceding statement incorrect?\n\nThat wouldn't work in the `jj describe qx -m 'new description'; jj new\nqx` example I used above, right? I think you're suggesting that when\nthe user runs `jj describe qx -m 'new description'`, we should compare\nthe reachable commits before the command to the reachable commits\nafter the command and then record in some storage that the new commit\nis part of the same \"change\" as the old commit. Is that what you\nmeant? In this particular case, the commit message obviously changed,\nso comparing the commit messages will obviously fail. We could of\ncourse make this command record the information itself, however.\n\n> > Given that we already have this stable ID, [...]\n>\n> \"We\" == jujutsu?\n\nYes, sorry :)\n\n> How is this stable ID constructed?\n\nIt's just random bytes (16 when using the Git backend, 32 in the\nGoogle backend).\n\n> How would things other than jj construct these?  We spent many messages\n> trying to work that out and in my estimate that wasn't settled.\n\nRandom bytes has worked well for jj.\n"},{"id":"517914","messageId":"aCJwgWaNoBVjvImJ@tapette.crustytoothpaste.net","threadId":"63245","inReplyTo":"CAESOdVD_Cse6AjwLb-4QKjdo4ESWwF3FzSS5JaHbE6ZrMjFeZw@mail.gmail.com","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2025-05-12T22:04:49Z","receivedAt":"2025-05-12T22:04:51Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2025-05-12 at 21:43:46, Martin von Zweigbergk wrote:\n> On Mon, 12 May 2025 at 14:07, Nico Williams <nico@cryptonector.com> wrote:\n> >\n> > How is this stable ID constructed?\n> \n> It's just random bytes (16 when using the Git backend, 32 in the\n> Google backend).\n> \n> > How would things other than jj construct these?  We spent many messages\n> > trying to work that out and in my estimate that wasn't settled.\n> \n> Random bytes has worked well for jj.\n\nI would like to suggest that we use a deterministic approach.  People\nrely on Git commits being deterministic, including in my stash\nimport/export series[0].  In addition, it's important to avoid any\nallegations of side channels or leaking information in commits, which\nwould be a concern in many environments and which a deterministic\napproach would avoid[1].\n\nI'd suggest a simple SHA-256 hash of the original commit data (for both\nSHA-1 and SHA-256 commits, but one that would change to a new hash if we\nadded one) or an HMAC-SHA-256 with a fixed and documented key.\n\nI would also recommend a config option to avoid creating these IDs for\nthose who don't want them included for privacy reasons.  I expect to set\nsuch an option, for instance.\n\n[0] That series will definitely require that they be disabled when\ncreating commits, since the goal is to ensure bit-for-bit\nreproducibility between different Git versions so that users can\nimmediately tell if the stash history is identical.\n\n[1] For instance, it's an easy way to leak keys or other credentials\nwithout people noticing just by pushing an innocuous-looking commit.\n-- \nbrian m. carlson (they/them)\nToronto, Ontario, CA\n"},{"id":"517984","messageId":"CALnO6CBSyCyJ_veinUndZNxBnDuwY4cn3RZu7Jcd3bM7pVV5xw@mail.gmail.com","threadId":"63245","inReplyTo":"CAESOdVD_Cse6AjwLb-4QKjdo4ESWwF3FzSS5JaHbE6ZrMjFeZw@mail.gmail.com","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"D. Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2025-05-13T21:22:09Z","receivedAt":"2025-05-13T21:22:22Z","isPatch":false,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"On Mon, May 12, 2025 at 5:43 PM Martin von Zweigbergk\n<martinvonz@google.com> wrote:\n>\n> On Mon, 12 May 2025 at 14:07, Nico Williams <nico@cryptonector.com> wrote:\n> >\n> > On Sat, May 10, 2025 at 01:31:32PM -0700, Martin von Zweigbergk wrote:\n> > > To me, the main benefit is being able to refer to an evolving change\n> > > by a stable ID. That enables things like `jj describe qx -m 'new\n> > > description'; jj new qx` (update commit message, then switch to it)\n> > > without having to look up the new commit ID after setting the\n> > > description.\n> >\n> > Notionally this is not different from renaming a file.  You have a name\n> > (file name, commit message subject) and you have the thing it refers to\n> > (file contents, tree object).\n> >\n> >   <insert sub-thread about why Git does not have inode numbers for\n> >    files, does not record rename/copy intent, and depends on file\n> >    content similarity checks to detect renames>\n> >\n> > If Git can do file content similarity checking to discover renames, then\n> > surely so can jj and other CR tools do commit similarity checking to\n> > discover commit message changes.  Is there anything that makes the\n> > preceding statement incorrect?\n>\n> That wouldn't work in the `jj describe qx -m 'new description'; jj new\n> qx` example I used above, right? I think you're suggesting that when\n> the user runs `jj describe qx -m 'new description'`, we should compare\n> the reachable commits before the command to the reachable commits\n> after the command and then record in some storage that the new commit\n> is part of the same \"change\" as the old commit. Is that what you\n> meant? In this particular case, the commit message obviously changed,\n> so comparing the commit messages will obviously fail. We could of\n> course make this command record the information itself, however.\n\nI didn't follow the entirety of the example (since I think you'd have\nto have change ID \"qx\" to start with in that case), but:\n\nI think you'd compare commit /contents/ (aka the trees they point to)\nrather than /messages/, just like how rename detection compares blobs\nor trees? (Although I seem to recall a recent thread where a heuristic\ninvolving the old/new name went wrong because it was too short, so a\nheuristic in the messages is probably also reasonable. Doesn't\nrange-diff do something similar?)\n\n>\n> > > Given that we already have this stable ID, [...]\n> >\n> > \"We\" == jujutsu?\n>\n> Yes, sorry :)\n>\n> > How is this stable ID constructed?\n>\n> It's just random bytes (16 when using the Git backend, 32 in the\n> Google backend).\n>\n> > How would things other than jj construct these?  We spent many messages\n> > trying to work that out and in my estimate that wasn't settled.\n>\n> Random bytes has worked well for jj.\n\n\n\n-- \nD. Ben Knoble\n"},{"id":"518042","messageId":"xmqqjz6jb6kd.fsf@gitster.g","threadId":"63245","inReplyTo":"CAESOdVD-8j9k2Dq9WgiR9WWO09mpfR9Xxe3pMUWg-KoTfELG8w@mail.gmail.com","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-05-14T14:38:58Z","receivedAt":"2025-05-14T14:39:01Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Martin von Zweigbergk <martinvonz@google.com> writes:\n\n> Sorry, what I meant that those two things (limiting it to 0 or 1\n> predecessor and transferring predecessors to the remote) would be\n> needed if you want to be able to use the predecessor pointers to infer\n> a \"change ID\" that allows you to do things like `jj describe qx -m\n> 'new\n> description'; jj new qx`.\n\nIf we limit the data model so that it cannot represent one commit\nbecoming split into two (or vice versa), and instead one old commit\nalways corresponds to one new commit, our design of convenience\noperations may become a lot easier and simpler. No question about\nthat. I am not sure that is not like a tail wagging a dog, though.\nIf most use cases are one-to-one without split/merge, then you can\nstill perform operations that require one-to-one correspondence in\nmost cases where \"predecessor pointer\" and \"change ID\" are moral\nequivalents, no?\n\nIn any case, that wasn't what I primarily was interested to talk\nabout.\n\nIt sounds like the \"change ID\" being discussed is a simple and\nuseful thing that can and should be a commit trailer that is carried\nforward automatically across \"git commit --amend\", \"git rebase\", and\n\"git cherry-pick\", i.e. those commands that duplicate a new copy, a\nrefinement, of an existing commit object, without any additional\nsupport from Git proper.  More importantly, unlike the \"predecessor\"\nthing, which may be helped if Git knew that it can optionally affect\nreachability, it does not look like it needs any meaning that needs\nstructural support from Git internals.\n\nSo I do appreciate that the wider Git ecosystem like Gerrit and JJ\nare talking to adopt \"Change ID\" with the same syntax and semantics\n(if they all can agree, that is), but I do not think it needs to\naffect Git at the object level.  From what I have heard so far, it\ncertainly does not fit in the commit and tag object headers, but\nmore like the usual trailer thing.\n"},{"id":"518045","messageId":"a07f1102-1596-45b0-bdef-346f4fdcc3fe@app.fastmail.com","threadId":"63245","inReplyTo":"xmqqwmbuybhg.fsf@gitster.g","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Kristoffer Haugsbakk","fromEmail":"kristofferhaugsbakk@fastmail.com","sentAt":"2025-05-14T15:08:53Z","receivedAt":"2025-05-14T15:09:15Z","isPatch":false,"sender":{"key":"kristofferhaugsbakk@fastmail.com","avatar":null},"body":"What a thread!\n\nOn Tue, Apr 8, 2025, at 16:27, Junio C Hamano wrote:\n> Martin von Zweigbergk <martinvonz@google.com> writes:\n>\n>>> A set of individual commits that share the same \"change ID\" is,\n>>> unlike reflog entries which is an ordered set of tip of topics, not\n>>> inherently ordered.  This is inevitable in the distributed world\n>>> where many people can simultaneously work on improving a single\n>>> \"change\" in many different ways, but making it difficult if not\n>>> impossible to see how things evolved, simply because you first need\n>>> to figure out the order of these commits that share the same \"change\n>>> ID\".  Some may be independently evolved from the same ancestor\n>>> iteration.  Some may be repeatedly worked on on a single strand of\n>>> pearls (much like how development recorded in reflog entries of a\n>>> single branch in a single user set-up goes).  I guess you would need\n>>> a way to record the predecessor vs successor relationship of various\n>>> commits that share the same \"change ID\", much like commits form DAG\n>>> to represent ancestor vs descendant relationship.\n>>\n>> That is correct. The change ID should be sufficient for handling\n>> simple distributed cases involving a single remote but it's not a full\n>> replacement for something like Mercurial's Changeset Evolution [1].\n>\n> Just a random thought.  We could very easily replace \"change ID\"\n> with a concept of predecessor-successor commits.\n>\n> Just like we can represent parents-children NxM transitive relation\n> only with 0 or more \"parent\" commit object headers, we can record\n> zero or more \"predecessor\" trailer in the commit log.\n>\n>  (1) a commit with no \"predecessor\" is like \"root commit\" in the\n>      commit history topology.  It is a brand new change that took\n>      inspiration from nobody else and that is not a polished form of\n>      any other existing commit.\n>\n>  (2) a commit created as a refinement for one or more existing\n>      commits record each of them as \"predecessor\" to it.  Having\n>      more than one of them is like a \"merge commit\" in the commit\n>      history topology and represents that two patches were squashed\n>      into one.\n>\n>  (3) Splitting an originally large change into multiple changes can\n>      be represented the same way.  They share the same commit as\n>      their \"predecessor\".  Perhaps you have originally two-commit\n>      series, A and B, and split them differently in such a way that\n>      C has half of a and D has the rest of A plus B.  In which case,\n>      C has A as its predecessor while D has both A and B as its\n>      predecessor.\n>\n>  (4) Just like we can use auxiliary data structures like bitmaps to\n>      figure out reachability without following all the links in the\n>      commit history topology, we should be able to learn how a new\n>      change was born, and trace how it evolved into newer iteration\n>      of the moral equivalent of the change, possibly as a series\n>      with mutiple commits, using auxiliary data structure, which\n>      would represent predecessor-successor NxM transitive relation\n>      in a similar way in a form that is efficient to access.\n>\n> Something like this should allow us avoid relying on \"change ID\"s\n> that can collide elsewhere in the world without having a central\n> authority to assign them.\n\nI have a few submissions where I recorded the commit hash and the\nprevious commits in the email headers.\n\nhttps://lore.kernel.org/git/0ab05a4cf09ba02016b4493936ad1b092b1326aa.1730979849.git.code@khaugsbakk.name/\n\nFor this one (v3):[1]\n\n```\nX-Commit-Hash: 0ab05a4cf09ba02016b4493936ad1b092b1326aa\nX-Previous-Commits: c50f9d405f9043a03cb5ca1855fbf27f9423c759 63a431537b78e2d84a172b5c837adba6184a1f1b\n```\n\n• `X-Commit-Hash`: my local commit for this patch\n• `X-Previous-Commits`: the two previous commits (v1 and v2 in arbitrary order)\n\nVersion 1 just has the hash:\n\nhttps://lore.kernel.org/git/63a431537b78e2d84a172b5c837adba6184a1f1b.1729451376.git.code@khaugsbakk.name/\n\n```\nX-Commit-Hash: 63a431537b78e2d84a172b5c837adba6184a1f1b\n```\n\nAnd v2:\n\nhttps://lore.kernel.org/git/c50f9d405f9043a03cb5ca1855fbf27f9423c759.1730234365.git.code@khaugsbakk.name/\n\n```\nX-Commit-Hash: c50f9d405f9043a03cb5ca1855fbf27f9423c759\nX-Previous-Commits: 63a431537b78e2d84a172b5c837adba6184a1f1b\n```\n\n† 1: The hash is in the message-id in my case.  But I wanted a dedicated\n    field instead of taking it out of the msg id.  And the msg id makeup\n    doesn’t seem documented.  I’ve already seen a thread where someone\n    relied on parsing data out of the msg id until it changed from under\n    them.\n\n>  (4) Just like we can use auxiliary data structures like bitmaps to\n>      figure out reachability without following all the links in the\n>      commit history topology, we should be able to learn how a new\n>      change was born, and trace how it evolved into newer iteration\n>      of the moral equivalent of the change, possibly as a series\n>      with mutiple commits, using auxiliary data structure, which\n>      would represent predecessor-successor NxM transitive relation\n>      in a similar way in a form that is efficient to access.\n\nI don’t know if this is related but it would be amazing if we users\ncould define custom indexes on the DB.  Maybe people won’t agree on what\na change-id should mean (judging by this thread?) but with custom\nindexes you could maybe get fast queries for whatever “id” you want to define.\n\nUnrelated example: defining an index on `git patch-id --stable` for\nquick *cherry* checks without making your own table with:\n\n```\n<rev list> | git diff-tree --patch --stdin \\\n    | git patch-id --stable\n```\n"},{"id":"518093","messageId":"aCXCgKYpEqxWxIT_@ugly","threadId":"63245","inReplyTo":"xmqqjz6jb6kd.fsf@gitster.g","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Oswald Buddenhagen","fromEmail":"oswald.buddenhagen@gmx.de","sentAt":"2025-05-15T10:31:28Z","receivedAt":"2025-05-15T10:31:39Z","isPatch":false,"sender":{"key":"oswald.buddenhagen@gmx.de","avatar":"https://avatars.githubusercontent.com/u/812380?v=4"},"body":"On Wed, May 14, 2025 at 07:38:58AM -0700, Junio C Hamano wrote:\n>So I do appreciate that the wider Git ecosystem like Gerrit and JJ\n>are talking to adopt \"Change ID\" with the same syntax and semantics\n>(if they all can agree, that is), but I do not think it needs to\n>affect Git at the object level.  From what I have heard so far, it\n>certainly does not fit in the commit and tag object headers, but\n>more like the usual trailer thing.\n>\nbut it doesn't really speak against it, either. it's a rather subjective \ncall whether 3rd party metadata should be included into the commit \nproper or bolted on via the commit message.\nthis is really quite similar to rfc822 headers, and no-one would \nseriously argue that \"auxiliary\" meta data should be stored somewhere \nelse. in fact, the current recommendation is even to omit X- prefixes \n(rfc6648).\n\nit certainly seems valuable to at least have an example that proves that \ngit doesn't corrupt extra headers. however, that would be obviously \ninsufficient for \"proper\" support in git itself, as creation, editing \nand display of these headers needs to be reliable and convenient as \nwell.\n\nthe use case for change-ids is addressing evolving changes (be it \nlocally or in online review systems), which requires a one-to-one \nmapping with a commit in the active working set (sans the fact that a \nchange can exist on multiple branches). as such, the ambiguity that \ntracking the complete evolution of commits would create would be even \nactively counter-productive for this use case.\n\nby extension, it is sometimes necessary to make a conscious decision \nwhich change-id a particular derived commit should have. for this, it's \nquite convenient to have it in the commit message, because this makes it \neasy to drop one (or both) change-ids when squashing (or splitting) \ncommits. however, this could also be achieved by having a defined format \nfor editing metadata from within the commit message editor (this can be \nalready achieved by having prepare-commit-msg and commit-msg hooks, but \nit's messy and lacks standardization).\n\none argument against change-id trailers is that they are eye sores, in \nparticular in small projects that don't use trailers otherwise. this is \nin fact a common argument against even optional use of gerrit for \nreviews. having a more \"subtle\" implementation in git upstream would \ncertainly alleviate this.\n"},{"id":"518148","messageId":"CA+P7+xrruw=NUJgzV4D6CQbmGJO4CEjhkU_+qFDruD5YMsidDw@mail.gmail.com","threadId":"63245","inReplyTo":"aCXCgKYpEqxWxIT_@ugly","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Jacob Keller","fromEmail":"jacob.keller@gmail.com","sentAt":"2025-05-15T16:32:23Z","receivedAt":"2025-05-15T16:32:36Z","isPatch":false,"sender":{"key":"jacob.keller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/874719?v=4"},"body":"On Thu, May 15, 2025 at 3:49 AM Oswald Buddenhagen\n<oswald.buddenhagen@gmx.de> wrote:\n> one argument against change-id trailers is that they are eye sores, in\n> particular in small projects that don't use trailers otherwise. this is\n> in fact a common argument against even optional use of gerrit for\n> reviews. having a more \"subtle\" implementation in git upstream would\n> certainly alleviate this.\n>\n\nAt one point, the driver team I work for wanted to include Change-Id\ntrailers to commits we submitted to the Linux kernel, for tracking\nagainst our own database (we used Gerrit at the time). They were\nrejected for this very reason of being an eye sore --  (possibly other\nreasons as well, I can't recall the full discussion). Of course, if\nthey aren't in the commit message but instead a header, it would have\nneeded some other way to specify them in emailed patch form if we\nwanted them to stick around when using the am-based workflows.\n"},{"id":"518172","messageId":"xmqqfrh5zlu5.fsf@gitster.g","threadId":"63245","inReplyTo":"CA+P7+xrruw=NUJgzV4D6CQbmGJO4CEjhkU_+qFDruD5YMsidDw@mail.gmail.com","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-05-15T19:59:46Z","receivedAt":"2025-05-15T19:59:49Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jacob Keller <jacob.keller@gmail.com> writes:\n\n> At one point, the driver team I work for wanted to include Change-Id\n> trailers to commits we submitted to the Linux kernel, for tracking\n> against our own database (we used Gerrit at the time). They were\n> rejected for this very reason of being an eye sore --  (possibly other\n> reasons as well, I can't recall the full discussion).\n\nFor something to be an eye sore, it also has to be of no use to\nthose who consider it an eye sore.  The signed-off-by trailer is\nnoisy and it becomes annoying after reading \"git log --no-merges\"\nfor a week worth of commits, but it serves useful purpose so nobody\nwould complain them as being an eye sore, even if they complain for\nother reasons.\n\nWhy weren't they seeing any benefit of having such trailer?  Would\nthey have found a good use of the information if it were hidden in\nthe header part?\n\nIf the answer is \"it is only useful to some people\", what is the\nreason why those other people find it useless?  Is it \"our own\ndatabase\" being closed and there were no federated catalog of\nchange-ids that can be used by all project participants?  Or does it\ngo beyond that, like what a Change-Id trailer means to project\nparticipants from one organization is different to those from\nanother, or something?\n\n"},{"id":"518190","messageId":"aCZKK2/OdrpEUqI3@ubby","threadId":"63245","inReplyTo":"xmqqfrh5zlu5.fsf@gitster.g","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Nico Williams","fromEmail":"nico@cryptonector.com","sentAt":"2025-05-15T20:10:19Z","receivedAt":"2025-05-15T21:50:11Z","isPatch":false,"sender":{"key":"nico@cryptonector.com","avatar":null},"body":"On Thu, May 15, 2025 at 12:59:46PM -0700, Junio C Hamano wrote:\n> For something to be an eye sore, it also has to be of no use to\n> those who consider it an eye sore.  The signed-off-by trailer is\n> noisy and it becomes annoying after reading \"git log --no-merges\"\n> for a week worth of commits, but it serves useful purpose so nobody\n> would complain them as being an eye sore, even if they complain for\n> other reasons.\n\nMaybe `git log` can have options for leaving out trailers of no interest\nto the user?  Email workflows still will see them, of course.\n\n> Why weren't they seeing any benefit of having such trailer?  Would\n> they have found a good use of the information if it were hidden in\n> the header part?\n> \n> If the answer is \"it is only useful to some people\", what is the\n> [...]\n\nI think the answer is \"they are only useful a tiny fraction of the times\nI/we/they look at the git log\".\n\nNico\n-- \n"},{"id":"519847","messageId":"87tt4t12c0.fsf@iotcl.com","threadId":"63245","inReplyTo":"aCJwgWaNoBVjvImJ@tapette.crustytoothpaste.net","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Toon Claes","fromEmail":"toon@iotcl.com","sentAt":"2025-06-06T12:28:31Z","receivedAt":"2025-06-06T12:28:49Z","isPatch":false,"sender":{"key":"toon@iotcl.com","avatar":"https://avatars.githubusercontent.com/u/121621?v=4"},"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> On 2025-05-12 at 21:43:46, Martin von Zweigbergk wrote:\n>> Random bytes has worked well for jj.\n>\n> I would like to suggest that we use a deterministic approach.  People\n> rely on Git commits being deterministic, including in my stash\n> import/export series[0].  In addition, it's important to avoid any\n> allegations of side channels or leaking information in commits, which\n> would be a concern in many environments and which a deterministic\n> approach would avoid[1].\n>\n> I'd suggest a simple SHA-256 hash of the original commit data (for both\n> SHA-1 and SHA-256 commits, but one that would change to a new hash if we\n> added one) or an HMAC-SHA-256 with a fixed and documented key.\n\nI was thinking: you cannot guarantee determinism, because the change-ID\nwould remain stable, even when if the underlaying data on which it was\ngenerated changes. But on second thought, _some_ determinisn *can* be\nuseful, for example when different tools try to generate a change-ID for\nthe same source commit.\n\n> I would also recommend a config option to avoid creating these IDs for\n> those who don't want them included for privacy reasons.  I expect to set\n> such an option, for instance.\n\nFair enough.\n\n-- \nCheers,\nToon\n"},{"id":"519849","messageId":"87plfh10nv.fsf@iotcl.com","threadId":"63245","inReplyTo":"CAESOdVCjc1kvQSKnxGfNNSTvFhLRjH_vzwMauP8ZWQ5hhfBnEw@mail.gmail.com","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Toon Claes","fromEmail":"toon@iotcl.com","sentAt":"2025-06-06T13:04:36Z","receivedAt":"2025-06-06T13:04:50Z","isPatch":false,"sender":{"key":"toon@iotcl.com","avatar":"https://avatars.githubusercontent.com/u/121621?v=4"},"body":"Martin von Zweigbergk <martinvonz@google.com> writes:\n\n> On Wed, 23 Apr 2025 at 08:51, Junio C Hamano <gitster@pobox.com> wrote:\n>> Would it make sense, though?  Imagine that a contributor in your\n>> project did not refactor code properly and instead made a\n>> copy-and-paste duplicates of a very similar code.  I find a bug in\n>> one of them, without realizing that the old mistake of duplicating\n>> code (instead of making it a shared helper that is called from the\n>> two places) and create a fix for it.  Later somebody else realizes\n>> the same fix is needed for the other copy---attempting to cherry\n>> pick the original fix may find that remaining copy of a buggy code\n>> as the logic to perform a three-way merge across renames that is\n>> sufficiently clever kicks in.  Shouldn't these two commits to fix\n>> the same bug in two places share the same change ID so that it is\n>> clear to the later developers that the latter fix was derived from\n>> the former one?\n>\n> Maybe it depends on how the forge uses the change id. If the forge is\n> Gerrit, it will use the (change id, target branch) to identify a\n> review (IIUC), so then it will require a new change ID because it\n> requires a new review. If it doesn't have that requirement (maybe it's\n> PR-style forge), then it could at least highlight to the reviewer that\n> there was an old version of the change that has already been merged.\n\nI've been thinking some more about duplicates.\n\nFirst, I don't think you can enforce uniqueness. Whenever you're working\nwith forks you can have commit A from fork I to have the same change-ID\nas commit B in fork II (because either of them might be rebased).\nEnforcing uniqueness on change-IDs would disallow the user to fetch from\nboth forks, or they would need to specify how to resolve the change-id\nconflict.\n\nI like Martin's idea of having forges ensuring uniqueness, but I was\nwondering if we should take it a step furter: a commit should not be\nable to reach another commit with the same change-id. Verifying this on\nthe client-side, helps the user detect issues sooner. Only seeing that\nerror when pushing to the forge might be annoying.\n\nAnyhow on the other hand, what about merges? If people merge the 'main'\nbranch into their feature branch, it would no longer be possible to\nmerge the feature branch into the 'main' branch.\n\nBut to circle back to Junio's example. It depends on what you want to\ntrack. You could consider both changes to be different, and thus\nrequiring different change-ids, because one didn't fix all of it.\n\n-- \nCheers,\nToon\n"},{"id":"519856","messageId":"xmqqmsak50zd.fsf@gitster.g","threadId":"63245","inReplyTo":"87tt4t12c0.fsf@iotcl.com","subject":"Re: Semantics of change IDs (Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-06-06T15:44:06Z","receivedAt":"2025-06-06T15:44:10Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Toon Claes <toon@iotcl.com> writes:\n\n> ... But on second thought, _some_ determinisn *can* be\n> useful, for example when different tools try to generate a change-ID for\n> the same source commit.\n\nYup.  And once we have such determinism for the change-ID that is\ngiven to a freshly written commit not derived from anything else, as\nlong as different tools use the same criteria to decide when to and\nnot to carry forward the existing change-IDs forward to a commit\nthey newly create,\n\n"},{"id":"524446","messageId":"20250819140449.730068-1-safinaskar@zohomail.com","threadId":"63245","inReplyTo":"CAESOdVAspxUJKGAA58i0tvks4ZOfoGf1Aa5gPr0FXzdcywqUUw@mail.gmail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Askar Safin","fromEmail":"safinaskar@zohomail.com","sentAt":"2025-08-19T14:04:49Z","receivedAt":"2025-08-19T14:05:26Z","isPatch":false,"sender":{"key":"safinaskar@zohomail.com","avatar":null},"body":"If your change-id proposal requires some incompatible changes in git itself,\nthen, please, do them now! Incompatible release of git (git 3.0) is near.\n\n--\nAskar Safin\n"},{"id":"524453","messageId":"35C37A8B-732C-4CD2-8177-666CB84E173E@gmail.com","threadId":"63245","inReplyTo":"20250819140449.730068-1-safinaskar@zohomail.com","subject":"Re: Gerrit, GitButler, and Jujutsu projects collaborating on change-id commit footer","fromName":"Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2025-08-19T16:44:59Z","receivedAt":"2025-08-19T16:45:11Z","isPatch":false,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"\n> Le 19 août 2025 à 10:10, Askar Safin <safinaskar@zohomail.com> a écrit :\n> \n> ﻿If your change-id proposal requires some incompatible changes in git itself,\n> then, please, do them now! Incompatible release of git (git 3.0) is near.\n\nI’m not the maintainer (so grain of salt), but as far as I know nobody has even picked a date for that release. I’m not sure that qualifies as near."}]}