{"thread":{"id":"58181","subject":"Feature request: provide a persistent IDs on a commit","startedAt":"2022-07-18T17:23:25Z","lastAt":"2024-12-15T17:09:20Z","messageCount":30,"participants":["Stephen Finucane","Konstantin Ryabitsev","Ævar Arnfjörð Bjarmason","Michal Suchánek","Glen Choo","Theodore Ts'o","Han-Wen Nienhuys","Phillip Susi","Hilco Wijbenga","Philip Oakley","Jacob Keller","Elijah Newren","Jason Pyeron","Martin von Zweigbergk"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"459303","messageId":"bdbe9b7c1123f70c0b4325d778af1df8fea2bb1b.camel@that.guru","threadId":"58181","inReplyTo":null,"subject":"Feature request: provide a persistent IDs on a commit","fromName":"Stephen Finucane","fromEmail":"stephen@that.guru","sentAt":"2022-07-18T17:18:11Z","receivedAt":"2022-07-18T17:23:25Z","isPatch":false,"sender":{"key":"stephen@that.guru","avatar":"https://gravatar.com/avatar/2f9716975d743b6d7764972ed20fecdad956a5d4e3e6cfe8282d10151f1b2df0?d=mp&s=160"},"body":"...to track evolution of a patch through time.\n\ntl;dr: How hard would it be to retrofit an 'ChangeID' concept à la the 'Change-\nID' trailer used by Gerrit into git core?\n\nFirstly, apologies in advance if this is the wrong forum to post a feature\nrequest. I help maintain the Patchwork project [1], which a web-based tool that\nprovides a mechanism to track the state of patches submitted to a mailing list\nand make sure stuff doesn't slip through the crack. One of our long-term goals\nhas been to track the evolution of an individual patch through multiple\nrevisions. This is surprisingly hard goal because oftentimes there isn't a whole\nlot to work with. One can try to guess whether things are the same by inspecting\nthe metadata of the commit (subject, author, commit message, and the diff\nitself) but each of these metadata items are subject to arbitrary changes and\nare therefore fallible.\n\nOne of the mechanisms I've seen used to address this is the 'Change-ID' trailer\nused by Gerrit. For anyone that hasn't seen this, the Gerrit server provides a\ngit commit hook that you can install locally. When installed, this appends a\n'Change-ID' trailer to each and every commit message. In this way, the evolution\nof a patch (or a \"change\", in Gerrit parlance) can be tracked through time since\nthe Change ID provides an authoritative answer to the question \"is this still\nthe same patch\". Unfortunately, there are still some obvious downside to this\napproach. Not only does this additional trailer clutter your commit messages but\nit's also something the user must install themselves. While Gerrit can insist\nthat this is installed before pushing a change, this isn't an option for any of\nthe common forges nor is it something git-send-email supports.\n\nI imagine most people working with mailing list based workflows have their own\nclient side tooling to support this while software forges like GitHub and GitLab\nsimply don't bother tracking version history between individual commits in a\npull/merge request. IMO though, it would be fantastic if third party tools\nweren't necessary though. What I suspect we want is a persistent ID (or rather\nUUID) that never changes regardless of how many times a patch is cherry-picked,\nrebased, or otherwise modified, similar to the Author and AuthorDate fields.\nLike Author and AuthorDate, it would be part of the core git commit metadata\nrather than something in the commit message like Signed-Off-By or Change-ID.\n\nHas such an idea ever been explored? Is it even possible? Would it be broadly\nuseful?\n\nCheers,\nStephen\n\n[1] github.com/getpatchwork/patchwork/\n"},{"id":"459305","messageId":"20220718173511.rje43peodwdprsid@meerkat.local","threadId":"58181","inReplyTo":"bdbe9b7c1123f70c0b4325d778af1df8fea2bb1b.camel@that.guru","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2022-07-18T17:35:11Z","receivedAt":"2022-07-18T17:35:20Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On Mon, Jul 18, 2022 at 06:18:11PM +0100, Stephen Finucane wrote:\n> ...to track evolution of a patch through time.\n> \n> tl;dr: How hard would it be to retrofit an 'ChangeID' concept à la the 'Change-\n> ID' trailer used by Gerrit into git core?\n\nI just started working on this for b4, with the notable difference that the\nchange-id trailer is used in the cover letter instead of in individual\ncommits, which moves the concept of \"change\" from a single commit to a series\nof commits. IMO, it's much more useful in that scope, because as series are\nreviewed and iterated, individual patches can get squashed, split up or\notherwise transformed.\n\nYou can see my test commits here:\nhttps://lore.kernel.org/linux-patches/20220707-my-new-branch-v1-0-8d355bae1bb5@linuxfoundation.org/\n\nYou will notice that each cover letter has the following in the basement:\n\n    ---\n    base-commit: 88084a3df1672e131ddc1b4e39eeacfd39864acf\n    change-id: 20220707-my-new-branch-[uniquerandomstr]\n\nThere are 3 revisions of the series and you can locate all of them by\nsearching for that trailer:\nhttps://lore.kernel.org/linux-patches/?q=%22change-id%3A+20220707-my-new-branch-1325e0e7fd1c%22\n\nNote, that \"b4 submit\" is in the early experimental stage and will likely\nundergo significant changes in the next few weeks, so I wouldn't treat it as\nany more than curiosity at this point.\n\n-K\n"},{"id":"459311","messageId":"220718.86ilnuw8jo.gmgdl@evledraar.gmail.com","threadId":"58181","inReplyTo":"bdbe9b7c1123f70c0b4325d778af1df8fea2bb1b.camel@that.guru","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-07-18T18:50:31Z","receivedAt":"2022-07-18T19:00:20Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Mon, Jul 18 2022, Stephen Finucane wrote:\n\n> ...to track evolution of a patch through time.\n>\n> tl;dr: How hard would it be to retrofit an 'ChangeID' concept à la the 'Change-\n> ID' trailer used by Gerrit into git core?\n>\n> Firstly, apologies in advance if this is the wrong forum to post a feature\n> request. I help maintain the Patchwork project [1], which a web-based tool that\n> provides a mechanism to track the state of patches submitted to a mailing list\n> and make sure stuff doesn't slip through the crack. One of our long-term goals\n> has been to track the evolution of an individual patch through multiple\n> revisions. This is surprisingly hard goal because oftentimes there isn't a whole\n> lot to work with. One can try to guess whether things are the same by inspecting\n> the metadata of the commit (subject, author, commit message, and the diff\n> itself) but each of these metadata items are subject to arbitrary changes and\n> are therefore fallible.\n>\n> One of the mechanisms I've seen used to address this is the 'Change-ID' trailer\n> used by Gerrit. For anyone that hasn't seen this, the Gerrit server provides a\n> git commit hook that you can install locally. When installed, this appends a\n> 'Change-ID' trailer to each and every commit message. In this way, the evolution\n> of a patch (or a \"change\", in Gerrit parlance) can be tracked through time since\n> the Change ID provides an authoritative answer to the question \"is this still\n> the same patch\". Unfortunately, there are still some obvious downside to this\n> approach. Not only does this additional trailer clutter your commit messages but\n> it's also something the user must install themselves. While Gerrit can insist\n> that this is installed before pushing a change, this isn't an option for any of\n> the common forges nor is it something git-send-email supports.\n\ngit format-patch+send-email will send your trailers along as-is, how\ndoesn't it support Change-Id. Does it need some support that any other\nmade-up trailer doesn't?\n\n> I imagine most people working with mailing list based workflows have their own\n> client side tooling to support this while software forges like GitHub and GitLab\n> simply don't bother tracking version history between individual commits in a\n> pull/merge request.\n\nIt's far from ideal, but at least GitLab shows a diff on a push to a MR,\nincluding if it's force-pushed. I'm not sure about GitHub.\n\n> IMO though, it would be fantastic if third party tools\n> weren't necessary though. What I suspect we want is a persistent ID (or rather\n> UUID) that never changes regardless of how many times a patch is cherry-picked,\n> rebased, or otherwise modified, similar to the Author and AuthorDate fields.\n> Like Author and AuthorDate, it would be part of the core git commit metadata\n> rather than something in the commit message like Signed-Off-By or Change-ID.\n>\n> Has such an idea ever been explored? Is it even possible? Would it be broadly\n> useful?\n\nThis has come up a bunch of times. I think that the thing git itself\nshould be doing is to lean into the same notion that we use for tracking\nrenames. I.e. we don't, we analyze history after-the-fact and spot the\nrenames for you.\n\nWe have some of that in git already, as git-patch-id, and more recently\ngit-range-diff. Both are flawed in a bunch of ways, and it's easy to run\ninto edge cases where they don't spot something that they \"should\"\nhave. Where \"should\" exists in the mind of the user.\n"},{"id":"459313","messageId":"20220718190403.GT17705@kitsune.suse.cz","threadId":"58181","inReplyTo":"20220718173511.rje43peodwdprsid@meerkat.local","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Michal Suchánek","fromEmail":"msuchanek@suse.de","sentAt":"2022-07-18T19:04:03Z","receivedAt":"2022-07-18T19:04:10Z","isPatch":false,"sender":{"key":"msuchanek@suse.de","avatar":"https://avatars.githubusercontent.com/u/787652?v=4"},"body":"On Mon, Jul 18, 2022 at 01:35:11PM -0400, Konstantin Ryabitsev wrote:\n> On Mon, Jul 18, 2022 at 06:18:11PM +0100, Stephen Finucane wrote:\n> > ...to track evolution of a patch through time.\n> > \n> > tl;dr: How hard would it be to retrofit an 'ChangeID' concept à la the 'Change-\n> > ID' trailer used by Gerrit into git core?\n> \n> I just started working on this for b4, with the notable difference that the\n> change-id trailer is used in the cover letter instead of in individual\n> commits, which moves the concept of \"change\" from a single commit to a series\n> of commits. IMO, it's much more useful in that scope, because as series are\n> reviewed and iterated, individual patches can get squashed, split up or\n> otherwise transformed.\n\nYou can turn that around and say that IDs of individual commits are more\npowerful because they are preserved as series are reviewed, split,\nmerged, and commits cherry-picked.\n\nThanks\n\nMichal\n"},{"id":"459332","messageId":"kl6lo7xmt8qw.fsf@chooglen-macbookpro.roam.corp.google.com","threadId":"58181","inReplyTo":"20220718173511.rje43peodwdprsid@meerkat.local","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Glen Choo","fromEmail":"chooglen@google.com","sentAt":"2022-07-18T21:24:07Z","receivedAt":"2022-07-18T21:24:14Z","isPatch":false,"sender":{"key":"glencbz@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58092771?v=4"},"body":"Konstantin Ryabitsev <konstantin@linuxfoundation.org> writes:\n\n> On Mon, Jul 18, 2022 at 06:18:11PM +0100, Stephen Finucane wrote:\n>> ...to track evolution of a patch through time.\n>> \n>> tl;dr: How hard would it be to retrofit an 'ChangeID' concept à la the 'Change-\n>> ID' trailer used by Gerrit into git core?\n>\n> I just started working on this for b4, with the notable difference that the\n> change-id trailer is used in the cover letter instead of in individual\n> commits, which moves the concept of \"change\" from a single commit to a series\n> of commits. IMO, it's much more useful in that scope, because as series are\n> reviewed and iterated, individual patches can get squashed, split up or\n> otherwise transformed.\n\nMy 2 cents, since I used to use Gerrit a lot :)\n\nI find persistent per-commit ids really useful, even when patches get\nmoved around. E.g. Gerrit can show and diff previous versions of the\npatch, which makes it really easy to tell how the patch has evolved\nover time.\n\nThat's not to say that we don't need per-topic ids though ;) E.g. Gerrit\nis pretty bad at handling whole topics - it does naive mapping on a\nper-commit level, so it has no concept of \"these (n - 1) patches should\nreplace these n patches\".\n\nI, for one, would love to see some kind of \"rewrite tracking\" in Git.\nOne use case that comes up often is downstream patches, where patches\nare continuously rebased onto a new upstream; in those cases, it's\npretty hard to keep track of how the patch has changed over time\n"},{"id":"459349","messageId":"61333be26339440d9bae8f12fd1a4faeb5e68ab6.camel@that.guru","threadId":"58181","inReplyTo":"220718.86ilnuw8jo.gmgdl@evledraar.gmail.com","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Stephen Finucane","fromEmail":"stephen@that.guru","sentAt":"2022-07-19T10:47:55Z","receivedAt":"2022-07-19T10:48:12Z","isPatch":false,"sender":{"key":"stephen@that.guru","avatar":"https://gravatar.com/avatar/2f9716975d743b6d7764972ed20fecdad956a5d4e3e6cfe8282d10151f1b2df0?d=mp&s=160"},"body":"On Mon, 2022-07-18 at 20:50 +0200, Ævar Arnfjörð Bjarmason wrote:\n> On Mon, Jul 18 2022, Stephen Finucane wrote:\n> \n> > ...to track evolution of a patch through time.\n> > \n> > tl;dr: How hard would it be to retrofit an 'ChangeID' concept à la the 'Change-\n> > ID' trailer used by Gerrit into git core?\n> > \n> > Firstly, apologies in advance if this is the wrong forum to post a feature\n> > request. I help maintain the Patchwork project [1], which a web-based tool that\n> > provides a mechanism to track the state of patches submitted to a mailing list\n> > and make sure stuff doesn't slip through the crack. One of our long-term goals\n> > has been to track the evolution of an individual patch through multiple\n> > revisions. This is surprisingly hard goal because oftentimes there isn't a whole\n> > lot to work with. One can try to guess whether things are the same by inspecting\n> > the metadata of the commit (subject, author, commit message, and the diff\n> > itself) but each of these metadata items are subject to arbitrary changes and\n> > are therefore fallible.\n> > \n> > One of the mechanisms I've seen used to address this is the 'Change-ID' trailer\n> > used by Gerrit. For anyone that hasn't seen this, the Gerrit server provides a\n> > git commit hook that you can install locally. When installed, this appends a\n> > 'Change-ID' trailer to each and every commit message. In this way, the evolution\n> > of a patch (or a \"change\", in Gerrit parlance) can be tracked through time since\n> > the Change ID provides an authoritative answer to the question \"is this still\n> > the same patch\". Unfortunately, there are still some obvious downside to this\n> > approach. Not only does this additional trailer clutter your commit messages but\n> > it's also something the user must install themselves. While Gerrit can insist\n> > that this is installed before pushing a change, this isn't an option for any of\n> > the common forges nor is it something git-send-email supports.\n> \n> git format-patch+send-email will send your trailers along as-is, how\n> doesn't it support Change-Id. Does it need some support that any other\n> made-up trailer doesn't?\n\nIt supports sending the trailers, sure. What it doesn't support is insisting you\nsend this specific trailer (Change-Id). Only Gerrit can do this (server side,\nthankfully, which means you don't need to ask all contributors to install this\nhook if you want to rely on it for tooling, CI, etc.).\n\n> > I imagine most people working with mailing list based workflows have their own\n> > client side tooling to support this while software forges like GitHub and GitLab\n> > simply don't bother tracking version history between individual commits in a\n> > pull/merge request.\n> \n> It's far from ideal, but at least GitLab shows a diff on a push to a MR,\n> including if it's force-pushed. I'm not sure about GitHub.\n\nGitHub does not. Simply piling multiple additional \"fix\" commits onto the PR\nbranch results in a less horrible review experience since you can maintain\ncontext, alas at the cost of a rotten git log. We don't need to debate the pros\nand cons of the various forges though :)\n\n> \n> > IMO though, it would be fantastic if third party tools\n> > weren't necessary though. What I suspect we want is a persistent ID (or rather\n> > UUID) that never changes regardless of how many times a patch is cherry-picked,\n> > rebased, or otherwise modified, similar to the Author and AuthorDate fields.\n> > Like Author and AuthorDate, it would be part of the core git commit metadata\n> > rather than something in the commit message like Signed-Off-By or Change-ID.\n> > \n> > Has such an idea ever been explored? Is it even possible? Would it be broadly\n> > useful?\n> \n> This has come up a bunch of times. I think that the thing git itself\n> should be doing is to lean into the same notion that we use for tracking\n> renames. I.e. we don't, we analyze history after-the-fact and spot the\n> renames for you.\n\nAny idea where I'd find previous discussions on this? I did look, and the only\nproposal I found was an old one that seemed to suggest including the Change-Id\ncommit-msg hook with git itself which is not what I'm suggesting here.\n\n> We have some of that in git already, as git-patch-id, and more recently\n> git-range-diff. Both are flawed in a bunch of ways, and it's easy to run\n> into edge cases where they don't spot something that they \"should\"\n> have. Where \"should\" exists in the mind of the user.\n\nThat's a fair point and is of course what we (Patchwork) have to do currently.\nPatchwork can track relations between individual patches but doesn't attempt to\ngenerate these relations itself. Instead, we rely on third-party tooling. The\nPaStA tool was one such example of a tool that could do this [1]. I can't\nimagine a tool like Gerrit would ever work without this concept of an\nauthoritative (and arbitrary) identifier to track a patch's identity through\ntime, hence its reliance on the Change-Id trailer.\n\nPerhaps we could flip this on its head. What would be the _downsides_ of\nproviding a persistent, arbitrary identifier on a commit similar to Author and\nAuthorDate fields? There's obviously some work involved in implementing it but\nassuming that was already done, what would break/be worse as a result?\n\nStephen\n\n[1] https://rsarky.github.io/2020/08/10/pasta-patchwork.html\n"},{"id":"459350","messageId":"e2d73c87ae798de6d1dd21fe14371b8cdc65228f.camel@that.guru","threadId":"58181","inReplyTo":"20220718190403.GT17705@kitsune.suse.cz","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Stephen Finucane","fromEmail":"stephen@that.guru","sentAt":"2022-07-19T10:57:39Z","receivedAt":"2022-07-19T10:57:57Z","isPatch":false,"sender":{"key":"stephen@that.guru","avatar":"https://gravatar.com/avatar/2f9716975d743b6d7764972ed20fecdad956a5d4e3e6cfe8282d10151f1b2df0?d=mp&s=160"},"body":"On Mon, 2022-07-18 at 21:04 +0200, Michal Suchánek wrote:\n> On Mon, Jul 18, 2022 at 01:35:11PM -0400, Konstantin Ryabitsev wrote:\n> > On Mon, Jul 18, 2022 at 06:18:11PM +0100, Stephen Finucane wrote:\n> > > ...to track evolution of a patch through time.\n> > > \n> > > tl;dr: How hard would it be to retrofit an 'ChangeID' concept à la the 'Change-\n> > > ID' trailer used by Gerrit into git core?\n> > \n> > I just started working on this for b4, with the notable difference that the\n> > change-id trailer is used in the cover letter instead of in individual\n> > commits, which moves the concept of \"change\" from a single commit to a series\n> > of commits. IMO, it's much more useful in that scope, because as series are\n> > reviewed and iterated, individual patches can get squashed, split up or\n> > otherwise transformed.\n> \n> You can turn that around and say that IDs of individual commits are more\n> powerful because they are preserved as series are reviewed, split,\n> merged, and commits cherry-picked.\n\nThere's also the fact that many communities insist on small, atomic commits:\nthey're much easier to review. It stands to reason that reviewing a series on a\npatch-by-patch basis is also much easier, as is reviewing a series _revision_ on\na patch-by-patch basis. To be able to do this though, you need to be able to map\npatch revisions to their predecessors/successors and well as the series\nrevisions. I don't see how you realistically rely on a series-only identifier.\n\nThere's no reason 'git-format-patch' couldn't allow you to set an\nAuthorID/ChangeID/<whatever we want to call this field> value for a cover\nletter, though it obviously would need to be done manually since cover letters\naren't git objects.\n\n  git send-email \\\n    --reroll-count 2 \\\n    --series-id 300628e5-8b27-45fe-be71-95417f7ccd6f\n    main\n\nStephen\n\n> \n> Thanks\n> \n> Michal\n\n"},{"id":"459352","messageId":"220719.86y1wpuy5o.gmgdl@evledraar.gmail.com","threadId":"58181","inReplyTo":"61333be26339440d9bae8f12fd1a4faeb5e68ab6.camel@that.guru","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-07-19T11:09:02Z","receivedAt":"2022-07-19T11:42:18Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Jul 19 2022, Stephen Finucane wrote:\n\n> On Mon, 2022-07-18 at 20:50 +0200, Ævar Arnfjörð Bjarmason wrote:\n>> On Mon, Jul 18 2022, Stephen Finucane wrote:\n>> \n>> > ...to track evolution of a patch through time.\n>> > \n>> > tl;dr: How hard would it be to retrofit an 'ChangeID' concept à la the 'Change-\n>> > ID' trailer used by Gerrit into git core?\n>> > \n>> > Firstly, apologies in advance if this is the wrong forum to post a feature\n>> > request. I help maintain the Patchwork project [1], which a web-based tool that\n>> > provides a mechanism to track the state of patches submitted to a mailing list\n>> > and make sure stuff doesn't slip through the crack. One of our long-term goals\n>> > has been to track the evolution of an individual patch through multiple\n>> > revisions. This is surprisingly hard goal because oftentimes there isn't a whole\n>> > lot to work with. One can try to guess whether things are the same by inspecting\n>> > the metadata of the commit (subject, author, commit message, and the diff\n>> > itself) but each of these metadata items are subject to arbitrary changes and\n>> > are therefore fallible.\n>> > \n>> > One of the mechanisms I've seen used to address this is the 'Change-ID' trailer\n>> > used by Gerrit. For anyone that hasn't seen this, the Gerrit server provides a\n>> > git commit hook that you can install locally. When installed, this appends a\n>> > 'Change-ID' trailer to each and every commit message. In this way, the evolution\n>> > of a patch (or a \"change\", in Gerrit parlance) can be tracked through time since\n>> > the Change ID provides an authoritative answer to the question \"is this still\n>> > the same patch\". Unfortunately, there are still some obvious downside to this\n>> > approach. Not only does this additional trailer clutter your commit messages but\n>> > it's also something the user must install themselves. While Gerrit can insist\n>> > that this is installed before pushing a change, this isn't an option for any of\n>> > the common forges nor is it something git-send-email supports.\n>> \n>> git format-patch+send-email will send your trailers along as-is, how\n>> doesn't it support Change-Id. Does it need some support that any other\n>> made-up trailer doesn't?\n>\n> It supports sending the trailers, sure. What it doesn't support is insisting you\n> send this specific trailer (Change-Id). Only Gerrit can do this (server side,\n> thankfully, which means you don't need to ask all contributors to install this\n> hook if you want to rely on it for tooling, CI, etc.).\n\nAh, it's still unclear to me what you're proposing here though. That\nsend-email always (generates?) or otherwise insists on the trailer, that\nit can be configured ot add it?\n\nThat send-email have some \"pre-send-email\" hook? Something else?\n\nI'd think for projects that care about this they're likely to have a\ncentralized enough workflow that it can be checked on the remote side,\nwhether that's some sanity check on the applier's \"git am\" pipeline, or\na \"pre-receive\" hook.\n\n>> > I imagine most people working with mailing list based workflows have their own\n>> > client side tooling to support this while software forges like GitHub and GitLab\n>> > simply don't bother tracking version history between individual commits in a\n>> > pull/merge request.\n>> \n>> It's far from ideal, but at least GitLab shows a diff on a push to a MR,\n>> including if it's force-pushed. I'm not sure about GitHub.\n>\n> GitHub does not. Simply piling multiple additional \"fix\" commits onto the PR\n> branch results in a less horrible review experience since you can maintain\n> context, alas at the cost of a rotten git log. We don't need to debate the pros\n> and cons of the various forges though :)\n\nYes, I'm only mentioning it because it's worth looking at existing\n\"solutions\" that are in use in the wild, however flawed those may be.\n\n>> > IMO though, it would be fantastic if third party tools\n>> > weren't necessary though. What I suspect we want is a persistent ID (or rather\n>> > UUID) that never changes regardless of how many times a patch is cherry-picked,\n>> > rebased, or otherwise modified, similar to the Author and AuthorDate fields.\n>> > Like Author and AuthorDate, it would be part of the core git commit metadata\n>> > rather than something in the commit message like Signed-Off-By or Change-ID.\n>> > \n>> > Has such an idea ever been explored? Is it even possible? Would it be broadly\n>> > useful?\n>> \n>> This has come up a bunch of times. I think that the thing git itself\n>> should be doing is to lean into the same notion that we use for tracking\n>> renames. I.e. we don't, we analyze history after-the-fact and spot the\n>> renames for you.\n>\n> Any idea where I'd find previous discussions on this? I did look, and the only\n> proposal I found was an old one that seemed to suggest including the Change-Id\n> commit-msg hook with git itself which is not what I'm suggesting here.\n\nAt the time I was punting on finding the links, and just working off\nvague recollection, and hoping you'd go list spelunking.\n\nBut I since recalled some details, I think the most relevant thing is\nthis discussion about a \"git evolve\":\n\n    https://lore.kernel.org/git/CAPL8ZivFmHqS2y+WmNR6faRMnuahiqwPVYsV99NiJ1QLHOs9fQ@mail.gmail.com/\n\nWhich I think you'll find useful, especially as mercurial has an\nexisting implementation. The wider context for that \"git evolve\" is (I\nbelieve) people at Google who maintain Gerrit trying to \"upstream\" the\nChange-Id.\n\nNow, it hasn't landed in git.git, and it's been a few years, but going\nthrough the details of why it fizzled out will be useful to you, if\nyou're interested in driving something like this forward.\n\nThere's also these two proposals from Eric Raymond:\n\n\thttps://lore.kernel.org/git/20190515191605.21D394703049@snark.thyrsus.com/\n\thttps://lore.kernel.org/git/20190521013250.3506B470485F@snark.thyrsus.com/\n\nWhich I'm linking to here not because I think they're viable, as you can\nsee from my participation in those threads I think what he suggested is\nan architectural dead end as far as git is concerned.\n\nBut rather because it's conceptually adjacent (you could in principle\nuse nanosecond timestamps as a poor man's UUID), and much of the\nfollow-up discussion is about format changes in general, and if/when\nthose might be viable.\n\n>> We have some of that in git already, as git-patch-id, and more recently\n>> git-range-diff. Both are flawed in a bunch of ways, and it's easy to run\n>> into edge cases where they don't spot something that they \"should\"\n>> have. Where \"should\" exists in the mind of the user.\n>\n> That's a fair point and is of course what we (Patchwork) have to do currently.\n> Patchwork can track relations between individual patches but doesn't attempt to\n> generate these relations itself. Instead, we rely on third-party tooling. The\n> PaStA tool was one such example of a tool that could do this [1]. I can't\n> imagine a tool like Gerrit would ever work without this concept of an\n> authoritative (and arbitrary) identifier to track a patch's identity through\n> time, hence its reliance on the Change-Id trailer.\n\nI haven't used Gerrit or Patchwork, so much of this is from ignorance on\nthat front, but I have spent a lot of time thinking about this in the\ncontext of git in general.\n\nI think as users of git go the git project itself makes very heavy use\nof this, i.e. sequences of patches are substantially rewritten, split,\nsquashed etc. all the time, or even split into two or more sets of\nsubmissions.\n\nHaving said all that I can't see how a Change-Id isn't a Bad Idea(TM)\nfor all the same reasons that pre-git SCMs file formats that track\nrenames explicitly were a bad idea.\n\nI.e. yes you can come up with cases where that's \"better\" than what git\ndoes, but they didn't handle splitting/merging files etc.\n\nSimilarly what happens when you have 3 patches each with their own\nChange-Id and you split them into 4 patches. Is the Change-Id 1=1 or\n1=many. I'm suggesting that you'd want a solution that can be many=many.\n\nAnd also, that those many=many should be dynamically configurable and\ninferred after the fact. E.g. range-diff will commits that are similar\nenough that two authors with no knowledge of each other independently\ncame up with.\n\nI think that range-diff is still lacking in a lot of ways, in particular:\n\n * It matches entire commits (log + diff) on a similarity score, I've\n   often wanted a way to \"weigh\" it, so e.g. a matching hunk would have\n   3x the matching score of a matching commit message.\n\n   Now it often \"gives up\", you can give it a higher --creation-factor,\n   but that's \"global\", so for a large range you'll often start\n   including irrelevant things as well.\n\n * It only does 1=1 attribution, and e.g. currently can't find/represent\n   a case where a commit with 3 hunks got split into two commits, with 2\n   and 1 hunks, respectively. It'll (usually) show a diff to the new 2\n   hunk commit, but the \"new\" 1 hunk will be shown as new.\n\n   We could continue to drill down and find such \"unattributed\" hunks.\n\n> Perhaps we could flip this on its head. What would be the _downsides_ of\n> providing a persistent, arbitrary identifier on a commit similar to Author and\n> AuthorDate fields? There's obviously some work involved in implementing it but\n> assuming that was already done, what would break/be worse as a result?\n\nThat \"Repository formats matter\", to borrow a phrase from a classic post\nabout git[1]. Once you provide a way to do something it will be used,\nand when that something has inherent limitations (think SCM rename\ntracking) used to the exclusion of others.\n\nYou can't provide something like that as an opt-in and \"upstream\" it\nwithout it inevetably trickling into a lot of areas of Git's UX.\n\nTo continue the rename example, now you can just re-arrange your source\ntree and not worry about micro-managing it with \"git mv\" (in the \"svn\nmv\" sense), git will figure it out after the fact.\n\nThat's a sinificant UX benefit, we can provide a *much simpler* UX as a\nresult.\n\nWhat would be the harm of an optional \"rename tracking\" header? After\nall the heuristic sometimes \"fails\".\n\nThe harm would be that if you really wanted to lean into that (even\noptionally) you'd be forced to add that to all sorts of tooling, not\njust the cheap convenience that is \"git mv\" currently.\n\nLikewise everything from \"cherry-pick\" to \"rebase\" to \"commit\" would\ninevitably have to learn some way to know about, carry forward and ask\nthe user about Change-Id's and their preservation. Don't you think so?\n\nOtherwise they'd be much too easy to lose track of, and if they only\nreason we did all that is because we didn't think enough about the \"work\nit out after\" approach that would be a bad investment of time.\n\nBut I may be wrong about all of that, I think one thing that would\nreally help clarify this & similar proposals is if people pushing it\nforward came up with some basic tests for it, i.e. just something like\na:\n\n    series-v1/\n    series-v2/\n\nWhere those two directories would be the \"git format-patch\" output (or\nwhatever) of two versions of a series that Gerrit or Patchwork are now\nmanaging, along with some (plain text?) manual mapping of which things\nin v1 correspond to v2.\n\nWe could then compare how that manual attribution performs v.s. trying\nto find which things match (range-diff) afterwards.\n\n1. https://keithp.com/blog/Repository_Formats_Matter/\n\n"},{"id":"459355","messageId":"20220719115729.GV17705@kitsune.suse.cz","threadId":"58181","inReplyTo":"220719.86y1wpuy5o.gmgdl@evledraar.gmail.com","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Michal Suchánek","fromEmail":"msuchanek@suse.de","sentAt":"2022-07-19T11:57:29Z","receivedAt":"2022-07-19T11:59:36Z","isPatch":false,"sender":{"key":"msuchanek@suse.de","avatar":"https://avatars.githubusercontent.com/u/787652?v=4"},"body":"On Tue, Jul 19, 2022 at 01:09:02PM +0200, Ævar Arnfjörð Bjarmason wrote:\n> \n> On Tue, Jul 19 2022, Stephen Finucane wrote:\n> \n> > On Mon, 2022-07-18 at 20:50 +0200, Ævar Arnfjörð Bjarmason wrote:\n> >> On Mon, Jul 18 2022, Stephen Finucane wrote:\n> >> \n> >> > ...to track evolution of a patch through time.\n> >> > \n> >> > tl;dr: How hard would it be to retrofit an 'ChangeID' concept à la the 'Change-\n> >> > ID' trailer used by Gerrit into git core?\n> >> > \n> >> > Firstly, apologies in advance if this is the wrong forum to post a feature\n> >> > request. I help maintain the Patchwork project [1], which a web-based tool that\n> >> > provides a mechanism to track the state of patches submitted to a mailing list\n> >> > and make sure stuff doesn't slip through the crack. One of our long-term goals\n> >> > has been to track the evolution of an individual patch through multiple\n> >> > revisions. This is surprisingly hard goal because oftentimes there isn't a whole\n> >> > lot to work with. One can try to guess whether things are the same by inspecting\n> >> > the metadata of the commit (subject, author, commit message, and the diff\n> >> > itself) but each of these metadata items are subject to arbitrary changes and\n> >> > are therefore fallible.\n> >> > \n> >> > One of the mechanisms I've seen used to address this is the 'Change-ID' trailer\n> >> > used by Gerrit. For anyone that hasn't seen this, the Gerrit server provides a\n> >> > git commit hook that you can install locally. When installed, this appends a\n> >> > 'Change-ID' trailer to each and every commit message. In this way, the evolution\n> >> > of a patch (or a \"change\", in Gerrit parlance) can be tracked through time since\n> >> > the Change ID provides an authoritative answer to the question \"is this still\n> >> > the same patch\". Unfortunately, there are still some obvious downside to this\n> >> > approach. Not only does this additional trailer clutter your commit messages but\n> >> > it's also something the user must install themselves. While Gerrit can insist\n> >> > that this is installed before pushing a change, this isn't an option for any of\n> >> > the common forges nor is it something git-send-email supports.\n> >> \n> >> git format-patch+send-email will send your trailers along as-is, how\n> >> doesn't it support Change-Id. Does it need some support that any other\n> >> made-up trailer doesn't?\n> >\n> > It supports sending the trailers, sure. What it doesn't support is insisting you\n> > send this specific trailer (Change-Id). Only Gerrit can do this (server side,\n> > thankfully, which means you don't need to ask all contributors to install this\n> > hook if you want to rely on it for tooling, CI, etc.).\n> \n> Ah, it's still unclear to me what you're proposing here though. That\n> send-email always (generates?) or otherwise insists on the trailer, that\n> it can be configured ot add it?\n\nAnd isn't send-email time too late?\n\nThat would mean that you get new ID for every version of the patch sent.\n\nThanks\n\nMichal\n"},{"id":"459535","messageId":"20220720192144.mxdemgcdjxb2klgl@nitro.local","threadId":"58181","inReplyTo":"kl6lo7xmt8qw.fsf@chooglen-macbookpro.roam.corp.google.com","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2022-07-20T19:21:44Z","receivedAt":"2022-07-20T19:21:49Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On Mon, Jul 18, 2022 at 02:24:07PM -0700, Glen Choo wrote:\n> > I just started working on this for b4, with the notable difference that the\n> > change-id trailer is used in the cover letter instead of in individual\n> > commits, which moves the concept of \"change\" from a single commit to a series\n> > of commits. IMO, it's much more useful in that scope, because as series are\n> > reviewed and iterated, individual patches can get squashed, split up or\n> > otherwise transformed.\n> \n> My 2 cents, since I used to use Gerrit a lot :)\n> \n> I find persistent per-commit ids really useful, even when patches get\n> moved around. E.g. Gerrit can show and diff previous versions of the\n> patch, which makes it really easy to tell how the patch has evolved\n> over time.\n\nThe kernel community has repeatedly rejected per-patch Change-id trailers\nbecause they carry no meaningful information outside of the gerrit system on\nwhich they were created. Seeing a Change-Id trailer in a commit tells you\nnothing about the history of that commit unless you know the gerrit system on\nwhich this patch was reviewed (and have access to it, which is not a given).\nThis is not as opaque as it used to be now that Gerrit provided ability to\nclone the underlying notedb, but this still fails on commits that were\ncontributed to an upstream that doesn't use Gerrit.\n\nThe current recommended strategy for the kernel is to put any historical\ninformation (including any links to archival sites, etc) into the merge\ncommit and only keep chain-of-custody and code-review trailers in actual\ncode commits. For this reason, I opted to use change-ids in the cover letter\nonly.\n\n-Konstantin\n"},{"id":"459536","messageId":"20220720193010.GZ17705@kitsune.suse.cz","threadId":"58181","inReplyTo":"20220720192144.mxdemgcdjxb2klgl@nitro.local","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Michal Suchánek","fromEmail":"msuchanek@suse.de","sentAt":"2022-07-20T19:30:10Z","receivedAt":"2022-07-20T19:30:19Z","isPatch":false,"sender":{"key":"msuchanek@suse.de","avatar":"https://avatars.githubusercontent.com/u/787652?v=4"},"body":"On Wed, Jul 20, 2022 at 03:21:44PM -0400, Konstantin Ryabitsev wrote:\n> On Mon, Jul 18, 2022 at 02:24:07PM -0700, Glen Choo wrote:\n> > > I just started working on this for b4, with the notable difference that the\n> > > change-id trailer is used in the cover letter instead of in individual\n> > > commits, which moves the concept of \"change\" from a single commit to a series\n> > > of commits. IMO, it's much more useful in that scope, because as series are\n> > > reviewed and iterated, individual patches can get squashed, split up or\n> > > otherwise transformed.\n> > \n> > My 2 cents, since I used to use Gerrit a lot :)\n> > \n> > I find persistent per-commit ids really useful, even when patches get\n> > moved around. E.g. Gerrit can show and diff previous versions of the\n> > patch, which makes it really easy to tell how the patch has evolved\n> > over time.\n> \n> The kernel community has repeatedly rejected per-patch Change-id trailers\n> because they carry no meaningful information outside of the gerrit system on\n> which they were created. Seeing a Change-Id trailer in a commit tells you\n> nothing about the history of that commit unless you know the gerrit system on\n\nUnless you happen to see another patch with the same ID, and for that to\nhappen the ID needs to be generated when the commit is created (not when\nit's uploaded to gerrit or sent to a mailing list), and preserved by\ndefault in all processing of the commit.\n\nThen you can actually track the commit as it evolves in tools like\npatchwork, in theory.\n\nThanks\n\nMichal\n"},{"id":"459553","messageId":"Yth9TCCEXfmagaaw@mit.edu","threadId":"58181","inReplyTo":"20220720192144.mxdemgcdjxb2klgl@nitro.local","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Theodore Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2022-07-20T22:10:20Z","receivedAt":"2022-07-20T22:10:34Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Wed, Jul 20, 2022 at 03:21:44PM -0400, Konstantin Ryabitsev wrote:\n> The kernel community has repeatedly rejected per-patch Change-id trailers\n> because they carry no meaningful information outside of the gerrit system on\n> which they were created. Seeing a Change-Id trailer in a commit tells you\n> nothing about the history of that commit unless you know the gerrit system on\n> which this patch was reviewed (and have access to it, which is not a given).\n\nThe \"no meaningful information outside of the gerrit system\" is the\nkey.  This was extensively discussed in the\nksummit-discuss@lists.linux-foundation.org mailing list in late August\n2019, subject line \"Allowing something Change-Id (or something like\nit) in kernel commits\".  Quoting from Linus Torvalds:\n\n    From: Linus Torvalds\n    Date: Thu, 22 Aug 2019 17:17:05 -0700\n    Message-Id: CAHk-=whFbgy4RXG11c_=S7O-248oWmwB_aZOcWzWMVh3w7=RCw@mail.gmail.com\n\n    No. That's not it at all. It's not \"dislike gerrit\".\n\n    It's \"dislike pointless garbage\".\n\n    If the gerrit database is public and searchable using the uuid, then\n    that would make the uuid useful to outsiders. And instead of just\n    putting a UUID (which is hard to look up unless you know where it came\n    from), make it be that \"Link:\" that gives not just the UUID, but also\n    gives you the metadata for that UUID to be looked up.\n\n    But so far, in every single case the uuid's I've ever seen have been\n    pointless garbage, that aren't useful in general to public open source\n    developers, and as such shouldn't be in the git tree.\n\n    See the difference?\n\n    So if you guys make the gerrit database actually public, and then\n    start adding \"Link: ...\" tags so that we can see what they point to, I\n    think people will be more than supportive of it.\n\n    But if it's some stupid and pointless UUID that is useful to nobody\n    outside of google (or special magical groups of people associated with\n    it), then I will personally continue to be very much against it.\n\nSo....  imagine if we had some kind of search service, maybe homed at\nlore.kernel.org, where given a particular \"Change Id\" --- and it could\nlook either like a Gerrit-style Change-Id or something else like a URL\nor URL-like (it matters not) the search service could give you a list\nof:\n\n  * All mailing list threads where the body contained the \"Change-Id:\n    XXX\" id, so we could find the previous versions of the commit, and\n    the reviews that took place on a mailing list.  (And this could be\n    either a pointer to lore.kernel.org and/or a patchwork URL.)\n\n  * All URL's to public gerrit servers where that patch may have been reviewed.\n\n  * A list of git Commit ID's from a set of \"interesting\" git trees\n    (e.g., the upstream Linux tree, the Long Term Stable trees, maybe\n    some other interesting trees ala Android Common, etc.\n\nIf we had such a thing, as opposed to something that only worked in a\nclosed private garden like an internal Gerrit server sitting behind a\ncorporate firewall, even if the patch initially was developed in a\nclosed private Gerrit ecosystem --- if the moment it was published for\nexternal upstream review, it would get captured by this search\nservice, then the Change ID would be useful.  And if that Change ID\ncould also be used to find out how the patch was ported to various\nstable or productg trees, then it would be even more useful --- and\nthen people would probably find it to be useful, and resistance to\nhaving a per-commit Change-ID would probably drop, or perhaps, even\nenthusiastically embraced, because people could actually see the\n*value* behind it.\n\nTo do this, we would need to have various tools, such as Patchwork,\nGerrit, Git, public-inbox, etc., treat Change-ID as a first-class\nindexed object, so that you could quickly map from a Change-ID to a\ngit commit in a particular git tree (if present), or to set of\npublic-inbox URL's, or a set of patchwork URL's, etc.\n\nAnd then we would need some kind of aggregation service which would\naggregate the information from all of the various sources\n(public-inbox, git, Gerrit, Patchwork, etc.) and then gave users a\nsingle \"front door\" where they could submit a Change-Id, and find all\nthe patch history, patch review comments, and later, patch backports\nand forward ports.\n\nThe question is ---- is this doable?   And who will do the work?   :-)\n\n    \t     \t     \t     \t       - Ted\n"},{"id":"459607","messageId":"CAFQ2z_Mc6Z-FyFkUURMCM11yQTH+Y2PLBrAKB-16BXv159=oOA@mail.gmail.com","threadId":"58181","inReplyTo":"Yth9TCCEXfmagaaw@mit.edu","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Han-Wen Nienhuys","fromEmail":"hanwen@google.com","sentAt":"2022-07-21T11:57:56Z","receivedAt":"2022-07-21T11:58:12Z","isPatch":false,"sender":{"key":"hanwen@google.com","avatar":"https://avatars.githubusercontent.com/u/31547?v=4"},"body":"On Thu, Jul 21, 2022 at 12:10 AM Theodore Ts'o <tytso@mit.edu> wrote:\n> On Wed, Jul 20, 2022 at 03:21:44PM -0400, Konstantin Ryabitsev wrote:\n> > The kernel community has repeatedly rejected per-patch Change-id trailers\n> > because they carry no meaningful information outside of the gerrit system on\n> > which they were created. Seeing a Change-Id trailer in a commit tells you\n> > nothing about the history of that commit unless you know the gerrit system on\n> > which this patch was reviewed (and have access to it, which is not a given).\n>\n> The \"no meaningful information outside of the gerrit system\" is the\n> key.  This was extensively discussed in the\n> ksummit-discuss@lists.linux-foundation.org mailing list in late August\n> 2019, subject line \"Allowing something Change-Id (or something like\n> it) in kernel commits\".  Quoting from Linus Torvalds:\n>\n>     From: Linus Torvalds\n>     Date: Thu, 22 Aug 2019 17:17:05 -0700\n>     Message-Id: CAHk-=whFbgy4RXG11c_=S7O-248oWmwB_aZOcWzWMVh3w7=RCw@mail.gmail.com\n>\n>     No. That's not it at all. It's not \"dislike gerrit\".\n>\n>     It's \"dislike pointless garbage\".\n>\n>     If the gerrit database is public and searchable using the uuid, then\n>     that would make the uuid useful to outsiders. And instead of just\n>     putting a UUID (which is hard to look up unless you know where it came\n>     from), make it be that \"Link:\" that gives not just the UUID, but also\n>     gives you the metadata for that UUID to be looked up.\n>..\n>     So if you guys make the gerrit database actually public, and then\n>     start adding \"Link: ...\" tags so that we can see what they point to, I\n>     think people will be more than supportive of it.\n\nSupport for the \"Link:\" footer as a change ID has been implemented in\nGerrit as of https://gerrit.googlesource.com/gerrit/+/8cab93302d9c35316d691e848b67e687a68182b5\n(available in Gerrit 3.3 and onwards).  I'm not sure if it has seen\nmuch use, though.\n\n-- \nHan-Wen Nienhuys - Google Munich\nI work 80%. Don't expect answers from me on Fridays.\n--\nGoogle Germany GmbH, Erika-Mann-Strasse 33, 80636 Munich\nRegistergericht und -nummer: Hamburg, HRB 86891\nSitz der Gesellschaft: Hamburg\nGeschäftsführer: Paul Manicle, Liana Sebastian\n"},{"id":"459656","messageId":"87a692e8vj.fsf@vps.thesusis.net","threadId":"58181","inReplyTo":"220718.86ilnuw8jo.gmgdl@evledraar.gmail.com","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Phillip Susi","fromEmail":"phill@thesusis.net","sentAt":"2022-07-21T16:18:17Z","receivedAt":"2022-07-21T16:28:13Z","isPatch":false,"sender":{"key":"phill@thesusis.net","avatar":null},"body":"\nÆvar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n\n> This has come up a bunch of times. I think that the thing git itself\n> should be doing is to lean into the same notion that we use for tracking\n> renames. I.e. we don't, we analyze history after-the-fact and spot the\n> renames for you.\n\nI've never been a big fan of that quality of git because it is\ninherently unreliable.\n"},{"id":"459682","messageId":"CAE1pOi1pS76iXU8j=A54wPGHC7qofxrPDAO4uyy0d6yMxeQwvw@mail.gmail.com","threadId":"58181","inReplyTo":"87a692e8vj.fsf@vps.thesusis.net","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Hilco Wijbenga","fromEmail":"hilco.wijbenga@gmail.com","sentAt":"2022-07-21T18:58:24Z","receivedAt":"2022-07-21T18:58:46Z","isPatch":false,"sender":{"key":"hilco.wijbenga@gmail.com","avatar":null},"body":"On Thu, Jul 21, 2022 at 9:39 AM Phillip Susi <phill@thesusis.net> wrote:\n> Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n>\n> > This has come up a bunch of times. I think that the thing git itself\n> > should be doing is to lean into the same notion that we use for tracking\n> > renames. I.e. we don't, we analyze history after-the-fact and spot the\n> > renames for you.\n>\n> I've never been a big fan of that quality of git because it is\n> inherently unreliable.\n\nIndeed, which would be fine ... if there were a way to tell Git, \"no\nthis is not a rename\" or \"hey, you missed this rename\" but there\nisn't.\n\nReading previous messages, it seems like the\nafter-the-fact-rename-heuristic makes the Git code simpler. That is a\nperfectly valid argument for not supporting \"explicit\" renames but I\nhave seen several messages from which I inferred that rename handling\nwas deemed a \"solved problem\". And _that_, at least in my experience,\nis definitely not the case.\n"},{"id":"459791","messageId":"6426b5c3-0a09-f641-9876-3534b0abd96d@iee.email","threadId":"58181","inReplyTo":"CAE1pOi1pS76iXU8j=A54wPGHC7qofxrPDAO4uyy0d6yMxeQwvw@mail.gmail.com","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2022-07-22T20:08:56Z","receivedAt":"2022-07-22T20:09:21Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"On 21/07/2022 19:58, Hilco Wijbenga wrote:\n> On Thu, Jul 21, 2022 at 9:39 AM Phillip Susi <phill@thesusis.net> wrote:\n>> Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n>>\n>>> This has come up a bunch of times. I think that the thing git itself\n>>> should be doing is to lean into the same notion that we use for tracking\n>>> renames. I.e. we don't, we analyze history after-the-fact and spot the\n>>> renames for you.\n>> I've never been a big fan of that quality of git because it is\n>> inherently unreliable.\n> Indeed, which would be fine ... if there were a way to tell Git, \"no\n> this is not a rename\" or \"hey, you missed this rename\" but there\n> isn't.\n>\n> Reading previous messages, it seems like the\n> after-the-fact-rename-heuristic makes the Git code simpler. That is a\n> perfectly valid argument for not supporting \"explicit\" renames but I\n> have seen several messages from which I inferred that rename handling\n> was deemed a \"solved problem\". And _that_, at least in my experience,\n> is definitely not the case.\n\nPart of the rename problem is that there can be many different routes to\nthe same result, and often the route used isn't the one 'specified' by\nthose who wish a complicated rename process to have happened 'their\nway', plus people forget to record what they actually did. Attempting to\ncapture what happened still results major gaps in the record.\n\nIt's nice to believe that in software we could perfectly capture the\ncopy/edit/rename processes between revisions, but with humans in the\nloop it just doesn't work as planned.\n\nHence the value of Git is that it does record faithfully the end points,\nand allows a moderately standardised way of viewing the perceived rename\nprocess. It also removes the external attempts at 'control' of the\nrevision record that bedevil other approaches.\n\nPhilip\n"},{"id":"459793","messageId":"20220722203642.GD17705@kitsune.suse.cz","threadId":"58181","inReplyTo":"6426b5c3-0a09-f641-9876-3534b0abd96d@iee.email","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Michal Suchánek","fromEmail":"msuchanek@suse.de","sentAt":"2022-07-22T20:36:42Z","receivedAt":"2022-07-22T20:36:49Z","isPatch":false,"sender":{"key":"msuchanek@suse.de","avatar":"https://avatars.githubusercontent.com/u/787652?v=4"},"body":"On Fri, Jul 22, 2022 at 09:08:56PM +0100, Philip Oakley wrote:\n> On 21/07/2022 19:58, Hilco Wijbenga wrote:\n> > On Thu, Jul 21, 2022 at 9:39 AM Phillip Susi <phill@thesusis.net> wrote:\n> >> Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n> >>\n> >>> This has come up a bunch of times. I think that the thing git itself\n> >>> should be doing is to lean into the same notion that we use for tracking\n> >>> renames. I.e. we don't, we analyze history after-the-fact and spot the\n> >>> renames for you.\n> >> I've never been a big fan of that quality of git because it is\n> >> inherently unreliable.\n> > Indeed, which would be fine ... if there were a way to tell Git, \"no\n> > this is not a rename\" or \"hey, you missed this rename\" but there\n> > isn't.\n> >\n> > Reading previous messages, it seems like the\n> > after-the-fact-rename-heuristic makes the Git code simpler. That is a\n> > perfectly valid argument for not supporting \"explicit\" renames but I\n> > have seen several messages from which I inferred that rename handling\n> > was deemed a \"solved problem\". And _that_, at least in my experience,\n> > is definitely not the case.\n> \n> Part of the rename problem is that there can be many different routes to\n> the same result, and often the route used isn't the one 'specified' by\n> those who wish a complicated rename process to have happened 'their\n> way', plus people forget to record what they actually did. Attempting to\n> capture what happened still results major gaps in the record.\n\nDoesn't git have rebase?\n\nIt is not required that the rename is captured perfectly every time so\nlong as it can be amended later.\n\nThanks\n\nMichal\n"},{"id":"459799","messageId":"CA+P7+xr+k35RXoGv-O96fsfOJ+sg65HrVvt-3JKYAzerA0TJRw@mail.gmail.com","threadId":"58181","inReplyTo":"20220722203642.GD17705@kitsune.suse.cz","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Jacob Keller","fromEmail":"jacob.keller@gmail.com","sentAt":"2022-07-22T22:46:22Z","receivedAt":"2022-07-22T22:46:36Z","isPatch":false,"sender":{"key":"jacob.keller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/874719?v=4"},"body":"On Fri, Jul 22, 2022 at 1:42 PM Michal Suchánek <msuchanek@suse.de> wrote:\n>\n> On Fri, Jul 22, 2022 at 09:08:56PM +0100, Philip Oakley wrote:\n> > On 21/07/2022 19:58, Hilco Wijbenga wrote:\n> > > On Thu, Jul 21, 2022 at 9:39 AM Phillip Susi <phill@thesusis.net> wrote:\n> > >> Ęvar Arnfjörš Bjarmason <avarab@gmail.com> writes:\n> > >>\n> > >>> This has come up a bunch of times. I think that the thing git itself\n> > >>> should be doing is to lean into the same notion that we use for tracking\n> > >>> renames. I.e. we don't, we analyze history after-the-fact and spot the\n> > >>> renames for you.\n> > >> I've never been a big fan of that quality of git because it is\n> > >> inherently unreliable.\n> > > Indeed, which would be fine ... if there were a way to tell Git, \"no\n> > > this is not a rename\" or \"hey, you missed this rename\" but there\n> > > isn't.\n> > >\n> > > Reading previous messages, it seems like the\n> > > after-the-fact-rename-heuristic makes the Git code simpler. That is a\n> > > perfectly valid argument for not supporting \"explicit\" renames but I\n> > > have seen several messages from which I inferred that rename handling\n> > > was deemed a \"solved problem\". And _that_, at least in my experience,\n> > > is definitely not the case.\n> >\n> > Part of the rename problem is that there can be many different routes to\n> > the same result, and often the route used isn't the one 'specified' by\n> > those who wish a complicated rename process to have happened 'their\n> > way', plus people forget to record what they actually did. Attempting to\n> > capture what happened still results major gaps in the record.\n>\n> Doesn't git have rebase?\n>\n> It is not required that the rename is captured perfectly every time so\n> long as it can be amended later.\n>\n> Thanks\n>\n> Michal\n\nRebase is typically reserved only to modify commits which are not yet\n\"permanent\". Once a commit starts being referenced by many others it\nbecomes more and more difficult to rebase it. Any rebase effectively\ncreates a new commit.\n\nThere are multiple threads discussing renames and handling them in git\nin the past which are worth re-reading, including at least\n\nhttps://public-inbox.org/git/Pine.LNX.4.58.0504141102430.7211@ppc970.osdl.org/\n\nA fuller analysis here too:\nhttps://public-inbox.org/git/Pine.LNX.4.64.0510221251330.10477@g5.osdl.org/\n\nAs mentioned above in this thread, depending on what context you are\nusing, a change to a commit could be many to many: i.e. a commit which\nsplits into 2, or 3 commits merging into one, or 3 commits splitting\napart and then becoming 2 commits. When that happens, what \"change id\"\ndo you use for each commit?\n"},{"id":"459831","messageId":"20220723070055.GE17705@kitsune.suse.cz","threadId":"58181","inReplyTo":"CA+P7+xr+k35RXoGv-O96fsfOJ+sg65HrVvt-3JKYAzerA0TJRw@mail.gmail.com","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Michal Suchánek","fromEmail":"msuchanek@suse.de","sentAt":"2022-07-23T07:00:55Z","receivedAt":"2022-07-23T07:01:04Z","isPatch":false,"sender":{"key":"msuchanek@suse.de","avatar":"https://avatars.githubusercontent.com/u/787652?v=4"},"body":"On Fri, Jul 22, 2022 at 03:46:22PM -0700, Jacob Keller wrote:\n> On Fri, Jul 22, 2022 at 1:42 PM Michal Suchánek <msuchanek@suse.de> wrote:\n> >\n> > On Fri, Jul 22, 2022 at 09:08:56PM +0100, Philip Oakley wrote:\n> > > On 21/07/2022 19:58, Hilco Wijbenga wrote:\n> > > > On Thu, Jul 21, 2022 at 9:39 AM Phillip Susi <phill@thesusis.net> wrote:\n> > > >> Ęvar Arnfjörš Bjarmason <avarab@gmail.com> writes:\n> > > >>\n> > > >>> This has come up a bunch of times. I think that the thing git itself\n> > > >>> should be doing is to lean into the same notion that we use for tracking\n> > > >>> renames. I.e. we don't, we analyze history after-the-fact and spot the\n> > > >>> renames for you.\n> > > >> I've never been a big fan of that quality of git because it is\n> > > >> inherently unreliable.\n> > > > Indeed, which would be fine ... if there were a way to tell Git, \"no\n> > > > this is not a rename\" or \"hey, you missed this rename\" but there\n> > > > isn't.\n> > > >\n> > > > Reading previous messages, it seems like the\n> > > > after-the-fact-rename-heuristic makes the Git code simpler. That is a\n> > > > perfectly valid argument for not supporting \"explicit\" renames but I\n> > > > have seen several messages from which I inferred that rename handling\n> > > > was deemed a \"solved problem\". And _that_, at least in my experience,\n> > > > is definitely not the case.\n> > >\n> > > Part of the rename problem is that there can be many different routes to\n> > > the same result, and often the route used isn't the one 'specified' by\n> > > those who wish a complicated rename process to have happened 'their\n> > > way', plus people forget to record what they actually did. Attempting to\n> > > capture what happened still results major gaps in the record.\n> >\n> > Doesn't git have rebase?\n> >\n> > It is not required that the rename is captured perfectly every time so\n> > long as it can be amended later.\n> >\n> > Thanks\n> >\n> > Michal\n> \n> Rebase is typically reserved only to modify commits which are not yet\n> \"permanent\". Once a commit starts being referenced by many others it\n> becomes more and more difficult to rebase it. Any rebase effectively\n> creates a new commit.\n> \n> There are multiple threads discussing renames and handling them in git\n> in the past which are worth re-reading, including at least\n> \n> https://public-inbox.org/git/Pine.LNX.4.58.0504141102430.7211@ppc970.osdl.org/\n> \n> A fuller analysis here too:\n> https://public-inbox.org/git/Pine.LNX.4.64.0510221251330.10477@g5.osdl.org/\n> \n> As mentioned above in this thread, depending on what context you are\n> using, a change to a commit could be many to many: i.e. a commit which\n> splits into 2, or 3 commits merging into one, or 3 commits splitting\n> apart and then becoming 2 commits. When that happens, what \"change id\"\n> do you use for each commit?\n\nSame as commit message and any trailers you might have - they are\npreserved, concatenated, and can be regenerated.\n\nThanks\n\nMichal\n"},{"id":"459843","messageId":"CABPp-BHNYbLEWeG+XSzGxcTxsQ2wA2COX6DqtvVZ6Nm1KG7CEQ@mail.gmail.com","threadId":"58181","inReplyTo":"kl6lo7xmt8qw.fsf@chooglen-macbookpro.roam.corp.google.com","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2022-07-24T05:09:02Z","receivedAt":"2022-07-24T05:09:18Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Mon, Jul 18, 2022 at 2:29 PM Glen Choo <chooglen@google.com> wrote:\n>\n> Konstantin Ryabitsev <konstantin@linuxfoundation.org> writes:\n>\n> > On Mon, Jul 18, 2022 at 06:18:11PM +0100, Stephen Finucane wrote:\n> >> ...to track evolution of a patch through time.\n> >>\n> >> tl;dr: How hard would it be to retrofit an 'ChangeID' concept à la the 'Change-\n> >> ID' trailer used by Gerrit into git core?\n> >\n> > I just started working on this for b4, with the notable difference that the\n> > change-id trailer is used in the cover letter instead of in individual\n> > commits, which moves the concept of \"change\" from a single commit to a series\n> > of commits. IMO, it's much more useful in that scope, because as series are\n> > reviewed and iterated, individual patches can get squashed, split up or\n> > otherwise transformed.\n>\n> My 2 cents, since I used to use Gerrit a lot :)\n>\n> I find persistent per-commit ids really useful, even when patches get\n> moved around. E.g. Gerrit can show and diff previous versions of the\n> patch, which makes it really easy to tell how the patch has evolved\n> over time.\n>\n> That's not to say that we don't need per-topic ids though ;) E.g. Gerrit\n> is pretty bad at handling whole topics - it does naive mapping on a\n> per-commit level, so it has no concept of \"these (n - 1) patches should\n> replace these n patches\".\n>\n> I, for one, would love to see some kind of \"rewrite tracking\" in Git.\n> One use case that comes up often is downstream patches, where patches\n> are continuously rebased onto a new upstream; in those cases, it's\n> pretty hard to keep track of how the patch has changed over time\n\nTwo angles I can think of that partially address this:\n\n1) If you have the old commits still around and know what they were,\nyou can run range-diff to see differences between any pair of versions\nof the commits.\n\n2) cherry-picks and reverts might already include a link to an \"old\"\ncommit for you in the commit message (\"cherry picked from commit\n<hash>\" or \"This reverts <hash>\").  Those could be used to show how\nthe new commit differs from what would have been done with an\nautomatic cherry-pick or automatic revert.  (By \"automatic\", I\nbasically mean what the state of files in the working tree would be\nwhen the operation stops to allow users to resolve conflicts.)  In\nfact, I wrote some patches to do precisely this quite a while ago\nwhich are up at https://github.com/gitgitgadget/git/pull/1151 if\nyou're curious.  But this approach is not useful for general rebasing,\nbecause there's no automated way to find out what the original commit\nwas so that you can take a look at such a difference.\n"},{"id":"459844","messageId":"CABPp-BFGLXRvwZGdF543me2qBXq3HB-TuzW6j7GVb6ATw3qNeQ@mail.gmail.com","threadId":"58181","inReplyTo":"20220722203642.GD17705@kitsune.suse.cz","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2022-07-24T05:10:11Z","receivedAt":"2022-07-24T05:10:28Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Fri, Jul 22, 2022 at 1:42 PM Michal Suchánek <msuchanek@suse.de> wrote:\n>\n> On Fri, Jul 22, 2022 at 09:08:56PM +0100, Philip Oakley wrote:\n> > On 21/07/2022 19:58, Hilco Wijbenga wrote:\n> > > On Thu, Jul 21, 2022 at 9:39 AM Phillip Susi <phill@thesusis.net> wrote:\n> > >> Ęvar Arnfjörš Bjarmason <avarab@gmail.com> writes:\n> > >>\n> > >>> This has come up a bunch of times. I think that the thing git itself\n> > >>> should be doing is to lean into the same notion that we use for tracking\n> > >>> renames. I.e. we don't, we analyze history after-the-fact and spot the\n> > >>> renames for you.\n> > >> I've never been a big fan of that quality of git because it is\n> > >> inherently unreliable.\n> > > Indeed, which would be fine ... if there were a way to tell Git, \"no\n> > > this is not a rename\" or \"hey, you missed this rename\" but there\n> > > isn't.\n> > >\n> > > Reading previous messages, it seems like the\n> > > after-the-fact-rename-heuristic makes the Git code simpler. That is a\n> > > perfectly valid argument for not supporting \"explicit\" renames but I\n> > > have seen several messages from which I inferred that rename handling\n> > > was deemed a \"solved problem\". And _that_, at least in my experience,\n> > > is definitely not the case.\n> >\n> > Part of the rename problem is that there can be many different routes to\n> > the same result, and often the route used isn't the one 'specified' by\n> > those who wish a complicated rename process to have happened 'their\n> > way', plus people forget to record what they actually did. Attempting to\n> > capture what happened still results major gaps in the record.\n>\n> Doesn't git have rebase?\n>\n> It is not required that the rename is captured perfectly every time so\n> long as it can be amended later.\n\n\"so long as\".  Therefore, since it can't be amended after the commit\nis accepted/merged, it is required that this auxiliary data be\ncaptured perfectly before that time if it's going to be captured at\nall.\n\nDid I read that right?\n"},{"id":"459845","messageId":"CABPp-BEYQOtr6EZmi4emKRKNVXS3071Ud90jiLycdGXGG+YqgQ@mail.gmail.com","threadId":"58181","inReplyTo":"20220723070055.GE17705@kitsune.suse.cz","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2022-07-24T05:23:09Z","receivedAt":"2022-07-24T05:23:37Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Sat, Jul 23, 2022 at 12:44 AM Michal Suchánek <msuchanek@suse.de> wrote:\n>\n> On Fri, Jul 22, 2022 at 03:46:22PM -0700, Jacob Keller wrote:\n> > On Fri, Jul 22, 2022 at 1:42 PM Michal Suchánek <msuchanek@suse.de> wrote:\n> > >\n> > > On Fri, Jul 22, 2022 at 09:08:56PM +0100, Philip Oakley wrote:\n[...]\n> > > > Part of the rename problem is that there can be many different routes to\n> > > > the same result, and often the route used isn't the one 'specified' by\n> > > > those who wish a complicated rename process to have happened 'their\n> > > > way', plus people forget to record what they actually did. Attempting to\n> > > > capture what happened still results major gaps in the record.\n> > >\n> > > Doesn't git have rebase?\n> > >\n> > > It is not required that the rename is captured perfectly every time so\n> > > long as it can be amended later.\n> > >\n> >\n> > Rebase is typically reserved only to modify commits which are not yet\n> > \"permanent\". Once a commit starts being referenced by many others it\n> > becomes more and more difficult to rebase it. Any rebase effectively\n> > creates a new commit.\n> >\n> > There are multiple threads discussing renames and handling them in git\n> > in the past which are worth re-reading, including at least\n> >\n> > https://public-inbox.org/git/Pine.LNX.4.58.0504141102430.7211@ppc970.osdl.org/\n> >\n> > A fuller analysis here too:\n> > https://public-inbox.org/git/Pine.LNX.4.64.0510221251330.10477@g5.osdl.org/\n> >\n> > As mentioned above in this thread, depending on what context you are\n> > using, a change to a commit could be many to many: i.e. a commit which\n> > splits into 2, or 3 commits merging into one, or 3 commits splitting\n> > apart and then becoming 2 commits. When that happens, what \"change id\"\n> > do you use for each commit?\n>\n> Same as commit message and any trailers you might have - they are\n> preserved, concatenated\n\nExactly how are they concatenated?  Is that a user operation, or\nsomething a Git command does automatically?  Which commands and which\ncircumstances?  If users do it, what's the UI for them to discover\nwhat the fields are, for them to discover whether such a thing might\nbe needed or beneficial, and the UI for them to change these fields?\nThis sounds like a massive UX/UI issue that I don't have a clue how to\ntackle (assuming I wanted to).\n\n> and can be regenerated.\n\n\"can be\".  But generally won't be even when it should be, right?\n\nCommitter name/email/date basically don't even exist as far as many\nGit users are concerned.  They aren't shown in the default log output\n(which greatly saddens me), and even after attempting to educate users\nfor well over a decade now, I still routinely find developers who are\nsurprised that these things exist.\n\nGiven that committer name/email/date aren't shown with --pretty=full\nbut with the lame option name --pretty=fuller, I can't see why it'd\nmake any sense to show Change-Ids in the log output by default.\n\nBut if it's not shown -- and by default -- then it doesn't exist for\nmany users.  And if it doesn't exist, users aren't going to fix it\nwhen they need to.\n\n(Even if it were shown by default, it's not clear to me that users\nwould know when to fix it, or how to fix it, or even care to fix it\nand instead view it as a pedantic requirement being foisted on them.)\n\nI think the \"many-to-many issue\" others have raised in this thread is\nan important, big, and thorny problem.  I think it has the potential\nto be a minefield of UX and a steady stream of bug reports.  And\nseeing proponents of Change-Id just dismissing the issue makes me all\nthe more suspicious of the proposal in the first place.\n"},{"id":"459847","messageId":"20220724085419.GG17705@kitsune.suse.cz","threadId":"58181","inReplyTo":"CABPp-BEYQOtr6EZmi4emKRKNVXS3071Ud90jiLycdGXGG+YqgQ@mail.gmail.com","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Michal Suchánek","fromEmail":"msuchanek@suse.de","sentAt":"2022-07-24T08:54:19Z","receivedAt":"2022-07-24T08:54:25Z","isPatch":false,"sender":{"key":"msuchanek@suse.de","avatar":"https://avatars.githubusercontent.com/u/787652?v=4"},"body":"On Sat, Jul 23, 2022 at 10:23:09PM -0700, Elijah Newren wrote:\n> On Sat, Jul 23, 2022 at 12:44 AM Michal Suchánek <msuchanek@suse.de> wrote:\n> >\n> > On Fri, Jul 22, 2022 at 03:46:22PM -0700, Jacob Keller wrote:\n> > > On Fri, Jul 22, 2022 at 1:42 PM Michal Suchánek <msuchanek@suse.de> wrote:\n> > > >\n> > > > On Fri, Jul 22, 2022 at 09:08:56PM +0100, Philip Oakley wrote:\n> [...]\n> > > > > Part of the rename problem is that there can be many different routes to\n> > > > > the same result, and often the route used isn't the one 'specified' by\n> > > > > those who wish a complicated rename process to have happened 'their\n> > > > > way', plus people forget to record what they actually did. Attempting to\n> > > > > capture what happened still results major gaps in the record.\n> > > >\n> > > > Doesn't git have rebase?\n> > > >\n> > > > It is not required that the rename is captured perfectly every time so\n> > > > long as it can be amended later.\n> > > >\n> > >\n> > > Rebase is typically reserved only to modify commits which are not yet\n> > > \"permanent\". Once a commit starts being referenced by many others it\n> > > becomes more and more difficult to rebase it. Any rebase effectively\n> > > creates a new commit.\n> > >\n> > > There are multiple threads discussing renames and handling them in git\n> > > in the past which are worth re-reading, including at least\n> > >\n> > > https://public-inbox.org/git/Pine.LNX.4.58.0504141102430.7211@ppc970.osdl.org/\n> > >\n> > > A fuller analysis here too:\n> > > https://public-inbox.org/git/Pine.LNX.4.64.0510221251330.10477@g5.osdl.org/\n> > >\n> > > As mentioned above in this thread, depending on what context you are\n> > > using, a change to a commit could be many to many: i.e. a commit which\n> > > splits into 2, or 3 commits merging into one, or 3 commits splitting\n> > > apart and then becoming 2 commits. When that happens, what \"change id\"\n> > > do you use for each commit?\n> >\n> > Same as commit message and any trailers you might have - they are\n> > preserved, concatenated\n> \n> Exactly how are they concatenated?  Is that a user operation, or\n> something a Git command does automatically?  Which commands and which\n> circumstances?  If users do it, what's the UI for them to discover\n> what the fields are, for them to discover whether such a thing might\n> be needed or beneficial, and the UI for them to change these fields?\n> This sounds like a massive UX/UI issue that I don't have a clue how to\n> tackle (assuming I wanted to).\n\nCurrently when you squash commits you get both commit messages\nconcatenated, including any trailers.\n\nYou are free to adjust as you see fit.\n\n> \n> > and can be regenerated.\n> \n> \"can be\".  But generally won't be even when it should be, right?\n\n\"when it should\" is not something that can be programmatically\nderrmined, so it's up to the user, and sure, there will be cases where\nsomebody thinks it \"should\" but it has not. Then they can complain, just\nlike with any other trailer we already have today.\n\n> I think the \"many-to-many issue\" others have raised in this thread is\n> an important, big, and thorny problem.  I think it has the potential\n> to be a minefield of UX and a steady stream of bug reports.  And\n> seeing proponents of Change-Id just dismissing the issue makes me all\n> the more suspicious of the proposal in the first place.\n\nAnd how do you get this many to many situation in the first place?\n\nYou reset to a base before your changes and create completely new series\nwith completely new messages and everything?\n\nThen of course you get completely new trailers as well unless you\nsomehow fish out some metadata from the old commits and manually apply\nthem.\n\nI don't see any functionality in git that does many to many commits\ntransform in one step. It's always just split/merge.\n\nThanks\n\nMichal\n"},{"id":"459848","messageId":"20220724085921.GH17705@kitsune.suse.cz","threadId":"58181","inReplyTo":"CABPp-BFGLXRvwZGdF543me2qBXq3HB-TuzW6j7GVb6ATw3qNeQ@mail.gmail.com","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Michal Suchánek","fromEmail":"msuchanek@suse.de","sentAt":"2022-07-24T08:59:21Z","receivedAt":"2022-07-24T08:59:26Z","isPatch":false,"sender":{"key":"msuchanek@suse.de","avatar":"https://avatars.githubusercontent.com/u/787652?v=4"},"body":"On Sat, Jul 23, 2022 at 10:10:11PM -0700, Elijah Newren wrote:\n> On Fri, Jul 22, 2022 at 1:42 PM Michal Suchánek <msuchanek@suse.de> wrote:\n> >\n> > On Fri, Jul 22, 2022 at 09:08:56PM +0100, Philip Oakley wrote:\n> > > On 21/07/2022 19:58, Hilco Wijbenga wrote:\n> > > > On Thu, Jul 21, 2022 at 9:39 AM Phillip Susi <phill@thesusis.net> wrote:\n> > > >> Ęvar Arnfjörš Bjarmason <avarab@gmail.com> writes:\n> > > >>\n> > > >>> This has come up a bunch of times. I think that the thing git itself\n> > > >>> should be doing is to lean into the same notion that we use for tracking\n> > > >>> renames. I.e. we don't, we analyze history after-the-fact and spot the\n> > > >>> renames for you.\n> > > >> I've never been a big fan of that quality of git because it is\n> > > >> inherently unreliable.\n> > > > Indeed, which would be fine ... if there were a way to tell Git, \"no\n> > > > this is not a rename\" or \"hey, you missed this rename\" but there\n> > > > isn't.\n> > > >\n> > > > Reading previous messages, it seems like the\n> > > > after-the-fact-rename-heuristic makes the Git code simpler. That is a\n> > > > perfectly valid argument for not supporting \"explicit\" renames but I\n> > > > have seen several messages from which I inferred that rename handling\n> > > > was deemed a \"solved problem\". And _that_, at least in my experience,\n> > > > is definitely not the case.\n> > >\n> > > Part of the rename problem is that there can be many different routes to\n> > > the same result, and often the route used isn't the one 'specified' by\n> > > those who wish a complicated rename process to have happened 'their\n> > > way', plus people forget to record what they actually did. Attempting to\n> > > capture what happened still results major gaps in the record.\n> >\n> > Doesn't git have rebase?\n> >\n> > It is not required that the rename is captured perfectly every time so\n> > long as it can be amended later.\n> \n> \"so long as\".  Therefore, since it can't be amended after the commit\n> is accepted/merged, it is required that this auxiliary data be\n> captured perfectly before that time if it's going to be captured at\n> all.\n> \n> Did I read that right?\n\nOr it will be broken after it is merged, just as many other things in\ncommits that are accepted into history that is not to be modified\nanymore.\n\nThe only point I can see here is that if there is any user-crafted\nmetadata that describes renames then it should be considered advisory,\nand an option to override it should exist because it may be wrong.\n\nNonetheless, if such feature existed users that are willing to generate\nsuch metadata and review it before it gets merged may get more out of\nthe rename tracking than can be done automatically today.\n\nThanks\n\nMichal\n"},{"id":"459929","messageId":"CA+P7+xoygnvi8_8JjOSftahKZFC3bZBkzA-LQ8-xAp9fkV79pw@mail.gmail.com","threadId":"58181","inReplyTo":"CABPp-BEYQOtr6EZmi4emKRKNVXS3071Ud90jiLycdGXGG+YqgQ@mail.gmail.com","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Jacob Keller","fromEmail":"jacob.keller@gmail.com","sentAt":"2022-07-25T21:47:24Z","receivedAt":"2022-07-25T21:47:39Z","isPatch":false,"sender":{"key":"jacob.keller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/874719?v=4"},"body":"On Sat, Jul 23, 2022 at 10:23 PM Elijah Newren <newren@gmail.com> wrote:\n>\n> On Sat, Jul 23, 2022 at 12:44 AM Michal Suchánek <msuchanek@suse.de> wrote:\n> >\n> > On Fri, Jul 22, 2022 at 03:46:22PM -0700, Jacob Keller wrote:\n> > > On Fri, Jul 22, 2022 at 1:42 PM Michal Suchánek <msuchanek@suse.de> wrote:\n> > > >\n> > > > On Fri, Jul 22, 2022 at 09:08:56PM +0100, Philip Oakley wrote:\n> [...]\n> > > > > Part of the rename problem is that there can be many different routes to\n> > > > > the same result, and often the route used isn't the one 'specified' by\n> > > > > those who wish a complicated rename process to have happened 'their\n> > > > > way', plus people forget to record what they actually did. Attempting to\n> > > > > capture what happened still results major gaps in the record.\n> > > >\n> > > > Doesn't git have rebase?\n> > > >\n> > > > It is not required that the rename is captured perfectly every time so\n> > > > long as it can be amended later.\n> > > >\n> > >\n> > > Rebase is typically reserved only to modify commits which are not yet\n> > > \"permanent\". Once a commit starts being referenced by many others it\n> > > becomes more and more difficult to rebase it. Any rebase effectively\n> > > creates a new commit.\n> > >\n> > > There are multiple threads discussing renames and handling them in git\n> > > in the past which are worth re-reading, including at least\n> > >\n> > > https://public-inbox.org/git/Pine.LNX.4.58.0504141102430.7211@ppc970.osdl.org/\n> > >\n> > > A fuller analysis here too:\n> > > https://public-inbox.org/git/Pine.LNX.4.64.0510221251330.10477@g5.osdl.org/\n> > >\n> > > As mentioned above in this thread, depending on what context you are\n> > > using, a change to a commit could be many to many: i.e. a commit which\n> > > splits into 2, or 3 commits merging into one, or 3 commits splitting\n> > > apart and then becoming 2 commits. When that happens, what \"change id\"\n> > > do you use for each commit?\n> >\n> > Same as commit message and any trailers you might have - they are\n> > preserved, concatenated\n>\n> Exactly how are they concatenated?  Is that a user operation, or\n> something a Git command does automatically?  Which commands and which\n> circumstances?  If users do it, what's the UI for them to discover\n> what the fields are, for them to discover whether such a thing might\n> be needed or beneficial, and the UI for them to change these fields?\n> This sounds like a massive UX/UI issue that I don't have a clue how to\n> tackle (assuming I wanted to).\n>\n> > and can be regenerated.\n>\n> \"can be\".  But generally won't be even when it should be, right?\n>\n> Committer name/email/date basically don't even exist as far as many\n> Git users are concerned.  They aren't shown in the default log output\n> (which greatly saddens me), and even after attempting to educate users\n> for well over a decade now, I still routinely find developers who are\n> surprised that these things exist.\n>\n> Given that committer name/email/date aren't shown with --pretty=full\n> but with the lame option name --pretty=fuller, I can't see why it'd\n> make any sense to show Change-Ids in the log output by default.\n>\n> But if it's not shown -- and by default -- then it doesn't exist for\n> many users.  And if it doesn't exist, users aren't going to fix it\n> when they need to.\n>\n> (Even if it were shown by default, it's not clear to me that users\n> would know when to fix it, or how to fix it, or even care to fix it\n> and instead view it as a pedantic requirement being foisted on them.)\n>\n> I think the \"many-to-many issue\" others have raised in this thread is\n> an important, big, and thorny problem.  I think it has the potential\n> to be a minefield of UX and a steady stream of bug reports.  And\n> seeing proponents of Change-Id just dismissing the issue makes me all\n> the more suspicious of the proposal in the first place.\n\nI do think there is some value in having a sort of generic id like\nchange-id, but I do think we want to be careful about how exactly we\nhandle it.\n\nAs you say, if we hide it then users may not be aware of it, and if we\nmake it visible users who don't care may be annoyed. I don't think we\ncan fully automate it because of the nature of combining changes and\nsplitting changes require humans to decide which change keeps which\nID. Its not even clear when rebasing whether a split is going to\nhappen. A combine operation is easier to detect in rebase\n(fixup/squash), but determining which id to keep is not. Would we even\nwant to have support for \"this commit merges two and is now one, but\nwe keep both IDs because it really is both commits\"? That gets messy\npretty fast.\n\nUsers such as gerrit already simply use the trailer with Change-id and\nmanage to make it work by enforcing some constraints and assuming\nusers will know what to do (because otherwise they fail to interact\nwith gerrit servers).\n\nFor cases where it helps, I think its very valuable. Being able to\ntrack revisions of a series or a patch is super useful. Getting\nexternal tooling like public-inbox, patchworks, etc to use this would\nalso be useful. But I think we would want to sort out the situation a\nbit for how and when are they generated, when are they\nreplaced/re-generated, how this interacts with mailing etc.\n\nShould rebase just always regenerate? that loses a lot of value. I\nguess squashing could offer users a choice of which to keep? Fixup\nwould always keep the same one. And otherwise it becomes up to users\nto know when they need to copy from an old commit or refresh an\nexisting commit... Thats pretty much what gerrit does these days, if a\ncommit doesn't have the trailer it gets added, and if it does, its up\nto the user to know when to remove it or regenerate it... Since its a\ncommit message trailer it gets sent implicitly through the mailing\nlist unless removed.\n"},{"id":"459957","messageId":"CABPp-BHWueqbFMpz4jMPK9G3xY5=jvWLELLQribU69CYD5DjKw@mail.gmail.com","threadId":"58181","inReplyTo":"CA+P7+xoygnvi8_8JjOSftahKZFC3bZBkzA-LQ8-xAp9fkV79pw@mail.gmail.com","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2022-07-26T03:49:30Z","receivedAt":"2022-07-26T03:49:53Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Mon, Jul 25, 2022 at 2:47 PM Jacob Keller <jacob.keller@gmail.com> wrote:\n>\n> On Sat, Jul 23, 2022 at 10:23 PM Elijah Newren <newren@gmail.com> wrote:\n> >\n> > On Sat, Jul 23, 2022 at 12:44 AM Michal Suchánek <msuchanek@suse.de> wrote:\n> > >\n> > > On Fri, Jul 22, 2022 at 03:46:22PM -0700, Jacob Keller wrote:\n> > > > On Fri, Jul 22, 2022 at 1:42 PM Michal Suchánek <msuchanek@suse.de> wrote:\n> > > > >\n> > > > > On Fri, Jul 22, 2022 at 09:08:56PM +0100, Philip Oakley wrote:\n> > [...]\n> > > > > > Part of the rename problem is that there can be many different routes to\n> > > > > > the same result, and often the route used isn't the one 'specified' by\n> > > > > > those who wish a complicated rename process to have happened 'their\n> > > > > > way', plus people forget to record what they actually did. Attempting to\n> > > > > > capture what happened still results major gaps in the record.\n> > > > >\n> > > > > Doesn't git have rebase?\n> > > > >\n> > > > > It is not required that the rename is captured perfectly every time so\n> > > > > long as it can be amended later.\n> > > > >\n> > > >\n> > > > Rebase is typically reserved only to modify commits which are not yet\n> > > > \"permanent\". Once a commit starts being referenced by many others it\n> > > > becomes more and more difficult to rebase it. Any rebase effectively\n> > > > creates a new commit.\n> > > >\n> > > > There are multiple threads discussing renames and handling them in git\n> > > > in the past which are worth re-reading, including at least\n> > > >\n> > > > https://public-inbox.org/git/Pine.LNX.4.58.0504141102430.7211@ppc970.osdl.org/\n> > > >\n> > > > A fuller analysis here too:\n> > > > https://public-inbox.org/git/Pine.LNX.4.64.0510221251330.10477@g5.osdl.org/\n> > > >\n> > > > As mentioned above in this thread, depending on what context you are\n> > > > using, a change to a commit could be many to many: i.e. a commit which\n> > > > splits into 2, or 3 commits merging into one, or 3 commits splitting\n> > > > apart and then becoming 2 commits. When that happens, what \"change id\"\n> > > > do you use for each commit?\n> > >\n> > > Same as commit message and any trailers you might have - they are\n> > > preserved, concatenated\n> >\n> > Exactly how are they concatenated?  Is that a user operation, or\n> > something a Git command does automatically?  Which commands and which\n> > circumstances?  If users do it, what's the UI for them to discover\n> > what the fields are, for them to discover whether such a thing might\n> > be needed or beneficial, and the UI for them to change these fields?\n> > This sounds like a massive UX/UI issue that I don't have a clue how to\n> > tackle (assuming I wanted to).\n> >\n> > > and can be regenerated.\n> >\n> > \"can be\".  But generally won't be even when it should be, right?\n> >\n> > Committer name/email/date basically don't even exist as far as many\n> > Git users are concerned.  They aren't shown in the default log output\n> > (which greatly saddens me), and even after attempting to educate users\n> > for well over a decade now, I still routinely find developers who are\n> > surprised that these things exist.\n> >\n> > Given that committer name/email/date aren't shown with --pretty=full\n> > but with the lame option name --pretty=fuller, I can't see why it'd\n> > make any sense to show Change-Ids in the log output by default.\n> >\n> > But if it's not shown -- and by default -- then it doesn't exist for\n> > many users.  And if it doesn't exist, users aren't going to fix it\n> > when they need to.\n> >\n> > (Even if it were shown by default, it's not clear to me that users\n> > would know when to fix it, or how to fix it, or even care to fix it\n> > and instead view it as a pedantic requirement being foisted on them.)\n> >\n> > I think the \"many-to-many issue\" others have raised in this thread is\n> > an important, big, and thorny problem.  I think it has the potential\n> > to be a minefield of UX and a steady stream of bug reports.  And\n> > seeing proponents of Change-Id just dismissing the issue makes me all\n> > the more suspicious of the proposal in the first place.\n>\n> I do think there is some value in having a sort of generic id like\n> change-id, but I do think we want to be careful about how exactly we\n> handle it.\n>\n> As you say, if we hide it then users may not be aware of it, and if we\n> make it visible users who don't care may be annoyed. I don't think we\n> can fully automate it because of the nature of combining changes and\n> splitting changes require humans to decide which change keeps which\n> ID. Its not even clear when rebasing whether a split is going to\n> happen. A combine operation is easier to detect in rebase\n> (fixup/squash), but determining which id to keep is not. Would we even\n> want to have support for \"this commit merges two and is now one, but\n> we keep both IDs because it really is both commits\"? That gets messy\n> pretty fast.\n>\n> Users such as gerrit already simply use the trailer with Change-id and\n> manage to make it work by enforcing some constraints and assuming\n> users will know what to do (because otherwise they fail to interact\n> with gerrit servers).\n>\n> For cases where it helps, I think its very valuable. Being able to\n> track revisions of a series or a patch is super useful. Getting\n> external tooling like public-inbox, patchworks, etc to use this would\n> also be useful. But I think we would want to sort out the situation a\n> bit for how and when are they generated, when are they\n> replaced/re-generated, how this interacts with mailing etc.\n>\n> Should rebase just always regenerate? that loses a lot of value. I\n> guess squashing could offer users a choice of which to keep? Fixup\n> would always keep the same one. And otherwise it becomes up to users\n> to know when they need to copy from an old commit or refresh an\n> existing commit... Thats pretty much what gerrit does these days, if a\n> commit doesn't have the trailer it gets added, and if it does, its up\n> to the user to know when to remove it or regenerate it... Since its a\n> commit message trailer it gets sent implicitly through the mailing\n> list unless removed.\n\nYes, I fully agree it needs to be spelled out a lot more.  And not\njust obvious commands (everyone seems to focus on commit, cherry-pick,\nand rebase), but what about e.g. `git merge --squash`?\n\nAlso, as far as value goes, I have an interesting story related to\nChange-Ids (read the last sentence if you only want the summary):\n\n\n<long story>\n\nI have used Gerrit fairly heavily.  I maintained an instance for a few\nhundred developers for several years (inheriting it from others), and\nwas responsible for various build & release stuff related to one of\nthe larger products tracked in it (an approximately\nlinux-kernel-sized) product.  I also attended a Gerrit conference or\ntwo and submitted a few patches for Gerrit that were accepted.  So,\nfor context, I'm clearly not a Gerrit developer since I only submitted\na few patches, but I was the clear expert within my company on Gerrit.\nSo that's my background.\n\nSome background on the (insane) project management of the time\n(unrelated to Git or Gerrit or Change-Ids) is also important to\nunderstand this story:  Years ago, this project had well over 100\nactive branches (!!).  And yes, branches were being aggressively\nretired, but I still remember when we finally managed to get the count\nunder 200.  Each branch had important patches, and hundreds of patches\non these branches was not uncommon.  Yes, it was insane, and yes, I\nand many others were really happy when we eventually reached the land\nof sanity with just a few active branches (and all but the main one\nonly gets backport fixes).\n\nOne of the things I did to help us move in the direction towards\nsanity was a \"snowflake report\" -- a report that would help people\ndetermine which patches had not already been upstreamed (i.e. not\nincluded in the main development branch of the same repository), and\nwhich still needed to be.  (The terminology for the report was that\nunique patches that weren't upstream were \"snowflakes\" which we were\ntrying to pick out of a \"blizzard of commits\").  Anyway, the immediate\nimpetus for this report was that after enough disasters from\nforgetting to upstream important patches, people started asking around\nabout how to avoid another repeat.  I thought this was a trivial\nquestion to answer at first, but...I was wrong.\n\nNow, as I said before, this product was tracked in Gerrit.  Also,\ndirect pushing to bypass reviews was disabled (with _very_ rare\nexceptions, that essentially didn't affect the quality of the report\nmentioned below at all), and Change-Ids were required for all commits\npushed up for code reviews.  So, yes, people had the Gerrit-suggested\nhook installed, and yes virtually all commits had Change-Ids.\n\nSo, what was implemented to answer \"which patches in this branch have\nbeen upstreamed (i.e. included in the main development branch)\"?  A\nvariety of checks.  One of which was fully reliable:\n\n    * \"git cherry\" to catch the \"100% certainty cherry-picks\" (only\ncaught maybe 5% of the upstreamed patches, but still useful).\n\nAnd a bunch of other checks that were just heuristics:\n\n    * (author name, author email, author date) triples matching\n    * commit message exactly matching\n    * \"(cherry-picked from commit <HASH>)\" footers from one commit\nreferring to another (or there being a transitive chain that wasn't\ntoo long, or there being a transitive tree with a path between the\ncommits -- think for example of both commits being a cherry-pick of\nthe same thing)\n    * patch-hunk-ids matching (instead of git-patch-id which computes\na patch-id for the overall patch, compute one for each hunk.  Then\nlook for other commits that have one of their patch hunks match. This\ncould result in a many-to-many relationship between downstream and\nupstream patches)\n\nWe generated a report based on this (on a wiki so folks could edit and\nadd notes), with html links and such.  For each downstream commit, we\nadded links to all potential upstream commit(s) and included reasons\nwhy each commit was thought to be a potential match.  Further, for any\npotential upstream commit that might match (and which wasn't a 100%\ncertainty pick from \"git cherry\"), we also looked at which filename(s)\nwere modified in both commits and reported on the number of matching\nand non-matching filename(s) between the pair in addition to the\nnumber of matching and non-matching patch hunks.\n\nBasically, it was a large amount of information to allow humans to\nreview how similar the commits were to help them determine if the\nchanges from the commit were already (partially or fully) included\nupstream.\n\nI got lots of requests to find as many links as possible, because with\nunfortunately frequent regularity, a new report would be created for\nsome branch and dozens of teams of developers were being forced to\nreview the reports and sign off on every single patch and whether it\nneeded to be upstreamed, or even finish being upstreamed (and to do\nthe upstreaming work if so).  I got a fair number of comments and\nquestions on the reports as a result.\n\nI'm glad we no longer need this report due to switching to a much more\nsane testing/backporting/delivery/branching/etc. story, but years ago\nthis report was helpful.  Anyway...\n\nA few surprises I found:\n  * Patches could be partially upstreamed.  For example, someone\ncherry-picked something from a newer version of upstream than their\ncurrent branch was based off of, then found and amended important\nfixes into that commit.  So although the patch might appear to be\nupstream (because the commit messages matched, the author\nname/email/dates match, etc.), important fixes in it were not.\n  * I was surprised by the number of cases that were not one-to-one\nmappings.  One patch might be upstreamed via several patches on the\nmain branch.  Or parts of several patches may have been upstreamed as\none big combination commit upstream.  I knew theoretically there could\nbe some like this, and perhaps wasn't too surprised that there would\nbe at least one, but there were quite a few more than I expected.\n  * One of the weirder cases I remember: Someone had created changes\nby going into the Gerrit UI, picking an existing but obsolete/closed\ncode review, editing the code in ways that had absolutely nothing to\ndo with the original commit they started on (and perhaps on some other\nbranch?), and left the Change-Id around because it looked valid and\nthey didn't know what it meant anyway.  (Something about their laptop\nbeing broken and being unable to edit code locally, so they just\nedited \"in the cloud\" and let the automated regression tests verify\ntheir changes; Gerrit either didn't have a way to start a new commit\nat the time or the developer didn't find it, so they just edited\nsomething else.)  I knew about it because I found weird pairs of\nnon-matching things and this developer left a note about it in their\nedited commit message just in case anything weird happened.\n\nBut let me be more explicit about Change-Ids: we didn't use them at\nall.  It came up multiple times as a question, but in my looking into\nuseful factors, to me they seemed to provide no extra value and I had\nfound multiple cases where they seemed to be misleading.  The purpose\nof the report was to avoid more disasters from forgetting to backport\ncommits.  This report seems like the kind of thing that Change-Ids\nwere invented for, and we already had Change-Ids in all commits due to\nusing Gerrit, but from my exploration I didn't trust the Change-Ids to\nprovide net-positive value, so I simply didn't use them at all.\n\n</long story>\n"},{"id":"459964","messageId":"20220726084324.GO17705@kitsune.suse.cz","threadId":"58181","inReplyTo":"CABPp-BHWueqbFMpz4jMPK9G3xY5=jvWLELLQribU69CYD5DjKw@mail.gmail.com","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Michal Suchánek","fromEmail":"msuchanek@suse.de","sentAt":"2022-07-26T08:43:24Z","receivedAt":"2022-07-26T08:43:32Z","isPatch":false,"sender":{"key":"msuchanek@suse.de","avatar":"https://avatars.githubusercontent.com/u/787652?v=4"},"body":"Hello,\n\nOn Mon, Jul 25, 2022 at 08:49:30PM -0700, Elijah Newren wrote:\n> On Mon, Jul 25, 2022 at 2:47 PM Jacob Keller <jacob.keller@gmail.com> wrote:\n> >\n> > On Sat, Jul 23, 2022 at 10:23 PM Elijah Newren <newren@gmail.com> wrote:\n> > >\n> > > On Sat, Jul 23, 2022 at 12:44 AM Michal Suchánek <msuchanek@suse.de> wrote:\n> > > >\n> > > > On Fri, Jul 22, 2022 at 03:46:22PM -0700, Jacob Keller wrote:\n> > > > > On Fri, Jul 22, 2022 at 1:42 PM Michal Suchánek <msuchanek@suse.de> wrote:\n> > > > > >\n> > > > > > On Fri, Jul 22, 2022 at 09:08:56PM +0100, Philip Oakley wrote:\n> > > [...]\n> > > > > > > Part of the rename problem is that there can be many different routes to\n> > > > > > > the same result, and often the route used isn't the one 'specified' by\n> > > > > > > those who wish a complicated rename process to have happened 'their\n> > > > > > > way', plus people forget to record what they actually did. Attempting to\n> > > > > > > capture what happened still results major gaps in the record.\n> > > > > >\n> > > > > > Doesn't git have rebase?\n> > > > > >\n> > > > > > It is not required that the rename is captured perfectly every time so\n> > > > > > long as it can be amended later.\n> > > > > >\n> > > > >\n> > > > > Rebase is typically reserved only to modify commits which are not yet\n> > > > > \"permanent\". Once a commit starts being referenced by many others it\n> > > > > becomes more and more difficult to rebase it. Any rebase effectively\n> > > > > creates a new commit.\n> > > > >\n> > > > > There are multiple threads discussing renames and handling them in git\n> > > > > in the past which are worth re-reading, including at least\n> > > > >\n> > > > > https://public-inbox.org/git/Pine.LNX.4.58.0504141102430.7211@ppc970.osdl.org/\n> > > > >\n> > > > > A fuller analysis here too:\n> > > > > https://public-inbox.org/git/Pine.LNX.4.64.0510221251330.10477@g5.osdl.org/\n> > > > >\n> > > > > As mentioned above in this thread, depending on what context you are\n> > > > > using, a change to a commit could be many to many: i.e. a commit which\n> > > > > splits into 2, or 3 commits merging into one, or 3 commits splitting\n> > > > > apart and then becoming 2 commits. When that happens, what \"change id\"\n> > > > > do you use for each commit?\n> > > >\n> > > > Same as commit message and any trailers you might have - they are\n> > > > preserved, concatenated\n> > >\n> > > Exactly how are they concatenated?  Is that a user operation, or\n> > > something a Git command does automatically?  Which commands and which\n> > > circumstances?  If users do it, what's the UI for them to discover\n> > > what the fields are, for them to discover whether such a thing might\n> > > be needed or beneficial, and the UI for them to change these fields?\n> > > This sounds like a massive UX/UI issue that I don't have a clue how to\n> > > tackle (assuming I wanted to).\n> > >\n> > > > and can be regenerated.\n> > >\n> > > \"can be\".  But generally won't be even when it should be, right?\n> > >\n> > > Committer name/email/date basically don't even exist as far as many\n> > > Git users are concerned.  They aren't shown in the default log output\n> > > (which greatly saddens me), and even after attempting to educate users\n> > > for well over a decade now, I still routinely find developers who are\n> > > surprised that these things exist.\n> > >\n> > > Given that committer name/email/date aren't shown with --pretty=full\n> > > but with the lame option name --pretty=fuller, I can't see why it'd\n> > > make any sense to show Change-Ids in the log output by default.\n> > >\n> > > But if it's not shown -- and by default -- then it doesn't exist for\n> > > many users.  And if it doesn't exist, users aren't going to fix it\n> > > when they need to.\n> > >\n> > > (Even if it were shown by default, it's not clear to me that users\n> > > would know when to fix it, or how to fix it, or even care to fix it\n> > > and instead view it as a pedantic requirement being foisted on them.)\n> > >\n> > > I think the \"many-to-many issue\" others have raised in this thread is\n> > > an important, big, and thorny problem.  I think it has the potential\n> > > to be a minefield of UX and a steady stream of bug reports.  And\n> > > seeing proponents of Change-Id just dismissing the issue makes me all\n> > > the more suspicious of the proposal in the first place.\n> >\n> > I do think there is some value in having a sort of generic id like\n> > change-id, but I do think we want to be careful about how exactly we\n> > handle it.\n> >\n> > As you say, if we hide it then users may not be aware of it, and if we\n> > make it visible users who don't care may be annoyed. I don't think we\n> > can fully automate it because of the nature of combining changes and\n> > splitting changes require humans to decide which change keeps which\n> > ID. Its not even clear when rebasing whether a split is going to\n> > happen. A combine operation is easier to detect in rebase\n> > (fixup/squash), but determining which id to keep is not. Would we even\n> > want to have support for \"this commit merges two and is now one, but\n> > we keep both IDs because it really is both commits\"? That gets messy\n> > pretty fast.\n> >\n> > Users such as gerrit already simply use the trailer with Change-id and\n> > manage to make it work by enforcing some constraints and assuming\n> > users will know what to do (because otherwise they fail to interact\n> > with gerrit servers).\n> >\n> > For cases where it helps, I think its very valuable. Being able to\n> > track revisions of a series or a patch is super useful. Getting\n> > external tooling like public-inbox, patchworks, etc to use this would\n> > also be useful. But I think we would want to sort out the situation a\n> > bit for how and when are they generated, when are they\n> > replaced/re-generated, how this interacts with mailing etc.\n> >\n> > Should rebase just always regenerate? that loses a lot of value. I\n> > guess squashing could offer users a choice of which to keep? Fixup\n> > would always keep the same one. And otherwise it becomes up to users\n> > to know when they need to copy from an old commit or refresh an\n> > existing commit... Thats pretty much what gerrit does these days, if a\n> > commit doesn't have the trailer it gets added, and if it does, its up\n> > to the user to know when to remove it or regenerate it... Since its a\n> > commit message trailer it gets sent implicitly through the mailing\n> > list unless removed.\n> \n> Yes, I fully agree it needs to be spelled out a lot more.  And not\n> just obvious commands (everyone seems to focus on commit, cherry-pick,\n> and rebase), but what about e.g. `git merge --squash`?\n> \n> Also, as far as value goes, I have an interesting story related to\n> Change-Ids (read the last sentence if you only want the summary):\n> \n> \n> <long story>\n> \n> I have used Gerrit fairly heavily.  I maintained an instance for a few\n> hundred developers for several years (inheriting it from others), and\n> was responsible for various build & release stuff related to one of\n> the larger products tracked in it (an approximately\n> linux-kernel-sized) product.  I also attended a Gerrit conference or\n> two and submitted a few patches for Gerrit that were accepted.  So,\n> for context, I'm clearly not a Gerrit developer since I only submitted\n> a few patches, but I was the clear expert within my company on Gerrit.\n> So that's my background.\n> \n> Some background on the (insane) project management of the time\n> (unrelated to Git or Gerrit or Change-Ids) is also important to\n> understand this story:  Years ago, this project had well over 100\n> active branches (!!).  And yes, branches were being aggressively\n> retired, but I still remember when we finally managed to get the count\n> under 200.  Each branch had important patches, and hundreds of patches\n> on these branches was not uncommon.  Yes, it was insane, and yes, I\n> and many others were really happy when we eventually reached the land\n> of sanity with just a few active branches (and all but the main one\n> only gets backport fixes).\n> \n> One of the things I did to help us move in the direction towards\n> sanity was a \"snowflake report\" -- a report that would help people\n> determine which patches had not already been upstreamed (i.e. not\n> included in the main development branch of the same repository), and\n> which still needed to be.  (The terminology for the report was that\n> unique patches that weren't upstream were \"snowflakes\" which we were\n> trying to pick out of a \"blizzard of commits\").  Anyway, the immediate\n> impetus for this report was that after enough disasters from\n> forgetting to upstream important patches, people started asking around\n> about how to avoid another repeat.  I thought this was a trivial\n> question to answer at first, but...I was wrong.\n> \n> Now, as I said before, this product was tracked in Gerrit.  Also,\n> direct pushing to bypass reviews was disabled (with _very_ rare\n> exceptions, that essentially didn't affect the quality of the report\n> mentioned below at all), and Change-Ids were required for all commits\n> pushed up for code reviews.  So, yes, people had the Gerrit-suggested\n> hook installed, and yes virtually all commits had Change-Ids.\n> \n> So, what was implemented to answer \"which patches in this branch have\n> been upstreamed (i.e. included in the main development branch)\"?  A\n> variety of checks.  One of which was fully reliable:\n> \n>     * \"git cherry\" to catch the \"100% certainty cherry-picks\" (only\n> caught maybe 5% of the upstreamed patches, but still useful).\n> \n> And a bunch of other checks that were just heuristics:\n> \n>     * (author name, author email, author date) triples matching\n>     * commit message exactly matching\n>     * \"(cherry-picked from commit <HASH>)\" footers from one commit\n> referring to another (or there being a transitive chain that wasn't\n> too long, or there being a transitive tree with a path between the\n> commits -- think for example of both commits being a cherry-pick of\n> the same thing)\n>     * patch-hunk-ids matching (instead of git-patch-id which computes\n> a patch-id for the overall patch, compute one for each hunk.  Then\n> look for other commits that have one of their patch hunks match. This\n> could result in a many-to-many relationship between downstream and\n> upstream patches)\n> \n> We generated a report based on this (on a wiki so folks could edit and\n> add notes), with html links and such.  For each downstream commit, we\n> added links to all potential upstream commit(s) and included reasons\n> why each commit was thought to be a potential match.  Further, for any\n> potential upstream commit that might match (and which wasn't a 100%\n> certainty pick from \"git cherry\"), we also looked at which filename(s)\n> were modified in both commits and reported on the number of matching\n> and non-matching filename(s) between the pair in addition to the\n> number of matching and non-matching patch hunks.\n> \n> Basically, it was a large amount of information to allow humans to\n> review how similar the commits were to help them determine if the\n> changes from the commit were already (partially or fully) included\n> upstream.\n> \n> I got lots of requests to find as many links as possible, because with\n> unfortunately frequent regularity, a new report would be created for\n> some branch and dozens of teams of developers were being forced to\n> review the reports and sign off on every single patch and whether it\n> needed to be upstreamed, or even finish being upstreamed (and to do\n> the upstreaming work if so).  I got a fair number of comments and\n> questions on the reports as a result.\n> \n> I'm glad we no longer need this report due to switching to a much more\n> sane testing/backporting/delivery/branching/etc. story, but years ago\n> this report was helpful.  Anyway...\n> \n> A few surprises I found:\n>   * Patches could be partially upstreamed.  For example, someone\n> cherry-picked something from a newer version of upstream than their\n> current branch was based off of, then found and amended important\n> fixes into that commit.  So although the patch might appear to be\n> upstream (because the commit messages matched, the author\n> name/email/dates match, etc.), important fixes in it were not.\n>   * I was surprised by the number of cases that were not one-to-one\n> mappings.  One patch might be upstreamed via several patches on the\n> main branch.  Or parts of several patches may have been upstreamed as\n> one big combination commit upstream.  I knew theoretically there could\n> be some like this, and perhaps wasn't too surprised that there would\n> be at least one, but there were quite a few more than I expected.\n>   * One of the weirder cases I remember: Someone had created changes\n> by going into the Gerrit UI, picking an existing but obsolete/closed\n> code review, editing the code in ways that had absolutely nothing to\n> do with the original commit they started on (and perhaps on some other\n> branch?), and left the Change-Id around because it looked valid and\n> they didn't know what it meant anyway.  (Something about their laptop\n> being broken and being unable to edit code locally, so they just\n> edited \"in the cloud\" and let the automated regression tests verify\n> their changes; Gerrit either didn't have a way to start a new commit\n> at the time or the developer didn't find it, so they just edited\n> something else.)  I knew about it because I found weird pairs of\n> non-matching things and this developer left a note about it in their\n> edited commit message just in case anything weird happened.\n> \n> But let me be more explicit about Change-Ids: we didn't use them at\n> all.  It came up multiple times as a question, but in my looking into\n> useful factors, to me they seemed to provide no extra value and I had\n> found multiple cases where they seemed to be misleading.  The purpose\n> of the report was to avoid more disasters from forgetting to backport\n> commits.  This report seems like the kind of thing that Change-Ids\n> were invented for, and we already had Change-Ids in all commits due to\n> using Gerrit, but from my exploration I didn't trust the Change-Ids to\n> provide net-positive value, so I simply didn't use them at all.\n> \n> </long story>\n\nif you are into long stories I have enother one.\n\n<long story>\nI am maintaining a code review system that I inherited that comes from\ntime when code review systems weren't cool, or even a thing at all. It's\nbasically a bunch of perls scripts that run as git hooks, cron jobs, and\nCGI pages.\n\nIt is used to maintain some backports and original development on top of\nupstream kernel releases.\n\nThe design is in some ways interesting. One particular interesting\ndecision which was made some time in ancient past is that the changes\nare maintained as quilt patch series rather than modified sources.\n\nNot sure what was the rationale behind the decision. Surely you could\nsave a lot of space if you had few changes. Also pre-git you would get\nthe release tarball, the point release patches, and any local changes as\npatches on top so it makes sense to track the files you have together\nwith the series file that says how to use them, and you cannot get\neasily away from tracking patches because that's what you get from\nupstream.\n\nThis leads to an interesting pattern - you are not tracking changes to\nsource code but changes to history of changes to source code. At any\ntime you can render the quilt series as a git history that starts with\nthe upstream release and adds commits on top, in order specified in the\nseries file. You can add and drop commits in the middle, update tags,\nwhatever. And it can all be nice linear development in the repository\nthat tracks the quilt series.\n\nInitially you would get a lot of big patches similar to the upstream\npoint releasee patches - like a patch that diffs the changes in a\nparticular driver between two kernel versions, with some additional\nfixes on top. Clearly an unamanageable, untrackable, many-to-many patch\nrelationship to the upstream development.\n\nHowever, over time this became a problem with more and more patches\npiling up, base kernel version updates causing regressions, etc, etc\n\nSo over time some rules emerged. When backporting something from\nupstream add each upstream git commit as separate patch file. Tag it\nwith the upstream git commit ID (change ID, yay). Add additional fixes\nas separate patch files, and send them upstream. etc, etc.\n\nIn the end 1:1 relationship emerged with some rare exceptions.\n\nSure, there are upstream subsystems that routinely cherry-pick patches\n\nSure, there are upstream merge commits that introduce random unrelated\nchanges.\n\nNonetheless, the saner the upstream subsystem maintenanace and the the\nsaner the downstream branch maintenance the closer you get to 1:1 patch\nrelationship.\n</long story>\n\nIf I can derive anything from these stories it is that while arbitrary\nmany to many relationsips between patches are possible they are not at\nall desirable. Then it is not a bug for change tracking to not support\nsuch relationship, you can even see it as a safety against people\nshooting themselves in the foot.\n\nIf you are developing a feature over many revisions before it gets\nmerged upstream it can happen that you squash some patches that were\ninitially separate, and later split them differently. However, it does\nnot happen as one operation. So the many to many problem is not\nsomething worth solving, at least for sane workflows.\n\nThanks\n\nMichal\n"},{"id":"460204","messageId":"bf871a430177ced6d628641eac9d478389fb6c2b.camel@that.guru","threadId":"58181","inReplyTo":"220719.86y1wpuy5o.gmgdl@evledraar.gmail.com","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Stephen Finucane","fromEmail":"stephen@that.guru","sentAt":"2022-07-29T12:11:24Z","receivedAt":"2022-07-29T12:11:33Z","isPatch":false,"sender":{"key":"stephen@that.guru","avatar":"https://gravatar.com/avatar/2f9716975d743b6d7764972ed20fecdad956a5d4e3e6cfe8282d10151f1b2df0?d=mp&s=160"},"body":"On Tue, 2022-07-19 at 13:09 +0200, Ævar Arnfjörð Bjarmason wrote:\n> On Tue, Jul 19 2022, Stephen Finucane wrote:\n> \n> > On Mon, 2022-07-18 at 20:50 +0200, Ævar Arnfjörð Bjarmason wrote:\n> > > On Mon, Jul 18 2022, Stephen Finucane wrote:\n> > > \n> > > > ...to track evolution of a patch through time.\n> > > > \n> > > > tl;dr: How hard would it be to retrofit an 'ChangeID' concept à la the 'Change-\n> > > > ID' trailer used by Gerrit into git core?\n> > > > \n> > > > Firstly, apologies in advance if this is the wrong forum to post a feature\n> > > > request. I help maintain the Patchwork project [1], which a web-based tool that\n> > > > provides a mechanism to track the state of patches submitted to a mailing list\n> > > > and make sure stuff doesn't slip through the crack. One of our long-term goals\n> > > > has been to track the evolution of an individual patch through multiple\n> > > > revisions. This is surprisingly hard goal because oftentimes there isn't a whole\n> > > > lot to work with. One can try to guess whether things are the same by inspecting\n> > > > the metadata of the commit (subject, author, commit message, and the diff\n> > > > itself) but each of these metadata items are subject to arbitrary changes and\n> > > > are therefore fallible.\n> > > > \n> > > > One of the mechanisms I've seen used to address this is the 'Change-ID' trailer\n> > > > used by Gerrit. For anyone that hasn't seen this, the Gerrit server provides a\n> > > > git commit hook that you can install locally. When installed, this appends a\n> > > > 'Change-ID' trailer to each and every commit message. In this way, the evolution\n> > > > of a patch (or a \"change\", in Gerrit parlance) can be tracked through time since\n> > > > the Change ID provides an authoritative answer to the question \"is this still\n> > > > the same patch\". Unfortunately, there are still some obvious downside to this\n> > > > approach. Not only does this additional trailer clutter your commit messages but\n> > > > it's also something the user must install themselves. While Gerrit can insist\n> > > > that this is installed before pushing a change, this isn't an option for any of\n> > > > the common forges nor is it something git-send-email supports.\n> > > \n> > > git format-patch+send-email will send your trailers along as-is, how\n> > > doesn't it support Change-Id. Does it need some support that any other\n> > > made-up trailer doesn't?\n> > \n> > It supports sending the trailers, sure. What it doesn't support is insisting you\n> > send this specific trailer (Change-Id). Only Gerrit can do this (server side,\n> > thankfully, which means you don't need to ask all contributors to install this\n> > hook if you want to rely on it for tooling, CI, etc.).\n> \n> Ah, it's still unclear to me what you're proposing here though. That\n> send-email always (generates?) or otherwise insists on the trailer, that\n> it can be configured ot add it?\n>\n> That send-email have some \"pre-send-email\" hook? Something else?\n \n(Apologies for the delayed response: I was on holiday).\n\nI'm afraid I don't have the correct terminology to describe what I'm suggesting\nso I'll show an example instead.\n\nI have configured the 'fuller' pretty formatter locally:\n\n   $ git config format.pretty\n   fuller\n\nWhen I do git log on e.g. the openstack nova repo, I see:\n\n   commit 2709e30956b53be1dca91eec801220f0efbaed93\n   Author:     Stephen Finucane <sfinucan@redhat.com>\n   AuthorDate: Thu Jul 14 15:43:40 2022 +0100\n   Commit:     Stephen Finucane <sfinucan@redhat.com>\n   CommitDate: Mon Jul 18 12:30:25 2022 +0100\n   \n       Fix compatibility with jsonschema 4.x\n       \n       This changed one of the error messages we depend on [1].\n       \n       [1] https://github.com/python-jsonschema/jsonschema/commit/641e9b8c\n       \n       Change-Id: I643ec568ee2eb2ec1a555f813fd2f1acff915afa\n       Signed-off-by: Stephen Finucane <sfinucan@redhat.com>\n\n(Side note: What *is the term for the \"Author\", \"AuthorDate\", \"Commit\" and\n\"CommitDate\" fields? Commit header? Commit metadata? Something else?)\n\nMy thinking is there are two types of information here: information that relates\nto the \"commiting\" of this change and information that relates to the\n\"authorship\" of the this change. The commit ID, 'Commit' and 'CommitDate' fields\nclearly form the commit parts. I'm arguing that it would be good to have an\nequivalent to the commit ID field for the authorship-type metadata.\n   \n   commit 2709e30956b53be1dca91eec801220f0efbaed93\n   Author:     Stephen Finucane <sfinucan@redhat.com>\n   AuthorDate: Thu Jul 14 15:43:40 2022 +0100\n   AuthorID:   I643ec568ee2eb2ec1a555f813fd2f1acff915afa\n   Commit:     Stephen Finucane <sfinucan@redhat.com>\n   CommitDate: Mon Jul 18 12:30:25 2022 +0100\n   \n       Fix compatibility with jsonschema 4.x\n       \n       This changed one of the error messages we depend on [1].\n       \n       [1] https://github.com/python-jsonschema/jsonschema/commit/641e9b8c\n       \n       Signed-off-by: Stephen Finucane <sfinucan@redhat.com>\n\nAt risk of repeating myself, I think this information would be valuable to allow\nme to answer the question \"is this the same[*] commit?\". During code review,\nthis would allow me to track the evolution of an individual patch. Once a patch\nis merged, it would allow me to track the backporting or cherry-picking of that\npatch between branches (in a more reliable fashion than the \"cherry picked from\"\ntrailer that one can add with the '-x' flag).\n\nNow I do realize that there will be issues with this. As has been noted\nelsewhere in the thread, people do split patches up or merge them together, and\na patch can change so drastically during review that it doesn't resemble the\noriginal patch in any way. However, I'd argue that in both cases the presence of\nthese persistent IDs would at least leave a breadcrumb trail for either tooling\nor humans to follow. Similarly, it is possible for users to mess things up by\nresetting or reusing the persistent ID fields, but as has been noted elsewhere\nin this thread this is already an issue with the existing Author* fields (which\nmany users likely don't know about) yet I couldn't imagine anyone wanting to get\nrid of these. It's an education thing.\n\n> I'd think for projects that care about this they're likely to have a\n> centralized enough workflow that it can be checked on the remote side,\n> whether that's some sanity check on the applier's \"git am\" pipeline, or\n> a \"pre-receive\" hook.\n\nYeah, as above I'm hoping this would form part of the core metadata of a commit\nrather than a trailer or something. Tools like Gerrit could of course do\nvalidation on this but that's outside the scope of what I'm looking at.\n\n> > > > I imagine most people working with mailing list based workflows have their own\n> > > > client side tooling to support this while software forges like GitHub and GitLab\n> > > > simply don't bother tracking version history between individual commits in a\n> > > > pull/merge request.\n> > > \n> > > It's far from ideal, but at least GitLab shows a diff on a push to a MR,\n> > > including if it's force-pushed. I'm not sure about GitHub.\n> > \n> > GitHub does not. Simply piling multiple additional \"fix\" commits onto the PR\n> > branch results in a less horrible review experience since you can maintain\n> > context, alas at the cost of a rotten git log. We don't need to debate the pros\n> > and cons of the various forges though :)\n> \n> Yes, I'm only mentioning it because it's worth looking at existing\n> \"solutions\" that are in use in the wild, however flawed those may be.\n> \n> > > > IMO though, it would be fantastic if third party tools\n> > > > weren't necessary though. What I suspect we want is a persistent ID (or rather\n> > > > UUID) that never changes regardless of how many times a patch is cherry-picked,\n> > > > rebased, or otherwise modified, similar to the Author and AuthorDate fields.\n> > > > Like Author and AuthorDate, it would be part of the core git commit metadata\n> > > > rather than something in the commit message like Signed-Off-By or Change-ID.\n> > > > \n> > > > Has such an idea ever been explored? Is it even possible? Would it be broadly\n> > > > useful?\n> > > \n> > > This has come up a bunch of times. I think that the thing git itself\n> > > should be doing is to lean into the same notion that we use for tracking\n> > > renames. I.e. we don't, we analyze history after-the-fact and spot the\n> > > renames for you.\n> > \n> > Any idea where I'd find previous discussions on this? I did look, and the only\n> > proposal I found was an old one that seemed to suggest including the Change-Id\n> > commit-msg hook with git itself which is not what I'm suggesting here.\n> \n> At the time I was punting on finding the links, and just working off\n> vague recollection, and hoping you'd go list spelunking.\n> \n> But I since recalled some details, I think the most relevant thing is\n> this discussion about a \"git evolve\":\n> \n>     https://lore.kernel.org/git/CAPL8ZivFmHqS2y+WmNR6faRMnuahiqwPVYsV99NiJ1QLHOs9fQ@mail.gmail.com/\n> \n> Which I think you'll find useful, especially as mercurial has an\n> existing implementation. The wider context for that \"git evolve\" is (I\n> believe) people at Google who maintain Gerrit trying to \"upstream\" the\n> Change-Id.\n> \n> Now, it hasn't landed in git.git, and it's been a few years, but going\n> through the details of why it fizzled out will be useful to you, if\n> you're interested in driving something like this forward.\n\nYeah to be clear I'm not suggesting tracking anything like this in Git core. My\nmain request is here is a persistent Author ID field. Commits as they are would\nremain the same: we'd just be able to show the evolution of a \"change\" in\nexternal tooling without the need for separate trailers.\n\n> There's also these two proposals from Eric Raymond:\n> \n> \thttps://lore.kernel.org/git/20190515191605.21D394703049@snark.thyrsus.com/\n\nThis however, looks more similar to what I'm proposing. If understand this\ncorrectly (I'm still reading the full thread), Eric is proposing allowing two\nways to reference a commit: the hash and a sort of alias. There would still be a\n1:1 mapping though, which is explicitly not what I want. I'm also not suggesting\ngenerating this stuff server-side. It should be part of the commit when\ninitially created, just like Author and AuthorDate.\n\n> \thttps://lore.kernel.org/git/20190521013250.3506B470485F@snark.thyrsus.com/\n> \n> Which I'm linking to here not because I think they're viable, as you can\n> see from my participation in those threads I think what he suggested is\n> an architectural dead end as far as git is concerned.\n> \n> But rather because it's conceptually adjacent (you could in principle\n> use nanosecond timestamps as a poor man's UUID), and much of the\n> follow-up discussion is about format changes in general, and if/when\n> those might be viable.\n> \n> > > We have some of that in git already, as git-patch-id, and more recently\n> > > git-range-diff. Both are flawed in a bunch of ways, and it's easy to run\n> > > into edge cases where they don't spot something that they \"should\"\n> > > have. Where \"should\" exists in the mind of the user.\n> > \n> > That's a fair point and is of course what we (Patchwork) have to do currently.\n> > Patchwork can track relations between individual patches but doesn't attempt to\n> > generate these relations itself. Instead, we rely on third-party tooling. The\n> > PaStA tool was one such example of a tool that could do this [1]. I can't\n> > imagine a tool like Gerrit would ever work without this concept of an\n> > authoritative (and arbitrary) identifier to track a patch's identity through\n> > time, hence its reliance on the Change-Id trailer.\n> \n> I haven't used Gerrit or Patchwork, so much of this is from ignorance on\n> that front, but I have spent a lot of time thinking about this in the\n> context of git in general.\n> \n> I think as users of git go the git project itself makes very heavy use\n> of this, i.e. sequences of patches are substantially rewritten, split,\n> squashed etc. all the time, or even split into two or more sets of\n> submissions.\n> \n> Having said all that I can't see how a Change-Id isn't a Bad Idea(TM)\n> for all the same reasons that pre-git SCMs file formats that track\n> renames explicitly were a bad idea.\n> \n> I.e. yes you can come up with cases where that's \"better\" than what git\n> does, but they didn't handle splitting/merging files etc.\n> \n> Similarly what happens when you have 3 patches each with their own\n> Change-Id and you split them into 4 patches. Is the Change-Id 1=1 or\n> 1=many. I'm suggesting that you'd want a solution that can be many=many.\n> \n> And also, that those many=many should be dynamically configurable and\n> inferred after the fact. E.g. range-diff will commits that are similar\n> enough that two authors with no knowledge of each other independently\n> came up with.\n\nI touched on the splitting/merging of changes above but just to reiterate, I\ndon't think this is an issue. I'm using Gerrit for OpenStack-related efforts nad\nmailing lists (with Patchwork tracking submissions) elsewhere. Patches are\nfrequently split and merged as part of a review process and can often be merged\nas part of a backport (I've yet to see a patch split up when backporting but it\ncould happen too). If a patch is split, the original patch retains the 'Change-\nID' as well as 'Author' and 'AuthorDate' fields while the split out patch(es)\nget new versions of these. If one or more patches are squashed, you get the\n'Change-ID' and 'Author'/'AuthorDate' of the first patch in the series of\nsquashed patches. In both cases though, some Change-ID persists which means you\ncan track the evolution of a patch or series through time. These are all\nextremely helpful breadcrumbs for reviewers.\n\nRegarding the rename issue, I agree that this isn't something Git should do\neither. As you note, it's too hard to do 100% reliably, which would be expected\ngoal. I'm not looking for 100% reliability here. I just want a better breadcrumb\nthan e.g. range-diff currently provides.\n\n> I think that range-diff is still lacking in a lot of ways, in particular:\n> \n>  * It matches entire commits (log + diff) on a similarity score, I've\n>    often wanted a way to \"weigh\" it, so e.g. a matching hunk would have\n>    3x the matching score of a matching commit message.\n> \n>    Now it often \"gives up\", you can give it a higher --creation-factor,\n>    but that's \"global\", so for a large range you'll often start\n>    including irrelevant things as well.\n> \n>  * It only does 1=1 attribution, and e.g. currently can't find/represent\n>    a case where a commit with 3 hunks got split into two commits, with 2\n>    and 1 hunks, respectively. It'll (usually) show a diff to the new 2\n>    hunk commit, but the \"new\" 1 hunk will be shown as new.\n> \n>    We could continue to drill down and find such \"unattributed\" hunks.\n> \n> > Perhaps we could flip this on its head. What would be the _downsides_ of\n> > providing a persistent, arbitrary identifier on a commit similar to Author and\n> > AuthorDate fields? There's obviously some work involved in implementing it but\n> > assuming that was already done, what would break/be worse as a result?\n> \n> That \"Repository formats matter\", to borrow a phrase from a classic post\n> about git[1]. Once you provide a way to do something it will be used,\n> and when that something has inherent limitations (think SCM rename\n> tracking) used to the exclusion of others.\n> \n> You can't provide something like that as an opt-in and \"upstream\" it\n> without it inevetably trickling into a lot of areas of Git's UX.\n> \n> To continue the rename example, now you can just re-arrange your source\n> tree and not worry about micro-managing it with \"git mv\" (in the \"svn\n> mv\" sense), git will figure it out after the fact.\n> \n> That's a sinificant UX benefit, we can provide a *much simpler* UX as a\n> result.\n> \n> What would be the harm of an optional \"rename tracking\" header? After\n> all the heuristic sometimes \"fails\".\n> \n> The harm would be that if you really wanted to lean into that (even\n> optionally) you'd be forced to add that to all sorts of tooling, not\n> just the cheap convenience that is \"git mv\" currently.\n> \n> Likewise everything from \"cherry-pick\" to \"rebase\" to \"commit\" would\n> inevitably have to learn some way to know about, carry forward and ask\n> the user about Change-Id's and their preservation. Don't you think so?\n\nThese are all valid points. Hopefully my points above regarding the similarity\nto the Author and AuthorDate fields helps though.\n\n> Otherwise they'd be much too easy to lose track of, and if they only\n> reason we did all that is because we didn't think enough about the \"work\n> it out after\" approach that would be a bad investment of time.\n> \n> But I may be wrong about all of that, I think one thing that would\n> really help clarify this & similar proposals is if people pushing it\n> forward came up with some basic tests for it, i.e. just something like\n> a:\n> \n>     series-v1/\n>     series-v2/\n> \n> Where those two directories would be the \"git format-patch\" output (or\n> whatever) of two versions of a series that Gerrit or Patchwork are now\n> managing, along with some (plain text?) manual mapping of which things\n> in v1 correspond to v2.\n> \n> We could then compare how that manual attribution performs v.s. trying\n> to find which things match (range-diff) afterwards.\n\nI hope my examples above helped with this, but I can prepare a sample series\n(including a sample 'git log' output) if you'd like. Just let me know where\nyou'd like it sent.\n\nCheers,\nStephen\n\n\n> \n> 1. https://keithp.com/blog/Repository_Formats_Matter/\n> \n\n"},{"id":"460209","messageId":"1f2c01d8a348$68701ba0$395052e0$@pdinc.us","threadId":"58181","inReplyTo":"bf871a430177ced6d628641eac9d478389fb6c2b.camel@that.guru","subject":"RE: Feature request: provide a persistent IDs on a commit","fromName":"Jason Pyeron","fromEmail":"jpyeron@pdinc.us","sentAt":"2022-07-29T12:40:38Z","receivedAt":"2022-07-29T12:49:57Z","isPatch":false,"sender":{"key":"jpyeron@pdinc.us","avatar":"https://gravatar.com/avatar/c2e53452caa53d940768a1ffc9cf76196d851b9b534b7a39cd39852a70a0508f?d=mp&s=160"},"body":"> From: Stephen Finucane\n> Sent: Friday, July 29, 2022 8:11 AM\n> \n> On Tue, 2022-07-19 at 13:09 +0200, Ævar Arnfjörð Bjarmason wrote:\n> > On Tue, Jul 19 2022, Stephen Finucane wrote:\n> >\n> > > On Mon, 2022-07-18 at 20:50 +0200, Ævar Arnfjörð Bjarmason wrote:\n> > > > On Mon, Jul 18 2022, Stephen Finucane wrote:\n> > > >\n> > > > > ...to track evolution of a patch through time.\n> > > > >\n> > > > > tl;dr: How hard would it be to retrofit an 'ChangeID' concept à la the 'Change-\n> > > > > ID' trailer used by Gerrit into git core?\n> > > > >\n> > > > > Firstly, apologies in advance if this is the wrong forum to post a feature\n> > > > > request. I help maintain the Patchwork project [1], which a web-based tool that\n> > > > > provides a mechanism to track the state of patches submitted to a mailing list\n> > > > > and make sure stuff doesn't slip through the crack. One of our long-term goals\n> > > > > has been to track the evolution of an individual patch through multiple\n> > > > > revisions. This is surprisingly hard goal because oftentimes there isn't a whole\n> > > > > lot to work with. One can try to guess whether things are the same by inspecting\n> > > > > the metadata of the commit (subject, author, commit message, and the diff\n> > > > > itself) but each of these metadata items are subject to arbitrary changes and\n> > > > > are therefore fallible.\n> > > > >\n> > > > > One of the mechanisms I've seen used to address this is the 'Change-ID' trailer\n> > > > > used by Gerrit. For anyone that hasn't seen this, the Gerrit server provides a\n> > > > > git commit hook that you can install locally. When installed, this appends a\n> > > > > 'Change-ID' trailer to each and every commit message. In this way, the evolution\n> > > > > of a patch (or a \"change\", in Gerrit parlance) can be tracked through time since\n> > > > > the Change ID provides an authoritative answer to the question \"is this still\n> > > > > the same patch\". Unfortunately, there are still some obvious downside to this\n> > > > > approach. Not only does this additional trailer clutter your commit messages but\n> > > > > it's also something the user must install themselves. While Gerrit can insist\n> > > > > that this is installed before pushing a change, this isn't an option for any of\n> > > > > the common forges nor is it something git-send-email supports.\n> > > >\n> > > > git format-patch+send-email will send your trailers along as-is, how\n> > > > doesn't it support Change-Id. Does it need some support that any other\n> > > > made-up trailer doesn't?\n> > >\n> > > It supports sending the trailers, sure. What it doesn't support is insisting you\n> > > send this specific trailer (Change-Id). Only Gerrit can do this (server side,\n> > > thankfully, which means you don't need to ask all contributors to install this\n> > > hook if you want to rely on it for tooling, CI, etc.).\n> >\n> > Ah, it's still unclear to me what you're proposing here though. That\n> > send-email always (generates?) or otherwise insists on the trailer, that\n> > it can be configured ot add it?\n> >\n> > That send-email have some \"pre-send-email\" hook? Something else?\n> \n> (Apologies for the delayed response: I was on holiday).\n> \n> I'm afraid I don't have the correct terminology to describe what I'm suggesting\n> so I'll show an example instead.\n> \n> I have configured the 'fuller' pretty formatter locally:\n> \n>    $ git config format.pretty\n>    fuller\n> \n> When I do git log on e.g. the openstack nova repo, I see:\n> \n>    commit 2709e30956b53be1dca91eec801220f0efbaed93\n>    Author:     Stephen Finucane <sfinucan@redhat.com>\n>    AuthorDate: Thu Jul 14 15:43:40 2022 +0100\n>    Commit:     Stephen Finucane <sfinucan@redhat.com>\n>    CommitDate: Mon Jul 18 12:30:25 2022 +0100\n> \n>        Fix compatibility with jsonschema 4.x\n> \n>        This changed one of the error messages we depend on [1].\n> \n>        [1] https://github.com/python-jsonschema/jsonschema/commit/641e9b8c\n> \n>        Change-Id: I643ec568ee2eb2ec1a555f813fd2f1acff915afa\n>        Signed-off-by: Stephen Finucane <sfinucan@redhat.com>\n> \n> (Side note: What *is the term for the \"Author\", \"AuthorDate\", \"Commit\" and\n> \"CommitDate\" fields? Commit header? Commit metadata? Something else?)\n> \n> My thinking is there are two types of information here: information that relates\n> to the \"commiting\" of this change and information that relates to the\n> \"authorship\" of the this change. The commit ID, 'Commit' and 'CommitDate' fields\n> clearly form the commit parts. I'm arguing that it would be good to have an\n> equivalent to the commit ID field for the authorship-type metadata.\n> \n>    commit 2709e30956b53be1dca91eec801220f0efbaed93\n>    Author:     Stephen Finucane <sfinucan@redhat.com>\n>    AuthorDate: Thu Jul 14 15:43:40 2022 +0100\n>    AuthorID:   I643ec568ee2eb2ec1a555f813fd2f1acff915afa\n>    Commit:     Stephen Finucane <sfinucan@redhat.com>\n>    CommitDate: Mon Jul 18 12:30:25 2022 +0100\n> \n>        Fix compatibility with jsonschema 4.x\n> \n>        This changed one of the error messages we depend on [1].\n> \n>        [1] https://github.com/python-jsonschema/jsonschema/commit/641e9b8c\n> \n>        Signed-off-by: Stephen Finucane <sfinucan@redhat.com>\n> \n> At risk of repeating myself, I think this information would be valuable to allow\n> me to answer the question \"is this the same[*] commit?\". During code review,\n> this would allow me to track the evolution of an individual patch. Once a patch\n> is merged, it would allow me to track the backporting or cherry-picking of that\n\nWe have been toying with this. We are looking at a field (behaves like parent) to track \"original commit\".\n\nThis value would be set on first rebase, amend, cherry-pick, etc.\n\nThe bonus for us will be when we patch gerrit to consume it and git log --graph --somenewoption to use it.\n\nIt would be nice if git core did add such value.\n\n-Jason\n\n"},{"id":"509134","messageId":"CANiSa6gwup5vXU235mG+Ybbc+P=SbwoNFEmuhg=iYu0yGvSXVA@mail.gmail.com","threadId":"58181","inReplyTo":"CANiSa6gou27CH-y8Yy7FX8xu2Kr66BM9ghvM9nDQkq0mkP0-QQ@mail.gmail.com","subject":"Re: Feature request: provide a persistent IDs on a commit","fromName":"Martin von Zweigbergk","fromEmail":"martinvonz@gmail.com","sentAt":"2024-12-15T17:09:08Z","receivedAt":"2024-12-15T17:09:20Z","isPatch":false,"sender":{"key":"martinvonz@gmail.com","avatar":"https://avatars.githubusercontent.com/u/891642?v=4"},"body":"I just noticed that my message from almost 2.5 years ago got blocked\nby the HTML filter, so here's a very late reply. I'd like to continue\nthis discussion, but maybe I'll start a separate thread later if you\nprefer.\n\nOn Wed, Jul 27, 2022 at 1:26 AM Martin von Zweigbergk\n<martinvonz@gmail.com> wrote:\n>\n>\n> On Wed, Jul 20, 2022, 21:47 Konstantin Ryabitsev <konstantin@linuxfoundation.org> wrote:\n>>\n>> On Mon, Jul 18, 2022 at 02:24:07PM -0700, Glen Choo wrote:\n>> > > I just started working on this for b4, with the notable difference that the\n>> > > change-id trailer is used in the cover letter instead of in individual\n>> > > commits, which moves the concept of \"change\" from a single commit to a series\n>> > > of commits. IMO, it's much more useful in that scope, because as series are\n>> > > reviewed and iterated, individual patches can get squashed, split up or\n>> > > otherwise transformed.\n>> >\n>> > My 2 cents, since I used to use Gerrit a lot :)\n>> >\n>> > I find persistent per-commit ids really useful, even when patches get\n>> > moved around. E.g. Gerrit can show and diff previous versions of the\n>> > patch, which makes it really easy to tell how the patch has evolved\n>> > over time.\n>>\n>> The kernel community has repeatedly rejected per-patch Change-id trailers\n>> because they carry no meaningful information outside of the gerrit system on\n>> which they were created.\n>\n>\n> Since I didn't see a mention of it yet in this thread, I thought I'd point of some possible ways that Git itself could use a change ID:\n>\n> * By allowing change IDs to be used for specifying a commit, you could do e.g. `git rebase main <change ID>; git checkout <change ID>` without requiring the user to look up the hash of the rewritten commit.\n> * Can be used for identifying rewritten commits that still have descendants (because you would have multiple reachable commits with the same change ID). That could be used for suggesting which descendants to rebase to where (a bit like Mercurial's `hg evolve`).\n> * Can be used by `git rebase` for skipping commits that are already in the destination.\n> * Can be used by `git cherry` for identifying commits that have been rebased. However, for `git cherry-pick`, we'd have to decide if we want to keep the change ID or generate a new one. We would not want to reuse the change ID if we consider multiple reachable commits with the same change ID a (soft) error.\n>\n>> Seeing a Change-Id trailer in a commit tells you\n>> nothing about the history of that commit unless you know the gerrit system on\n>> which this patch was reviewed (and have access to it, which is not a given).\n>> This is not as opaque as it used to be now that Gerrit provided ability to\n>> clone the underlying notedb, but this still fails on commits that were\n>> contributed to an upstream that doesn't use Gerrit.\n>\n>\n> Yes, that's a good point. There's the old \"git evolve\" proposal [1] that may be a good solution for the problem of tracking rewrites across multi-remote exchanges.\n>\n> [1] https://public-inbox.org/git/86d0spzoi1.fsf@gmail.com/T/\n"}]}