git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [RFC] origin link for cherry-pick and revert

From
Theodore Tso <tytso@mit.edu>
Date
Sep 10, 2008, 16:18 UTC
Message-ID
<20080910161852.GR21071@mit.edu>
In-Reply-To
<20080910141630.GB7397@cuci.nl>
On Wed, Sep 10, 2008 at 04:16:30PM +0200, Stephen R. van den Berg wrote:
Show 5 quoted lines
> The renumbering is not a problem, renumbering is a rare operation since
> a project's history is supposed to be stable.  And even if renumbering
> is performed, it is a well understood operation of which the renumbering
> of the origin links imposes a negligible overhead on top of the existing
> renumbering overhead.

Well *you* were the one using this as an argument for using the origin link. But I'll note that in some workflows, rebasing happens all the time when a patch is being developed and moved around. Sometimes patches are created in git, exported as a patch, and then it re-enters git again later (which is another reason why using an external UUID or bug tracking identifier is a good thing).

Show 11 quoted lines
> >Addresses-Bug: Red_Hat/149480, Sourceforge_Feature/120167
> >or
> >Addresses-Bug: Debian/432865, Launchpad/203323, Sourceforge_Bug/1926023
> 
> >Once you have this information, it is not difficult to maintain a
> >berk_db database which maps a particular Bug identifier (i.e.,
> >Red_Hat/149480, or Debian/471977, or Launchpad/203323) to a series of
> >commits.
> 
> This is nice, I admit, but it has the following downsides:
> - It is nontrivial to automate this on execution of "git cherry-pick".

It's trivial if it's in the free-form text. In fact, it happens automatically. If it's stored within the git commit object, then it will be done in the C code (if you've updated to the latest git; again, one of the advantages of doing it in free-form text).

Show 6 quoted lines
> - In a distributed environment this requires a network-reachable bug
>   database.
> - A network-reachable bug database means that suddenly git needs network
>   access for e.g. cherry-pick, revert, gitk, log --graph, blame.
> - Network queries for commits containing references kind of kills
>   performance.

No, because you don't need to look up the bug identifier unless you want to, you know, actually look at the bug. Otherwise, we are just using something like "debian/432865" as an identifier; you only need to look them up if you want to look up the bug. Any time you have a collaborative development environment, you will need either a centralized, network accessible bug tracking system, or use a distributed bug tracking system. Either way, though, if it's just matter of seeing whether or not a bug fix such as debian/432865 is fixed by some commit in some branch, using the bug identifier actually makes this *easier*, not harder.

> - Some backports don't have entries in a bug database because they
>   weren't bugs to begin with, in which case it becomes impossible to add
>   an identifier to the commit message after the fact.

This is true. The transition is a little easier if you are pointing to a pre-existing commit, whereas if you need some kind of rendevous identifer (whether it is a bug ID or some UUID). On the other hand, you've cherry-picked some bug fix using a git that didn't support the origin link, you'd also be screwed, so

> - It relies heavily on tools outside of git-core, which raises the
>   threshold for using it.

Well, it relies on changes to git --- just like the origin link requires changes to git. If the it is implemented using free-form text, which is a great way to prototype it, you have the *option* of implementing it via either git porcelain changes or outside tools like emacs or vi macros (just as most of us who are kernel developers have editor macros that insert Signed-off-by: into git commit messages, as well as changes in git porcelain such that "git am -s" automatically adds the Signed-off-by header). But given the wildly successful use of Signed-off-by in the kernel sources, this objection seems not very credible, to say the least.

Show 9 quoted lines
> The recommended practice here is quite simple:
> 
> - Origin links should only be created pointing to stable commits (i.e.
>   commits which you'd be willing to publish or already have published).
> 
> - This implies that pointing an origin link at a commit in a strain that
>   you still want to rebase is asking for trouble.  Doing this is akin to
>   doing a merge between two branches and then you start rebasing 4
>   commits *below* the mergepoint.  Don't do that.

Right. And if we use a UUID to identify commits, then we don't have to have these restrictions.

> - The only special case I'd allow is if you rebase a strain and the
>   origin link points from one of the commits in the strain to be rebased
>   back *into* the same strain being rebased (most likely a revert).
>   Rebase can be bothered to renumber the origin link in this case.

Nope, because you might have a branch to the original origin link, and some body else may have already done a cherry-pick to the original origin commit. You've hand-waved around the problem by saying, "don't do that", but it just points out how **fragile** the origin link scheme really is. It's just not robust.

In contrast, generating a UUID per commit is much more robust, since you can now export it out of git in a patch, and then re-import it later, and have the right thing happen.

Show 7 quoted lines
> >(and I am not convinced that you do), the ***much*** better approach
> >is to use the same approach as the bug tracking identifier, and add a
> >level of indirection.  How would that work in practice?  Whenever you
> >create a new commit, create a UUID which is assigned to the patch.
> 
> This only works if you know at time of commit that you want to backport
> it at some later date.

I'm suggesting that all commits (once you upgrade to a version of git that supports this --- and you've already handwaved away the question on whether you can get all of developers for a project to upgrade to the latest git, remember) would have a UUID generated. That UUID could be stored internal to git, or (perhaps as an initial prototyping as a proof of concept, before we add something into the git commit record **forever**) could be in the free-form text.

Show 13 quoted lines
> >Yes, it means that you have to maintain a separate database so you can
> >easily find the list of commits that contain a particular UUID, but I
> >suspect you would need this in the case of the origin link concept
> >anyway, since sooner or later some of the more useful uses of said
> >link would require you to be able to find the commits which had origin
> >links to the original commit, which means you would need to create and
> >maintain this database anyway.
> 
> That isn't true.  Finding commits which have origin links to a certain
> commit is just as hard as finding all children of a certain commit.
> It's not exactly instant, but it is not a big problem, and depending on
> the amount of repositorytraversal you already are doing, it might even
> be a negligible amount of extra overhead.

My point is you'll need this separate database anyway, in order to deal with the cases where you have two commits that point to the same (non-existent) origin link, one in maint1, and one in maint2, and given the commit in the maint1 branch, you want to see if there is related comit in the maint2 branch you'll need this database anyway (or you do a brute force search of the repository, which isn't too bad for modest datbases). It's identical in both cases --- but having a UUID field in the commit is much *cleaner*, since it merely states that these two commits introduce the same semantic change; it doesn't imply some kind of parent/child relationship which an origin link implies.

Show 8 quoted lines
> The database needs to be available to anyone doing a clone of the
> repository, which implies that:
> - It needs to be network based.
> - It needs controlled write access (which is a mess).
> - It is slow during blame/gitk operations.
> - It is rather nontrivial to get things setup such that someone (after
>   cloning the repository) is able to run cherry-pick/gitk/blame/revert
>   and have those commands use the database transparently.

No it doesn't, since the database can be inferred from the objects in the repository. So you can generate it locally if you need it, merely as an optimization. The same is true for the origin link proposal, as I've said.

    		    	    	     - Ted
Previous: Nicolas PitreNext: Petr Baudis
Message 129 of 137 in “[RFC] origin link for cherry-pick and revert”
  1. Stephen R. van den BergSep 9, 2008
  2. Paolo BonziniSep 9, 2008
  3. Stephen R. van den BergSep 9, 2008
  4. Stephen R. van den BergSep 9, 2008
  5. Jakub NarebskiSep 9, 2008
  6. Steven GrimmSep 9, 2008
  7. Stephen R. van den BergSep 9, 2008
  8. Jeff KingSep 9, 2008
  9. Stephen R. van den BergSep 9, 2008
  10. Junio C HamanoSep 9, 2008
  11. Shawn O. PearceSep 9, 2008
  12. Jeff KingSep 9, 2008
  13. Jakub NarebskiSep 9, 2008
  14. Jakub NarebskiSep 9, 2008
  15. Paolo BonziniSep 10, 2008
  16. Stephen R. van den BergSep 10, 2008
  17. Junio C HamanoSep 10, 2008
  18. Stephen R. van den BergSep 10, 2008
  19. Junio C HamanoSep 9, 2008
  20. Jeff KingSep 9, 2008
  21. Stephen R. van den BergSep 9, 2008
  22. Jakub NarebskiSep 9, 2008
  23. Stephen R. van den BergSep 9, 2008
  24. Linus TorvaldsSep 9, 2008
  25. Stephen R. van den BergSep 9, 2008
  26. Linus TorvaldsSep 10, 2008
  27. Stephen R. van den BergSep 10, 2008
  28. Linus TorvaldsSep 10, 2008
  29. Stephen R. van den BergSep 10, 2008
  30. Linus TorvaldsSep 11, 2008
  31. Stephen R. van den BergSep 11, 2008
  32. Jakub NarebskiSep 11, 2008
  33. Stephen R. van den BergSep 11, 2008
  34. Theodore TsoSep 11, 2008
  35. Stephen R. van den BergSep 11, 2008
  36. Theodore TsoSep 11, 2008
  37. Stephen R. van den BergSep 11, 2008
  38. Nicolas PitreSep 11, 2008
  39. Stephen R. van den BergSep 11, 2008
  40. Nicolas PitreSep 11, 2008
  41. Stephen R. van den BergSep 11, 2008
  42. Jakub NarebskiSep 11, 2008
  43. Stephen R. van den BergSep 11, 2008
  44. A Large Angry SCMSep 12, 2008
  45. Stephen R. van den BergSep 12, 2008
  46. Theodore TsoSep 11, 2008
  47. Jeff KingSep 11, 2008
  48. Stephen R. van den BergSep 11, 2008
  49. Jeff KingSep 11, 2008
  50. Stephen R. van den BergSep 11, 2008
  51. Linus TorvaldsSep 11, 2008
  52. Jeff KingSep 11, 2008
  53. Stephen R. van den BergSep 11, 2008
  54. Nicolas PitreSep 11, 2008
  55. Stephen R. van den BergSep 11, 2008
  56. Nicolas PitreSep 11, 2008
  57. Stephen R. van den BergSep 11, 2008
  58. Nicolas PitreSep 11, 2008
  59. Junio C HamanoSep 11, 2008
  60. Stephen R. van den BergSep 11, 2008
  61. Stephen R. van den BergSep 11, 2008
  62. A Large Angry SCMSep 11, 2008
  63. Stephen R. van den BergSep 11, 2008
  64. A Large Angry SCMSep 12, 2008
  65. Stephen R. van den BergSep 12, 2008
  66. Linus TorvaldsSep 11, 2008
  67. Paolo BonziniSep 11, 2008
  68. Linus TorvaldsSep 11, 2008
  69. Stephen R. van den BergSep 11, 2008
  70. Jakub NarebskiSep 11, 2008
  71. Stephen R. van den BergSep 11, 2008
  72. Nicolas PitreSep 11, 2008
  73. Stephen R. van den BergSep 11, 2008
  74. Nicolas PitreSep 11, 2008
  75. Stephen R. van den BergSep 12, 2008
  76. Theodore TsoSep 11, 2008
  77. Stephen R. van den BergSep 12, 2008
  78. Paolo BonziniSep 10, 2008
  79. Linus TorvaldsSep 10, 2008
  80. Paolo BonziniSep 10, 2008
  81. Linus TorvaldsSep 10, 2008
  82. Linus TorvaldsSep 10, 2008
  83. Paolo BonziniSep 10, 2008
  84. Stephen R. van den BergSep 10, 2008
  85. Jakub NarebskiSep 10, 2008
  86. Sam VilainSep 11, 2008
  87. Linus TorvaldsSep 11, 2008
  88. Sam VilainSep 12, 2008
  89. Stephen R. van den BergSep 12, 2008
  90. Rogan DawesSep 12, 2008
  91. Stephen R. van den BergSep 12, 2008
  92. Theodore TsoSep 12, 2008
  93. Paolo BonziniSep 12, 2008
  94. Jakub NarebskiSep 12, 2008
  95. Paolo BonziniSep 12, 2008
  96. Theodore TsoSep 12, 2008
  97. Stephen R. van den BergSep 12, 2008
  98. Jeff KingSep 12, 2008
  99. Stephen R. van den BergSep 12, 2008
  100. Theodore TsoSep 12, 2008
  101. Stephen R. van den BergSep 12, 2008
  102. Sam VilainSep 15, 2008
  103. Jakub NarebskiSep 9, 2008
  104. Petr BaudisSep 9, 2008
  105. Stephen R. van den BergSep 9, 2008
  106. Petr BaudisSep 9, 2008
  107. Stephen R. van den BergSep 9, 2008
  108. Paolo BonziniSep 10, 2008
  109. Petr BaudisSep 10, 2008
  110. Stephen R. van den BergSep 10, 2008
  111. Petr BaudisSep 10, 2008
  112. Stephen R. van den BergSep 10, 2008
  113. Dmitry PotapovSep 10, 2008
  114. Stephen R. van den BergSep 10, 2008
  115. Paolo BonziniSep 10, 2008
  116. Theodore TsoSep 10, 2008
  117. Stephen R. van den BergSep 10, 2008
  118. Jeff KingSep 10, 2008
  119. Stephen R. van den BergSep 10, 2008
  120. Jeff KingSep 10, 2008
  121. Stephen R. van den BergSep 10, 2008
  122. Jeff KingSep 10, 2008
  123. Stephen R. van den BergSep 10, 2008
  124. Paolo BonziniSep 11, 2008
  125. Stephen R. van den BergSep 11, 2008
  126. Paolo BonziniSep 11, 2008
  127. A Large Angry SCMSep 11, 2008
  128. Nicolas PitreSep 11, 2008
  129. Theodore TsoSep 10, 2008
  130. Petr BaudisSep 10, 2008
  131. Paolo BonziniSep 10, 2008
  132. Stephen R. van den BergSep 10, 2008
  133. Paolo BonziniSep 10, 2008
  134. Recording "partial merges" (was: Re: [RFC] origin link for cherry-pick and revert)Peter Krefting, Sep 23, 2008
  135. Miklos VajnaSep 10, 2008
  136. Nicolas PitreSep 10, 2008
  137. Miklos VajnaSep 10, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.