git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH 0/3] Teach Git about the patience diff algorithm

From
Johannes Schindelin <johannes.schindelin@gmx.de>
Date
Jan 2, 2009, 19:22 UTC
Message-ID
<alpine.DEB.1.00.0901022008490.30769@pacific.mpi-cbg.de>
In-Reply-To
<alpine.LFD.2.00.0901021050450.5086@localhost.localdomain>
Hi,
On Fri, 2 Jan 2009, Linus Torvalds wrote:
Show 12 quoted lines
> On Fri, 2 Jan 2009, Johannes Schindelin wrote:
> 
> > And the worst part: one can only _guess_ what motivated patience diff.  
> > I imagine it came from the observation that function headers are 
> > unique, and that you usually want to preserve as much context around 
> > them.
> 
> Well, I do like the notion of giving more weight to unique lines - I think 
> it makes sense. That said, I suspect it would make almost as much sense to 
> give more weight simply to _longer_ lines, and I suspect the standard 
> Myers' algorithm could possibly be simply extended to take line size into 
> account when calculating the weights.

I think that it makes more sense with the common unique lines, because they are _unambiguously_ common. That is, if possible, we would like to keep them as common lines.

BTW this also opens the door for more automatic code movement detection.

As for the longer lines: what exactly would you want to put "weight" on? The edit distance is the number of plusses and minusses, and this is the thing that actually is pretty critical for the performance: the larger the distance, the longer it takes. So if you want to put a different "weight" on a line, i.e. something else than a "1", you are risking a substantially worse performance.

And I am still not convinced that longer lines should get more weight. A line starting with "exit:" can be much more important than 8 tabs followed by a curly bracket.

> Because the problem with diffs for C doesn't really tend to be as much 
> about non-unique lines as about just _trivial_ lines that are mostly empty 
> or contain just braces etc. Those are quite arguably almost totally 
> worthless for equality testing.
Oh, but then we get very C specific.  Let's not go there.
> And btw, don't get me wrong - I don't think there is anything wrong with 
> the patience diff. I think it's a worthy thing to try, and I'm not at 
> all arguing against it.

I never took you to be opposed. I myself mainly implemented it because I wanted to play with it, and have something more useful than what I found in the whole wide web.

> However, I do think that the people arguing for it often do so based on 
> less-than-very-logical arguments, and it's entirely possible that other 
> approaches are better (eg the "weight by size" thing rather than "weight 
> by uniqueness").

Actually, I think it is very possible that merging hunks if there are not enough alnums between them could make a whole lot of sense.

But IIRC it was shot down and replaced by the not-so-specific-to-C algorithm to merge hunks when the output is shorter than having separate hunks.

Show 7 quoted lines
> The thing about unique lines is that there are no guarantees at all that 
> they even exist, so a uniqueness-based thing will always have to fall 
> back on anything else. That, to me, implies that the whole notion is 
> somewhat mis-designed: it's clearly not a generic concept.
> 
> (In contrast, taking the length of the matching lines into account would 
> not have that kind of bad special case)

However, we know that humans like to start from the unique features they see, and continue from there.

And while it is quite possible that it is the wrong approach (see all that security theater at the airport, and some nice rational analyses about the cost/benefit equations there), imitating humans is still the thing that will convince most humans.

Ciao, Dscho

Previous: Linus TorvaldsNext: Jeff King
Message 29 of 67 in “libxdiff and patience diff”
  1. Pierre HabouzitNov 4, 2008
  2. Davide LibenziNov 4, 2008
  3. Pierre HabouzitNov 4, 2008
  4. Johannes SchindelinNov 4, 2008
  5. Pierre HabouzitNov 4, 2008
  6. Johannes SchindelinNov 4, 2008
  7. Pierre HabouzitNov 4, 2008
  8. Johannes SchindelinNov 4, 2008
  9. Pierre HabouzitNov 4, 2008
  10. 0/3 Teach Git about the patience diff algorithmJohannes Schindelin, Jan 1, 2009
  11. 1/3 Implement the patience diff algorithmJohannes Schindelin, Jan 1, 2009
  12. 2/3 Introduce the diff option '--patience'Johannes Schindelin, Jan 1, 2009
  13. 3/3 bash completions: Add the --patience optionJohannes Schindelin, Jan 1, 2009
  14. Linus TorvaldsJan 1, 2009
  15. Linus TorvaldsJan 1, 2009
  16. Johannes SchindelinJan 2, 2009
  17. Linus TorvaldsJan 2, 2009
  18. Johannes SchindelinJan 2, 2009
  19. Jeff KingJan 2, 2009
  20. 1/3 Implement the patience diff algorithmJohannes Schindelin, Jan 2, 2009
  21. Johannes SchindelinJan 2, 2009
  22. Adeodato SimóJan 1, 2009
  23. Linus TorvaldsJan 2, 2009
  24. Clemens BuchacherJan 2, 2009
  25. Clemens BuchacherJan 2, 2009
  26. Linus TorvaldsJan 2, 2009
  27. Johannes SchindelinJan 2, 2009
  28. Linus TorvaldsJan 2, 2009
  29. Johannes SchindelinJan 2, 2009
  30. Jeff KingJan 2, 2009
  31. Jeff KingJan 2, 2009
  32. Jeff KingJan 2, 2009
  33. Linus TorvaldsJan 2, 2009
  34. Bazaar's patience diff as GIT_EXTERNAL_DIFFAdeodato Simó, Jan 3, 2009
  35. Johannes SchindelinJan 2, 2009
  36. Junio C HamanoJan 2, 2009
  37. Adeodato SimóJan 2, 2009
  38. Pierre HabouzitJan 6, 2009
  39. Pierre HabouzitJan 6, 2009
  40. Johannes SchindelinJan 6, 2009
  41. Pierre HabouzitJan 7, 2009
  42. Johannes SchindelinJan 7, 2009
  43. 1/3 Implement the patience diff algorithmJohannes Schindelin, Jan 7, 2009
  44. Davide LibenziJan 7, 2009
  45. Johannes SchindelinJan 7, 2009
  46. Davide LibenziJan 7, 2009
  47. Johannes SchindelinJan 7, 2009
  48. Linus TorvaldsJan 7, 2009
  49. Johannes SchindelinJan 7, 2009
  50. Davide LibenziJan 7, 2009
  51. Sam VilainJan 7, 2009
  52. Linus TorvaldsJan 7, 2009
  53. Sam VilainJan 8, 2009
  54. Johannes SchindelinJan 7, 2009
  55. Junio C HamanoJan 7, 2009
  56. Johannes SchindelinJan 7, 2009
  57. Pierre HabouzitJan 7, 2009
  58. Johannes SchindelinJan 7, 2009
  59. Adeodato SimóJan 8, 2009
  60. Adeodato SimóJan 8, 2009
  61. Junio C HamanoJan 9, 2009
  62. Johannes SchindelinJan 9, 2009
  63. Adeodato SimóJan 9, 2009
  64. Linus TorvaldsJan 9, 2009
  65. Linus TorvaldsJan 9, 2009
  66. Junio C HamanoJan 9, 2009
  67. Johannes SchindelinJan 10, 2009

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.