git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: libxdiff and patience diff

From
Pierre Habouzit <madcoder@debian.org>
Date
Nov 4, 2008, 16:15 UTC
Message-ID
<20081104161548.GB21842@artemis.corp>
In-Reply-To
<alpine.DEB.1.00.0811041650510.24407@pacific.mpi-cbg.de>
On Tue, Nov 04, 2008 at 03:57:44PM +0000, Johannes Schindelin wrote:
Show 10 quoted lines
> Hi,
> 
> On Tue, 4 Nov 2008, Pierre Habouzit wrote:
> 
> > The nasty thing about the patience diff is that it still needs the usual
> > diff algorithm once it has split the file into chunks separated by
> > "unique lines".
> 
> Actually, it should try to apply patience diff again in those chunks, 
> separately.

yes it's what I do, but this has a fixed point as soon as you don't find unique lines between the new found ones, or that that space is "empty". E.g. you could have the two following hunks:

File A File B 1 2 2 1 1 2 2 1 1 2 2 1

The simple leading/trailing reduction will do nothing, and you don't have any shared unique lines, on that you must apply the usual diff algorithm.

Show 8 quoted lines
> > So you cannot make really independant stuff. What I could do is put most 
> > of the xpatience diff into xpatience.c but it would still have to use 
> > some functions from xdiffi.c that are currently private, so it messes 
> > somehow the files more than it's worth IMHO.
> 
> I think it is better that you use the stuff from xdiffi.c through a well 
> defined interface, i.e. _not_ mess up the code by mingling it together 
> with the code in xdiffi.c.  The code is hard enough to read already.
Hmmm. I'll see to that later, once I have something that works.
> Oh, BTW, "ha" is a hash of the lines which is used to make the line 
> matching more performant.  You will see a lot of "ha" comparisons before 
> actually calling xdl_recmatch() for that reason.  Incidentally, this is 
> also the hash that I'd use for the hash multi-set I was referring to.
Yeah, that's what I assumed it would be.
> Oh, and I am not sure that it is worth your time trying to get it to run 
> with the linear list, since you cannot reuse that code afterwards, and 
> have to spend the same amount of time to redo it with the hash set.

Having the linear list (actually an array) work would show me I hook at the proper place. Replacing a data structure doesn't makes me afraid because I've split the functions properly.

> I am awfully short on time, so it will take some days until I can review 
> what you have already, unfortunately.

NP, it was just in case, because I'm horribly stuck with that code right now ;)

-- 
·O·  Pierre Habouzit
··O                                                madcoder@debian.org
OOO                                                http://www.madism.org
Previous: Johannes SchindelinNext: Johannes Schindelin
Message 9 of 67 in “libxdiff and patience diff”
  1. Pierre HabouzitNov 4, 2008
  2. Davide LibenziNov 4, 2008
  3. Pierre HabouzitNov 4, 2008
  4. Johannes SchindelinNov 4, 2008
  5. Pierre HabouzitNov 4, 2008
  6. Johannes SchindelinNov 4, 2008
  7. Pierre HabouzitNov 4, 2008
  8. Johannes SchindelinNov 4, 2008
  9. Pierre HabouzitNov 4, 2008
  10. 0/3 Teach Git about the patience diff algorithmJohannes Schindelin, Jan 1, 2009
  11. 1/3 Implement the patience diff algorithmJohannes Schindelin, Jan 1, 2009
  12. 2/3 Introduce the diff option '--patience'Johannes Schindelin, Jan 1, 2009
  13. 3/3 bash completions: Add the --patience optionJohannes Schindelin, Jan 1, 2009
  14. Linus TorvaldsJan 1, 2009
  15. Linus TorvaldsJan 1, 2009
  16. Johannes SchindelinJan 2, 2009
  17. Linus TorvaldsJan 2, 2009
  18. Johannes SchindelinJan 2, 2009
  19. Jeff KingJan 2, 2009
  20. 1/3 Implement the patience diff algorithmJohannes Schindelin, Jan 2, 2009
  21. Johannes SchindelinJan 2, 2009
  22. Adeodato SimóJan 1, 2009
  23. Linus TorvaldsJan 2, 2009
  24. Clemens BuchacherJan 2, 2009
  25. Clemens BuchacherJan 2, 2009
  26. Linus TorvaldsJan 2, 2009
  27. Johannes SchindelinJan 2, 2009
  28. Linus TorvaldsJan 2, 2009
  29. Johannes SchindelinJan 2, 2009
  30. Jeff KingJan 2, 2009
  31. Jeff KingJan 2, 2009
  32. Jeff KingJan 2, 2009
  33. Linus TorvaldsJan 2, 2009
  34. Bazaar's patience diff as GIT_EXTERNAL_DIFFAdeodato Simó, Jan 3, 2009
  35. Johannes SchindelinJan 2, 2009
  36. Junio C HamanoJan 2, 2009
  37. Adeodato SimóJan 2, 2009
  38. Pierre HabouzitJan 6, 2009
  39. Pierre HabouzitJan 6, 2009
  40. Johannes SchindelinJan 6, 2009
  41. Pierre HabouzitJan 7, 2009
  42. Johannes SchindelinJan 7, 2009
  43. 1/3 Implement the patience diff algorithmJohannes Schindelin, Jan 7, 2009
  44. Davide LibenziJan 7, 2009
  45. Johannes SchindelinJan 7, 2009
  46. Davide LibenziJan 7, 2009
  47. Johannes SchindelinJan 7, 2009
  48. Linus TorvaldsJan 7, 2009
  49. Johannes SchindelinJan 7, 2009
  50. Davide LibenziJan 7, 2009
  51. Sam VilainJan 7, 2009
  52. Linus TorvaldsJan 7, 2009
  53. Sam VilainJan 8, 2009
  54. Johannes SchindelinJan 7, 2009
  55. Junio C HamanoJan 7, 2009
  56. Johannes SchindelinJan 7, 2009
  57. Pierre HabouzitJan 7, 2009
  58. Johannes SchindelinJan 7, 2009
  59. Adeodato SimóJan 8, 2009
  60. Adeodato SimóJan 8, 2009
  61. Junio C HamanoJan 9, 2009
  62. Johannes SchindelinJan 9, 2009
  63. Adeodato SimóJan 9, 2009
  64. Linus TorvaldsJan 9, 2009
  65. Linus TorvaldsJan 9, 2009
  66. Junio C HamanoJan 9, 2009
  67. Johannes SchindelinJan 10, 2009

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.