git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: A better approach to diffing and merging

From
Karl Hasselström <kha@treskal.com>
Date
Dec 1, 2008, 09:54 UTC
Message-ID
<20081201095449.GA30857@diana.vm.bytemark.co.uk>
In-Reply-To
<4931F2DC.CE9B1E35@dessent.net>
On 2008-11-29 17:56:44 -0800, Brian Dessent wrote:
Show 11 quoted lines
> Ian Clarke wrote:
>
> > Provide the merge algorithm with the grammar of the programming
> > language, perhaps in the form of a Bison grammar file, or some
> > other standardized way to represent a grammar.
>
> There's a huge flaw in that approach for C/C++: in order to parse
> C/C++ you have to first preprocess it -- consider the twisty mazes
> that #ifdef/#else/#endif can create. But in order to preprocess
> source code you need a whole heap of extra information that is not
> in the repository (or if it is, cannot be automatically extracted.)

But it's probably not necessary to parse the input files exactly. All you have to do is parse it well enough that the diff of the parse trees is interesting.

And in practice, you'd probably also generate the "normal" diff, and then fall back to that one if the parse tree diff was worse.

> The idea may have value for langauges that are easy to parse and do
> not have all this preprocessor cruft, but I just don't see how it
> would be able to provide anything useful for non-trivial changes to
> real world C/C++, which require human eyes to decipher.

I think it could work. But there would be quite a bit of heuristics involved to get the "approximate" parsing right, so I'm pretty sure there's no way to find out without actually trying to build the thing.

-- 
Karl Hasselström, kha@treskal.com
      www.treskal.com/kalle
Previous: Brian DessentNext: Jakub Narebski
Message 5 of 7 in “A better approach to diffing and merging”
  1. Ian ClarkeNov 29, 2008
  2. Boyd Stephen Smith Jr.Nov 29, 2008
  3. Miklos VajnaNov 30, 2008
  4. Brian DessentNov 30, 2008
  5. Karl HasselströmDec 1, 2008
  6. Jakub NarebskiDec 1, 2008
  7. Karl HasselströmDec 2, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.