# Searching explanation of different diff algorithms

3 messages from 2013-09-25 to 2013-09-25. Participants: Thomas Koch, Ondřej Bílka, Peter Oberndorfer.
Thread: https://gitlist.dev/t/35014

## Thomas Koch, 2013-09-25 07:24

Subject: Searching explanation of different diff algorithms
Message-ID: <201309250924.15741.thomas@koch.ro>
URL: https://gitlist.dev/e/201309250924.15741.thomas%40koch.ro

```
Is there any explanation available of the different merrits and drawbacks of 
the diff algorithms that Git supports?

I'm not satisfied with the default diff but have enough processing power for a 
slower algorithm that might produce diffs that better show the intention of the 
edit.

Thank you, Thomas Koch

```

## Ondřej Bílka, 2013-09-25 08:55

Subject: Re: Searching explanation of different diff algorithms
Message-ID: <20130925085557.GA11402@domone.kolej.mff.cuni.cz>
URL: https://gitlist.dev/e/20130925085557.GA11402%40domone.kolej.mff.cuni.cz
In-Reply-To: <201309250924.15741.thomas@koch.ro>

```
On Wed, Sep 25, 2013 at 09:24:15AM +0200, Thomas Koch wrote:
> Is there any explanation available of the different merrits and drawbacks of 
> the diff algorithms that Git supports?
> 
> I'm not satisfied with the default diff but have enough processing power for a 
> slower algorithm that might produce diffs that better show the intention of the 
> edit.
> 
It is not just question of algorithm, even definition how should most
readable diff look like is problematic, for example when large block is
rewritten and one line is unchanged then you get diff like

if (x){
- foo
+ bar
} else {
- foo
+ bar
}

but it is better to create following diff as it does not break flow of code.

if (x) {
- foo
-} else {
- foo
+ bar
+} else {
+ bar
}

```

## Peter Oberndorfer, 2013-09-25 15:30

Subject: Re: Searching explanation of different diff algorithms
Message-ID: <524301A5.2060401@arcor.de>
URL: https://gitlist.dev/e/524301A5.2060401%40arcor.de
In-Reply-To: <20130925085557.GA11402@domone.kolej.mff.cuni.cz>

```
On 2013-09-25 10:55, Ondřej Bílka wrote:
> On Wed, Sep 25, 2013 at 09:24:15AM +0200, Thomas Koch wrote:
>> Is there any explanation available of the different merrits and drawbacks of 
>> the diff algorithms that Git supports?
>>
>> I'm not satisfied with the default diff but have enough processing power for a 
>> slower algorithm that might produce diffs that better show the intention of the 
>> edit.
>>
> It is not just question of algorithm, even definition how should most
> readable diff look like is problematic, for example when large block is
> rewritten and one line is unchanged then you get diff like
> 
> if (x){
> - foo
> + bar
> } else {
> - foo
> + bar
> }
> 
> but it is better to create following diff as it does not break flow of code.
> 
> if (x) {
> - foo
> -} else {
> - foo
> + bar
> +} else {
> + bar
> }

I already asked the list for such a feature in the past[1].
I might be able to provide a rough/unfinished hack
that does exactly this in a few days after cleaning it up a bit.

It works like this:
If 2 hunks are separated by less than a certain count of lines and
those lines are identified as containing no "interesting information"
like {, }, /*, */, <whitespace> then the 2 hunks are fused together.

The hack is mainly lacking the following things:
* A way to identify boring lines.
(a like a list of boring keywords?, per filetype?)
* Configuration/commandline options to turn it on/off
* Tests
* Cleanup the code

Greetings Peter

[1] http://article.gmane.org/gmane.comp.version-control.git/207239/

```
