Re: [ANNOUNCE] Cogito-0.8 (former git-pasky, big changes!)
- From
- Paul Jackson <pj@sgi.com>
- Date
- Apr 26, 2005, 17:40 UTC
- Message-ID
- <20050426104022.0c53167d.pj@sgi.com>
- In-Reply-To
- <20050426122304.GD18971@pasky.ji.cz>
> some way to deal with renames sensibly.
For a Linux kernel merge tool that I did inside SGI, I come to the conclusion the following heuristics identified renames fairly accurately, coming up with the same renames as a careful manual examination.
The key heuristic was to consider one file A to be a rename of another B if the size (number of lines) of the diff of A and B (diff -auw A B | wc -l) is less than 50% of the combined size of A and B (cat A B | wc -l).
Only pathname pairs with the same basename were considered - I was focused here on renames due to directory restructuring. I'm not sure now if this is a good assumption - but it sure saved some computation.
Any file with the string "Makefile" in its name had to be excluded from consideration.
In case of multiple potential renames (say one file is copied to two places, removing the original and modifying each copy a little), the 'best' rename was selected, where 'best' meant the lowest ratio of diff 'diff -auw' size to combined 'cat' size.
The end result of the above was a fairly natural identification of renames. If say I moved kernel/cpuset.c to mm/cpuset.c and changed it a little bit, the above heuristics would show a rename, plus a few changes.
--
I won't rest till it's the best ...
Programmer, Linux Scalability
Paul Jackson <pj@engr.sgi.com> 1.650.933.1373, 1.925.600.0401