git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Fix a pathological case in git detecting proper renames

From
Linus Torvalds <torvalds@linux-foundation.org>
Date
Nov 29, 2007, 23:03 UTC
Message-ID
<alpine.LFD.0.9999.0711291442300.8458@woody.linux-foundation.org>
In-Reply-To
<alpine.LFD.0.9999.0711291303000.8458@woody.linux-foundation.org>
On Thu, 29 Nov 2007, Linus Torvalds wrote:
Show 20 quoted lines
> 
> It's worth noting a few gotchas:
> 
>  - this scoring is currently only done for the "exact match" case. 
> 
>    In particular, in Kumar's example, even after this patch, the inexact
>    match case is still done as a copy+delete rather than as two renames:
> 
> 	 delete mode 100644 board/cds/mpc8555cds/u-boot.lds
> 	 copy board/{cds => freescale}/mpc8541cds/u-boot.lds (97%)
> 	 rename board/{cds/mpc8541cds => freescale/mpc8555cds}/u-boot.lds (97%)
> 
>    because apparently the "cds/mpc8541cds/u-boot.lds" copy looked 
>    a bit more similar to both end results. That said, I *suspect* we just 
>    have the exact same issue there - the similarity analysis just gave 
>    identical (or at least very _close_ to identical) similarity points, 
>    and we do not have any logic to prefer multiple renames over a 
>    copy/delete there.
> 
>    That is a separate patch.

Side note: just in case people were expecting me to actually _ship_ that separate patch that handles the fuzzy matches too.. I wasn't planning on doing that patch. The way the fuzzy rename detection is currently done, that's actually quite painful.

For the fuzzy rename detection, we generate the full score matrix, and sort it by the score, up front. So all the scoring - and more importantly, all the sorting - has actually been done before we actually start looking at *any* renames at all, so we cannot easily do the same thing I did for the exact renames, namely to take into account _earlier_ renames in the scoring. Because those earlier renames have simply not been done when the score is calculated.

This would probably become easier to do with the linear-time hash-based similarity engine (the stuff Jeff King was working on), but the way the code is currently structured - with no incremental rename detection at all, and with all the scoring in one global table - it's pretty painful.

			Linus
Previous: Linus TorvaldsNext: Jeff King
Message 8 of 14 in “problem with git detecting proper renames”
  1. Kumar GalaNov 29, 2007
  2. Linus TorvaldsNov 29, 2007
  3. Kumar GalaNov 29, 2007
  4. Linus TorvaldsNov 29, 2007
  5. Kumar GalaNov 29, 2007
  6. Kumar GalaNov 29, 2007
  7. Fix a pathological case in git detecting proper renamesLinus Torvalds, Nov 29, 2007
  8. Linus TorvaldsNov 29, 2007
  9. Jeff KingNov 29, 2007
  10. Linus TorvaldsNov 30, 2007
  11. Jeff KingNov 30, 2007
  12. Kumar GalaNov 30, 2007
  13. Junio C HamanoNov 30, 2007
  14. Jakub NarebskiNov 30, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.