git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Following renames

From
Jakub Narebski <jnareb@gmail.com>
Date
Mar 27, 2006, 06:55 UTC
Message-ID
<e0827k$7tk$1@sea.gmane.org>
In-Reply-To
<Pine.LNX.4.64.0603260947100.15714@g5.osdl.org>
Linus Torvalds wrote:
Show 9 quoted lines
> On Sun, 26 Mar 2006, Jakub Narebski wrote:
>> 
>> If (2) is common enough then discussed improvements to rename detection,
>> namely comparing basenames as a base for candidate selection is a good
>> idea.
> 
> BK had this "renametool" which got started automatically when you applied
> a patch that removed one or more files and added one or more files, so
> that you could then pair up the files manually.
[...]
> The thing is, the fast rename detection that is in the "next" branch
> really does a lot better, and it's fast enough.
I was thinking about the fast ename detection algorithm in "next" branch.

That is the question if recording additional (helper) information about contents copying and moving like the mentioned "renametool" did is worth the effort, both in coding it and from user's point of view. Or would better contents copying and moving detection ("renames detection") for whatchanged and similar suffice.

I am of opinion that voluntary information about contents moving and copying in the commits would help.

Purposes:
1.) Record contents moving and similarity information which cannot or cannot
be easily calculated; see Paul Jakma response in this thread
  MessageID: <Pine.LNX.4.64.0603270642090.5276@sheen.jakma.org>
for example copying fragment of code, small fragment of the whole file,
creating documentation or header file from code, or code skeleton from
template, or rewrite of code in different language (e.g. shell script to
perl, script to compiled code e.g. Perl or Python to C).
2.) Caching the results of similarity algorithm/rename detection tool (also
Paul Jakma post), including remembering false positives and undetected
renames, for efficiency. Calculated automatically parts might be
throw-away.

Sources of information: 1.) Manually entered information *at commit*, including *-rm, *-mv, *-cp like commands (which nobody likes) and systematized (pseudolanguage?) for copying and moving contents in the log messages. 2.) Semi-manual tools like the mentioned "renametool" of BK. 3.) Support from editor (remebering where copied and pasted, or cut and pasted fragment came from, and providing prefilled command to record contents moving ("renames") or prefilled commit log containing this information. Hard to get, probably most useful. 4.) Information from resolved merges and results of diagnosis (pickaxe like) tools, especially recording "renames" which were not detected, and removing "renames" which were detected falsily.

Is that the place where I should provide code (patch) for testing the idea :) ?

Show 24 quoted lines
>> I wonder how common is (2) compared to (1)+(2) i.e. move to other dir
>> and rename, old-dir/old-file.c to new-dir/new-subdir/new-file.c
>
> For example, one common case was a directory structure like
> 
> ..
> type-file1.c
> type-file2.c
> otherfiles.c
> yet-more.c
> ..
> 
> being split up into a subdirectory
> 
> ..
> type/file1.c
> type/file2.c
> otherfiles.c
> yet-more.c
> ..
> 
> (eg drivers/scsi/aic7xx-* being given a subdirectory of it's own, as
> drivers/scsi/aic7xx/*). So the basename wouldn't stay the same, because it
> contained some piece of data that became redundant with the move.

Perhaps fast rename detection algorithm needs some smart similarity estimate for names, which would put more weight in the parts closer to basename, and would detect */type-file1.c and */type/file1.c as similar.

-- 
Jakub Narebski
Warsaw, Poland
Previous: Andreas EricssonNext: David Lang
Message 23 of 41 in “Following renames”
  1. Petr BaudisMar 26, 2006
  2. Junio C HamanoMar 26, 2006
  3. Jakub NarebskiMar 26, 2006
  4. Paul JakmaMar 27, 2006
  5. Petr BaudisMar 26, 2006
  6. Petr BaudisMar 26, 2006
  7. Timo HirvonenMar 26, 2006
  8. Linus TorvaldsMar 26, 2006
  9. Jakub NarebskiMar 26, 2006
  10. Linus TorvaldsMar 26, 2006
  11. Jakub NarebskiMar 26, 2006
  12. Linus TorvaldsMar 26, 2006
  13. Marco CostalbaMar 26, 2006
  14. Linus TorvaldsMar 26, 2006
  15. Marco CostalbaMar 27, 2006
  16. Junio C HamanoMar 27, 2006
  17. Linus TorvaldsMar 27, 2006
  18. Marco CostalbaMar 27, 2006
  19. Johannes SchindelinMar 27, 2006
  20. Linus TorvaldsMar 27, 2006
  21. Marco CostalbaMar 27, 2006
  22. Andreas EricssonMar 27, 2006
  23. Jakub NarebskiMar 27, 2006
  24. David LangMar 27, 2006
  25. Jakub NarebskiMar 27, 2006
  26. Linus TorvaldsMar 26, 2006
  27. Ryan AndersonMar 26, 2006
  28. Petr BaudisMar 26, 2006
  29. Fredrik KuivinenMar 26, 2006
  30. Linus TorvaldsMar 26, 2006
  31. Petr BaudisMar 26, 2006
  32. Petr BaudisMar 26, 2006
  33. Linus TorvaldsMar 26, 2006
  34. Petr BaudisMar 26, 2006
  35. Junio C HamanoMar 26, 2006
  36. Linus TorvaldsMar 26, 2006
  37. Junio C HamanoMar 27, 2006
  38. Linus TorvaldsMar 26, 2006
  39. Petr BaudisMar 26, 2006
  40. Petr BaudisMar 27, 2006
  41. Petr BaudisMar 26, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.