git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Following renames

From
PJPaul Jakma <paul@clubi.ie>
Date
Mar 27, 2006, 06:00 UTC
Message-ID
<Pine.LNX.4.64.0603270642090.5276@sheen.jakma.org>
In-Reply-To
<e05354$cc9$1@sea.gmane.org>
On Sun, 26 Mar 2006, Jakub Narebski wrote:
> I think one of the better ideas/suggestions about *recording* filenames was
> in the "impure renames / history tracking" thread
> http://marc.theaimsgroup.com/?l=git&m=114122175216489&w=2
> <Pine.LNX.4.64.0603011343170.13612@sheen.jakma.org>

For the record, the responses I received were educational ;). Sufficiently so I no longer think renames should be recorded. At least, definitely not as renames.

I now grok the reasoning for doing it by 'similarity' - it is indeed a *much* more useful concept. (E.g. the 'pickaxe' idea people keep alluding though sounds amazingly useful).

So the question really is what, if any, weaknesses does the current similarity estimation have, and how to solve them. I can think of two weaknesses:

1. the similarity algorithms can be expensive potentially, and they
    essentially get run a lot with the same inputs, to produce the
    same results - over and over as one works with a git repo. (there
    was a thread a while ago on this I think).
2. Some 'similarities' are just not deducible by current software
    state of the art. E.g. where some code is rewritten in another
    language:
 	foo.X -> foo.Y
    The high-level algorithms may remain the exact same, but the code
    may be unrecognisable as similar except to a human. However,
    tracking history back across this rewrite probably would still be
    valuable to the human.

So I think what /might/ be interesting is to have a 'similarity cache', which would help 1, and to allow for manual injection of such hints (into a seperate and stronger cache most likely) - which would help 2.

Something to record the following information:
(tree1,tree2)[1]:
 	Id1 <-> Id1'
 	.
 	.
 	.
 	Idn <-> Idn'
That would allow:
1. Performance repercussions of similarity estimation to be one-time,
    cached there-after. (throw-away information, if a better
    similarity estimation heuristic comes along, you can rebuild this
    cache)
2. The user to inject their own 'hints' into similarity estimation,
    particularly for cases that just aren't obvious and probably never
    will be to software estimators (e.g. the rewrite cases), but where
    the user sees value in being able to follow back the history.
Avoids:
- encoding anything permanently into the repository (which was
   something I was thinking of, and others before me apparently, but
   which I now accept would be an awful idea ;) ).
1. I'm not sure if it should be indexed by (commit ID) or
    (tree1,tree2) tuple. ??
regards,
-- 
Paul Jakma	paul@clubi.ie	paul@jakma.org	Key ID: 64A2FF6A
Fortune:
Men take only their needs into consideration -- never their abilities.
 		-- Napoleon Bonaparte
Previous: Jakub NarebskiNext: Petr Baudis
Message 4 of 41 in “Following renames”
  1. Petr BaudisMar 26, 2006
  2. Junio C HamanoMar 26, 2006
  3. Jakub NarebskiMar 26, 2006
  4. Paul JakmaMar 27, 2006
  5. Petr BaudisMar 26, 2006
  6. Petr BaudisMar 26, 2006
  7. Timo HirvonenMar 26, 2006
  8. Linus TorvaldsMar 26, 2006
  9. Jakub NarebskiMar 26, 2006
  10. Linus TorvaldsMar 26, 2006
  11. Jakub NarebskiMar 26, 2006
  12. Linus TorvaldsMar 26, 2006
  13. Marco CostalbaMar 26, 2006
  14. Linus TorvaldsMar 26, 2006
  15. Marco CostalbaMar 27, 2006
  16. Junio C HamanoMar 27, 2006
  17. Linus TorvaldsMar 27, 2006
  18. Marco CostalbaMar 27, 2006
  19. Johannes SchindelinMar 27, 2006
  20. Linus TorvaldsMar 27, 2006
  21. Marco CostalbaMar 27, 2006
  22. Andreas EricssonMar 27, 2006
  23. Jakub NarebskiMar 27, 2006
  24. David LangMar 27, 2006
  25. Jakub NarebskiMar 27, 2006
  26. Linus TorvaldsMar 26, 2006
  27. Ryan AndersonMar 26, 2006
  28. Petr BaudisMar 26, 2006
  29. Fredrik KuivinenMar 26, 2006
  30. Linus TorvaldsMar 26, 2006
  31. Petr BaudisMar 26, 2006
  32. Petr BaudisMar 26, 2006
  33. Linus TorvaldsMar 26, 2006
  34. Petr BaudisMar 26, 2006
  35. Junio C HamanoMar 26, 2006
  36. Linus TorvaldsMar 26, 2006
  37. Junio C HamanoMar 27, 2006
  38. Linus TorvaldsMar 26, 2006
  39. Petr BaudisMar 26, 2006
  40. Petr BaudisMar 27, 2006
  41. Petr BaudisMar 26, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.