git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: impure renames / history tracking

From
PJPaul Jakma <paul@clubi.ie>
Date
Mar 1, 2006, 21:25 UTC
Message-ID
<Pine.LNX.4.64.0603012105230.13612@sheen.jakma.org>
In-Reply-To
<7v3bi2ey63.fsf@assigned-by-dhcp.cox.net>
Hi Junio,
On Wed, 1 Mar 2006, Junio C Hamano wrote:
Show 6 quoted lines
> Interestingly enough, there are two levels of "rename tracking" the 
> current git does.  Whey you run "git whatchanged -M", you are 
> looking at renames between each commit in the commit chain, one 
> step at a time.  There as long as the rename+rewrite does not 
> amount to too much rewrite, you would see what should be detected 
> as rename to be detected as renames.
Right.
> I found the current default threshold parameters to be about right, 
> maybe a bit too tight sometimes, though.  If you want to loosen the 
> default, you can specify similiarity index after -M.
That's one option.

I'm wondering though if we couldn't also allow for users to additionally encode naming 'hints', to aid this 'similarity' detection process.

Show 10 quoted lines
> The way recursive merge strategy uses the rename detection, unlike 
> what whatchanged shows you, does not use chains of commits down to 
> the common merge base in order to detect renames (my recollection 
> may be wrong here -- it's a while since I looked at the recursive 
> merge the last time).  It just looks at the two heads being merged, 
> and detects similarility between them.  So it does not make _any_ 
> difference with the current implementation of recursive merge if 
> you kept a history full of "honest but disgusting" commits or 
> collapsed them into a history with small number of "cleaned up" 
> commits.

I'm going to have to stare at this paragraph a lot longer and harder to understand it :).

> One thing it _could_ do (and you _could_ implement as another merge 
> strategy and call it "pauls-rename" merge) is to follow the commit 
> chain one by one down to the common merge base from both heads 
> being merged, and analyze rename history on the both commit chains.

Right, I was just thinking that while making tea actually. This could be part of the 'collapsing' process. (or call it "coalesce too-detailed commits" process if that is less offensive to ones sense of process ;) ).

Actually, you're sort of suggesting following the chains in parallel, right? Ie in wall-clock time order, rather than chain order. And doing name resolution across the 'to-be-merged' chains at each step of the way? Sort of a lesser subset of how other SCMs maintain state for names globally?

It's not so much /resolving/ names I'm worried about in the first place. It's there simply being no information in the first place to indicate (from one single-parent commit to the next) which names were renamed.

> Then, you would get better rename+rewrite detection than what it 
> currently does.
But if I follow the commit chain in order to try extract
> HOWEVER.
Show 5 quoted lines
> If you have that kind of rename-following merge, a workflow that 
> collapses a useful history into a single huge commit "Ok, this 
> commit is a roll-up patch between version 2.6.14 and 2.6.15" 
> becomes far less attractive than it currently already is.  At that 
> point, you _are_ throwing away useful history.

Yes, I agree. And I am, as part of arguing git's case (several SCMs are being evaluated and considered, I'm the git proponent at the moment), I'm going to suggest workflow ought to be re-evaluated to ensure it is generally reasonable, rather than be kept for the sake of it keeping (particularly as it may be tailored to the needs/limitations of $TRADITIONAL_SCM).

However, I suspect at least some level of collapsing will be desired (just as it is with Linux and git).

The workflow issue is seperate from the 'impure rename' issue though, even if the workflow I gave as an example excerbates the issue, "rename and rewrite half of it" and hard-to-detect renames can still occur in the detailed git/linux workflows, surely?

regards,
-- 
Paul Jakma	paul@clubi.ie	paul@jakma.org	Key ID: 64A2FF6A
Fortune:
If you really knew C++, you wouldn't even joke about putting it
in the kernel.

 	- Richard Johnson on linux-kernel
Previous: Junio C HamanoNext: Andreas Ericsson
Message 12 of 15 in “impure renames / history tracking”
  1. Paul JakmaMar 1, 2006
  2. Andreas EricssonMar 1, 2006
  3. Paul JakmaMar 1, 2006
  4. Linus TorvaldsMar 1, 2006
  5. Paul JakmaMar 1, 2006
  6. Andreas EricssonMar 1, 2006
  7. Paul JakmaMar 2, 2006
  8. Andreas EricssonMar 2, 2006
  9. Martin LanghoffMar 1, 2006
  10. Paul JakmaMar 1, 2006
  11. Junio C HamanoMar 1, 2006
  12. Paul JakmaMar 1, 2006
  13. Andreas EricssonMar 1, 2006
  14. Paul JakmaMar 1, 2006
  15. Junio C HamanoMar 1, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.