git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: git-mv redux: there must be something else going on

From
RGRon Garret <ron1@flownet.com>
Date
Feb 3, 2010, 20:30 UTC
Message-ID
<ron1-DFA9D6.12301403022010@news.gmane.org>
In-Reply-To
<32541b131002031147r367ee08fxc64c4c54165953a3@mail.gmail.com>
In article 
<32541b131002031147r367ee08fxc64c4c54165953a3@mail.gmail.com>,
 Avery Pennarun <apenwarr@gmail.com> wrote:
Show 38 quoted lines
> On Wed, Feb 3, 2010 at 2:23 PM, Ron Garret <ron1@flownet.com> wrote:
> > In article
> > Ah.  That explains everything.  Thanks.  (I thought git mv was
> > equivalent to git rm followed by git add.  But it's not.)
> 
> I suppose in this case it's not.  The only difference is when your
> work tree differs from your index, though, and it's to be expected
> that 'git rm', in removing things from the index, would lose your
> ability to track those differences.
> 
> > So... how *does* git decide when two blobs are different blobs and when
> > they are the same blob with mods?  I asked this question before and was
> > pointed to the diffcore docs, but that didn't really clear things up.
> > That just describes all the different ways git can do diffs, not the
> > actual heuristics that git uses to track content.
> 
> If you really want to know the details, looking at the code really is
> probably the best solution; it's not even that long.
> 
> The short version is that git chooses a set of candidate blobs, then
> diffs them and figures out a percentage similarity between each pair.
> (A simple way to think of the similarity index is "how long is the
> diff compared to the file itself?"  If the diff is of length zero, the
> similarity is 100%, and so on.) If the similarity is greater than a
> certain threshold, then it's considered to be the same file.
> 
> Choosing the set of candidates is actually the more interesting
> problem, since detecting moves using the above algorithm is O(n^2)
> with the number of candidates.  That's why 'git diff' and 'git log'
> don't do it at all by default.
> 
> If you provide -M, the set of candidates is the set of files that were
> removed/modified and the set of files that were added.  (Added files
> are compared against removed/modified files, iirc.)  Normally that's a
> very short list.  With -C, you need to compare all
> added/removed/modified files with all others, which is slightly more
> work.  With --find-copies-harder, it becomes potentially a *lot* of
> work.
Thanks!  That clarifies a lot.
rg
Previous: Avery PennarunNext: Nicolas Pitre
Message 5 of 20 in “git-mv redux: there must be something else going on”
  1. Ron GarretFeb 3, 2010
  2. Avery PennarunFeb 3, 2010
  3. Ron GarretFeb 3, 2010
  4. Avery PennarunFeb 3, 2010
  5. Ron GarretFeb 3, 2010
  6. Nicolas PitreFeb 3, 2010
  7. Ron GarretFeb 3, 2010
  8. Ron GarretFeb 3, 2010
  9. Avery PennarunFeb 3, 2010
  10. Ron GarretFeb 3, 2010
  11. Avery PennarunFeb 3, 2010
  12. Jay SoffianFeb 3, 2010
  13. Ron GarretFeb 4, 2010
  14. Ron GarretFeb 4, 2010
  15. Junio C HamanoFeb 4, 2010
  16. Nicolas PitreFeb 3, 2010
  17. Pete HarlanFeb 3, 2010
  18. Ron GarretFeb 3, 2010
  19. Documentation: clarify git-mv behaviour wrt dirty filesThomas Rast, Feb 3, 2010
  20. Junio C HamanoFeb 3, 2010

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.