git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [ANNOUNCE] Git wiki

From
Petr Baudis <pasky@suse.cz>
Date
May 5, 2006, 18:54 UTC
Message-ID
<20060505185445.GD27689@pasky.or.cz>
In-Reply-To
<Pine.LNX.4.64.0605051123420.3622@g5.osdl.org>

Dear diary, on Fri, May 05, 2006 at 08:31:06PM CEST, I got a letter where Linus Torvalds <torvalds@osdl.org> said that...

> Moving data around happens with a whole lot more than "mv".

Let's keep this on the per-file level - if you want to go below the file granularity, I already _DID_ say that I agree that explicit tracking is not a way. (If sub-file tracking would end up having any usable reliability in real-world cases, which is something I do not take for granted.)

Another thing is, the sub-file content tracking would end up being a lot more "magic" than the simple per-file content tracking, and you stated several times that you prefer simple merge over better but magic merge - so why do you prefer sub-file content tracking anyway?

> It happens with patches (somebody _else_ may have done an "mv", without 
> using git at all),

_Here_ is the place for automated renames detection. Between applying and committing the patch, the user can verify that it got the renames right. That's impossible when guessing the renames later.

> and it happens with editors (moving data around until 
> most of it exists in another file).

I doubt this in fact happens that often (to a degree the automatic rename detection would catch). And if it happens, then the user has to tell Git - I have never heard that _this_ would be any problem in other version control systems. You could make it more foolproof by running the automatic rename detection on the diff being committed and suggesting the user that other yet unrecorded renames did happen.

The point is, the user stays in control and can override any stupid guess.
> So doing "*mv" is just a special case.
> 
> And supporting special cases is _wrong_. If you start depending on data 
> that isn't actually dependable, that's WRONG.

I prefer making this data dependable to having to resort to guessing on dependable less amount of data.

Show 12 quoted lines
> There's another reason why encoding movement information in the commit is 
> totally broken, namely the fact that a lot of the actions DO NOT WALK THE 
> COMMIT CHAIN!
> 
> Try doing
> 
> 	git diff v1.3.0..
> 
> and think about what that actually _means_. Think about the fact that it 
> doesn't actually walk the commit chain at all: it diffs the trees between 
> v1.3.0 and the current one. What if the rename happened in a commit in the 
> middle?

Then the automated renames detection will miss it given that the other accumulated differences are large enough, and the suggested workarounds _are_ precisely walking the commit chain.

If you use persistent file ids, you never miss it _AND_ you DO NOT WALK THE COMMIT CHAIN! You still just match file ids in the two trees.

> The "track contents, not intentions" approach avoids both these things. 
> The end result is _reliable_, not a "random guess".

No, the end result is whichever some heuristic randomly guessed, and it's not reliable either since the heuristic can change.

-- 
				Petr "Pasky" Baudis
Stuff: http://pasky.or.cz/
Right now I am having amnesia and deja-vu at the same time.  I think
I have forgotten this before.
Previous: Linus TorvaldsNext: Jakub Narebski
Message 16 of 25 in “Re: [ANNOUNCE] Git wiki”
  1. linux@horizon.comMay 5, 2006
  2. Fredrik KuivinenMay 5, 2006
  3. Jakub NarebskiMay 5, 2006
  4. Petr BaudisMay 5, 2006
  5. Junio C HamanoMay 5, 2006
  6. Petr BaudisMay 5, 2006
  7. Jakub NarebskiMay 5, 2006
  8. Jakub NarebskiMay 5, 2006
  9. Petr BaudisMay 5, 2006
  10. Linus TorvaldsMay 5, 2006
  11. Dave JonesMay 5, 2006
  12. Petr BaudisMay 5, 2006
  13. Petr BaudisMay 5, 2006
  14. Jakub NarebskiMay 5, 2006
  15. Linus TorvaldsMay 5, 2006
  16. Petr BaudisMay 5, 2006
  17. Jakub NarebskiMay 5, 2006
  18. Jakub NarebskiMay 6, 2006
  19. Junio C HamanoMay 5, 2006
  20. Martin LanghoffMay 6, 2006
  21. Junio C HamanoMay 6, 2006
  22. Jakub NarebskiMay 6, 2006
  23. Junio C HamanoMay 6, 2006
  24. Bertrand JacquinMay 6, 2006
  25. Olivier GalibertMay 5, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.