threads / announce / 4064

Re: [ANNOUNCE] Git wiki

Subject: Re: [ANNOUNCE] Git wiki

## tl;dr

25 messages between May 5, 2006 and May 6, 2006.

replies: 24people: 10as markdown or json

linux@horizon.com· May 5, 2006, 00:56 UTC · lore

Actually, AFAICT from looking at the mailing list history, it's not dirty politics: the tie-breaker was the support and enthusiasm of the mercurial developers. It passed with only minor comment on the git mailing list, but it was a Big Thing to the hg folks.

There are ups and downs. OpenSolaris is definitely the big fish in the mercurial pond (that wasn't *meant* to sound like a recipe for heavy metal toxicity), and will get lots of attention, but git has more real-world experience. The big fish in the git pond is Linus and Linux.

In any case, mercurial and git are really very similar, far closer to each other than any third system, so it's not like the decision is a descent into heresy. Hopefully some useful cross-pollination can occur, and converting history from one to the other would be simple if anyone ever wanted to.

As for explicit renames, people are confused on the subject. IMHO, the two most revolutionary things about git are:

- Finally, a complete break from file-oriented history.  History is made
  of trees, and trees are made of files.  There is no direct connection
  between files in different commits.
- An explicit representation of an in-progress merge.
  This is what makes multiple merge strategies easily implementable.
Third, I suppose, is the raw diff format and the diffcore pipeline.

But finally getting away from the SCCS & RCS idea that the file is the unit of history is one of git's Great Features, and it shouldn't be thrown away.

What people who are asking for explicit rename tracking actually want is automatic rename merging. If branch A renames a file, and branch B corrects a typo on a comment somewhere, they'd like the merge to both patch and rename the file. If you can do that, you have met the need, even if your solution isn't the one the feature requester imagined.

(This is the general consulting problem: a client calls when they've been trying a solution and can't get past some problem. Usually, this is because they've wandered into a blind alley, and what they're asking for is either far more difficult than necessary, or will just lead them into greater problems. The first thing you have to determine is what they actually want to do, as distinct from how they've decided to do it.)

But, as Linus has pointed out, this is a very partial solution which introduces a lot of difficulties elsewhere. File renaming is a subset of the general class of code reorganizations. Source files will be split, merged, and have functions moved back and forth. You want the patch to find the code it applies to even if that code was moved.

And that can be done by taking a more global view of the patch. Identical file names is only a heuristic. If the hunk on branch A can't find a place to apply on the same file in branch B, then you have to look a little harder, either at changes from branch B that introduce matching code elsewhere, or perhaps looking through history for a change that removed the match from the obvious place to see if it added a match elsewhere.

The one thing that makes this difficult is git-read-tree's automatic collapse of "trivial" merges. If branch B moves foo() unchanged from x.c to y.c, while branch A doesn't touch y.c, but edits foo() in x.c, git-read-tree will collapse the changes to y.c before even invoking the advanced resolve script.

(The solution might be to keep *four* versions of the file in the index: the three pre-merge, *and* the post-merge. Then git-write-tree makes sure everything has a stage 0 entry and strips out the stage 1, 2 and 3 entries. This way, one merge algorithm can use another as a subroutine but decide not to accept something it did.)

But anyway, it's the merging that's the desired feature. Explicitly recording renames is only the means to that end, and is superfluous if there's another way of getting there. (And the place to look for interesting new ideas in that area Darcs.)

Fredrik Kuivinen· May 5, 2006, 06:22 UTC · re: linux@horizon.com · lore
On Thu, May 04, 2006 at 08:56:59PM -0400, linux@horizon.com wrote:
Show 6 quoted lines
> What people who are asking for explicit rename tracking actually want
> is automatic rename merging.  If branch A renames a file, and branch B
> corrects a typo on a comment somewhere, they'd like the merge to
> both patch and rename the file.  If you can do that, you have met the
> need, even if your solution isn't the one the feature requester
> imagined.

I don't know if you already know this, if you do it might be valuable for other readers.

If the rename is detected by the current rename detection code (git-diff-tree -M) then the merge case described above is handled perfectly fine by the current git. That is, the rename is followed and the patch fixing the typo is applied to the renamed file. This assumes that the default merge strategy (recursive) is used.

- Fredrik
Jakub Narebski· May 5, 2006, 06:26 UTC · re: Fredrik Kuivinen · lore
Fredrik Kuivinen wrote:
Show 16 quoted lines
> On Thu, May 04, 2006 at 08:56:59PM -0400, linux@horizon.com wrote:
>> What people who are asking for explicit rename tracking actually want
>> is automatic rename merging.  If branch A renames a file, and branch B
>> corrects a typo on a comment somewhere, they'd like the merge to
>> both patch and rename the file.  If you can do that, you have met the
>> need, even if your solution isn't the one the feature requester
>> imagined.
> 
> I don't know if you already know this, if you do it might be valuable
> for other readers.
> 
> If the rename is detected by the current rename detection code
> (git-diff-tree -M) then the merge case described above is handled
> perfectly fine by the current git. That is, the rename is followed and
> the patch fixing the typo is applied to the renamed file. This assumes
> that the default merge strategy (recursive) is used.

And if you do 'commit - rename, no changes - commit' sequence then rename will be detected.

-- 
Jakub Narebski
Warsaw, Poland
Petr Baudis· May 5, 2006, 09:23 UTC · re: Fredrik Kuivinen · lore

Dear diary, on Fri, May 05, 2006 at 08:22:36AM CEST, I got a letter where Fredrik Kuivinen <freku045@student.liu.se> said that...

Show 16 quoted lines
> On Thu, May 04, 2006 at 08:56:59PM -0400, linux@horizon.com wrote:
> > What people who are asking for explicit rename tracking actually want
> > is automatic rename merging.  If branch A renames a file, and branch B
> > corrects a typo on a comment somewhere, they'd like the merge to
> > both patch and rename the file.  If you can do that, you have met the
> > need, even if your solution isn't the one the feature requester
> > imagined.
> 
> I don't know if you already know this, if you do it might be valuable
> for other readers.
> 
> If the rename is detected by the current rename detection code
> (git-diff-tree -M) then the merge case described above is handled
> perfectly fine by the current git. That is, the rename is followed and
> the patch fixing the typo is applied to the renamed file. This assumes
> that the default merge strategy (recursive) is used.

But the non-obviously important part here to note is that the branch B merely "corrects a typo on a comment somewhere" - the latest versions in branch A and branch B are always compared for renames, therefore if branch A renamed the file and branch B sums up to some larger-scale changes in the file, it still won't be merged properly.

-- 
				Petr "Pasky" Baudis
Stuff: http://pasky.or.cz/
Right now I am having amnesia and deja-vu at the same time.  I think
I have forgotten this before.
Junio C Hamano· May 5, 2006, 09:51 UTC · re: Petr Baudis · lore
Petr Baudis <pasky@suse.cz> writes:
Show 5 quoted lines
> But the non-obviously important part here to note is that the branch B
> merely "corrects a typo on a comment somewhere" - the latest versions in
> branch A and branch B are always compared for renames, therefore if
> branch A renamed the file and branch B sums up to some larger-scale
> changes in the file, it still won't be merged properly.

I probably am guilty of starting this misinformation, but the code does not compare the latest in A and B for rename detection; it compares (O, A) and (O, B).

But the end result is the same - what you say is correct. If a path (say O to A) that renamed has too big a change, then no matter how small the changes are on the other path (O to B), rename detection can be fooled. We could perhaps alleviate it by following the whole commit chain.

Petr Baudis· May 5, 2006, 16:40 UTC · re: Junio C Hamano · lore

Dear diary, on Fri, May 05, 2006 at 11:51:01AM CEST, I got a letter where Junio C Hamano <junkio@cox.net> said that...

Show 11 quoted lines
> Petr Baudis <pasky@suse.cz> writes:
> 
> > But the non-obviously important part here to note is that the branch B
> > merely "corrects a typo on a comment somewhere" - the latest versions in
> > branch A and branch B are always compared for renames, therefore if
> > branch A renamed the file and branch B sums up to some larger-scale
> > changes in the file, it still won't be merged properly.
> 
> I probably am guilty of starting this misinformation, but the
> code does not compare the latest in A and B for rename
> detection; it compares (O, A) and (O, B).

Where O = LCA(A,B) (modulo recursiveness)? Yes, that is what I meant to say but I phrased it wrong, sorry.

-- 
				Petr "Pasky" Baudis
Stuff: http://pasky.or.cz/
Right now I am having amnesia and deja-vu at the same time.  I think
I have forgotten this before.
Jakub Narebski· May 5, 2006, 16:47 UTC · re: Junio C Hamano · lore
Junio C Hamano wrote:
Show 17 quoted lines
> Petr Baudis <pasky@suse.cz> writes:
> 
>> But the non-obviously important part here to note is that the branch B
>> merely "corrects a typo on a comment somewhere" - the latest versions in
>> branch A and branch B are always compared for renames, therefore if
>> branch A renamed the file and branch B sums up to some larger-scale
>> changes in the file, it still won't be merged properly.
> 
> I probably am guilty of starting this misinformation, but the
> code does not compare the latest in A and B for rename
> detection; it compares (O, A) and (O, B).
> 
> But the end result is the same - what you say is correct.  If a
> path (say O to A) that renamed has too big a change, then no
> matter how small the changes are on the other path (O to B),
> rename detection can be fooled.  We could perhaps alleviate it
> by following the whole commit chain.

Or perhaps by helper information about renames, entered either by git-mv (and git-cp) or rename detection at commit, e.g. in the following form

        note at <commit-sha1> was-in <pathname>
        note at <commit-sha1> was-in <pathname>

(with the obvious limit of this "note header" solution is that it wouldn't work for filenames and directory name containing "\n"). I'm not sure if <pathname> should be just basename, of full pathname.

-- 
Jakub Narebski
Warsaw, Poland
Jakub Narebski· May 5, 2006, 18:49 UTC · re: Jakub Narebski · lore
Jakub Narebski wrote:
Show 29 quoted lines
> Junio C Hamano wrote:
> 
>> Petr Baudis <pasky@suse.cz> writes:
>> 
>>> But the non-obviously important part here to note is that the branch B
>>> merely "corrects a typo on a comment somewhere" - the latest versions in
>>> branch A and branch B are always compared for renames, therefore if
>>> branch A renamed the file and branch B sums up to some larger-scale
>>> changes in the file, it still won't be merged properly.
>> 
>> I probably am guilty of starting this misinformation, but the
>> code does not compare the latest in A and B for rename
>> detection; it compares (O, A) and (O, B).
>> 
>> But the end result is the same - what you say is correct.  If a
>> path (say O to A) that renamed has too big a change, then no
>> matter how small the changes are on the other path (O to B),
>> rename detection can be fooled.  We could perhaps alleviate it
>> by following the whole commit chain.
> 
> Or perhaps by helper information about renames, entered either by git-mv
> (and git-cp) or rename detection at commit, e.g. in the following form
> 
>         note at <commit-sha1> was-in <pathname>
>         note at <commit-sha1> was-in <pathname>
> 
> (with the obvious limit of this "note header" solution is that it wouldn't
> work for filenames and directory name containing "\n"). I'm not sure if
> <pathname> should be just basename, of full pathname.

Erm, I'm sorry, forget the implementation which wouldn't work. The idea was to accumulate renames and contents moving information, and remember at which commit it occured. But it's place (as a _helper_ information) is perhaps in separate structure.

-- 
Jakub Narebski
Warsaw, Poland
Petr Baudis· May 5, 2006, 16:36 UTC · re: linux@horizon.com · lore

Dear diary, on Fri, May 05, 2006 at 02:56:59AM CEST, I got a letter where linux@horizon.com said that...

Show 15 quoted lines
> Actually, AFAICT from looking at the mailing list history, it's not dirty
> politics: the tie-breaker was the support and enthusiasm of the mercurial
> developers.  It passed with only minor comment on the git mailing list,
> but it was a Big Thing to the hg folks.
> 
> There are ups and downs.  OpenSolaris is definitely the big fish in
> the mercurial pond (that wasn't *meant* to sound like a recipe for
> heavy metal toxicity), and will get lots of attention, but git has more
> real-world experience.  The big fish in the git pond is Linus and Linux.
> 
> In any case, mercurial and git are really very similar, far closer
> to each other than any third system, so it's not like the decision is
> a descent into heresy.  Hopefully some useful cross-pollination
> can occur, and converting history from one to the other would be
> simple if anyone ever wanted to.

It's a philosophical question here, but I'd say that Git is much closer to Monotone than to any other version control system - I think it can be described as Monotone model with more elegant implementation (for some, at least ;), no certificates and restriction of one head per branch. And another important difference is that Monotone has persistent file identifiers, but I think that's about the only thing that would make Monotone more "file orientated".

I'm not much of a Mercurial pro but it appears to me that the architectural differences there are larger, especially wrt. the revlogs and wholly quite a more file-oriented model.

-- 
				Petr "Pasky" Baudis
Stuff: http://pasky.or.cz/
Right now I am having amnesia and deja-vu at the same time.  I think
I have forgotten this before.
Linus Torvalds· May 5, 2006, 17:48 UTC · re: Petr Baudis · lore
On Fri, 5 May 2006, Petr Baudis wrote:
> 
> It's a philosophical question here, but I'd say that Git is much closer
> to Monotone than to any other version control system
Some historical background..

Before I dropped BK, I ended up being involved in trying to get Larry and Tridge to come to some agreement about how to solve the issues Tridge had with BK not being open-source. That actually went on for maybe two months or so, and I kept on hoping that we'd find some acceptably middle ground.

I thought we could find somethign that would actually work for everybody: to hopefully both make BK technically better, _and_ to make the end result more palatable to the "free software or bust" contingency.

One of the suggestions that I tried to push as an acceptable middle ground was to make a "generic" BK repository export format, so that people who didn't want to use BK could still get all the information, and not in a broken format like CVS (yes, CVS makes sense as an interchange format, since _everybody_ speaks CVS, but it's a horrible, horrible, horrible format from any technical standpoint).

My example export format was really a strange mixture of patches with parenthood information, where the history information was described with hashes (MD5 rather than SHA1, but that was just an implementation thing, and mostly because BK used MD5 sums). Not something really useful as a real SCM, but it wasn't designed for that - it was just meant to be a useful and unambiguous interoperability format.

Now, that didn't work out, and I was a little bummed. I thought it would have made both sides happy, because it would actually have been a better format than CVS (and yes, I'm somewhat biased: in my opinion, having a million monkeys throwing crap at the walls and encoding the information in the patterns on monkey shit is a better format than CVS), so it would actually have improved BK, while also making it possible to interoperate if you didn't want to use BK itself.

But Tridge didn't believe that it would actually have exported all the information in a BK tree, even if both I and Larry told him it would. I'm not a hundred percent sure that Larry would have gone for the export format either, but hey, one sign of a good compromise is that neither side really gets what they really want. Whatever. It didn't work.

So it didn't actually resolve the deadlock, but when it became clear that I couldn't work with BK any more, I thought I might use something like that "patch + parenthood" representation as a way to maintain my tree while looking at other alternatives.

So in many ways, when I started looking around for distributed SCM's, I came into the game with the background of keeping the history around as chains of hashes describing it, and then just having patches to describe the differences between versions.

So that was really my "fallback" position: if nothing out there worked, I'd rather go back to lists of patches than use CVS.

Now, if you keep track of just patches, one of the issues is that you can't afford to re-create the tree every time by walking patches forward from the beginning, so I also was planning to have an "cache" that maintained the current state of the tree as a separate state from the working tree, so that I would always have the "working tree" and the "result of patches up to this moment" as two separate things (so that I could do the "bk diff" that I was used to doing to see the difference between my last state and the current state of the working tree).

In other words, I was already working on the git "index" file. And I was planning to just have a patch-based system behind it, with a hashed history. Kind of "quilt with history and an index to speed things up".

The index itself would be backed-up with whole files (all hidden in the ".dircache" directory), and the patch series would thus normally never actually be _used_. So the inefficiency of working with patches would never be much of an issue. A "commit" would create a new patch from the current working directory and the previous shadow tree, and update the shadow tree and add a new entry to the history list.

And then I found Monotone.

Now, monotone was slow. Monotone was so _horrendously_ slow that I had to do special hacks just to import _one_ version of Linux into it in less than two hours. It was something stupid like an O(N**3) algorithm in the number of filenames (and the kernel had 17,291 files at that time: v2.6.12-rc2), and it was just totally unusable for me.

I also thought (and still think) that the whole signing thing was a waste of time and misdesigned, and I obviously am not a huge fan of databases. So in many ways I disliked the monotone implementation decisions (and some of its design decisions). But at the same time, I immediately liked the SHA1 object naming concept of Monotone.

It also already matched how I had conceptually planned on doing on the history anyway, and had some ideas for, but it took that whole "history hashing" all the way.

And thus git was born. 

So git really has three parents. In a very real sense, BK (or, perhaps more appropriately - the way I personally used BK, which is not necessarily how others have used it) was the biggest thing from the standpoint of what I wanted my _workflow_ to be like. It was simply how I had done things for the last few years, so a lot of my mental model for how things are supposed to _work_ came from BK.

I still don't think people give Larry enough credit for actually pushing this whole distributed SCM thing as a _usable_ model. Very few of the open-source distributed SCM's are actually usable even today, and as far as I've been able to gather, the commercial ones aren't really any closer either. Larry didn't have the kind of examples of what _can_ work that I had.

The other parent was the stupid "series of patches" model, which was what really resulted in the "index" thing. I realize that people don't always much like the index, but it's really a pretty central part of git history, and one of the distinguising marks of git. It may be trivial, and to some degree it's been overshadowed by all the tree operations we do (the combination of revision walking and tree diffing), but it was very central to how git came to be.

The index also ended up being central to how we did merges - even if some day we may end up doing more of that on a pure tree level (ie the current git-merge-tree model), I think the way we ended up doing merges owes a lot to the index as a staging area.

(Historically, the "index" was called the "cache". Exactly because it came from the notion of "caching" the top commit state in a patch series, and then working with patches either backwards or forwards from that top cached state. Similarly, we didn't have a ".git" directory: it was called ".dircache", exactly because it was all about caching the state of the previous commit directory layout).

And finally, Monotone for the "everything is an object named by its SHA1" model, which to some degree is perhaps the central - or at least the most obvious - part of git. It largely was designed really just to be the "backing store" for the "cache", and to not be _that_ important. That also explains why I didn't worry too much about disk usage etc initially: the object store wasn't even the most important part, and I envisioned just moving old objects that weren't needed into some "backup storage" kind of thing.

			Linus
Dave Jones· May 5, 2006, 19:04 UTC · re: Linus Torvalds · lore
On Fri, May 05, 2006 at 10:48:38AM -0700, Linus Torvalds wrote:
 > (and yes, I'm somewhat biased: in my opinion, having a 
 > million monkeys throwing crap at the walls and encoding the information in 
 > the patterns on monkey shit is a better format than CVS), so it would 
 > actually have improved BK, while also making it possible to interoperate 
 > if you didn't want to use BK itself.
 >  ...
 > So that was really my "fallback" position: if nothing out there worked, 
 > I'd rather go back to lists of patches than use CVS. 

I've encountered managing kernel trees in CVS both during my tenure at SuSE, and to a more involved extent as Fedora/RHEL maintainer, and I'd just like to echo how much it _completely sucks_ at times.

Rebasing to a newer release is a *nightmare* that usually takes up most of an afternoon compared to rebasing my git based projects.

In the event I can't persuade the powers at be to switch to git at some point for managing our packages, I'll be sure to bring up your suggestion of a million monkeys. I believe you can pick them up fairly cheap these days.

		Dave
-- 
http://www.codemonkey.org.uk
Petr Baudis· May 5, 2006, 18:15 UTC · re: linux@horizon.com · lore

Dear diary, on Fri, May 05, 2006 at 02:56:59AM CEST, I got a letter where linux@horizon.com said that...

Show 13 quoted lines
> But, as Linus has pointed out, this is a very partial solution which
> introduces a lot of difficulties elsewhere.  File renaming is a subset of
> the general class of code reorganizations.  Source files will be split,
> merged, and have functions moved back and forth.  You want the patch to
> find the code it applies to even if that code was moved.
> 
> And that can be done by taking a more global view of the patch.
> Identical file names is only a heuristic.  If the hunk on branch A
> can't find a place to apply on the same file in branch B, then
> you have to look a little harder, either at changes from branch B
> that introduce matching code elsewhere, or perhaps looking
> through history for a change that removed the match from the
> obvious place to see if it added a match elsewhere.

There are really two distinctions here which should be kept separate: automatic vs. explicit movement tracking and file-level vs. subfile-level movement tracking.

The automatic vs. explicit movement tracking is a lot more
controversial. Explicit movement tracking is pretty easy to provide for
file-level movements, it's just that the user says "I _did_ move file
A to file B" (I never got the Linus' argument that the user has no idea
- he just _performed_ the move, also explicitly, by calling *mv).

However, I guess the explicit movement tracking completely fails if you go sub-file (without being extremely bothersome for the user) - you would have to have control over the editor and the clipboard and even then I'm not sure if you could reach any sensible results.

I still dislike automated movement tracking for whole files, but I'm conciliated with it. Because it is probably the only really sensible way to implement subfile-level tracking. It would not be hard to implement using pickaxe (actually, I believe it was near the top of Junio's TODO few weeks ago) and a similarity detector comparing new and old version (if it's dissimilar enough, check if that or a similar hunk was not added somewhere else in the same commit; well, at least the idea sounds simple).

One obvious problem are ambiguities - several similar files are renamed to other similar files and now how do you decide which version to choose? Merge the change to all the new files? Only to some? Panic? I wonder how does the current recursive strategy deal with that. Of course, this case sounds quite artificial and rare for whole files, but I suspect that it will be much more common once you do not deal with files but just hunks, moving bits of code around.

-- 
				Petr "Pasky" Baudis
Stuff: http://pasky.or.cz/
Right now I am having amnesia and deja-vu at the same time.  I think
I have forgotten this before.
Petr Baudis· May 5, 2006, 18:20 UTC · re: Petr Baudis · lore

Dear diary, on Fri, May 05, 2006 at 08:15:41PM CEST, I got a letter where Petr Baudis <pasky@suse.cz> said that...

> There are really two distinctions here which should be kept separate:
> automatic vs. explicit movement tracking and file-level vs.
> subfile-level movement tracking.

I should have revised this paragraph before sending the mail out, I ended up sorting out my thoughts on the subject as I wrote the mail. The two aspects end up so tied that it makes sense to mingle them. Examining them separately here still hopefully shed some light on possible reasoning behind the Git design decisions.

-- 
				Petr "Pasky" Baudis
Stuff: http://pasky.or.cz/
Right now I am having amnesia and deja-vu at the same time.  I think
I have forgotten this before.
Jakub Narebski· May 5, 2006, 18:27 UTC · re: Petr Baudis · lore
Petr Baudis wrote:
Show 10 quoted lines
> The automatic vs. explicit movement tracking is a lot more
> controversial. Explicit movement tracking is pretty easy to provide for
> file-level movements, it's just that the user says "I _did_ move file
> A to file B" (I never got the Linus' argument that the user has no idea
> - he just _performed_ the move, also explicitly, by calling *mv).
> 
> However, I guess the explicit movement tracking completely fails if you
> go sub-file (without being extremely bothersome for the user) - you
> would have to have control over the editor and the clipboard and even
> then I'm not sure if you could reach any sensible results.

If I remember correctly there are some problems if the explicit file-level contents movement tracking (aka. file rename tracking) is done via equivalent of file-id, inodes, or persistent names. Although it works for many (most?) cases.

-- 
Jakub Narebski
Warsaw, Poland
Linus Torvalds· May 5, 2006, 18:31 UTC · re: Petr Baudis · lore
On Fri, 5 May 2006, Petr Baudis wrote:
Show 6 quoted lines
> 
> The automatic vs. explicit movement tracking is a lot more
> controversial. Explicit movement tracking is pretty easy to provide for
> file-level movements, it's just that the user says "I _did_ move file
> A to file B" (I never got the Linus' argument that the user has no idea
> - he just _performed_ the move, also explicitly, by calling *mv).
THE USER DID NO SUCH THING.
Moving data around happens with a whole lot more than "mv".

It happens with patches (somebody _else_ may have done an "mv", without using git at all), and it happens with editors (moving data around until most of it exists in another file).

So doing "*mv" is just a special case.

And supporting special cases is _wrong_. If you start depending on data that isn't actually dependable, that's WRONG.

There's another reason why encoding movement information in the commit is totally broken, namely the fact that a lot of the actions DO NOT WALK THE COMMIT CHAIN!

Try doing
	git diff v1.3.0..

and think about what that actually _means_. Think about the fact that it doesn't actually walk the commit chain at all: it diffs the trees between v1.3.0 and the current one. What if the rename happened in a commit in the middle?

The "track contents, not intentions" approach avoids both these things. The end result is _reliable_, not a "random guess".

Adding file movement note to commits is simply WRONG.

Why does this come up every three months or so? I was right the first time. You'd think that as time passes, people would just notice more and more how right I was and am, instead of forgetting and bringing this idiotic idea up over and over and over again.

		Linus
Petr Baudis· May 5, 2006, 18:54 UTC · re: Linus Torvalds · lore

Dear diary, on Fri, May 05, 2006 at 08:31:06PM CEST, I got a letter where Linus Torvalds <torvalds@osdl.org> said that...

> Moving data around happens with a whole lot more than "mv".

Let's keep this on the per-file level - if you want to go below the file granularity, I already _DID_ say that I agree that explicit tracking is not a way. (If sub-file tracking would end up having any usable reliability in real-world cases, which is something I do not take for granted.)

Another thing is, the sub-file content tracking would end up being a lot more "magic" than the simple per-file content tracking, and you stated several times that you prefer simple merge over better but magic merge - so why do you prefer sub-file content tracking anyway?

> It happens with patches (somebody _else_ may have done an "mv", without 
> using git at all),

_Here_ is the place for automated renames detection. Between applying and committing the patch, the user can verify that it got the renames right. That's impossible when guessing the renames later.

> and it happens with editors (moving data around until 
> most of it exists in another file).

I doubt this in fact happens that often (to a degree the automatic rename detection would catch). And if it happens, then the user has to tell Git - I have never heard that _this_ would be any problem in other version control systems. You could make it more foolproof by running the automatic rename detection on the diff being committed and suggesting the user that other yet unrecorded renames did happen.

The point is, the user stays in control and can override any stupid guess.
> So doing "*mv" is just a special case.
> 
> And supporting special cases is _wrong_. If you start depending on data 
> that isn't actually dependable, that's WRONG.

I prefer making this data dependable to having to resort to guessing on dependable less amount of data.

Show 12 quoted lines
> There's another reason why encoding movement information in the commit is 
> totally broken, namely the fact that a lot of the actions DO NOT WALK THE 
> COMMIT CHAIN!
> 
> Try doing
> 
> 	git diff v1.3.0..
> 
> and think about what that actually _means_. Think about the fact that it 
> doesn't actually walk the commit chain at all: it diffs the trees between 
> v1.3.0 and the current one. What if the rename happened in a commit in the 
> middle?

Then the automated renames detection will miss it given that the other accumulated differences are large enough, and the suggested workarounds _are_ precisely walking the commit chain.

If you use persistent file ids, you never miss it _AND_ you DO NOT WALK THE COMMIT CHAIN! You still just match file ids in the two trees.

> The "track contents, not intentions" approach avoids both these things. 
> The end result is _reliable_, not a "random guess".

No, the end result is whichever some heuristic randomly guessed, and it's not reliable either since the heuristic can change.

-- 
				Petr "Pasky" Baudis
Stuff: http://pasky.or.cz/
Right now I am having amnesia and deja-vu at the same time.  I think
I have forgotten this before.
Jakub Narebski· May 5, 2006, 19:39 UTC · re: Petr Baudis · lore
Petr Baudis wrote:
> Dear diary, on Fri, May 05, 2006 at 08:31:06PM CEST, I got a letter
> where Linus Torvalds <torvalds@osdl.org> said that...
Show 22 quoted lines
> I prefer making this [rename detection] data dependable to having to
> resort to guessing on dependable less amount of data.
> 
>> There's another reason why encoding movement information in the commit is
>> totally broken, namely the fact that a lot of the actions DO NOT WALK THE
>> COMMIT CHAIN!
>> 
>> Try doing
>> 
>> git diff v1.3.0..
>> 
>> and think about what that actually _means_. Think about the fact that it
>> doesn't actually walk the commit chain at all: it diffs the trees between
>> v1.3.0 and the current one. What if the rename happened in a commit in
>> the middle?
> 
> Then the automated renames detection will miss it given that the other
> accumulated differences are large enough, and the suggested workarounds
> _are_ precisely walking the commit chain.
> 
> If you use persistent file ids, you never miss it _AND_ you DO NOT WALK
> THE COMMIT CHAIN! You still just match file ids in the two trees.

Let not jump to the one of the possible solution. The detecting and noting renames and content moving (with user interaction) at commit is nice... unless does something which cannot allow interactiveness (like applying patchbomb), but even then detecting and saving info at commit would be good idea.

What we need is to for two given linked revisions (with a path between them) to easily extract information about renames (content moving). Perhaps using additional structure... best if we could do this without walking the chain. The rest is details... ;-P

-- 
Jakub Narebski
Warsaw, Poland
Jakub Narebski· May 6, 2006, 13:37 UTC · re: Jakub Narebski · lore
Jakub Narebski wrote:
> Petr Baudis wrote:
Show 13 quoted lines
>> If you use persistent file ids, you never miss it _AND_ you DO NOT WALK
>> THE COMMIT CHAIN! You still just match file ids in the two trees.
> 
> Let not jump to the one of the possible solution. The detecting and noting
> renames and content moving (with user interaction) at commit is nice...
> unless does something which cannot allow interactiveness (like applying
> patchbomb), but even then detecting and saving info at commit would be
> good idea.
> 
> What we need is to for two given linked revisions (with a path between
> them) to easily extract information about renames (content moving).
> Perhaps using additional structure... best if we could do this without
> walking the chain. The rest is details... ;-P

Or rather structure, which for given file F in given revision A, for given other revision B would tell ALL the files in the revision B which are source of contents (via history/commit tree) of the file F.

-- 
Jakub Narebski
Warsaw, Poland
Junio C Hamano· May 5, 2006, 19:49 UTC · re: Petr Baudis · lore
Petr Baudis <pasky@suse.cz> writes:
> I doubt this in fact happens that often (to a degree the automatic
> rename detection would catch). And if it happens, then the user has to
> tell Git - I have never heard that _this_ would be any problem in other
> version control systems.

It does not become an issue only because users accept it as a fact of life. When Linus was moving most of the contents in rev-list.c to create a new revision.c, I already had some tweaks to rev-list.c published before he sent me a patch for the code movement, and I am sure he needed to re-roll the patch by merging the change I did to rev-list.c back into his revision.c file. No SCM may handle that automatically, and no user accustomed to existing SCM (including git) expect that to work automatically. But that does not necessarily mean a tool that notices it and tells user what is going on is a bad thing.

However it is a different story to try recording "what is going on" whether it comes from the tool's guess or directly from the user.

Having a way to affect the inprecise "guess" the tool makes when that guesswork is needed might make sense. If you (think you) know arch/i386/foo.h was copied to create arch/x86-64/foo.h but the detector does not detect it and seeing a creation patch for arch/x86-64/foo.h frustrates you, you may want to have a way to explicitly say "compare arch/i386/foo.h with arch/x86-64/foo.h in that commit -- I want to examine the change needed to adjust foo to x86-64 architecture".

But we have "git diff v2.6.14:arch/i386/foo.h v2.6.14:arch/x86-64/foo.h" for that ;-).

> Then the automated renames detection will miss it given that the other
> accumulated differences are large enough, and the suggested workarounds
> _are_ precisely walking the commit chain.

The HEAD may _not_ have anything to do with v1.3.0 in which case you would get nothing from walking the ancestry.

> If you use persistent file ids, you never miss it _AND_ you DO NOT WALK
> THE COMMIT CHAIN! You still just match file ids in the two trees.
It is unworkable.

Which one should inherit the persistent id of the old rev-list.c? New rev-list.c, or revision.c that has most of the old contents split out?

Oh, and did you know there was a different revision.h that is not related to the current revision.h in the history of git? Should its persistent id have any relation with the persistent id of the current revision.h? When would you decide to make the id inherited and when not to? If I remove revision.h by mistake in a commit and resurrect it in the next commit, should it get the same id back? If I forget to tell the tool that those two "disappeared and then reappeared" are related and should get the same persistent id when I make the resurrection commit, and keep piling other commits on top, do I have to rewind the ancestry chain all the way to correct the mistake?

Martin Langhoff· May 6, 2006, 06:53 UTC · re: Junio C Hamano · lore
On 5/6/06, Junio C Hamano <junkio@cox.net> wrote:
> > If you use persistent file ids, you never miss it _AND_ you DO NOT WALK
> > THE COMMIT CHAIN! You still just match file ids in the two trees.
>
> It is unworkable.

+1 -- explicit file ids are evil. Arch/TLA demonstrated that amply... they are a serious annoyance to the end user, they have a lot of not-elegantly solvable cases (same file created with the same contents in several repos -- say via an emailed patch) that git gets right _today_.

They _are_ useful in a very small set of cases -- namely in the case of a naive mv, which git handles correctly today. Subtler things git sometimes does right, sometimes fails, but it can be made to be much smarter by interpreting content changes better, for instance all this talk about getting pickaxe to guess where the patch should be applied for a file that got split into 3.

But those subtler cases are totally impossible with explicit id tracking. I used Arch for a long time with very large trees, and renames coming left, right and centre. Explicit ids didn't help much, and the number of manual fixups we had to do was awful.

I am using GIT with the very same project, and just now, typing this, I realised that there are still many renames happening in the project. I had forgotten about it -- well, not really: I do use git-merge instead of cg-merge when I suspect there may be interesting cases ;-)

Of course, YMMV, and I have to confess I was a sceptic for a while... but now as an end-user dealing with messy projects, I say LIRAR: Linus Is Right About Renames.

OTOH,
Show 12 quoted lines
>> Try doing
>>
>> git diff v1.3.0..
>>
>> and think about what that actually _means_. Think about the fact that it
>> doesn't actually walk the commit chain at all: it diffs the trees between
>> v1.3.0 and the current one. What if the rename happened in a commit in
>> the middle?
>
> Then the automated renames detection will miss it given that the other
> accumulated differences are large enough, and the suggested workarounds
> _are_ precisely walking the commit chain.

I agree here with Pasky that after a while the automated renames/copy/splitup detection will miss the operation in cases where it would be interesting to note it to the user. IIRC git-rerere is the tool that knows about this (still voodoo to me how) and could be used to help here. At what (runtime) cost, I don't know, but that kind of walking history to tell me more interesting things about the diff is something that is usually worthwhile.

Usual disclaimers apply.
martin
Junio C Hamano· May 6, 2006, 07:14 UTC · re: Martin Langhoff · lore
"Martin Langhoff" <martin.langhoff@gmail.com> writes:
Show 7 quoted lines
> I agree here with Pasky that after a while the automated
> renames/copy/splitup detection will miss the operation in cases where
> it would be interesting to note it to the user. IIRC git-rerere is the
> tool that knows about this (still voodoo to me how) and could be used
> to help here. At what (runtime) cost, I don't know, but that kind of
> walking history to tell me more interesting things about the diff is
> something that is usually worthwhile.
FYI rerere is a totally unrelated voodoo.

It remembers the conflict marker pattern <<< === >>> immediately after it runs "merge" (ah, that reminds me -- I should replace them with diff3), and then remembers the result of the manual resolution just before the user makes a commit. Then, when next time it runs "merge" for something and notices <<< === >>> pattern it has seen before, it runs a three-way merge between the previous resolution result and the current conflicted state, using the previous conflicted state as the common origin.

Jakub Narebski· May 6, 2006, 07:33 UTC · re: Martin Langhoff · lore
Martin Langhoff wrote:
> On 5/6/06, Junio C Hamano <junkio@cox.net> wrote:
Show 16 quoted lines
>>> Try doing
>>>
>>> git diff v1.3.0..
>>>
>>> and think about what that actually _means_. Think about the fact that it
>>> doesn't actually walk the commit chain at all: it diffs the trees
>>> between v1.3.0 and the current one. What if the rename happened in a
>>> commit in the middle?
>>
>> Then the automated renames detection will miss it given that the other
>> accumulated differences are large enough, and the suggested workarounds
>> _are_ precisely walking the commit chain.
> 
> I agree here with Pasky that after a while the automated
> renames/copy/splitup detection will miss the operation in cases where
> it would be interesting to note it to the user.
Perhaps an option to do rename detection with walking the commit chain?
-- 
Jakub Narebski
Warsaw, Poland
Bertrand Jacquin· May 6, 2006, 12:46 UTC · re: Junio C Hamano · lore
On 5/6/06, Junio C Hamano <junkio@cox.net> wrote:
Show 5 quoted lines
> Jakub Narebski <jnareb@gmail.com> writes:
>
> > Perhaps an option to do rename detection with walking the commit chain?
>
> Have fun implementing that ;-).

I agree that it could be interesting to have a such thing. But that's increndibly stupid and moreover a rare case.

-- Beber #e.fr@freenode

Olivier Galibert· May 5, 2006, 20:45 UTC · re: Petr Baudis · lore
On Fri, May 05, 2006 at 08:15:41PM +0200, Petr Baudis wrote:
Show 5 quoted lines
> The automatic vs. explicit movement tracking is a lot more
> controversial. Explicit movement tracking is pretty easy to provide for
> file-level movements, it's just that the user says "I _did_ move file
> A to file B" (I never got the Linus' argument that the user has no idea
> - he just _performed_ the move, also explicitly, by calling *mv).

In one of my projects 99% or the renames are "done" when unzipping the source release of the next version. Explicit tracking would be unbearable, frankly.

And once you have a good enough implicit tracking, why bother with an explicit one?

  OG.

← back to recent threads