{"thread":{"id":"4064","subject":"Re: [ANNOUNCE] Git wiki","startedAt":"2006-05-05T00:56:59Z","lastAt":"2006-05-06T13:37:47Z","messageCount":25,"participants":["linux@horizon.com","Fredrik Kuivinen","Jakub Narebski","Petr Baudis","Junio C Hamano","Linus Torvalds","Dave Jones","Olivier Galibert","Martin Langhoff","Bertrand Jacquin"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"19538","messageId":"20060505005659.9092.qmail@science.horizon.com","threadId":"4064","inReplyTo":null,"subject":"Re: [ANNOUNCE] Git wiki","fromName":"","fromEmail":"linux@horizon.com","sentAt":"2006-05-05T00:56:59Z","receivedAt":"2006-05-05T00:56:59Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"Actually, AFAICT from looking at the mailing list history, it's not dirty\npolitics: the tie-breaker was the support and enthusiasm of the mercurial\ndevelopers.  It passed with only minor comment on the git mailing list,\nbut it was a Big Thing to the hg folks.\n\nThere are ups and downs.  OpenSolaris is definitely the big fish in\nthe mercurial pond (that wasn't *meant* to sound like a recipe for\nheavy metal toxicity), and will get lots of attention, but git has more\nreal-world experience.  The big fish in the git pond is Linus and Linux.\n\nIn any case, mercurial and git are really very similar, far closer\nto each other than any third system, so it's not like the decision is\na descent into heresy.  Hopefully some useful cross-pollination\ncan occur, and converting history from one to the other would be\nsimple if anyone ever wanted to.\n\n\nAs for explicit renames, people are confused on the subject.\nIMHO, the two most revolutionary things about git are:\n\n- Finally, a complete break from file-oriented history.  History is made\n  of trees, and trees are made of files.  There is no direct connection\n  between files in different commits.\n- An explicit representation of an in-progress merge.\n  This is what makes multiple merge strategies easily implementable.\n\nThird, I suppose, is the raw diff format and the diffcore pipeline.\n\nBut finally getting away from the SCCS & RCS idea that the file is the\nunit of history is one of git's Great Features, and it shouldn't be\nthrown away.\n\n\nWhat people who are asking for explicit rename tracking actually want\nis automatic rename merging.  If branch A renames a file, and branch B\ncorrects a typo on a comment somewhere, they'd like the merge to\nboth patch and rename the file.  If you can do that, you have met the\nneed, even if your solution isn't the one the feature requester\nimagined.\n\n(This is the general consulting problem: a client calls when they've\nbeen trying a solution and can't get past some problem.  Usually, this\nis because they've wandered into a blind alley, and what they're asking\nfor is either far more difficult than necessary, or will just lead them\ninto greater problems.  The first thing you have to determine is what\nthey actually want to do, as distinct from how they've decided to do it.)\n\n\nBut, as Linus has pointed out, this is a very partial solution which\nintroduces a lot of difficulties elsewhere.  File renaming is a subset of\nthe general class of code reorganizations.  Source files will be split,\nmerged, and have functions moved back and forth.  You want the patch to\nfind the code it applies to even if that code was moved.\n\nAnd that can be done by taking a more global view of the patch.\nIdentical file names is only a heuristic.  If the hunk on branch A\ncan't find a place to apply on the same file in branch B, then\nyou have to look a little harder, either at changes from branch B\nthat introduce matching code elsewhere, or perhaps looking\nthrough history for a change that removed the match from the\nobvious place to see if it added a match elsewhere.\n\nThe one thing that makes this difficult is git-read-tree's automatic\ncollapse of \"trivial\" merges.  If branch B moves foo() unchanged from\nx.c to y.c, while branch A doesn't touch y.c, but edits foo() in x.c,\ngit-read-tree will collapse the changes to y.c before even invoking\nthe advanced resolve script.\n\n(The solution might be to keep *four* versions of the file in the index:\nthe three pre-merge, *and* the post-merge.  Then git-write-tree makes\nsure everything has a stage 0 entry and strips out the stage 1, 2 and\n3 entries.  This way, one merge algorithm can use another as a\nsubroutine but decide not to accept something it did.)\n\n\nBut anyway, it's the merging that's the desired feature.  Explicitly\nrecording renames is only the means to that end, and is superfluous\nif there's another way of getting there.  (And the place to look for\ninteresting new ideas in that area Darcs.)\n"},{"id":"19547","messageId":"20060505062236.GA4544@c165.ib.student.liu.se","threadId":"4064","inReplyTo":"20060505005659.9092.qmail@science.horizon.com","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Fredrik Kuivinen","fromEmail":"freku045@student.liu.se","sentAt":"2006-05-05T06:22:36Z","receivedAt":"2006-05-05T06:22:36Z","isPatch":false,"sender":{"key":"frekui@gmail.com","avatar":"https://avatars.githubusercontent.com/u/13770967?v=4"},"body":"On Thu, May 04, 2006 at 08:56:59PM -0400, linux@horizon.com wrote:\n> What people who are asking for explicit rename tracking actually want\n> is automatic rename merging.  If branch A renames a file, and branch B\n> corrects a typo on a comment somewhere, they'd like the merge to\n> both patch and rename the file.  If you can do that, you have met the\n> need, even if your solution isn't the one the feature requester\n> imagined.\n\nI don't know if you already know this, if you do it might be valuable\nfor other readers.\n\nIf the rename is detected by the current rename detection code\n(git-diff-tree -M) then the merge case described above is handled\nperfectly fine by the current git. That is, the rename is followed and\nthe patch fixing the typo is applied to the renamed file. This assumes\nthat the default merge strategy (recursive) is used.\n\n\n- Fredrik\n"},{"id":"19548","messageId":"e3er79$6s4$1@sea.gmane.org","threadId":"4064","inReplyTo":"20060505062236.GA4544@c165.ib.student.liu.se","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-05-05T06:26:52Z","receivedAt":"2006-05-05T06:26:52Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Fredrik Kuivinen wrote:\n\n> On Thu, May 04, 2006 at 08:56:59PM -0400, linux@horizon.com wrote:\n>> What people who are asking for explicit rename tracking actually want\n>> is automatic rename merging.  If branch A renames a file, and branch B\n>> corrects a typo on a comment somewhere, they'd like the merge to\n>> both patch and rename the file.  If you can do that, you have met the\n>> need, even if your solution isn't the one the feature requester\n>> imagined.\n> \n> I don't know if you already know this, if you do it might be valuable\n> for other readers.\n> \n> If the rename is detected by the current rename detection code\n> (git-diff-tree -M) then the merge case described above is handled\n> perfectly fine by the current git. That is, the rename is followed and\n> the patch fixing the typo is applied to the renamed file. This assumes\n> that the default merge strategy (recursive) is used.\n\nAnd if you do 'commit - rename, no changes - commit' sequence then rename\nwill be detected. \n\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"19552","messageId":"20060505092332.GY27689@pasky.or.cz","threadId":"4064","inReplyTo":"20060505062236.GA4544@c165.ib.student.liu.se","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2006-05-05T09:23:32Z","receivedAt":"2006-05-05T09:23:32Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Fri, May 05, 2006 at 08:22:36AM CEST, I got a letter\nwhere Fredrik Kuivinen <freku045@student.liu.se> said that...\n> On Thu, May 04, 2006 at 08:56:59PM -0400, linux@horizon.com wrote:\n> > What people who are asking for explicit rename tracking actually want\n> > is automatic rename merging.  If branch A renames a file, and branch B\n> > corrects a typo on a comment somewhere, they'd like the merge to\n> > both patch and rename the file.  If you can do that, you have met the\n> > need, even if your solution isn't the one the feature requester\n> > imagined.\n> \n> I don't know if you already know this, if you do it might be valuable\n> for other readers.\n> \n> If the rename is detected by the current rename detection code\n> (git-diff-tree -M) then the merge case described above is handled\n> perfectly fine by the current git. That is, the rename is followed and\n> the patch fixing the typo is applied to the renamed file. This assumes\n> that the default merge strategy (recursive) is used.\n\nBut the non-obviously important part here to note is that the branch B\nmerely \"corrects a typo on a comment somewhere\" - the latest versions in\nbranch A and branch B are always compared for renames, therefore if\nbranch A renamed the file and branch B sums up to some larger-scale\nchanges in the file, it still won't be merged properly.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nRight now I am having amnesia and deja-vu at the same time.  I think\nI have forgotten this before.\n"},{"id":"19553","messageId":"7vejz8241m.fsf@assigned-by-dhcp.cox.net","threadId":"4064","inReplyTo":"20060505092332.GY27689@pasky.or.cz","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-05-05T09:51:01Z","receivedAt":"2006-05-05T09:51:01Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Petr Baudis <pasky@suse.cz> writes:\n\n> But the non-obviously important part here to note is that the branch B\n> merely \"corrects a typo on a comment somewhere\" - the latest versions in\n> branch A and branch B are always compared for renames, therefore if\n> branch A renamed the file and branch B sums up to some larger-scale\n> changes in the file, it still won't be merged properly.\n\nI probably am guilty of starting this misinformation, but the\ncode does not compare the latest in A and B for rename\ndetection; it compares (O, A) and (O, B).\n\nBut the end result is the same - what you say is correct.  If a\npath (say O to A) that renamed has too big a change, then no\nmatter how small the changes are on the other path (O to B),\nrename detection can be fooled.  We could perhaps alleviate it\nby following the whole commit chain.\n"},{"id":"19563","messageId":"20060505163629.GZ27689@pasky.or.cz","threadId":"4064","inReplyTo":"20060505005659.9092.qmail@science.horizon.com","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2006-05-05T16:36:29Z","receivedAt":"2006-05-05T16:36:29Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Fri, May 05, 2006 at 02:56:59AM CEST, I got a letter\nwhere linux@horizon.com said that...\n> Actually, AFAICT from looking at the mailing list history, it's not dirty\n> politics: the tie-breaker was the support and enthusiasm of the mercurial\n> developers.  It passed with only minor comment on the git mailing list,\n> but it was a Big Thing to the hg folks.\n> \n> There are ups and downs.  OpenSolaris is definitely the big fish in\n> the mercurial pond (that wasn't *meant* to sound like a recipe for\n> heavy metal toxicity), and will get lots of attention, but git has more\n> real-world experience.  The big fish in the git pond is Linus and Linux.\n> \n> In any case, mercurial and git are really very similar, far closer\n> to each other than any third system, so it's not like the decision is\n> a descent into heresy.  Hopefully some useful cross-pollination\n> can occur, and converting history from one to the other would be\n> simple if anyone ever wanted to.\n\nIt's a philosophical question here, but I'd say that Git is much closer\nto Monotone than to any other version control system - I think it can be\ndescribed as Monotone model with more elegant implementation (for some,\nat least ;), no certificates and restriction of one head per branch.\nAnd another important difference is that Monotone has persistent file\nidentifiers, but I think that's about the only thing that would make\nMonotone more \"file orientated\".\n\nI'm not much of a Mercurial pro but it appears to me that the\narchitectural differences there are larger, especially wrt. the revlogs\nand wholly quite a more file-oriented model.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nRight now I am having amnesia and deja-vu at the same time.  I think\nI have forgotten this before.\n"},{"id":"19564","messageId":"20060505164045.GA27689@pasky.or.cz","threadId":"4064","inReplyTo":"7vejz8241m.fsf@assigned-by-dhcp.cox.net","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2006-05-05T16:40:45Z","receivedAt":"2006-05-05T16:40:45Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Fri, May 05, 2006 at 11:51:01AM CEST, I got a letter\nwhere Junio C Hamano <junkio@cox.net> said that...\n> Petr Baudis <pasky@suse.cz> writes:\n> \n> > But the non-obviously important part here to note is that the branch B\n> > merely \"corrects a typo on a comment somewhere\" - the latest versions in\n> > branch A and branch B are always compared for renames, therefore if\n> > branch A renamed the file and branch B sums up to some larger-scale\n> > changes in the file, it still won't be merged properly.\n> \n> I probably am guilty of starting this misinformation, but the\n> code does not compare the latest in A and B for rename\n> detection; it compares (O, A) and (O, B).\n\nWhere O = LCA(A,B) (modulo recursiveness)? Yes, that is what I meant to\nsay but I phrased it wrong, sorry.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nRight now I am having amnesia and deja-vu at the same time.  I think\nI have forgotten this before.\n"},{"id":"19565","messageId":"e3fvj2$779$1@sea.gmane.org","threadId":"4064","inReplyTo":"7vejz8241m.fsf@assigned-by-dhcp.cox.net","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-05-05T16:47:35Z","receivedAt":"2006-05-05T16:47:35Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Junio C Hamano wrote:\n\n> Petr Baudis <pasky@suse.cz> writes:\n> \n>> But the non-obviously important part here to note is that the branch B\n>> merely \"corrects a typo on a comment somewhere\" - the latest versions in\n>> branch A and branch B are always compared for renames, therefore if\n>> branch A renamed the file and branch B sums up to some larger-scale\n>> changes in the file, it still won't be merged properly.\n> \n> I probably am guilty of starting this misinformation, but the\n> code does not compare the latest in A and B for rename\n> detection; it compares (O, A) and (O, B).\n> \n> But the end result is the same - what you say is correct.  If a\n> path (say O to A) that renamed has too big a change, then no\n> matter how small the changes are on the other path (O to B),\n> rename detection can be fooled.  We could perhaps alleviate it\n> by following the whole commit chain.\n\nOr perhaps by helper information about renames, entered either by git-mv\n(and git-cp) or rename detection at commit, e.g. in the following form\n\n        note at <commit-sha1> was-in <pathname>\n        note at <commit-sha1> was-in <pathname>\n\n(with the obvious limit of this \"note header\" solution is that it wouldn't\nwork for filenames and directory name containing \"\\n\"). I'm not sure if\n<pathname> should be just basename, of full pathname.\n\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"19567","messageId":"Pine.LNX.4.64.0605050944200.3622@g5.osdl.org","threadId":"4064","inReplyTo":"20060505163629.GZ27689@pasky.or.cz","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-05T17:48:38Z","receivedAt":"2006-05-05T17:48:38Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 5 May 2006, Petr Baudis wrote:\n> \n> It's a philosophical question here, but I'd say that Git is much closer\n> to Monotone than to any other version control system\n\nSome historical background..\n\nBefore I dropped BK, I ended up being involved in trying to get Larry and \nTridge to come to some agreement about how to solve the issues Tridge had \nwith BK not being open-source. That actually went on for maybe two months \nor so, and I kept on hoping that we'd find some acceptably middle ground. \n\nI thought we could find somethign that would actually work for everybody: \nto hopefully both make BK technically better, _and_ to make the end result \nmore palatable to the \"free software or bust\" contingency.\n\nOne of the suggestions that I tried to push as an acceptable middle ground \nwas to make a \"generic\" BK repository export format, so that people who \ndidn't want to use BK could still get all the information, and not in a \nbroken format like CVS (yes, CVS makes sense as an interchange format, \nsince _everybody_ speaks CVS, but it's a horrible, horrible, horrible \nformat from any technical standpoint).\n\nMy example export format was really a strange mixture of patches with \nparenthood information, where the history information was described with \nhashes (MD5 rather than SHA1, but that was just an implementation thing, \nand mostly because BK used MD5 sums). Not something really useful as a \nreal SCM, but it wasn't designed for that - it was just meant to be a \nuseful and unambiguous interoperability format.\n\nNow, that didn't work out, and I was a little bummed. I thought it would \nhave made both sides happy, because it would actually have been a better \nformat than CVS (and yes, I'm somewhat biased: in my opinion, having a \nmillion monkeys throwing crap at the walls and encoding the information in \nthe patterns on monkey shit is a better format than CVS), so it would \nactually have improved BK, while also making it possible to interoperate \nif you didn't want to use BK itself.\n\nBut Tridge didn't believe that it would actually have exported all the \ninformation in a BK tree, even if both I and Larry told him it would. I'm \nnot a hundred percent sure that Larry would have gone for the export \nformat either, but hey, one sign of a good compromise is that neither side \nreally gets what they really want. Whatever. It didn't work.\n\nSo it didn't actually resolve the deadlock, but when it became clear that \nI couldn't work with BK any more, I thought I might use something like \nthat \"patch + parenthood\" representation as a way to maintain my tree \nwhile looking at other alternatives.\n\nSo in many ways, when I started looking around for distributed SCM's, I \ncame into the game with the background of keeping the history around as \nchains of hashes describing it, and then just having patches to describe \nthe differences between versions.\n\nSo that was really my \"fallback\" position: if nothing out there worked, \nI'd rather go back to lists of patches than use CVS. \n\nNow, if you keep track of just patches, one of the issues is that you \ncan't afford to re-create the tree every time by walking patches forward \nfrom the beginning, so I also was planning to have an \"cache\" that \nmaintained the current state of the tree as a separate state from the \nworking tree, so that I would always have the \"working tree\" and the \n\"result of patches up to this moment\" as two separate things (so that I \ncould do the \"bk diff\" that I was used to doing to see the difference \nbetween my last state and the current state of the working tree).\n\nIn other words, I was already working on the git \"index\" file. And I was \nplanning to just have a patch-based system behind it, with a hashed \nhistory. Kind of \"quilt with history and an index to speed things up\".\n\nThe index itself would be backed-up with whole files (all hidden in the \n\".dircache\" directory), and the patch series would thus normally never \nactually be _used_. So the inefficiency of working with patches would \nnever be much of an issue. A \"commit\" would create a new patch from the \ncurrent working directory and the previous shadow tree, and update the \nshadow tree and add a new entry to the history list.\n\nAnd then I found Monotone.\n\nNow, monotone was slow. Monotone was so _horrendously_ slow that I had to \ndo special hacks just to import _one_ version of Linux into it in less \nthan two hours. It was something stupid like an O(N**3) algorithm in the \nnumber of filenames (and the kernel had 17,291 files at that time: \nv2.6.12-rc2), and it was just totally unusable for me.\n\nI also thought (and still think) that the whole signing thing was a waste \nof time and misdesigned, and I obviously am not a huge fan of databases. \nSo in many ways I disliked the monotone implementation decisions (and some \nof its design decisions). But at the same time, I immediately liked the \nSHA1 object naming concept of Monotone.\n\nIt also already matched how I had conceptually planned on doing on the \nhistory anyway, and had some ideas for, but it took that whole \"history \nhashing\" all the way.\n\nAnd thus git was born. \n\nSo git really has three parents. In a very real sense, BK (or, perhaps \nmore appropriately - the way I personally used BK, which is not \nnecessarily how others have used it) was the biggest thing from the \nstandpoint of what I wanted my _workflow_ to be like. It was simply how I \nhad done things for the last few years, so a lot of my mental model for \nhow things are supposed to _work_ came from BK. \n\nI still don't think people give Larry enough credit for actually pushing \nthis whole distributed SCM thing as a _usable_ model. Very few of the \nopen-source distributed SCM's are actually usable even today, and as far \nas I've been able to gather, the commercial ones aren't really any closer \neither. Larry didn't have the kind of examples of what _can_ work that I \nhad.\n\nThe other parent was the stupid \"series of patches\" model, which was what \nreally resulted in the \"index\" thing. I realize that people don't always \nmuch like the index, but it's really a pretty central part of git history, \nand one of the distinguising marks of git. It may be trivial, and to some \ndegree it's been overshadowed by all the tree operations we do (the \ncombination of revision walking and tree diffing), but it was very central \nto how git came to be.\n\nThe index also ended up being central to how we did merges - even if some \nday we may end up doing more of that on a pure tree level (ie the current \ngit-merge-tree model), I think the way we ended up doing merges owes a lot \nto the index as a staging area.\n\n(Historically, the \"index\" was called the \"cache\". Exactly because it came \nfrom the notion of \"caching\" the top commit state in a patch series, and \nthen working with patches either backwards or forwards from that top \ncached state. Similarly, we didn't have a \".git\" directory: it was \ncalled \".dircache\", exactly because it was all about caching the state \nof the previous commit directory layout).\n\nAnd finally, Monotone for the \"everything is an object named by its SHA1\" \nmodel, which to some degree is perhaps the central - or at least the most \nobvious - part of git. It largely was designed really just to be the \n\"backing store\" for the \"cache\", and to not be _that_ important. That also \nexplains why I didn't worry too much about disk usage etc initially: the \nobject store wasn't even the most important part, and I envisioned just \nmoving old objects that weren't needed into some \"backup storage\" kind of \nthing.\n\n\t\t\tLinus\n"},{"id":"19568","messageId":"20060505181540.GB27689@pasky.or.cz","threadId":"4064","inReplyTo":"20060505005659.9092.qmail@science.horizon.com","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2006-05-05T18:15:41Z","receivedAt":"2006-05-05T18:15:41Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Fri, May 05, 2006 at 02:56:59AM CEST, I got a letter\nwhere linux@horizon.com said that...\n> But, as Linus has pointed out, this is a very partial solution which\n> introduces a lot of difficulties elsewhere.  File renaming is a subset of\n> the general class of code reorganizations.  Source files will be split,\n> merged, and have functions moved back and forth.  You want the patch to\n> find the code it applies to even if that code was moved.\n> \n> And that can be done by taking a more global view of the patch.\n> Identical file names is only a heuristic.  If the hunk on branch A\n> can't find a place to apply on the same file in branch B, then\n> you have to look a little harder, either at changes from branch B\n> that introduce matching code elsewhere, or perhaps looking\n> through history for a change that removed the match from the\n> obvious place to see if it added a match elsewhere.\n\nThere are really two distinctions here which should be kept separate:\nautomatic vs. explicit movement tracking and file-level vs.\nsubfile-level movement tracking.\n\nThe automatic vs. explicit movement tracking is a lot more\ncontroversial. Explicit movement tracking is pretty easy to provide for\nfile-level movements, it's just that the user says \"I _did_ move file\nA to file B\" (I never got the Linus' argument that the user has no idea\n- he just _performed_ the move, also explicitly, by calling *mv).\n\nHowever, I guess the explicit movement tracking completely fails if you\ngo sub-file (without being extremely bothersome for the user) - you\nwould have to have control over the editor and the clipboard and even\nthen I'm not sure if you could reach any sensible results.\n\nI still dislike automated movement tracking for whole files, but I'm\nconciliated with it. Because it is probably the only really sensible way\nto implement subfile-level tracking.  It would not be hard to implement\nusing pickaxe (actually, I believe it was near the top of Junio's TODO\nfew weeks ago) and a similarity detector comparing new and old version\n(if it's dissimilar enough, check if that or a similar hunk was not\nadded somewhere else in the same commit; well, at least the idea\nsounds simple).\n\nOne obvious problem are ambiguities - several similar files are renamed\nto other similar files and now how do you decide which version to\nchoose? Merge the change to all the new files? Only to some? Panic?\nI wonder how does the current recursive strategy deal with that.\nOf course, this case sounds quite artificial and rare for whole files,\nbut I suspect that it will be much more common once you do not deal with\nfiles but just hunks, moving bits of code around.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nRight now I am having amnesia and deja-vu at the same time.  I think\nI have forgotten this before.\n"},{"id":"19569","messageId":"20060505182001.GC27689@pasky.or.cz","threadId":"4064","inReplyTo":"20060505181540.GB27689@pasky.or.cz","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2006-05-05T18:20:01Z","receivedAt":"2006-05-05T18:20:01Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Fri, May 05, 2006 at 08:15:41PM CEST, I got a letter\nwhere Petr Baudis <pasky@suse.cz> said that...\n> There are really two distinctions here which should be kept separate:\n> automatic vs. explicit movement tracking and file-level vs.\n> subfile-level movement tracking.\n\nI should have revised this paragraph before sending the mail out, I\nended up sorting out my thoughts on the subject as I wrote the mail. The\ntwo aspects end up so tied that it makes sense to mingle them. Examining\nthem separately here still hopefully shed some light on possible\nreasoning behind the Git design decisions.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nRight now I am having amnesia and deja-vu at the same time.  I think\nI have forgotten this before.\n"},{"id":"19570","messageId":"e3g5em$q0f$1@sea.gmane.org","threadId":"4064","inReplyTo":"20060505181540.GB27689@pasky.or.cz","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-05-05T18:27:40Z","receivedAt":"2006-05-05T18:27:40Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Petr Baudis wrote:\n\n> The automatic vs. explicit movement tracking is a lot more\n> controversial. Explicit movement tracking is pretty easy to provide for\n> file-level movements, it's just that the user says \"I _did_ move file\n> A to file B\" (I never got the Linus' argument that the user has no idea\n> - he just _performed_ the move, also explicitly, by calling *mv).\n> \n> However, I guess the explicit movement tracking completely fails if you\n> go sub-file (without being extremely bothersome for the user) - you\n> would have to have control over the editor and the clipboard and even\n> then I'm not sure if you could reach any sensible results.\n\nIf I remember correctly there are some problems if the explicit file-level\ncontents movement tracking (aka. file rename tracking) is done via\nequivalent of file-id, inodes, or persistent names. Although it works for\nmany (most?) cases.\n\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"19571","messageId":"Pine.LNX.4.64.0605051123420.3622@g5.osdl.org","threadId":"4064","inReplyTo":"20060505181540.GB27689@pasky.or.cz","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-05T18:31:06Z","receivedAt":"2006-05-05T18:31:06Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 5 May 2006, Petr Baudis wrote:\n> \n> The automatic vs. explicit movement tracking is a lot more\n> controversial. Explicit movement tracking is pretty easy to provide for\n> file-level movements, it's just that the user says \"I _did_ move file\n> A to file B\" (I never got the Linus' argument that the user has no idea\n> - he just _performed_ the move, also explicitly, by calling *mv).\n\nTHE USER DID NO SUCH THING.\n\nMoving data around happens with a whole lot more than \"mv\".\n\nIt happens with patches (somebody _else_ may have done an \"mv\", without \nusing git at all), and it happens with editors (moving data around until \nmost of it exists in another file).\n\nSo doing \"*mv\" is just a special case.\n\nAnd supporting special cases is _wrong_. If you start depending on data \nthat isn't actually dependable, that's WRONG.\n\nThere's another reason why encoding movement information in the commit is \ntotally broken, namely the fact that a lot of the actions DO NOT WALK THE \nCOMMIT CHAIN!\n\nTry doing\n\n\tgit diff v1.3.0..\n\nand think about what that actually _means_. Think about the fact that it \ndoesn't actually walk the commit chain at all: it diffs the trees between \nv1.3.0 and the current one. What if the rename happened in a commit in the \nmiddle?\n\nThe \"track contents, not intentions\" approach avoids both these things. \nThe end result is _reliable_, not a \"random guess\".\n\nAdding file movement note to commits is simply WRONG.\n\nWhy does this come up every three months or so? I was right the first \ntime. You'd think that as time passes, people would just notice more and \nmore how right I was and am, instead of forgetting and bringing this \nidiotic idea up over and over and over again.\n\n\t\tLinus\n"},{"id":"19573","messageId":"e3g6mq$uoq$1@sea.gmane.org","threadId":"4064","inReplyTo":"e3fvj2$779$1@sea.gmane.org","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-05-05T18:49:03Z","receivedAt":"2006-05-05T18:49:03Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Jakub Narebski wrote:\n\n> Junio C Hamano wrote:\n> \n>> Petr Baudis <pasky@suse.cz> writes:\n>> \n>>> But the non-obviously important part here to note is that the branch B\n>>> merely \"corrects a typo on a comment somewhere\" - the latest versions in\n>>> branch A and branch B are always compared for renames, therefore if\n>>> branch A renamed the file and branch B sums up to some larger-scale\n>>> changes in the file, it still won't be merged properly.\n>> \n>> I probably am guilty of starting this misinformation, but the\n>> code does not compare the latest in A and B for rename\n>> detection; it compares (O, A) and (O, B).\n>> \n>> But the end result is the same - what you say is correct.  If a\n>> path (say O to A) that renamed has too big a change, then no\n>> matter how small the changes are on the other path (O to B),\n>> rename detection can be fooled.  We could perhaps alleviate it\n>> by following the whole commit chain.\n> \n> Or perhaps by helper information about renames, entered either by git-mv\n> (and git-cp) or rename detection at commit, e.g. in the following form\n> \n>         note at <commit-sha1> was-in <pathname>\n>         note at <commit-sha1> was-in <pathname>\n> \n> (with the obvious limit of this \"note header\" solution is that it wouldn't\n> work for filenames and directory name containing \"\\n\"). I'm not sure if\n> <pathname> should be just basename, of full pathname.\n\nErm, I'm sorry, forget the implementation which wouldn't work. The idea was\nto accumulate renames and contents moving information, and remember at\nwhich commit it occured. But it's place (as a _helper_ information) is\nperhaps in separate structure.\n\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"19574","messageId":"20060505185445.GD27689@pasky.or.cz","threadId":"4064","inReplyTo":"Pine.LNX.4.64.0605051123420.3622@g5.osdl.org","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2006-05-05T18:54:45Z","receivedAt":"2006-05-05T18:54:45Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Fri, May 05, 2006 at 08:31:06PM CEST, I got a letter\nwhere Linus Torvalds <torvalds@osdl.org> said that...\n> Moving data around happens with a whole lot more than \"mv\".\n\nLet's keep this on the per-file level - if you want to go below the file\ngranularity, I already _DID_ say that I agree that explicit tracking is\nnot a way. (If sub-file tracking would end up having any usable\nreliability in real-world cases, which is something I do not take for\ngranted.)\n\nAnother thing is, the sub-file content tracking would end up being a lot\nmore \"magic\" than the simple per-file content tracking, and you stated\nseveral times that you prefer simple merge over better but magic merge -\nso why do you prefer sub-file content tracking anyway?\n\n> It happens with patches (somebody _else_ may have done an \"mv\", without \n> using git at all),\n\n_Here_ is the place for automated renames detection. Between applying\nand committing the patch, the user can verify that it got the renames\nright. That's impossible when guessing the renames later.\n\n> and it happens with editors (moving data around until \n> most of it exists in another file).\n\nI doubt this in fact happens that often (to a degree the automatic\nrename detection would catch). And if it happens, then the user has to\ntell Git - I have never heard that _this_ would be any problem in other\nversion control systems. You could make it more foolproof by running the\nautomatic rename detection on the diff being committed and suggesting\nthe user that other yet unrecorded renames did happen.\n\nThe point is, the user stays in control and can override any stupid guess.\n\n> So doing \"*mv\" is just a special case.\n> \n> And supporting special cases is _wrong_. If you start depending on data \n> that isn't actually dependable, that's WRONG.\n\nI prefer making this data dependable to having to resort to guessing on\ndependable less amount of data.\n\n> There's another reason why encoding movement information in the commit is \n> totally broken, namely the fact that a lot of the actions DO NOT WALK THE \n> COMMIT CHAIN!\n> \n> Try doing\n> \n> \tgit diff v1.3.0..\n> \n> and think about what that actually _means_. Think about the fact that it \n> doesn't actually walk the commit chain at all: it diffs the trees between \n> v1.3.0 and the current one. What if the rename happened in a commit in the \n> middle?\n\nThen the automated renames detection will miss it given that the other\naccumulated differences are large enough, and the suggested workarounds\n_are_ precisely walking the commit chain.\n\nIf you use persistent file ids, you never miss it _AND_ you DO NOT WALK\nTHE COMMIT CHAIN! You still just match file ids in the two trees.\n\n> The \"track contents, not intentions\" approach avoids both these things. \n> The end result is _reliable_, not a \"random guess\".\n\nNo, the end result is whichever some heuristic randomly guessed, and\nit's not reliable either since the heuristic can change.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nRight now I am having amnesia and deja-vu at the same time.  I think\nI have forgotten this before.\n"},{"id":"19576","messageId":"20060505190409.GC9937@redhat.com","threadId":"4064","inReplyTo":"Pine.LNX.4.64.0605050944200.3622@g5.osdl.org","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Dave Jones","fromEmail":"davej@redhat.com","sentAt":"2006-05-05T19:04:09Z","receivedAt":"2006-05-05T19:04:09Z","isPatch":false,"sender":{"key":"davej@redhat.com","avatar":null},"body":"On Fri, May 05, 2006 at 10:48:38AM -0700, Linus Torvalds wrote:\n\n > (and yes, I'm somewhat biased: in my opinion, having a \n > million monkeys throwing crap at the walls and encoding the information in \n > the patterns on monkey shit is a better format than CVS), so it would \n > actually have improved BK, while also making it possible to interoperate \n > if you didn't want to use BK itself.\n >  ...\n > So that was really my \"fallback\" position: if nothing out there worked, \n > I'd rather go back to lists of patches than use CVS. \n\nI've encountered managing kernel trees in CVS both during my tenure at SuSE,\nand to a more involved extent as Fedora/RHEL maintainer, and I'd just like\nto echo how much it _completely sucks_ at times.\n\nRebasing to a newer release is a *nightmare* that usually takes\nup most of an afternoon compared to rebasing my git based projects.\n\nIn the event I can't persuade the powers at be to switch to git at some point\nfor managing our packages, I'll be sure to bring up your suggestion of\na million monkeys. I believe you can pick them up fairly cheap these days.\n\n\t\tDave\n\n-- \nhttp://www.codemonkey.org.uk\n"},{"id":"19582","messageId":"e3g9m6$8i0$1@sea.gmane.org","threadId":"4064","inReplyTo":"20060505185445.GD27689@pasky.or.cz","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-05-05T19:39:56Z","receivedAt":"2006-05-05T19:39:56Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Petr Baudis wrote:\n\n> Dear diary, on Fri, May 05, 2006 at 08:31:06PM CEST, I got a letter\n> where Linus Torvalds <torvalds@osdl.org> said that...\n\n> I prefer making this [rename detection] data dependable to having to\n> resort to guessing on dependable less amount of data.\n> \n>> There's another reason why encoding movement information in the commit is\n>> totally broken, namely the fact that a lot of the actions DO NOT WALK THE\n>> COMMIT CHAIN!\n>> \n>> Try doing\n>> \n>> git diff v1.3.0..\n>> \n>> and think about what that actually _means_. Think about the fact that it\n>> doesn't actually walk the commit chain at all: it diffs the trees between\n>> v1.3.0 and the current one. What if the rename happened in a commit in\n>> the middle?\n> \n> Then the automated renames detection will miss it given that the other\n> accumulated differences are large enough, and the suggested workarounds\n> _are_ precisely walking the commit chain.\n> \n> If you use persistent file ids, you never miss it _AND_ you DO NOT WALK\n> THE COMMIT CHAIN! You still just match file ids in the two trees.\n\nLet not jump to the one of the possible solution. The detecting and noting\nrenames and content moving (with user interaction) at commit is nice...\nunless does something which cannot allow interactiveness (like applying\npatchbomb), but even then detecting and saving info at commit would be good\nidea.\n\nWhat we need is to for two given linked revisions (with a path between them)\nto easily extract information about renames (content moving). Perhaps using\nadditional structure... best if we could do this without walking the chain.\nThe rest is details... ;-P\n\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"19584","messageId":"7vr738w8t4.fsf@assigned-by-dhcp.cox.net","threadId":"4064","inReplyTo":"20060505185445.GD27689@pasky.or.cz","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-05-05T19:49:59Z","receivedAt":"2006-05-05T19:49:59Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Petr Baudis <pasky@suse.cz> writes:\n\n> I doubt this in fact happens that often (to a degree the automatic\n> rename detection would catch). And if it happens, then the user has to\n> tell Git - I have never heard that _this_ would be any problem in other\n> version control systems.\n\nIt does not become an issue only because users accept it as a\nfact of life.  When Linus was moving most of the contents in\nrev-list.c to create a new revision.c, I already had some tweaks\nto rev-list.c published before he sent me a patch for the code\nmovement, and I am sure he needed to re-roll the patch by\nmerging the change I did to rev-list.c back into his revision.c\nfile.  No SCM may handle that automatically, and no user\naccustomed to existing SCM (including git) expect that to work\nautomatically.  But that does not necessarily mean a tool that\nnotices it and tells user what is going on is a bad thing.\n\nHowever it is a different story to try recording \"what is going on\"\nwhether it comes from the tool's guess or directly from the user.\n\nHaving a way to affect the inprecise \"guess\" the tool makes when\nthat guesswork is needed might make sense.  If you (think you)\nknow arch/i386/foo.h was copied to create arch/x86-64/foo.h but\nthe detector does not detect it and seeing a creation patch for\narch/x86-64/foo.h frustrates you, you may want to have a way to\nexplicitly say \"compare arch/i386/foo.h with arch/x86-64/foo.h\nin that commit -- I want to examine the change needed to adjust\nfoo to x86-64 architecture\".\n\nBut we have \"git diff v2.6.14:arch/i386/foo.h v2.6.14:arch/x86-64/foo.h\"\nfor that ;-).\n\n> Then the automated renames detection will miss it given that the other\n> accumulated differences are large enough, and the suggested workarounds\n> _are_ precisely walking the commit chain.\n\nThe HEAD may _not_ have anything to do with v1.3.0 in which case\nyou would get nothing from walking the ancestry.\n\n> If you use persistent file ids, you never miss it _AND_ you DO NOT WALK\n> THE COMMIT CHAIN! You still just match file ids in the two trees.\n\nIt is unworkable.\n\nWhich one should inherit the persistent id of the old\nrev-list.c?  New rev-list.c, or revision.c that has most of the\nold contents split out?\n\nOh, and did you know there was a different revision.h that is\nnot related to the current revision.h in the history of git?\nShould its persistent id have any relation with the persistent\nid of the current revision.h?  When would you decide to make the\nid inherited and when not to?  If I remove revision.h by mistake\nin a commit and resurrect it in the next commit, should it get\nthe same id back?  If I forget to tell the tool that those two\n\"disappeared and then reappeared\" are related and should get the\nsame persistent id when I make the resurrection commit, and keep\npiling other commits on top, do I have to rewind the ancestry\nchain all the way to correct the mistake?\n"},{"id":"19587","messageId":"20060505204516.GA82888@dspnet.fr.eu.org","threadId":"4064","inReplyTo":"20060505181540.GB27689@pasky.or.cz","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Olivier Galibert","fromEmail":"galibert@pobox.com","sentAt":"2006-05-05T20:45:16Z","receivedAt":"2006-05-05T20:45:16Z","isPatch":false,"sender":{"key":"galibert@pobox.com","avatar":null},"body":"On Fri, May 05, 2006 at 08:15:41PM +0200, Petr Baudis wrote:\n> The automatic vs. explicit movement tracking is a lot more\n> controversial. Explicit movement tracking is pretty easy to provide for\n> file-level movements, it's just that the user says \"I _did_ move file\n> A to file B\" (I never got the Linus' argument that the user has no idea\n> - he just _performed_ the move, also explicitly, by calling *mv).\n\nIn one of my projects 99% or the renames are \"done\" when unzipping the\nsource release of the next version.  Explicit tracking would be\nunbearable, frankly.\n\nAnd once you have a good enough implicit tracking, why bother with an\nexplicit one?\n\n  OG.\n"},{"id":"19596","messageId":"46a038f90605052353m2d2aca11weac7efee80c6fb35@mail.gmail.com","threadId":"4064","inReplyTo":"7vr738w8t4.fsf@assigned-by-dhcp.cox.net","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-06T06:53:36Z","receivedAt":"2006-05-06T06:53:36Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/6/06, Junio C Hamano <junkio@cox.net> wrote:\n> > If you use persistent file ids, you never miss it _AND_ you DO NOT WALK\n> > THE COMMIT CHAIN! You still just match file ids in the two trees.\n>\n> It is unworkable.\n\n+1 -- explicit file ids are evil. Arch/TLA demonstrated that amply...\nthey are a serious annoyance to the end user, they have a lot of\nnot-elegantly solvable cases (same file created with the same contents\nin several repos -- say via an emailed patch) that git gets right\n_today_.\n\nThey _are_ useful in a very small set of cases -- namely in the case\nof a naive mv, which git handles correctly today. Subtler things git\nsometimes does right, sometimes fails, but it can be made to be much\nsmarter by interpreting content changes better, for instance all this\ntalk about getting pickaxe to guess where the patch should be applied\nfor a file that got split into 3.\n\nBut those subtler cases are totally impossible with explicit id\ntracking. I used Arch for a long time with very large trees, and\nrenames coming left, right and centre. Explicit ids didn't help much,\nand the number of manual fixups we had to do was awful.\n\nI am using GIT with the very same project, and just now, typing this,\nI realised that there are still many renames happening in the project.\nI had forgotten about it -- well, not really: I do use git-merge\ninstead of cg-merge when I suspect there may be interesting cases ;-)\n\nOf course, YMMV, and I have to confess I was a sceptic for a while...\nbut now as an end-user dealing with messy projects, I say LIRAR: Linus\nIs Right About Renames.\n\nOTOH,\n\n>> Try doing\n>>\n>> git diff v1.3.0..\n>>\n>> and think about what that actually _means_. Think about the fact that it\n>> doesn't actually walk the commit chain at all: it diffs the trees between\n>> v1.3.0 and the current one. What if the rename happened in a commit in\n>> the middle?\n>\n> Then the automated renames detection will miss it given that the other\n> accumulated differences are large enough, and the suggested workarounds\n> _are_ precisely walking the commit chain.\n\nI agree here with Pasky that after a while the automated\nrenames/copy/splitup detection will miss the operation in cases where\nit would be interesting to note it to the user. IIRC git-rerere is the\ntool that knows about this (still voodoo to me how) and could be used\nto help here. At what (runtime) cost, I don't know, but that kind of\nwalking history to tell me more interesting things about the diff is\nsomething that is usually worthwhile.\n\nUsual disclaimers apply.\n\n\nmartin\n"},{"id":"19598","messageId":"7v7j4zvd4x.fsf@assigned-by-dhcp.cox.net","threadId":"4064","inReplyTo":"46a038f90605052353m2d2aca11weac7efee80c6fb35@mail.gmail.com","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-05-06T07:14:06Z","receivedAt":"2006-05-06T07:14:06Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Martin Langhoff\" <martin.langhoff@gmail.com> writes:\n\n> I agree here with Pasky that after a while the automated\n> renames/copy/splitup detection will miss the operation in cases where\n> it would be interesting to note it to the user. IIRC git-rerere is the\n> tool that knows about this (still voodoo to me how) and could be used\n> to help here. At what (runtime) cost, I don't know, but that kind of\n> walking history to tell me more interesting things about the diff is\n> something that is usually worthwhile.\n\nFYI rerere is a totally unrelated voodoo.\n\nIt remembers the conflict marker pattern <<< === >>> immediately\nafter it runs \"merge\" (ah, that reminds me -- I should replace\nthem with diff3), and then remembers the result of the manual\nresolution just before the user makes a commit.  Then, when next\ntime it runs \"merge\" for something and notices <<< === >>>\npattern it has seen before, it runs a three-way merge between\nthe previous resolution result and the current conflicted state,\nusing the previous conflicted state as the common origin.\n"},{"id":"19599","messageId":"e3hjfk$bjn$1@sea.gmane.org","threadId":"4064","inReplyTo":"46a038f90605052353m2d2aca11weac7efee80c6fb35@mail.gmail.com","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-05-06T07:33:15Z","receivedAt":"2006-05-06T07:33:15Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Martin Langhoff wrote:\n\n> On 5/6/06, Junio C Hamano <junkio@cox.net> wrote:\n\n>>> Try doing\n>>>\n>>> git diff v1.3.0..\n>>>\n>>> and think about what that actually _means_. Think about the fact that it\n>>> doesn't actually walk the commit chain at all: it diffs the trees\n>>> between v1.3.0 and the current one. What if the rename happened in a\n>>> commit in the middle?\n>>\n>> Then the automated renames detection will miss it given that the other\n>> accumulated differences are large enough, and the suggested workarounds\n>> _are_ precisely walking the commit chain.\n> \n> I agree here with Pasky that after a while the automated\n> renames/copy/splitup detection will miss the operation in cases where\n> it would be interesting to note it to the user.\n\nPerhaps an option to do rename detection with walking the commit chain?\n\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"19601","messageId":"7vslnntxay.fsf@assigned-by-dhcp.cox.net","threadId":"4064","inReplyTo":"e3hjfk$bjn$1@sea.gmane.org","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-05-06T07:41:25Z","receivedAt":"2006-05-06T07:41:25Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> writes:\n\n> Perhaps an option to do rename detection with walking the commit chain?\n\nHave fun implementing that ;-).\n"},{"id":"19603","messageId":"4fb292fa0605060546g54d8226ake4a91e8a30e0b2f1@mail.gmail.com","threadId":"4064","inReplyTo":"7vslnntxay.fsf@assigned-by-dhcp.cox.net","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Bertrand Jacquin","fromEmail":"beber.mailing@gmail.com","sentAt":"2006-05-06T12:46:31Z","receivedAt":"2006-05-06T12:46:31Z","isPatch":false,"sender":{"key":"beber.mailing@gmail.com","avatar":null},"body":"On 5/6/06, Junio C Hamano <junkio@cox.net> wrote:\n> Jakub Narebski <jnareb@gmail.com> writes:\n>\n> > Perhaps an option to do rename detection with walking the commit chain?\n>\n> Have fun implementing that ;-).\n\nI agree that it could be interesting to have a such thing. But that's\nincrendibly stupid and moreover a rare case.\n\n--\nBeber\n#e.fr@freenode\n"},{"id":"19604","messageId":"e3i8r1$ve2$1@sea.gmane.org","threadId":"4064","inReplyTo":"e3g9m6$8i0$1@sea.gmane.org","subject":"Re: [ANNOUNCE] Git wiki","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-05-06T13:37:47Z","receivedAt":"2006-05-06T13:37:47Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Jakub Narebski wrote:\n\n> Petr Baudis wrote:\n\n>> If you use persistent file ids, you never miss it _AND_ you DO NOT WALK\n>> THE COMMIT CHAIN! You still just match file ids in the two trees.\n> \n> Let not jump to the one of the possible solution. The detecting and noting\n> renames and content moving (with user interaction) at commit is nice...\n> unless does something which cannot allow interactiveness (like applying\n> patchbomb), but even then detecting and saving info at commit would be\n> good idea.\n> \n> What we need is to for two given linked revisions (with a path between\n> them) to easily extract information about renames (content moving).\n> Perhaps using additional structure... best if we could do this without\n> walking the chain. The rest is details... ;-P\n\nOr rather structure, which for given file F in given revision A, for given\nother revision B would tell ALL the files in the revision B which are\nsource of contents (via history/commit tree) of the file F.\n\n-- \nJakub Narebski\nWarsaw, Poland\n"}]}