{"thread":{"id":"3502","subject":"impure renames / history tracking","startedAt":"2006-03-01T14:01:28Z","lastAt":"2006-03-02T22:06:00Z","messageCount":15,"participants":["Paul Jakma","Andreas Ericsson","Linus Torvalds","Junio C Hamano","Martin Langhoff"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"16942","messageId":"Pine.LNX.4.64.0603011343170.13612@sheen.jakma.org","threadId":"3502","inReplyTo":null,"subject":"impure renames / history tracking","fromName":"Paul Jakma","fromEmail":"paul@clubi.ie","sentAt":"2006-03-01T14:01:28Z","receivedAt":"2006-03-01T14:01:28Z","isPatch":false,"sender":{"key":"paul@clubi.ie","avatar":null},"body":"Hi,\n\nI'm trying to understand git better (so I can explain it better to \nothers, with an eye to them considering switching to git), one \nquestion I have is about renames.\n\n- git obviously detects pure renames perfectly well\n\n- git doesn't however record renames, so 'impure' renames may not be\n   detected\n\nMy question is:\n\n- why not record rename information explicitely in the commit object?\n\nI.e. so as to be able to follow history information through 'impure' \nrenames without having to resort to heuristics.\n\nE.g. imagine a project where development typically occurs through:\n\no: commit\nm: merge\n\n    o---o-m--o-o-o--o----m <- project\n   /     /              /\no-o-o-o-o--o-o-o--o-o-o <- main branch\n\nThe project merge back to main in one 'big' combined merge \n(collapsing all of the commits on 'project' into one commit). This \nleads to 'impure renames' being not uncommon. The desired end-result \nof merging back to 'main' being to rebase 'project' as one commit \nagainst 'main', and merge that single commit back, a la:\n\n    o---o-m--o-o-o--o----m <- project\n   /     /              /\no-o-o-o-o--o-o-o--o-o-o---m <- main branch\n                        \\ /\n                         o <- project_collapsed\n\nSo that 'm' on 'main' is that one commit[1].\n\nThe merits or demerits of such merging practice aside, what reason \nwould there be /against/ recording explicit rename information in the \ncommit object, so as to help browsers follow history (particularly \nimpure renames) better in a commit?\n\nI.e. would there be resistance to adding meta-info rename headers \ncommit objects, and having diffcore and other tools to use those \nheaders to /augment/ their existing heuristics in detecting renames?\n\nThanks!\n\n1. Git currently doesn't have 'porcelain' to do this, presumably \nthere'd be no objection to one?\n\nregards,\n-- \nPaul Jakma\tpaul@clubi.ie\tpaul@jakma.org\tKey ID: 64A2FF6A\nFortune:\nIt is the quality rather than the quantity that matters.\n- Lucius Annaeus Seneca (4 B.C. - A.D. 65)\n"},{"id":"16946","messageId":"4405C012.6080407@op5.se","threadId":"3502","inReplyTo":"Pine.LNX.4.64.0603011343170.13612@sheen.jakma.org","subject":"Re: impure renames / history tracking","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-03-01T15:38:58Z","receivedAt":"2006-03-01T15:38:58Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Paul Jakma wrote:\n> \n> - git obviously detects pure renames perfectly well\n> \n> - git doesn't however record renames, so 'impure' renames may not be\n>   detected\n> \n> My question is:\n> \n> - why not record rename information explicitely in the commit object?\n> \n\nMainly for two reasons, iirc:\n1. Extensive metadata is evil.\n2. Backwards compatibility. Old repos should always work with new tools. \nOld tools should work with new repos, at least until a new major-release \nis released.\n\n\n> I.e. so as to be able to follow history information through 'impure' \n> renames without having to resort to heuristics.\n> \n> E.g. imagine a project where development typically occurs through:\n> \n> o: commit\n> m: merge\n> \n>    o---o-m--o-o-o--o----m <- project\n>   /     /              /\n> o-o-o-o-o--o-o-o--o-o-o <- main branch\n> \n> The project merge back to main in one 'big' combined merge (collapsing \n> all of the commits on 'project' into one commit). This leads to 'impure \n> renames' being not uncommon. The desired end-result of merging back to \n> 'main' being to rebase 'project' as one commit against 'main', and merge \n> that single commit back, a la:\n> \n>    o---o-m--o-o-o--o----m <- project\n>   /     /              /\n> o-o-o-o-o--o-o-o--o-o-o---m <- main branch\n>                        \\ /\n>                         o <- project_collapsed\n> \n> So that 'm' on 'main' is that one commit[1].\n> \n\nI think you're misunderstanding the git meaning of rebase here. \"git \nrebase\" moves all commits since \"project\" forked from \"main branch\" to \nthe tip of \"main branch\".\n\nOther than that, this is the recommended workflow, and exactly how Linux \nand git both are managed (i.e. topic branches eventually merged into \n'master').\n\nIn your drawings, 'main branch' would be 'master' and 'project' would be \nany amount of topic-branches (or just one, if you like that better).\n\nI'm not sure what you mean by 'project_collapsed' though. If I \nunderstand you correctly, each branch-head represents one 'collapse'. I \nsuggest you clone the git repo and do\n\n\t$ gitk master\n\t$ gitk next\n\t$ gitk pu\n\ngitk is great for visualizing what you've done and what the repo looks \nlike. Use and abuse it frequently every time you're unsure what was you \njust did. It's the best way to quickly learn what happens, really.\n\nIf you just want to distribute snapshots I suggest you do take a look at \ngit-tar-tree. Junio makes nice use of it in the git Makefile (the dist: \ntarget).\n\n\n> The merits or demerits of such merging practice aside, what reason would \n> there be /against/ recording explicit rename information in the commit \n> object, so as to help browsers follow history (particularly impure \n> renames) better in a commit?\n> \n> I.e. would there be resistance to adding meta-info rename headers commit \n> objects, and having diffcore and other tools to use those headers to \n> /augment/ their existing heuristics in detecting renames?\n> \n\nPersonally I think metadata is evil. Renames will still be auto-detected \nanyway, and with the distributed repo setup the only reason git \nshouldn't be able to detect a rename is if you rename a file and hack it \nup so it doesn't even come close to matching its origin (close in this \ncase is 80% by default, I think). In those cases it isn't so much a \nrename as a rewrite. If you find the commit where the file was renamed \nit should be listed in that commit, like so:\n\n\tsimilarity index 92%\n\trename from Documentation/git-log-script.txt\n\trename to Documentation/git-log.txt\n\n(this is gitk output from the git repo. Search for \"Big tool rename\")\n\nIMO this is far better than having to tell git \"I renamed this file to \nthat\", since it also detects code-copying with modifications, and it's \nusually quick enough to find those renames as well.\n\n> Thanks!\n> \n> 1. Git currently doesn't have 'porcelain' to do this, presumably there'd \n> be no objection to one?\n> \n\n\t$ git checkout master\n\t$ git pull . project\n\nThe dot means \"pull from the local repo\". \"project\" is the branch you \nwant to merge into master. You can pull an arbitrary amount of branches \nin one go (\"octopus\" merge). The current tested limit is 12 (thanks, Len \n;) ).\n\nIf, for some reason, you want to combine lots of commits into a single \nmega-patch (like Linus does for each release of the kernel), you can do:\n\n\t$ git diff $(git merge-base main project) project > patch-file\n\nThen you can apply patch-file to whatever branch you want and make the \ncommit as if it was a single change-set. I'd recommend against it unless \nyou're just toying around though. It's a bad idea to lie in a projects \nhistory.\n\nHope that helps.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"16954","messageId":"Pine.LNX.4.64.0603011558390.13612@sheen.jakma.org","threadId":"3502","inReplyTo":"4405C012.6080407@op5.se","subject":"Re: impure renames / history tracking","fromName":"Paul Jakma","fromEmail":"paul@clubi.ie","sentAt":"2006-03-01T16:27:20Z","receivedAt":"2006-03-01T16:27:20Z","isPatch":false,"sender":{"key":"paul@clubi.ie","avatar":null},"body":"On Wed, 1 Mar 2006, Andreas Ericsson wrote:\n\n> Mainly for two reasons, iirc:\n\n> 1. Extensive metadata is evil.\n\nOnly if /required/. I wouldn't argue for rename meta-data to be \n'core', only as an additional hint into the rename-detection process.\n\nFWIW, I think git's rename handling is really nice. It's just I \nsuspect, being a heuristic, it won't be able to follow history \nreliably across 'very impure' renames.\n\n> 2. Backwards compatibility. Old repos should always work with new \n> tools. Old tools should work with new repos, at least until a new \n> major-release is released.\n\nAbsolutely.\n\n>> o: commit\n>> m: merge\n>>\n>>    o---o-m--o-o-o--o----m <- project\n>>   /     /              /\n>> o-o-o-o-o--o-o-o--o-o-o <- main branch\n>> \n>> The project merge back to main in one 'big' combined merge (collapsing all \n>> of the commits on 'project' into one commit). This leads to 'impure \n>> renames' being not uncommon. The desired end-result of merging back to \n>> 'main' being to rebase 'project' as one commit against 'main', and merge \n>> that single commit back, a la:\n>>\n>>    o---o-m--o-o-o--o----m <- project\n>>   /     /              /\n>> o-o-o-o-o--o-o-o--o-o-o---m <- main branch\n>>                        \\ /\n>>                         o <- project_collapsed\n>> \n>> So that 'm' on 'main' is that one commit[1].\n\n> I think you're misunderstanding the git meaning of rebase here. \n> \"git rebase\" moves all commits since \"project\" forked from \"main \n> branch\" to the tip of \"main branch\".\n\nRight, I'm referring to 'rebase' generally, as a concept, not to \ngit-rebase specifically. E.g. git diff main..project is another way \nof rebasing I think.\n\n> Other than that, this is the recommended workflow, and exactly how Linux and \n> git both are managed (i.e. topic branches eventually merged into 'master').\n\nThey're not rebased though, generally. They're pulled. Ie, in Linux \nand git when 'project' is merged, things look like:\n\n     o---o-m--o-o-o--o----m   <- project\n    /     /              / \\\no-o-o-o-o--o-o-o--o-o-o----m <- main branch\n\nThe rest of the world sees /all/ the individual commits of 'project' \nright? The traditional process for the case I'm thinking of results \nin the 'main' tree seeing only /one/ single commit for the project.\n\n> I'm not sure what you mean by 'project_collapsed' though.\n\nAll the commits on the project branch are 'collapsed' into one single \ncommit/delta, and then that /single/ commit is merged to 'main'. Rest \nof the world sees:\n\no-o-o-o-o--o-o-o--o-o-o---m <- main branch\n                        \\ /\n                         o <- project\n\n> correctly, each branch-head represents one 'collapse'.\n\nNot quite. It represents a branch with one or more commits. In the \nLinux and git work flow, multiple commits are left as is.\n\n> gitk is great for visualizing what you've done and what the repo \n> looks like. Use and abuse it frequently every time you're unsure \n> what was you just did. It's the best way to quickly learn what \n> happens, really.\n\nI do. It rocks! :)\n\n> If you just want to distribute snapshots I suggest you do take a \n> look at git-tar-tree. Junio makes nice use of it in the git \n> Makefile (the dist: target).\n\nNeat.\n\nThough, I probably should stay away from the git Makefile for now. \n<cough>.\n\n> Personally I think metadata is evil.\n\nNot sure I agree. Silly/redundant meta-data can be evil alright. But \nI'm talking about meta-data which is not there and potentially not \nreconstructable.\n\n> Renames will still be auto-detected anyway,\n\nChances are so, yes. Definitely with the git and Linux workflows.\n\nThe traditional workflow for the software project I'm thinking of is \ndifferent though. One commit may encompass multiple renames and edits \nof a file (discouraged, but it's possible).\n\nIf my understanding is correct, following back history for such cases \nwould be difficult.\n\nThere is an argument that that 'traditional' process should be \nchanged. However, leaving aside that argument, I'd like to know if \ngit could accomodate that process.\n\n> be able to detect a rename is if you rename a file and hack it up \n> so it doesn't even come close to matching its origin (close in this \n> case is 80% by default, I think). In those cases it isn't so much a \n> rename as a rewrite.\n\nExactly - this is the case I'm concerned about. Imagine that you'd \nlike to be follow the history back through the rewrite and through to \nthe original file.\n\n> IMO this is far better than having to tell git \"I renamed this file \n> to that\", since it also detects code-copying with modifications, \n> and it's usually quick enough to find those renames as well.\n\nI think so too, but that involves arguing that very very \nlong-standing workflows should be changed to accomodate git. I intend \nto make that argument to the 'project' concerned, however I would \nalso like to be say git could equally well deal with the \n'traditional' workflow, modulo having to explicitely use (say) \ngit-mv.\n\n>> 1. Git currently doesn't have 'porcelain' to do this, presumably there'd be \n>> no objection to one?\n>> \n>\n> \t$ git checkout master\n> \t$ git pull . project\n\nRight, but 'pull' isn't what I mean :).\n\nI mean:\n\n \t$ git checkout project\n \t$ git pull . master\n \t$ git checkout -b tmp project\n \t$ git diff project..master | <git apply I think>\n\n> If, for some reason, you want to combine lots of commits into a single \n> mega-patch (like Linus does for each release of the kernel), you can do:\n>\n> \t$ git diff $(git merge-base main project) project > patch-file\n\nRight.\n\n> Then you can apply patch-file to whatever branch you want and make \n> the commit as if it was a single change-set. I'd recommend against \n> it unless you're just toying around though. It's a bad idea to lie \n> in a projects history.\n\nPresume that 'project' in the workflow is defined as\n\n \t\"achieve one goal with one commit to the master\"\n\nSo by definition, it always correct that the project only ever has \none commit.\n\nThe trouble is that /sometimes/ projects do indeed 'rename and \nrewrite' a file. At present, chances are git might not notice this, \nand ability to follow history through the rename+rewrite would be \nlost.\n\nI'm wondering whether:\n\n- this could be solved?\n- how? (some additional advisory-only meta-data in the\n   index-cache and commit?)\n\nIf there is consensus on an acceptable way, I'm willing to implement \nit. (I was thinking of just adding 'rename' headers to the commit \nobjects, then teaching diffcore to consider them in addition to \ncurrent heuristics).\n\nregards,\n-- \nPaul Jakma\tpaul@clubi.ie\tpaul@jakma.org\tKey ID: 64A2FF6A\nFortune:\nBe nice to people on the way up, because you'll meet them on your way down.\n \t\t-- Wilson Mizner\n"},{"id":"16956","messageId":"Pine.LNX.4.64.0603010859200.22647@g5.osdl.org","threadId":"3502","inReplyTo":"Pine.LNX.4.64.0603011558390.13612@sheen.jakma.org","subject":"Re: impure renames / history tracking","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-03-01T17:13:53Z","receivedAt":"2006-03-01T17:13:53Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 1 Mar 2006, Paul Jakma wrote:\n> \n> FWIW, I think git's rename handling is really nice. It's just I suspect, being\n> a heuristic, it won't be able to follow history reliably across 'very impure'\n> renames.\n\nThe thing is, it does better than anything that _tries_ to be \"reliable\".\n\nI can pretty much _guarantee_ that you can't do it better.\n\nTracking \"inodes\" - aka file identities - (which is what BK does, and I \nassume what SVN does) is fundamentally problematic. I particular, it's a \nhorrible problem when two inodes \"meet\" under the same name. You now have \ntwo identities for the same file, and you're fundamentally screwed.\n\nAnd don't tell me it doesn't happen. It _does_ happen, and it did happen \nwith the kernel under BK.\n\nIt doesn't even need renames to be a problem. JUST THE FACT THAT YOU TRY \nTO TRACK FILE \"IDENTITY\" HISTORY IS BROKEN. For example, take CVS, which \ndoesn't actually try to do renames, but _does_ try to track the identity \nof a file, since all the history is tied into that identity: think about \nwhat happens in Attic when a file is deleted. Completely broken model.\n\nNow, CVS doesn't tend to show the problems very much, because people don't \nactually use branches that much (they are a pain in the neck), and they \nsure as hell try to avoid deleting and creating the same filename under a \nbranch and on HEAD. I'm sure you can do it, but I'm also pretty sure \nthere's a lot of old projects around that have ended up moving the ,v \nfiles around to play rename/delete games.\n\nAnd that's really fundamental. CVS doesn't show the problems so much, \nbecause CVS actively tries to make it hard to do these things.\n\nWith renames-tracking-file-identities, it's _really_ easy to get some \nmajor confusion going. What happens when one branch creates a file, and \nanother one renames a file to that same name, and they merge?\n\nDon't tell me it doesn't happen. It happened under BK. The way BK \"solved\" \nit was to keep the two separate identities: one of them got resolved to \nthe new filename, the other one went into the \"deleted\" directory. Guess \nwhat happens when the side that got merged into \"deleted\" continues to \nedit the file? That's right - their edits happen on the deleted file, and \nnever show up in the real tree in a subsequent merge ever again.\n\nAnd as far as I can tell, BK really did the best you can do. Following \nfile identities really _is_ fundamentally broken. It sounds like a nice \nidea, but while you migth solve a few problems, you create a whole raft of \nmuch more fundamental problems.\n\nSo next time you think about a merge that migt have been improved by \ntracking renames, please also think about a merge where one of the \nfilenames came from two or more different sources through an earlier \nmerge, and thank your benevolent Gods that they instructed me to make git \nbe based purely on file contents.\n\n\t\tLinus\n"},{"id":"16963","messageId":"4405DD35.8060804@op5.se","threadId":"3502","inReplyTo":"Pine.LNX.4.64.0603011558390.13612@sheen.jakma.org","subject":"Re: impure renames / history tracking","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-03-01T17:43:17Z","receivedAt":"2006-03-01T17:43:17Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Paul Jakma wrote:\n> On Wed, 1 Mar 2006, Andreas Ericsson wrote:\n> \n>>> o: commit\n>>> m: merge\n>>>\n>>>    o---o-m--o-o-o--o----m <- project\n>>>   /     /              /\n>>> o-o-o-o-o--o-o-o--o-o-o <- main branch\n>>>\n>>> The project merge back to main in one 'big' combined merge \n>>> (collapsing all of the commits on 'project' into one commit). This \n>>> leads to 'impure renames' being not uncommon. The desired end-result \n>>> of merging back to 'main' being to rebase 'project' as one commit \n>>> against 'main', and merge that single commit back, a la:\n>>>\n>>>    o---o-m--o-o-o--o----m <- project\n>>>   /     /              /\n>>> o-o-o-o-o--o-o-o--o-o-o---m <- main branch\n>>>                        \\ /\n>>>                         o <- project_collapsed\n>>>\n>>> So that 'm' on 'main' is that one commit[1].\n> \n> \n>> I think you're misunderstanding the git meaning of rebase here. \"git \n>> rebase\" moves all commits since \"project\" forked from \"main branch\" to \n>> the tip of \"main branch\".\n> \n> \n> Right, I'm referring to 'rebase' generally, as a concept, not to \n> git-rebase specifically. E.g. git diff main..project is another way of \n> rebasing I think.\n> \n\nYes, but imo a poor one, as you're losing all the history. git *can* do \nwhat you want, but it was designed to maintain a long history so that \neveryone can see it and improve on the code with many chains of small \nand simultanous changes.\n\n\n>> Other than that, this is the recommended workflow, and exactly how \n>> Linux and git both are managed (i.e. topic branches eventually merged \n>> into 'master').\n> \n> \n> They're not rebased though, generally. They're pulled. Ie, in Linux and \n> git when 'project' is merged, things look like:\n> \n>     o---o-m--o-o-o--o----m   <- project\n>    /     /              / \\\n> o-o-o-o-o--o-o-o--o-o-o----m <- main branch\n> \n> The rest of the world sees /all/ the individual commits of 'project' \n> right? The traditional process for the case I'm thinking of results in \n> the 'main' tree seeing only /one/ single commit for the project.\n> \n\nPerhpas we have a nomenclature clash here. When you say \"one single \ncommit\", I can't help but thinking \"snapshot\". It's completely \nimpossible to fold *ALL* the history into a single commit, and since you \nwant heuristics I would imagine you wouldn't want that either.\n\n\n>> I'm not sure what you mean by 'project_collapsed' though.\n> \n> \n> All the commits on the project branch are 'collapsed' into one single \n> commit/delta, and then that /single/ commit is merged to 'main'. Rest of \n> the world sees:\n> \n> o-o-o-o-o--o-o-o--o-o-o---m <- main branch\n>                        \\ /\n>                         o <- project\n> \n\nThe only sane way to represent this is by doing a mega-patch and \napplying it with a new commit message. That way renamed files will show \nup as\n\n\trenamed from /path/to/foo\n\trenamed to /path/to/some/where/else\n\nSince you're removing all the history in between one mega-patch and the \nnext (as if Linus would have v2.6.12 one day and in the next commit it \nwould be v2.6.13... strange thought), the history for that tree can't \nwell know about renames that doesn't exist in its history. Again, if you \nwan't to keep \"master\" (can we please call it that? I can't keep up with \nwhat you call \"project\" and \"main branch\") to a single commit you'll \nhave no history in it. In essence, that's a snapshot (or a release, \nwhich is just a snapshot with a tag).\n\n>> Personally I think metadata is evil.\n> \n> \n> Not sure I agree. Silly/redundant meta-data can be evil alright. But I'm \n> talking about meta-data which is not there and potentially not \n> reconstructable.\n> \n>> Renames will still be auto-detected anyway,\n> \n> \n> Chances are so, yes. Definitely with the git and Linux workflows.\n> \n> The traditional workflow for the software project I'm thinking of is \n> different though. One commit may encompass multiple renames and edits of \n> a file (discouraged, but it's possible).\n> \n> If my understanding is correct, following back history for such cases \n> would be difficult.\n> \n\nIt would be impossible. At best you can get \"before mega-patch 64, the \ntree looked like this\", \"after mega-patch 64, it looked like this, and \nhere are the files with 80% of above similarity index\".\n\n\n> There is an argument that that 'traditional' process should be changed. \n> However, leaving aside that argument, I'd like to know if git could \n> accomodate that process.\n> \n>> be able to detect a rename is if you rename a file and hack it up so \n>> it doesn't even come close to matching its origin (close in this case \n>> is 80% by default, I think). In those cases it isn't so much a rename \n>> as a rewrite.\n> \n> \n> Exactly - this is the case I'm concerned about. Imagine that you'd like \n> to be follow the history back through the rewrite and through to the \n> original file.\n> \n\nI'm confused. First you say you want to have one single mega-patch for \neach commit, then you say you want to be able to follow history back. \nIt's like deciding to throw away your wallet and then trying to get \nsomeone to pick it up and carry it around for you.\n\n>> IMO this is far better than having to tell git \"I renamed this file to \n>> that\", since it also detects code-copying with modifications, and it's \n>> usually quick enough to find those renames as well.\n> \n> \n> I think so too, but that involves arguing that very very long-standing \n> workflows should be changed to accomodate git. I intend to make that \n> argument to the 'project' concerned, however I would also like to be say \n> git could equally well deal with the 'traditional' workflow, modulo \n> having to explicitely use (say) git-mv.\n> \n\nThe simple fact is that once you start juggling 12MB patches instead of \nkeeping the commits, your history is out the window anyway. Adding \nmeta-data to accommodate for the lack of history when you throw it away \nis, to be honest, an approach that leaves \"insane\" in the dust.\n\nAs for convincing others, shove git-bisect under their noses and ask \nthem if they'd like a tool to find their bugs for them.\n\n\n>>\n>>     $ git checkout master\n>>     $ git pull . project\n> \n> \n> Right, but 'pull' isn't what I mean :).\n> \n> I mean:\n> \n>     $ git checkout project\n>     $ git pull . master\n>     $ git checkout -b tmp project\n>     $ git diff project..master | <git apply I think>\n>\n\nThis way, 'project' and 'tmp' both would hold all patches since you \nmerge 'master' into 'project' before creating the 'tmp' branch at the \nhead of 'project'. As such, 'project' is ahead of 'master' (it has its \nown changes, those in master and the merge between 'project' and \n'master'), so the diff will be empty.\n\nIf 'master' is where you commit regularly (i.e. not mega-patches), you \ncan do these two steps to create the mega-patch branch\n\n\t$ git checkout -b mega; # create the mega-patch branch\n\t$ # rewind the mega-patch branch to the dawn of time\n\t$ git reset --hard $(git rev-list HEAD | tail -n 1)\n\nAnd for each mega-patch, do this:\n\n\t$ # create and apply mega-patch 1\n\t$ git diff project..master | git apply\n\t$ # commit the changes we just applied\n\t$ git commit -s -a -m \"mega-patch 1\"\n\t$ git checkout project; # back to project branch\n\t$ # Merge with 'master', or the next mega-patch won't apply\n\t$ git pull . master\n\n\n>> Then you can apply patch-file to whatever branch you want and make the \n>> commit as if it was a single change-set. I'd recommend against it \n>> unless you're just toying around though. It's a bad idea to lie in a \n>> projects history.\n> \n> \n> Presume that 'project' in the workflow is defined as\n> \n>     \"achieve one goal with one commit to the master\"\n> \n> So by definition, it always correct that the project only ever has one \n> commit.\n> \n\nBut that can't be true either, unless you intend to stop working at the \nproject. At \"best\", you could be able to get a chain of commits in \n'master' where each commit hold several tons of changes.\n\nThe topic-branch approach to this would be to\na) Implement all changes required for a certain feature in one go and \ncommit all of them. do \"git pull . topic-branch\" when on master branch. \nThis will result in a \"fast-forward\" (i.e. top of 'master' is the \nmerge-base between 'master' and 'topic-branch'), so no merge will happen.\n\nb) Implement all changes required for a certain feature in small steps \nand then apply the diff between 'master..topic-branch' to master. The \ntopic-branch has to be thrown away, since it can't ever be merged back \ninto master, and master can't be merged into the topic-branch (that's \nok, topic-branches are made to throw away).\n\nFor small changes, or one change and some stupid bugfixes, I'd say b) is \na viable option. The kind of changes you talk about, with several \nrenames of files and sometimes near-complete rewrite of them, would \ncertainly warrant a merge (or a fast-forward).\n\n\n> The trouble is that /sometimes/ projects do indeed 'rename and rewrite' \n> a file. At present, chances are git might not notice this, and ability \n> to follow history through the rename+rewrite would be lost.\n> \n> I'm wondering whether:\n> \n> - this could be solved?\n\n\nNot with the mega-patch approach.\n\n> - how? (some additional advisory-only meta-data in the\n>   index-cache and commit?)\n> \n\nYou could maintain that data yourself in either an external or versioned \nfile. I've never heard of anyone employing the workflow you describe so \nI doubt it's very common. I also shudder to think that git will be made \nless efficient for the benefit of throwing history away, when tracking \nhistory efficiently is what it's all about in the first place.\n\n\n> If there is consensus on an acceptable way, I'm willing to implement it. \n> (I was thinking of just adding 'rename' headers to the commit objects, \n> then teaching diffcore to consider them in addition to current heuristics).\n> \n\nThe code is mightier than the mail. Perhaps if I see an implementation \nof this I could wrap my head around what you really mean. I'm sure I \nmust misunderstand you one way or another.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"295293","messageId":"46a038f90603011005m68af7485qfdfffb9f82717427@mail.gmail.com","threadId":"3502","inReplyTo":"Pine.LNX.4.64.0603011558390.13612@sheen.jakma.org","subject":"Re: impure renames / history tracking","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-03-01T18:05:44Z","receivedAt":"2006-03-01T18:05:44Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 3/2/06, Paul Jakma <paul@clubi.ie> wrote:\n> I mean:\n>\n>         $ git checkout project\n>         $ git pull . master\n>         $ git checkout -b tmp project\n>         $ git diff project..master | <git apply I think>\n\nThe moment you 'merge' by using git-diff | patch you lose all the\nsupport git gives you, because you are discarding all of git's\nmetadata! git's metadata is about all the commits you are merging, and\nis good enough that it will help future merges across renames.\n\nYou should really use git-pull/git-merge at that point.\n\nMy guess is that you do this to achieve what you describe later:\n\n> Presume that 'project' in the workflow is defined as\n>\n>         \"achieve one goal with one commit to the master\"\n>\n> So by definition, it always correct that the project only ever has\n> one commit.\n\nWhat happens if you rephrase that to read: \"achieve one goal with one\nmerge to the master\"? Long term, it gives you much better support from\nthe SCM. If a particular commit broke something, you can use\nwhatchanged, log, annotate and bisect to figure out in which /small/\ncommit things went astray.\n\nAnd you can modify your practices ever so slightly to match the\nbenefits of the old model:\n\n - force merge message editing in git-merge, and prepare appropriate\ncommit messages for your merges\n - write a modified git-log that displays only the merges to master\n\nthat way, you get the best of both worlds.\n\n> The trouble is that /sometimes/ projects do indeed 'rename and\n> rewrite' a file. At present, chances are git might not notice this,\n\nIt will, if you preserve git's metadata.\n\nThe thing is that with any scm that tracks metadata of some kind, the\nmoment you bypass its tools and do diff|patch to discard the\nmetadata... well, you lose its benefits...\n\nAnd what I've found, managing a project with 13K files, is that in\npractice git does far better tracking renames than several SCMs that\ndo explicit tracking. Don't be distracted by the 'we don't track\nrenames posturing'. We do, and it's so magic that it just works.\n\ncheers,\n\n\n"},{"id":"16969","messageId":"Pine.LNX.4.64.0603011815150.13612@sheen.jakma.org","threadId":"3502","inReplyTo":"Pine.LNX.4.64.0603010859200.22647@g5.osdl.org","subject":"Re: impure renames / history tracking","fromName":"Paul Jakma","fromEmail":"paul@clubi.ie","sentAt":"2006-03-01T18:50:21Z","receivedAt":"2006-03-01T18:50:21Z","isPatch":false,"sender":{"key":"paul@clubi.ie","avatar":null},"body":"Hi Linus,\n\nOn Wed, 1 Mar 2006, Linus Torvalds wrote:\n\n> The thing is, it does better than anything that _tries_ to be \n> \"reliable\".\n>\n> I can pretty much _guarantee_ that you can't do it better.\n\nI'm willing to take that argument to the 'project' concerned, I just \nneed to be pretty sure of it.\n\n> Tracking \"inodes\" - aka file identities - (which is what BK does, \n> and I assume what SVN does) is fundamentally problematic. I \n> particular, it's a horrible problem when two inodes \"meet\" under \n> the same name. You now have two identities for the same file, and \n> you're fundamentally screwed.\n\nYes, in that model it is. This interestingly, is not the BK model, I \nsuspect (see below).\n\n> It doesn't even need renames to be a problem. JUST THE FACT THAT \n> YOU TRY TO TRACK FILE \"IDENTITY\" HISTORY IS BROKEN.\n\nIf it's \"file identity\" globally across the lifetime of the project, \nI agree 100% per cent. The 'traditional' SCM concerned does this.\n\nThat's not what a solution I'd want to explore either, I'm only \ninterested in the identity of files for any one /one/ commit. In \nsaying that, I recognise it's pointless to try annotate file-change \ninformation in multi-parent commits (merges).\n\n> For example, take CVS, which doesn't actually try to do renames, \n> but _does_ try to track the identity of a file, since all the \n> history is tied into that identity: think about what happens in \n> Attic when a file is deleted. Completely broken model.\n\nACK, {Attic,deleted_files}/ is just horrid.\n\n> And that's really fundamental. CVS doesn't show the problems so \n> much, because CVS actively tries to make it hard to do these \n> things.\n\nACK.\n\n> With renames-tracking-file-identities, it's _really_ easy to get \n> some major confusion going. What happens when one branch creates a \n> file, and another one renames a file to that same name, and they \n> merge?\n\nWell, the conflict has to be resolved somehow, even today.\n\n> Don't tell me it doesn't happen. It happened under BK. The way BK \n> \"solved\" it was to keep the two separate identities: one of them \n> got resolved to the new filename, the other one went into the \n> \"deleted\" directory.\n\nRight. That's what the 'traditional workflow' SCM I'm thinking of \ndoes - not BK funnily enough, but an SCM predating BK which also \nhappens to use SCCS files, and with some of the same high-level \npush/pull constructs as BK (interestingly).\n\nIt also tracks name history globally using a deleted_files/ history, \nwhich is maintained, but I don't think it does this for name merges \nlike the above.\n\nIn the one I'm thinking of, it does (I /think/, I'm not an expert in \nit) the following:\n\nGiven two files, say:\n\n'old:\n\n1.1---1.2---1.3\n\nnew:\n\n1.1\n\n- constructs a 'fake' base SCCS revision, empty\n- adds the top 'old' version as a branch\n- adds the top new version as a new delta\n\n    1.1.1.1\n   /\n1.1---------1.2\n\nWhere in the merged file:\n\n \t1.1: empty\n \t1.1.1.1: was 1.3 from 'old'\n \t1.2: is 1.1 from 'new'\n\nHowever, it does /not/ create a deleted_files entry for the 'old' \nfile. (AFAICT - I may not have a sufficiently full understanding of \nthis SCM)\n\n> Guess what happens when the side that got merged into \"deleted\" \n> continues to edit the file? That's right - their edits happen on \n> the deleted file, and never show up in the real tree in a \n> subsequent merge ever again.\n\nIndeed - horrid.\n\n> And as far as I can tell, BK really did the best you can do. \n> Following file identities really _is_ fundamentally broken. It \n> sounds like a nice idea, but while you migth solve a few problems, \n> you create a whole raft of much more fundamental problems.\n\nFor tracking identity across more than one commit - I fully agree.\n\nThat's not what quite I'm thinking of though. Is it worth going on \nwith the discussion on a:\n\n \t 'track identities *only* from context of /the/ parent to\n           this commit'\n\n> So next time you think about a merge that migt have been improved \n> by tracking renames, please also think about a merge where one of \n> the filenames came from two or more different sources through an \n> earlier merge, and thank your benevolent Gods that they instructed \n> me to make git be based purely on file contents.\n\nOh, I agree muchely here.\n\nI wouldn't change git. I only wonder if it give its rename-heuristics \nan additional advisory-only hint? (for single-parent commits at least \n- never merges - and only on a per-commit basis).\n\nI probably should first explore how git deals with rename clashes..\n\nregards,\n-- \nPaul Jakma\tpaul@clubi.ie\tpaul@jakma.org\tKey ID: 64A2FF6A\nFortune:\nI'm glad I was not born before tea.\n \t\t-- Sidney Smith (1771-1845)\n"},{"id":"16971","messageId":"Pine.LNX.4.64.0603011851430.13612@sheen.jakma.org","threadId":"3502","inReplyTo":"46a038f90603011005m68af7485qfdfffb9f82717427@mail.gmail.com","subject":"Re: impure renames / history tracking","fromName":"Paul Jakma","fromEmail":"paul@clubi.ie","sentAt":"2006-03-01T19:13:36Z","receivedAt":"2006-03-01T19:13:36Z","isPatch":false,"sender":{"key":"paul@clubi.ie","avatar":null},"body":"On Thu, 2 Mar 2006, Martin Langhoff wrote:\n\n> The moment you 'merge' by using git-diff | patch you lose all the \n> support git gives you, because you are discarding all of git's \n> metadata! git's metadata is about all the commits you are merging, \n> and is good enough that it will help future merges across renames.\n\n> You should really use git-pull/git-merge at that point.\n\nLet's try not get stuck on the workflow.\n\nI probably shouldn't have brought it up. However, just assume it's \nbeen decided that 'detail' of the project implementation is too much \nclutter for the 'master'. I note that people do this already even in \nthe \"keep all the details\" Linux and Git workflows, where they \nrejiggle commits in order to cut-out 'oops, made a typo' type of \ncommits.\n\nSo the level of detail that is suitable is for 'merging upstream' \nclearly is arbitrary and subjective, and even with git and Linux that \nknob already is set past 0 (all detail), maybe to 1 - the workflow \nI'm thinking of has it set to (say) 2.\n\nFor sake of argument assume the workflow corresponds to:\n\n     o-o-o-o---o--o\n    /              \\\n--o----------------m->\n\nAnd collapsing just the 'oops, made a typo' commits so it looks like:\n\n     o-----o------o\n    /              \\\n--o----------------m->\n\n\nThe /real/ point, other than workflow, is:\n\n- can we track 'rename and rewrite'?\n\n> And you can modify your practices ever so slightly to match the\n> benefits of the old model:\n\nI agree completely on the workflow argument, I intend to make it to \nthe project concerned ;).\n\n> And what I've found, managing a project with 13K files, is that in \n> practice git does far better tracking renames than several SCMs \n> that do explicit tracking. Don't be distracted by the 'we don't \n> track renames posturing'. We do, and it's so magic that it just \n> works.\n\nYep, I know. :).\n\nI just wonder if that magic could use additional hints (*not* Attic/ \ntype stuff, ick ye gods no! Agree fully there!). Cause 'rename and \nrewrite' it just does not get right.\n\nSimplest test-case (simulating 'rename and rewrite half the file') \nis:\n\n- create a one-line file\n- commit to git\n- mv it and add a line\n\nTo show:\n\n$ git status\nnothing to commit\n$ cat test\nfoo\n$ git-mv test toast\n$ echo bar >> toast\n$ git-update-index toast\n$ git status\n#\n# Updated but not checked in:\n#   (will commit)\n#\n#       deleted:  test\n#       new file: toast\n#\n\nA year later, someone comes along and looks at the history for \n'toast', they'll never know they can look back further by following \n'test'.\n\nI'd like to fix the above somehow, possibly by adding 'renamed test \ntoast' meta-data to index cache and commit objects. Having git-mv / \ngit-cp add that meta-data.\n\nThen diffcore using that meta-data as /advisory/ and auxilliary \ninformation *only* in /helping/ to determining renames, as an \nadditional input to its existing heuristics. This meta-data would not \nbe intrinsic to the operation git, it would /only/ be to aid humans \n(or their tools rather) in tracking back/forward through history.\n\nWould that be the best way to explore solving the above problem?\n\nregards,\n-- \nPaul Jakma\tpaul@clubi.ie\tpaul@jakma.org\tKey ID: 64A2FF6A\nFortune:\nHuman resources are human first, and resources second.\n \t\t-- J. Garbers\n"},{"id":"16972","messageId":"7v3bi2ey63.fsf@assigned-by-dhcp.cox.net","threadId":"3502","inReplyTo":"Pine.LNX.4.64.0603011851430.13612@sheen.jakma.org","subject":"Re: impure renames / history tracking","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-03-01T19:56:52Z","receivedAt":"2006-03-01T19:56:52Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Paul Jakma <paul@clubi.ie> writes:\n\n> For sake of argument assume the workflow corresponds to:\n>\n>     o-o-o-o---o--o\n>    /              \\\n> --o----------------m->\n>\n> And collapsing just the 'oops, made a typo' commits so it looks like:\n>\n>     o-----o------o\n>    /              \\\n> --o----------------m->\n>\n>\n> The /real/ point, other than workflow, is:\n>\n> - can we track 'rename and rewrite'?\n\nYes.  Especially the collapsing is 'oops, made a typo' kind.\n\nInterestingly enough, there are two levels of \"rename tracking\"\nthe current git does.  Whey you run \"git whatchanged -M\", you\nare looking at renames between each commit in the commit chain,\none step at a time.  There as long as the rename+rewrite does\nnot amount to too much rewrite, you would see what should be\ndetected as rename to be detected as renames.  I found the\ncurrent default threshold parameters to be about right, maybe a\nbit too tight sometimes, though.  If you want to loosen the\ndefault, you can specify similiarity index after -M.\n\nThe way recursive merge strategy uses the rename detection,\nunlike what whatchanged shows you, does not use chains of\ncommits down to the common merge base in order to detect renames\n(my recollection may be wrong here -- it's a while since I\nlooked at the recursive merge the last time).  It just looks at\nthe two heads being merged, and detects similarility between\nthem.  So it does not make _any_ difference with the current\nimplementation of recursive merge if you kept a history full of\n\"honest but disgusting\" commits or collapsed them into a history\nwith small number of \"cleaned up\" commits.\n\nOne thing it _could_ do (and you _could_ implement as another\nmerge strategy and call it \"pauls-rename\" merge) is to follow\nthe commit chain one by one down to the common merge base from\nboth heads being merged, and analyze rename history on the both\ncommit chains.  Then, you would get better rename+rewrite\ndetection than what it currently does.\n\nHOWEVER.\n\nIf you have that kind of rename-following merge, a workflow that\ncollapses a useful history into a single huge commit \"Ok, this\ncommit is a roll-up patch between version 2.6.14 and 2.6.15\"\nbecomes far less attractive than it currently already is.  At\nthat point, you _are_ throwing away useful history.\n"},{"id":"16979","messageId":"Pine.LNX.4.64.0603012105230.13612@sheen.jakma.org","threadId":"3502","inReplyTo":"7v3bi2ey63.fsf@assigned-by-dhcp.cox.net","subject":"Re: impure renames / history tracking","fromName":"Paul Jakma","fromEmail":"paul@clubi.ie","sentAt":"2006-03-01T21:25:08Z","receivedAt":"2006-03-01T21:25:08Z","isPatch":false,"sender":{"key":"paul@clubi.ie","avatar":null},"body":"Hi Junio,\n\nOn Wed, 1 Mar 2006, Junio C Hamano wrote:\n\n> Interestingly enough, there are two levels of \"rename tracking\" the \n> current git does.  Whey you run \"git whatchanged -M\", you are \n> looking at renames between each commit in the commit chain, one \n> step at a time.  There as long as the rename+rewrite does not \n> amount to too much rewrite, you would see what should be detected \n> as rename to be detected as renames.\n\nRight.\n\n> I found the current default threshold parameters to be about right, \n> maybe a bit too tight sometimes, though.  If you want to loosen the \n> default, you can specify similiarity index after -M.\n\nThat's one option.\n\nI'm wondering though if we couldn't also allow for users to \nadditionally encode naming 'hints', to aid this 'similarity' \ndetection process.\n\n> The way recursive merge strategy uses the rename detection, unlike \n> what whatchanged shows you, does not use chains of commits down to \n> the common merge base in order to detect renames (my recollection \n> may be wrong here -- it's a while since I looked at the recursive \n> merge the last time).  It just looks at the two heads being merged, \n> and detects similarility between them.  So it does not make _any_ \n> difference with the current implementation of recursive merge if \n> you kept a history full of \"honest but disgusting\" commits or \n> collapsed them into a history with small number of \"cleaned up\" \n> commits.\n\nI'm going to have to stare at this paragraph a lot longer and harder \nto understand it :).\n\n> One thing it _could_ do (and you _could_ implement as another merge \n> strategy and call it \"pauls-rename\" merge) is to follow the commit \n> chain one by one down to the common merge base from both heads \n> being merged, and analyze rename history on the both commit chains.\n\nRight, I was just thinking that while making tea actually. This could \nbe part of the 'collapsing' process. (or call it \"coalesce \ntoo-detailed commits\" process if that is less offensive to ones sense \nof process ;) ).\n\nActually, you're sort of suggesting following the chains in parallel, \nright? Ie in wall-clock time order, rather than chain order. And \ndoing name resolution across the 'to-be-merged' chains at each step \nof the way? Sort of a lesser subset of how other SCMs maintain state \nfor names globally?\n\nIt's not so much /resolving/ names I'm worried about in the first \nplace. It's there simply being no information in the first place to \nindicate (from one single-parent commit to the next) which names were \nrenamed.\n\n> Then, you would get better rename+rewrite detection than what it \n> currently does.\n\nBut if I follow the commit chain in order to try extract\n\n> HOWEVER.\n\n> If you have that kind of rename-following merge, a workflow that \n> collapses a useful history into a single huge commit \"Ok, this \n> commit is a roll-up patch between version 2.6.14 and 2.6.15\" \n> becomes far less attractive than it currently already is.  At that \n> point, you _are_ throwing away useful history.\n\nYes, I agree. And I am, as part of arguing git's case (several SCMs \nare being evaluated and considered, I'm the git proponent at the \nmoment), I'm going to suggest workflow ought to be re-evaluated to \nensure it is generally reasonable, rather than be kept for the sake \nof it keeping (particularly as it may be tailored to the \nneeds/limitations of $TRADITIONAL_SCM).\n\nHowever, I suspect at least some level of collapsing will be desired \n(just as it is with Linux and git).\n\nThe workflow issue is seperate from the 'impure rename' issue though, \neven if the workflow I gave as an example excerbates the issue, \n\"rename and rewrite half of it\" and hard-to-detect renames can still \noccur in the detailed git/linux workflows, surely?\n\nregards,\n-- \nPaul Jakma\tpaul@clubi.ie\tpaul@jakma.org\tKey ID: 64A2FF6A\nFortune:\nIf you really knew C++, you wouldn't even joke about putting it\nin the kernel.\n\n \t- Richard Johnson on linux-kernel\n"},{"id":"16985","messageId":"44061C59.20204@op5.se","threadId":"3502","inReplyTo":"Pine.LNX.4.64.0603012105230.13612@sheen.jakma.org","subject":"Re: impure renames / history tracking","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-03-01T22:12:41Z","receivedAt":"2006-03-01T22:12:41Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Just to cap off my own engagement in this discussion, here's the last \ntime rename detection was seriously discussed on the list:\n\nhttp://www.gelato.unsw.edu.au/archives/git/0504/0147.html\n\nIf you're going to implement something you might benefit from the \nsuggestions made there.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"16987","messageId":"Pine.LNX.4.64.0603012225560.13612@sheen.jakma.org","threadId":"3502","inReplyTo":"44061C59.20204@op5.se","subject":"Re: impure renames / history tracking","fromName":"Paul Jakma","fromEmail":"paul@clubi.ie","sentAt":"2006-03-01T22:28:25Z","receivedAt":"2006-03-01T22:28:25Z","isPatch":false,"sender":{"key":"paul@clubi.ie","avatar":null},"body":"On Wed, 1 Mar 2006, Andreas Ericsson wrote:\n\n> http://www.gelato.unsw.edu.au/archives/git/0504/0147.html\n\nIn terms of format, that's pretty much exactly what I was thinking, \nexcept it's been vetoed.\n\n> If you're going to implement something you might benefit from the \n> suggestions made there.\n\nCheers.\n\nIs there a correct way to extend the git header? To add meta-data \nthat normal git porcelain won't display? (there doesn't appear to \nbe..)\n\nregards,\n-- \nPaul Jakma\tpaul@clubi.ie\tpaul@jakma.org\tKey ID: 64A2FF6A\nFortune:\nZombie processes haunting the computer\n"},{"id":"16988","messageId":"7vk6bdeqb8.fsf@assigned-by-dhcp.cox.net","threadId":"3502","inReplyTo":"44061C59.20204@op5.se","subject":"Re: impure renames / history tracking","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-03-01T22:46:35Z","receivedAt":"2006-03-01T22:46:35Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Andreas Ericsson <ae@op5.se> writes:\n\n> Just to cap off my own engagement in this discussion, here's the last\n> time rename detection was seriously discussed on the list:\n>\n> http://www.gelato.unsw.edu.au/archives/git/0504/0147.html\n>\n> If you're going to implement something you might benefit from the\n> suggestions made there.\n\nAlso, today's #git log has some interesting material.\n\n\thttp://colabti.de/irclogger/irclogger_logs/git\n\nFor anybody who wants to discuss rename recording (not\ntracking), the following is a must-read:\n\n\thttp://article.gmane.org/gmane.comp.version-control.git/217\n"},{"id":"17069","messageId":"Pine.LNX.4.64.0603012129310.13612@sheen.jakma.org","threadId":"3502","inReplyTo":"4405DD35.8060804@op5.se","subject":"Re: impure renames / history tracking","fromName":"Paul Jakma","fromEmail":"paul@clubi.ie","sentAt":"2006-03-02T21:10:37Z","receivedAt":"2006-03-02T21:10:37Z","isPatch":false,"sender":{"key":"paul@clubi.ie","avatar":null},"body":"On Wed, 1 Mar 2006, Andreas Ericsson wrote:\n\n> Yes, but imo a poor one, as you're losing all the history.\n\nWell, not per se. You might keep the original 'detail' branch. It's a \nterminal branch obviously, you can't pull master's changes to it once \nthe aggregate patch goes into master. But you can keep it around.\n\n> git *can* do what you want, but it was designed to maintain a long \n> history so that everyone can see it and improve on the code with \n> many chains of small and simultanous changes.\n\nIndeed, and I appreciate that.\n\n> Perhpas we have a nomenclature clash here. When you say \"one single \n> commit\", I can't help but thinking \"snapshot\".\n\nI mean:\n\n \tgit diff upstream..bugfix_xyz\n\nor:\n\n \tgit diff upstream..project_foo_phase1\n\ntype of thing.\n\n> It's completely impossible to fold *ALL* the history into a single \n> commit, and since you want heuristics I would imagine you wouldn't \n> want that either.\n\nI want to know whether additional meta-data to help the existing \nheuristics would be acceptable. From a discussion on #git yesterday I \ngather the best way forward would to be to first prototype something \nkeeping state in a file in .git.\n\nAll that's needed really is something that relates the following 3 \nthings:\n\n \tcommit-id obj1-id obj2-id\n\nIe: For <commit-id>, <obj1-id> is similar to <obj2-id>.\n\nMaintaining this state could be done via the git-mv/rename wrappers \nand an additional git-edit wrapper. Those who are quite happy with \nthe existing diff-input only similarity heuristics wouldn't have to \nbother using a git-edit wrapper obviously, those who want to let git \ngather additional 'similarity hint' in this way could.\n\nAside:\n\nGit might be easier to extend generally if it adopted just /one/ new \ncore header, say \"see-also\" - that could serve as a pointer to \narbitrary commit-related meta-info objects that aren't of immediate \ninterest to either:\n\na) core git\n\nor\n\nb) the user\n\nFormat:\n\n \tsee-also <word> <obj-id>\n\nE.g.:\n\n \tsee-also similars <obj-id>\n\nWhere <obj-id> would list the 'commit obj1 obj2', but just as:\n\n \tobj1 obj2\n\nWould ultimately be neater than fishing around in .git/, and would \nallow other extensions in the future too.\n\nThe <word> identifier preferably would need to be centrally \nco-ordinated.\n\n> I'm confused. First you say you want to have one single mega-patch \n> for each commit, then you say you want to be able to follow history \n> back. It's like deciding to throw away your wallet and then trying \n> to get someone to pick it up and carry it around for you.\n\nI'm not sure why think mega-patch. Collapsing a bunch of commits \nrelated to one project need not result in a big patch relative to the \nrepository as a whole.\n\nIn Linux terms think project == \"Add ATAPI support to SATA\" or \n\"Change the foo VFS method and update its filesystem users\" type of \nthing (ok, the latter would be big enough, but still not /that/ big \nin terms of the whole Linux source base). Where the project concerned \nis like BSD, not just a kernel but a complete userland (so 1.1GB of \nsource code).\n\nI'm aware of the workflow arguments, I /do/ intend to make those but \nelsewhere ;).\n\n> As for convincing others, shove git-bisect under their noses and \n> ask them if they'd like a tool to find their bugs for them.\n\n;)\n\n[snip - thanks, interesting]\n\n> The code is mightier than the mail. Perhaps if I see an implementation of \n> this I could wrap my head around what you really mean. I'm sure I must \n> misunderstand you one way or another.\n\nYes, you're right. I think Junio gave me the required hints on \ndirections last night on #git.\n\nI think now at least it's quite possible to achieve without violating \ngit's \"track the /content/\" philosophy, via .git.\n\nThanks!\n\nregards,\n-- \nPaul Jakma\tpaul@clubi.ie\tpaul@jakma.org\tKey ID: 64A2FF6A\nFortune:\nFactorials were someone's attempt to make math LOOK exciting.\n"},{"id":"17075","messageId":"44076C48.7010207@op5.se","threadId":"3502","inReplyTo":"Pine.LNX.4.64.0603012129310.13612@sheen.jakma.org","subject":"Re: impure renames / history tracking","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-03-02T22:06:00Z","receivedAt":"2006-03-02T22:06:00Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Paul Jakma wrote:\n> On Wed, 1 Mar 2006, Andreas Ericsson wrote:\n> \n>> It's completely impossible to fold *ALL* the history into a single \n>> commit, and since you want heuristics I would imagine you wouldn't \n>> want that either.\n> \n> \n> I want to know whether additional meta-data to help the existing \n> heuristics would be acceptable. From a discussion on #git yesterday I \n> gather the best way forward would to be to first prototype something \n> keeping state in a file in .git.\n> \n> All that's needed really is something that relates the following 3 things:\n> \n>     commit-id obj1-id obj2-id\n> \n> Ie: For <commit-id>, <obj1-id> is similar to <obj2-id>.\n> \n> Maintaining this state could be done via the git-mv/rename wrappers and \n> an additional git-edit wrapper. Those who are quite happy with the \n> existing diff-input only similarity heuristics wouldn't have to bother \n> using a git-edit wrapper obviously, those who want to let git gather \n> additional 'similarity hint' in this way could.\n> \n> Aside:\n> \n> Git might be easier to extend generally if it adopted just /one/ new \n> core header, say \"see-also\" - that could serve as a pointer to arbitrary \n> commit-related meta-info objects that aren't of immediate interest to \n> either:\n> \n> a) core git\n> \n> or\n> \n> b) the user\n> \n\nThings that aren't of interest to either core git or the user is already \nhandled properly. It's called \"cruft\". ;)\n\nHowever, I see what you're trying for here. Something like the X-* \nheaders inside a mailer. Not all MUA's understand them, but if they do \nthey can make use of them to the users benefit.\n\n\n> Format:\n> \n>     see-also <word> <obj-id>\n> \n> E.g.:\n> \n>     see-also similars <obj-id>\n> \n> Where <obj-id> would list the 'commit obj1 obj2', but just as:\n> \n>     obj1 obj2\n> \n> Would ultimately be neater than fishing around in .git/, and would allow \n> other extensions in the future too.\n> \n> The <word> identifier preferably would need to be centrally co-ordinated.\n> \n\nWith X-* headers I don't see why it should have to be. Only the X-* part \nis mentioned in the RFC, so with a proper format Junio won't have to \ncoordinate cross-SCM tools, git-tortoise, etc, etc...\n\n\n>> I'm confused. First you say you want to have one single mega-patch for \n>> each commit, then you say you want to be able to follow history back. \n>> It's like deciding to throw away your wallet and then trying to get \n>> someone to pick it up and carry it around for you.\n> \n> \n> I'm not sure why think mega-patch. Collapsing a bunch of commits related \n> to one project need not result in a big patch relative to the repository \n> as a whole.\n> \n\nMainly I think it's because you mentioned several renames of a single \nfile and many files renamed + rewritten (beyond gits current ability of \nrecognizing it). That's definitely a mega-patch in my book.\n\n\n> Where the project concerned is like BSD, not \n> just a kernel but a complete userland (so 1.1GB of source code).\n> \n\n<just curious>\nSuch a large project surely must be split in several smaller \nsub-projects? GNU is, after all, several small (and not so small) \ncomponents. X works the same way. Linux is a large project, but each \ncompartment of code can be managed on its own, so long as they adhere to \nthe ABI hooking them back in to the kernel core.\n</just curious>\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"}]}