{"thread":{"id":"22","subject":"another perspective on renames.","startedAt":"2005-04-14T22:22:46Z","lastAt":"2005-04-15T14:47:01Z","messageCount":4,"participants":["C. Scott Ananian","Paul Jackson","Ingo Molnar"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"137","messageId":"Pine.LNX.4.61.0504141759440.7261@cag.csail.mit.edu","threadId":"22","inReplyTo":null,"subject":"another perspective on renames.","fromName":"C. Scott Ananian","fromEmail":"cscott@cscott.net","sentAt":"2005-04-14T22:22:46Z","receivedAt":"2005-04-14T22:22:46Z","isPatch":false,"sender":{"key":"cscott@cscott.net","avatar":"https://gravatar.com/avatar/3551c2aefb299a0c45807f7677f5b26d8a5be4a4af359b4bf4fabbdd1f2b990e?d=mp&s=160"},"body":"Perhaps our thinking is being clouded by 'how other SCMs do things' ---\ndo we *really* need extra rename metadata?  As Linus pointed out, as long \nas a commit is done immediately after a rename (ie before the renamed file \nis changed) the tree object contains all the information one needs: you \ncan notice that a given object's content-hash is named 'foo' in the first \nversion and 'bar' in the second version.\n\nIngo thought that this was insufficient because two *different* objects \n(ie having different revision histories) might be mutated to a point where \nthey had a *same* contents (and then would be condensed into a single \nblob).  But isn't that a feature of the git-fs history generally (ie not a \nrenaming-specific issue)?\n\nOne solution would be to invent a new 'file-revision-history' annotation \non top of git-fs in order to keep these derivation paths seperate...\n\n...but perhaps we might think of this as a 'feature' of our SCM instead?\nThe 'history' of a file may have join points where a single 'content' may \nhave been derived by two or more completely different paths.  Explicit \nguidance to the front-end tools is required to 'unmerge' these files after \nthis occurs (ie updating the directory cache for one, but not the others). \nThis makes sense for include/arch/{foo,bar}/baz.h, but maybe not so much \nfor (say) the empty file.\n\nAnyway, maybe it's worth thinking a little about an SCM in which this is a \nfeature, instead of (or in addition to) automatically assuming this is a \nbug we need to add infrastructure to work around.\n  --scott\n\nPBFORTUNE Soviet  cryptographic D5 SLBM MI5 CIA postcard WASHTUB [Hello to all my fans in domestic surveillance] \nexplosion Sigint Bush ODEARL FJHOPEFUL assassination Uzi Hussein Nader\n                          ( http://cscott.net/ )\n"},{"id":"186","messageId":"20050414221626.10c6c0e7.pj@engr.sgi.com","threadId":"22","inReplyTo":"Pine.LNX.4.61.0504141759440.7261@cag.csail.mit.edu","subject":"Re: another perspective on renames.","fromName":"Paul Jackson","fromEmail":"pj@engr.sgi.com","sentAt":"2005-04-15T05:16:26Z","receivedAt":"2005-04-15T05:16:26Z","isPatch":false,"sender":{"key":"pj@engr.sgi.com","avatar":null},"body":"Scott wrote:\n> Anyway, maybe it's worth thinking a little about an SCM in which this is a \n> feature, instead of (or in addition to) automatically assuming this is a \n> bug we need to add infrastructure to work around.\n\nAgreed.\n\nTo me, the main purpose in tracking renames is to obtain a deeper\nhistory of the line-by-line changes in a file.\n\n  ==> But that doesn't seem relevant here.\n\nLast I looked, git has no such history.  A given file contents\nis the indivisable atom of the git world, with no fine structure.\n\nThis is quite unlike classic SCM's, built on file formats that\ntrack source lines, not files, as the atomic unit.\n\nTo me, rename is a special case of the more general case of a\nbig chunk of code (a portion of a file) that was in one place\neither being moved or copied to another place.\n\nI wonder if there might be someway to use the tools that biologists use\nto analyze DNA sequences, to track the evolution of source code,\nidentifying things like common chunks of code that differ in just a few\nmutations, and presenting the history of the evolution, at selectable\nlevels of detail.\n\n-- \n                  I won't rest till it's the best ...\n                  Programmer, Linux Scalability\n                  Paul Jackson <pj@engr.sgi.com> 1.650.933.1373, 1.925.600.0401\n"},{"id":"195","messageId":"20050415082759.GA26112@elte.hu","threadId":"22","inReplyTo":"20050414221626.10c6c0e7.pj@engr.sgi.com","subject":"Re: another perspective on renames.","fromName":"Ingo Molnar","fromEmail":"mingo@elte.hu","sentAt":"2005-04-15T08:27:59Z","receivedAt":"2005-04-15T08:27:59Z","isPatch":false,"sender":{"key":"mingo@elte.hu","avatar":null},"body":"\n* Paul Jackson <pj@engr.sgi.com> wrote:\n\n> Scott wrote:\n> > Anyway, maybe it's worth thinking a little about an SCM in which this is a \n> > feature, instead of (or in addition to) automatically assuming this is a \n> > bug we need to add infrastructure to work around.\n> \n> Agreed.\n> \n> To me, the main purpose in tracking renames is to obtain a deeper\n> history of the line-by-line changes in a file.\n> \n>   ==> But that doesn't seem relevant here.\n> \n> Last I looked, git has no such history.  A given file contents is the \n> indivisable atom of the git world, with no fine structure.\n> \n> This is quite unlike classic SCM's, built on file formats that track \n> source lines, not files, as the atomic unit.\n\ni believe the fundamental thing to think about is not file or line or \nnamespace, but 'tracking developer intent'. While keeping in mind that \nGIT is not an SCM, all SCMs boil down to this single thing: being able \nto track what the developer did and why he did it - to be a useful tool \nlater on. (SCMs are for humans with bad limitations, who have this \nfundamental design bug and keep forgetting things.)\n\nthe basic question is, how much to track. The most extreme form of \ntracking (just for the sake of visualizing it) would be to have an \neye-position recognizing software attached to a webcam looking at the \ndeveloper, and then exactly mapping what he did, how long did he look at \none particular line of code and exactly what did he type when doing \nthat. [ Perhaps also a thought-reader module in addition, once one is \navailable. (combined with another module that removes all the swearing)]\n\nbut i think Linus is on the right track to suggest that \"the file names \ndont matter all that much, it's all about the content\". Global diffs \nmight track most types of plain renames, and if it gets it wrong - do we \ncare? Misdetection of renames can happen, but realistically only with \nsmall files and trivial code, which wont have alot of history.\n\nThe only serious type of misdetection would be if two large modules in \ntwo different places in the namespace happen to have exactly the same \ncontent but have a different history (because e.g. they were merged in \nvia two separate trees, one came from one tree, the other from the other \ntree), and the developer renamed both of them in the same commit: in \nsuch a case the global diff would have no way to figure out what the \nproper thread of history is. But is this a realistic scenario?  If the \ntwo files are nontrivial and have the same content, why werent they \nmerged in the namespace in the first place?\n\nthe moment we allow 'namespace' into the picture, things get complex and \nugly. Directory recursion is already a complexity that would have been \nnice to avoid.\n\n\tIngo\n"},{"id":"212","messageId":"Pine.LNX.4.61.0504151031330.27637@cag.csail.mit.edu","threadId":"22","inReplyTo":"20050414221626.10c6c0e7.pj@engr.sgi.com","subject":"Re: another perspective on renames.","fromName":"C. Scott Ananian","fromEmail":"cscott@cscott.net","sentAt":"2005-04-15T14:47:01Z","receivedAt":"2005-04-15T14:47:01Z","isPatch":false,"sender":{"key":"cscott@cscott.net","avatar":"https://gravatar.com/avatar/3551c2aefb299a0c45807f7677f5b26d8a5be4a4af359b4bf4fabbdd1f2b990e?d=mp&s=160"},"body":"On Thu, 14 Apr 2005, Paul Jackson wrote:\n\n> To me, rename is a special case of the more general case of a\n> big chunk of code (a portion of a file) that was in one place\n> either being moved or copied to another place.\n>\n> I wonder if there might be someway to use the tools that biologists use\n> to analyze DNA sequences, to track the evolution of source code,\n> identifying things like common chunks of code that differ in just a few\n> mutations, and presenting the history of the evolution, at selectable\n> levels of detail.\n\nThe rsync algorithm (http://samba.anu.edu.au/rsync/tech_report/node2.html) \nis probably a good place to start, although it is relatively sensitive to \nmutations.  It will be able to efficiently detect identical blocks larger \nthan some block size N (512 bytes or so for rsync).  You might well \nconsider smaller blocks to be irrelevant.  The data can be made \nconsiderably more useful to developers by canonicalizing before searching \n(ie, compressing whitespace to ' ', etc)[*].  Note that the identical \nregions do *not* have to line up on block boundaries; see the rsync \nalgorithm for more detail.\n\nI think Linus has made a persuasive case that the 'developer-friendly' \nfeatures of an SCM (ie annotate, log, and friends) can be built *on top* \nof GIT.   This is a perfect example.  Since the computation is non-trivial \n(although linear in the number of lines of code involved in the history of \na file; ie doesn't depend on the unrelated size of the archive), it might \nmake sense for the front-end SCM to maintain its own caches --- for \nexample, of the block and rolling checksums for each file required by the \nrsync algorithm.  The key point being that these are just *caches*, not \nessential history information, and can always be wiped and regenerated.\n\nThe nice 'feature' of this system (some may disagree, I guess) is that it \ndoes *not* depend on extensive programmer annotation of file changes (ie, \nchunk A in file B came from lines C-D of file D, or file E was once named \nF, etc).  By inferring history from content-similar files and blocks, it \nseems that it would be more able to generate useful results after \nimporting third-party sources, which may come in distinct 'releases' but \nlack explicit history annotations.\n   --scott\n\n[*] in general, i will be *glad* to see source-management move away from \nCVS' line-oriented style; there's no good reason we should still be worrying\nabout whitespace changes, etc.  When we build 'developer-friendly' tools \nwe should make every effort to auto-detect source code, image formats, \netc, and automatically perform appropriate canonicalization and \nbeautification of diffs, because this can be/should be/is entirely \nseparate from git's underlying storage representation.\n\nMk 48 PANCHO ZPSECANT MKDELTA SCRANTON D5 SLBM JMTRAX Delta Force \nMI6 SGUAT Khaddafi SMOTH interception mail drop SECANT PBSUCCESS Cocaine\n                          ( http://cscott.net/ )\n"}]}