{"thread":{"id":"52","subject":"Introductions","startedAt":"2005-04-16T16:50:15Z","lastAt":"2005-04-16T16:50:15Z","messageCount":1,"participants":["Zed A. Shaw"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"325","messageId":"1113670215.6025.36.camel@thamachine","threadId":"52","inReplyTo":null,"subject":"Introductions","fromName":"Zed A. Shaw","fromEmail":"zedshaw@zedshaw.com","sentAt":"2005-04-16T16:50:15Z","receivedAt":"2005-04-16T16:50:15Z","isPatch":false,"sender":{"key":"zedshaw@zedshaw.com","avatar":null},"body":"Hi,\n\nJust a short message to introduce myself and give a shameless plug.  I'm\nZed A. Shaw and I'm the author of a little unknown SCM called FastCST\n(http://www.zedshaw.com/projects/fastcst ).  While I doubt that Linus\nwould ever adopt fastcst as his tool (and I probably wouldn't want him\ntoo since it's not quite ready for prime time) I did find many of the\ndiscussions on the list so far very interesting.\n\nSome sent me Linus' message about wanting to do a diff on the whole\nsource tree, and just thought I'd mention that I already tried this in\nFastCST.  FastCST uses a suffix array to construct a delta (not a diff),\nso I thought it might be possible to simply apply the delta algorithm to\nthe entire source tree and get very small changesets.\n\nIt worked on small source trees, but when it came to the Linux 2.6 tree\nit choked hard.  Even with an efficient suffix array implementation,\nyou're talking about performing a diff/delta on 225M of source.  Added\nto the problem is that you have to track file locations within the\nmassive blob.  In the end, it also wasn't much more efficient from a\nsize/space/time perspective so I dropped it.\n\nMy current solution to Linus' problem is to use an inverted index to\nprocess all the sources and revisions on the fly as they are created.\nUsing the inverted index, I'm able to VERY quickly find any chunk of\nsource in files or revisions.  This lets me track things like how\nfunctions move through the files, where chunks of code moved to, etc.\nIn the end this turns out to be much more efficient (7 seconds on my\ncomputer to find all references to \"sprintf\" in the Linux 2.6 source) as\nI can use the super small deltas for distributing changes, and give\ndevelopers a means tracking content changes across \"the world\" in a\nsimple search format.\n\nAnyway, just thought I'd throw in my experiences attempting what Linus\nis talking about.  I actually agree with him that rename tracking isn't\nthat great, but I've come to the conclusion that tracking renames is\nactually a specific case of just a general search problem.  Different\nstrokes for different folks I guess.\n\nOther than that, I'm mostly interested in reading the messages and\nprobably won't write anything unless people ask me directly for\nsomething.  Thanks!\n\nZed A. Shaw\n\n"}]}