{"thread":{"id":"352","subject":"Finding file revisions","startedAt":"2005-04-27T16:50:59Z","lastAt":"2005-04-28T22:27:54Z","messageCount":25,"participants":["Chris Mason","Linus Torvalds","Simon Fowler","Thomas Gleixner","David Woodhouse","Daniel Barkalow","Kay Sievers","Tony Luck","Thomas Glanzmann"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"1875","messageId":"200504271251.00635.mason@suse.com","threadId":"352","inReplyTo":null,"subject":"Finding file revisions","fromName":"Chris Mason","fromEmail":"mason@suse.com","sentAt":"2005-04-27T16:50:59Z","receivedAt":"2005-04-27T16:50:59Z","isPatch":false,"sender":{"key":"mason@suse.com","avatar":null},"body":"Hello everyone,\n\nI haven't seen a tool yet to find which changeset modified a given file, so \nI whipped up something.  The basic idea is to:\n\nfor each changeset in rev-list\n\tfor each file in diff-tree -r parent changeset\n\t\tmatch against desired files\n\nIs there a faster way?  This will scale pretty badly as the tree grows, but \nI usually only want to search back a few months in the history.  So, it \nmight make sense to limit the results by date or commit/tag.\n\nUsage:\nfile-changes [-c commit id] file1 ...\n\nThe file names can be perl regular expressions, and it will match any file \nstarting with the expression listed.  So \"file-changes fs/ext\" will show \neverything in ext2 and ext3.\n\nExample output:\n\ndiff-tree -r 56022b4d00cae3ff816d3ff05d9f8a80e1517c60 9bd104d712d710d53c35166e40bd5fe24caf893e\n8a796b48e757e56b50802c28abf28e0199c45ad9->2db368df614de4799be2d1baffb6563dbe1b8926 fs/ext2/inode.c\ndbc8fd9bab639b84b8cc94fdbbf850b1e4bf1b2b->a4cd819734ba2eea9d5d21039deca62057f72d44 fs/ext3/inode.c\ncat-file commit 9bd104d712d710d53c35166e40bd5fe24caf893e\n    tree cd4e40eae003e29c0d3be2aa769c3b572ab1b488\n    parent 56022b4d00cae3ff816d3ff05d9f8a80e1517c60\n    author mason <mason@coffee> 1114617717 -0400\n    committer mason <mason@coffee> 1114617717 -0400\n\n    comments go here\n\nThis is meant for cut n' paste.  If you find a changeset comment you like, \nrun the diff-tree -r command on the first line to see a diff of the \nchangeset (maybe I should add | diff-tree-helper here?)\n\n-chris\n\n\n#!/usr/bin/perl\n\nuse strict;\n\nmy $last;\nmy $ret;\nmy $i;\nmy @wanted = ();\nmy $matched;\nmy $argc = scalar(@ARGV);\nmy $commit;\n\nsub print_usage() {\n    print STDERR \"usage: file-changes [-c commit] file_list\\n\";\n    exit(1);\n}\n\nif ($argc < 1) {\n    print_usage();\n}\n\nfor ($i = 0 ; $i < $argc ; $i++)  {\n    if ($ARGV[$i] eq \"-c\") {\n    \tif ($i == $argc - 1) {\n\t    print_usage();\n\t}\n\t$commit = $ARGV[++$i];\n    } else {\n\tpush @wanted, $ARGV[$i];\n    }\n}\n\nif (!defined($commit)) {\n    $commit = `commit-id`;\n    if ($?) {\n    \tprint STDERR \"commit-id failed, try using -c to specify a commit\\n\";\n\texit(1);\n    }\n    chomp $commit;\n}\n\n$last = $commit;\n\nopen(RL, \"rev-list $commit|\") || die \"rev-list failed\";\nwhile(<RL>) {\n    chomp;\n    my $cur = $_;\n    $matched = 0;\n    if ($cur eq $last) {\n        next;\n    }\n    # rev-list gives us the commits from newest to oldest\n    open(DT, \"diff-tree -r $cur $last|\") || die \"diff-tree failed\";\n    while(<DT>) {\n        chomp;\n\tmy @words = split;\n\tmy $file = $words[3];\n\t# if the filename has whitespace, suck it in\n\tif (scalar(@words) > 4) {\n\t    if (m/$file(.*)/) {\n\t        $file .= $1;\n\t    }\n\t}\n\tforeach my $m (@wanted) {\n\t    if ($file =~ m/^$m/) {\n\t\tif (!$matched) {\n\t\t    print \"diff-tree -r $cur $last\\n\";\n\t\t}\n\t\tprint \"$words[2] $file\\n\";\n\t\t$matched = 1;\n\t    }\n\t}\n    }\n    close(DT);\n    if ($?) {\n\t$ret = $? >> 8;\n\tdie \"diff-tree failed with $ret\";\n    }\n    if ($matched) {\n\tprint \"cat-file commit $last\\n\";\n\topen(COMMIT, \"cat-file commit $last|\") || die \"cat-file $last failed\";\n\twhile(<COMMIT>) {\n\t    print \"    $_\";\n\t}\n\tclose(COMMIT);\n\tif ($?) {\n\t    $ret = $? >> 8;\n\t    die \"cat-file failed with $ret\";\n\t}\n\tprint \"\\n\";\n    }\n    $last = $cur;\n}\n\nclose(RL);\nif ($?) {\n    $ret = $? >> 8;\n    die \"rev-list failed with $ret\";\n}\n"},{"id":"1878","messageId":"Pine.LNX.4.58.0504271027460.18901@ppc970.osdl.org","threadId":"352","inReplyTo":"200504271251.00635.mason@suse.com","subject":"Re: Finding file revisions","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-27T17:34:55Z","receivedAt":"2005-04-27T17:34:55Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 27 Apr 2005, Chris Mason wrote:\n> \n> I haven't seen a tool yet to find which changeset modified a given file, so \n> I whipped up something.  The basic idea is to:\n> \n> for each changeset in rev-list\n> \tfor each file in diff-tree -r parent changeset\n> \t\tmatch against desired files\n> \n> Is there a faster way? \n\nYes. Tell \"diff-tree\" what your desired files are, and it will cut down \nthe amount of work by a _lot_ (because then diff-tree doesn't need to \nrecurse into subdirectories that don't matter).\n\nSo you should just do\n\n\tfor each changeset in rev-list\n\tdo \n\t\tdiff-tree -r parent changeset <file-list>\n\t...\n\ninstead. \n\n> This will scale pretty badly as the tree grows, but \n> I usually only want to search back a few months in the history.  So, it \n> might make sense to limit the results by date or commit/tag.\n\nWith more history, \"rev-list\" should do basically the right thing: it will\nbe constant-time for _recent_ commits, and it is linear time in how far\nback you want to go. Which seems quite reasonable.\n\nAnd diff-tree is obviously constant-time (and very fast at that, \nespecially if you limit it to just a few files, since then it won't even \nbother with any other subdirectories).\n\n\t\tLinus\n"},{"id":"1884","messageId":"200504271423.37433.mason@suse.com","threadId":"352","inReplyTo":"Pine.LNX.4.58.0504271027460.18901@ppc970.osdl.org","subject":"Re: Finding file revisions","fromName":"Chris Mason","fromEmail":"mason@suse.com","sentAt":"2005-04-27T18:23:37Z","receivedAt":"2005-04-27T18:23:37Z","isPatch":false,"sender":{"key":"mason@suse.com","avatar":null},"body":"On Wednesday 27 April 2005 13:34, Linus Torvalds wrote:\n> On Wed, 27 Apr 2005, Chris Mason wrote:\n> > Is there a faster way?\n>\n> Yes. Tell \"diff-tree\" what your desired files are, and it will cut down\n> the amount of work by a _lot_ (because then diff-tree doesn't need to\n> recurse into subdirectories that don't matter).\n\nThanks.  I originally called diff-tree without the file list so that I could \ndo the regexp matching, but this is probably one of those features that will \nnever get used.\n\nMy test case here is a tree with 400 commits, giving diff-tree the file list \nbrings us down from 16s to 9s on a cold cache.  Hot cache is about 1.5 \nseconds on both.\n\n>\n> > This will scale pretty badly as the tree grows, but\n> > I usually only want to search back a few months in the history.  So, it\n> > might make sense to limit the results by date or commit/tag.\n>\n> With more history, \"rev-list\" should do basically the right thing: it will\n> be constant-time for _recent_ commits, and it is linear time in how far\n> back you want to go. Which seems quite reasonable.\n>\n> And diff-tree is obviously constant-time (and very fast at that,\n> especially if you limit it to just a few files, since then it won't even\n> bother with any other subdirectories).\n\nUsually the question I will want to ask is \"how did foo.c change since tag X\", \nwhich usually won't go back more then a few months.   This should be \nreasonable, and I'd rather not slow down common operations adding extra \nindexing for the uncommon file-changes run.\n\nSo, new prog attached.  New usage:\n\nfile-changes [-c commit_id] [-s commit_id] file ...\n\n-c is the commit where you want to start searching\n-s is the commit where you want to stop searching\n\n-chris\n"},{"id":"2023","messageId":"1114627268.20916.8.camel@tglx.tec.linutronix.de","threadId":"352","inReplyTo":"Pine.LNX.4.58.0504271027460.18901@ppc970.osdl.org","subject":"Re: Finding file revisions","fromName":"Thomas Gleixner","fromEmail":"tglx@linutronix.de","sentAt":"2005-04-27T18:41:08Z","receivedAt":"2005-04-27T18:41:08Z","isPatch":false,"sender":{"key":"tglx@linutronix.de","avatar":null},"body":"On Wed, 2005-04-27 at 10:34 -0700, Linus Torvalds wrote:\n\n> > This will scale pretty badly as the tree grows, but \n> > I usually only want to search back a few months in the history.  So, it \n> > might make sense to limit the results by date or commit/tag.\n> \n> With more history, \"rev-list\" should do basically the right thing: it will\n> be constant-time for _recent_ commits, and it is linear time in how far\n> back you want to go. Which seems quite reasonable.\n\nWhich is quite horrible, if you have a 500k+ blobs repo.\n\nI know you are database allergic, but there a database is the correct\nsolution. Having stored all the relations of those file/tree/commit\nblobs in a database it takes <20ms to have a list of all those file\nblobs in historical order with some context information retrieved. Thats\nnot on a monster machine, its on an ordinary wallmart pc\n\ntglx\n\n\n"},{"id":"1924","messageId":"Pine.LNX.4.58.0504271506290.18901@ppc970.osdl.org","threadId":"352","inReplyTo":"200504271423.37433.mason@suse.com","subject":"Re: Finding file revisions","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-27T22:19:23Z","receivedAt":"2005-04-27T22:19:23Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 27 Apr 2005, Chris Mason wrote:\n> \n> So, new prog attached.  New usage:\n> \n> file-changes [-c commit_id] [-s commit_id] file ...\n> \n> -c is the commit where you want to start searching\n> -s is the commit where you want to stop searching\n\nYour script will do some funky stuff, because you incorrectly think that\nthe rev-list is sorted linearly. It's not. It's sorted in a rough\nchronological order, but you really can't do the \"last\" vs \"cur\" thing\nthat you do, because two commits after each other in the rev-list listing\nmay well be from two totally different branches, so when you compare one\ntree against the other, you're really doing something pretty nonsensical.\n\ndiff-tree will happily compare trees that aren't related, so it will \n\"work\" in a sense, but it doesn't actually do what you think it does ;)\n\nSo what you should do is basically something like\n\n\topen(RL, \"rev-list $commit|\") || die \"rev-list failed\";\n\twhile(<RL>) {\n\t\tchomp;\n\t\tmy $cur = $_;\n\n(so far so good) but then you should look at the _parents_ of that \ncommit, ie do (NOTE NOTE NOTE! I'm a total perl idiot, so I'm not going to \ndo this right):\n\n\t\topen(PARENT, \"cat-file commit $cur\") || die \"cat-file failed\");\n\t\twhile(<PARENT>) {\n\t\t\tchomp;\n\t\t\tmy @words = split;\n\t\t\tif ($words[1] == \"tree\")\n\t\t\t\tcontinue;\n\t\t\tif ($words[1] != \"parent\")\n\t\t\t\tbreak;\n\t\t\ttest_diff($cur, $words[2]);\n\t\t}\n\t\tclose(PARENT);\n\t}\n\tclose(RL);\n\nand now your \"test_diff()\" thing can do the tree diff.\n\nThat way you actually do \"tree-diff\" on the thing you should do, and it \nwill show you _which_ way it changed in a merge (ie if you hit a \nmerge-point, it will do a tree-diff against both parents, and show you \nwhich one had the difference - then you'll obviously usually see that same \ndifference later on when you dig down to the actual changeset that did it \ntoo).\n\nRemember: time is not a nice linear stream.\n\n\t\tLinus\n"},{"id":"1920","messageId":"200504271831.47830.mason@suse.com","threadId":"352","inReplyTo":"Pine.LNX.4.58.0504271506290.18901@ppc970.osdl.org","subject":"Re: Finding file revisions","fromName":"Chris Mason","fromEmail":"mason@suse.com","sentAt":"2005-04-27T22:31:47Z","receivedAt":"2005-04-27T22:31:47Z","isPatch":false,"sender":{"key":"mason@suse.com","avatar":null},"body":"On Wednesday 27 April 2005 18:19, Linus Torvalds wrote:\n> On Wed, 27 Apr 2005, Chris Mason wrote:\n> > So, new prog attached.  New usage:\n> >\n> > file-changes [-c commit_id] [-s commit_id] file ...\n> >\n> > -c is the commit where you want to start searching\n> > -s is the commit where you want to stop searching\n>\n> Your script will do some funky stuff, because you incorrectly think that\n> the rev-list is sorted linearly. It's not. It's sorted in a rough\n> chronological order, but you really can't do the \"last\" vs \"cur\" thing\n> that you do, because two commits after each other in the rev-list listing\n> may well be from two totally different branches, so when you compare one\n> tree against the other, you're really doing something pretty nonsensical.\n\nAha, didn't realize that one.  Thanks, I'll rework things here.\n\n-chris\n"},{"id":"2019","messageId":"20050428084156.GK17682@himi.org","threadId":"352","inReplyTo":"200504271831.47830.mason@suse.com","subject":"Re: Finding file revisions","fromName":"Simon Fowler","fromEmail":"simon@hollie.ento.csiro.au","sentAt":"2005-04-28T08:41:57Z","receivedAt":"2005-04-28T08:41:57Z","isPatch":false,"sender":{"key":"simon@hollie.ento.csiro.au","avatar":null},"body":"On Wed, Apr 27, 2005 at 06:31:47PM -0400, Chris Mason wrote:\n> On Wednesday 27 April 2005 18:19, Linus Torvalds wrote:\n> > On Wed, 27 Apr 2005, Chris Mason wrote:\n> > > So, new prog attached.  New usage:\n> > >\n> > > file-changes [-c commit_id] [-s commit_id] file ...\n> > >\n> > > -c is the commit where you want to start searching\n> > > -s is the commit where you want to stop searching\n> >\n> > Your script will do some funky stuff, because you incorrectly think that\n> > the rev-list is sorted linearly. It's not. It's sorted in a rough\n> > chronological order, but you really can't do the \"last\" vs \"cur\" thing\n> > that you do, because two commits after each other in the rev-list listing\n> > may well be from two totally different branches, so when you compare one\n> > tree against the other, you're really doing something pretty nonsensical.\n> \n> Aha, didn't realize that one.  Thanks, I'll rework things here.\n> \nI've got a version of this written in C that I've been working on\nfor a bit - some example output:\n\n+040000 tree    bfb75011c32589b282dd9c86621dadb0f0bb3866        ppc\n+100644 blob    5ba4fc5259b063dab6417c142938d987ee894fc0        ppc/sha1.c\n+100644 blob    c3c51aa4d487f2e85c02b0257c1f0b57d6158d76        ppc/sha1.h\n+100644 blob    e85611a4ef0598f45911357d0d2f1fc354039de4        ppc/sha1ppc.S\ncommit b5af9107270171b79d46b099ee0b198e653f3a24->a6ef3518f9ac8a1c46a36c8d27173b1f73d839c4\n\nYou run it as:\nfind-changes commit_id file_prefix ...\n\nThe file_prefix is a path prefix to match - it's not as flexible as\nregexes, but it shouldn't be too much less useful.\n\nSimon\n\n-- \nPGP public key Id 0x144A991C, or http://himi.org/stuff/himi.asc\n(crappy) Homepage: http://himi.org\ndoe #237 (see http://www.lemuria.org/DeCSS) \nMy DeCSS mirror: ftp://himi.org/pub/mirrors/css/ \n\n\nFind commits that changed files matching the prefix given on the command line.\n\nSigned-off-by: Simon Fowler <simon@dreamcraft.com.au>\n---\n\nIndex: Makefile\n===================================================================\n--- c3aa1e6b53cc59d5fbe261f3f859584904ae3a63/Makefile  (mode:100644 sha1:d73bea1cbb9451a89b03d6066bf2ed7fec32fd31)\n+++ uncommitted/Makefile  (mode:100644)\n@@ -38,7 +38,7 @@\n \tcat-file fsck-cache checkout-cache diff-tree rev-tree show-files \\\n \tcheck-files ls-tree merge-base merge-cache unpack-file git-export \\\n \tdiff-cache convert-cache http-pull rpush rpull rev-list git-mktag \\\n-\tdiff-tree-helper\n+\tdiff-tree-helper find-changes\n \n SCRIPT=\tcommit-id tree-id parent-id cg-Xdiffdo cg-Xmergefile \\\n \tcg-add cg-admin-lsobj cg-cancel cg-clone cg-commit cg-diff \\\nIndex: find-changes.c\n===================================================================\n--- /dev/null  (tree:c3aa1e6b53cc59d5fbe261f3f859584904ae3a63)\n+++ uncommitted/find-changes.c  (mode:100644 sha1:64c0c3627d84969ee1596b05f97705455fba1871)\n@@ -0,0 +1,279 @@\n+/*\n+ * find-changes.c - find the commits that changed a particular file.\n+ */\n+\n+#include \"cache.h\"\n+//#include \"revision.h\"\n+#include \"commit.h\"\n+#include <sys/param.h>\n+\n+/* \n+ * This is a simple tool that walks through the revisions cache and\n+ * checks the parent-child diffs to see if they include the given\n+ * filename. \n+ */\n+\n+static int recursive = 1;\n+static int found = 0;\n+\n+static char *malloc_base(const char *base, const char *path, int pathlen)\n+{\n+\tint baselen = strlen(base);\n+\tchar *newbase = malloc(baselen + pathlen + 2);\n+\tmemcpy(newbase, base, baselen);\n+\tmemcpy(newbase + baselen, path, pathlen);\n+\tmemcpy(newbase + baselen + pathlen, \"/\", 2);\n+\treturn newbase;\n+}\n+\n+static void update_tree_entry(void **bufp, unsigned long *sizep)\n+{\n+\tvoid *buf = *bufp;\n+\tunsigned long size = *sizep;\n+\tint len = strlen(buf) + 1 + 20;\n+\n+\tif (size < len)\n+\t\tdie(\"corrupt tree file\");\n+\t*bufp = buf + len;\n+\t*sizep = size - len;\n+}\n+\n+static const unsigned char *extract(void *tree, unsigned long size, const char **pathp, unsigned int *modep)\n+{\n+\tint len = strlen(tree)+1;\n+\tconst unsigned char *sha1 = tree + len;\n+\tconst char *path = strchr(tree, ' ');\n+\n+\tif (!path || size < len + 20 || sscanf(tree, \"%o\", modep) != 1)\n+\t\tdie(\"corrupt tree file\");\n+\t*pathp = path+1;\n+\treturn sha1;\n+}\n+\n+static int check_file(void *tree, unsigned long size, const char *base, const char *target);\n+\n+/* A whole sub-tree went away or appeared */\n+static int check_tree(void *tree, unsigned long size, const char *base, const char *target)\n+{\n+\tint retval = 0;\n+\n+\twhile (size && !retval) {\n+\t\tretval = check_file(tree, size, base, target);\n+\t\tupdate_tree_entry(&tree, &size);\n+\t}\n+\treturn retval;\n+}\n+\n+/* A file entry went away or appeared.\n+ * Check the entire subtree under this, and long_jmp() back to the parse_diffs()\n+ * function if we find the target. */\n+static int check_file(void *tree, unsigned long size, const char *base, const char *target)\n+{\n+\tunsigned mode;\n+\tconst char *path;\n+\tchar full_path[MAXPATHLEN + 1];\n+\tint pathlen, retval;\n+\tconst unsigned char *sha1 = extract(tree, size, &path, &mode);\n+\n+\tpathlen = snprintf(full_path, MAXPATHLEN, \"%s%s\", base, path);\n+\tif (!cache_name_compare(full_path, pathlen, target, strlen(target)))\n+\t\tfound = 1;\n+\n+\tif (recursive && S_ISDIR(mode)) {\n+\t\tchar type[20];\n+\t\tunsigned long size;\n+\t\tchar *newbase = malloc_base(base, path, strlen(path));\n+\t\tvoid *tree;\n+\n+\t\ttree = read_sha1_file(sha1, type, &size);\n+\t\tif (!tree || strcmp(type, \"tree\"))\n+\t\t\tdie(\"corrupt tree sha %s\", sha1_to_hex(sha1));\n+\n+\t\tretval = check_tree(tree, size, newbase, target);\n+\t\t\n+\t\tfree(tree);\n+\t\tfree(newbase);\n+\t\treturn retval;\n+\t}\n+\treturn 0;\n+}\n+\t\n+static int diff_tree_sha1(const unsigned char *old, const unsigned char *new, const char *base, const char *target);\n+\n+/* the diff-tree algorithm depends on compare_tree_entry returning basically\n+ * the same thing that memcmp() would on the filenames - this is important\n+ * because the directories are sorted, and hence you need to decide what */\n+static int compare_tree_entry(void *tree1, unsigned long size1, \n+\t\t\t      void *tree2, unsigned long size2, \n+\t\t\t      const char *base, const char *target)\n+{\n+\tunsigned mode1, mode2;\n+\tconst char *path1, *path2;\n+\tconst unsigned char *sha1, *sha2;\n+\tint cmp, pathlen1, pathlen2;\n+\n+\tif (found)\n+\t\treturn 0;\n+\n+\tsha1 = extract(tree1, size1, &path1, &mode1);\n+\tsha2 = extract(tree2, size2, &path2, &mode2);\n+\n+\tpathlen1 = strlen(path1);\n+\tpathlen2 = strlen(path2);\n+\tcmp = cache_name_compare(path1, pathlen1, path2, pathlen2);\n+\t/* these files are different - if this is a directory then the\n+\t * contents of the subtree are all different. So, we need to\n+\t * run over the subtree and see if our target is in there\n+\t * . . . */\n+\tif (cmp) {\n+\t\tcheck_file(tree1, size1, base, target);\n+\t\tcheck_file(tree2, size2, base, target);\n+\t\treturn cmp;\n+\t}\n+\n+\tif (!memcmp(sha1, sha2, 20) && mode1 == mode2)\n+\t\treturn 0;\n+\n+\t/*\n+\t * If the filemode has changed to/from a directory from/to a regular\n+\t * file, we need to consider it a remove and an add.\n+\t */\n+\tif (S_ISDIR(mode1) != S_ISDIR(mode2)) {\n+\t\tcheck_file(tree1, size1, base, target);\n+\t\tcheck_file(tree2, size2, base, target);\n+\t\treturn 0;\n+\t}\n+\n+\tif (recursive && S_ISDIR(mode1)) {\n+\t\tint retval;\n+\t\tchar *newbase = malloc_base(base, path1, pathlen1);\n+\t\tretval = diff_tree_sha1(sha1, sha2, newbase, target);\n+\t\tfree(newbase);\n+\t\treturn retval;\n+\t}\n+\t\n+\tcheck_file(tree1, size1, base, target);\n+\tcheck_file(tree2, size2, base, target);\n+\treturn 0;\n+}\n+\n+static int diff_tree(void *tree1, unsigned long size1, void *tree2, unsigned long size2, \n+\t\t     const char *base, const char *target)\n+{\n+\twhile (size1 | size2) {\n+\t\tif (!size1) {\n+\t\t\tcheck_file(tree2, size2, base, target);\n+\t\t\tupdate_tree_entry(&tree2, &size2);\n+\t\t\tcontinue;\n+\t\t}\n+\t\tif (!size2) {\n+\t\t\tcheck_file(tree1, size1, base, target);\n+\t\t\tupdate_tree_entry(&tree1, &size1);\n+\t\t\tcontinue;\n+\t\t}\n+\t\tswitch (compare_tree_entry(tree1, size1, tree2, size2, base, target)) {\n+\t\tcase -1:\n+\t\t\tupdate_tree_entry(&tree1, &size1);\n+\t\t\tcontinue;\n+\t\tcase 0:\n+\t\t\tupdate_tree_entry(&tree1, &size1);\n+\t\t\t/* Fallthrough */\n+\t\tcase 1:\n+\t\t\tupdate_tree_entry(&tree2, &size2);\n+\t\t\tcontinue;\n+\t\t}\n+\t\tdie(\"diff-tree: internal error\");\n+\t}\n+\treturn 0;\n+}\n+\n+static int diff_tree_sha1(const unsigned char *old, const unsigned char *new, const char *base,\n+\t\t\t  const char *target)\n+{\n+\tvoid *tree1, *tree2;\n+\tunsigned long size1, size2;\n+\tchar type[20];\n+\tint retval;\n+\n+\ttree1 = read_sha1_file(old, type, &size1);\n+\tif (!tree1 || strcmp(type, \"tree\"))\n+\t\tdie(\"unable to read source tree %s\", sha1_to_hex(old));\n+\ttree2 = read_sha1_file(new, type, &size2);\n+\tif (!tree2 || strcmp(type, \"tree\"))\n+\t\tdie(\"unable to read destination tree %s\", sha1_to_hex(new));\n+\tretval = diff_tree(tree1, size1, tree2, size2, base, target);\n+\tfree(tree1);\n+\tfree(tree2);\n+\treturn retval;\n+}\n+\n+static int process_diffs(struct commit *parent, struct commit *commit, const char *target)\n+{\n+\tfound = 0;\n+\tdiff_tree_sha1(parent->tree->object.sha1, commit->tree->object.sha1, \"\", target);\n+\tif (found)\n+\t\tprintf(\"%s\\n\", sha1_to_hex(commit->object.sha1));\n+\treturn 0;\n+}\n+\n+/*\n+ * Walk the set of parents, and collect a list of the objects. \n+ */\n+void process_commit(struct commit *item)\n+{\n+\tstruct commit_list *parents;\n+\n+\tif (parse_commit(item))\n+\t\tdie(\"unable to parse commit %s\", sha1_to_hex(item->object.sha1));\n+\t\n+\tparents = item->parents;\n+\twhile (parents) {\n+\t\tprocess_commit(parents->item);\n+\t\tparents = parents->next;\n+\t}\n+}\n+\n+/*\n+ * Usage: find-changes <parent-id> <filename>\n+ *\n+ * Note that this code will find the commits that change the given\n+ * file in the set of commits that are parents of the one given on the\n+ * command line.\n+ */ \n+\n+int main(int argc, char **argv)\n+{\n+\tint i;\n+\tchar sha1[20];\n+\tstruct commit *orig;\n+\n+\tif (argc != 3) \n+\t\tusage(\"find-changes <parent-id> <filename>\");\n+\t\t\n+\tget_sha1_hex(argv[1], sha1);\n+\torig = lookup_commit(sha1);\n+\tprocess_commit(orig);\n+\tmark_reachable(&lookup_commit(argv[1])->object, 1);\n+\n+\t/* this code needs to use tree.c to do most of the work - this\n+\t * will simplify things a lot. \n+\t * XXX: rewrite diff-tree.c to do the same. */\n+\t\n+\tfor (i = 0; i < nr_objs; i++) {\n+\t\tstruct object *obj = objs[i];\n+\t\tstruct commit *commit;\n+\t\tstruct commit_list *p;\n+\n+\t\tif (obj->type != commit_type)\n+\t\t\tcontinue;\n+\n+\t\tcommit = (struct commit *) obj;\n+\n+\t\tp = commit->parents;\n+\t\twhile (p) {\n+\t\t\tprocess_diffs(p->item, commit, argv[2]);\n+\t\t\tp = p->next;\n+\t\t}\n+\t}\n+\treturn 0;\n+}\n"},{"id":"2028","messageId":"200504280745.05505.mason@suse.com","threadId":"352","inReplyTo":"Pine.LNX.4.58.0504271506290.18901@ppc970.osdl.org","subject":"Re: Finding file revisions","fromName":"Chris Mason","fromEmail":"mason@suse.com","sentAt":"2005-04-28T11:45:04Z","receivedAt":"2005-04-28T11:45:04Z","isPatch":false,"sender":{"key":"mason@suse.com","avatar":null},"body":"On Wednesday 27 April 2005 18:19, Linus Torvalds wrote:\n> On Wed, 27 Apr 2005, Chris Mason wrote:\n> > So, new prog attached.  New usage:\n> >\n> > file-changes [-c commit_id] [-s commit_id] file ...\n> >\n> > -c is the commit where you want to start searching\n> > -s is the commit where you want to stop searching\n>\n> Your script will do some funky stuff, because you incorrectly think that\n> the rev-list is sorted linearly. It's not. It's sorted in a rough\n> chronological order, but you really can't do the \"last\" vs \"cur\" thing\n> that you do, because two commits after each other in the rev-list listing\n> may well be from two totally different branches, so when you compare one\n> tree against the other, you're really doing something pretty nonsensical.\n\nOne more rev that should work as you suggested Here's the example output \nfrom a cogito changeset with merges.  I print the diff-tree lines once for each \nmatching parent and then print the commit once.  It's very primitive, but\nhopefully some day someone will make a gui with happy clicky buttons\nfor changesets and filerevs.\n\ndiff-tree -r 2544d7558f0ce94ab9c163f5b67244f71d8c85b8 69eeae031bf5447e99b9274761e2361e8c5a944e\n618fdb616cebbd2fc9f1cddc0b6b75fd575250a1->3579b5fd1182679a39b83eaaa9dd0e7c970f4545 diff-tree.c\ndiff-tree -r 9831d8f86095edde393e495d7a55cab9d35d5d05 69eeae031bf5447e99b9274761e2361e8c5a944e\n2d2913b6b98ac836b43755b1304d2a838dad87dd->4f01bbbbb3fd0e53e9ce968f167b6dae68fcfa92 Makefile\ncat-file commit 69eeae031bf5447e99b9274761e2361e8c5a944e\n    tree 7510dc1b63e9e690ec73952e40a31e43af4b55bc\n    parent 2544d7558f0ce94ab9c163f5b67244f71d8c85b8\n    parent 9831d8f86095edde393e495d7a55cab9d35d5d05\n    author Petr Baudis <pasky@ucw.cz> 1114544917 +0200\n    committer Petr Baudis <xpasky@machine.sinus.cz> 1114544917 +0200\n\n    Merge with rsync://www.kernel.org/pub/linux/kernel/people/torvalds/git.git\n\n-chris\n"},{"id":"2029","messageId":"200504280756.58293.mason@suse.com","threadId":"352","inReplyTo":"20050428084156.GK17682@himi.org","subject":"Re: Finding file revisions","fromName":"Chris Mason","fromEmail":"mason@suse.com","sentAt":"2005-04-28T11:56:57Z","receivedAt":"2005-04-28T11:56:57Z","isPatch":false,"sender":{"key":"mason@suse.com","avatar":null},"body":"On Thursday 28 April 2005 04:41, Simon Fowler wrote:\n> I've got a version of this written in C that I've been working on\n> for a bit - some example output:\n>\n> +040000 tree    bfb75011c32589b282dd9c86621dadb0f0bb3866        ppc\n> +100644 blob    5ba4fc5259b063dab6417c142938d987ee894fc0        ppc/sha1.c\n> +100644 blob    c3c51aa4d487f2e85c02b0257c1f0b57d6158d76        ppc/sha1.h\n> +100644 blob    e85611a4ef0598f45911357d0d2f1fc354039de4       \n> ppc/sha1ppc.S commit\n> b5af9107270171b79d46b099ee0b198e653f3a24->a6ef3518f9ac8a1c46a36c8d27173b1f7\n>3d839c4\n>\n> You run it as:\n> find-changes commit_id file_prefix ...\n>\n> The file_prefix is a path prefix to match - it's not as flexible as\n> regexes, but it shouldn't be too much less useful.\n\nI dropped the regexes for speed with diff-tree, they weren't that important to \nme...The features I was going for are:\n\n1) ability to see the changeset comments in the output.\n2) ability to look for revs on more than one file at a time.  The single file \nlimit in bk revtool always bugged me.\n3) Some quick cut n' paste method to generate the changeset diff.  This is why \nI do diff-tree -r in the output, so I can just copy into a different window \nand go.\n\nYour c version would hopefully end up faster on cpu time by limiting the \nnumber of times we read/decompress the commit files.\n\n-chris\n"},{"id":"2031","messageId":"1114693318.27227.111.camel@hades.cambridge.redhat.com","threadId":"352","inReplyTo":"200504271423.37433.mason@suse.com","subject":"Re: Finding file revisions","fromName":"David Woodhouse","fromEmail":"dwmw2@infradead.org","sentAt":"2005-04-28T13:01:58Z","receivedAt":"2005-04-28T13:01:58Z","isPatch":false,"sender":{"key":"dwmw2@infradead.org","avatar":"https://gravatar.com/avatar/7afd4f07e0cf7d7e046ae2d23678296b37777c96488e6f3451e78a5514154ebd?d=mp&s=160"},"body":"On Wed, 2005-04-27 at 14:23 -0400, Chris Mason wrote:\n> Thanks.  I originally called diff-tree without the file list so that I could \n> do the regexp matching, but this is probably one of those features that will \n> never get used.\n\nWhen I added this functionality to diff-tree I didn't want to add regexp\nsupport, but I did make sure it could handle the simple case of \"changes\nwithin directory xxx/yyy\". It can also take _multiple_ names. \n\nAt the same time, I also posted a primitive script which attempted to do\nsomething similar to what you're doing. The output of rev-tree is\nuseless, as Linus pointed out. Chronological sorting is\ncounterproductive in all cases and should be avoided _everywhere_.\n\nMy script is based on the original 'gitlog.sh' script, which walks the\ncommit tree from the head to its parents. It lists only those commits\nwhere the file(s) in question actually changed, giving the commit ID and\nthe changes.\n\nThere's one problem with that already documented in my (attached) mail\n-- we don't print merge changesets where the file in the child is\nidentical to the file in all the parents, but the changeset in question\n_is_ relevant to the history because it's merging two branches on which\nthe file _independently_ changed.\n\nThe other problem is that we still don't have enough information to\npiece together the full tree. With each commit we print, we're also\nprinting the last _relevant_ child (see $lastprinted in the script). \n\nThat allows us to piece together most of the graph, but when we\neventually reach a commit which has already been processed (but not\nnecessarily _printed_, we just stop -- so we don't have useful parent\ninformation for the oldset change in each branch and can't tie it back\nto the point at which it branched. We know the _immediate_ parent, but\nthat parent isn't necessarily going to have been one of the commits we\nactually printed.\n\nI suspect the best way to do this is to start with a copy of rev-tree\nand do something like..\n\n\t1. Add a 'struct commit_list children' to 'struct commit'\n\n\t2. Make process_commit() set it correctly:\n@@ wherever @@ process_commit\n\t        while (parents) {\n\t                process_commit(parents->item->object.sha1);\n+\t                commit_list_insert(obj, &parents->item->children);\n\t                parents = parents->next;\n\t        }\n\n\t3. Check each 'interesting' commit to see if it affects the\n\t   file(s) in question.\n\t   \n\t4. Prune the tree: For each commit which isn't a merge and which\n\t   doesn't touch the file(s), just dump it from the tree,\n\t   changing the child pointer of its parent and the parent\n\t   pointer of its child accordingly to maintain the tree.\n\t   For each merge where there are no changes to the file(s)\n\t   between the merge point and the point at which the branch was\n\t   taken, drop that too.\n\n\t5. Print the remaining commits.\n\n\n-- \ndwmw2\n\n\nOn Wed, 2005-04-13 at 14:57 +0100, David Woodhouse wrote:\n> The plan is that this will also form the basis of a tool which will report the\n> revision tree for a given file, which is why I really want to avoid the\n> unnecessary recursion rather than just post-processing the output.\n\nScript attached. Its output is something like this:\n\ncommit 97c9a63e76bf667c21f24a5cfa8172aff0dd1294 child\n*100664->100644 blob    6e4064e920792d5b0219b9f8f55a38ab4a1af856->c1091cd15e2ed1be65b50eaa910f7b45c08d93ac      rev-tree.c\n\n--------------------------\ncommit 13b6f29ac1686955e15f0250f796362460b4992e child 97c9a63e76bf667c21f24a5cfa8172aff0dd1294\n*100644->100644 blob    5b3090780d49cc610339a19f070a5954dce9a8bc->c1091cd15e2ed1be65b50eaa910f7b45c08d93ac      rev-tree.c\n\n--------------------------\ncommit 6420f0732f695269c0e3f28e62ed4b9aa6578d9f child 13b6f29ac1686955e15f0250f796362460b4992e\n*100644->100644 blob    7429b9c4d0aab2e4a494eb4b65129a59da138106->5b3090780d49cc610339a19f070a5954dce9a8bc      rev-tree.c\n*100664->100644 blob    28a980482bf2053e022409cc3e50b2ad8adafd55->5b3090780d49cc610339a19f070a5954dce9a8bc      rev-tree.c\n\n <...>\n\nAs we walk the tree from the HEAD to its parents, we print only those\ncommits which modify the file(s) in question. We remember the last\ncommit we printed as we recurse, so that we can generate a complete\ngraph. The SHA-1 of the blobs themselves aren't good enough on their own\nbecause they're not guaranteed to be unique -- if the same change\nhappens on two different branches, the SHA-1 will be the same, and we\nwon't know how it fits together.\n\nAs it is, it's not quite perfect because I'm still omitting merge\ncommits where the resulting file is identical to the same file in _all_\nof the parents. So if we have the following tree (for the _file):\n\n       ----- (AB) ----,\n      /                \\ \n  (A) ------ (AB) ----- (AB) --,\n      \\                         \\\n       ----- (AC) --------------(ABC)\n\n(Where the delta A->AB is a trivial one-line fix which two people\nindependently reproduce, then they merge their trees together)\n\n.. the point where the two independent instances of (AB) are merged\ntogether won't be shown in the output of the attached script. The output\nwould show only this:\n\n       ----- (AB) ----,\n      /                \\ \n  (A) ------ (AB) ----- (ABC)\n      \\                /           \n       ----- (AC) ----'\n\nDo we care about this? Or is it good enough? I don't really want to emit\noutput for _every_ merge commit we traverse, just in _case_ it happens\nto be relevant later. Should just give in to the voices in my head which\nare telling me I should through the damn thing away and rewrite it in C?\n\nGiven this output, it should be possible to display a pretty graph of\nthe history of the file, and easily find both diffs and whole files.\nCreating a graphical tool which does this is left as an exercise for the\nreader.\n \n-- \ndwmw2\n"},{"id":"2032","messageId":"1114693800.27227.115.camel@hades.cambridge.redhat.com","threadId":"352","inReplyTo":"Pine.LNX.4.58.0504271506290.18901@ppc970.osdl.org","subject":"Re: Finding file revisions","fromName":"David Woodhouse","fromEmail":"dwmw2@infradead.org","sentAt":"2005-04-28T13:09:59Z","receivedAt":"2005-04-28T13:09:59Z","isPatch":false,"sender":{"key":"dwmw2@infradead.org","avatar":"https://gravatar.com/avatar/7afd4f07e0cf7d7e046ae2d23678296b37777c96488e6f3451e78a5514154ebd?d=mp&s=160"},"body":"On Wed, 2005-04-27 at 15:19 -0700, Linus Torvalds wrote:\n> Remember: time is not a nice linear stream.\n\nTime is neither nice nor linear. Time is a complete illusion.\n\nIf _any_ of your tools are using the time for _any_ purpose other than\nto display it to the user along with the author/committer information,\nthen you are probably making a mistake. \n\nRelative time does not represent the revision history of a distributed\nsystem which supports merges. Any correlation you think you see is\n_purely_ a coincidence.\n\n-- \ndwmw2\n\n"},{"id":"2033","messageId":"20050428131350.GL17682@himi.org","threadId":"352","inReplyTo":"200504280756.58293.mason@suse.com","subject":"Re: Finding file revisions","fromName":"Simon Fowler","fromEmail":"simon@dreamcraft.com.au","sentAt":"2005-04-28T13:13:50Z","receivedAt":"2005-04-28T13:13:50Z","isPatch":false,"sender":{"key":"simon@dreamcraft.com.au","avatar":null},"body":"On Thu, Apr 28, 2005 at 07:56:57AM -0400, Chris Mason wrote:\n> On Thursday 28 April 2005 04:41, Simon Fowler wrote:\n> > I've got a version of this written in C that I've been working on\n> > for a bit - some example output:\n> >\n> > +040000 tree    bfb75011c32589b282dd9c86621dadb0f0bb3866        ppc\n> > +100644 blob    5ba4fc5259b063dab6417c142938d987ee894fc0        ppc/sha1.c\n> > +100644 blob    c3c51aa4d487f2e85c02b0257c1f0b57d6158d76        ppc/sha1.h\n> > +100644 blob    e85611a4ef0598f45911357d0d2f1fc354039de4       \n> > ppc/sha1ppc.S commit\n> > b5af9107270171b79d46b099ee0b198e653f3a24->a6ef3518f9ac8a1c46a36c8d27173b1f7\n> >3d839c4\n> >\n> > You run it as:\n> > find-changes commit_id file_prefix ...\n> >\n> > The file_prefix is a path prefix to match - it's not as flexible as\n> > regexes, but it shouldn't be too much less useful.\n> \n> I dropped the regexes for speed with diff-tree, they weren't that important to \n> me...The features I was going for are:\n> \n> 1) ability to see the changeset comments in the output.\n\nI'll add a -v option tomorrow (it's 11pm here) to show the commit\ncomments in the output, if you like the idea.\n\n> 3) Some quick cut n' paste method to generate the changeset diff.  This is why \n> I do diff-tree -r in the output, so I can just copy into a different window \n> and go.\n> \nWould the two commit sha1s space seperated, rather than with the\n'->', be better for that? I'm a little reluctant to have it output\nthe full 'diff-tree -r' thing, since it's inconsistent with the\nother tools. \n\nSimon\n\n-- \nPGP public key Id 0x144A991C, or http://himi.org/stuff/himi.asc\n(crappy) Homepage: http://himi.org\ndoe #237 (see http://www.lemuria.org/DeCSS) \nMy DeCSS mirror: ftp://himi.org/pub/mirrors/css/ \n"},{"id":"2041","messageId":"Pine.LNX.4.58.0504280815120.18901@ppc970.osdl.org","threadId":"352","inReplyTo":"1114627268.20916.8.camel@tglx.tec.linutronix.de","subject":"Re: Finding file revisions","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-28T15:24:54Z","receivedAt":"2005-04-28T15:24:54Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 27 Apr 2005, Thomas Gleixner wrote:\n>\n> On Wed, 2005-04-27 at 10:34 -0700, Linus Torvalds wrote:\n> > \n> > With more history, \"rev-list\" should do basically the right thing: it will\n> > be constant-time for _recent_ commits, and it is linear time in how far\n> > back you want to go. Which seems quite reasonable.\n> \n> Which is quite horrible, if you have a 500k+ blobs repo.\n\nIt's _not_ linear in blobs. It doesn't care at all about them, in fact. \n\nIt's linear in how many revisions you go backwards. And I claim that you \ncan't do any better than that, without doing _really_ bad things.\n\n> I know you are database allergic, but there a database is the correct\n> solution.\n\nI disagree. I'm not database allergic, I just don't believe in the notion \nthat databases solve all the worlds problems.\n\n> Having stored all the relations of those file/tree/commit\n> blobs in a database it takes <20ms to have a list of all those file\n> blobs in historical order with some context information retrieved.\n\n.. and such an SCM will _suck_ for anything else.\n\nYou just made creating a commit etc much slower. You now have to update \nper-file information that you never updated before, and look at \ninformation that git simply doesn't _care_ about. \n\nRight now, when we create a new version, it's pretty much instantaneous.  \nExactly becaue we do not look at a _single_ file, and we don't care how\nthey changed from the \"previous\" version. We just write out the knowledge\nabout what the files are now.\n\nDoing a database of file changes would absolutely _suck_. Anybody who\nthinks that databases are magically faster than not using a database\ndoesn't understand basic physics. Things don't go faster just because you\ncall it a database. Things go faster by _doing_less_.\n\nNormally, a database does less by keeping indexes etc around, and the \nindexes require less work than the data itself. But git _does_ all of that \nalready. Git very much _is_ a database, it's just a specialized one.\n\nI dare you to show me wrong. I don't _care_ of you can show the revision\nhistory of a single file in 20ms. The easiest way to do that is with a\ndelta format, where the file information basically is single-file in the\nfirst place, and you just open the file and print out the results. Guess\nwhat? We've had that. It's called RCS/SCCS/CVS, and it's a piece of total\nand absolute crap. Exactly because single-file revisions simply do not\nmatter.\n\nIf you want to use a database, go wild. But use it as a _cache_. Then you \ncan build up the database of file revisions \"after the fact\", and always \nknow that your database is not the real data, it's just an index, and can \nbe thrown away and regenerated at will.\n\nThat way you don't add overhead to the stuff that actually matters, and \nthat git does a lot better than a general-purpose database could ever do.\n\n\t\t\tLinus\n"},{"id":"2045","messageId":"Pine.LNX.4.21.0504281147500.30848-100000@iabervon.org","threadId":"352","inReplyTo":"200504271251.00635.mason@suse.com","subject":"Re: Finding file revisions","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-04-28T16:08:09Z","receivedAt":"2005-04-28T16:08:09Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Wed, 27 Apr 2005, Chris Mason wrote:\n\n> I haven't seen a tool yet to find which changeset modified a given file, so \n> I whipped up something.  The basic idea is to:\n\nWhat is the answer supposed to be in the presence of merges? It seems like\nyou shouldn't report the merge that brought in the change, but rather\n(assuming it's available) the changeset that originally made it.\n\nThat is:\n\ngo through the history tree:\n  if a commit has a parent with a different version:\n    if it also has a parent with the same version as the child, ignore the\n      different parent(s) and enqueue the same parent(s)\n    otherwise, report it (for a single head, it's the original change; for\n      a merge, it merged two changes to the file)\n  otherwise, enqueue all the parents\n\nSorting by time is probably not useful, because there must be some source\nof the current version, and all paths going back, after ignoring versions\nthat were replaced by it in a merge, must go back to that source, so\ndepth-first search is fastest. (If there are multiple possible solutions,\nthen it means that multiple people applied the same patch, and any of them\nshould do).\n\nThis should be easy in C, but difficult in something that isn't generating\nthe history info itself.\n\n\t-Daniel\n*This .sig left intentionally blank*\n\n"},{"id":"2049","messageId":"1114706099.4212.25.camel@localhost.localdomain","threadId":"352","inReplyTo":"200504280745.05505.mason@suse.com","subject":"Re: Finding file revisions","fromName":"Kay Sievers","fromEmail":"kay.sievers@vrfy.org","sentAt":"2005-04-28T16:34:59Z","receivedAt":"2005-04-28T16:34:59Z","isPatch":false,"sender":{"key":"kay.sievers@vrfy.org","avatar":null},"body":"On Thu, 2005-04-28 at 07:45 -0400, Chris Mason wrote:\n> On Wednesday 27 April 2005 18:19, Linus Torvalds wrote:\n> > On Wed, 27 Apr 2005, Chris Mason wrote:\n> > > So, new prog attached.  New usage:\n> > >\n> > > file-changes [-c commit_id] [-s commit_id] file ...\n> > >\n> > > -c is the commit where you want to start searching\n> > > -s is the commit where you want to stop searching\n> >\n> > Your script will do some funky stuff, because you incorrectly think that\n> > the rev-list is sorted linearly. It's not. It's sorted in a rough\n> > chronological order, but you really can't do the \"last\" vs \"cur\" thing\n> > that you do, because two commits after each other in the rev-list listing\n> > may well be from two totally different branches, so when you compare one\n> > tree against the other, you're really doing something pretty nonsensical.\n> \n> One more rev that should work as you suggested Here's the example output \n> from a cogito changeset with merges.  I print the diff-tree lines once for each \n> matching parent and then print the commit once.  It's very primitive, but\n> hopefully some day someone will make a gui with happy clicky buttons\n> for changesets and filerevs.\n\nNot really happy clicky, but ... :)\n\nLook at the (history) link:\n  http://ehlo.org/~kay/gitweb.cgi?p=linux/kernel/git/torvalds/linux-2.6.git;a=commit;h=fb3b4ebc0be618dbcc2326482a83c920d51af7de\n\nKay\n\n"},{"id":"2044","messageId":"1114706876.20916.18.camel@tglx.tec.linutronix.de","threadId":"352","inReplyTo":"Pine.LNX.4.58.0504280815120.18901@ppc970.osdl.org","subject":"Re: Finding file revisions","fromName":"Thomas Gleixner","fromEmail":"tglx@linutronix.de","sentAt":"2005-04-28T16:47:56Z","receivedAt":"2005-04-28T16:47:56Z","isPatch":false,"sender":{"key":"tglx@linutronix.de","avatar":null},"body":"On Thu, 2005-04-28 at 08:24 -0700, Linus Torvalds wrote:\n> I disagree. I'm not database allergic, I just don't believe in the notion \n> that databases solve all the worlds problems.\n\nI never claimed, they did\n\n> You just made creating a commit etc much slower. You now have to update \n> per-file information that you never updated before, and look at \n> information that git simply doesn't _care_ about. \n\nI did not say, that such a fetaure should be included into git itself.\nThat was never my intention.\n\n> what? We've had that. It's called RCS/SCCS/CVS, and it's a piece of total\n> and absolute crap. Exactly because single-file revisions simply do not\n> matter.\n\nI agree that RCS is crap for distributed development, but seeing a\nchange in a file in the correct context is quite helpful at times.\n\n> If you want to use a database, go wild. But use it as a _cache_. Then you \n> can build up the database of file revisions \"after the fact\", and always \n> know that your database is not the real data, it's just an index, and can \n> be thrown away and regenerated at will.\n\nThats all I want to use it for. Exactly for of tracking information over\nvarious repos and longer time intervals.\n\ntglx\n\n\n"},{"id":"2051","messageId":"200504281305.23063.mason@suse.com","threadId":"352","inReplyTo":"Pine.LNX.4.21.0504281147500.30848-100000@iabervon.org","subject":"Re: Finding file revisions","fromName":"Chris Mason","fromEmail":"mason@suse.com","sentAt":"2005-04-28T17:05:21Z","receivedAt":"2005-04-28T17:05:21Z","isPatch":false,"sender":{"key":"mason@suse.com","avatar":null},"body":"On Thursday 28 April 2005 12:08, Daniel Barkalow wrote:\n> On Wed, 27 Apr 2005, Chris Mason wrote:\n> > I haven't seen a tool yet to find which changeset modified a given file,\n> > so I whipped up something.  The basic idea is to:\n>\n> What is the answer supposed to be in the presence of merges? It seems like\n> you shouldn't report the merge that brought in the change, but rather\n> (assuming it's available) the changeset that originally made it.\n\nBased on comments from Linus I did make it a little more merge aware.  But \nsince my tool was just to tide me over until someone fixed things in gui \nform, I didn't want to kill off too many brain cells coding it.\n\nIt sounds as though David's script is already has more merge brains then mine, \nand the git web stuff is pretty slick.  So it seems I didn't look hard enough \nbefore...\n\n-chris\n"},{"id":"2053","messageId":"12c511ca050428101070e12e74@mail.gmail.com","threadId":"352","inReplyTo":"1114706099.4212.25.camel@localhost.localdomain","subject":"Re: Finding file revisions","fromName":"Tony Luck","fromEmail":"tony.luck@gmail.com","sentAt":"2005-04-28T17:10:04Z","receivedAt":"2005-04-28T17:10:04Z","isPatch":false,"sender":{"key":"tony.luck@gmail.com","avatar":null},"body":"> Not really happy clicky, but ... :)\n> \n> Look at the (history) link:\n>   http://ehlo.org/~kay/gitweb.cgi?p=linux/kernel/git/torvalds/linux-2.6.git;a=commit;h=fb3b4ebc0be618dbcc2326482a83c920d51af7de\n\nLooks very useful.  Would it be possible to display the date (from the\ncommit) instead of\nthe 40-hex-char blobname (but have the link still point to the blob). \nLike this:\n\n2005-04-27 [PATCH] USB: MODALIAS change for bcdDevice\n2005-04-26 Merge with\nkernel.org:/pub/scm/linux/kernel/git/gregkh/driver-2.6.git/\n2005-04-26 Merge with kernel.org:/pub/scm/linux/kernel/git/gregkh/aoe-2.6.git/\n\nThat way you'd trade some screen space that is filled with hex numbers for some\nuseful information.  Dates could either be absolute (as in my\nexample), or relative\n(\"4 hours ago\", \"2 weeks ago\", etc.)\n\n-Tony\n"},{"id":"2058","messageId":"20050428172222.GF20834@cip.informatik.uni-erlangen.de","threadId":"352","inReplyTo":"12c511ca050428101070e12e74@mail.gmail.com","subject":"Re: Finding file revisions","fromName":"Thomas Glanzmann","fromEmail":"sithglan@stud.uni-erlangen.de","sentAt":"2005-04-28T17:22:22Z","receivedAt":"2005-04-28T17:22:22Z","isPatch":false,"sender":{"key":"sithglan@stud.uni-erlangen.de","avatar":null},"body":"Hello,\n\n> Looks very useful.  Would it be possible to display the date (from the\n> commit) instead of the 40-hex-char blobname (but have the link still\n> point to the blob).  Like this:\n\nFirst of all there is a date on the site and second I think the sha1\nhash much more useful than the date.\n\n\tThomas\n"},{"id":"2066","messageId":"1114715496.4212.36.camel@localhost.localdomain","threadId":"352","inReplyTo":"200504280745.05505.mason@suse.com","subject":"Re: Finding file revisions","fromName":"Kay Sievers","fromEmail":"kay.sievers@vrfy.org","sentAt":"2005-04-28T19:11:36Z","receivedAt":"2005-04-28T19:11:36Z","isPatch":false,"sender":{"key":"kay.sievers@vrfy.org","avatar":null},"body":"On Thu, 2005-04-28 at 07:45 -0400, Chris Mason wrote:\n> On Wednesday 27 April 2005 18:19, Linus Torvalds wrote:\n> > On Wed, 27 Apr 2005, Chris Mason wrote:\n> > > So, new prog attached.  New usage:\n> > >\n> > > file-changes [-c commit_id] [-s commit_id] file ...\n> > >\n> > > -c is the commit where you want to start searching\n> > > -s is the commit where you want to stop searching\n> >\n> > Your script will do some funky stuff, because you incorrectly think that\n> > the rev-list is sorted linearly. It's not. It's sorted in a rough\n> > chronological order, but you really can't do the \"last\" vs \"cur\" thing\n> > that you do, because two commits after each other in the rev-list listing\n> > may well be from two totally different branches, so when you compare one\n> > tree against the other, you're really doing something pretty nonsensical.\n> \n> One more rev that should work as you suggested Here's the example output \n> from a cogito changeset with merges.  I print the diff-tree lines once for each \n> matching parent and then print the commit once.  It's very primitive, but\n> hopefully some day someone will make a gui with happy clicky buttons\n> for changesets and filerevs.\n> \n> diff-tree -r 2544d7558f0ce94ab9c163f5b67244f71d8c85b8 69eeae031bf5447e99b9274761e2361e8c5a944e\n> 618fdb616cebbd2fc9f1cddc0b6b75fd575250a1->3579b5fd1182679a39b83eaaa9dd0e7c970f4545 diff-tree.c\n> diff-tree -r 9831d8f86095edde393e495d7a55cab9d35d5d05 69eeae031bf5447e99b9274761e2361e8c5a944e\n> 2d2913b6b98ac836b43755b1304d2a838dad87dd->4f01bbbbb3fd0e53e9ce968f167b6dae68fcfa92 Makefile\n> cat-file commit 69eeae031bf5447e99b9274761e2361e8c5a944e\n>     tree 7510dc1b63e9e690ec73952e40a31e43af4b55bc\n>     parent 2544d7558f0ce94ab9c163f5b67244f71d8c85b8\n>     parent 9831d8f86095edde393e495d7a55cab9d35d5d05\n>     author Petr Baudis <pasky@ucw.cz> 1114544917 +0200\n>     committer Petr Baudis <xpasky@machine.sinus.cz> 1114544917 +0200\n\nCan you confirm this with the kernel tree? \n  file-changes -c 9acf6597c533f3d5c991f730c6a1be296679018e drivers/usb/core/usb.c\n\nlists the commit:\n  diff-tree -r 1d66c64c3cee10a465cd3f8bd9191bbeb718f650 c79bea07ec4d3ef087962699fe8b2f6dc5ca7754\n  f0534ee064901d0108eb7b2b1fcb59a98bb53c2b->c231b4bef314284a168fedb6c5f6c47aec5084fc drivers/usb/core/usb.c\n  cat-file commit c79bea07ec4d3ef087962699fe8b2f6dc5ca7754\n\nwhich seems not to have changed the file asked for.\n\nThanks,\nKay\n\n"},{"id":"2077","messageId":"200504281658.39300.mason@suse.com","threadId":"352","inReplyTo":"1114715496.4212.36.camel@localhost.localdomain","subject":"Re: Finding file revisions","fromName":"Chris Mason","fromEmail":"mason@suse.com","sentAt":"2005-04-28T20:58:38Z","receivedAt":"2005-04-28T20:58:38Z","isPatch":false,"sender":{"key":"mason@suse.com","avatar":null},"body":"On Thursday 28 April 2005 15:11, Kay Sievers wrote:\n>\n> Can you confirm this with the kernel tree?\n>   file-changes -c 9acf6597c533f3d5c991f730c6a1be296679018e\n> drivers/usb/core/usb.c\n>\n> lists the commit:\n>   diff-tree -r 1d66c64c3cee10a465cd3f8bd9191bbeb718f650\n> c79bea07ec4d3ef087962699fe8b2f6dc5ca7754\n> f0534ee064901d0108eb7b2b1fcb59a98bb53c2b->c231b4bef314284a168fedb6c5f6c47ae\n>c5084fc drivers/usb/core/usb.c cat-file commit\n> c79bea07ec4d3ef087962699fe8b2f6dc5ca7754\n>\n> which seems not to have changed the file asked for.\n\nHmmm, that does work here:\n\ncoffee:/src/git # diff-tree -r 1d66c64c3cee10a465cd3f8bd9191bbeb718f650 c79bea07ec4d3ef087962699fe8b2f6dc5ca7754 | grep usb.core.usb.c\n*100644->100644 blob    f0534ee064901d0108eb7b2b1fcb59a98bb53c2b->c231b4bef314284a168fedb6c5f6c47aec5084fc      drivers/usb/core/usb.c\n\n-chris\n"},{"id":"2084","messageId":"Pine.LNX.4.58.0504281428300.18901@ppc970.osdl.org","threadId":"352","inReplyTo":"200504281658.39300.mason@suse.com","subject":"Re: Finding file revisions","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-28T21:32:34Z","receivedAt":"2005-04-28T21:32:34Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 28 Apr 2005, Chris Mason wrote:\n\n> On Thursday 28 April 2005 15:11, Kay Sievers wrote:\n> >\n> > Can you confirm this with the kernel tree?\n> >   file-changes -c 9acf6597c533f3d5c991f730c6a1be296679018e drivers/usb/core/usb.c\n> >\n> > lists the commit:\n> >   diff-tree -r 1d66c64c3cee10a465cd3f8bd9191bbeb718f650 c79bea07ec4d3ef087962699fe8b2f6dc5ca7754\n> > f0534ee064901d0108eb7b2b1fcb59a98bb53c2b->c231b4bef314284a168fedb6c5f6c47aec5084fc drivers/usb/core/usb.c\n> >\n> >  cat-file commit c79bea07ec4d3ef087962699fe8b2f6dc5ca7754\n> >\n> > which seems not to have changed the file asked for.\n> \n> Hmmm, that does work here:\n> \n> coffee:/src/git # diff-tree -r 1d66c64c3cee10a465cd3f8bd9191bbeb718f650 c79bea07ec4d3ef087962699fe8b2f6dc5ca7754 | grep usb.core.usb.c\n> *100644->100644 blob    f0534ee064901d0108eb7b2b1fcb59a98bb53c2b->c231b4bef314284a168fedb6c5f6c47aec5084fc      drivers/usb/core/usb.c\n\nI think Key is confused by the fact that the commit is a -merge- commit, \nand the first parent has _not_ changed that file - it got changed through \nthe merge.\n\nIe:\n\n\tcat-file commit c79bea07ec4d3ef087962699fe8b2f6dc5ca7754\n\ngives\n\n\ttree 3fbdc4745cfde60df7d05815b343e4a253020530\n\tparent a9e4820c4c170b3df0d2185f7b4130b0b2daed2c\n\tparent 1d66c64c3cee10a465cd3f8bd9191bbeb718f650\n\tauthor Linus Torvalds <torvalds@ppc970.osdl.org.(none)> 1113921100 -0700\n\tcommitter Linus Torvalds <torvalds@ppc970.osdl.org.(none)> 1113921100 -0700\n\n\tMerge master.kernel.org:/pub/scm/linux/kernel/git/gregkh/i2c-2.6.git/\n\nand if you do a diff against the _first_ parent you don't see anything \nchanging in USB..\n\n\t\tLinus\n"},{"id":"2086","messageId":"1114723987.4212.51.camel@localhost.localdomain","threadId":"352","inReplyTo":"200504281658.39300.mason@suse.com","subject":"Re: Finding file revisions","fromName":"Kay Sievers","fromEmail":"kay.sievers@vrfy.org","sentAt":"2005-04-28T21:33:06Z","receivedAt":"2005-04-28T21:33:06Z","isPatch":false,"sender":{"key":"kay.sievers@vrfy.org","avatar":null},"body":"On Thu, 2005-04-28 at 16:58 -0400, Chris Mason wrote:\n> On Thursday 28 April 2005 15:11, Kay Sievers wrote:\n> >\n> > Can you confirm this with the kernel tree?\n> >   file-changes -c 9acf6597c533f3d5c991f730c6a1be296679018e\n> > drivers/usb/core/usb.c\n> >\n> > lists the commit:\n> >   diff-tree -r 1d66c64c3cee10a465cd3f8bd9191bbeb718f650\n> > c79bea07ec4d3ef087962699fe8b2f6dc5ca7754\n> > f0534ee064901d0108eb7b2b1fcb59a98bb53c2b->c231b4bef314284a168fedb6c5f6c47ae\n> >c5084fc drivers/usb/core/usb.c cat-file commit\n> > c79bea07ec4d3ef087962699fe8b2f6dc5ca7754\n> >\n> > which seems not to have changed the file asked for.\n> \n> Hmmm, that does work here:\n> \n> coffee:/src/git # diff-tree -r 1d66c64c3cee10a465cd3f8bd9191bbeb718f650 c79bea07ec4d3ef087962699fe8b2f6dc5ca7754 | grep usb.core.usb.c\n> *100644->100644 blob    f0534ee064901d0108eb7b2b1fcb59a98bb53c2b->c231b4bef314284a168fedb6c5f6c47aec5084fc      drivers/usb/core/usb.c\n> \n> -chris\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n\nSure. But file-changes lists the commit:\n  c79bea07ec4d3ef087962699fe8b2f6dc5ca7754\n\nwhen asked for:\n  \"drivers/usb/core/usb.c\"\n\nand that file isn't touched there. Actually it lists merge-commits which\nare not related to the file.\n\nThanks,\nKay\n\n\n"},{"id":"2091","messageId":"Pine.LNX.4.58.0504281445280.18901@ppc970.osdl.org","threadId":"352","inReplyTo":"1114723987.4212.51.camel@localhost.localdomain","subject":"Re: Finding file revisions","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-28T21:50:08Z","receivedAt":"2005-04-28T21:50:08Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 28 Apr 2005, Kay Sievers wrote:\n> \n> Sure. But file-changes lists the commit:\n>   c79bea07ec4d3ef087962699fe8b2f6dc5ca7754\n> \n> when asked for:\n>   \"drivers/usb/core/usb.c\"\n> \n> and that file isn't touched there. Actually it lists merge-commits which\n> are not related to the file.\n\nIt really _is_ touched by that commit. Look closer.\n\nIt has two parents: one that had already merged with Greg's USB tree, and \none that had _not_ done so.\n\nSo whether it \"modifies\" the USB files or not really depends on which \nparent you go back. \n\nIn general, you tend to want to ignore merge-nodes for looking at \ndifferences, but the differences are definitely there, and they are often \nvital (ie it's often _very_ important to know which side of a merge didn't \nchange something).\n\n\t\tLinus\n"},{"id":"2104","messageId":"200504281827.54553.mason@suse.com","threadId":"352","inReplyTo":"1114723987.4212.51.camel@localhost.localdomain","subject":"Re: Finding file revisions","fromName":"Chris Mason","fromEmail":"mason@suse.com","sentAt":"2005-04-28T22:27:54Z","receivedAt":"2005-04-28T22:27:54Z","isPatch":false,"sender":{"key":"mason@suse.com","avatar":null},"body":"On Thursday 28 April 2005 17:33, Kay Sievers wrote:\n> Sure. But file-changes lists the commit:\n>   c79bea07ec4d3ef087962699fe8b2f6dc5ca7754\n>\n> when asked for:\n>   \"drivers/usb/core/usb.c\"\n>\n> and that file isn't touched there. Actually it lists merge-commits which\n> are not related to the file.\n\nOk, this is what Daniel and David were talking about.  When we've got commit \nwith multiple parents, we'll find the file at least one more time than it was \nreally changed.  Looking at the results on git web, it's easy it ignore the \nmerge sets as noise, but it would be nice if we only printed the merge set \nwhen it made some change to the file the original cset being merged did not.\n\nI had misread your first mail, thinking that you had developed this \nindependently and solved these issues ;)  The problem is that if we do a true \ndepth first search, it seems like we'll have to keep a potentially unbounded \namount of data in order to find the first changeset that happened to create a \ngiven sha1.  I'd really rather print the mergeset and let the user figure it \nout.\n\nBut, we're not really printing a merge set so much as we're printing the \ncomplete diff of what was merged.  Is there some way to see what changes had \nto be done in order to resolve conflicts during a merge?\n\n-chris\n"}]}