{"thread":{"id":"13925","subject":"[PATCH v2 2/3] filter-branch --blob-filter: speed/flexibility improvements.","startedAt":"2008-06-13T00:52:22Z","lastAt":"2008-06-13T16:10:16Z","messageCount":5,"participants":["Avery Pennarun","Jeff King"],"isPatch":true,"patchVersion":2,"patchTotal":3},"messages":[{"id":"79661","messageId":"1213318344-26013-1-git-send-email-apenwarr@gmail.com","threadId":"13925","inReplyTo":null,"subject":"[PATCH v2 1/3] filter-branch: add new --blob-filter option.","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2008-06-13T00:52:22Z","receivedAt":"2008-06-13T00:52:22Z","isPatch":true,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"From: Jeff King <peff@peff.net>\n\nOn Tue, Apr 22, 2008 at 12:51:14PM -0400, Avery Pennarun wrote:\n\n> Do you think git would benefit from having a generalized version of\n> this script?  Basically, the user provides a \"munge\" script on the\n> command line, and there's a git-filter-branch mode for auto-munging\n> (with a cache) every file in every checkin.  Even if it's *only* ever\n> used for CRLF, I can imagine this being useful to a lot of people.\n\nIt was easy enough to work up the patch below, which allows\n\n  git filter-branch --blob-filter 'tr a-z A-Z'\n\nHowever, it's _still_ horribly slow. Shell script is nice and flexible,\nbut running a tight loop like this is just painful. I suspect\nfilter-branch in something like perl would be a lot faster and just as\nflexible (you could even do it in C, but you'd probably have to invent a\nlittle domain-specific scripting language).\n\nIt is still much better performance than a tree filter, though:\n\n  $ cd git && time git filter-branch --tree-filter '\n      find . -type f | while read f; do\n        tr a-z A-Z <\"$f\" >tmp\n        mv tmp \"$f\"\n      done\n    ' HEAD~10..HEAD\n\n  real    4m38.626s\n  user    1m32.726s\n  sys     2m51.163s\n\n  $ cd git && git filter-branch --blob-filter 'tr a-z A-Z' HEAD~10..HEAD\n  real    1m40.809s\n  user    0m36.822s\n  sys     1m14.273s\n\nLots of system time in both. I'm sure we spend a fair bit of time\nhitting our very large map and blob-cache directories, which would be\nmuch more nicely implemented as associative arrays in memory (if we were\nusing a more featureful language).\n\nAnyway, here is the patch. I don't know if it is even worth applying,\nsince it is still painfully slow.\n\nAcked-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n git-filter-branch.sh |   30 ++++++++++++++++++++++++++++++\n 1 files changed, 30 insertions(+), 0 deletions(-)\n\ndiff --git a/git-filter-branch.sh b/git-filter-branch.sh\nindex d04c346..a0d9a79 100755\n--- a/git-filter-branch.sh\n+++ b/git-filter-branch.sh\n@@ -54,6 +54,23 @@ EOF\n \n eval \"$functions\"\n \n+munge_blobs() {\n+\twhile read mode sha1 stage path\n+\tdo\n+\t\tif ! test -r \"$workdir/../blob-cache/$sha1\"\n+\t\tthen\n+\t\t\tnew=`git cat-file blob $sha1 |\n+\t\t\t     eval \"$filter_blob\" |\n+\t\t\t     git hash-object -w --stdin`\n+\t\t\tprintf $new >$workdir/../blob-cache/$sha1\n+\t\tfi\n+\t\tprintf \"%s %s\\t%s\\n\" \\\n+\t\t\t\"$mode\" \\\n+\t\t\t$(cat \"$workdir/../blob-cache/$sha1\") \\\n+\t\t\t\"$path\"\n+\tdone\n+}\n+\n # When piped a commit, output a script to set the ident of either\n # \"author\" or \"committer\n \n@@ -105,6 +122,7 @@ tempdir=.git-rewrite\n filter_env=\n filter_tree=\n filter_index=\n+filter_blob=\n filter_parent=\n filter_msg=cat\n filter_commit='git commit-tree \"$@\"'\n@@ -150,6 +168,9 @@ do\n \t--index-filter)\n \t\tfilter_index=\"$OPTARG\"\n \t\t;;\n+\t--blob-filter)\n+\t\tfilter_blob=\"$OPTARG\"\n+\t\t;;\n \t--parent-filter)\n \t\tfilter_parent=\"$OPTARG\"\n \t\t;;\n@@ -227,6 +248,9 @@ ret=0\n # map old->new commit ids for rewriting parents\n mkdir ../map || die \"Could not create map/ directory\"\n \n+# cache rewritten blobs for blob filter\n+mkdir ../blob-cache || die \"Could not create blob-cache/ directory\"\n+\n case \"$filter_subdir\" in\n \"\")\n \tgit rev-list --reverse --topo-order --default HEAD \\\n@@ -295,6 +319,12 @@ while read commit parents; do\n \teval \"$filter_index\" < /dev/null ||\n \t\tdie \"index filter failed: $filter_index\"\n \n+\tif test -n \"$filter_blob\"; then\n+\t\tgit ls-files --stage |\n+\t\tmunge_blobs |\n+\t\tgit update-index --index-info\n+\tfi\n+\n \tparentstr=\n \tfor parent in $parents; do\n \t\tfor reparent in $(map \"$parent\"); do\n-- \n1.5.6.rc2.29.g4717e\n"},{"id":"79659","messageId":"1213318344-26013-2-git-send-email-apenwarr@gmail.com","threadId":"13925","inReplyTo":"1213318344-26013-1-git-send-email-apenwarr@gmail.com","subject":"[PATCH v2 2/3] filter-branch --blob-filter: speed/flexibility improvements.","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2008-06-13T00:52:23Z","receivedAt":"2008-06-13T00:52:23Z","isPatch":true,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"Export the current file path as $GIT_BLOB_PATH, so we can filter a blob\ndifferently based on its path, and change the caching mechanism to re-filter\na particular blob if its path changes.\n\nAlso, make it much faster by not calling 'cat'. The main loop of\nmunge_blobs() had to fork-exec \"cat\" every time through the loop, even when\na blob was already cached.  Let's use the sh builtin 'read' instead for a\nhuge speedup.\n\ncd git\ntime git filter-branch --blob-filter 'tr a-z A-Z' HEAD~10..HEAD\n\n(original --blob-filter)\nreal    3m58.569s\nuser    0m22.900s\nsys     3m32.030s\n\n(with 'cat' calls removed)\nreal\t1m11.931s\nuser\t0m8.520s\nsys\t1m2.900s\n\n(with 'cat' calls removed and blob cache already filled)\nreal\t0m19.660s\nuser\t0m3.930s\nsys\t0m15.720s\n\nSigned-off-by: Avery Pennarun <apenwarr@gmail.com>\n---\n Documentation/git-filter-branch.txt |   27 +++++++++++++++++++++++++++\n git-filter-branch.sh                |   27 +++++++++++++++++----------\n 2 files changed, 44 insertions(+), 10 deletions(-)\n\ndiff --git a/Documentation/git-filter-branch.txt b/Documentation/git-filter-branch.txt\nindex ea77f1f..0c5cd0f 100644\n--- a/Documentation/git-filter-branch.txt\n+++ b/Documentation/git-filter-branch.txt\n@@ -12,6 +12,7 @@ SYNOPSIS\n \t[--index-filter <command>] [--parent-filter <command>]\n \t[--msg-filter <command>] [--commit-filter <command>]\n \t[--tag-name-filter <command>] [--subdirectory-filter <directory>]\n+\t[--blob-filter <command]\n \t[--original <namespace>] [-d <directory>] [-f | --force]\n \t[<rev-list options>...]\n \n@@ -149,6 +150,16 @@ to other tags will be rewritten to point to the underlying commit.\n \tThe result will contain that directory (and only that) as its\n \tproject root.\n \n+--blob-filter <command>::\n+\tThis is the filter for modifying the contents of each file (blob) in\n+\tthe tree.  The contents of a file are provided on stdin, and the new\n+\tfile contents should be provided on stdout.  The pathname of the\n+\tblob in the current revision is in $GIT_BLOB_PATH. For efficiency,\n+\tthe before/after results of a given blob+filename are only\n+\tcalculated once and then cached, so your filter must always return\n+\tthe same output blob for any given input blob.  You might use this\n+\tfilter for converting CRLF to LF in all your files, for example.\n+\n --original <namespace>::\n \tUse this option to set the namespace where the original commits\n \twill be stored. The default value is 'refs/original'.\n@@ -196,6 +207,22 @@ git filter-branch --index-filter 'git update-index --remove filename' HEAD\n \n Now, you will get the rewritten history saved in HEAD.\n \n+To convert CRLF to LF in all your files using the \"fromdos\" program (be\n+careful: this will attempt to modify binary files too!):\n+\n+----------------------------------------------\n+git filter-branch --blob-filter 'fromdos' HEAD\n+----------------------------------------------\n+\n+To convert CRLF to LF in all your *.c and *.cpp files:\n+\n+---------------------------------------------------------\n+git filter-branch --blob-filter 'case \"$GIT_BLOB_PATH\" in\n+\t*.c|*.cpp) fromdos;;\n+\t*) cat;;\n+esac' HEAD\n+---------------------------------------------------------\n+\n To set a commit (which typically is at the tip of another\n history) to be the parent of the current initial commit, in\n order to paste the other history behind the current history:\ndiff --git a/git-filter-branch.sh b/git-filter-branch.sh\nindex a0d9a79..f1ee263 100755\n--- a/git-filter-branch.sh\n+++ b/git-filter-branch.sh\n@@ -55,19 +55,24 @@ EOF\n eval \"$functions\"\n \n munge_blobs() {\n-\twhile read mode sha1 stage path\n+\twhile read GIT_BLOB_MODE GIT_BLOB_SHA1 stage GIT_BLOB_PATH\n \tdo\n-\t\tif ! test -r \"$workdir/../blob-cache/$sha1\"\n+\t\texport GIT_BLOB_MODE GIT_BLOB_SHA1 GIT_BLOB_PATH\n+\t\tcachefile=\"$cachedir/$GIT_BLOB_SHA1/$GIT_BLOB_PATH\"\n+\t\tif ! test -r \"$cachefile\"\n \t\tthen\n-\t\t\tnew=`git cat-file blob $sha1 |\n-\t\t\t     eval \"$filter_blob\" |\n-\t\t\t     git hash-object -w --stdin`\n-\t\t\tprintf $new >$workdir/../blob-cache/$sha1\n+\t\t\tnew=$(git cat-file blob $GIT_BLOB_SHA1 |\n+\t\t\t      eval \"$filter_blob\" |\n+\t\t\t      git hash-object -w --stdin)\n+\t\t\tmkdir -p \"$(dirname \"$cachefile\")\"\n+\t\t\techo -n $new >\"$cachefile\"\n+\t\telse\n+\t\t\tread new <\"$cachefile\"\n \t\tfi\n \t\tprintf \"%s %s\\t%s\\n\" \\\n-\t\t\t\"$mode\" \\\n-\t\t\t$(cat \"$workdir/../blob-cache/$sha1\") \\\n-\t\t\t\"$path\"\n+\t\t\t\"$GIT_BLOB_MODE\" \\\n+\t\t\t\"$new\" \\\n+\t\t\t\"$GIT_BLOB_PATH\"\n \tdone\n }\n \n@@ -108,6 +113,7 @@ USAGE=\"[--env-filter <command>] [--tree-filter <command>] \\\n [--index-filter <command>] [--parent-filter <command>] \\\n [--msg-filter <command>] [--commit-filter <command>] \\\n [--tag-name-filter <command>] [--subdirectory-filter <directory>] \\\n+[--blob-filter <command>] \\\n [--original <namespace>] [-d <directory>] [-f | --force] \\\n [<rev-list options>...]\"\n \n@@ -249,7 +255,8 @@ ret=0\n mkdir ../map || die \"Could not create map/ directory\"\n \n # cache rewritten blobs for blob filter\n-mkdir ../blob-cache || die \"Could not create blob-cache/ directory\"\n+cachedir=\"$workdir/../blob-cache\"\n+mkdir \"$cachedir\" || die \"Could not create blob-cache/ directory\"\n \n case \"$filter_subdir\" in\n \"\")\n-- \n1.5.6.rc2.29.g4717e\n"},{"id":"79660","messageId":"1213318344-26013-3-git-send-email-apenwarr@gmail.com","threadId":"13925","inReplyTo":"1213318344-26013-2-git-send-email-apenwarr@gmail.com","subject":"[PATCH v2 3/3] filter-branch --blob-filter: add tests.","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2008-06-13T00:52:24Z","receivedAt":"2008-06-13T00:52:24Z","isPatch":true,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"Signed-off-by: Avery Pennarun <apenwarr@gmail.com>\n---\n t/t7003-filter-branch-blob.sh |   94 +++++++++++++++++++++++++++++++++++++++++\n 1 files changed, 94 insertions(+), 0 deletions(-)\n create mode 100755 t/t7003-filter-branch-blob.sh\n\ndiff --git a/t/t7003-filter-branch-blob.sh b/t/t7003-filter-branch-blob.sh\nnew file mode 100755\nindex 0000000..f18031e\n--- /dev/null\n+++ b/t/t7003-filter-branch-blob.sh\n@@ -0,0 +1,94 @@\n+#!/bin/sh\n+\n+test_description='git-filter-branch --blob-filter'\n+. ./test-lib.sh\n+\n+make_commit () {\n+\techo -n \"$2\" >\"$1\"\n+\tgit add \"$1\"\n+\tgit commit -a -m \"$1\"\n+}\n+\n+match_file() {\n+\tf=\"$(cat \"$1\")\"\n+\techo \"'$f'\" = \"'$2'\"\n+\ttest \"$f\" = \"$2\"\n+}\n+\n+myfilter() {\n+\tgit-filter-branch --blob-filter 'case \"$GIT_BLOB_PATH\" in '\"$1\"') echo -n REPLACEMENT;; *) cat;; esac' HEAD &&\n+\trm -rf .git/refs/original\n+}\n+\n+test_expect_success 'setup' '\n+\tmake_commit A \"textA\" &&\n+\tmake_commit \"Space file\" \"Space text\" &&\n+\tmake_commit B.txt \"textB\" &&\n+\tmake_commit C.jpg \"jpgC\" &&\n+\tgit checkout -b caching &&\n+\tmake_commit AA \"textA\" &&\n+\tmake_commit A \"textA2\" &&\n+\tgit checkout -b renames master &&\n+\tmkdir dir &&\n+\trm -f B.txt &&\n+\tmake_commit dir/B.jpg \"textB\" &&\n+\trm -f C.jpg &&\n+\tmake_commit dir/C.txt \"jpgC\"\n+'\n+\n+test_expect_success 'rewrite all' '\n+\tgit checkout -b rewrite1 master &&\n+\tgit-filter-branch --blob-filter echo\\ -n\\ \\$GIT_BLOB_PATH HEAD\n+\trm -rf .git/refs/original\n+'\n+\n+test_expect_success 'rewrite all - result' '\n+\tmatch_file A \"A\" &&\n+\tmatch_file \"Space file\" \"Space file\" &&\n+\tmatch_file B.txt \"B.txt\" &&\n+\tmatch_file C.jpg \"C.jpg\"\n+'\n+\n+countfilter() {\n+\trm -f counter\n+\texport P=\"$PWD\"\n+\tgit-filter-branch --blob-filter 'echo tick >>$P/counter; echo -n $GIT_BLOB_PATH' HEAD &&\n+\trm -rf .git/refs/original\n+}\n+\n+test_expect_success 'caching' '\n+\tgit checkout -b rewrite1b caching &&\n+\tcountfilter\n+'\n+\n+test_expect_success 'caching - result' '\n+\tmatch_file A \"A\" &&\n+\tmatch_file AA \"AA\" &&\n+\tmatch_file B.txt \"B.txt\" &&\n+\tmatch_file C.jpg \"C.jpg\" &&\n+\ttest \"$(cat counter | wc -l)\" = 6\n+'\n+\n+test_expect_success 'rewrite .txt only' '\n+\tgit checkout -b rewrite2 master &&\n+\tmyfilter \\*.txt\n+'\n+\n+test_expect_success 'rewrite .txt only - result' '\n+\tmatch_file A \"textA\" &&\n+\tmatch_file B.txt \"REPLACEMENT\" &&\n+\tmatch_file C.jpg \"jpgC\"\n+'\n+\n+test_expect_success 'rewrite with renames' '\n+\tgit checkout -b rewrite3 renames &&\n+\tmyfilter \\*.txt\n+'\n+\n+test_expect_success 'rewrite with renames - result' '\n+\tmatch_file A \"textA\" &&\n+\tmatch_file dir/B.jpg \"textB\" &&\n+\tmatch_file dir/C.txt \"REPLACEMENT\"\n+'\n+\n+test_done\n-- \n1.5.6.rc2.29.g4717e\n"},{"id":"79681","messageId":"20080613062546.GD26768@sigill.intra.peff.net","threadId":"13925","inReplyTo":"1213318344-26013-1-git-send-email-apenwarr@gmail.com","subject":"Re: [PATCH v2 1/3] filter-branch: add new --blob-filter option.","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-06-13T06:25:46Z","receivedAt":"2008-06-13T06:25:46Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jun 12, 2008 at 08:52:22PM -0400, Avery Pennarun wrote:\n\n> It was easy enough to work up the patch below, which allows\n> \n>   git filter-branch --blob-filter 'tr a-z A-Z'\n\nFirst, two procedural complaints:\n\n  1. We're supposed to be in rc freeze, so this is not a great time to\n     publish a new feature. ;)\n\n  2. When bringing back an old patch, please please please give at least\n     a little bit of cover letter context. \"Here is what happened last\n     time, here are the reasons this patch was not accepted before, and\n     here is {why I think it that decision was wrong, what I have done\n     to improve the patch, etc}.\n\nIIRC, the situation last time had two issues:\n\n  1. it was a one-off \"we're not sure if this is really useful\" patch\n\n  2. it was unclear whether paths should be available, and if they were,\n     there was an issue of encountering the same hash at two different\n     paths.\n\nI assume your answer to '1' is \"I have been using this and it is\nuseful\". And for '2', it looks like you have extended the cache\nmechanism to take into account the sha1 and the path, which I think is\nthe right solution (and I am pleased to see it looks like the final test\ncovers the exact situation I was concerned about).\n\nSo:\n\n(for 1/3):\nSigned-off-by: Jeff King <peff@peff.net>\n\n(for the others (and for 1/3, do I get to ack my own patch?)):\nAcked-by: Jeff King <peff@peff.net>\n\n-Peff\n"},{"id":"79736","messageId":"32541b130806130910w1975e092y192785fbab5c908@mail.gmail.com","threadId":"13925","inReplyTo":"20080613062546.GD26768@sigill.intra.peff.net","subject":"Re: [PATCH v2 1/3] filter-branch: add new --blob-filter option.","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2008-06-13T16:10:16Z","receivedAt":"2008-06-13T16:10:16Z","isPatch":true,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On 6/13/08, Jeff King <peff@peff.net> wrote:\n>   1. We're supposed to be in rc freeze, so this is not a great time to\n>      publish a new feature. ;)\n\nI thought that was what branches were for :)  Anyway, I only rarely\nget the chance to work on this stuff lately, so I guess I got excited.\n\n>   2. When bringing back an old patch, please please please give at least\n>      a little bit of cover letter context. \"Here is what happened last\n>      time, here are the reasons this patch was not accepted before, and\n>      here is {why I think it that decision was wrong, what I have done\n>      to improve the patch, etc}.\n\nWill do next time.\n\n>  IIRC, the situation last time had two issues:\n>\n>   1. it was a one-off \"we're not sure if this is really useful\" patch\n>\n>   2. it was unclear whether paths should be available, and if they were,\n>      there was an issue of encountering the same hash at two different\n>      paths.\n>\n>  I assume your answer to '1' is \"I have been using this and it is\n>  useful\". And for '2', it looks like you have extended the cache\n>  mechanism to take into account the sha1 and the path, which I think is\n>  the right solution (and I am pleased to see it looks like the final test\n>  covers the exact situation I was concerned about).\n\nYes, for #1 it is indeed useful.  I'm using git-svn on Windows with an\nIDE that auto-generates files with CRLF in them, and the translation\nof that is something roughly like \"ARRGH!\"  I have to re-fix the\nnewlines on various different branches at various times and this is\nthe best way I've found.  (Although I can also imagine using it for\nwhitespace fixes, etc.)\n\nYou are correct about #2.  I believe I've covered all the complaints\nthat were brought up at the time.\n\n>  (for 1/3):\n>  Signed-off-by: Jeff King <peff@peff.net>\n>\n>  (for the others (and for 1/3, do I get to ack my own patch?)):\n>  Acked-by: Jeff King <peff@peff.net>\n\nThanks!\n\nAvery\n"}]}