{"thread":{"id":"10459","subject":"[RESEND PATCH 0/9] Make git-svn fetch ~1.7x faster","startedAt":"2007-10-25T10:25:18Z","lastAt":"2007-10-27T01:02:05Z","messageCount":17,"participants":["Adam Roben","Eric Wong","Junio C Hamano","Brian Downing"],"isPatch":true,"patchVersion":1,"patchTotal":9},"messages":[{"id":"57159","messageId":"1193307927-3592-1-git-send-email-aroben@apple.com","threadId":"10459","inReplyTo":null,"subject":"[RESEND PATCH 0/9] Make git-svn fetch ~1.7x faster","fromName":"Adam Roben","fromEmail":"aroben@apple.com","sentAt":"2007-10-25T10:25:18Z","receivedAt":"2007-10-25T10:25:18Z","isPatch":true,"sender":{"key":"aroben@apple.com","avatar":"https://gravatar.com/avatar/9d3697e1de53890adf241331f4b970bdd2b18962b2ff0b8028ebb00e085807f8?d=mp&s=160"},"body":"\nThis is a resend of my previous patch series to speed up git-svn, taking into\naccount comments from Eric, Johannes, and Brian.\n\n--\n Documentation/git-cat-file.txt    |    6 +-\n Documentation/git-hash-object.txt |    5 +-\n builtin-cat-file.c                |   87 +++++++++++++++++----\n git-svn.perl                      |   40 +++++-----\n hash-object.c                     |   29 +++++++-\n perl/Git.pm                       |  153 ++++++++++++++++++++++++++++++++++++-\n t/t1005-cat-file.sh               |  126 ++++++++++++++++++++++++++++++\n t/t1006-hash-object.sh            |   49 ++++++++++++\n 8 files changed, 452 insertions(+), 43 deletions(-)\n"},{"id":"57160","messageId":"1193307927-3592-2-git-send-email-aroben@apple.com","threadId":"10459","inReplyTo":"1193307927-3592-1-git-send-email-aroben@apple.com","subject":"[PATCH 1/9] Add tests for git cat-file","fromName":"Adam Roben","fromEmail":"aroben@apple.com","sentAt":"2007-10-25T10:25:19Z","receivedAt":"2007-10-25T10:25:19Z","isPatch":true,"sender":{"key":"aroben@apple.com","avatar":"https://gravatar.com/avatar/9d3697e1de53890adf241331f4b970bdd2b18962b2ff0b8028ebb00e085807f8?d=mp&s=160"},"body":"\nSigned-off-by: Adam Roben <aroben@apple.com>\n---\nJohannes Sixt wrote:\n> Adam Roben schrieb:\n> > +    test_expect_success \\\n> > +        \"$type exists\" \\\n> > +        \"git cat-file -e $hello_sha1\"\n> \n> You mean $sha1 here, right?\n\nI most definitely did!\n\n> > +    test_expect_success \\\n> > +        \"Type of $type is correct\" \\\n> > +        \"test $type = \\\"$(git cat-file -t $sha1)\\\"\"\n> \n> This should escape the $(...) in all the tests. Like this:\n> \n>         \"test $type = \\\"\\$(git cat-file -t $sha1)\\\"\"\n> \n> > +test_expect_success \\\n> > +    \"Reach a blob from a tag pointing to it\" \\\n> > +    \"test \\\"$hello_content\\\" = \\\"$(git cat-file blob $tag_sha1)\\\"\"\n> \n> And use single quotes without escaping the double-quotes here. \n\nDone.\n\n t/t1005-cat-file.sh |   91 +++++++++++++++++++++++++++++++++++++++++++++++++++\n 1 files changed, 91 insertions(+), 0 deletions(-)\n create mode 100755 t/t1005-cat-file.sh\n\ndiff --git a/t/t1005-cat-file.sh b/t/t1005-cat-file.sh\nnew file mode 100755\nindex 0000000..697354d\n--- /dev/null\n+++ b/t/t1005-cat-file.sh\n@@ -0,0 +1,91 @@\n+#!/bin/sh\n+\n+test_description='git cat-file'\n+\n+. ./test-lib.sh\n+\n+function maybe_remove_timestamp()\n+{\n+    if test -z \"$2\"; then\n+        echo \"$1\"\n+    else\n+        echo \"$1\" | sed -e 's/ [0-9]\\{10\\} [+-][0-9]\\{4\\}$//'\n+    fi\n+}\n+\n+function run_tests()\n+{\n+    type=$1\n+    sha1=$2\n+    size=$3\n+    content=$4\n+    pretty_content=$5\n+    no_timestamp=$6\n+\n+    test_expect_success \\\n+        \"$type exists\" \\\n+        \"git cat-file -e $sha1\"\n+    test_expect_success \\\n+        \"Type of $type is correct\" \\\n+        \"test $type = \\\"\\$(git cat-file -t $sha1)\\\"\"\n+    test_expect_success \\\n+        \"Size of $type is correct\" \\\n+        \"test $size = \\\"\\$(git cat-file -s $sha1)\\\"\"\n+    test -z \"$content\" || test_expect_success \\\n+        \"Content of $type is correct\" \\\n+        \"test \\\"\\$(maybe_remove_timestamp '$content' $no_timestamp)\\\" = \\\"\\$(maybe_remove_timestamp \\\"\\$(git cat-file $type $sha1)\\\" $no_timestamp)\\\"\"\n+    test_expect_success \\\n+        \"Pretty content of $type is correct\" \\\n+        \"test \\\"\\$(maybe_remove_timestamp '$pretty_content' $no_timestamp)\\\" = \\\"\\$(maybe_remove_timestamp \\\"\\$(git cat-file -p $sha1)\\\" $no_timestamp)\\\"\"\n+}\n+\n+hello_content=\"Hello World\"\n+hello_size=$(echo \"$hello_content\" | wc -c)\n+hello_sha1=557db03de997c86a4a028e1ebd3a1ceb225be238\n+\n+echo \"$hello_content\" > hello\n+\n+git update-index --add hello\n+\n+run_tests 'blob' $hello_sha1 $hello_size \"$hello_content\" \"$hello_content\"\n+\n+tree_sha1=$(git write-tree)\n+tree_size=33\n+tree_pretty_content=\"100644 blob $hello_sha1\thello\"\n+\n+run_tests 'tree' $tree_sha1 $tree_size \"\" \"$tree_pretty_content\"\n+\n+commit_message=\"Intial commit\"\n+commit_sha1=$(echo \"$commit_message\" | git commit-tree $tree_sha1)\n+commit_size=177\n+commit_content=\"tree $tree_sha1\n+author $GIT_AUTHOR_NAME <$GIT_AUTHOR_EMAIL> 0000000000 +0000\n+committer $GIT_COMMITTER_NAME <$GIT_COMMITTER_EMAIL> 0000000000 +0000\n+\n+$commit_message\"\n+\n+run_tests 'commit' $commit_sha1 $commit_size \"$commit_content\" \"$commit_content\" 1\n+\n+tag_header=\"object $hello_sha1\n+type blob\n+tag hellotag\n+tagger $GIT_COMMITTER_NAME <$GIT_COMMITTER_EMAIL>\"\n+tag_description=\"This is a tag\"\n+tag_content=\"$tag_header\n+\n+$tag_description\"\n+tag_pretty_content=\"$tag_header\n+Thu Jan 1 00:00:00 1970 +0000\n+\n+$tag_description\"\n+\n+tag_sha1=$(echo \"$tag_content\" | git mktag)\n+tag_size=$(echo \"$tag_content\" | wc -c)\n+\n+run_tests 'tag' $tag_sha1 $tag_size \"$tag_content\" \"$tag_pretty_content\"\n+\n+test_expect_success \\\n+    \"Reach a blob from a tag pointing to it\" \\\n+    \"test '$hello_content' = \\\"\\$(git cat-file blob $tag_sha1)\\\"\"\n+\n+test_done\n-- \n1.5.3.4.1337.g8e67d-dirty\n"},{"id":"57161","messageId":"1193307927-3592-3-git-send-email-aroben@apple.com","threadId":"10459","inReplyTo":"1193307927-3592-2-git-send-email-aroben@apple.com","subject":"[PATCH 2/9] git-cat-file: Small refactor of cmd_cat_file","fromName":"Adam Roben","fromEmail":"aroben@apple.com","sentAt":"2007-10-25T10:25:20Z","receivedAt":"2007-10-25T10:25:20Z","isPatch":true,"sender":{"key":"aroben@apple.com","avatar":"https://gravatar.com/avatar/9d3697e1de53890adf241331f4b970bdd2b18962b2ff0b8028ebb00e085807f8?d=mp&s=160"},"body":"I separated the logic of parsing the arguments from the logic of fetching and\noutputting the data. cat_one_file now does the latter.\n\nSigned-off-by: Adam Roben <aroben@apple.com>\n---\n builtin-cat-file.c |   38 ++++++++++++++++++++++----------------\n 1 files changed, 22 insertions(+), 16 deletions(-)\n\ndiff --git a/builtin-cat-file.c b/builtin-cat-file.c\nindex f132d58..34a63d1 100644\n--- a/builtin-cat-file.c\n+++ b/builtin-cat-file.c\n@@ -76,31 +76,16 @@ static void pprint_tag(const unsigned char *sha1, const char *buf, unsigned long\n \t\twrite_or_die(1, cp, endp - cp);\n }\n \n-int cmd_cat_file(int argc, const char **argv, const char *prefix)\n+static int cat_one_file(int opt, const char *exp_type, const char *obj_name)\n {\n \tunsigned char sha1[20];\n \tenum object_type type;\n \tvoid *buf;\n \tunsigned long size;\n-\tint opt;\n-\tconst char *exp_type, *obj_name;\n-\n-\tgit_config(git_default_config);\n-\tif (argc != 3)\n-\t\tusage(\"git-cat-file [-t|-s|-e|-p|<type>] <sha1>\");\n-\texp_type = argv[1];\n-\tobj_name = argv[2];\n \n \tif (get_sha1(obj_name, sha1))\n \t\tdie(\"Not a valid object name %s\", obj_name);\n \n-\topt = 0;\n-\tif ( exp_type[0] == '-' ) {\n-\t\topt = exp_type[1];\n-\t\tif ( !opt || exp_type[2] )\n-\t\t\topt = -1; /* Not a single character option */\n-\t}\n-\n \tbuf = NULL;\n \tswitch (opt) {\n \tcase 't':\n@@ -157,3 +142,24 @@ int cmd_cat_file(int argc, const char **argv, const char *prefix)\n \twrite_or_die(1, buf, size);\n \treturn 0;\n }\n+\n+int cmd_cat_file(int argc, const char **argv, const char *prefix)\n+{\n+\tint opt;\n+\tconst char *exp_type, *obj_name;\n+\n+\tgit_config(git_default_config);\n+\tif (argc != 3)\n+\t\tusage(\"git-cat-file [-t|-s|-e|-p|<type>] <sha1>\");\n+\texp_type = argv[1];\n+\tobj_name = argv[2];\n+\n+\topt = 0;\n+\tif ( exp_type[0] == '-' ) {\n+\t\topt = exp_type[1];\n+\t\tif ( !opt || exp_type[2] )\n+\t\t\topt = -1; /* Not a single character option */\n+\t}\n+\n+\treturn cat_one_file(opt, exp_type, obj_name);\n+}\n-- \n1.5.3.4.1337.g8e67d-dirty\n"},{"id":"57163","messageId":"1193307927-3592-4-git-send-email-aroben@apple.com","threadId":"10459","inReplyTo":"1193307927-3592-3-git-send-email-aroben@apple.com","subject":"[PATCH 3/9] git-cat-file: Make option parsing a little more flexible","fromName":"Adam Roben","fromEmail":"aroben@apple.com","sentAt":"2007-10-25T10:25:21Z","receivedAt":"2007-10-25T10:25:21Z","isPatch":true,"sender":{"key":"aroben@apple.com","avatar":"https://gravatar.com/avatar/9d3697e1de53890adf241331f4b970bdd2b18962b2ff0b8028ebb00e085807f8?d=mp&s=160"},"body":"This will make it easier to add newer options later.\n\nSigned-off-by: Adam Roben <aroben@apple.com>\n---\n builtin-cat-file.c |   42 ++++++++++++++++++++++++++++++------------\n 1 files changed, 30 insertions(+), 12 deletions(-)\n\ndiff --git a/builtin-cat-file.c b/builtin-cat-file.c\nindex 34a63d1..3a0be4a 100644\n--- a/builtin-cat-file.c\n+++ b/builtin-cat-file.c\n@@ -143,23 +143,41 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name)\n \treturn 0;\n }\n \n+static const char cat_file_usage[] = \"git-cat-file [-t|-s|-e|-p|<type>] <sha1>\";\n+\n int cmd_cat_file(int argc, const char **argv, const char *prefix)\n {\n-\tint opt;\n-\tconst char *exp_type, *obj_name;\n+\tint i, opt = 0;\n+\tconst char *exp_type = 0, *obj_name = 0;\n \n \tgit_config(git_default_config);\n-\tif (argc != 3)\n-\t\tusage(\"git-cat-file [-t|-s|-e|-p|<type>] <sha1>\");\n-\texp_type = argv[1];\n-\tobj_name = argv[2];\n-\n-\topt = 0;\n-\tif ( exp_type[0] == '-' ) {\n-\t\topt = exp_type[1];\n-\t\tif ( !opt || exp_type[2] )\n-\t\t\topt = -1; /* Not a single character option */\n+\n+\tfor (i = 1; i < argc; ++i) {\n+\t\tconst char *arg = argv[i];\n+\n+\t\tif (!strcmp(arg, \"-t\") || !strcmp(arg, \"-s\") || !strcmp(arg, \"-e\") || !strcmp(arg, \"-p\")) {\n+\t\t\texp_type = arg;\n+\t\t\topt = exp_type[1];\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\tif (arg[0] == '-')\n+\t\t\tusage(cat_file_usage);\n+\n+\t\tif (!exp_type) {\n+\t\t\texp_type = arg;\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\tif (obj_name)\n+\t\t\tusage(cat_file_usage);\n+\n+\t\tobj_name = arg;\n+\t\tbreak;\n \t}\n \n+\tif (!exp_type || !obj_name)\n+\t\tusage(cat_file_usage);\n+\n \treturn cat_one_file(opt, exp_type, obj_name);\n }\n-- \n1.5.3.4.1337.g8e67d-dirty\n"},{"id":"57164","messageId":"1193307927-3592-5-git-send-email-aroben@apple.com","threadId":"10459","inReplyTo":"1193307927-3592-4-git-send-email-aroben@apple.com","subject":"[PATCH 4/9] git-cat-file: Add --stdin option","fromName":"Adam Roben","fromEmail":"aroben@apple.com","sentAt":"2007-10-25T10:25:22Z","receivedAt":"2007-10-25T10:25:22Z","isPatch":true,"sender":{"key":"aroben@apple.com","avatar":"https://gravatar.com/avatar/9d3697e1de53890adf241331f4b970bdd2b18962b2ff0b8028ebb00e085807f8?d=mp&s=160"},"body":"This lets you specify object names on stdin instead of on the command line.\nWhen printing object contents or pretty-printing, objects will be printed\npreceded by their size:\n\n<size>LF\n<content>LF\n\nSigned-off-by: Adam Roben <aroben@apple.com>\n---\nBrian Downing wrote:\n> I think a far more reasonable output format for multiple objects would\n> be something like:\n> \n> <count> LF\n> <raw data> LF\n> \n> Where <count> is the number of bytes in the <raw data> as an ASCII\n> decimal integer.\n\nAgreed.\n\n Documentation/git-cat-file.txt |    6 ++++-\n builtin-cat-file.c             |   43 ++++++++++++++++++++++++++++++++++-----\n t/t1005-cat-file.sh            |   35 ++++++++++++++++++++++++++++++++\n 3 files changed, 77 insertions(+), 7 deletions(-)\n\ndiff --git a/Documentation/git-cat-file.txt b/Documentation/git-cat-file.txt\nindex afa095c..588d71a 100644\n--- a/Documentation/git-cat-file.txt\n+++ b/Documentation/git-cat-file.txt\n@@ -8,7 +8,7 @@ git-cat-file - Provide content or type/size information for repository objects\n \n SYNOPSIS\n --------\n-'git-cat-file' [-t | -s | -e | -p | <type>] <object>\n+'git-cat-file' [-t | -s | -e | -p | <type>] [--stdin | <object>]\n \n DESCRIPTION\n -----------\n@@ -23,6 +23,10 @@ OPTIONS\n \tFor a more complete list of ways to spell object names, see\n \t\"SPECIFYING REVISIONS\" section in gitlink:git-rev-parse[1].\n \n+--stdin::\n+\tRead object names from stdin instead of specifying one on the\n+\tcommand line.\n+\n -t::\n \tInstead of the content, show the object type identified by\n \t<object>.\ndiff --git a/builtin-cat-file.c b/builtin-cat-file.c\nindex 3a0be4a..ee46ba4 100644\n--- a/builtin-cat-file.c\n+++ b/builtin-cat-file.c\n@@ -76,7 +76,7 @@ static void pprint_tag(const unsigned char *sha1, const char *buf, unsigned long\n \t\twrite_or_die(1, cp, endp - cp);\n }\n \n-static int cat_one_file(int opt, const char *exp_type, const char *obj_name)\n+static int cat_one_file(int opt, const char *exp_type, const char *obj_name, int print_size)\n {\n \tunsigned char sha1[20];\n \tenum object_type type;\n@@ -139,16 +139,26 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name)\n \tif (!buf)\n \t\tdie(\"git-cat-file %s: bad file\", obj_name);\n \n+\tif (print_size) {\n+\t\tprintf(\"%lu\\n\", size);\n+\t\tfflush(stdout);\n+\t}\n \twrite_or_die(1, buf, size);\n+\tif (print_size) {\n+\t\tprintf(\"\\n\");\n+\t\tfflush(stdout);\n+\t}\n \treturn 0;\n }\n \n-static const char cat_file_usage[] = \"git-cat-file [-t|-s|-e|-p|<type>] <sha1>\";\n+static const char cat_file_usage[] = \"git-cat-file [-t|-s|-e|-p|<type>] [--stdin | <sha1>]\";\n \n int cmd_cat_file(int argc, const char **argv, const char *prefix)\n {\n-\tint i, opt = 0;\n+\tint i, opt = 0, print_size = 0;\n+\tint read_stdin = 0;\n \tconst char *exp_type = 0, *obj_name = 0;\n+\tstruct strbuf buf;\n \n \tgit_config(git_default_config);\n \n@@ -161,6 +171,11 @@ int cmd_cat_file(int argc, const char **argv, const char *prefix)\n \t\t\tcontinue;\n \t\t}\n \n+\t\tif (!strcmp(arg, \"--stdin\")) {\n+\t\t    read_stdin = 1;\n+\t\t    continue;\n+\t\t}\n+\n \t\tif (arg[0] == '-')\n \t\t\tusage(cat_file_usage);\n \n@@ -169,15 +184,31 @@ int cmd_cat_file(int argc, const char **argv, const char *prefix)\n \t\t\tcontinue;\n \t\t}\n \n-\t\tif (obj_name)\n+\t\tif (obj_name || read_stdin)\n \t\t\tusage(cat_file_usage);\n \n \t\tobj_name = arg;\n \t\tbreak;\n \t}\n \n-\tif (!exp_type || !obj_name)\n+\tif (!exp_type)\n \t\tusage(cat_file_usage);\n \n-\treturn cat_one_file(opt, exp_type, obj_name);\n+\tif (!read_stdin) {\n+\t\tif (!obj_name)\n+\t\t\tusage(cat_file_usage);\n+\t\treturn cat_one_file(opt, exp_type, obj_name, 0);\n+\t}\n+\n+\tprint_size = !opt || opt == 'p';\n+\n+\tstrbuf_init(&buf, 0);\n+\twhile (strbuf_getline(&buf, stdin, '\\n') != EOF) {\n+\t\tint error = cat_one_file(opt, exp_type, buf.buf, print_size);\n+\t\tif (error)\n+\t\t\treturn error;\n+\t}\n+\tstrbuf_release(&buf);\n+\n+\treturn 0;\n }\ndiff --git a/t/t1005-cat-file.sh b/t/t1005-cat-file.sh\nindex 697354d..2b2d386 100755\n--- a/t/t1005-cat-file.sh\n+++ b/t/t1005-cat-file.sh\n@@ -88,4 +88,39 @@ test_expect_success \\\n     \"Reach a blob from a tag pointing to it\" \\\n     \"test '$hello_content' = \\\"\\$(git cat-file blob $tag_sha1)\\\"\"\n \n+sha1s=\"$hello_sha1\n+$tree_sha1\n+$commit_sha1\n+$tag_sha1\"\n+\n+sizes=\"$hello_size\n+$tree_size\n+$commit_size\n+$tag_size\"\n+\n+test_expect_success \\\n+    \"Pass object hashes on stdin to retrieve sizes\" \\\n+    \"test '$sizes' = \\\"\\$(echo '$sha1s' | git cat-file -s --stdin)\\\"\"\n+\n+example_content=\"Silly example\"\n+example_size=$(echo \"$example_content\" | wc -c)\n+example_sha1=f24c74a2e500f5ee1332c86b94199f52b1d1d962\n+\n+echo \"$example_content\" > example\n+\n+git update-index --add example\n+\n+sha1s=\"$hello_sha1\n+$example_sha1\"\n+\n+contents=\"$hello_size\n+$hello_content\n+\n+$example_size\n+$example_content\"\n+\n+test_expect_success \\\n+    \"Pass object hashes on stdin to retrieve contents\" \\\n+    \"test '$contents' = \\\"\\$(echo '$sha1s' | git cat-file blob --stdin)\\\"\"\n+\n test_done\n-- \n1.5.3.4.1337.g8e67d-dirty\n"},{"id":"57162","messageId":"1193307927-3592-6-git-send-email-aroben@apple.com","threadId":"10459","inReplyTo":"1193307927-3592-5-git-send-email-aroben@apple.com","subject":"[PATCH 5/9] Add tests for git hash-object","fromName":"Adam Roben","fromEmail":"aroben@apple.com","sentAt":"2007-10-25T10:25:23Z","receivedAt":"2007-10-25T10:25:23Z","isPatch":true,"sender":{"key":"aroben@apple.com","avatar":"https://gravatar.com/avatar/9d3697e1de53890adf241331f4b970bdd2b18962b2ff0b8028ebb00e085807f8?d=mp&s=160"},"body":"\nSigned-off-by: Adam Roben <aroben@apple.com>\n---\nJohannes Sixt wrote:\n> Adam Roben schrieb:\n> > +test_expect_success \\\n> > +    'hash a file' \\\n> > +    \"test $hello_sha1 = $(git hash-object hello)\"\n> \n> Put tests in double-quotes; otherwise, the substitutions happen before the test begins, and not as part of the test. \n\nI think escaping the $(...) is enough to delay command execution.\n\n t/t1006-hash-object.sh |   27 +++++++++++++++++++++++++++\n 1 files changed, 27 insertions(+), 0 deletions(-)\n create mode 100755 t/t1006-hash-object.sh\n\ndiff --git a/t/t1006-hash-object.sh b/t/t1006-hash-object.sh\nnew file mode 100755\nindex 0000000..12f95f0\n--- /dev/null\n+++ b/t/t1006-hash-object.sh\n@@ -0,0 +1,27 @@\n+#!/bin/sh\n+\n+test_description='git hash-object'\n+\n+. ./test-lib.sh\n+\n+hello_content=\"Hello World\"\n+hello_sha1=557db03de997c86a4a028e1ebd3a1ceb225be238\n+echo \"$hello_content\" > hello\n+\n+test_expect_success \\\n+    'hash a file' \\\n+    \"test $hello_sha1 = \\$(git hash-object hello)\"\n+\n+test_expect_success \\\n+    'hash from stdin' \\\n+    \"test $hello_sha1 = \\$(echo '$hello_content' | git hash-object --stdin)\"\n+\n+test_expect_success \\\n+    'hash a file and write to database' \\\n+    \"test $hello_sha1 = \\$(git hash-object -w hello)\"\n+\n+test_expect_success \\\n+    'hash from stdin and write to database' \\\n+    \"test $hello_sha1 = \\$(echo '$hello_content' | git hash-object -w --stdin)\"\n+\n+test_done\n-- \n1.5.3.4.1337.g8e67d-dirty\n"},{"id":"57167","messageId":"1193307927-3592-7-git-send-email-aroben@apple.com","threadId":"10459","inReplyTo":"1193307927-3592-6-git-send-email-aroben@apple.com","subject":"[PATCH 6/9] git-hash-object: Add --stdin-paths option","fromName":"Adam Roben","fromEmail":"aroben@apple.com","sentAt":"2007-10-25T10:25:24Z","receivedAt":"2007-10-25T10:25:24Z","isPatch":true,"sender":{"key":"aroben@apple.com","avatar":"https://gravatar.com/avatar/9d3697e1de53890adf241331f4b970bdd2b18962b2ff0b8028ebb00e085807f8?d=mp&s=160"},"body":"This allows multiple paths to be specified on stdin.\n\nSigned-off-by: Adam Roben <aroben@apple.com>\n---\n Documentation/git-hash-object.txt |    5 ++++-\n hash-object.c                     |   29 ++++++++++++++++++++++++++++-\n t/t1006-hash-object.sh            |   22 ++++++++++++++++++++++\n 3 files changed, 54 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/git-hash-object.txt b/Documentation/git-hash-object.txt\nindex 616f196..50fc401 100644\n--- a/Documentation/git-hash-object.txt\n+++ b/Documentation/git-hash-object.txt\n@@ -8,7 +8,7 @@ git-hash-object - Compute object ID and optionally creates a blob from a file\n \n SYNOPSIS\n --------\n-'git-hash-object' [-t <type>] [-w] [--stdin] [--] <file>...\n+'git-hash-object' [-t <type>] [-w] [--stdin | --stdin-paths] [--] <file>...\n \n DESCRIPTION\n -----------\n@@ -32,6 +32,9 @@ OPTIONS\n --stdin::\n \tRead the object from standard input instead of from a file.\n \n+--stdin-paths::\n+\tRead file names from stdin instead of from the command-line.\n+\n Author\n ------\n Written by Junio C Hamano <junkio@cox.net>\ndiff --git a/hash-object.c b/hash-object.c\nindex 18f5017..fd96d50 100644\n--- a/hash-object.c\n+++ b/hash-object.c\n@@ -20,6 +20,7 @@ static void hash_object(const char *path, enum object_type type, int write_objec\n \t\t    ? \"Unable to add %s to database\"\n \t\t    : \"Unable to hash %s\", path);\n \tprintf(\"%s\\n\", sha1_to_hex(sha1));\n+\tmaybe_flush_or_die(stdout, \"hash to stdout\");\n }\n \n static void hash_stdin(const char *type, int write_object)\n@@ -31,7 +32,7 @@ static void hash_stdin(const char *type, int write_object)\n }\n \n static const char hash_object_usage[] =\n-\"git-hash-object [-t <type>] [-w] [--stdin] <file>...\";\n+\"git-hash-object [-t <type>] [-w] [--stdin | --stdin-paths] <file>...\";\n \n int main(int argc, char **argv)\n {\n@@ -41,6 +42,7 @@ int main(int argc, char **argv)\n \tconst char *prefix = NULL;\n \tint prefix_length = -1;\n \tint no_more_flags = 0;\n+\tint found_stdin_flag = 0;\n \n \tfor (i = 1 ; i < argc; i++) {\n \t\tif (!no_more_flags && argv[i][0] == '-') {\n@@ -62,7 +64,32 @@ int main(int argc, char **argv)\n \t\t\t}\n \t\t\telse if (!strcmp(argv[i], \"--help\"))\n \t\t\t\tusage(hash_object_usage);\n+\t\t\telse if (!strcmp(argv[i], \"--stdin-paths\")) {\n+\t\t\t\tstruct strbuf buf, nbuf;\n+\n+\t\t\t\tif (found_stdin_flag)\n+\t\t\t\t\tdie(\"Can't use both --stdin and --stdin-paths\");\n+\t\t\t\tfound_stdin_flag = 1;\n+\n+\t\t\t\tstrbuf_init(&buf, 0);\n+\t\t\t\tstrbuf_init(&nbuf, 0);\n+\t\t\t\twhile (strbuf_getline(&buf, stdin, '\\n') != EOF) {\n+\t\t\t\t\tif (buf.buf[0] == '\"') {\n+\t\t\t\t\t\tstrbuf_reset(&nbuf);\n+\t\t\t\t\t\tif (unquote_c_style(&nbuf, buf.buf, NULL))\n+\t\t\t\t\t\t\tdie(\"line is badly quoted\");\n+\t\t\t\t\t\tstrbuf_swap(&buf, &nbuf);\n+\t\t\t\t\t}\n+\t\t\t\t\thash_object(buf.buf, type_from_string(type), write_object);\n+\t\t\t\t}\n+\t\t\t\tstrbuf_release(&buf);\n+\t\t\t\tstrbuf_release(&nbuf);\n+\t\t\t}\n \t\t\telse if (!strcmp(argv[i], \"--stdin\")) {\n+\t\t\t\tif (found_stdin_flag)\n+\t\t\t\t\tdie(\"Can't use both --stdin and --stdin-paths\");\n+\t\t\t\tfound_stdin_flag = 1;\n+\n \t\t\t\thash_stdin(type, write_object);\n \t\t\t}\n \t\t\telse\ndiff --git a/t/t1006-hash-object.sh b/t/t1006-hash-object.sh\nindex 12f95f0..e747004 100755\n--- a/t/t1006-hash-object.sh\n+++ b/t/t1006-hash-object.sh\n@@ -24,4 +24,26 @@ test_expect_success \\\n     'hash from stdin and write to database' \\\n     \"test $hello_sha1 = \\$(echo '$hello_content' | git hash-object -w --stdin)\"\n \n+example_content=\"Silly example\"\n+example_sha1=f24c74a2e500f5ee1332c86b94199f52b1d1d962\n+echo \"$example_content\" > example\n+\n+filenames=\"hello\n+example\"\n+\n+sha1s=\"$hello_sha1\n+$example_sha1\"\n+\n+test_expect_success \\\n+    'hash two files with names on stdin' \\\n+    \"test '$sha1s' = \\\"\\$(echo '$filenames' | git hash-object --stdin-paths)\\\"\"\n+\n+test_expect_success \\\n+    'hash two files with names on stdin and write to database' \\\n+    \"test '$sha1s' = \\\"\\$(echo '$filenames' | git hash-object --stdin-paths)\\\"\"\n+\n+test_expect_failure \\\n+    \"Can't use --stdin and --stdin-paths together\" \\\n+    \"echo '$filenames' | git hash-object --stdin --stdin-paths\"\n+\n test_done\n-- \n1.5.3.4.1337.g8e67d-dirty\n"},{"id":"57165","messageId":"1193307927-3592-8-git-send-email-aroben@apple.com","threadId":"10459","inReplyTo":"1193307927-3592-7-git-send-email-aroben@apple.com","subject":"[PATCH 7/9] Git.pm: Add command_bidi_pipe and command_close_bidi_pipe","fromName":"Adam Roben","fromEmail":"aroben@apple.com","sentAt":"2007-10-25T10:25:25Z","receivedAt":"2007-10-25T10:25:25Z","isPatch":true,"sender":{"key":"aroben@apple.com","avatar":"https://gravatar.com/avatar/9d3697e1de53890adf241331f4b970bdd2b18962b2ff0b8028ebb00e085807f8?d=mp&s=160"},"body":"command_bidi_pipe hands back the stdin and stdout file handles from the\nexecuted command. command_close_bidi_pipe closes these handles and terminates\nthe process.\n\nSigned-off-by: Adam Roben <aroben@apple.com>\n---\n perl/Git.pm |   56 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n 1 files changed, 56 insertions(+), 0 deletions(-)\n\ndiff --git a/perl/Git.pm b/perl/Git.pm\nindex 3f4080c..46c5d10 100644\n--- a/perl/Git.pm\n+++ b/perl/Git.pm\n@@ -51,6 +51,7 @@ require Exporter;\n # Methods which can be called as standalone functions as well:\n @EXPORT_OK = qw(command command_oneline command_noisy\n                 command_output_pipe command_input_pipe command_close_pipe\n+                command_bidi_pipe command_close_bidi_pipe\n                 version exec_path hash_object git_cmd_try);\n \n \n@@ -92,6 +93,7 @@ increate nonwithstanding).\n use Carp qw(carp croak); # but croak is bad - throw instead\n use Error qw(:try);\n use Cwd qw(abs_path);\n+use IPC::Open2 qw(open2);\n \n }\n \n@@ -375,6 +377,60 @@ sub command_close_pipe {\n \t_cmd_close($fh, $ctx);\n }\n \n+=item command_bidi_pipe ( COMMAND [, ARGUMENTS... ] )\n+\n+Execute the given C<COMMAND> in the same way as command_output_pipe()\n+does but return both an input pipe filehandle and an output pipe filehandle.\n+\n+The function will return return C<($pid, $pipe_in, $pipe_out, $ctx)>.\n+See C<command_close_bidi_pipe()> for details.\n+\n+=cut\n+\n+sub command_bidi_pipe {\n+\tmy ($pid, $in, $out);\n+\t$pid = open2($in, $out, 'git', @_);\n+\treturn ($pid, $in, $out, join(' ', @_));\n+}\n+\n+=item command_close_bidi_pipe ( PID, PIPE_IN, PIPE_OUT [, CTX] )\n+\n+Close the C<PIPE_IN> and C<PIPE_OUT> as returned from C<command_bidi_pipe()>,\n+checking whether the command finished successfully. The optional C<CTX>\n+argument is required if you want to see the command name in the error message,\n+and it is the fourth value returned by C<command_bidi_pipe()>.  The call idiom\n+is:\n+\n+\tmy ($pid, $in, $out, $ctx) = $r->command_bidi_pipe('cat-file --stdin');\n+\tprint \"000000000\\n\" $out;\n+\twhile (<$in>) { ... }\n+\t$r->command_close_bidi_pipe($pid, $in, $out, $ctx);\n+\n+Note that you should not rely on whatever actually is in C<CTX>;\n+currently it is simply the command name but in future the context might\n+have more complicated structure.\n+\n+=cut\n+\n+sub command_close_bidi_pipe {\n+\tmy ($pid, $in, $out, $ctx) = @_;\n+\tforeach my $fh ($in, $out) {\n+\t\tif (not close $fh) {\n+\t\t\tif ($!) {\n+\t\t\t\tcarp \"error closing pipe: $!\";\n+\t\t\t} elsif ($? >> 8) {\n+\t\t\t\tthrow Git::Error::Command($ctx, $? >>8);\n+\t\t\t}\n+\t\t}\n+\t}\n+\n+\twaitpid $pid, 0;\n+\n+\tif ($? >> 8) {\n+\t\tthrow Git::Error::Command($ctx, $? >>8);\n+\t}\n+}\n+\n \n =item command_noisy ( COMMAND [, ARGUMENTS... ] )\n \n-- \n1.5.3.4.1337.g8e67d-dirty\n"},{"id":"57168","messageId":"1193307927-3592-9-git-send-email-aroben@apple.com","threadId":"10459","inReplyTo":"1193307927-3592-8-git-send-email-aroben@apple.com","subject":"[PATCH 8/9] Git.pm: Add hash_and_insert_object and cat_blob","fromName":"Adam Roben","fromEmail":"aroben@apple.com","sentAt":"2007-10-25T10:25:26Z","receivedAt":"2007-10-25T10:25:26Z","isPatch":true,"sender":{"key":"aroben@apple.com","avatar":"https://gravatar.com/avatar/9d3697e1de53890adf241331f4b970bdd2b18962b2ff0b8028ebb00e085807f8?d=mp&s=160"},"body":"These functions are more efficient ways of executing `git hash-object -w` and\n`git cat-file blob` when you are dealing with many files/objects.\n\nSigned-off-by: Adam Roben <aroben@apple.com>\n---\nEric Wong wrote:\n> > +package Git::Commands;\n> \n> Can this be a separate file, or a part of Git.pm?  I'm sure other\n> scripts can eventually use this and I've been meaning to split\n> git-svn.perl into separate files so it's easier to follow.\n\nI ended up making it part of Git.pm, because I realized that made far more\nsense than splitting it into a separate file.\n\n perl/Git.pm |   97 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++-\n 1 files changed, 95 insertions(+), 2 deletions(-)\n\ndiff --git a/perl/Git.pm b/perl/Git.pm\nindex 46c5d10..f23edef 100644\n--- a/perl/Git.pm\n+++ b/perl/Git.pm\n@@ -39,6 +39,9 @@ $VERSION = '0.01';\n   my $lastrev = $repo->command_oneline( [ 'rev-list', '--all' ],\n                                         STDERR => 0 );\n \n+  my $sha1 = $repo->hash_and_insert_object('file.txt');\n+  my $contents = $repo->cat_blob($sha1);\n+\n =cut\n \n \n@@ -218,7 +221,6 @@ sub repository {\n \tbless $self, $class;\n }\n \n-\n =back\n \n =head1 METHODS\n@@ -675,6 +677,93 @@ sub hash_object {\n }\n \n \n+=item hash_and_insert_object ( FILENAME )\n+\n+Compute the SHA1 object id of the given C<FILENAME> and add the object to the\n+object database.\n+\n+The function returns the SHA1 hash.\n+\n+=cut\n+\n+# TODO: Support for passing FILEHANDLE instead of FILENAME\n+sub hash_and_insert_object {\n+\tmy ($self, $filename) = @_;\n+\n+\t$self->_open_hash_and_insert_object_if_needed();\n+\tmy ($in, $out) = ($self->{hash_object_in}, $self->{hash_object_out});\n+\n+\tprint $out $filename, \"\\n\";\n+\tchomp(my $hash = <$in>);\n+\treturn $hash;\n+}\n+\n+sub _open_hash_and_insert_object_if_needed {\n+\tmy ($self) = @_;\n+\n+\treturn if defined($self->{hash_object_pid});\n+\n+\t($self->{hash_object_pid}, $self->{hash_object_in},\n+\t $self->{hash_object_out}, $self->{hash_object_ctx}) =\n+\t\tcommand_bidi_pipe(qw(hash-object -w --stdin-paths));\n+}\n+\n+sub _close_hash_and_insert_object {\n+\tmy ($self) = @_;\n+\n+\treturn unless defined($self->{hash_object_pid});\n+\n+\tmy @vars = map { 'hash_object' . $_ } qw(pid in out ctx);\n+\n+\tcommand_close_bidi_pipe($self->{@vars});\n+\tdelete $self->{@vars};\n+}\n+\n+=item cat_blob ( SHA1 )\n+\n+Returns the contents of the blob identified by C<SHA1>.\n+\n+=cut\n+\n+sub cat_blob {\n+\tmy ($self, $sha1) = @_;\n+\n+\t$self->_open_cat_blob_if_needed();\n+\tmy ($in, $out) = ($self->{cat_blob_in}, $self->{cat_blob_out});\n+\n+\tprint $out $sha1, \"\\n\";\n+\tchomp(my $size = <$in>);\n+\n+\tmy $blob;\n+\tmy $result = read($in, $blob, $size);\n+\tdefined $result or carp $!;\n+\n+\t# Skip past the trailing newline.\n+\tread($in, my $newline, 1);\n+\n+\treturn $blob;\n+}\n+\n+sub _open_cat_blob_if_needed {\n+\tmy ($self) = @_;\n+\n+\treturn if defined($self->{cat_blob_pid});\n+\n+\t($self->{cat_blob_pid}, $self->{cat_blob_in},\n+\t $self->{cat_blob_out}, $self->{cat_blob_ctx}) =\n+\t\tcommand_bidi_pipe(qw(cat-file blob --stdin));\n+}\n+\n+sub _close_cat_blob {\n+\tmy ($self) = @_;\n+\n+\treturn unless defined($self->{cat_blob_pid});\n+\n+\tmy @vars = map { 'cat_blob' . $_ } qw(pid in out ctx);\n+\n+\tcommand_close_bidi_pipe($self->{@vars});\n+\tdelete $self->{@vars};\n+}\n \n =back\n \n@@ -892,7 +981,11 @@ sub _cmd_close {\n }\n \n \n-sub DESTROY { }\n+sub DESTROY {\n+\tmy ($self) = @_;\n+\t$self->_close_hash_and_insert_object();\n+\t$self->_close_cat_blob();\n+}\n \n \n # Pipe implementation for ActiveState Perl.\n-- \n1.5.3.4.1342.g32de\n"},{"id":"57166","messageId":"1193307927-3592-10-git-send-email-aroben@apple.com","threadId":"10459","inReplyTo":"1193307927-3592-9-git-send-email-aroben@apple.com","subject":"[PATCH 9/9] git-svn: Make fetch ~1.7x faster","fromName":"Adam Roben","fromEmail":"aroben@apple.com","sentAt":"2007-10-25T10:25:27Z","receivedAt":"2007-10-25T10:25:27Z","isPatch":true,"sender":{"key":"aroben@apple.com","avatar":"https://gravatar.com/avatar/9d3697e1de53890adf241331f4b970bdd2b18962b2ff0b8028ebb00e085807f8?d=mp&s=160"},"body":"We were spending a lot of time forking/execing git-cat-file and\ngit-hash-object. We now maintain a global Git repository object in order to use\nGit.pm's more efficient hash_and_insert_object and cat_blob methods.\n\nSigned-off-by: Adam Roben <aroben@apple.com>\n---\nEric Wong wrote:\n> > +sub hash_object {\n> > +   my (undef, $fh) = @_;\n> > +\n> > +   my ($tmp_fh, $tmp_filename) = tempfile(UNLINK => 1);\n> > +   while (my $line = <$fh>) {\n> > +           print $tmp_fh $line;\n> > +   }\n> > +   close($tmp_fh);\n> \n> Related to the above.  It's better to sysread()/syswrite() or\n> read()/print() in a loop with a predefined buffer size rather than to\n> use a readline() since you could be dealing with files with very long\n> lines or binaries with no newline characters in them at all.\n\nFixed.\n\n> > +   _open_hash_object_if_needed();\n> > +   print $_hash_object_out $tmp_filename . \"\\n\";\n> \n> Minor, but\n> \n>         print $_hash_object_out $tmp_filename, \"\\n\";\n> \n> avoids creating a new string.\n\nFixed.\n\n git-svn.perl |   40 ++++++++++++++++++----------------------\n 1 files changed, 18 insertions(+), 22 deletions(-)\n\ndiff --git a/git-svn.perl b/git-svn.perl\nindex 22bb47b..fcb07f5 100755\n--- a/git-svn.perl\n+++ b/git-svn.perl\n@@ -4,7 +4,7 @@\n use warnings;\n use strict;\n use vars qw/\t$AUTHOR $VERSION\n-\t\t$sha1 $sha1_short $_revision\n+\t\t$sha1 $sha1_short $_revision $_repository\n \t\t$_q $_authors %users/;\n $AUTHOR = 'Eric Wong <normalperson@yhbt.net>';\n $VERSION = '@@GIT_VERSION@@';\n@@ -225,6 +225,7 @@ unless ($cmd =~ /(?:clone|init|multi-init)$/) {\n \t\t}\n \t\t$ENV{GIT_DIR} = $git_dir;\n \t}\n+\t$_repository = Git->repository(Repository => $ENV{GIT_DIR});\n }\n unless ($cmd =~ /^(?:clone|init|multi-init|commit-diff)$/) {\n \tGit::SVN::Migration::migration_check();\n@@ -332,6 +333,7 @@ sub cmd_init {\n \t                       \"as a command-line argument\\n\";\n \tinit_subdir(@_);\n \tdo_git_init_db();\n+\t$_repository = Git->repository(Repository => $ENV{GIT_DIR});\n \n \tGit::SVN->init($url);\n }\n@@ -2541,6 +2543,7 @@ use vars qw/@ISA/;\n use strict;\n use warnings;\n use Carp qw/croak/;\n+use File::Temp qw/tempfile/;\n use IO::File qw//;\n use Digest::MD5;\n \n@@ -2683,14 +2686,8 @@ sub apply_textdelta {\n \tmy $base = IO::File->new_tmpfile;\n \t$base->autoflush(1);\n \tif ($fb->{blob}) {\n-\t\tdefined (my $pid = fork) or croak $!;\n-\t\tif (!$pid) {\n-\t\t\topen STDOUT, '>&', $base or croak $!;\n-\t\t\tprint STDOUT 'link ' if ($fb->{mode_a} == 120000);\n-\t\t\texec qw/git-cat-file blob/, $fb->{blob} or croak $!;\n-\t\t}\n-\t\twaitpid $pid, 0;\n-\t\tcroak $? if $?;\n+\t\tmy $contents = $::_repository->cat_blob($fb->{blob});\n+\t\tprint $base $contents;\n \n \t\tif (defined $exp) {\n \t\t\tseek $base, 0, 0 or croak $!;\n@@ -2729,14 +2726,18 @@ sub close_file {\n \t\t\t$buf eq 'link ' or die \"$path has mode 120000\",\n \t\t\t                       \"but is not a link\\n\";\n \t\t}\n-\t\tdefined(my $pid = open my $out,'-|') or die \"Can't fork: $!\\n\";\n-\t\tif (!$pid) {\n-\t\t\topen STDIN, '<&', $fh or croak $!;\n-\t\t\texec qw/git-hash-object -w --stdin/ or croak $!;\n+\n+\t\tmy ($tmp_fh, $tmp_filename) = File::Temp::tempfile(UNLINK => 1);\n+\t\tmy $result;\n+\t\twhile ($result = sysread($fh, my $string, 1024)) {\n+\t\t\tsyswrite($tmp_fh, $string, $result);\n \t\t}\n-\t\tchomp($hash = do { local $/; <$out> });\n-\t\tclose $out or croak $!;\n+\t\tdefined $result or croak $!;\n+\t\tclose $tmp_fh or croak $!;\n+\n \t\tclose $fh or croak $!;\n+\n+\t\t$hash = $::_repository->hash_and_insert_object($tmp_filename);\n \t\t$hash =~ /^[a-f\\d]{40}$/ or die \"not a sha1: $hash\\n\";\n \t\tclose $fb->{base} or croak $!;\n \t} else {\n@@ -3063,13 +3064,8 @@ sub chg_file {\n \t} elsif ($m->{mode_a} =~ /^120/ && $m->{mode_b} !~ /^120/) {\n \t\t$self->change_file_prop($fbat,'svn:special',undef);\n \t}\n-\tdefined(my $pid = fork) or croak $!;\n-\tif (!$pid) {\n-\t\topen STDOUT, '>&', $fh or croak $!;\n-\t\texec qw/git-cat-file blob/, $m->{sha1_b} or croak $!;\n-\t}\n-\twaitpid $pid, 0;\n-\tcroak $? if $?;\n+\tmy $blob = $::_repository->cat_blob($m->{sha1_b});\n+\tprint $fh $blob;\n \t$fh->flush == 0 or croak $!;\n \tseek $fh, 0, 0 or croak $!;\n \n-- \n1.5.3.4.1337.g8e67d-dirty\n"},{"id":"57278","messageId":"20071026151107.GA29522@muzzle","threadId":"10459","inReplyTo":"1193307927-3592-9-git-send-email-aroben@apple.com","subject":"Re: [PATCH 8/9] Git.pm: Add hash_and_insert_object and cat_blob","fromName":"Eric Wong","fromEmail":"normalperson@yhbt.net","sentAt":"2007-10-26T15:11:07Z","receivedAt":"2007-10-26T15:11:07Z","isPatch":true,"sender":{"key":"e@80x24.org","avatar":null},"body":"Adam Roben <aroben@apple.com> wrote:\n> These functions are more efficient ways of executing `git hash-object -w` and\n> `git cat-file blob` when you are dealing with many files/objects.\n> \n> Signed-off-by: Adam Roben <aroben@apple.com>\n> ---\n> Eric Wong wrote:\n> > > +package Git::Commands;\n> > \n> > Can this be a separate file, or a part of Git.pm?  I'm sure other\n> > scripts can eventually use this and I've been meaning to split\n> > git-svn.perl into separate files so it's easier to follow.\n> \n> I ended up making it part of Git.pm, because I realized that made far more\n> sense than splitting it into a separate file.\n> \n>  perl/Git.pm |   97 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++-\n>  1 files changed, 95 insertions(+), 2 deletions(-)\n\nHi Adam,\n\nThanks.\n\n> diff --git a/perl/Git.pm b/perl/Git.pm\n> index 46c5d10..f23edef 100644\n> --- a/perl/Git.pm\n> +++ b/perl/Git.pm\n> @@ -39,6 +39,9 @@ $VERSION = '0.01';\n>    my $lastrev = $repo->command_oneline( [ 'rev-list', '--all' ],\n>                                          STDERR => 0 );\n>  \n> +  my $sha1 = $repo->hash_and_insert_object('file.txt');\n> +  my $contents = $repo->cat_blob($sha1);\n\nI missed this the first time around.  But I'd rather be able to pass a\nfile handle to cat_blob for writing, instead of returning a potentially\nhuge string in memory.\n\n> @@ -675,6 +677,93 @@ sub hash_object {\n>  }\n>  \n>  \n> +=item hash_and_insert_object ( FILENAME )\n> +\n> +Compute the SHA1 object id of the given C<FILENAME> and add the object to the\n> +object database.\n> +\n> +The function returns the SHA1 hash.\n> +\n> +=cut\n> +\n> +# TODO: Support for passing FILEHANDLE instead of FILENAME\n\nFilenames are fine for this input since they (are/should be) generated\nby File::Temp and not from an untrusted repo.\n\nWe should, however assert that the caller of this function\nisn't using a stupid filename with \"\\n\" in it.\n\n> +sub hash_and_insert_object {\n> +\tmy ($self, $filename) = @_;\n> +\n> +\t$self->_open_hash_and_insert_object_if_needed();\n> +\tmy ($in, $out) = ($self->{hash_object_in}, $self->{hash_object_out});\n> +\n> +\tprint $out $filename, \"\\n\";\n> +\tchomp(my $hash = <$in>);\n> +\treturn $hash;\n> +}\n> +\n> +sub _open_hash_and_insert_object_if_needed {\n> +\tmy ($self) = @_;\n> +\n> +\treturn if defined($self->{hash_object_pid});\n> +\n> +\t($self->{hash_object_pid}, $self->{hash_object_in},\n> +\t $self->{hash_object_out}, $self->{hash_object_ctx}) =\n> +\t\tcommand_bidi_pipe(qw(hash-object -w --stdin-paths));\n> +}\n> +\n> +sub _close_hash_and_insert_object {\n> +\tmy ($self) = @_;\n> +\n> +\treturn unless defined($self->{hash_object_pid});\n> +\n> +\tmy @vars = map { 'hash_object' . $_ } qw(pid in out ctx);\n\nIt looks like you're missing a '_' in there.\n\n> +\tcommand_close_bidi_pipe($self->{@vars});\n> +\tdelete $self->{@vars};\n> +}\n> +\n\n\n> +=item cat_blob ( SHA1 )\n> +\n> +Returns the contents of the blob identified by C<SHA1>.\n> +\n> +=cut\n> +\n> +sub cat_blob {\n> +\tmy ($self, $sha1) = @_;\n> +\n> +\t$self->_open_cat_blob_if_needed();\n> +\tmy ($in, $out) = ($self->{cat_blob_in}, $self->{cat_blob_out});\n> +\n> +\tprint $out $sha1, \"\\n\";\n> +\tchomp(my $size = <$in>);\n> +\n> +\tmy $blob;\n> +\tmy $result = read($in, $blob, $size);\n> +\tdefined $result or carp $!;\n> +\n> +\t# Skip past the trailing newline.\n> +\tread($in, my $newline, 1);\n> +\n> +\treturn $blob;\n> +}\n\nHowever, I'd very much like to be able to pass a file handle to this\nfunction.  This should read()/print() to a file handle passed to it in a\nloop rather than slurping all of $size at once, since the files we're\nreceiving can be huge.\n\nI'd also be happier if we checked that we actually read $size bytes in\nthe loop, and that $newline is actually \"\\n\" to safeguard against bugs\nin cat-blob.\n\n> +sub _open_cat_blob_if_needed {\n> +\tmy ($self) = @_;\n> +\n> +\treturn if defined($self->{cat_blob_pid});\n> +\n> +\t($self->{cat_blob_pid}, $self->{cat_blob_in},\n> +\t $self->{cat_blob_out}, $self->{cat_blob_ctx}) =\n> +\t\tcommand_bidi_pipe(qw(cat-file blob --stdin));\n> +}\n> +\n> +sub _close_cat_blob {\n> +\tmy ($self) = @_;\n> +\n> +\treturn unless defined($self->{cat_blob_pid});\n> +\n> +\tmy @vars = map { 'cat_blob' . $_ } qw(pid in out ctx);\n\nIt looks like you're missing a '_' here, too.\n\n> +\tcommand_close_bidi_pipe($self->{@vars});\n> +\tdelete $self->{@vars};\n> +}\n>  \n\nOne more nit, I'm a bit paranoid, but I personally like to die/croak if\nthe result of every print()/syswrite() to make sure the pipe we're\nwriting to didn't die or if there were other error indicators.\n\nHopefully that's the last of tweaks I'd like to see :)\n\n-- \nEric Wong\n"},{"id":"57304","messageId":"7vlk9py4r9.fsf@gitster.siamese.dyndns.org","threadId":"10459","inReplyTo":"1193307927-3592-4-git-send-email-aroben@apple.com","subject":"Re: [PATCH 3/9] git-cat-file: Make option parsing a little more flexible","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-10-26T20:56:10Z","receivedAt":"2007-10-26T20:56:10Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Adam Roben <aroben@apple.com> writes:\n\n> This will make it easier to add newer options later.\n\nA good change in principle.\n\n> diff --git a/builtin-cat-file.c b/builtin-cat-file.c\n> index 34a63d1..3a0be4a 100644\n> --- a/builtin-cat-file.c\n> +++ b/builtin-cat-file.c\n> @@ -143,23 +143,41 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name)\n>  \treturn 0;\n>  }\n>  \n> +static const char cat_file_usage[] = \"git-cat-file [-t|-s|-e|-p|<type>] <sha1>\";\n> +\n>  int cmd_cat_file(int argc, const char **argv, const char *prefix)\n>  {\n> -\tint opt;\n> -\tconst char *exp_type, *obj_name;\n> +\tint i, opt = 0;\n> +\tconst char *exp_type = 0, *obj_name = 0;\n\nNULL pointer constants in git sources are spelled \"NULL\", not\n\"0\".\n"},{"id":"57305","messageId":"7vd4v1y4lv.fsf@gitster.siamese.dyndns.org","threadId":"10459","inReplyTo":"1193307927-3592-5-git-send-email-aroben@apple.com","subject":"Re: [PATCH 4/9] git-cat-file: Add --stdin option","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-10-26T20:59:24Z","receivedAt":"2007-10-26T20:59:24Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Adam Roben <aroben@apple.com> writes:\n\n> @@ -23,6 +23,10 @@ OPTIONS\n>  \tFor a more complete list of ways to spell object names, see\n>  \t\"SPECIFYING REVISIONS\" section in gitlink:git-rev-parse[1].\n>  \n> +--stdin::\n> +\tRead object names from stdin instead of specifying one on the\n> +\tcommand line.\n> +\n\nThis does not talk about modified output format: what the format\nis, nor when that modified format is used.\n\n> @@ -139,16 +139,26 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name)\n>  \tif (!buf)\n>  \t\tdie(\"git-cat-file %s: bad file\", obj_name);\n>  \n> +\tif (print_size) {\n> +\t\tprintf(\"%lu\\n\", size);\n> +\t\tfflush(stdout);\n> +\t}\n>  \twrite_or_die(1, buf, size);\n> +\tif (print_size) {\n> +\t\tprintf(\"\\n\");\n> +\t\tfflush(stdout);\n> +\t}\n>  \treturn 0;\n>  }\n>  \n\nNot that I object strongly to it, but do we need extra LF after\nthe contents?\n\n  - \"It would help readers written in typical scripting\n    languages\" is an acceptable answer, but I doubt that is the\n    case --- the reader is given the number of bytes and is\n    going to \"read($pipe, $buf, $that_size)\" anyway.\n\n  - \"The reader can assert that one-byte past the content is a\n    LF to catch errors, and this LF would help re-synchronize\n    after such an error\" would be another acceptable answer, but\n    for the re-synchronization to work, the output needs to tell\n    which record each chunk is about (i.e. if the output were\n    \"<type> <sha1> <size>LF<contents>LF\", the \"re-sync\" argument\n    would make a bit more sense).\n\n> +\tprint_size = !opt || opt == 'p';\n\nNeeds a bit of comment here, and in the documentation.  E.g.\n\n\tgit-cat-file --stdin -t <list-of-sha1\n        git-cat-file --stdin -s <list-of-sha1\n\n\tare ways to check types and sizes of the objects in the\n\tlist.\n\nHow does --stdin interact with -e?\n\nHow does --stdin interact with -p when printing a tree or a tag\nobject?\n\nHow does \"blob --stdin\" do when input sequence contains a non\nblob SHA1?\n\nIt almost feels that --stdin should be named something else,\nsuch as --batch or --bulk, as it is not just affecting the\ninput.\n\nHere is an alternative suggestion.\n\n   Two new options, --batch and --batch-check, are introduced.\n   These options are incompatible with -[tsep] or an object type\n   given as the first parameter to git-cat-file.\n\n   * git-cat-file --batch-check <list-of-sha1\n\n     outputs a record of this form\n\n          <sha1> SP <type> SP <size> LF\n\n     for each of the input lines.\n\n   * git-cat-file --batch <list-of-sha1\n\n     outputs a record of this form\n\n          <sha1> SP <type> SP <size> LF <contents> LF\n\n     for each of the input lines.\n\n  For a missing object, either option gives a record of form:\n\n          <sha1> SP missing LF\n"},{"id":"57306","messageId":"7v640ty4kp.fsf@gitster.siamese.dyndns.org","threadId":"10459","inReplyTo":"1193307927-3592-6-git-send-email-aroben@apple.com","subject":"Re: [PATCH 5/9] Add tests for git hash-object","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-10-26T21:00:06Z","receivedAt":"2007-10-26T21:00:06Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Adam Roben <aroben@apple.com> writes:\n\n> +test_expect_success \\\n> +    'hash a file' \\\n> +    \"test $hello_sha1 = \\$(git hash-object hello)\"\n> +\n> +test_expect_success \\\n> +    'hash from stdin' \\\n> +    \"test $hello_sha1 = \\$(echo '$hello_content' | git hash-object --stdin)\"\n\nNeeds to make sure no object has been written to the object\ndatabase at this point?\n\n> +test_expect_success \\\n> +    'hash a file and write to database' \\\n> +    \"test $hello_sha1 = \\$(git hash-object -w hello)\"\n\n... and make sure the objectis written here?\n\n> +test_expect_success \\\n> +    'hash from stdin and write to database' \\\n> +    \"test $hello_sha1 = \\$(echo '$hello_content' | git hash-object -w --stdin)\"\n> +\n> +test_done\n\n... and/or here?\n"},{"id":"57307","messageId":"7vy7dpwpz4.fsf@gitster.siamese.dyndns.org","threadId":"10459","inReplyTo":"1193307927-3592-7-git-send-email-aroben@apple.com","subject":"Re: [PATCH 6/9] git-hash-object: Add --stdin-paths option","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-10-26T21:00:47Z","receivedAt":"2007-10-26T21:00:47Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Adam Roben <aroben@apple.com> writes:\n\n> This allows multiple paths to be specified on stdin.\n\nOk.  List of paths is certainly a good thing to have.\n\nIn addition, if you are enhancing cat-file to spew chunked\noutput out, I suspect that there should be a mode of operation\nfor hash-object that eats that data format.  IOW, this pipe\n\n\tgit-cat-file --batch <list-of-sha1 |\n        git-hash-object --batch\n\nshould be an intuitive no-op, shouldn't it?\n"},{"id":"57316","messageId":"20071026231902.GC2519@lavos.net","threadId":"10459","inReplyTo":"7vy7dpwpz4.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH 6/9] git-hash-object: Add --stdin-paths option","fromName":"Brian Downing","fromEmail":"bdowning@lavos.net","sentAt":"2007-10-26T23:19:02Z","receivedAt":"2007-10-26T23:19:02Z","isPatch":true,"sender":{"key":"bdowning@lavos.net","avatar":"https://avatars.githubusercontent.com/u/366426?v=4"},"body":"On Fri, Oct 26, 2007 at 02:00:47PM -0700, Junio C Hamano wrote:\n> In addition, if you are enhancing cat-file to spew chunked\n> output out, I suspect that there should be a mode of operation\n> for hash-object that eats that data format.  IOW, this pipe\n> \n> \tgit-cat-file --batch <list-of-sha1 |\n>         git-hash-object --batch\n> \n> should be an intuitive no-op, shouldn't it?\n\nI think that's an obviously good thing to do.  However, given your\nsuggested output format (which I also like):\n\n>    * git-cat-file --batch <list-of-sha1\n> \n>      outputs a record of this form\n> \n>           <sha1> SP <type> SP <size> LF <contents> LF\n> \n>      for each of the input lines.\n\nWhat should the input behavior be?  Obviously the sha1 will probably\nnot be known on the input side.  Should that simply be optional (i.e.\nit will accept either \"<sha1> SP <type> SP <size>\" or \"<type> SP <size>\"\nor should it only accept the latter, and a dummy sha1 will need to be\nfilled in if the sha1 is not known (presumably \"000...000\")?\n\n-bcd\n"},{"id":"57323","messageId":"7vlk9pv08i.fsf@gitster.siamese.dyndns.org","threadId":"10459","inReplyTo":"20071026231902.GC2519@lavos.net","subject":"Re: [PATCH 6/9] git-hash-object: Add --stdin-paths option","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-10-27T01:02:05Z","receivedAt":"2007-10-27T01:02:05Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"bdowning@lavos.net (Brian Downing) writes:\n\n> On Fri, Oct 26, 2007 at 02:00:47PM -0700, Junio C Hamano wrote:\n>> In addition, if you are enhancing cat-file to spew chunked\n>> output out, I suspect that there should be a mode of operation\n>> for hash-object that eats that data format.  IOW, this pipe\n>> \n>> \tgit-cat-file --batch <list-of-sha1 |\n>>         git-hash-object --batch\n>> \n>> should be an intuitive no-op, shouldn't it?\n>\n> I think that's an obviously good thing to do.  However, given your\n> suggested output format (which I also like):\n>\n>>    * git-cat-file --batch <list-of-sha1\n>> \n>>      outputs a record of this form\n>> \n>>           <sha1> SP <type> SP <size> LF <contents> LF\n>> \n>>      for each of the input lines.\n>\n> What should the input behavior be?  Obviously the sha1 will probably\n> not be known on the input side.  Should that simply be optional (i.e.\n> it will accept either \"<sha1> SP <type> SP <size>\" or \"<type> SP <size>\"\n> or should it only accept the latter, and a dummy sha1 will need to be\n> filled in if the sha1 is not known (presumably \"000...000\")?\n\nYeah, you caught me ;-)\n\nEither making it optional or requiring a dummy value would work.\nIf a non-dummy value is given, we could use it to validate it.\n\nBut that would not be a useful application anyway.  So perhaps\njust the sequence of \"<type> SP <size> LF <contents> LF\" would\nbe the most sensible thing to do.\n"}]}