{"thread":{"id":"29753","subject":"[PATCH 00/11] Large blob fixes","startedAt":"2012-02-27T07:55:04Z","lastAt":"2012-03-06T00:59:13Z","messageCount":48,"participants":["Nguyễn Thái Ngọc Duy","Junio C Hamano","Peter Baumann","Nguyen Thai Ngoc Duy"],"isPatch":true,"patchVersion":1,"patchTotal":11},"messages":[{"id":"185498","messageId":"1330329315-11407-1-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":null,"subject":"[PATCH 00/11] Large blob fixes","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-27T07:55:04Z","receivedAt":"2012-02-27T07:55:04Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"These patches make sure we avoid keeping whole blob in memory, at\nleast in common cases. Blob-only streaming code paths are opened to\naccomplish that.\n\nI don't quite like having three different implementations for\nchecking sha-1 signature (one on git_istream, one on packed_git and\nthe other one in index-pack) but I failed to see how to unify them.\n\nMaking archive-zip work with stream can be difficult. But at least tar\nformat works. Good enough for me.\n\nNguyễn Thái Ngọc Duy (11):\n  Add more large blob test cases\n  Factor out and export large blob writing code to arbitrary file\n    handle\n  cat-file: use streaming interface to print blobs\n  parse_object: special code path for blobs to avoid putting whole\n    object in memory\n  show: use streaming interface for showing blobs\n  index-pack --verify: skip sha-1 collision test\n  index-pack: split second pass obj handling into own function\n  index-pack: reduce memory usage when the pack has large blobs\n  pack-check: do not unpack blobs\n  archive: support streaming large files to a tar archive\n  fsck: use streaming interface for writing lost-found blobs\n\n archive-tar.c        |   35 +++++++++++++---\n archive-zip.c        |    9 ++--\n archive.c            |   51 ++++++++++++++++--------\n archive.h            |   11 ++++-\n builtin/cat-file.c   |   22 ++++++++++\n builtin/fsck.c       |    8 +---\n builtin/index-pack.c |  108 +++++++++++++++++++++++++++++++++++++------------\n builtin/log.c        |    9 ++++-\n cache.h              |    5 ++-\n entry.c              |   39 ++++++++++++------\n fast-import.c        |    2 +-\n object.c             |   11 +++++\n pack-check.c         |   21 +++++++++-\n sha1_file.c          |   78 +++++++++++++++++++++++++++++++-----\n t/t1050-large.sh     |   59 +++++++++++++++++++++++++++-\n wrapper.c            |   27 +++++++++++-\n 16 files changed, 400 insertions(+), 95 deletions(-)\n\n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"185499","messageId":"1330329315-11407-2-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330329315-11407-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 01/11] Add more large blob test cases","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-27T07:55:05Z","receivedAt":"2012-02-27T07:55:05Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"New test cases list commands that should work when memory is\nlimited. All memory allocation functions (*) learn to reject any\nallocation larger than $GIT_ALLOC_LIMIT if set.\n\n(*) Not exactly all. Some places do not use x* functions, but\nmalloc/calloc directly, notably diff-delta. These could path should\nnever be run on large blobs.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n t/t1050-large.sh |   59 +++++++++++++++++++++++++++++++++++++++++++++++++++++-\n wrapper.c        |   27 ++++++++++++++++++++++--\n 2 files changed, 82 insertions(+), 4 deletions(-)\n\ndiff --git a/t/t1050-large.sh b/t/t1050-large.sh\nindex 29d6024..f245e59 100755\n--- a/t/t1050-large.sh\n+++ b/t/t1050-large.sh\n@@ -10,7 +10,9 @@ test_expect_success setup '\n \techo X | dd of=large1 bs=1k seek=2000 &&\n \techo X | dd of=large2 bs=1k seek=2000 &&\n \techo X | dd of=large3 bs=1k seek=2000 &&\n-\techo Y | dd of=huge bs=1k seek=2500\n+\techo Y | dd of=huge bs=1k seek=2500 &&\n+\tGIT_ALLOC_LIMIT=1500 &&\n+\texport GIT_ALLOC_LIMIT\n '\n \n test_expect_success 'add a large file or two' '\n@@ -100,4 +102,59 @@ test_expect_success 'packsize limit' '\n \t)\n '\n \n+test_expect_success 'diff --raw' '\n+\tgit commit -q -m initial &&\n+\techo modified >>large1 &&\n+\tgit add large1 &&\n+\tgit commit -q -m modified &&\n+\tgit diff --raw HEAD^\n+'\n+\n+test_expect_success 'hash-object' '\n+\tgit hash-object large1\n+'\n+\n+test_expect_failure 'cat-file a large file' '\n+\tgit cat-file blob :large1 >/dev/null\n+'\n+\n+test_expect_failure 'git-show a large file' '\n+\tgit show :large1 >/dev/null\n+\n+'\n+\n+test_expect_failure 'clone' '\n+\tgit clone -n file://\"$PWD\"/.git new &&\n+\t(\n+\tcd new &&\n+\tgit config core.bigfilethreshold 200k &&\n+\tgit checkout master\n+\t)\n+'\n+\n+test_expect_failure 'fetch updates' '\n+\techo modified >> large1 &&\n+\tgit commit -q -a -m updated &&\n+\t(\n+\tcd new &&\n+\tgit fetch --keep # FIXME should not need --keep\n+\t)\n+'\n+\n+test_expect_failure 'fsck' '\n+\tgit fsck --full\n+'\n+\n+test_expect_success 'repack' '\n+\tgit repack -ad\n+'\n+\n+test_expect_failure 'tar achiving' '\n+\tgit archive --format=tar HEAD >/dev/null\n+'\n+\n+test_expect_failure 'zip achiving' '\n+\tgit archive --format=zip HEAD >/dev/null\n+'\n+\n test_done\ndiff --git a/wrapper.c b/wrapper.c\nindex 85f09df..d4c0972 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -9,6 +9,18 @@ static void do_nothing(size_t size)\n \n static void (*try_to_free_routine)(size_t size) = do_nothing;\n \n+static void memory_limit_check(size_t size)\n+{\n+\tstatic int limit = -1;\n+\tif (limit == -1) {\n+\t\tconst char *env = getenv(\"GIT_ALLOC_LIMIT\");\n+\t\tlimit = env ? atoi(env) * 1024 : 0;\n+\t}\n+\tif (limit && size > limit)\n+\t\tdie(\"attempting to allocate %d over limit %d\",\n+\t\t    size, limit);\n+}\n+\n try_to_free_t set_try_to_free_routine(try_to_free_t routine)\n {\n \ttry_to_free_t old = try_to_free_routine;\n@@ -32,7 +44,10 @@ char *xstrdup(const char *str)\n \n void *xmalloc(size_t size)\n {\n-\tvoid *ret = malloc(size);\n+\tvoid *ret;\n+\n+\tmemory_limit_check(size);\n+\tret = malloc(size);\n \tif (!ret && !size)\n \t\tret = malloc(1);\n \tif (!ret) {\n@@ -79,7 +94,10 @@ char *xstrndup(const char *str, size_t len)\n \n void *xrealloc(void *ptr, size_t size)\n {\n-\tvoid *ret = realloc(ptr, size);\n+\tvoid *ret;\n+\n+\tmemory_limit_check(size);\n+\tret = realloc(ptr, size);\n \tif (!ret && !size)\n \t\tret = realloc(ptr, 1);\n \tif (!ret) {\n@@ -95,7 +113,10 @@ void *xrealloc(void *ptr, size_t size)\n \n void *xcalloc(size_t nmemb, size_t size)\n {\n-\tvoid *ret = calloc(nmemb, size);\n+\tvoid *ret;\n+\n+\tmemory_limit_check(size * nmemb);\n+\tret = calloc(nmemb, size);\n \tif (!ret && (!nmemb || !size))\n \t\tret = calloc(1, 1);\n \tif (!ret) {\n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"185500","messageId":"1330329315-11407-3-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330329315-11407-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 02/11] Factor out and export large blob writing code to arbitrary file handle","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-27T07:55:06Z","receivedAt":"2012-02-27T07:55:06Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n cache.h |    3 +++\n entry.c |   39 ++++++++++++++++++++++++++-------------\n 2 files changed, 29 insertions(+), 13 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex e12b15f..6ce691b 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -937,6 +937,9 @@ struct checkout {\n \t\t refresh_cache:1;\n };\n \n+extern int streaming_write_sha1(int fd, int seekable, const unsigned char *sha1,\n+\t\t\t\tenum object_type exp_type,\n+\t\t\t\tstruct stream_filter *filter);\n extern int checkout_entry(struct cache_entry *ce, const struct checkout *state, char *topath);\n \n struct cache_def {\ndiff --git a/entry.c b/entry.c\nindex 852fea1..dde0d17 100644\n--- a/entry.c\n+++ b/entry.c\n@@ -115,26 +115,20 @@ static int fstat_output(int fd, const struct checkout *state, struct stat *st)\n \treturn 0;\n }\n \n-static int streaming_write_entry(struct cache_entry *ce, char *path,\n-\t\t\t\t struct stream_filter *filter,\n-\t\t\t\t const struct checkout *state, int to_tempfile,\n-\t\t\t\t int *fstat_done, struct stat *statbuf)\n+int streaming_write_sha1(int fd, int seekable, const unsigned char *sha1,\n+\t\t\t enum object_type exp_type,\n+\t\t\t struct stream_filter *filter)\n {\n \tstruct git_istream *st;\n \tenum object_type type;\n \tunsigned long sz;\n \tint result = -1;\n \tssize_t kept = 0;\n-\tint fd = -1;\n \n-\tst = open_istream(ce->sha1, &type, &sz, filter);\n+\tst = open_istream(sha1, &type, &sz, filter);\n \tif (!st)\n \t\treturn -1;\n-\tif (type != OBJ_BLOB)\n-\t\tgoto close_and_exit;\n-\n-\tfd = open_output_fd(path, ce, to_tempfile);\n-\tif (fd < 0)\n+\tif (exp_type != OBJ_ANY && type != exp_type)\n \t\tgoto close_and_exit;\n \n \tfor (;;) {\n@@ -144,7 +138,7 @@ static int streaming_write_entry(struct cache_entry *ce, char *path,\n \n \t\tif (!readlen)\n \t\t\tbreak;\n-\t\tif (sizeof(buf) == readlen) {\n+\t\tif (seekable && sizeof(buf) == readlen) {\n \t\t\tfor (holeto = 0; holeto < readlen; holeto++)\n \t\t\t\tif (buf[holeto])\n \t\t\t\t\tbreak;\n@@ -166,10 +160,29 @@ static int streaming_write_entry(struct cache_entry *ce, char *path,\n \tif (kept && (lseek(fd, kept - 1, SEEK_CUR) == (off_t) -1 ||\n \t\t     write(fd, \"\", 1) != 1))\n \t\tgoto close_and_exit;\n-\t*fstat_done = fstat_output(fd, state, statbuf);\n+\tresult = 0;\n \n close_and_exit:\n \tclose_istream(st);\n+\treturn result;\n+}\n+\n+static int streaming_write_entry(struct cache_entry *ce, char *path,\n+\t\t\t\t struct stream_filter *filter,\n+\t\t\t\t const struct checkout *state, int to_tempfile,\n+\t\t\t\t int *fstat_done, struct stat *statbuf)\n+{\n+\tint result = -1;\n+\tint fd = open_output_fd(path, ce, to_tempfile);\n+\tif (fd < 0)\n+\t\tgoto close_and_exit;\n+\n+\tif (streaming_write_sha1(fd, 1, ce->sha1, OBJ_BLOB, filter))\n+\t\tgoto close_and_exit;\n+\n+\t*fstat_done = fstat_output(fd, state, statbuf);\n+\n+close_and_exit:\n \tif (0 <= fd)\n \t\tresult = close(fd);\n \tif (result && 0 <= fd)\n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"185501","messageId":"1330329315-11407-4-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330329315-11407-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 03/11] cat-file: use streaming interface to print blobs","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-27T07:55:07Z","receivedAt":"2012-02-27T07:55:07Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/cat-file.c |   22 ++++++++++++++++++++++\n t/t1050-large.sh   |    2 +-\n 2 files changed, 23 insertions(+), 1 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 8ed501f..3f3b558 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -82,6 +82,24 @@ static void pprint_tag(const unsigned char *sha1, const char *buf, unsigned long\n \t\twrite_or_die(1, cp, endp - cp);\n }\n \n+static int write_blob(const unsigned char *sha1)\n+{\n+\tunsigned char new_sha1[20];\n+\n+\tif (sha1_object_info(sha1, NULL) == OBJ_TAG) {\n+\t\tenum object_type type;\n+\t\tunsigned long size;\n+\t\tchar *buffer = read_sha1_file(sha1, &type, &size);\n+\t\tif (memcmp(buffer, \"object \", 7) ||\n+\t\t    get_sha1_hex(buffer + 7, new_sha1))\n+\t\t\tdie(\"%s not a valid tag\", sha1_to_hex(sha1));\n+\t\tsha1 = new_sha1;\n+\t\tfree(buffer);\n+\t}\n+\n+\treturn streaming_write_sha1(1, 0, sha1, OBJ_BLOB, NULL);\n+}\n+\n static int cat_one_file(int opt, const char *exp_type, const char *obj_name)\n {\n \tunsigned char sha1[20];\n@@ -127,6 +145,8 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name)\n \t\t\treturn cmd_ls_tree(2, ls_args, NULL);\n \t\t}\n \n+\t\tif (type == OBJ_BLOB)\n+\t\t\treturn write_blob(sha1);\n \t\tbuf = read_sha1_file(sha1, &type, &size);\n \t\tif (!buf)\n \t\t\tdie(\"Cannot read object %s\", obj_name);\n@@ -149,6 +169,8 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name)\n \t\tbreak;\n \n \tcase 0:\n+\t\tif (type_from_string(exp_type) == OBJ_BLOB)\n+\t\t\treturn write_blob(sha1);\n \t\tbuf = read_object_with_reference(sha1, exp_type, &size, NULL);\n \t\tbreak;\n \ndiff --git a/t/t1050-large.sh b/t/t1050-large.sh\nindex f245e59..39a3e77 100755\n--- a/t/t1050-large.sh\n+++ b/t/t1050-large.sh\n@@ -114,7 +114,7 @@ test_expect_success 'hash-object' '\n \tgit hash-object large1\n '\n \n-test_expect_failure 'cat-file a large file' '\n+test_expect_success 'cat-file a large file' '\n \tgit cat-file blob :large1 >/dev/null\n '\n \n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"185502","messageId":"1330329315-11407-5-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330329315-11407-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 04/11] parse_object: special code path for blobs to avoid putting whole object in memory","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-27T07:55:08Z","receivedAt":"2012-02-27T07:55:08Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n object.c    |   11 +++++++++++\n sha1_file.c |   33 ++++++++++++++++++++++++++++++++-\n 2 files changed, 43 insertions(+), 1 deletions(-)\n\ndiff --git a/object.c b/object.c\nindex 6b06297..0498b18 100644\n--- a/object.c\n+++ b/object.c\n@@ -198,6 +198,17 @@ struct object *parse_object(const unsigned char *sha1)\n \tif (obj && obj->parsed)\n \t\treturn obj;\n \n+\tif ((obj && obj->type == OBJ_BLOB) ||\n+\t    (!obj && has_sha1_file(sha1) &&\n+\t     sha1_object_info(sha1, NULL) == OBJ_BLOB)) {\n+\t\tif (check_sha1_signature(repl, NULL, 0, NULL) < 0) {\n+\t\t\terror(\"sha1 mismatch %s\\n\", sha1_to_hex(repl));\n+\t\t\treturn NULL;\n+\t\t}\n+\t\tparse_blob_buffer(lookup_blob(sha1), NULL, 0);\n+\t\treturn lookup_object(sha1);\n+\t}\n+\n \tbuffer = read_sha1_file(sha1, &type, &size);\n \tif (buffer) {\n \t\tif (check_sha1_signature(repl, buffer, size, typename(type)) < 0) {\ndiff --git a/sha1_file.c b/sha1_file.c\nindex f9f8d5e..a77ef0a 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -19,6 +19,7 @@\n #include \"pack-revindex.h\"\n #include \"sha1-lookup.h\"\n #include \"bulk-checkin.h\"\n+#include \"streaming.h\"\n \n #ifndef O_NOATIME\n #if defined(__linux__) && (defined(__i386__) || defined(__PPC__))\n@@ -1149,7 +1150,37 @@ static const struct packed_git *has_packed_and_bad(const unsigned char *sha1)\n int check_sha1_signature(const unsigned char *sha1, void *map, unsigned long size, const char *type)\n {\n \tunsigned char real_sha1[20];\n-\thash_sha1_file(map, size, type, real_sha1);\n+\tenum object_type obj_type;\n+\tstruct git_istream *st;\n+\tgit_SHA_CTX c;\n+\tchar hdr[32];\n+\tint hdrlen;\n+\n+\tif (map) {\n+\t\thash_sha1_file(map, size, type, real_sha1);\n+\t\treturn hashcmp(sha1, real_sha1) ? -1 : 0;\n+\t}\n+\n+\tst = open_istream(sha1, &obj_type, &size, NULL);\n+\tif (!st)\n+\t\treturn -1;\n+\n+\t/* Generate the header */\n+\thdrlen = sprintf(hdr, \"%s %lu\", typename(obj_type), size) + 1;\n+\n+\t/* Sha1.. */\n+\tgit_SHA1_Init(&c);\n+\tgit_SHA1_Update(&c, hdr, hdrlen);\n+\tfor (;;) {\n+\t\tchar buf[1024 * 16];\n+\t\tssize_t readlen = read_istream(st, buf, sizeof(buf));\n+\n+\t\tif (!readlen)\n+\t\t\tbreak;\n+\t\tgit_SHA1_Update(&c, buf, readlen);\n+\t}\n+\tgit_SHA1_Final(real_sha1, &c);\n+\tclose_istream(st);\n \treturn hashcmp(sha1, real_sha1) ? -1 : 0;\n }\n \n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"185503","messageId":"1330329315-11407-6-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330329315-11407-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 05/11] show: use streaming interface for showing blobs","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-27T07:55:09Z","receivedAt":"2012-02-27T07:55:09Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/log.c    |    9 ++++++++-\n t/t1050-large.sh |    2 +-\n 2 files changed, 9 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/log.c b/builtin/log.c\nindex 7d1f6f8..4c4b17a 100644\n--- a/builtin/log.c\n+++ b/builtin/log.c\n@@ -386,13 +386,20 @@ static int show_object(const unsigned char *sha1, int show_tag_object,\n {\n \tunsigned long size;\n \tenum object_type type;\n-\tchar *buf = read_sha1_file(sha1, &type, &size);\n+\tchar *buf;\n \tint offset = 0;\n \n+\tif (!show_tag_object) {\n+\t\tfflush(stdout);\n+\t\treturn streaming_write_sha1(1, 0, sha1, OBJ_ANY, NULL);\n+\t}\n+\n+\tbuf = read_sha1_file(sha1, &type, &size);\n \tif (!buf)\n \t\treturn error(_(\"Could not read object %s\"), sha1_to_hex(sha1));\n \n \tif (show_tag_object)\n+\t\tassert(type == OBJ_TAG);\n \t\twhile (offset < size && buf[offset] != '\\n') {\n \t\t\tint new_offset = offset + 1;\n \t\t\twhile (new_offset < size && buf[new_offset++] != '\\n')\ndiff --git a/t/t1050-large.sh b/t/t1050-large.sh\nindex 39a3e77..66acb3b 100755\n--- a/t/t1050-large.sh\n+++ b/t/t1050-large.sh\n@@ -118,7 +118,7 @@ test_expect_success 'cat-file a large file' '\n \tgit cat-file blob :large1 >/dev/null\n '\n \n-test_expect_failure 'git-show a large file' '\n+test_expect_success 'git-show a large file' '\n \tgit show :large1 >/dev/null\n \n '\n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"185504","messageId":"1330329315-11407-7-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330329315-11407-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 06/11] index-pack --verify: skip sha-1 collision test","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-27T07:55:10Z","receivedAt":"2012-02-27T07:55:10Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"index-pack --verify (or verify-pack) is about verifying the pack\nitself. SHA-1 collision test is about outside (probably malicious)\nobjects with the same SHA-1 entering current repo.\n\nSHA-1 collision test is currently done unconditionally. Which means if\nyou verify an in-repo pack, all objects from the pack will be checked\nagainst objects in repo, which are themselves.\n\nSkip this test for --verify, unless --strict is also specified.\n\nlinux-2.6 $ ls -sh .git/objects/pack/pack-e7732c98a8d54840add294c3c562840f78764196.pack\n401M .git/objects/pack/pack-e7732c98a8d54840add294c3c562840f78764196.pack\n\nWithout the patch (and with another patch to cut out second pass in\nindex-pack):\n\nlinux-2.6 $ time ~/w/git/old index-pack -v --verify .git/objects/pack/pack-e7732c98a8d54840add294c3c562840f78764196.pack\nIndexing objects: 100% (1944656/1944656), done.\nfatal: pack has 1617280 unresolved deltas\n\nreal    1m1.223s\nuser    0m55.028s\nsys     0m0.828s\n\nWith the patch:\n\nlinux-2.6 $ time ~/w/git/git index-pack -v --verify .git/objects/pack/pack-e7732c98a8d54840add294c3c562840f78764196.pack\nIndexing objects: 100% (1944656/1944656), done.\nfatal: pack has 1617280 unresolved deltas\n\nreal    0m41.714s\nuser    0m40.994s\nsys     0m0.550s\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/index-pack.c |    5 +++--\n 1 files changed, 3 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex dd1c5c9..cee83b9 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -62,6 +62,7 @@ static int nr_resolved_deltas;\n \n static int from_stdin;\n static int strict;\n+static int verify;\n static int verbose;\n \n static struct progress *progress;\n@@ -461,7 +462,7 @@ static void sha1_object(const void *data, unsigned long size,\n \t\t\tenum object_type type, unsigned char *sha1)\n {\n \thash_sha1_file(data, size, typename(type), sha1);\n-\tif (has_sha1_file(sha1)) {\n+\tif ((strict || !verify) && has_sha1_file(sha1)) {\n \t\tvoid *has_data;\n \t\tenum object_type has_type;\n \t\tunsigned long has_size;\n@@ -1078,7 +1079,7 @@ static void show_pack_info(int stat_only)\n \n int cmd_index_pack(int argc, const char **argv, const char *prefix)\n {\n-\tint i, fix_thin_pack = 0, verify = 0, stat_only = 0, stat = 0;\n+\tint i, fix_thin_pack = 0, stat_only = 0, stat = 0;\n \tconst char *curr_pack, *curr_index;\n \tconst char *index_name = NULL, *pack_name = NULL;\n \tconst char *keep_name = NULL, *keep_msg = NULL;\n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"185507","messageId":"1330329315-11407-8-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330329315-11407-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 07/11] index-pack: split second pass obj handling into own function","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-27T07:55:11Z","receivedAt":"2012-02-27T07:55:11Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/index-pack.c |   31 ++++++++++++++++++-------------\n 1 files changed, 18 insertions(+), 13 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex cee83b9..e3cb684 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -683,6 +683,23 @@ static int compare_delta_entry(const void *a, const void *b)\n \t\t\t\t   objects[delta_b->obj_no].type);\n }\n \n+/*\n+ * Second pass:\n+ * - for all non-delta objects, look if it is used as a base for\n+ *   deltas;\n+ * - if used as a base, uncompress the object and apply all deltas,\n+ *   recursively checking if the resulting object is used as a base\n+ *   for some more deltas.\n+ */\n+static void second_pass(struct object_entry *obj)\n+{\n+\tstruct base_data *base_obj = alloc_base_data();\n+\tbase_obj->obj = obj;\n+\tbase_obj->data = NULL;\n+\tfind_unresolved_deltas(base_obj);\n+\tdisplay_progress(progress, nr_resolved_deltas);\n+}\n+\n /* Parse all objects and return the pack content SHA1 hash */\n static void parse_pack_objects(unsigned char *sha1)\n {\n@@ -737,26 +754,14 @@ static void parse_pack_objects(unsigned char *sha1)\n \tqsort(deltas, nr_deltas, sizeof(struct delta_entry),\n \t      compare_delta_entry);\n \n-\t/*\n-\t * Second pass:\n-\t * - for all non-delta objects, look if it is used as a base for\n-\t *   deltas;\n-\t * - if used as a base, uncompress the object and apply all deltas,\n-\t *   recursively checking if the resulting object is used as a base\n-\t *   for some more deltas.\n-\t */\n \tif (verbose)\n \t\tprogress = start_progress(\"Resolving deltas\", nr_deltas);\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tstruct object_entry *obj = &objects[i];\n-\t\tstruct base_data *base_obj = alloc_base_data();\n \n \t\tif (is_delta_type(obj->type))\n \t\t\tcontinue;\n-\t\tbase_obj->obj = obj;\n-\t\tbase_obj->data = NULL;\n-\t\tfind_unresolved_deltas(base_obj);\n-\t\tdisplay_progress(progress, nr_resolved_deltas);\n+\t\tsecond_pass(obj);\n \t}\n }\n \n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"185506","messageId":"1330329315-11407-9-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330329315-11407-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 08/11] index-pack: reduce memory usage when the pack has large blobs","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-27T07:55:12Z","receivedAt":"2012-02-27T07:55:12Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"This command unpacks every non-delta objects in order to:\n\n1. calculate sha-1\n2. do byte-to-byte sha-1 collision test if we happen to have objects\n   with the same sha-1\n3. validate object content in strict mode\n\nAll this requires the entire object to stay in memory, a bad news for\ngiant blobs. This patch lowers memory consumption by not saving the\nobject in memory whenever possible, calculating SHA-1 while unpacking\nthe object.\n\nThis patch assumes that the collision test is rarely needed. The\ncollision test will be done later in second pass if necessary, which\nputs the entire object back to memory again (We could even do the\ncollision test without putting the entire object back in memory, by\ncomparing as we unpack it).\n\nIn strict mode, it always keeps non-blob objects in memory for\nvalidation (blobs do not need data validation). \"--strict --verify\"\nalso keeps blobs in memory.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/index-pack.c |   74 +++++++++++++++++++++++++++++++++++++++++---------\n t/t1050-large.sh     |    4 +-\n 2 files changed, 63 insertions(+), 15 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex e3cb684..86de813 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -277,30 +277,60 @@ static void unlink_base_data(struct base_data *c)\n \tfree_base_data(c);\n }\n \n-static void *unpack_entry_data(unsigned long offset, unsigned long size)\n+static void *unpack_entry_data(unsigned long offset, unsigned long size,\n+\t\t\t       enum object_type type, unsigned char *sha1)\n {\n+\tstatic char fixed_buf[8192];\n \tint status;\n \tgit_zstream stream;\n-\tvoid *buf = xmalloc(size);\n+\tvoid *buf;\n+\tgit_SHA_CTX c;\n+\n+\tif (sha1) {\t\t/* do hash_sha1_file internally */\n+\t\tchar hdr[32];\n+\t\tint hdrlen = sprintf(hdr, \"%s %lu\", typename(type), size)+1;\n+\t\tgit_SHA1_Init(&c);\n+\t\tgit_SHA1_Update(&c, hdr, hdrlen);\n+\n+\t\tbuf = fixed_buf;\n+\t} else {\n+\t\tbuf = xmalloc(size);\n+\t}\n \n \tmemset(&stream, 0, sizeof(stream));\n \tgit_inflate_init(&stream);\n \tstream.next_out = buf;\n-\tstream.avail_out = size;\n+\tstream.avail_out = buf == fixed_buf ? sizeof(fixed_buf) : size;\n \n \tdo {\n \t\tstream.next_in = fill(1);\n \t\tstream.avail_in = input_len;\n \t\tstatus = git_inflate(&stream, 0);\n \t\tuse(input_len - stream.avail_in);\n+\t\tif (sha1) {\n+\t\t\tgit_SHA1_Update(&c, buf, stream.next_out - (unsigned char *)buf);\n+\t\t\tstream.next_out = buf;\n+\t\t\tstream.avail_out = sizeof(fixed_buf);\n+\t\t}\n \t} while (status == Z_OK);\n \tif (stream.total_out != size || status != Z_STREAM_END)\n \t\tbad_object(offset, \"inflate returned %d\", status);\n \tgit_inflate_end(&stream);\n+\tif (sha1) {\n+\t\tgit_SHA1_Final(sha1, &c);\n+\t\tbuf = NULL;\n+\t}\n \treturn buf;\n }\n \n-static void *unpack_raw_entry(struct object_entry *obj, union delta_base *delta_base)\n+static int is_delta_type(enum object_type type)\n+{\n+\treturn (type == OBJ_REF_DELTA || type == OBJ_OFS_DELTA);\n+}\n+\n+static void *unpack_raw_entry(struct object_entry *obj,\n+\t\t\t      union delta_base *delta_base,\n+\t\t\t      unsigned char *sha1)\n {\n \tunsigned char *p;\n \tunsigned long size, c;\n@@ -360,7 +390,17 @@ static void *unpack_raw_entry(struct object_entry *obj, union delta_base *delta_\n \t}\n \tobj->hdr_size = consumed_bytes - obj->idx.offset;\n \n-\tdata = unpack_entry_data(obj->idx.offset, obj->size);\n+\t/*\n+\t * --verify --strict: sha1_object() does all collision test\n+\t *          --strict: sha1_object() does all except blobs,\n+\t *                    blobs tested in second pass\n+\t * --verify         : no collision test\n+\t *                  : all in second pass\n+\t */\n+\tif (is_delta_type(obj->type) ||\n+\t    (strict && (verify || obj->type != OBJ_BLOB)))\n+\t\tsha1 = NULL;\t/* save unpacked object */\n+\tdata = unpack_entry_data(obj->idx.offset, obj->size, obj->type, sha1);\n \tobj->idx.crc32 = input_crc32;\n \treturn data;\n }\n@@ -461,8 +501,9 @@ static void find_delta_children(const union delta_base *base,\n static void sha1_object(const void *data, unsigned long size,\n \t\t\tenum object_type type, unsigned char *sha1)\n {\n-\thash_sha1_file(data, size, typename(type), sha1);\n-\tif ((strict || !verify) && has_sha1_file(sha1)) {\n+\tif (data)\n+\t\thash_sha1_file(data, size, typename(type), sha1);\n+\tif (data && (strict || !verify) && has_sha1_file(sha1)) {\n \t\tvoid *has_data;\n \t\tenum object_type has_type;\n \t\tunsigned long has_size;\n@@ -511,11 +552,6 @@ static void sha1_object(const void *data, unsigned long size,\n \t}\n }\n \n-static int is_delta_type(enum object_type type)\n-{\n-\treturn (type == OBJ_REF_DELTA || type == OBJ_OFS_DELTA);\n-}\n-\n /*\n  * This function is part of find_unresolved_deltas(). There are two\n  * walkers going in the opposite ways.\n@@ -690,10 +726,22 @@ static int compare_delta_entry(const void *a, const void *b)\n  * - if used as a base, uncompress the object and apply all deltas,\n  *   recursively checking if the resulting object is used as a base\n  *   for some more deltas.\n+ * - if the same object exists in repository and we're not in strict\n+ *   mode, we skipped the sha-1 collision test in the first pass.\n+ *   Do it now.\n  */\n static void second_pass(struct object_entry *obj)\n {\n \tstruct base_data *base_obj = alloc_base_data();\n+\n+\tif (((!strict && !verify) ||\n+\t     (strict && !verify && obj->type == OBJ_BLOB)) &&\n+\t    has_sha1_file(obj->idx.sha1)) {\n+\t\tvoid *data = get_data_from_pack(obj);\n+\t\tsha1_object(data, obj->size, obj->type, obj->idx.sha1);\n+\t\tfree(data);\n+\t}\n+\n \tbase_obj->obj = obj;\n \tbase_obj->data = NULL;\n \tfind_unresolved_deltas(base_obj);\n@@ -719,7 +767,7 @@ static void parse_pack_objects(unsigned char *sha1)\n \t\t\t\tnr_objects);\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tstruct object_entry *obj = &objects[i];\n-\t\tvoid *data = unpack_raw_entry(obj, &delta->base);\n+\t\tvoid *data = unpack_raw_entry(obj, &delta->base, obj->idx.sha1);\n \t\tobj->real_type = obj->type;\n \t\tif (is_delta_type(obj->type)) {\n \t\t\tnr_deltas++;\ndiff --git a/t/t1050-large.sh b/t/t1050-large.sh\nindex 66acb3b..7e78c72 100755\n--- a/t/t1050-large.sh\n+++ b/t/t1050-large.sh\n@@ -123,7 +123,7 @@ test_expect_success 'git-show a large file' '\n \n '\n \n-test_expect_failure 'clone' '\n+test_expect_success 'clone' '\n \tgit clone -n file://\"$PWD\"/.git new &&\n \t(\n \tcd new &&\n@@ -132,7 +132,7 @@ test_expect_failure 'clone' '\n \t)\n '\n \n-test_expect_failure 'fetch updates' '\n+test_expect_success 'fetch updates' '\n \techo modified >> large1 &&\n \tgit commit -q -a -m updated &&\n \t(\n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"185505","messageId":"1330329315-11407-10-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330329315-11407-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 09/11] pack-check: do not unpack blobs","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-27T07:55:13Z","receivedAt":"2012-02-27T07:55:13Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"blob content is not used by verify_pack caller (currently only fsck),\nwe only need to make sure blob sha-1 signature matches its\ncontent. unpack_entry() is taught to hash pack entry as it is\nunpacked, eliminating the need to keep whole blob in memory.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n cache.h          |    2 +-\n fast-import.c    |    2 +-\n pack-check.c     |   21 ++++++++++++++++++++-\n sha1_file.c      |   45 +++++++++++++++++++++++++++++++++++----------\n t/t1050-large.sh |    2 +-\n 5 files changed, 58 insertions(+), 14 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex 6ce691b..33bfb69 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1065,7 +1065,7 @@ extern const unsigned char *nth_packed_object_sha1(struct packed_git *, uint32_t\n extern off_t nth_packed_object_offset(const struct packed_git *, uint32_t);\n extern off_t find_pack_entry_one(const unsigned char *, struct packed_git *);\n extern int is_pack_valid(struct packed_git *);\n-extern void *unpack_entry(struct packed_git *, off_t, enum object_type *, unsigned long *);\n+extern void *unpack_entry(struct packed_git *, off_t, enum object_type *, unsigned long *, unsigned char *);\n extern unsigned long unpack_object_header_buffer(const unsigned char *buf, unsigned long len, enum object_type *type, unsigned long *sizep);\n extern unsigned long get_size_from_delta(struct packed_git *, struct pack_window **, off_t);\n extern int unpack_object_header(struct packed_git *, struct pack_window **, off_t *, unsigned long *);\ndiff --git a/fast-import.c b/fast-import.c\nindex 6cd19e5..5e94a64 100644\n--- a/fast-import.c\n+++ b/fast-import.c\n@@ -1303,7 +1303,7 @@ static void *gfi_unpack_entry(\n \t\t */\n \t\tp->pack_size = pack_size + 20;\n \t}\n-\treturn unpack_entry(p, oe->idx.offset, &type, sizep);\n+\treturn unpack_entry(p, oe->idx.offset, &type, sizep, NULL);\n }\n \n static const char *get_mode(const char *str, uint16_t *modep)\ndiff --git a/pack-check.c b/pack-check.c\nindex 63a595c..1920bdb 100644\n--- a/pack-check.c\n+++ b/pack-check.c\n@@ -105,6 +105,7 @@ static int verify_packfile(struct packed_git *p,\n \t\tvoid *data;\n \t\tenum object_type type;\n \t\tunsigned long size;\n+\t\toff_t curpos = entries[i].offset;\n \n \t\tif (p->index_version > 1) {\n \t\t\toff_t offset = entries[i].offset;\n@@ -116,7 +117,25 @@ static int verify_packfile(struct packed_git *p,\n \t\t\t\t\t    sha1_to_hex(entries[i].sha1),\n \t\t\t\t\t    p->pack_name, (uintmax_t)offset);\n \t\t}\n-\t\tdata = unpack_entry(p, entries[i].offset, &type, &size);\n+\t\ttype = unpack_object_header(p, w_curs, &curpos, &size);\n+\t\tunuse_pack(w_curs);\n+\t\tif (type == OBJ_BLOB) {\n+\t\t\tunsigned char sha1[20];\n+\t\t\tdata = unpack_entry(p, entries[i].offset, &type, &size, sha1);\n+\t\t\tif (!data) {\n+\t\t\t\tif (hashcmp(entries[i].sha1, sha1))\n+\t\t\t\t\terr = error(\"packed %s from %s is corrupt\",\n+\t\t\t\t\t\t    sha1_to_hex(entries[i].sha1), p->pack_name);\n+\t\t\t\telse if (fn) {\n+\t\t\t\t\tint eaten = 0;\n+\t\t\t\t\tfn(entries[i].sha1, type, size, NULL, &eaten);\n+\t\t\t\t}\n+\t\t\t\tif (((base_count + i) & 1023) == 0)\n+\t\t\t\t\tdisplay_progress(progress, base_count + i);\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t}\n+\t\tdata = unpack_entry(p, entries[i].offset, &type, &size, NULL);\n \t\tif (!data)\n \t\t\terr = error(\"cannot unpack %s from %s at offset %\"PRIuMAX\"\",\n \t\t\t\t    sha1_to_hex(entries[i].sha1), p->pack_name,\ndiff --git a/sha1_file.c b/sha1_file.c\nindex a77ef0a..d68a5b0 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -1653,28 +1653,51 @@ static int packed_object_info(struct packed_git *p, off_t obj_offset,\n }\n \n static void *unpack_compressed_entry(struct packed_git *p,\n-\t\t\t\t    struct pack_window **w_curs,\n-\t\t\t\t    off_t curpos,\n-\t\t\t\t    unsigned long size)\n+\t\t\t\t     struct pack_window **w_curs,\n+\t\t\t\t     off_t curpos,\n+\t\t\t\t     unsigned long size,\n+\t\t\t\t     enum object_type type,\n+\t\t\t\t     unsigned char *sha1)\n {\n+\tstatic unsigned char fixed_buf[8192];\n \tint st;\n \tgit_zstream stream;\n \tunsigned char *buffer, *in;\n+\tgit_SHA_CTX c;\n+\n+\tif (sha1) {\t\t/* do hash_sha1_file internally */\n+\t\tchar hdr[32];\n+\t\tint hdrlen = sprintf(hdr, \"%s %lu\", typename(type), size)+1;\n+\t\tgit_SHA1_Init(&c);\n+\t\tgit_SHA1_Update(&c, hdr, hdrlen);\n+\n+\t\tbuffer = fixed_buf;\n+\t} else {\n+\t\tbuffer = xmallocz(size);\n+\t}\n \n-\tbuffer = xmallocz(size);\n \tmemset(&stream, 0, sizeof(stream));\n \tstream.next_out = buffer;\n-\tstream.avail_out = size + 1;\n+\tstream.avail_out = buffer == fixed_buf ? sizeof(fixed_buf) : size + 1;\n \n \tgit_inflate_init(&stream);\n \tdo {\n \t\tin = use_pack(p, w_curs, curpos, &stream.avail_in);\n \t\tstream.next_in = in;\n \t\tst = git_inflate(&stream, Z_FINISH);\n-\t\tif (!stream.avail_out)\n+\t\tif (sha1) {\n+\t\t\tgit_SHA1_Update(&c, buffer, stream.next_out - (unsigned char *)buffer);\n+\t\t\tstream.next_out = buffer;\n+\t\t\tstream.avail_out = sizeof(fixed_buf);\n+\t\t}\n+\t\telse if (!stream.avail_out)\n \t\t\tbreak; /* the payload is larger than it should be */\n \t\tcurpos += stream.next_in - in;\n \t} while (st == Z_OK || st == Z_BUF_ERROR);\n+\tif (sha1) {\n+\t\tgit_SHA1_Final(sha1, &c);\n+\t\tbuffer = NULL;\n+\t}\n \tgit_inflate_end(&stream);\n \tif ((st != Z_STREAM_END) || stream.total_out != size) {\n \t\tfree(buffer);\n@@ -1727,7 +1750,7 @@ static void *cache_or_unpack_entry(struct packed_git *p, off_t base_offset,\n \n \tret = ent->data;\n \tif (!ret || ent->p != p || ent->base_offset != base_offset)\n-\t\treturn unpack_entry(p, base_offset, type, base_size);\n+\t\treturn unpack_entry(p, base_offset, type, base_size, NULL);\n \n \tif (!keep_cache) {\n \t\tent->data = NULL;\n@@ -1844,7 +1867,7 @@ static void *unpack_delta_entry(struct packed_git *p,\n \t\t\treturn NULL;\n \t}\n \n-\tdelta_data = unpack_compressed_entry(p, w_curs, curpos, delta_size);\n+\tdelta_data = unpack_compressed_entry(p, w_curs, curpos, delta_size, OBJ_NONE, NULL);\n \tif (!delta_data) {\n \t\terror(\"failed to unpack compressed delta \"\n \t\t      \"at offset %\"PRIuMAX\" from %s\",\n@@ -1883,7 +1906,8 @@ static void write_pack_access_log(struct packed_git *p, off_t obj_offset)\n int do_check_packed_object_crc;\n \n void *unpack_entry(struct packed_git *p, off_t obj_offset,\n-\t\t   enum object_type *type, unsigned long *sizep)\n+\t\t   enum object_type *type, unsigned long *sizep,\n+\t\t   unsigned char *sha1)\n {\n \tstruct pack_window *w_curs = NULL;\n \toff_t curpos = obj_offset;\n@@ -1917,7 +1941,8 @@ void *unpack_entry(struct packed_git *p, off_t obj_offset,\n \tcase OBJ_TREE:\n \tcase OBJ_BLOB:\n \tcase OBJ_TAG:\n-\t\tdata = unpack_compressed_entry(p, &w_curs, curpos, *sizep);\n+\t\tdata = unpack_compressed_entry(p, &w_curs, curpos,\n+\t\t\t\t\t       *sizep, *type, sha1);\n \t\tbreak;\n \tdefault:\n \t\tdata = NULL;\ndiff --git a/t/t1050-large.sh b/t/t1050-large.sh\nindex 7e78c72..c749ecb 100755\n--- a/t/t1050-large.sh\n+++ b/t/t1050-large.sh\n@@ -141,7 +141,7 @@ test_expect_success 'fetch updates' '\n \t)\n '\n \n-test_expect_failure 'fsck' '\n+test_expect_success 'fsck' '\n \tgit fsck --full\n '\n \n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"185508","messageId":"1330329315-11407-11-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330329315-11407-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 10/11] archive: support streaming large files to a tar archive","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-27T07:55:14Z","receivedAt":"2012-02-27T07:55:14Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n archive-tar.c    |   35 ++++++++++++++++++++++++++++-------\n archive-zip.c    |    9 +++++----\n archive.c        |   51 ++++++++++++++++++++++++++++++++++-----------------\n archive.h        |   11 +++++++++--\n t/t1050-large.sh |    2 +-\n 5 files changed, 77 insertions(+), 31 deletions(-)\n\ndiff --git a/archive-tar.c b/archive-tar.c\nindex 20af005..5bffe49 100644\n--- a/archive-tar.c\n+++ b/archive-tar.c\n@@ -5,6 +5,7 @@\n #include \"tar.h\"\n #include \"archive.h\"\n #include \"run-command.h\"\n+#include \"streaming.h\"\n \n #define RECORDSIZE\t(512)\n #define BLOCKSIZE\t(RECORDSIZE * 20)\n@@ -123,9 +124,29 @@ static size_t get_path_prefix(const char *path, size_t pathlen, size_t maxlen)\n \treturn i;\n }\n \n+static void write_file(struct git_istream *stream, const void *buffer,\n+\t\t       unsigned long size)\n+{\n+\tif (!stream) {\n+\t\twrite_blocked(buffer, size);\n+\t\treturn;\n+\t}\n+\tfor (;;) {\n+\t\tchar buf[1024 * 16];\n+\t\tssize_t readlen;\n+\n+\t\treadlen = read_istream(stream, buf, sizeof(buf));\n+\n+\t\tif (!readlen)\n+\t\t\tbreak;\n+\t\twrite_blocked(buf, readlen);\n+\t}\n+}\n+\n static int write_tar_entry(struct archiver_args *args,\n-\t\tconst unsigned char *sha1, const char *path, size_t pathlen,\n-\t\tunsigned int mode, void *buffer, unsigned long size)\n+\t\t\t   const unsigned char *sha1, const char *path,\n+\t\t\t   size_t pathlen, unsigned int mode, void *buffer,\n+\t\t\t   struct git_istream *stream, unsigned long size)\n {\n \tstruct ustar_header header;\n \tstruct strbuf ext_header = STRBUF_INIT;\n@@ -200,14 +221,14 @@ static int write_tar_entry(struct archiver_args *args,\n \n \tif (ext_header.len > 0) {\n \t\terr = write_tar_entry(args, sha1, NULL, 0, 0, ext_header.buf,\n-\t\t\t\text_header.len);\n+\t\t\t\t      NULL, ext_header.len);\n \t\tif (err)\n \t\t\treturn err;\n \t}\n \tstrbuf_release(&ext_header);\n \twrite_blocked(&header, sizeof(header));\n-\tif (S_ISREG(mode) && buffer && size > 0)\n-\t\twrite_blocked(buffer, size);\n+\tif (S_ISREG(mode) && size > 0)\n+\t\twrite_file(stream, buffer, size);\n \treturn err;\n }\n \n@@ -219,7 +240,7 @@ static int write_global_extended_header(struct archiver_args *args)\n \n \tstrbuf_append_ext_header(&ext_header, \"comment\", sha1_to_hex(sha1), 40);\n \terr = write_tar_entry(args, NULL, NULL, 0, 0, ext_header.buf,\n-\t\t\text_header.len);\n+\t\t\t      NULL, ext_header.len);\n \tstrbuf_release(&ext_header);\n \treturn err;\n }\n@@ -308,7 +329,7 @@ static int write_tar_archive(const struct archiver *ar,\n \tif (args->commit_sha1)\n \t\terr = write_global_extended_header(args);\n \tif (!err)\n-\t\terr = write_archive_entries(args, write_tar_entry);\n+\t\terr = write_archive_entries(args, write_tar_entry, 1);\n \tif (!err)\n \t\twrite_trailer();\n \treturn err;\ndiff --git a/archive-zip.c b/archive-zip.c\nindex 02d1f37..4a1e917 100644\n--- a/archive-zip.c\n+++ b/archive-zip.c\n@@ -120,9 +120,10 @@ static void *zlib_deflate(void *data, unsigned long size,\n \treturn buffer;\n }\n \n-static int write_zip_entry(struct archiver_args *args,\n-\t\tconst unsigned char *sha1, const char *path, size_t pathlen,\n-\t\tunsigned int mode, void *buffer, unsigned long size)\n+int write_zip_entry(struct archiver_args *args,\n+\t\t\t   const unsigned char *sha1, const char *path,\n+\t\t\t   size_t pathlen, unsigned int mode, void *buffer,\n+\t\t\t   struct git_istream *stream, unsigned long size)\n {\n \tstruct zip_local_header header;\n \tstruct zip_dir_header dirent;\n@@ -271,7 +272,7 @@ static int write_zip_archive(const struct archiver *ar,\n \tzip_dir = xmalloc(ZIP_DIRECTORY_MIN_SIZE);\n \tzip_dir_size = ZIP_DIRECTORY_MIN_SIZE;\n \n-\terr = write_archive_entries(args, write_zip_entry);\n+\terr = write_archive_entries(args, write_zip_entry, 0);\n \tif (!err)\n \t\twrite_zip_trailer(args->commit_sha1);\n \ndiff --git a/archive.c b/archive.c\nindex 1ee837d..257eadf 100644\n--- a/archive.c\n+++ b/archive.c\n@@ -5,6 +5,7 @@\n #include \"archive.h\"\n #include \"parse-options.h\"\n #include \"unpack-trees.h\"\n+#include \"streaming.h\"\n \n static char const * const archive_usage[] = {\n \t\"git archive [options] <tree-ish> [<path>...]\",\n@@ -59,26 +60,35 @@ static void format_subst(const struct commit *commit,\n \tfree(to_free);\n }\n \n-static void *sha1_file_to_archive(const char *path, const unsigned char *sha1,\n-\t\tunsigned int mode, enum object_type *type,\n-\t\tunsigned long *sizep, const struct commit *commit)\n+void sha1_file_to_archive(void **buffer, struct git_istream **stream,\n+\t\t\t  const char *path, const unsigned char *sha1,\n+\t\t\t  unsigned int mode, enum object_type *type,\n+\t\t\t  unsigned long *sizep,\n+\t\t\t  const struct commit *commit)\n {\n-\tvoid *buffer;\n+\tif (stream) {\n+\t\tstruct stream_filter *filter;\n+\t\tfilter = get_stream_filter(path, sha1);\n+\t\tif (!commit && S_ISREG(mode) && is_null_stream_filter(filter)) {\n+\t\t\t*buffer = NULL;\n+\t\t\t*stream = open_istream(sha1, type, sizep, NULL);\n+\t\t\treturn;\n+\t\t}\n+\t\t*stream = NULL;\n+\t}\n \n-\tbuffer = read_sha1_file(sha1, type, sizep);\n-\tif (buffer && S_ISREG(mode)) {\n+\t*buffer = read_sha1_file(sha1, type, sizep);\n+\tif (*buffer && S_ISREG(mode)) {\n \t\tstruct strbuf buf = STRBUF_INIT;\n \t\tsize_t size = 0;\n \n-\t\tstrbuf_attach(&buf, buffer, *sizep, *sizep + 1);\n+\t\tstrbuf_attach(&buf, *buffer, *sizep, *sizep + 1);\n \t\tconvert_to_working_tree(path, buf.buf, buf.len, &buf);\n \t\tif (commit)\n \t\t\tformat_subst(commit, buf.buf, buf.len, &buf);\n-\t\tbuffer = strbuf_detach(&buf, &size);\n+\t\t*buffer = strbuf_detach(&buf, &size);\n \t\t*sizep = size;\n \t}\n-\n-\treturn buffer;\n }\n \n static void setup_archive_check(struct git_attr_check *check)\n@@ -97,6 +107,7 @@ static void setup_archive_check(struct git_attr_check *check)\n struct archiver_context {\n \tstruct archiver_args *args;\n \twrite_archive_entry_fn_t write_entry;\n+\tint stream_ok;\n };\n \n static int write_archive_entry(const unsigned char *sha1, const char *base,\n@@ -109,6 +120,7 @@ static int write_archive_entry(const unsigned char *sha1, const char *base,\n \twrite_archive_entry_fn_t write_entry = c->write_entry;\n \tstruct git_attr_check check[2];\n \tconst char *path_without_prefix;\n+\tstruct git_istream *stream = NULL;\n \tint convert = 0;\n \tint err;\n \tenum object_type type;\n@@ -133,25 +145,29 @@ static int write_archive_entry(const unsigned char *sha1, const char *base,\n \t\tstrbuf_addch(&path, '/');\n \t\tif (args->verbose)\n \t\t\tfprintf(stderr, \"%.*s\\n\", (int)path.len, path.buf);\n-\t\terr = write_entry(args, sha1, path.buf, path.len, mode, NULL, 0);\n+\t\terr = write_entry(args, sha1, path.buf, path.len, mode, NULL, NULL, 0);\n \t\tif (err)\n \t\t\treturn err;\n \t\treturn (S_ISDIR(mode) ? READ_TREE_RECURSIVE : 0);\n \t}\n \n-\tbuffer = sha1_file_to_archive(path_without_prefix, sha1, mode,\n-\t\t\t&type, &size, convert ? args->commit : NULL);\n-\tif (!buffer)\n+\tsha1_file_to_archive(&buffer, c->stream_ok ? &stream : NULL,\n+\t\t\t     path_without_prefix, sha1, mode,\n+\t\t\t     &type, &size, convert ? args->commit : NULL);\n+\tif (!buffer && !stream)\n \t\treturn error(\"cannot read %s\", sha1_to_hex(sha1));\n \tif (args->verbose)\n \t\tfprintf(stderr, \"%.*s\\n\", (int)path.len, path.buf);\n-\terr = write_entry(args, sha1, path.buf, path.len, mode, buffer, size);\n+\terr = write_entry(args, sha1, path.buf, path.len, mode, buffer, stream, size);\n+\tif (stream)\n+\t\tclose_istream(stream);\n \tfree(buffer);\n \treturn err;\n }\n \n int write_archive_entries(struct archiver_args *args,\n-\t\twrite_archive_entry_fn_t write_entry)\n+\t\t\t  write_archive_entry_fn_t write_entry,\n+\t\t\t  int stream_ok)\n {\n \tstruct archiver_context context;\n \tstruct unpack_trees_options opts;\n@@ -167,13 +183,14 @@ int write_archive_entries(struct archiver_args *args,\n \t\tif (args->verbose)\n \t\t\tfprintf(stderr, \"%.*s\\n\", (int)len, args->base);\n \t\terr = write_entry(args, args->tree->object.sha1, args->base,\n-\t\t\t\tlen, 040777, NULL, 0);\n+\t\t\t\t  len, 040777, NULL, NULL, 0);\n \t\tif (err)\n \t\t\treturn err;\n \t}\n \n \tcontext.args = args;\n \tcontext.write_entry = write_entry;\n+\tcontext.stream_ok = stream_ok;\n \n \t/*\n \t * Setup index and instruct attr to read index only\ndiff --git a/archive.h b/archive.h\nindex 2b0884f..370cca9 100644\n--- a/archive.h\n+++ b/archive.h\n@@ -27,9 +27,16 @@ extern void register_archiver(struct archiver *);\n extern void init_tar_archiver(void);\n extern void init_zip_archiver(void);\n \n-typedef int (*write_archive_entry_fn_t)(struct archiver_args *args, const unsigned char *sha1, const char *path, size_t pathlen, unsigned int mode, void *buffer, unsigned long size);\n+struct git_istream;\n+typedef int (*write_archive_entry_fn_t)(struct archiver_args *args,\n+\t\t\t\t\tconst unsigned char *sha1,\n+\t\t\t\t\tconst char *path, size_t pathlen,\n+\t\t\t\t\tunsigned int mode,\n+\t\t\t\t\tvoid *buffer,\n+\t\t\t\t\tstruct git_istream *stream,\n+\t\t\t\t\tunsigned long size);\n \n-extern int write_archive_entries(struct archiver_args *args, write_archive_entry_fn_t write_entry);\n+extern int write_archive_entries(struct archiver_args *args, write_archive_entry_fn_t write_entry, int stream_ok);\n extern int write_archive(int argc, const char **argv, const char *prefix, int setup_prefix, const char *name_hint, int remote);\n \n const char *archive_format_from_filename(const char *filename);\ndiff --git a/t/t1050-large.sh b/t/t1050-large.sh\nindex c749ecb..1e64692 100755\n--- a/t/t1050-large.sh\n+++ b/t/t1050-large.sh\n@@ -149,7 +149,7 @@ test_expect_success 'repack' '\n \tgit repack -ad\n '\n \n-test_expect_failure 'tar achiving' '\n+test_expect_success 'tar achiving' '\n \tgit archive --format=tar HEAD >/dev/null\n '\n \n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"185509","messageId":"1330329315-11407-12-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330329315-11407-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 11/11] fsck: use streaming interface for writing lost-found blobs","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-27T07:55:15Z","receivedAt":"2012-02-27T07:55:15Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/fsck.c |    8 ++------\n 1 files changed, 2 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/fsck.c b/builtin/fsck.c\nindex 8c479a7..319b5c7 100644\n--- a/builtin/fsck.c\n+++ b/builtin/fsck.c\n@@ -236,13 +236,9 @@ static void check_unreachable_object(struct object *obj)\n \t\t\tif (!(f = fopen(filename, \"w\")))\n \t\t\t\tdie_errno(\"Could not open '%s'\", filename);\n \t\t\tif (obj->type == OBJ_BLOB) {\n-\t\t\t\tenum object_type type;\n-\t\t\t\tunsigned long size;\n-\t\t\t\tchar *buf = read_sha1_file(obj->sha1,\n-\t\t\t\t\t\t&type, &size);\n-\t\t\t\tif (buf && fwrite(buf, 1, size, f) != size)\n+\t\t\t\tif (streaming_write_sha1(fileno(f), 1,\n+\t\t\t\t\t\t\t obj->sha1, OBJ_BLOB, NULL))\n \t\t\t\t\tdie_errno(\"Could not write '%s'\", filename);\n-\t\t\t\tfree(buf);\n \t\t\t} else\n \t\t\t\tfprintf(f, \"%s\\n\", sha1_to_hex(obj->sha1));\n \t\t\tif (fclose(f))\n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"185525","messageId":"7v4nucb2xl.fsf@alter.siamese.dyndns.org","threadId":"29753","inReplyTo":"1330329315-11407-3-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH 02/11] Factor out and export large blob writing code to arbitrary file handle","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-02-27T17:29:10Z","receivedAt":"2012-02-27T17:29:10Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n>  cache.h |    3 +++\n>  entry.c |   39 ++++++++++++++++++++++++++-------------\n>  2 files changed, 29 insertions(+), 13 deletions(-)\n\nIt was the goal of the original streaming output topic to helping more\ncallers stream the data out directly from the object store in order to\nreduce memory pressure, and this series is very much in line with its\nspirit.\n\nThe static version of streaming_write_entry() in entry.c was very specific\nto writing out an index entry out to the working tree, and it made perfect\nsense to have the function in that file, but its interface was limited to\nthe original context the function was used in.\n\nThe whole point of your refactoring in this patch is to make it available\nfor callers outside that original context; e.g. archive that finds blob\nSHA-1 from a tree and writes the blob out to its standard output.  They\nshould not have to work with an API that takes a cache-entry and writes to\na working tree file.  And your result is much more generic.\n\nSo I think the external declaration and the definition should move to a\nmore generic place, namely streaming.[ch].  It does not belong to entry.c\nanymore.\n\nThanks for working on this.\n"},{"id":"185529","messageId":"7vzkc49nnu.fsf@alter.siamese.dyndns.org","threadId":"29753","inReplyTo":"1330329315-11407-4-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH 03/11] cat-file: use streaming interface to print blobs","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-02-27T17:44:21Z","receivedAt":"2012-02-27T17:44:21Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n>  builtin/cat-file.c |   22 ++++++++++++++++++++++\n>  t/t1050-large.sh   |    2 +-\n>  2 files changed, 23 insertions(+), 1 deletions(-)\n>\n> diff --git a/builtin/cat-file.c b/builtin/cat-file.c\n> index 8ed501f..3f3b558 100644\n> --- a/builtin/cat-file.c\n> +++ b/builtin/cat-file.c\n> @@ -82,6 +82,24 @@ static void pprint_tag(const unsigned char *sha1, const char *buf, unsigned long\n>  \t\twrite_or_die(1, cp, endp - cp);\n>  }\n>  \n> +static int write_blob(const unsigned char *sha1)\n> +{\n> +\tunsigned char new_sha1[20];\n> +\n> +\tif (sha1_object_info(sha1, NULL) == OBJ_TAG) {\n\nThis smells bad.  Why in the world could an API be sane if lets a caller\ncall \"write_blob()\" with something that can be a tag?\n\nBoth of your callsites call this function when (type == OBJ_BLOB), but the\n\"case 0:\" arm in the large switch in cat_one_file() only checks \"expected\ntype\" which may not match the real type at all, so it is wrong to switch\non that in the first place.  In addition, that call site alone needs to\nderef tag to the requested/expected type.\n\nThis block does not belong to this function, but to only one of its\ncallers among two.\n\n> +\t\tenum object_type type;\n> +\t\tunsigned long size;\n> +\t\tchar *buffer = read_sha1_file(sha1, &type, &size);\n> +\t\tif (memcmp(buffer, \"object \", 7) ||\n> +\t\t    get_sha1_hex(buffer + 7, new_sha1))\n> +\t\t\tdie(\"%s not a valid tag\", sha1_to_hex(sha1));\n> +\t\tsha1 = new_sha1;\n> +\t\tfree(buffer);\n> +\t}\n> +\n> +\treturn streaming_write_sha1(1, 0, sha1, OBJ_BLOB, NULL);\n\nI do not think your previous refactoring added a fall-back codepath to the\nfunction you are calling here.  In the original context, the caller of\nstreaming_write_entry() made sure that the blob is suitable for streaming\nwrite by getting an istream, and called the function only when that is the\ncase.  Blobs unsuitable for streaming (e.g. an deltified object in a pack)\nwere handled by the caller that decided not to call\nstreaming_write_entry() with the conventional \"read to core and then write\nit out\" codepath.\n\nAnd I do not think your updated caller in cat_one_file() is equipped to do\nso at all.\n\nSo it looks to me that this patch totally breaks the cat-file.  What am I\nmissing?\n\n> +}\n> +\n>  static int cat_one_file(int opt, const char *exp_type, const char *obj_name)\n>  {\n>  \tunsigned char sha1[20];\n> @@ -127,6 +145,8 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name)\n>  \t\t\treturn cmd_ls_tree(2, ls_args, NULL);\n>  \t\t}\n>  \n> +\t\tif (type == OBJ_BLOB)\n> +\t\t\treturn write_blob(sha1);\n>  \t\tbuf = read_sha1_file(sha1, &type, &size);\n>  \t\tif (!buf)\n>  \t\t\tdie(\"Cannot read object %s\", obj_name);\n> @@ -149,6 +169,8 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name)\n>  \t\tbreak;\n>  \n>  \tcase 0:\n> +\t\tif (type_from_string(exp_type) == OBJ_BLOB)\n> +\t\t\treturn write_blob(sha1);\n>  \t\tbuf = read_object_with_reference(sha1, exp_type, &size, NULL);\n>  \t\tbreak;\n>  \n> diff --git a/t/t1050-large.sh b/t/t1050-large.sh\n> index f245e59..39a3e77 100755\n> --- a/t/t1050-large.sh\n> +++ b/t/t1050-large.sh\n> @@ -114,7 +114,7 @@ test_expect_success 'hash-object' '\n>  \tgit hash-object large1\n>  '\n>  \n> -test_expect_failure 'cat-file a large file' '\n> +test_expect_success 'cat-file a large file' '\n>  \tgit cat-file blob :large1 >/dev/null\n>  '\n"},{"id":"185531","messageId":"7vvcms9mw6.fsf@alter.siamese.dyndns.org","threadId":"29753","inReplyTo":"1330329315-11407-6-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH 05/11] show: use streaming interface for showing blobs","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-02-27T18:00:57Z","receivedAt":"2012-02-27T18:00:57Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n>  builtin/log.c    |    9 ++++++++-\n>  t/t1050-large.sh |    2 +-\n>  2 files changed, 9 insertions(+), 2 deletions(-)\n>\n> diff --git a/builtin/log.c b/builtin/log.c\n> index 7d1f6f8..4c4b17a 100644\n> --- a/builtin/log.c\n> +++ b/builtin/log.c\n> @@ -386,13 +386,20 @@ static int show_object(const unsigned char *sha1, int show_tag_object,\n>  {\n>  \tunsigned long size;\n>  \tenum object_type type;\n> -\tchar *buf = read_sha1_file(sha1, &type, &size);\n> +\tchar *buf;\n>  \tint offset = 0;\n>  \n> +\tif (!show_tag_object) {\n> +\t\tfflush(stdout);\n> +\t\treturn streaming_write_sha1(1, 0, sha1, OBJ_ANY, NULL);\n> +\t}\n> +\n> +\tbuf = read_sha1_file(sha1, &type, &size);\n>  \tif (!buf)\n>  \t\treturn error(_(\"Could not read object %s\"), sha1_to_hex(sha1));\n>  \n>  \tif (show_tag_object)\n> +\t\tassert(type == OBJ_TAG);\n>  \t\twhile (offset < size && buf[offset] != '\\n') {\n>  \t\t\tint new_offset = offset + 1;\n>  \t\t\twhile (new_offset < size && buf[new_offset++] != '\\n')\n\nYuck.\n\nThe two callsites to this static function are to do BLOB to do TAG.  And\nafter you start handing all the blob handling to streaming_write_sha1(),\nthere is no shared code between the two callers for this function.\n\nSo why not remove this function, create one show_blob_object() and the\nother show_tag_object(), and update the callers to call the appropriate\none?\n\n> diff --git a/t/t1050-large.sh b/t/t1050-large.sh\n> index 39a3e77..66acb3b 100755\n> --- a/t/t1050-large.sh\n> +++ b/t/t1050-large.sh\n> @@ -118,7 +118,7 @@ test_expect_success 'cat-file a large file' '\n>  \tgit cat-file blob :large1 >/dev/null\n>  '\n>  \n> -test_expect_failure 'git-show a large file' '\n> +test_expect_success 'git-show a large file' '\n>  \tgit show :large1 >/dev/null\n>  \n>  '\n"},{"id":"185540","messageId":"7v7gz89kws.fsf@alter.siamese.dyndns.org","threadId":"29753","inReplyTo":"1330329315-11407-1-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH 00/11] Large blob fixes","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-02-27T18:43:47Z","receivedAt":"2012-02-27T18:43:47Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> These patches make sure we avoid keeping whole blob in memory, at\n> least in common cases. Blob-only streaming code paths are opened to\n> accomplish that.\n\nSome in the series seem to be unrelated to the above, namely, the\nindex-pack ones.\n"},{"id":"185554","messageId":"20120227201805.GA10195@m62s10.vlinux.de","threadId":"29753","inReplyTo":"1330329315-11407-2-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH 01/11] Add more large blob test cases","fromName":"Peter Baumann","fromEmail":"waste.manager@gmx.de","sentAt":"2012-02-27T20:18:05Z","receivedAt":"2012-02-27T20:18:05Z","isPatch":true,"sender":{"key":"waste.manager@gmx.de","avatar":null},"body":"A minor spelling error in the text.\n\nOn Mon, Feb 27, 2012 at 02:55:05PM +0700, Nguyễn Thái Ngọc Duy wrote:\n> New test cases list commands that should work when memory is\n> limited. All memory allocation functions (*) learn to reject any\n> allocation larger than $GIT_ALLOC_LIMIT if set.\n> \n> (*) Not exactly all. Some places do not use x* functions, but\n> malloc/calloc directly, notably diff-delta. These could path should\n                                                    ^code\n> never be run on large blobs.\n> \n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n\n-Peter\n"},{"id":"185570","messageId":"7vaa4454kt.fsf@alter.siamese.dyndns.org","threadId":"29753","inReplyTo":"7v4nucb2xl.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH 02/11] Factor out and export large blob writing code to arbitrary file handle","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-02-27T21:50:10Z","receivedAt":"2012-02-27T21:50:10Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> So I think the external declaration and the definition should move to a\n> more generic place, namely streaming.[ch].  It does not belong to entry.c\n> anymore.\n>\n> Thanks for working on this.\n\nIn other words, I think the result should look more like this.\n\nThe original logic in entry.c is that the caller should try to get a\nfilter and call streaming_write_entry(), but either of them is allowed to\nreturn a failure when the blob is not suitable for the streaming codepath\nto tell the caller to try their traditional codepath.\n\nWe might want to add another helper function for callers to use to decide\nif they should use the streaming interface, or the traditional one, before\nactually making a call to streaming_write_entry().  With the original (and\ncurrent) API, they have to retry even when the streaming codepath truly\nfailed (e.g. no such blob object), in which case it is very likely that\nthe traditional codepath in the caller will fail the same way. Retrying is\na wasted effort in such a case.\n\n-- >8 --\nSubject: [PATCH] streaming: make streaming-write-entry to be more reusable\n\nThe static function in entry.c takes a cache entry and streams its blob\ncontents to a file in the working tree.  Refactor the logic to a new API\nfunction stream_blob_to_fd() that takes an object name and an open file\ndescriptor, so that it can be reused by other callers.\n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n entry.c     |   53 +++++------------------------------------------------\n streaming.c |   55 +++++++++++++++++++++++++++++++++++++++++++++++++++++++\n streaming.h |    2 ++\n 3 files changed, 62 insertions(+), 48 deletions(-)\n\ndiff --git a/entry.c b/entry.c\nindex 852fea1..17a6bcc 100644\n--- a/entry.c\n+++ b/entry.c\n@@ -120,58 +120,15 @@ static int streaming_write_entry(struct cache_entry *ce, char *path,\n \t\t\t\t const struct checkout *state, int to_tempfile,\n \t\t\t\t int *fstat_done, struct stat *statbuf)\n {\n-\tstruct git_istream *st;\n-\tenum object_type type;\n-\tunsigned long sz;\n \tint result = -1;\n-\tssize_t kept = 0;\n-\tint fd = -1;\n-\n-\tst = open_istream(ce->sha1, &type, &sz, filter);\n-\tif (!st)\n-\t\treturn -1;\n-\tif (type != OBJ_BLOB)\n-\t\tgoto close_and_exit;\n+\tint fd;\n \n \tfd = open_output_fd(path, ce, to_tempfile);\n-\tif (fd < 0)\n-\t\tgoto close_and_exit;\n-\n-\tfor (;;) {\n-\t\tchar buf[1024 * 16];\n-\t\tssize_t wrote, holeto;\n-\t\tssize_t readlen = read_istream(st, buf, sizeof(buf));\n-\n-\t\tif (!readlen)\n-\t\t\tbreak;\n-\t\tif (sizeof(buf) == readlen) {\n-\t\t\tfor (holeto = 0; holeto < readlen; holeto++)\n-\t\t\t\tif (buf[holeto])\n-\t\t\t\t\tbreak;\n-\t\t\tif (readlen == holeto) {\n-\t\t\t\tkept += holeto;\n-\t\t\t\tcontinue;\n-\t\t\t}\n-\t\t}\n-\n-\t\tif (kept && lseek(fd, kept, SEEK_CUR) == (off_t) -1)\n-\t\t\tgoto close_and_exit;\n-\t\telse\n-\t\t\tkept = 0;\n-\t\twrote = write_in_full(fd, buf, readlen);\n-\n-\t\tif (wrote != readlen)\n-\t\t\tgoto close_and_exit;\n-\t}\n-\tif (kept && (lseek(fd, kept - 1, SEEK_CUR) == (off_t) -1 ||\n-\t\t     write(fd, \"\", 1) != 1))\n-\t\tgoto close_and_exit;\n-\t*fstat_done = fstat_output(fd, state, statbuf);\n-\n-close_and_exit:\n-\tclose_istream(st);\n-\tif (0 <= fd)\n+\tif (0 <= fd) {\n+\t\tresult = stream_blob_to_fd(fd, ce->sha1, filter, 1);\n+\t\t*fstat_done = fstat_output(fd, state, statbuf);\n \t\tresult = close(fd);\n+\t}\n \tif (result && 0 <= fd)\n \t\tunlink(path);\n \treturn result;\ndiff --git a/streaming.c b/streaming.c\nindex 71072e1..7e7ee2b 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -489,3 +489,58 @@ static open_method_decl(incore)\n \n \treturn st->u.incore.buf ? 0 : -1;\n }\n+\n+\n+/****************************************************************\n+ * Users of streaming interface\n+ ****************************************************************/\n+\n+int stream_blob_to_fd(int fd, unsigned const char *sha1, struct stream_filter *filter,\n+\t\t      int can_seek)\n+{\n+\tstruct git_istream *st;\n+\tenum object_type type;\n+\tunsigned long sz;\n+\tssize_t kept = 0;\n+\tint result = -1;\n+\n+\tst = open_istream(sha1, &type, &sz, filter);\n+\tif (!st)\n+\t\treturn result;\n+\tif (type != OBJ_BLOB)\n+\t\tgoto close_and_exit;\n+\tfor (;;) {\n+\t\tchar buf[1024 * 16];\n+\t\tssize_t wrote, holeto;\n+\t\tssize_t readlen = read_istream(st, buf, sizeof(buf));\n+\n+\t\tif (!readlen)\n+\t\t\tbreak;\n+\t\tif (can_seek && sizeof(buf) == readlen) {\n+\t\t\tfor (holeto = 0; holeto < readlen; holeto++)\n+\t\t\t\tif (buf[holeto])\n+\t\t\t\t\tbreak;\n+\t\t\tif (readlen == holeto) {\n+\t\t\t\tkept += holeto;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t}\n+\n+\t\tif (kept && lseek(fd, kept, SEEK_CUR) == (off_t) -1)\n+\t\t\tgoto close_and_exit;\n+\t\telse\n+\t\t\tkept = 0;\n+\t\twrote = write_in_full(fd, buf, readlen);\n+\n+\t\tif (wrote != readlen)\n+\t\t\tgoto close_and_exit;\n+\t}\n+\tif (kept && (lseek(fd, kept - 1, SEEK_CUR) == (off_t) -1 ||\n+\t\t     write(fd, \"\", 1) != 1))\n+\t\tgoto close_and_exit;\n+\tresult = 0;\n+\n+ close_and_exit:\n+\tclose_istream(st);\n+\treturn result;\n+}\ndiff --git a/streaming.h b/streaming.h\nindex 589e857..3e82770 100644\n--- a/streaming.h\n+++ b/streaming.h\n@@ -12,4 +12,6 @@ extern struct git_istream *open_istream(const unsigned char *, enum object_type\n extern int close_istream(struct git_istream *);\n extern ssize_t read_istream(struct git_istream *, char *, size_t);\n \n+extern int stream_blob_to_fd(int fd, const unsigned char *, struct stream_filter *, int can_seek);\n+\n #endif /* STREAMING_H */\n-- \n1.7.9.2.312.g1abc3\n"},{"id":"185584","messageId":"CACsJy8Aa_KRTYVMy8SgB91D3g_=DwP_noxWAO7NWYuQupKPcyA@mail.gmail.com","threadId":"29753","inReplyTo":"7vzkc49nnu.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH 03/11] cat-file: use streaming interface to print blobs","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-28T01:08:42Z","receivedAt":"2012-02-28T01:08:42Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"2012/2/28 Junio C Hamano <gitster@pobox.com>:\n>> +             enum object_type type;\n>> +             unsigned long size;\n>> +             char *buffer = read_sha1_file(sha1, &type, &size);\n>> +             if (memcmp(buffer, \"object \", 7) ||\n>> +                 get_sha1_hex(buffer + 7, new_sha1))\n>> +                     die(\"%s not a valid tag\", sha1_to_hex(sha1));\n>> +             sha1 = new_sha1;\n>> +             free(buffer);\n>> +     }\n>> +\n>> +     return streaming_write_sha1(1, 0, sha1, OBJ_BLOB, NULL);\n>\n> I do not think your previous refactoring added a fall-back codepath to the\n> function you are calling here.  In the original context, the caller of\n> streaming_write_entry() made sure that the blob is suitable for streaming\n> write by getting an istream, and called the function only when that is the\n> case.  Blobs unsuitable for streaming (e.g. an deltified object in a pack)\n> were handled by the caller that decided not to call\n> streaming_write_entry() with the conventional \"read to core and then write\n> it out\" codepath.\n>\n> And I do not think your updated caller in cat_one_file() is equipped to do\n> so at all.\n>\n> So it looks to me that this patch totally breaks the cat-file.  What am I\n> missing?\n\nI think open_istream can deal with unsuitable for streaming objects\ntoo. There's a fallback \"incore\" backend that does\nread_sha1_file_extended.\n-- \nDuy\n"},{"id":"185585","messageId":"CACsJy8D9Mfs5OvpLsoEqMAoJ3PR9Hka_V+4Mzmc3GpsZtYtm-g@mail.gmail.com","threadId":"29753","inReplyTo":"7v7gz89kws.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH 00/11] Large blob fixes","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-28T01:23:01Z","receivedAt":"2012-02-28T01:23:01Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"2012/2/28 Junio C Hamano <gitster@pobox.com>:\n> Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n>\n>> These patches make sure we avoid keeping whole blob in memory, at\n>> least in common cases. Blob-only streaming code paths are opened to\n>> accomplish that.\n>\n> Some in the series seem to be unrelated to the above, namely, the\n> index-pack ones.\n\nindex-pack patches in this series can make \"index-pack --verify\"\nworse, but it's already not so good. Will take the --verify patch out.\nI will need better strategy than blindly skipping sha-1 collision test\nwhen --verify is specified.\n-- \nDuy\n"},{"id":"186029","messageId":"1330865996-2069-1-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330329315-11407-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 00/10] Large blob fixes","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-04T12:59:46Z","receivedAt":"2012-03-04T12:59:46Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"These patches make sure we avoid keeping whole blob in memory, at\nleast in common cases. Blob-only streaming code paths are opened to\naccomplish that.\n\nThere are a few things I'd like to see addressed, perhaps as part of\nGSoC if any student steps up.\n\n - somehow avoid unpack-objects and keep the pack if it contains large\n   blobs. I guess we could just save the pack, then decide to\n   unpack-objects later. I've updated GSoC ideas page about this.\n \n - pack-objects still puts large blobs in memory if they are in loose\n   format. This should not happen if we fix the above. But if anyone\n   has spare energy, (s)he can try to stream large loose blobs in the\n   pack too. Not sure how ugly the end result could be.\n\n - archive-zip with large blobs. I think two phases are required\n   because we need to calculate crc32 in advance. I have a feeling\n   that we could just stream compressed blobs (either in loose or\n   packed format) to the zip file, i.e. no decompressing then\n   compresssing, which makes two phases nearly as good as one.\n\n - not really large blob related, but it'd be great to see\n   pack-check.c and index-pack.c share as much pack reading code as\n   possible, even bettere if sha1_file.c could join the party.\n\n - I've been thinking whether we could just drop pack-check.c, which\n   is only used by fsck, and make fsck run index-pack instead. The\n   pros is we can run index-pack in parallel. The cons is, how to\n   return marked object list to fsck efficiently.\n\nAnyway changes from v1:\n\n - use stream_blob_to_fd() patch from Junio (better factoring)\n - split show_object() in \"git show\" in two separate functions, one\n   for tag and one for blob, as they do not share much in the end\n - get rid of \"index-pack --verify\" patch. It'll come back separately\n\nJunio C Hamano (1):\n  streaming: make streaming-write-entry to be more reusable\n\nNguyễn Thái Ngọc Duy (9):\n  Add more large blob test cases\n  cat-file: use streaming interface to print blobs\n  parse_object: special code path for blobs to avoid putting whole\n    object in memory\n  show: use streaming interface for showing blobs\n  index-pack: split second pass obj handling into own function\n  index-pack: reduce memory usage when the pack has large blobs\n  pack-check: do not unpack blobs\n  archive: support streaming large files to a tar archive\n  fsck: use streaming interface for writing lost-found blobs\n\n archive-tar.c        |   35 +++++++++++++++----\n archive-zip.c        |    9 +++--\n archive.c            |   51 ++++++++++++++++++---------\n archive.h            |   11 +++++-\n builtin/cat-file.c   |   23 ++++++++++++\n builtin/fsck.c       |    8 +---\n builtin/index-pack.c |   95 ++++++++++++++++++++++++++++++++++++--------------\n builtin/log.c        |   34 ++++++++++-------\n cache.h              |    2 +-\n entry.c              |   53 +++-------------------------\n fast-import.c        |    2 +-\n object.c             |   11 ++++++\n pack-check.c         |   21 ++++++++++-\n sha1_file.c          |   78 +++++++++++++++++++++++++++++++++++------\n streaming.c          |   55 +++++++++++++++++++++++++++++\n streaming.h          |    2 +\n t/t1050-large.sh     |   59 ++++++++++++++++++++++++++++++-\n wrapper.c            |   27 ++++++++++++--\n 18 files changed, 434 insertions(+), 142 deletions(-)\n\n-- \n1.7.8.36.g69ee2\n"},{"id":"186030","messageId":"1330865996-2069-2-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330865996-2069-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 01/10] Add more large blob test cases","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-04T12:59:47Z","receivedAt":"2012-03-04T12:59:47Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"New test cases list commands that should work when memory is\nlimited. All memory allocation functions (*) learn to reject any\nallocation larger than $GIT_ALLOC_LIMIT if set.\n\n(*) Not exactly all. Some places do not use x* functions, but\nmalloc/calloc directly, notably diff-delta. These code path should\nnever be run on large blobs.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n t/t1050-large.sh |   59 +++++++++++++++++++++++++++++++++++++++++++++++++++++-\n wrapper.c        |   27 ++++++++++++++++++++++--\n 2 files changed, 82 insertions(+), 4 deletions(-)\n\ndiff --git a/t/t1050-large.sh b/t/t1050-large.sh\nindex 29d6024..f245e59 100755\n--- a/t/t1050-large.sh\n+++ b/t/t1050-large.sh\n@@ -10,7 +10,9 @@ test_expect_success setup '\n \techo X | dd of=large1 bs=1k seek=2000 &&\n \techo X | dd of=large2 bs=1k seek=2000 &&\n \techo X | dd of=large3 bs=1k seek=2000 &&\n-\techo Y | dd of=huge bs=1k seek=2500\n+\techo Y | dd of=huge bs=1k seek=2500 &&\n+\tGIT_ALLOC_LIMIT=1500 &&\n+\texport GIT_ALLOC_LIMIT\n '\n \n test_expect_success 'add a large file or two' '\n@@ -100,4 +102,59 @@ test_expect_success 'packsize limit' '\n \t)\n '\n \n+test_expect_success 'diff --raw' '\n+\tgit commit -q -m initial &&\n+\techo modified >>large1 &&\n+\tgit add large1 &&\n+\tgit commit -q -m modified &&\n+\tgit diff --raw HEAD^\n+'\n+\n+test_expect_success 'hash-object' '\n+\tgit hash-object large1\n+'\n+\n+test_expect_failure 'cat-file a large file' '\n+\tgit cat-file blob :large1 >/dev/null\n+'\n+\n+test_expect_failure 'git-show a large file' '\n+\tgit show :large1 >/dev/null\n+\n+'\n+\n+test_expect_failure 'clone' '\n+\tgit clone -n file://\"$PWD\"/.git new &&\n+\t(\n+\tcd new &&\n+\tgit config core.bigfilethreshold 200k &&\n+\tgit checkout master\n+\t)\n+'\n+\n+test_expect_failure 'fetch updates' '\n+\techo modified >> large1 &&\n+\tgit commit -q -a -m updated &&\n+\t(\n+\tcd new &&\n+\tgit fetch --keep # FIXME should not need --keep\n+\t)\n+'\n+\n+test_expect_failure 'fsck' '\n+\tgit fsck --full\n+'\n+\n+test_expect_success 'repack' '\n+\tgit repack -ad\n+'\n+\n+test_expect_failure 'tar achiving' '\n+\tgit archive --format=tar HEAD >/dev/null\n+'\n+\n+test_expect_failure 'zip achiving' '\n+\tgit archive --format=zip HEAD >/dev/null\n+'\n+\n test_done\ndiff --git a/wrapper.c b/wrapper.c\nindex 85f09df..d4c0972 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -9,6 +9,18 @@ static void do_nothing(size_t size)\n \n static void (*try_to_free_routine)(size_t size) = do_nothing;\n \n+static void memory_limit_check(size_t size)\n+{\n+\tstatic int limit = -1;\n+\tif (limit == -1) {\n+\t\tconst char *env = getenv(\"GIT_ALLOC_LIMIT\");\n+\t\tlimit = env ? atoi(env) * 1024 : 0;\n+\t}\n+\tif (limit && size > limit)\n+\t\tdie(\"attempting to allocate %d over limit %d\",\n+\t\t    size, limit);\n+}\n+\n try_to_free_t set_try_to_free_routine(try_to_free_t routine)\n {\n \ttry_to_free_t old = try_to_free_routine;\n@@ -32,7 +44,10 @@ char *xstrdup(const char *str)\n \n void *xmalloc(size_t size)\n {\n-\tvoid *ret = malloc(size);\n+\tvoid *ret;\n+\n+\tmemory_limit_check(size);\n+\tret = malloc(size);\n \tif (!ret && !size)\n \t\tret = malloc(1);\n \tif (!ret) {\n@@ -79,7 +94,10 @@ char *xstrndup(const char *str, size_t len)\n \n void *xrealloc(void *ptr, size_t size)\n {\n-\tvoid *ret = realloc(ptr, size);\n+\tvoid *ret;\n+\n+\tmemory_limit_check(size);\n+\tret = realloc(ptr, size);\n \tif (!ret && !size)\n \t\tret = realloc(ptr, 1);\n \tif (!ret) {\n@@ -95,7 +113,10 @@ void *xrealloc(void *ptr, size_t size)\n \n void *xcalloc(size_t nmemb, size_t size)\n {\n-\tvoid *ret = calloc(nmemb, size);\n+\tvoid *ret;\n+\n+\tmemory_limit_check(size * nmemb);\n+\tret = calloc(nmemb, size);\n \tif (!ret && (!nmemb || !size))\n \t\tret = calloc(1, 1);\n \tif (!ret) {\n-- \n1.7.8.36.g69ee2\n"},{"id":"186032","messageId":"1330865996-2069-3-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330865996-2069-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 02/10] streaming: make streaming-write-entry to be more reusable","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-04T12:59:48Z","receivedAt":"2012-03-04T12:59:48Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"From: Junio C Hamano <gitster@pobox.com>\n\nThe static function in entry.c takes a cache entry and streams its blob\ncontents to a file in the working tree.  Refactor the logic to a new API\nfunction stream_blob_to_fd() that takes an object name and an open file\ndescriptor, so that it can be reused by other callers.\n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n entry.c     |   53 +++++------------------------------------------------\n streaming.c |   55 +++++++++++++++++++++++++++++++++++++++++++++++++++++++\n streaming.h |    2 ++\n 3 files changed, 62 insertions(+), 48 deletions(-)\n\ndiff --git a/entry.c b/entry.c\nindex 852fea1..17a6bcc 100644\n--- a/entry.c\n+++ b/entry.c\n@@ -120,58 +120,15 @@ static int streaming_write_entry(struct cache_entry *ce, char *path,\n \t\t\t\t const struct checkout *state, int to_tempfile,\n \t\t\t\t int *fstat_done, struct stat *statbuf)\n {\n-\tstruct git_istream *st;\n-\tenum object_type type;\n-\tunsigned long sz;\n \tint result = -1;\n-\tssize_t kept = 0;\n-\tint fd = -1;\n-\n-\tst = open_istream(ce->sha1, &type, &sz, filter);\n-\tif (!st)\n-\t\treturn -1;\n-\tif (type != OBJ_BLOB)\n-\t\tgoto close_and_exit;\n+\tint fd;\n \n \tfd = open_output_fd(path, ce, to_tempfile);\n-\tif (fd < 0)\n-\t\tgoto close_and_exit;\n-\n-\tfor (;;) {\n-\t\tchar buf[1024 * 16];\n-\t\tssize_t wrote, holeto;\n-\t\tssize_t readlen = read_istream(st, buf, sizeof(buf));\n-\n-\t\tif (!readlen)\n-\t\t\tbreak;\n-\t\tif (sizeof(buf) == readlen) {\n-\t\t\tfor (holeto = 0; holeto < readlen; holeto++)\n-\t\t\t\tif (buf[holeto])\n-\t\t\t\t\tbreak;\n-\t\t\tif (readlen == holeto) {\n-\t\t\t\tkept += holeto;\n-\t\t\t\tcontinue;\n-\t\t\t}\n-\t\t}\n-\n-\t\tif (kept && lseek(fd, kept, SEEK_CUR) == (off_t) -1)\n-\t\t\tgoto close_and_exit;\n-\t\telse\n-\t\t\tkept = 0;\n-\t\twrote = write_in_full(fd, buf, readlen);\n-\n-\t\tif (wrote != readlen)\n-\t\t\tgoto close_and_exit;\n-\t}\n-\tif (kept && (lseek(fd, kept - 1, SEEK_CUR) == (off_t) -1 ||\n-\t\t     write(fd, \"\", 1) != 1))\n-\t\tgoto close_and_exit;\n-\t*fstat_done = fstat_output(fd, state, statbuf);\n-\n-close_and_exit:\n-\tclose_istream(st);\n-\tif (0 <= fd)\n+\tif (0 <= fd) {\n+\t\tresult = stream_blob_to_fd(fd, ce->sha1, filter, 1);\n+\t\t*fstat_done = fstat_output(fd, state, statbuf);\n \t\tresult = close(fd);\n+\t}\n \tif (result && 0 <= fd)\n \t\tunlink(path);\n \treturn result;\ndiff --git a/streaming.c b/streaming.c\nindex 71072e1..7e7ee2b 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -489,3 +489,58 @@ static open_method_decl(incore)\n \n \treturn st->u.incore.buf ? 0 : -1;\n }\n+\n+\n+/****************************************************************\n+ * Users of streaming interface\n+ ****************************************************************/\n+\n+int stream_blob_to_fd(int fd, unsigned const char *sha1, struct stream_filter *filter,\n+\t\t      int can_seek)\n+{\n+\tstruct git_istream *st;\n+\tenum object_type type;\n+\tunsigned long sz;\n+\tssize_t kept = 0;\n+\tint result = -1;\n+\n+\tst = open_istream(sha1, &type, &sz, filter);\n+\tif (!st)\n+\t\treturn result;\n+\tif (type != OBJ_BLOB)\n+\t\tgoto close_and_exit;\n+\tfor (;;) {\n+\t\tchar buf[1024 * 16];\n+\t\tssize_t wrote, holeto;\n+\t\tssize_t readlen = read_istream(st, buf, sizeof(buf));\n+\n+\t\tif (!readlen)\n+\t\t\tbreak;\n+\t\tif (can_seek && sizeof(buf) == readlen) {\n+\t\t\tfor (holeto = 0; holeto < readlen; holeto++)\n+\t\t\t\tif (buf[holeto])\n+\t\t\t\t\tbreak;\n+\t\t\tif (readlen == holeto) {\n+\t\t\t\tkept += holeto;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t}\n+\n+\t\tif (kept && lseek(fd, kept, SEEK_CUR) == (off_t) -1)\n+\t\t\tgoto close_and_exit;\n+\t\telse\n+\t\t\tkept = 0;\n+\t\twrote = write_in_full(fd, buf, readlen);\n+\n+\t\tif (wrote != readlen)\n+\t\t\tgoto close_and_exit;\n+\t}\n+\tif (kept && (lseek(fd, kept - 1, SEEK_CUR) == (off_t) -1 ||\n+\t\t     write(fd, \"\", 1) != 1))\n+\t\tgoto close_and_exit;\n+\tresult = 0;\n+\n+ close_and_exit:\n+\tclose_istream(st);\n+\treturn result;\n+}\ndiff --git a/streaming.h b/streaming.h\nindex 589e857..3e82770 100644\n--- a/streaming.h\n+++ b/streaming.h\n@@ -12,4 +12,6 @@ extern struct git_istream *open_istream(const unsigned char *, enum object_type\n extern int close_istream(struct git_istream *);\n extern ssize_t read_istream(struct git_istream *, char *, size_t);\n \n+extern int stream_blob_to_fd(int fd, const unsigned char *, struct stream_filter *, int can_seek);\n+\n #endif /* STREAMING_H */\n-- \n1.7.8.36.g69ee2\n"},{"id":"186031","messageId":"1330865996-2069-4-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330865996-2069-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 03/10] cat-file: use streaming interface to print blobs","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-04T12:59:49Z","receivedAt":"2012-03-04T12:59:49Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/cat-file.c |   23 +++++++++++++++++++++++\n t/t1050-large.sh   |    2 +-\n 2 files changed, 24 insertions(+), 1 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 8ed501f..bc6cc9f 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -11,6 +11,7 @@\n #include \"parse-options.h\"\n #include \"diff.h\"\n #include \"userdiff.h\"\n+#include \"streaming.h\"\n \n #define BATCH 1\n #define BATCH_CHECK 2\n@@ -82,6 +83,24 @@ static void pprint_tag(const unsigned char *sha1, const char *buf, unsigned long\n \t\twrite_or_die(1, cp, endp - cp);\n }\n \n+static int write_blob(const unsigned char *sha1)\n+{\n+\tunsigned char new_sha1[20];\n+\n+\tif (sha1_object_info(sha1, NULL) == OBJ_TAG) {\n+\t\tenum object_type type;\n+\t\tunsigned long size;\n+\t\tchar *buffer = read_sha1_file(sha1, &type, &size);\n+\t\tif (memcmp(buffer, \"object \", 7) ||\n+\t\t    get_sha1_hex(buffer + 7, new_sha1))\n+\t\t\tdie(\"%s not a valid tag\", sha1_to_hex(sha1));\n+\t\tsha1 = new_sha1;\n+\t\tfree(buffer);\n+\t}\n+\n+\treturn stream_blob_to_fd(1, sha1, NULL, 0);\n+}\n+\n static int cat_one_file(int opt, const char *exp_type, const char *obj_name)\n {\n \tunsigned char sha1[20];\n@@ -127,6 +146,8 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name)\n \t\t\treturn cmd_ls_tree(2, ls_args, NULL);\n \t\t}\n \n+\t\tif (type == OBJ_BLOB)\n+\t\t\treturn write_blob(sha1);\n \t\tbuf = read_sha1_file(sha1, &type, &size);\n \t\tif (!buf)\n \t\t\tdie(\"Cannot read object %s\", obj_name);\n@@ -149,6 +170,8 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name)\n \t\tbreak;\n \n \tcase 0:\n+\t\tif (type_from_string(exp_type) == OBJ_BLOB)\n+\t\t\treturn write_blob(sha1);\n \t\tbuf = read_object_with_reference(sha1, exp_type, &size, NULL);\n \t\tbreak;\n \ndiff --git a/t/t1050-large.sh b/t/t1050-large.sh\nindex f245e59..39a3e77 100755\n--- a/t/t1050-large.sh\n+++ b/t/t1050-large.sh\n@@ -114,7 +114,7 @@ test_expect_success 'hash-object' '\n \tgit hash-object large1\n '\n \n-test_expect_failure 'cat-file a large file' '\n+test_expect_success 'cat-file a large file' '\n \tgit cat-file blob :large1 >/dev/null\n '\n \n-- \n1.7.8.36.g69ee2\n"},{"id":"186034","messageId":"1330865996-2069-5-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330865996-2069-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 04/10] parse_object: special code path for blobs to avoid putting whole object in memory","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-04T12:59:50Z","receivedAt":"2012-03-04T12:59:50Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n object.c    |   11 +++++++++++\n sha1_file.c |   33 ++++++++++++++++++++++++++++++++-\n 2 files changed, 43 insertions(+), 1 deletions(-)\n\ndiff --git a/object.c b/object.c\nindex 6b06297..0498b18 100644\n--- a/object.c\n+++ b/object.c\n@@ -198,6 +198,17 @@ struct object *parse_object(const unsigned char *sha1)\n \tif (obj && obj->parsed)\n \t\treturn obj;\n \n+\tif ((obj && obj->type == OBJ_BLOB) ||\n+\t    (!obj && has_sha1_file(sha1) &&\n+\t     sha1_object_info(sha1, NULL) == OBJ_BLOB)) {\n+\t\tif (check_sha1_signature(repl, NULL, 0, NULL) < 0) {\n+\t\t\terror(\"sha1 mismatch %s\\n\", sha1_to_hex(repl));\n+\t\t\treturn NULL;\n+\t\t}\n+\t\tparse_blob_buffer(lookup_blob(sha1), NULL, 0);\n+\t\treturn lookup_object(sha1);\n+\t}\n+\n \tbuffer = read_sha1_file(sha1, &type, &size);\n \tif (buffer) {\n \t\tif (check_sha1_signature(repl, buffer, size, typename(type)) < 0) {\ndiff --git a/sha1_file.c b/sha1_file.c\nindex f9f8d5e..a77ef0a 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -19,6 +19,7 @@\n #include \"pack-revindex.h\"\n #include \"sha1-lookup.h\"\n #include \"bulk-checkin.h\"\n+#include \"streaming.h\"\n \n #ifndef O_NOATIME\n #if defined(__linux__) && (defined(__i386__) || defined(__PPC__))\n@@ -1149,7 +1150,37 @@ static const struct packed_git *has_packed_and_bad(const unsigned char *sha1)\n int check_sha1_signature(const unsigned char *sha1, void *map, unsigned long size, const char *type)\n {\n \tunsigned char real_sha1[20];\n-\thash_sha1_file(map, size, type, real_sha1);\n+\tenum object_type obj_type;\n+\tstruct git_istream *st;\n+\tgit_SHA_CTX c;\n+\tchar hdr[32];\n+\tint hdrlen;\n+\n+\tif (map) {\n+\t\thash_sha1_file(map, size, type, real_sha1);\n+\t\treturn hashcmp(sha1, real_sha1) ? -1 : 0;\n+\t}\n+\n+\tst = open_istream(sha1, &obj_type, &size, NULL);\n+\tif (!st)\n+\t\treturn -1;\n+\n+\t/* Generate the header */\n+\thdrlen = sprintf(hdr, \"%s %lu\", typename(obj_type), size) + 1;\n+\n+\t/* Sha1.. */\n+\tgit_SHA1_Init(&c);\n+\tgit_SHA1_Update(&c, hdr, hdrlen);\n+\tfor (;;) {\n+\t\tchar buf[1024 * 16];\n+\t\tssize_t readlen = read_istream(st, buf, sizeof(buf));\n+\n+\t\tif (!readlen)\n+\t\t\tbreak;\n+\t\tgit_SHA1_Update(&c, buf, readlen);\n+\t}\n+\tgit_SHA1_Final(real_sha1, &c);\n+\tclose_istream(st);\n \treturn hashcmp(sha1, real_sha1) ? -1 : 0;\n }\n \n-- \n1.7.8.36.g69ee2\n"},{"id":"186035","messageId":"1330865996-2069-6-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330865996-2069-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 05/10] show: use streaming interface for showing blobs","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-04T12:59:51Z","receivedAt":"2012-03-04T12:59:51Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/log.c    |   34 ++++++++++++++++++++--------------\n t/t1050-large.sh |    2 +-\n 2 files changed, 21 insertions(+), 15 deletions(-)\n\ndiff --git a/builtin/log.c b/builtin/log.c\nindex 7d1f6f8..d1702e7 100644\n--- a/builtin/log.c\n+++ b/builtin/log.c\n@@ -20,6 +20,7 @@\n #include \"string-list.h\"\n #include \"parse-options.h\"\n #include \"branch.h\"\n+#include \"streaming.h\"\n \n /* Set a default date-time format for git log (\"log.date\" config variable) */\n static const char *default_date_mode = NULL;\n@@ -381,8 +382,13 @@ static void show_tagger(char *buf, int len, struct rev_info *rev)\n \tstrbuf_release(&out);\n }\n \n-static int show_object(const unsigned char *sha1, int show_tag_object,\n-\tstruct rev_info *rev)\n+static int show_blob_object(const unsigned char *sha1, struct rev_info *rev)\n+{\n+\tfflush(stdout);\n+\treturn stream_blob_to_fd(1, sha1, NULL, 0);\n+}\n+\n+static int show_tag_object(const unsigned char *sha1, struct rev_info *rev)\n {\n \tunsigned long size;\n \tenum object_type type;\n@@ -392,16 +398,16 @@ static int show_object(const unsigned char *sha1, int show_tag_object,\n \tif (!buf)\n \t\treturn error(_(\"Could not read object %s\"), sha1_to_hex(sha1));\n \n-\tif (show_tag_object)\n-\t\twhile (offset < size && buf[offset] != '\\n') {\n-\t\t\tint new_offset = offset + 1;\n-\t\t\twhile (new_offset < size && buf[new_offset++] != '\\n')\n-\t\t\t\t; /* do nothing */\n-\t\t\tif (!prefixcmp(buf + offset, \"tagger \"))\n-\t\t\t\tshow_tagger(buf + offset + 7,\n-\t\t\t\t\t    new_offset - offset - 7, rev);\n-\t\t\toffset = new_offset;\n-\t\t}\n+\tassert(type == OBJ_TAG);\n+\twhile (offset < size && buf[offset] != '\\n') {\n+\t\tint new_offset = offset + 1;\n+\t\twhile (new_offset < size && buf[new_offset++] != '\\n')\n+\t\t\t; /* do nothing */\n+\t\tif (!prefixcmp(buf + offset, \"tagger \"))\n+\t\t\tshow_tagger(buf + offset + 7,\n+\t\t\t\t    new_offset - offset - 7, rev);\n+\t\toffset = new_offset;\n+\t}\n \n \tif (offset < size)\n \t\tfwrite(buf + offset, size - offset, 1, stdout);\n@@ -459,7 +465,7 @@ int cmd_show(int argc, const char **argv, const char *prefix)\n \t\tconst char *name = objects[i].name;\n \t\tswitch (o->type) {\n \t\tcase OBJ_BLOB:\n-\t\t\tret = show_object(o->sha1, 0, NULL);\n+\t\t\tret = show_blob_object(o->sha1, NULL);\n \t\t\tbreak;\n \t\tcase OBJ_TAG: {\n \t\t\tstruct tag *t = (struct tag *)o;\n@@ -470,7 +476,7 @@ int cmd_show(int argc, const char **argv, const char *prefix)\n \t\t\t\t\tdiff_get_color_opt(&rev.diffopt, DIFF_COMMIT),\n \t\t\t\t\tt->tag,\n \t\t\t\t\tdiff_get_color_opt(&rev.diffopt, DIFF_RESET));\n-\t\t\tret = show_object(o->sha1, 1, &rev);\n+\t\t\tret = show_tag_object(o->sha1, &rev);\n \t\t\trev.shown_one = 1;\n \t\t\tif (ret)\n \t\t\t\tbreak;\ndiff --git a/t/t1050-large.sh b/t/t1050-large.sh\nindex 39a3e77..66acb3b 100755\n--- a/t/t1050-large.sh\n+++ b/t/t1050-large.sh\n@@ -118,7 +118,7 @@ test_expect_success 'cat-file a large file' '\n \tgit cat-file blob :large1 >/dev/null\n '\n \n-test_expect_failure 'git-show a large file' '\n+test_expect_success 'git-show a large file' '\n \tgit show :large1 >/dev/null\n \n '\n-- \n1.7.8.36.g69ee2\n"},{"id":"186033","messageId":"1330865996-2069-7-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330865996-2069-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 06/10] index-pack: split second pass obj handling into own function","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-04T12:59:52Z","receivedAt":"2012-03-04T12:59:52Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/index-pack.c |   31 ++++++++++++++++++-------------\n 1 files changed, 18 insertions(+), 13 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex dd1c5c9..918684f 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -682,6 +682,23 @@ static int compare_delta_entry(const void *a, const void *b)\n \t\t\t\t   objects[delta_b->obj_no].type);\n }\n \n+/*\n+ * Second pass:\n+ * - for all non-delta objects, look if it is used as a base for\n+ *   deltas;\n+ * - if used as a base, uncompress the object and apply all deltas,\n+ *   recursively checking if the resulting object is used as a base\n+ *   for some more deltas.\n+ */\n+static void second_pass(struct object_entry *obj)\n+{\n+\tstruct base_data *base_obj = alloc_base_data();\n+\tbase_obj->obj = obj;\n+\tbase_obj->data = NULL;\n+\tfind_unresolved_deltas(base_obj);\n+\tdisplay_progress(progress, nr_resolved_deltas);\n+}\n+\n /* Parse all objects and return the pack content SHA1 hash */\n static void parse_pack_objects(unsigned char *sha1)\n {\n@@ -736,26 +753,14 @@ static void parse_pack_objects(unsigned char *sha1)\n \tqsort(deltas, nr_deltas, sizeof(struct delta_entry),\n \t      compare_delta_entry);\n \n-\t/*\n-\t * Second pass:\n-\t * - for all non-delta objects, look if it is used as a base for\n-\t *   deltas;\n-\t * - if used as a base, uncompress the object and apply all deltas,\n-\t *   recursively checking if the resulting object is used as a base\n-\t *   for some more deltas.\n-\t */\n \tif (verbose)\n \t\tprogress = start_progress(\"Resolving deltas\", nr_deltas);\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tstruct object_entry *obj = &objects[i];\n-\t\tstruct base_data *base_obj = alloc_base_data();\n \n \t\tif (is_delta_type(obj->type))\n \t\t\tcontinue;\n-\t\tbase_obj->obj = obj;\n-\t\tbase_obj->data = NULL;\n-\t\tfind_unresolved_deltas(base_obj);\n-\t\tdisplay_progress(progress, nr_resolved_deltas);\n+\t\tsecond_pass(obj);\n \t}\n }\n \n-- \n1.7.8.36.g69ee2\n"},{"id":"186036","messageId":"1330865996-2069-8-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330865996-2069-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 07/10] index-pack: reduce memory usage when the pack has large blobs","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-04T12:59:53Z","receivedAt":"2012-03-04T12:59:53Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"This command unpacks every non-delta objects in order to:\n\n1. calculate sha-1\n2. do byte-to-byte sha-1 collision test if we happen to have objects\n   with the same sha-1\n3. validate object content in strict mode\n\nAll this requires the entire object to stay in memory, a bad news for\ngiant blobs. This patch lowers memory consumption by not saving the\nobject in memory whenever possible, calculating SHA-1 while unpacking\nthe object.\n\nThis patch assumes that the collision test is rarely needed. The\ncollision test will be done later in second pass if necessary, which\nputs the entire object back to memory again (We could even do the\ncollision test without putting the entire object back in memory, by\ncomparing as we unpack it).\n\nIn strict mode, it always keeps non-blob objects in memory for\nvalidation (blobs do not need data validation).\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/index-pack.c |   64 +++++++++++++++++++++++++++++++++++++++----------\n t/t1050-large.sh     |    4 +-\n 2 files changed, 53 insertions(+), 15 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 918684f..db27133 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -276,30 +276,60 @@ static void unlink_base_data(struct base_data *c)\n \tfree_base_data(c);\n }\n \n-static void *unpack_entry_data(unsigned long offset, unsigned long size)\n+static void *unpack_entry_data(unsigned long offset, unsigned long size,\n+\t\t\t       enum object_type type, unsigned char *sha1)\n {\n+\tstatic char fixed_buf[8192];\n \tint status;\n \tgit_zstream stream;\n-\tvoid *buf = xmalloc(size);\n+\tvoid *buf;\n+\tgit_SHA_CTX c;\n+\n+\tif (sha1) {\t\t/* do hash_sha1_file internally */\n+\t\tchar hdr[32];\n+\t\tint hdrlen = sprintf(hdr, \"%s %lu\", typename(type), size)+1;\n+\t\tgit_SHA1_Init(&c);\n+\t\tgit_SHA1_Update(&c, hdr, hdrlen);\n+\n+\t\tbuf = fixed_buf;\n+\t} else {\n+\t\tbuf = xmalloc(size);\n+\t}\n \n \tmemset(&stream, 0, sizeof(stream));\n \tgit_inflate_init(&stream);\n \tstream.next_out = buf;\n-\tstream.avail_out = size;\n+\tstream.avail_out = buf == fixed_buf ? sizeof(fixed_buf) : size;\n \n \tdo {\n \t\tstream.next_in = fill(1);\n \t\tstream.avail_in = input_len;\n \t\tstatus = git_inflate(&stream, 0);\n \t\tuse(input_len - stream.avail_in);\n+\t\tif (sha1) {\n+\t\t\tgit_SHA1_Update(&c, buf, stream.next_out - (unsigned char *)buf);\n+\t\t\tstream.next_out = buf;\n+\t\t\tstream.avail_out = sizeof(fixed_buf);\n+\t\t}\n \t} while (status == Z_OK);\n \tif (stream.total_out != size || status != Z_STREAM_END)\n \t\tbad_object(offset, \"inflate returned %d\", status);\n \tgit_inflate_end(&stream);\n+\tif (sha1) {\n+\t\tgit_SHA1_Final(sha1, &c);\n+\t\tbuf = NULL;\n+\t}\n \treturn buf;\n }\n \n-static void *unpack_raw_entry(struct object_entry *obj, union delta_base *delta_base)\n+static int is_delta_type(enum object_type type)\n+{\n+\treturn (type == OBJ_REF_DELTA || type == OBJ_OFS_DELTA);\n+}\n+\n+static void *unpack_raw_entry(struct object_entry *obj,\n+\t\t\t      union delta_base *delta_base,\n+\t\t\t      unsigned char *sha1)\n {\n \tunsigned char *p;\n \tunsigned long size, c;\n@@ -359,7 +389,9 @@ static void *unpack_raw_entry(struct object_entry *obj, union delta_base *delta_\n \t}\n \tobj->hdr_size = consumed_bytes - obj->idx.offset;\n \n-\tdata = unpack_entry_data(obj->idx.offset, obj->size);\n+\tif (is_delta_type(obj->type) || strict)\n+\t\tsha1 = NULL;\t/* save unpacked object */\n+\tdata = unpack_entry_data(obj->idx.offset, obj->size, obj->type, sha1);\n \tobj->idx.crc32 = input_crc32;\n \treturn data;\n }\n@@ -460,8 +492,9 @@ static void find_delta_children(const union delta_base *base,\n static void sha1_object(const void *data, unsigned long size,\n \t\t\tenum object_type type, unsigned char *sha1)\n {\n-\thash_sha1_file(data, size, typename(type), sha1);\n-\tif (has_sha1_file(sha1)) {\n+\tif (data)\n+\t\thash_sha1_file(data, size, typename(type), sha1);\n+\tif (data && has_sha1_file(sha1)) {\n \t\tvoid *has_data;\n \t\tenum object_type has_type;\n \t\tunsigned long has_size;\n@@ -510,11 +543,6 @@ static void sha1_object(const void *data, unsigned long size,\n \t}\n }\n \n-static int is_delta_type(enum object_type type)\n-{\n-\treturn (type == OBJ_REF_DELTA || type == OBJ_OFS_DELTA);\n-}\n-\n /*\n  * This function is part of find_unresolved_deltas(). There are two\n  * walkers going in the opposite ways.\n@@ -689,10 +717,20 @@ static int compare_delta_entry(const void *a, const void *b)\n  * - if used as a base, uncompress the object and apply all deltas,\n  *   recursively checking if the resulting object is used as a base\n  *   for some more deltas.\n+ * - if the same object exists in repository and we're not in strict\n+ *   mode, we skipped the sha-1 collision test in the first pass.\n+ *   Do it now.\n  */\n static void second_pass(struct object_entry *obj)\n {\n \tstruct base_data *base_obj = alloc_base_data();\n+\n+\tif (!strict && has_sha1_file(obj->idx.sha1)) {\n+\t\tvoid *data = get_data_from_pack(obj);\n+\t\tsha1_object(data, obj->size, obj->type, obj->idx.sha1);\n+\t\tfree(data);\n+\t}\n+\n \tbase_obj->obj = obj;\n \tbase_obj->data = NULL;\n \tfind_unresolved_deltas(base_obj);\n@@ -718,7 +756,7 @@ static void parse_pack_objects(unsigned char *sha1)\n \t\t\t\tnr_objects);\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tstruct object_entry *obj = &objects[i];\n-\t\tvoid *data = unpack_raw_entry(obj, &delta->base);\n+\t\tvoid *data = unpack_raw_entry(obj, &delta->base, obj->idx.sha1);\n \t\tobj->real_type = obj->type;\n \t\tif (is_delta_type(obj->type)) {\n \t\t\tnr_deltas++;\ndiff --git a/t/t1050-large.sh b/t/t1050-large.sh\nindex 66acb3b..7e78c72 100755\n--- a/t/t1050-large.sh\n+++ b/t/t1050-large.sh\n@@ -123,7 +123,7 @@ test_expect_success 'git-show a large file' '\n \n '\n \n-test_expect_failure 'clone' '\n+test_expect_success 'clone' '\n \tgit clone -n file://\"$PWD\"/.git new &&\n \t(\n \tcd new &&\n@@ -132,7 +132,7 @@ test_expect_failure 'clone' '\n \t)\n '\n \n-test_expect_failure 'fetch updates' '\n+test_expect_success 'fetch updates' '\n \techo modified >> large1 &&\n \tgit commit -q -a -m updated &&\n \t(\n-- \n1.7.8.36.g69ee2\n"},{"id":"186037","messageId":"1330865996-2069-9-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330865996-2069-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 08/10] pack-check: do not unpack blobs","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-04T12:59:54Z","receivedAt":"2012-03-04T12:59:54Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"blob content is not used by verify_pack caller (currently only fsck),\nwe only need to make sure blob sha-1 signature matches its\ncontent. unpack_entry() is taught to hash pack entry as it is\nunpacked, eliminating the need to keep whole blob in memory.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n cache.h          |    2 +-\n fast-import.c    |    2 +-\n pack-check.c     |   21 ++++++++++++++++++++-\n sha1_file.c      |   45 +++++++++++++++++++++++++++++++++++----------\n t/t1050-large.sh |    2 +-\n 5 files changed, 58 insertions(+), 14 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex e12b15f..3365f89 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1062,7 +1062,7 @@ extern const unsigned char *nth_packed_object_sha1(struct packed_git *, uint32_t\n extern off_t nth_packed_object_offset(const struct packed_git *, uint32_t);\n extern off_t find_pack_entry_one(const unsigned char *, struct packed_git *);\n extern int is_pack_valid(struct packed_git *);\n-extern void *unpack_entry(struct packed_git *, off_t, enum object_type *, unsigned long *);\n+extern void *unpack_entry(struct packed_git *, off_t, enum object_type *, unsigned long *, unsigned char *);\n extern unsigned long unpack_object_header_buffer(const unsigned char *buf, unsigned long len, enum object_type *type, unsigned long *sizep);\n extern unsigned long get_size_from_delta(struct packed_git *, struct pack_window **, off_t);\n extern int unpack_object_header(struct packed_git *, struct pack_window **, off_t *, unsigned long *);\ndiff --git a/fast-import.c b/fast-import.c\nindex 6cd19e5..5e94a64 100644\n--- a/fast-import.c\n+++ b/fast-import.c\n@@ -1303,7 +1303,7 @@ static void *gfi_unpack_entry(\n \t\t */\n \t\tp->pack_size = pack_size + 20;\n \t}\n-\treturn unpack_entry(p, oe->idx.offset, &type, sizep);\n+\treturn unpack_entry(p, oe->idx.offset, &type, sizep, NULL);\n }\n \n static const char *get_mode(const char *str, uint16_t *modep)\ndiff --git a/pack-check.c b/pack-check.c\nindex 63a595c..1920bdb 100644\n--- a/pack-check.c\n+++ b/pack-check.c\n@@ -105,6 +105,7 @@ static int verify_packfile(struct packed_git *p,\n \t\tvoid *data;\n \t\tenum object_type type;\n \t\tunsigned long size;\n+\t\toff_t curpos = entries[i].offset;\n \n \t\tif (p->index_version > 1) {\n \t\t\toff_t offset = entries[i].offset;\n@@ -116,7 +117,25 @@ static int verify_packfile(struct packed_git *p,\n \t\t\t\t\t    sha1_to_hex(entries[i].sha1),\n \t\t\t\t\t    p->pack_name, (uintmax_t)offset);\n \t\t}\n-\t\tdata = unpack_entry(p, entries[i].offset, &type, &size);\n+\t\ttype = unpack_object_header(p, w_curs, &curpos, &size);\n+\t\tunuse_pack(w_curs);\n+\t\tif (type == OBJ_BLOB) {\n+\t\t\tunsigned char sha1[20];\n+\t\t\tdata = unpack_entry(p, entries[i].offset, &type, &size, sha1);\n+\t\t\tif (!data) {\n+\t\t\t\tif (hashcmp(entries[i].sha1, sha1))\n+\t\t\t\t\terr = error(\"packed %s from %s is corrupt\",\n+\t\t\t\t\t\t    sha1_to_hex(entries[i].sha1), p->pack_name);\n+\t\t\t\telse if (fn) {\n+\t\t\t\t\tint eaten = 0;\n+\t\t\t\t\tfn(entries[i].sha1, type, size, NULL, &eaten);\n+\t\t\t\t}\n+\t\t\t\tif (((base_count + i) & 1023) == 0)\n+\t\t\t\t\tdisplay_progress(progress, base_count + i);\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t}\n+\t\tdata = unpack_entry(p, entries[i].offset, &type, &size, NULL);\n \t\tif (!data)\n \t\t\terr = error(\"cannot unpack %s from %s at offset %\"PRIuMAX\"\",\n \t\t\t\t    sha1_to_hex(entries[i].sha1), p->pack_name,\ndiff --git a/sha1_file.c b/sha1_file.c\nindex a77ef0a..d68a5b0 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -1653,28 +1653,51 @@ static int packed_object_info(struct packed_git *p, off_t obj_offset,\n }\n \n static void *unpack_compressed_entry(struct packed_git *p,\n-\t\t\t\t    struct pack_window **w_curs,\n-\t\t\t\t    off_t curpos,\n-\t\t\t\t    unsigned long size)\n+\t\t\t\t     struct pack_window **w_curs,\n+\t\t\t\t     off_t curpos,\n+\t\t\t\t     unsigned long size,\n+\t\t\t\t     enum object_type type,\n+\t\t\t\t     unsigned char *sha1)\n {\n+\tstatic unsigned char fixed_buf[8192];\n \tint st;\n \tgit_zstream stream;\n \tunsigned char *buffer, *in;\n+\tgit_SHA_CTX c;\n+\n+\tif (sha1) {\t\t/* do hash_sha1_file internally */\n+\t\tchar hdr[32];\n+\t\tint hdrlen = sprintf(hdr, \"%s %lu\", typename(type), size)+1;\n+\t\tgit_SHA1_Init(&c);\n+\t\tgit_SHA1_Update(&c, hdr, hdrlen);\n+\n+\t\tbuffer = fixed_buf;\n+\t} else {\n+\t\tbuffer = xmallocz(size);\n+\t}\n \n-\tbuffer = xmallocz(size);\n \tmemset(&stream, 0, sizeof(stream));\n \tstream.next_out = buffer;\n-\tstream.avail_out = size + 1;\n+\tstream.avail_out = buffer == fixed_buf ? sizeof(fixed_buf) : size + 1;\n \n \tgit_inflate_init(&stream);\n \tdo {\n \t\tin = use_pack(p, w_curs, curpos, &stream.avail_in);\n \t\tstream.next_in = in;\n \t\tst = git_inflate(&stream, Z_FINISH);\n-\t\tif (!stream.avail_out)\n+\t\tif (sha1) {\n+\t\t\tgit_SHA1_Update(&c, buffer, stream.next_out - (unsigned char *)buffer);\n+\t\t\tstream.next_out = buffer;\n+\t\t\tstream.avail_out = sizeof(fixed_buf);\n+\t\t}\n+\t\telse if (!stream.avail_out)\n \t\t\tbreak; /* the payload is larger than it should be */\n \t\tcurpos += stream.next_in - in;\n \t} while (st == Z_OK || st == Z_BUF_ERROR);\n+\tif (sha1) {\n+\t\tgit_SHA1_Final(sha1, &c);\n+\t\tbuffer = NULL;\n+\t}\n \tgit_inflate_end(&stream);\n \tif ((st != Z_STREAM_END) || stream.total_out != size) {\n \t\tfree(buffer);\n@@ -1727,7 +1750,7 @@ static void *cache_or_unpack_entry(struct packed_git *p, off_t base_offset,\n \n \tret = ent->data;\n \tif (!ret || ent->p != p || ent->base_offset != base_offset)\n-\t\treturn unpack_entry(p, base_offset, type, base_size);\n+\t\treturn unpack_entry(p, base_offset, type, base_size, NULL);\n \n \tif (!keep_cache) {\n \t\tent->data = NULL;\n@@ -1844,7 +1867,7 @@ static void *unpack_delta_entry(struct packed_git *p,\n \t\t\treturn NULL;\n \t}\n \n-\tdelta_data = unpack_compressed_entry(p, w_curs, curpos, delta_size);\n+\tdelta_data = unpack_compressed_entry(p, w_curs, curpos, delta_size, OBJ_NONE, NULL);\n \tif (!delta_data) {\n \t\terror(\"failed to unpack compressed delta \"\n \t\t      \"at offset %\"PRIuMAX\" from %s\",\n@@ -1883,7 +1906,8 @@ static void write_pack_access_log(struct packed_git *p, off_t obj_offset)\n int do_check_packed_object_crc;\n \n void *unpack_entry(struct packed_git *p, off_t obj_offset,\n-\t\t   enum object_type *type, unsigned long *sizep)\n+\t\t   enum object_type *type, unsigned long *sizep,\n+\t\t   unsigned char *sha1)\n {\n \tstruct pack_window *w_curs = NULL;\n \toff_t curpos = obj_offset;\n@@ -1917,7 +1941,8 @@ void *unpack_entry(struct packed_git *p, off_t obj_offset,\n \tcase OBJ_TREE:\n \tcase OBJ_BLOB:\n \tcase OBJ_TAG:\n-\t\tdata = unpack_compressed_entry(p, &w_curs, curpos, *sizep);\n+\t\tdata = unpack_compressed_entry(p, &w_curs, curpos,\n+\t\t\t\t\t       *sizep, *type, sha1);\n \t\tbreak;\n \tdefault:\n \t\tdata = NULL;\ndiff --git a/t/t1050-large.sh b/t/t1050-large.sh\nindex 7e78c72..c749ecb 100755\n--- a/t/t1050-large.sh\n+++ b/t/t1050-large.sh\n@@ -141,7 +141,7 @@ test_expect_success 'fetch updates' '\n \t)\n '\n \n-test_expect_failure 'fsck' '\n+test_expect_success 'fsck' '\n \tgit fsck --full\n '\n \n-- \n1.7.8.36.g69ee2\n"},{"id":"186038","messageId":"1330865996-2069-10-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330865996-2069-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 09/10] archive: support streaming large files to a tar archive","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-04T12:59:55Z","receivedAt":"2012-03-04T12:59:55Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n archive-tar.c    |   35 ++++++++++++++++++++++++++++-------\n archive-zip.c    |    9 +++++----\n archive.c        |   51 ++++++++++++++++++++++++++++++++++-----------------\n archive.h        |   11 +++++++++--\n t/t1050-large.sh |    2 +-\n 5 files changed, 77 insertions(+), 31 deletions(-)\n\ndiff --git a/archive-tar.c b/archive-tar.c\nindex 20af005..5bffe49 100644\n--- a/archive-tar.c\n+++ b/archive-tar.c\n@@ -5,6 +5,7 @@\n #include \"tar.h\"\n #include \"archive.h\"\n #include \"run-command.h\"\n+#include \"streaming.h\"\n \n #define RECORDSIZE\t(512)\n #define BLOCKSIZE\t(RECORDSIZE * 20)\n@@ -123,9 +124,29 @@ static size_t get_path_prefix(const char *path, size_t pathlen, size_t maxlen)\n \treturn i;\n }\n \n+static void write_file(struct git_istream *stream, const void *buffer,\n+\t\t       unsigned long size)\n+{\n+\tif (!stream) {\n+\t\twrite_blocked(buffer, size);\n+\t\treturn;\n+\t}\n+\tfor (;;) {\n+\t\tchar buf[1024 * 16];\n+\t\tssize_t readlen;\n+\n+\t\treadlen = read_istream(stream, buf, sizeof(buf));\n+\n+\t\tif (!readlen)\n+\t\t\tbreak;\n+\t\twrite_blocked(buf, readlen);\n+\t}\n+}\n+\n static int write_tar_entry(struct archiver_args *args,\n-\t\tconst unsigned char *sha1, const char *path, size_t pathlen,\n-\t\tunsigned int mode, void *buffer, unsigned long size)\n+\t\t\t   const unsigned char *sha1, const char *path,\n+\t\t\t   size_t pathlen, unsigned int mode, void *buffer,\n+\t\t\t   struct git_istream *stream, unsigned long size)\n {\n \tstruct ustar_header header;\n \tstruct strbuf ext_header = STRBUF_INIT;\n@@ -200,14 +221,14 @@ static int write_tar_entry(struct archiver_args *args,\n \n \tif (ext_header.len > 0) {\n \t\terr = write_tar_entry(args, sha1, NULL, 0, 0, ext_header.buf,\n-\t\t\t\text_header.len);\n+\t\t\t\t      NULL, ext_header.len);\n \t\tif (err)\n \t\t\treturn err;\n \t}\n \tstrbuf_release(&ext_header);\n \twrite_blocked(&header, sizeof(header));\n-\tif (S_ISREG(mode) && buffer && size > 0)\n-\t\twrite_blocked(buffer, size);\n+\tif (S_ISREG(mode) && size > 0)\n+\t\twrite_file(stream, buffer, size);\n \treturn err;\n }\n \n@@ -219,7 +240,7 @@ static int write_global_extended_header(struct archiver_args *args)\n \n \tstrbuf_append_ext_header(&ext_header, \"comment\", sha1_to_hex(sha1), 40);\n \terr = write_tar_entry(args, NULL, NULL, 0, 0, ext_header.buf,\n-\t\t\text_header.len);\n+\t\t\t      NULL, ext_header.len);\n \tstrbuf_release(&ext_header);\n \treturn err;\n }\n@@ -308,7 +329,7 @@ static int write_tar_archive(const struct archiver *ar,\n \tif (args->commit_sha1)\n \t\terr = write_global_extended_header(args);\n \tif (!err)\n-\t\terr = write_archive_entries(args, write_tar_entry);\n+\t\terr = write_archive_entries(args, write_tar_entry, 1);\n \tif (!err)\n \t\twrite_trailer();\n \treturn err;\ndiff --git a/archive-zip.c b/archive-zip.c\nindex 02d1f37..4a1e917 100644\n--- a/archive-zip.c\n+++ b/archive-zip.c\n@@ -120,9 +120,10 @@ static void *zlib_deflate(void *data, unsigned long size,\n \treturn buffer;\n }\n \n-static int write_zip_entry(struct archiver_args *args,\n-\t\tconst unsigned char *sha1, const char *path, size_t pathlen,\n-\t\tunsigned int mode, void *buffer, unsigned long size)\n+int write_zip_entry(struct archiver_args *args,\n+\t\t\t   const unsigned char *sha1, const char *path,\n+\t\t\t   size_t pathlen, unsigned int mode, void *buffer,\n+\t\t\t   struct git_istream *stream, unsigned long size)\n {\n \tstruct zip_local_header header;\n \tstruct zip_dir_header dirent;\n@@ -271,7 +272,7 @@ static int write_zip_archive(const struct archiver *ar,\n \tzip_dir = xmalloc(ZIP_DIRECTORY_MIN_SIZE);\n \tzip_dir_size = ZIP_DIRECTORY_MIN_SIZE;\n \n-\terr = write_archive_entries(args, write_zip_entry);\n+\terr = write_archive_entries(args, write_zip_entry, 0);\n \tif (!err)\n \t\twrite_zip_trailer(args->commit_sha1);\n \ndiff --git a/archive.c b/archive.c\nindex 1ee837d..257eadf 100644\n--- a/archive.c\n+++ b/archive.c\n@@ -5,6 +5,7 @@\n #include \"archive.h\"\n #include \"parse-options.h\"\n #include \"unpack-trees.h\"\n+#include \"streaming.h\"\n \n static char const * const archive_usage[] = {\n \t\"git archive [options] <tree-ish> [<path>...]\",\n@@ -59,26 +60,35 @@ static void format_subst(const struct commit *commit,\n \tfree(to_free);\n }\n \n-static void *sha1_file_to_archive(const char *path, const unsigned char *sha1,\n-\t\tunsigned int mode, enum object_type *type,\n-\t\tunsigned long *sizep, const struct commit *commit)\n+void sha1_file_to_archive(void **buffer, struct git_istream **stream,\n+\t\t\t  const char *path, const unsigned char *sha1,\n+\t\t\t  unsigned int mode, enum object_type *type,\n+\t\t\t  unsigned long *sizep,\n+\t\t\t  const struct commit *commit)\n {\n-\tvoid *buffer;\n+\tif (stream) {\n+\t\tstruct stream_filter *filter;\n+\t\tfilter = get_stream_filter(path, sha1);\n+\t\tif (!commit && S_ISREG(mode) && is_null_stream_filter(filter)) {\n+\t\t\t*buffer = NULL;\n+\t\t\t*stream = open_istream(sha1, type, sizep, NULL);\n+\t\t\treturn;\n+\t\t}\n+\t\t*stream = NULL;\n+\t}\n \n-\tbuffer = read_sha1_file(sha1, type, sizep);\n-\tif (buffer && S_ISREG(mode)) {\n+\t*buffer = read_sha1_file(sha1, type, sizep);\n+\tif (*buffer && S_ISREG(mode)) {\n \t\tstruct strbuf buf = STRBUF_INIT;\n \t\tsize_t size = 0;\n \n-\t\tstrbuf_attach(&buf, buffer, *sizep, *sizep + 1);\n+\t\tstrbuf_attach(&buf, *buffer, *sizep, *sizep + 1);\n \t\tconvert_to_working_tree(path, buf.buf, buf.len, &buf);\n \t\tif (commit)\n \t\t\tformat_subst(commit, buf.buf, buf.len, &buf);\n-\t\tbuffer = strbuf_detach(&buf, &size);\n+\t\t*buffer = strbuf_detach(&buf, &size);\n \t\t*sizep = size;\n \t}\n-\n-\treturn buffer;\n }\n \n static void setup_archive_check(struct git_attr_check *check)\n@@ -97,6 +107,7 @@ static void setup_archive_check(struct git_attr_check *check)\n struct archiver_context {\n \tstruct archiver_args *args;\n \twrite_archive_entry_fn_t write_entry;\n+\tint stream_ok;\n };\n \n static int write_archive_entry(const unsigned char *sha1, const char *base,\n@@ -109,6 +120,7 @@ static int write_archive_entry(const unsigned char *sha1, const char *base,\n \twrite_archive_entry_fn_t write_entry = c->write_entry;\n \tstruct git_attr_check check[2];\n \tconst char *path_without_prefix;\n+\tstruct git_istream *stream = NULL;\n \tint convert = 0;\n \tint err;\n \tenum object_type type;\n@@ -133,25 +145,29 @@ static int write_archive_entry(const unsigned char *sha1, const char *base,\n \t\tstrbuf_addch(&path, '/');\n \t\tif (args->verbose)\n \t\t\tfprintf(stderr, \"%.*s\\n\", (int)path.len, path.buf);\n-\t\terr = write_entry(args, sha1, path.buf, path.len, mode, NULL, 0);\n+\t\terr = write_entry(args, sha1, path.buf, path.len, mode, NULL, NULL, 0);\n \t\tif (err)\n \t\t\treturn err;\n \t\treturn (S_ISDIR(mode) ? READ_TREE_RECURSIVE : 0);\n \t}\n \n-\tbuffer = sha1_file_to_archive(path_without_prefix, sha1, mode,\n-\t\t\t&type, &size, convert ? args->commit : NULL);\n-\tif (!buffer)\n+\tsha1_file_to_archive(&buffer, c->stream_ok ? &stream : NULL,\n+\t\t\t     path_without_prefix, sha1, mode,\n+\t\t\t     &type, &size, convert ? args->commit : NULL);\n+\tif (!buffer && !stream)\n \t\treturn error(\"cannot read %s\", sha1_to_hex(sha1));\n \tif (args->verbose)\n \t\tfprintf(stderr, \"%.*s\\n\", (int)path.len, path.buf);\n-\terr = write_entry(args, sha1, path.buf, path.len, mode, buffer, size);\n+\terr = write_entry(args, sha1, path.buf, path.len, mode, buffer, stream, size);\n+\tif (stream)\n+\t\tclose_istream(stream);\n \tfree(buffer);\n \treturn err;\n }\n \n int write_archive_entries(struct archiver_args *args,\n-\t\twrite_archive_entry_fn_t write_entry)\n+\t\t\t  write_archive_entry_fn_t write_entry,\n+\t\t\t  int stream_ok)\n {\n \tstruct archiver_context context;\n \tstruct unpack_trees_options opts;\n@@ -167,13 +183,14 @@ int write_archive_entries(struct archiver_args *args,\n \t\tif (args->verbose)\n \t\t\tfprintf(stderr, \"%.*s\\n\", (int)len, args->base);\n \t\terr = write_entry(args, args->tree->object.sha1, args->base,\n-\t\t\t\tlen, 040777, NULL, 0);\n+\t\t\t\t  len, 040777, NULL, NULL, 0);\n \t\tif (err)\n \t\t\treturn err;\n \t}\n \n \tcontext.args = args;\n \tcontext.write_entry = write_entry;\n+\tcontext.stream_ok = stream_ok;\n \n \t/*\n \t * Setup index and instruct attr to read index only\ndiff --git a/archive.h b/archive.h\nindex 2b0884f..370cca9 100644\n--- a/archive.h\n+++ b/archive.h\n@@ -27,9 +27,16 @@ extern void register_archiver(struct archiver *);\n extern void init_tar_archiver(void);\n extern void init_zip_archiver(void);\n \n-typedef int (*write_archive_entry_fn_t)(struct archiver_args *args, const unsigned char *sha1, const char *path, size_t pathlen, unsigned int mode, void *buffer, unsigned long size);\n+struct git_istream;\n+typedef int (*write_archive_entry_fn_t)(struct archiver_args *args,\n+\t\t\t\t\tconst unsigned char *sha1,\n+\t\t\t\t\tconst char *path, size_t pathlen,\n+\t\t\t\t\tunsigned int mode,\n+\t\t\t\t\tvoid *buffer,\n+\t\t\t\t\tstruct git_istream *stream,\n+\t\t\t\t\tunsigned long size);\n \n-extern int write_archive_entries(struct archiver_args *args, write_archive_entry_fn_t write_entry);\n+extern int write_archive_entries(struct archiver_args *args, write_archive_entry_fn_t write_entry, int stream_ok);\n extern int write_archive(int argc, const char **argv, const char *prefix, int setup_prefix, const char *name_hint, int remote);\n \n const char *archive_format_from_filename(const char *filename);\ndiff --git a/t/t1050-large.sh b/t/t1050-large.sh\nindex c749ecb..1e64692 100755\n--- a/t/t1050-large.sh\n+++ b/t/t1050-large.sh\n@@ -149,7 +149,7 @@ test_expect_success 'repack' '\n \tgit repack -ad\n '\n \n-test_expect_failure 'tar achiving' '\n+test_expect_success 'tar achiving' '\n \tgit archive --format=tar HEAD >/dev/null\n '\n \n-- \n1.7.8.36.g69ee2\n"},{"id":"186039","messageId":"1330865996-2069-11-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330865996-2069-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v2 10/10] fsck: use streaming interface for writing lost-found blobs","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-04T12:59:56Z","receivedAt":"2012-03-04T12:59:56Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/fsck.c |    8 ++------\n 1 files changed, 2 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/fsck.c b/builtin/fsck.c\nindex 8c479a7..7fcb33e 100644\n--- a/builtin/fsck.c\n+++ b/builtin/fsck.c\n@@ -12,6 +12,7 @@\n #include \"parse-options.h\"\n #include \"dir.h\"\n #include \"progress.h\"\n+#include \"streaming.h\"\n \n #define REACHABLE 0x0001\n #define SEEN      0x0002\n@@ -236,13 +237,8 @@ static void check_unreachable_object(struct object *obj)\n \t\t\tif (!(f = fopen(filename, \"w\")))\n \t\t\t\tdie_errno(\"Could not open '%s'\", filename);\n \t\t\tif (obj->type == OBJ_BLOB) {\n-\t\t\t\tenum object_type type;\n-\t\t\t\tunsigned long size;\n-\t\t\t\tchar *buf = read_sha1_file(obj->sha1,\n-\t\t\t\t\t\t&type, &size);\n-\t\t\t\tif (buf && fwrite(buf, 1, size, f) != size)\n+\t\t\t\tif (stream_blob_to_fd(fileno(f), obj->sha1, NULL, 1))\n \t\t\t\t\tdie_errno(\"Could not write '%s'\", filename);\n-\t\t\t\tfree(buf);\n \t\t\t} else\n \t\t\t\tfprintf(f, \"%s\\n\", sha1_to_hex(obj->sha1));\n \t\t\tif (fclose(f))\n-- \n1.7.8.36.g69ee2\n"},{"id":"186056","messageId":"7vling6jw2.fsf@alter.siamese.dyndns.org","threadId":"29753","inReplyTo":"1330865996-2069-4-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH v2 03/10] cat-file: use streaming interface to print blobs","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-03-04T23:12:13Z","receivedAt":"2012-03-04T23:12:13Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> +static int write_blob(const unsigned char *sha1)\n> +{\n> +\tunsigned char new_sha1[20];\n> +\n> +\tif (sha1_object_info(sha1, NULL) == OBJ_TAG) {\n\nHrm, didn't I say that it tastes bad for a function write_blob() to have\nto worry about OBJ_TAG already?\n"},{"id":"186065","messageId":"CACsJy8BFW_LhsoL_ickLopnnMBj10rn-9QcoqmXML8Uau74S5A@mail.gmail.com","threadId":"29753","inReplyTo":"7vling6jw2.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH v2 03/10] cat-file: use streaming interface to print blobs","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-05T02:42:16Z","receivedAt":"2012-03-05T02:42:16Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"2012/3/5 Junio C Hamano <gitster@pobox.com>:\n> Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n>\n>> +static int write_blob(const unsigned char *sha1)\n>> +{\n>> +     unsigned char new_sha1[20];\n>> +\n>> +     if (sha1_object_info(sha1, NULL) == OBJ_TAG) {\n>\n> Hrm, didn't I say that it tastes bad for a function write_blob() to have\n> to worry about OBJ_TAG already?\n\nMy bad. Reworked, added another test case for the dereference case,\nand clone exceeded memory limit again due to new test case :( Will\nneed some more work on this.\n-- \nDuy\n"},{"id":"186069","messageId":"1330919028-6611-1-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330865996-2069-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v3 00/11] Large blob fixes","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-05T03:43:37Z","receivedAt":"2012-03-05T03:43:37Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Changes from v2:\n\n - set core.bigfilethreshold globally in t1050 to make git-clone happy\n   because there's currently no way to specify this in git-clone (or\n   is there?)\n - fix the bad coding taste in builtin/cat-file.c\n - make update-server-info respect core.bigfilethreshold,\n   which makes repack pass on repositories that have tags\n\nJunio C Hamano (1):\n  streaming: make streaming-write-entry to be more reusable\n\nNguyễn Thái Ngọc Duy (10):\n  Add more large blob test cases\n  cat-file: use streaming interface to print blobs\n  parse_object: special code path for blobs to avoid putting whole\n    object in memory\n  show: use streaming interface for showing blobs\n  index-pack: split second pass obj handling into own function\n  index-pack: reduce memory usage when the pack has large blobs\n  pack-check: do not unpack blobs\n  archive: support streaming large files to a tar archive\n  fsck: use streaming interface for writing lost-found blobs\n  update-server-info: respect core.bigfilethreshold\n\n archive-tar.c                |   35 ++++++++++++---\n archive-zip.c                |    9 ++--\n archive.c                    |   51 +++++++++++++++-------\n archive.h                    |   11 ++++-\n builtin/cat-file.c           |   24 +++++++++++\n builtin/fsck.c               |    8 +---\n builtin/index-pack.c         |   95 ++++++++++++++++++++++++++++++-----------\n builtin/log.c                |   34 +++++++++------\n builtin/update-server-info.c |    1 +\n cache.h                      |    2 +-\n entry.c                      |   53 ++---------------------\n fast-import.c                |    2 +-\n object.c                     |   11 +++++\n pack-check.c                 |   21 +++++++++-\n sha1_file.c                  |   78 +++++++++++++++++++++++++++++-----\n streaming.c                  |   55 ++++++++++++++++++++++++\n streaming.h                  |    2 +\n t/t1050-large.sh             |   63 +++++++++++++++++++++++++++-\n wrapper.c                    |   27 +++++++++++-\n 19 files changed, 439 insertions(+), 143 deletions(-)\n\n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"186070","messageId":"1330919028-6611-2-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330865996-2069-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v3 01/11] Add more large blob test cases","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-05T03:43:38Z","receivedAt":"2012-03-05T03:43:38Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"New test cases list commands that should work when memory is\nlimited. All memory allocation functions (*) learn to reject any\nallocation larger than $GIT_ALLOC_LIMIT if set.\n\n(*) Not exactly all. Some places do not use x* functions, but\nmalloc/calloc directly, notably diff-delta. These code path should\nnever be run on large blobs.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n t/t1050-large.sh |   63 ++++++++++++++++++++++++++++++++++++++++++++++++++++-\n wrapper.c        |   27 ++++++++++++++++++++--\n 2 files changed, 85 insertions(+), 5 deletions(-)\n\ndiff --git a/t/t1050-large.sh b/t/t1050-large.sh\nindex 29d6024..80f157a 100755\n--- a/t/t1050-large.sh\n+++ b/t/t1050-large.sh\n@@ -6,11 +6,15 @@ test_description='adding and checking out large blobs'\n . ./test-lib.sh\n \n test_expect_success setup '\n-\tgit config core.bigfilethreshold 200k &&\n+\t# clone does not allow us to pass core.bigfilethreshold to\n+\t# new repos, so set core.bigfilethreshold globally\n+\tgit config --global core.bigfilethreshold 200k &&\n \techo X | dd of=large1 bs=1k seek=2000 &&\n \techo X | dd of=large2 bs=1k seek=2000 &&\n \techo X | dd of=large3 bs=1k seek=2000 &&\n-\techo Y | dd of=huge bs=1k seek=2500\n+\techo Y | dd of=huge bs=1k seek=2500 &&\n+\tGIT_ALLOC_LIMIT=1500 &&\n+\texport GIT_ALLOC_LIMIT\n '\n \n test_expect_success 'add a large file or two' '\n@@ -100,4 +104,59 @@ test_expect_success 'packsize limit' '\n \t)\n '\n \n+test_expect_success 'diff --raw' '\n+\tgit commit -q -m initial &&\n+\techo modified >>large1 &&\n+\tgit add large1 &&\n+\tgit commit -q -m modified &&\n+\tgit diff --raw HEAD^\n+'\n+\n+test_expect_success 'hash-object' '\n+\tgit hash-object large1\n+'\n+\n+test_expect_failure 'cat-file a large file' '\n+\tgit cat-file blob :large1 >/dev/null\n+'\n+\n+test_expect_failure 'cat-file a large file from a tag' '\n+\tgit tag -m largefile largefiletag :large1 &&\n+\tgit cat-file blob largefiletag >/dev/null\n+'\n+\n+test_expect_failure 'git-show a large file' '\n+\tgit show :large1 >/dev/null\n+\n+'\n+\n+test_expect_failure 'clone' '\n+\tgit clone file://\"$PWD\"/.git new\n+'\n+\n+test_expect_failure 'fetch updates' '\n+\techo modified >> large1 &&\n+\tgit commit -q -a -m updated &&\n+\t(\n+\tcd new &&\n+\tgit fetch --keep # FIXME should not need --keep\n+\t)\n+'\n+\n+test_expect_failure 'fsck' '\n+\tgit fsck --full\n+'\n+\n+test_expect_failure 'repack' '\n+\tgit repack -ad\n+'\n+\n+test_expect_failure 'tar achiving' '\n+\tgit archive --format=tar HEAD >/dev/null\n+'\n+\n+test_expect_failure 'zip achiving' '\n+\tgit archive --format=zip HEAD >/dev/null\n+'\n+\n test_done\ndiff --git a/wrapper.c b/wrapper.c\nindex 85f09df..d4c0972 100644\n--- a/wrapper.c\n+++ b/wrapper.c\n@@ -9,6 +9,18 @@ static void do_nothing(size_t size)\n \n static void (*try_to_free_routine)(size_t size) = do_nothing;\n \n+static void memory_limit_check(size_t size)\n+{\n+\tstatic int limit = -1;\n+\tif (limit == -1) {\n+\t\tconst char *env = getenv(\"GIT_ALLOC_LIMIT\");\n+\t\tlimit = env ? atoi(env) * 1024 : 0;\n+\t}\n+\tif (limit && size > limit)\n+\t\tdie(\"attempting to allocate %d over limit %d\",\n+\t\t    size, limit);\n+}\n+\n try_to_free_t set_try_to_free_routine(try_to_free_t routine)\n {\n \ttry_to_free_t old = try_to_free_routine;\n@@ -32,7 +44,10 @@ char *xstrdup(const char *str)\n \n void *xmalloc(size_t size)\n {\n-\tvoid *ret = malloc(size);\n+\tvoid *ret;\n+\n+\tmemory_limit_check(size);\n+\tret = malloc(size);\n \tif (!ret && !size)\n \t\tret = malloc(1);\n \tif (!ret) {\n@@ -79,7 +94,10 @@ char *xstrndup(const char *str, size_t len)\n \n void *xrealloc(void *ptr, size_t size)\n {\n-\tvoid *ret = realloc(ptr, size);\n+\tvoid *ret;\n+\n+\tmemory_limit_check(size);\n+\tret = realloc(ptr, size);\n \tif (!ret && !size)\n \t\tret = realloc(ptr, 1);\n \tif (!ret) {\n@@ -95,7 +113,10 @@ void *xrealloc(void *ptr, size_t size)\n \n void *xcalloc(size_t nmemb, size_t size)\n {\n-\tvoid *ret = calloc(nmemb, size);\n+\tvoid *ret;\n+\n+\tmemory_limit_check(size * nmemb);\n+\tret = calloc(nmemb, size);\n \tif (!ret && (!nmemb || !size))\n \t\tret = calloc(1, 1);\n \tif (!ret) {\n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"186071","messageId":"1330919028-6611-3-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330865996-2069-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v3 02/11] streaming: make streaming-write-entry to be more reusable","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-05T03:43:39Z","receivedAt":"2012-03-05T03:43:39Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"From: Junio C Hamano <gitster@pobox.com>\n\nThe static function in entry.c takes a cache entry and streams its blob\ncontents to a file in the working tree.  Refactor the logic to a new API\nfunction stream_blob_to_fd() that takes an object name and an open file\ndescriptor, so that it can be reused by other callers.\n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n entry.c     |   53 +++++------------------------------------------------\n streaming.c |   55 +++++++++++++++++++++++++++++++++++++++++++++++++++++++\n streaming.h |    2 ++\n 3 files changed, 62 insertions(+), 48 deletions(-)\n\ndiff --git a/entry.c b/entry.c\nindex 852fea1..17a6bcc 100644\n--- a/entry.c\n+++ b/entry.c\n@@ -120,58 +120,15 @@ static int streaming_write_entry(struct cache_entry *ce, char *path,\n \t\t\t\t const struct checkout *state, int to_tempfile,\n \t\t\t\t int *fstat_done, struct stat *statbuf)\n {\n-\tstruct git_istream *st;\n-\tenum object_type type;\n-\tunsigned long sz;\n \tint result = -1;\n-\tssize_t kept = 0;\n-\tint fd = -1;\n-\n-\tst = open_istream(ce->sha1, &type, &sz, filter);\n-\tif (!st)\n-\t\treturn -1;\n-\tif (type != OBJ_BLOB)\n-\t\tgoto close_and_exit;\n+\tint fd;\n \n \tfd = open_output_fd(path, ce, to_tempfile);\n-\tif (fd < 0)\n-\t\tgoto close_and_exit;\n-\n-\tfor (;;) {\n-\t\tchar buf[1024 * 16];\n-\t\tssize_t wrote, holeto;\n-\t\tssize_t readlen = read_istream(st, buf, sizeof(buf));\n-\n-\t\tif (!readlen)\n-\t\t\tbreak;\n-\t\tif (sizeof(buf) == readlen) {\n-\t\t\tfor (holeto = 0; holeto < readlen; holeto++)\n-\t\t\t\tif (buf[holeto])\n-\t\t\t\t\tbreak;\n-\t\t\tif (readlen == holeto) {\n-\t\t\t\tkept += holeto;\n-\t\t\t\tcontinue;\n-\t\t\t}\n-\t\t}\n-\n-\t\tif (kept && lseek(fd, kept, SEEK_CUR) == (off_t) -1)\n-\t\t\tgoto close_and_exit;\n-\t\telse\n-\t\t\tkept = 0;\n-\t\twrote = write_in_full(fd, buf, readlen);\n-\n-\t\tif (wrote != readlen)\n-\t\t\tgoto close_and_exit;\n-\t}\n-\tif (kept && (lseek(fd, kept - 1, SEEK_CUR) == (off_t) -1 ||\n-\t\t     write(fd, \"\", 1) != 1))\n-\t\tgoto close_and_exit;\n-\t*fstat_done = fstat_output(fd, state, statbuf);\n-\n-close_and_exit:\n-\tclose_istream(st);\n-\tif (0 <= fd)\n+\tif (0 <= fd) {\n+\t\tresult = stream_blob_to_fd(fd, ce->sha1, filter, 1);\n+\t\t*fstat_done = fstat_output(fd, state, statbuf);\n \t\tresult = close(fd);\n+\t}\n \tif (result && 0 <= fd)\n \t\tunlink(path);\n \treturn result;\ndiff --git a/streaming.c b/streaming.c\nindex 71072e1..7e7ee2b 100644\n--- a/streaming.c\n+++ b/streaming.c\n@@ -489,3 +489,58 @@ static open_method_decl(incore)\n \n \treturn st->u.incore.buf ? 0 : -1;\n }\n+\n+\n+/****************************************************************\n+ * Users of streaming interface\n+ ****************************************************************/\n+\n+int stream_blob_to_fd(int fd, unsigned const char *sha1, struct stream_filter *filter,\n+\t\t      int can_seek)\n+{\n+\tstruct git_istream *st;\n+\tenum object_type type;\n+\tunsigned long sz;\n+\tssize_t kept = 0;\n+\tint result = -1;\n+\n+\tst = open_istream(sha1, &type, &sz, filter);\n+\tif (!st)\n+\t\treturn result;\n+\tif (type != OBJ_BLOB)\n+\t\tgoto close_and_exit;\n+\tfor (;;) {\n+\t\tchar buf[1024 * 16];\n+\t\tssize_t wrote, holeto;\n+\t\tssize_t readlen = read_istream(st, buf, sizeof(buf));\n+\n+\t\tif (!readlen)\n+\t\t\tbreak;\n+\t\tif (can_seek && sizeof(buf) == readlen) {\n+\t\t\tfor (holeto = 0; holeto < readlen; holeto++)\n+\t\t\t\tif (buf[holeto])\n+\t\t\t\t\tbreak;\n+\t\t\tif (readlen == holeto) {\n+\t\t\t\tkept += holeto;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t}\n+\n+\t\tif (kept && lseek(fd, kept, SEEK_CUR) == (off_t) -1)\n+\t\t\tgoto close_and_exit;\n+\t\telse\n+\t\t\tkept = 0;\n+\t\twrote = write_in_full(fd, buf, readlen);\n+\n+\t\tif (wrote != readlen)\n+\t\t\tgoto close_and_exit;\n+\t}\n+\tif (kept && (lseek(fd, kept - 1, SEEK_CUR) == (off_t) -1 ||\n+\t\t     write(fd, \"\", 1) != 1))\n+\t\tgoto close_and_exit;\n+\tresult = 0;\n+\n+ close_and_exit:\n+\tclose_istream(st);\n+\treturn result;\n+}\ndiff --git a/streaming.h b/streaming.h\nindex 589e857..3e82770 100644\n--- a/streaming.h\n+++ b/streaming.h\n@@ -12,4 +12,6 @@ extern struct git_istream *open_istream(const unsigned char *, enum object_type\n extern int close_istream(struct git_istream *);\n extern ssize_t read_istream(struct git_istream *, char *, size_t);\n \n+extern int stream_blob_to_fd(int fd, const unsigned char *, struct stream_filter *, int can_seek);\n+\n #endif /* STREAMING_H */\n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"186072","messageId":"1330919028-6611-4-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330865996-2069-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v3 03/11] cat-file: use streaming interface to print blobs","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-05T03:43:40Z","receivedAt":"2012-03-05T03:43:40Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/cat-file.c |   24 ++++++++++++++++++++++++\n t/t1050-large.sh   |    4 ++--\n 2 files changed, 26 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 8ed501f..ce68a20 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -11,6 +11,7 @@\n #include \"parse-options.h\"\n #include \"diff.h\"\n #include \"userdiff.h\"\n+#include \"streaming.h\"\n \n #define BATCH 1\n #define BATCH_CHECK 2\n@@ -127,6 +128,8 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name)\n \t\t\treturn cmd_ls_tree(2, ls_args, NULL);\n \t\t}\n \n+\t\tif (type == OBJ_BLOB)\n+\t\t\treturn stream_blob_to_fd(1, sha1, NULL, 0);\n \t\tbuf = read_sha1_file(sha1, &type, &size);\n \t\tif (!buf)\n \t\t\tdie(\"Cannot read object %s\", obj_name);\n@@ -149,6 +152,27 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name)\n \t\tbreak;\n \n \tcase 0:\n+\t\tif (type_from_string(exp_type) == OBJ_BLOB) {\n+\t\t\tunsigned char blob_sha1[20];\n+\t\t\tif (sha1_object_info(sha1, NULL) == OBJ_TAG) {\n+\t\t\t\tenum object_type type;\n+\t\t\t\tunsigned long size;\n+\t\t\t\tchar *buffer = read_sha1_file(sha1, &type, &size);\n+\t\t\t\tif (memcmp(buffer, \"object \", 7) ||\n+\t\t\t\t    get_sha1_hex(buffer + 7, blob_sha1))\n+\t\t\t\t\tdie(\"%s not a valid tag\", sha1_to_hex(sha1));\n+\t\t\t\tfree(buffer);\n+\t\t\t} else\n+\t\t\t\thashcpy(blob_sha1, sha1);\n+\n+\t\t\tif (sha1_object_info(blob_sha1, NULL) == OBJ_BLOB)\n+\t\t\t\treturn stream_blob_to_fd(1, blob_sha1, NULL, 0);\n+\t\t\t/* we attempted to dereference a tag to a blob\n+\t\t\t   and failed, perhaps there are new dereference\n+\t\t\t   mechanisms this code is not aware of,\n+\t\t\t   fallthrough and let read_object_with_reference\n+\t\t\t   deal with it */\n+\t\t}\n \t\tbuf = read_object_with_reference(sha1, exp_type, &size, NULL);\n \t\tbreak;\n \ndiff --git a/t/t1050-large.sh b/t/t1050-large.sh\nindex 80f157a..97ad5b3 100755\n--- a/t/t1050-large.sh\n+++ b/t/t1050-large.sh\n@@ -116,11 +116,11 @@ test_expect_success 'hash-object' '\n \tgit hash-object large1\n '\n \n-test_expect_failure 'cat-file a large file' '\n+test_expect_success 'cat-file a large file' '\n \tgit cat-file blob :large1 >/dev/null\n '\n \n-test_expect_failure 'cat-file a large file from a tag' '\n+test_expect_success 'cat-file a large file from a tag' '\n \tgit tag -m largefile largefiletag :large1 &&\n \tgit cat-file blob largefiletag >/dev/null\n '\n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"186074","messageId":"1330919028-6611-5-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330865996-2069-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v3 04/11] parse_object: special code path for blobs to avoid putting whole object in memory","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-05T03:43:41Z","receivedAt":"2012-03-05T03:43:41Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n object.c    |   11 +++++++++++\n sha1_file.c |   33 ++++++++++++++++++++++++++++++++-\n 2 files changed, 43 insertions(+), 1 deletions(-)\n\ndiff --git a/object.c b/object.c\nindex 6b06297..0498b18 100644\n--- a/object.c\n+++ b/object.c\n@@ -198,6 +198,17 @@ struct object *parse_object(const unsigned char *sha1)\n \tif (obj && obj->parsed)\n \t\treturn obj;\n \n+\tif ((obj && obj->type == OBJ_BLOB) ||\n+\t    (!obj && has_sha1_file(sha1) &&\n+\t     sha1_object_info(sha1, NULL) == OBJ_BLOB)) {\n+\t\tif (check_sha1_signature(repl, NULL, 0, NULL) < 0) {\n+\t\t\terror(\"sha1 mismatch %s\\n\", sha1_to_hex(repl));\n+\t\t\treturn NULL;\n+\t\t}\n+\t\tparse_blob_buffer(lookup_blob(sha1), NULL, 0);\n+\t\treturn lookup_object(sha1);\n+\t}\n+\n \tbuffer = read_sha1_file(sha1, &type, &size);\n \tif (buffer) {\n \t\tif (check_sha1_signature(repl, buffer, size, typename(type)) < 0) {\ndiff --git a/sha1_file.c b/sha1_file.c\nindex f9f8d5e..a77ef0a 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -19,6 +19,7 @@\n #include \"pack-revindex.h\"\n #include \"sha1-lookup.h\"\n #include \"bulk-checkin.h\"\n+#include \"streaming.h\"\n \n #ifndef O_NOATIME\n #if defined(__linux__) && (defined(__i386__) || defined(__PPC__))\n@@ -1149,7 +1150,37 @@ static const struct packed_git *has_packed_and_bad(const unsigned char *sha1)\n int check_sha1_signature(const unsigned char *sha1, void *map, unsigned long size, const char *type)\n {\n \tunsigned char real_sha1[20];\n-\thash_sha1_file(map, size, type, real_sha1);\n+\tenum object_type obj_type;\n+\tstruct git_istream *st;\n+\tgit_SHA_CTX c;\n+\tchar hdr[32];\n+\tint hdrlen;\n+\n+\tif (map) {\n+\t\thash_sha1_file(map, size, type, real_sha1);\n+\t\treturn hashcmp(sha1, real_sha1) ? -1 : 0;\n+\t}\n+\n+\tst = open_istream(sha1, &obj_type, &size, NULL);\n+\tif (!st)\n+\t\treturn -1;\n+\n+\t/* Generate the header */\n+\thdrlen = sprintf(hdr, \"%s %lu\", typename(obj_type), size) + 1;\n+\n+\t/* Sha1.. */\n+\tgit_SHA1_Init(&c);\n+\tgit_SHA1_Update(&c, hdr, hdrlen);\n+\tfor (;;) {\n+\t\tchar buf[1024 * 16];\n+\t\tssize_t readlen = read_istream(st, buf, sizeof(buf));\n+\n+\t\tif (!readlen)\n+\t\t\tbreak;\n+\t\tgit_SHA1_Update(&c, buf, readlen);\n+\t}\n+\tgit_SHA1_Final(real_sha1, &c);\n+\tclose_istream(st);\n \treturn hashcmp(sha1, real_sha1) ? -1 : 0;\n }\n \n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"186073","messageId":"1330919028-6611-6-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330865996-2069-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v3 05/11] show: use streaming interface for showing blobs","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-05T03:43:42Z","receivedAt":"2012-03-05T03:43:42Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/log.c    |   34 ++++++++++++++++++++--------------\n t/t1050-large.sh |    2 +-\n 2 files changed, 21 insertions(+), 15 deletions(-)\n\ndiff --git a/builtin/log.c b/builtin/log.c\nindex 7d1f6f8..d1702e7 100644\n--- a/builtin/log.c\n+++ b/builtin/log.c\n@@ -20,6 +20,7 @@\n #include \"string-list.h\"\n #include \"parse-options.h\"\n #include \"branch.h\"\n+#include \"streaming.h\"\n \n /* Set a default date-time format for git log (\"log.date\" config variable) */\n static const char *default_date_mode = NULL;\n@@ -381,8 +382,13 @@ static void show_tagger(char *buf, int len, struct rev_info *rev)\n \tstrbuf_release(&out);\n }\n \n-static int show_object(const unsigned char *sha1, int show_tag_object,\n-\tstruct rev_info *rev)\n+static int show_blob_object(const unsigned char *sha1, struct rev_info *rev)\n+{\n+\tfflush(stdout);\n+\treturn stream_blob_to_fd(1, sha1, NULL, 0);\n+}\n+\n+static int show_tag_object(const unsigned char *sha1, struct rev_info *rev)\n {\n \tunsigned long size;\n \tenum object_type type;\n@@ -392,16 +398,16 @@ static int show_object(const unsigned char *sha1, int show_tag_object,\n \tif (!buf)\n \t\treturn error(_(\"Could not read object %s\"), sha1_to_hex(sha1));\n \n-\tif (show_tag_object)\n-\t\twhile (offset < size && buf[offset] != '\\n') {\n-\t\t\tint new_offset = offset + 1;\n-\t\t\twhile (new_offset < size && buf[new_offset++] != '\\n')\n-\t\t\t\t; /* do nothing */\n-\t\t\tif (!prefixcmp(buf + offset, \"tagger \"))\n-\t\t\t\tshow_tagger(buf + offset + 7,\n-\t\t\t\t\t    new_offset - offset - 7, rev);\n-\t\t\toffset = new_offset;\n-\t\t}\n+\tassert(type == OBJ_TAG);\n+\twhile (offset < size && buf[offset] != '\\n') {\n+\t\tint new_offset = offset + 1;\n+\t\twhile (new_offset < size && buf[new_offset++] != '\\n')\n+\t\t\t; /* do nothing */\n+\t\tif (!prefixcmp(buf + offset, \"tagger \"))\n+\t\t\tshow_tagger(buf + offset + 7,\n+\t\t\t\t    new_offset - offset - 7, rev);\n+\t\toffset = new_offset;\n+\t}\n \n \tif (offset < size)\n \t\tfwrite(buf + offset, size - offset, 1, stdout);\n@@ -459,7 +465,7 @@ int cmd_show(int argc, const char **argv, const char *prefix)\n \t\tconst char *name = objects[i].name;\n \t\tswitch (o->type) {\n \t\tcase OBJ_BLOB:\n-\t\t\tret = show_object(o->sha1, 0, NULL);\n+\t\t\tret = show_blob_object(o->sha1, NULL);\n \t\t\tbreak;\n \t\tcase OBJ_TAG: {\n \t\t\tstruct tag *t = (struct tag *)o;\n@@ -470,7 +476,7 @@ int cmd_show(int argc, const char **argv, const char *prefix)\n \t\t\t\t\tdiff_get_color_opt(&rev.diffopt, DIFF_COMMIT),\n \t\t\t\t\tt->tag,\n \t\t\t\t\tdiff_get_color_opt(&rev.diffopt, DIFF_RESET));\n-\t\t\tret = show_object(o->sha1, 1, &rev);\n+\t\t\tret = show_tag_object(o->sha1, &rev);\n \t\t\trev.shown_one = 1;\n \t\t\tif (ret)\n \t\t\t\tbreak;\ndiff --git a/t/t1050-large.sh b/t/t1050-large.sh\nindex 97ad5b3..4e08e02 100755\n--- a/t/t1050-large.sh\n+++ b/t/t1050-large.sh\n@@ -125,7 +125,7 @@ test_expect_success 'cat-file a large file from a tag' '\n \tgit cat-file blob largefiletag >/dev/null\n '\n \n-test_expect_failure 'git-show a large file' '\n+test_expect_success 'git-show a large file' '\n \tgit show :large1 >/dev/null\n \n '\n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"186075","messageId":"1330919028-6611-7-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330865996-2069-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v3 06/11] index-pack: split second pass obj handling into own function","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-05T03:43:43Z","receivedAt":"2012-03-05T03:43:43Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/index-pack.c |   31 ++++++++++++++++++-------------\n 1 files changed, 18 insertions(+), 13 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex dd1c5c9..918684f 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -682,6 +682,23 @@ static int compare_delta_entry(const void *a, const void *b)\n \t\t\t\t   objects[delta_b->obj_no].type);\n }\n \n+/*\n+ * Second pass:\n+ * - for all non-delta objects, look if it is used as a base for\n+ *   deltas;\n+ * - if used as a base, uncompress the object and apply all deltas,\n+ *   recursively checking if the resulting object is used as a base\n+ *   for some more deltas.\n+ */\n+static void second_pass(struct object_entry *obj)\n+{\n+\tstruct base_data *base_obj = alloc_base_data();\n+\tbase_obj->obj = obj;\n+\tbase_obj->data = NULL;\n+\tfind_unresolved_deltas(base_obj);\n+\tdisplay_progress(progress, nr_resolved_deltas);\n+}\n+\n /* Parse all objects and return the pack content SHA1 hash */\n static void parse_pack_objects(unsigned char *sha1)\n {\n@@ -736,26 +753,14 @@ static void parse_pack_objects(unsigned char *sha1)\n \tqsort(deltas, nr_deltas, sizeof(struct delta_entry),\n \t      compare_delta_entry);\n \n-\t/*\n-\t * Second pass:\n-\t * - for all non-delta objects, look if it is used as a base for\n-\t *   deltas;\n-\t * - if used as a base, uncompress the object and apply all deltas,\n-\t *   recursively checking if the resulting object is used as a base\n-\t *   for some more deltas.\n-\t */\n \tif (verbose)\n \t\tprogress = start_progress(\"Resolving deltas\", nr_deltas);\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tstruct object_entry *obj = &objects[i];\n-\t\tstruct base_data *base_obj = alloc_base_data();\n \n \t\tif (is_delta_type(obj->type))\n \t\t\tcontinue;\n-\t\tbase_obj->obj = obj;\n-\t\tbase_obj->data = NULL;\n-\t\tfind_unresolved_deltas(base_obj);\n-\t\tdisplay_progress(progress, nr_resolved_deltas);\n+\t\tsecond_pass(obj);\n \t}\n }\n \n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"186076","messageId":"1330919028-6611-8-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330865996-2069-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v3 07/11] index-pack: reduce memory usage when the pack has large blobs","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-05T03:43:44Z","receivedAt":"2012-03-05T03:43:44Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"This command unpacks every non-delta objects in order to:\n\n1. calculate sha-1\n2. do byte-to-byte sha-1 collision test if we happen to have objects\n   with the same sha-1\n3. validate object content in strict mode\n\nAll this requires the entire object to stay in memory, a bad news for\ngiant blobs. This patch lowers memory consumption by not saving the\nobject in memory whenever possible, calculating SHA-1 while unpacking\nthe object.\n\nThis patch assumes that the collision test is rarely needed. The\ncollision test will be done later in second pass if necessary, which\nputs the entire object back to memory again (We could even do the\ncollision test without putting the entire object back in memory, by\ncomparing as we unpack it).\n\nIn strict mode, it always keeps non-blob objects in memory for\nvalidation (blobs do not need data validation).\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/index-pack.c |   64 +++++++++++++++++++++++++++++++++++++++----------\n t/t1050-large.sh     |    4 +-\n 2 files changed, 53 insertions(+), 15 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 918684f..db27133 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -276,30 +276,60 @@ static void unlink_base_data(struct base_data *c)\n \tfree_base_data(c);\n }\n \n-static void *unpack_entry_data(unsigned long offset, unsigned long size)\n+static void *unpack_entry_data(unsigned long offset, unsigned long size,\n+\t\t\t       enum object_type type, unsigned char *sha1)\n {\n+\tstatic char fixed_buf[8192];\n \tint status;\n \tgit_zstream stream;\n-\tvoid *buf = xmalloc(size);\n+\tvoid *buf;\n+\tgit_SHA_CTX c;\n+\n+\tif (sha1) {\t\t/* do hash_sha1_file internally */\n+\t\tchar hdr[32];\n+\t\tint hdrlen = sprintf(hdr, \"%s %lu\", typename(type), size)+1;\n+\t\tgit_SHA1_Init(&c);\n+\t\tgit_SHA1_Update(&c, hdr, hdrlen);\n+\n+\t\tbuf = fixed_buf;\n+\t} else {\n+\t\tbuf = xmalloc(size);\n+\t}\n \n \tmemset(&stream, 0, sizeof(stream));\n \tgit_inflate_init(&stream);\n \tstream.next_out = buf;\n-\tstream.avail_out = size;\n+\tstream.avail_out = buf == fixed_buf ? sizeof(fixed_buf) : size;\n \n \tdo {\n \t\tstream.next_in = fill(1);\n \t\tstream.avail_in = input_len;\n \t\tstatus = git_inflate(&stream, 0);\n \t\tuse(input_len - stream.avail_in);\n+\t\tif (sha1) {\n+\t\t\tgit_SHA1_Update(&c, buf, stream.next_out - (unsigned char *)buf);\n+\t\t\tstream.next_out = buf;\n+\t\t\tstream.avail_out = sizeof(fixed_buf);\n+\t\t}\n \t} while (status == Z_OK);\n \tif (stream.total_out != size || status != Z_STREAM_END)\n \t\tbad_object(offset, \"inflate returned %d\", status);\n \tgit_inflate_end(&stream);\n+\tif (sha1) {\n+\t\tgit_SHA1_Final(sha1, &c);\n+\t\tbuf = NULL;\n+\t}\n \treturn buf;\n }\n \n-static void *unpack_raw_entry(struct object_entry *obj, union delta_base *delta_base)\n+static int is_delta_type(enum object_type type)\n+{\n+\treturn (type == OBJ_REF_DELTA || type == OBJ_OFS_DELTA);\n+}\n+\n+static void *unpack_raw_entry(struct object_entry *obj,\n+\t\t\t      union delta_base *delta_base,\n+\t\t\t      unsigned char *sha1)\n {\n \tunsigned char *p;\n \tunsigned long size, c;\n@@ -359,7 +389,9 @@ static void *unpack_raw_entry(struct object_entry *obj, union delta_base *delta_\n \t}\n \tobj->hdr_size = consumed_bytes - obj->idx.offset;\n \n-\tdata = unpack_entry_data(obj->idx.offset, obj->size);\n+\tif (is_delta_type(obj->type) || strict)\n+\t\tsha1 = NULL;\t/* save unpacked object */\n+\tdata = unpack_entry_data(obj->idx.offset, obj->size, obj->type, sha1);\n \tobj->idx.crc32 = input_crc32;\n \treturn data;\n }\n@@ -460,8 +492,9 @@ static void find_delta_children(const union delta_base *base,\n static void sha1_object(const void *data, unsigned long size,\n \t\t\tenum object_type type, unsigned char *sha1)\n {\n-\thash_sha1_file(data, size, typename(type), sha1);\n-\tif (has_sha1_file(sha1)) {\n+\tif (data)\n+\t\thash_sha1_file(data, size, typename(type), sha1);\n+\tif (data && has_sha1_file(sha1)) {\n \t\tvoid *has_data;\n \t\tenum object_type has_type;\n \t\tunsigned long has_size;\n@@ -510,11 +543,6 @@ static void sha1_object(const void *data, unsigned long size,\n \t}\n }\n \n-static int is_delta_type(enum object_type type)\n-{\n-\treturn (type == OBJ_REF_DELTA || type == OBJ_OFS_DELTA);\n-}\n-\n /*\n  * This function is part of find_unresolved_deltas(). There are two\n  * walkers going in the opposite ways.\n@@ -689,10 +717,20 @@ static int compare_delta_entry(const void *a, const void *b)\n  * - if used as a base, uncompress the object and apply all deltas,\n  *   recursively checking if the resulting object is used as a base\n  *   for some more deltas.\n+ * - if the same object exists in repository and we're not in strict\n+ *   mode, we skipped the sha-1 collision test in the first pass.\n+ *   Do it now.\n  */\n static void second_pass(struct object_entry *obj)\n {\n \tstruct base_data *base_obj = alloc_base_data();\n+\n+\tif (!strict && has_sha1_file(obj->idx.sha1)) {\n+\t\tvoid *data = get_data_from_pack(obj);\n+\t\tsha1_object(data, obj->size, obj->type, obj->idx.sha1);\n+\t\tfree(data);\n+\t}\n+\n \tbase_obj->obj = obj;\n \tbase_obj->data = NULL;\n \tfind_unresolved_deltas(base_obj);\n@@ -718,7 +756,7 @@ static void parse_pack_objects(unsigned char *sha1)\n \t\t\t\tnr_objects);\n \tfor (i = 0; i < nr_objects; i++) {\n \t\tstruct object_entry *obj = &objects[i];\n-\t\tvoid *data = unpack_raw_entry(obj, &delta->base);\n+\t\tvoid *data = unpack_raw_entry(obj, &delta->base, obj->idx.sha1);\n \t\tobj->real_type = obj->type;\n \t\tif (is_delta_type(obj->type)) {\n \t\t\tnr_deltas++;\ndiff --git a/t/t1050-large.sh b/t/t1050-large.sh\nindex 4e08e02..e4b77a2 100755\n--- a/t/t1050-large.sh\n+++ b/t/t1050-large.sh\n@@ -130,11 +130,11 @@ test_expect_success 'git-show a large file' '\n \n '\n \n-test_expect_failure 'clone' '\n+test_expect_success 'clone' '\n \tgit clone file://\"$PWD\"/.git new\n '\n \n-test_expect_failure 'fetch updates' '\n+test_expect_success 'fetch updates' '\n \techo modified >> large1 &&\n \tgit commit -q -a -m updated &&\n \t(\n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"186077","messageId":"1330919028-6611-9-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330865996-2069-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v3 08/11] pack-check: do not unpack blobs","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-05T03:43:45Z","receivedAt":"2012-03-05T03:43:45Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"blob content is not used by verify_pack caller (currently only fsck),\nwe only need to make sure blob sha-1 signature matches its\ncontent. unpack_entry() is taught to hash pack entry as it is\nunpacked, eliminating the need to keep whole blob in memory.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n cache.h          |    2 +-\n fast-import.c    |    2 +-\n pack-check.c     |   21 ++++++++++++++++++++-\n sha1_file.c      |   45 +++++++++++++++++++++++++++++++++++----------\n t/t1050-large.sh |    2 +-\n 5 files changed, 58 insertions(+), 14 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex e12b15f..3365f89 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1062,7 +1062,7 @@ extern const unsigned char *nth_packed_object_sha1(struct packed_git *, uint32_t\n extern off_t nth_packed_object_offset(const struct packed_git *, uint32_t);\n extern off_t find_pack_entry_one(const unsigned char *, struct packed_git *);\n extern int is_pack_valid(struct packed_git *);\n-extern void *unpack_entry(struct packed_git *, off_t, enum object_type *, unsigned long *);\n+extern void *unpack_entry(struct packed_git *, off_t, enum object_type *, unsigned long *, unsigned char *);\n extern unsigned long unpack_object_header_buffer(const unsigned char *buf, unsigned long len, enum object_type *type, unsigned long *sizep);\n extern unsigned long get_size_from_delta(struct packed_git *, struct pack_window **, off_t);\n extern int unpack_object_header(struct packed_git *, struct pack_window **, off_t *, unsigned long *);\ndiff --git a/fast-import.c b/fast-import.c\nindex 6cd19e5..5e94a64 100644\n--- a/fast-import.c\n+++ b/fast-import.c\n@@ -1303,7 +1303,7 @@ static void *gfi_unpack_entry(\n \t\t */\n \t\tp->pack_size = pack_size + 20;\n \t}\n-\treturn unpack_entry(p, oe->idx.offset, &type, sizep);\n+\treturn unpack_entry(p, oe->idx.offset, &type, sizep, NULL);\n }\n \n static const char *get_mode(const char *str, uint16_t *modep)\ndiff --git a/pack-check.c b/pack-check.c\nindex 63a595c..1920bdb 100644\n--- a/pack-check.c\n+++ b/pack-check.c\n@@ -105,6 +105,7 @@ static int verify_packfile(struct packed_git *p,\n \t\tvoid *data;\n \t\tenum object_type type;\n \t\tunsigned long size;\n+\t\toff_t curpos = entries[i].offset;\n \n \t\tif (p->index_version > 1) {\n \t\t\toff_t offset = entries[i].offset;\n@@ -116,7 +117,25 @@ static int verify_packfile(struct packed_git *p,\n \t\t\t\t\t    sha1_to_hex(entries[i].sha1),\n \t\t\t\t\t    p->pack_name, (uintmax_t)offset);\n \t\t}\n-\t\tdata = unpack_entry(p, entries[i].offset, &type, &size);\n+\t\ttype = unpack_object_header(p, w_curs, &curpos, &size);\n+\t\tunuse_pack(w_curs);\n+\t\tif (type == OBJ_BLOB) {\n+\t\t\tunsigned char sha1[20];\n+\t\t\tdata = unpack_entry(p, entries[i].offset, &type, &size, sha1);\n+\t\t\tif (!data) {\n+\t\t\t\tif (hashcmp(entries[i].sha1, sha1))\n+\t\t\t\t\terr = error(\"packed %s from %s is corrupt\",\n+\t\t\t\t\t\t    sha1_to_hex(entries[i].sha1), p->pack_name);\n+\t\t\t\telse if (fn) {\n+\t\t\t\t\tint eaten = 0;\n+\t\t\t\t\tfn(entries[i].sha1, type, size, NULL, &eaten);\n+\t\t\t\t}\n+\t\t\t\tif (((base_count + i) & 1023) == 0)\n+\t\t\t\t\tdisplay_progress(progress, base_count + i);\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t}\n+\t\tdata = unpack_entry(p, entries[i].offset, &type, &size, NULL);\n \t\tif (!data)\n \t\t\terr = error(\"cannot unpack %s from %s at offset %\"PRIuMAX\"\",\n \t\t\t\t    sha1_to_hex(entries[i].sha1), p->pack_name,\ndiff --git a/sha1_file.c b/sha1_file.c\nindex a77ef0a..d68a5b0 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -1653,28 +1653,51 @@ static int packed_object_info(struct packed_git *p, off_t obj_offset,\n }\n \n static void *unpack_compressed_entry(struct packed_git *p,\n-\t\t\t\t    struct pack_window **w_curs,\n-\t\t\t\t    off_t curpos,\n-\t\t\t\t    unsigned long size)\n+\t\t\t\t     struct pack_window **w_curs,\n+\t\t\t\t     off_t curpos,\n+\t\t\t\t     unsigned long size,\n+\t\t\t\t     enum object_type type,\n+\t\t\t\t     unsigned char *sha1)\n {\n+\tstatic unsigned char fixed_buf[8192];\n \tint st;\n \tgit_zstream stream;\n \tunsigned char *buffer, *in;\n+\tgit_SHA_CTX c;\n+\n+\tif (sha1) {\t\t/* do hash_sha1_file internally */\n+\t\tchar hdr[32];\n+\t\tint hdrlen = sprintf(hdr, \"%s %lu\", typename(type), size)+1;\n+\t\tgit_SHA1_Init(&c);\n+\t\tgit_SHA1_Update(&c, hdr, hdrlen);\n+\n+\t\tbuffer = fixed_buf;\n+\t} else {\n+\t\tbuffer = xmallocz(size);\n+\t}\n \n-\tbuffer = xmallocz(size);\n \tmemset(&stream, 0, sizeof(stream));\n \tstream.next_out = buffer;\n-\tstream.avail_out = size + 1;\n+\tstream.avail_out = buffer == fixed_buf ? sizeof(fixed_buf) : size + 1;\n \n \tgit_inflate_init(&stream);\n \tdo {\n \t\tin = use_pack(p, w_curs, curpos, &stream.avail_in);\n \t\tstream.next_in = in;\n \t\tst = git_inflate(&stream, Z_FINISH);\n-\t\tif (!stream.avail_out)\n+\t\tif (sha1) {\n+\t\t\tgit_SHA1_Update(&c, buffer, stream.next_out - (unsigned char *)buffer);\n+\t\t\tstream.next_out = buffer;\n+\t\t\tstream.avail_out = sizeof(fixed_buf);\n+\t\t}\n+\t\telse if (!stream.avail_out)\n \t\t\tbreak; /* the payload is larger than it should be */\n \t\tcurpos += stream.next_in - in;\n \t} while (st == Z_OK || st == Z_BUF_ERROR);\n+\tif (sha1) {\n+\t\tgit_SHA1_Final(sha1, &c);\n+\t\tbuffer = NULL;\n+\t}\n \tgit_inflate_end(&stream);\n \tif ((st != Z_STREAM_END) || stream.total_out != size) {\n \t\tfree(buffer);\n@@ -1727,7 +1750,7 @@ static void *cache_or_unpack_entry(struct packed_git *p, off_t base_offset,\n \n \tret = ent->data;\n \tif (!ret || ent->p != p || ent->base_offset != base_offset)\n-\t\treturn unpack_entry(p, base_offset, type, base_size);\n+\t\treturn unpack_entry(p, base_offset, type, base_size, NULL);\n \n \tif (!keep_cache) {\n \t\tent->data = NULL;\n@@ -1844,7 +1867,7 @@ static void *unpack_delta_entry(struct packed_git *p,\n \t\t\treturn NULL;\n \t}\n \n-\tdelta_data = unpack_compressed_entry(p, w_curs, curpos, delta_size);\n+\tdelta_data = unpack_compressed_entry(p, w_curs, curpos, delta_size, OBJ_NONE, NULL);\n \tif (!delta_data) {\n \t\terror(\"failed to unpack compressed delta \"\n \t\t      \"at offset %\"PRIuMAX\" from %s\",\n@@ -1883,7 +1906,8 @@ static void write_pack_access_log(struct packed_git *p, off_t obj_offset)\n int do_check_packed_object_crc;\n \n void *unpack_entry(struct packed_git *p, off_t obj_offset,\n-\t\t   enum object_type *type, unsigned long *sizep)\n+\t\t   enum object_type *type, unsigned long *sizep,\n+\t\t   unsigned char *sha1)\n {\n \tstruct pack_window *w_curs = NULL;\n \toff_t curpos = obj_offset;\n@@ -1917,7 +1941,8 @@ void *unpack_entry(struct packed_git *p, off_t obj_offset,\n \tcase OBJ_TREE:\n \tcase OBJ_BLOB:\n \tcase OBJ_TAG:\n-\t\tdata = unpack_compressed_entry(p, &w_curs, curpos, *sizep);\n+\t\tdata = unpack_compressed_entry(p, &w_curs, curpos,\n+\t\t\t\t\t       *sizep, *type, sha1);\n \t\tbreak;\n \tdefault:\n \t\tdata = NULL;\ndiff --git a/t/t1050-large.sh b/t/t1050-large.sh\nindex e4b77a2..52acae5 100755\n--- a/t/t1050-large.sh\n+++ b/t/t1050-large.sh\n@@ -143,7 +143,7 @@ test_expect_success 'fetch updates' '\n \t)\n '\n \n-test_expect_failure 'fsck' '\n+test_expect_success 'fsck' '\n \tgit fsck --full\n '\n \n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"186078","messageId":"1330919028-6611-10-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330865996-2069-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v3 09/11] archive: support streaming large files to a tar archive","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-05T03:43:46Z","receivedAt":"2012-03-05T03:43:46Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n archive-tar.c    |   35 ++++++++++++++++++++++++++++-------\n archive-zip.c    |    9 +++++----\n archive.c        |   51 ++++++++++++++++++++++++++++++++++-----------------\n archive.h        |   11 +++++++++--\n t/t1050-large.sh |    2 +-\n 5 files changed, 77 insertions(+), 31 deletions(-)\n\ndiff --git a/archive-tar.c b/archive-tar.c\nindex 20af005..5bffe49 100644\n--- a/archive-tar.c\n+++ b/archive-tar.c\n@@ -5,6 +5,7 @@\n #include \"tar.h\"\n #include \"archive.h\"\n #include \"run-command.h\"\n+#include \"streaming.h\"\n \n #define RECORDSIZE\t(512)\n #define BLOCKSIZE\t(RECORDSIZE * 20)\n@@ -123,9 +124,29 @@ static size_t get_path_prefix(const char *path, size_t pathlen, size_t maxlen)\n \treturn i;\n }\n \n+static void write_file(struct git_istream *stream, const void *buffer,\n+\t\t       unsigned long size)\n+{\n+\tif (!stream) {\n+\t\twrite_blocked(buffer, size);\n+\t\treturn;\n+\t}\n+\tfor (;;) {\n+\t\tchar buf[1024 * 16];\n+\t\tssize_t readlen;\n+\n+\t\treadlen = read_istream(stream, buf, sizeof(buf));\n+\n+\t\tif (!readlen)\n+\t\t\tbreak;\n+\t\twrite_blocked(buf, readlen);\n+\t}\n+}\n+\n static int write_tar_entry(struct archiver_args *args,\n-\t\tconst unsigned char *sha1, const char *path, size_t pathlen,\n-\t\tunsigned int mode, void *buffer, unsigned long size)\n+\t\t\t   const unsigned char *sha1, const char *path,\n+\t\t\t   size_t pathlen, unsigned int mode, void *buffer,\n+\t\t\t   struct git_istream *stream, unsigned long size)\n {\n \tstruct ustar_header header;\n \tstruct strbuf ext_header = STRBUF_INIT;\n@@ -200,14 +221,14 @@ static int write_tar_entry(struct archiver_args *args,\n \n \tif (ext_header.len > 0) {\n \t\terr = write_tar_entry(args, sha1, NULL, 0, 0, ext_header.buf,\n-\t\t\t\text_header.len);\n+\t\t\t\t      NULL, ext_header.len);\n \t\tif (err)\n \t\t\treturn err;\n \t}\n \tstrbuf_release(&ext_header);\n \twrite_blocked(&header, sizeof(header));\n-\tif (S_ISREG(mode) && buffer && size > 0)\n-\t\twrite_blocked(buffer, size);\n+\tif (S_ISREG(mode) && size > 0)\n+\t\twrite_file(stream, buffer, size);\n \treturn err;\n }\n \n@@ -219,7 +240,7 @@ static int write_global_extended_header(struct archiver_args *args)\n \n \tstrbuf_append_ext_header(&ext_header, \"comment\", sha1_to_hex(sha1), 40);\n \terr = write_tar_entry(args, NULL, NULL, 0, 0, ext_header.buf,\n-\t\t\text_header.len);\n+\t\t\t      NULL, ext_header.len);\n \tstrbuf_release(&ext_header);\n \treturn err;\n }\n@@ -308,7 +329,7 @@ static int write_tar_archive(const struct archiver *ar,\n \tif (args->commit_sha1)\n \t\terr = write_global_extended_header(args);\n \tif (!err)\n-\t\terr = write_archive_entries(args, write_tar_entry);\n+\t\terr = write_archive_entries(args, write_tar_entry, 1);\n \tif (!err)\n \t\twrite_trailer();\n \treturn err;\ndiff --git a/archive-zip.c b/archive-zip.c\nindex 02d1f37..4a1e917 100644\n--- a/archive-zip.c\n+++ b/archive-zip.c\n@@ -120,9 +120,10 @@ static void *zlib_deflate(void *data, unsigned long size,\n \treturn buffer;\n }\n \n-static int write_zip_entry(struct archiver_args *args,\n-\t\tconst unsigned char *sha1, const char *path, size_t pathlen,\n-\t\tunsigned int mode, void *buffer, unsigned long size)\n+int write_zip_entry(struct archiver_args *args,\n+\t\t\t   const unsigned char *sha1, const char *path,\n+\t\t\t   size_t pathlen, unsigned int mode, void *buffer,\n+\t\t\t   struct git_istream *stream, unsigned long size)\n {\n \tstruct zip_local_header header;\n \tstruct zip_dir_header dirent;\n@@ -271,7 +272,7 @@ static int write_zip_archive(const struct archiver *ar,\n \tzip_dir = xmalloc(ZIP_DIRECTORY_MIN_SIZE);\n \tzip_dir_size = ZIP_DIRECTORY_MIN_SIZE;\n \n-\terr = write_archive_entries(args, write_zip_entry);\n+\terr = write_archive_entries(args, write_zip_entry, 0);\n \tif (!err)\n \t\twrite_zip_trailer(args->commit_sha1);\n \ndiff --git a/archive.c b/archive.c\nindex 1ee837d..257eadf 100644\n--- a/archive.c\n+++ b/archive.c\n@@ -5,6 +5,7 @@\n #include \"archive.h\"\n #include \"parse-options.h\"\n #include \"unpack-trees.h\"\n+#include \"streaming.h\"\n \n static char const * const archive_usage[] = {\n \t\"git archive [options] <tree-ish> [<path>...]\",\n@@ -59,26 +60,35 @@ static void format_subst(const struct commit *commit,\n \tfree(to_free);\n }\n \n-static void *sha1_file_to_archive(const char *path, const unsigned char *sha1,\n-\t\tunsigned int mode, enum object_type *type,\n-\t\tunsigned long *sizep, const struct commit *commit)\n+void sha1_file_to_archive(void **buffer, struct git_istream **stream,\n+\t\t\t  const char *path, const unsigned char *sha1,\n+\t\t\t  unsigned int mode, enum object_type *type,\n+\t\t\t  unsigned long *sizep,\n+\t\t\t  const struct commit *commit)\n {\n-\tvoid *buffer;\n+\tif (stream) {\n+\t\tstruct stream_filter *filter;\n+\t\tfilter = get_stream_filter(path, sha1);\n+\t\tif (!commit && S_ISREG(mode) && is_null_stream_filter(filter)) {\n+\t\t\t*buffer = NULL;\n+\t\t\t*stream = open_istream(sha1, type, sizep, NULL);\n+\t\t\treturn;\n+\t\t}\n+\t\t*stream = NULL;\n+\t}\n \n-\tbuffer = read_sha1_file(sha1, type, sizep);\n-\tif (buffer && S_ISREG(mode)) {\n+\t*buffer = read_sha1_file(sha1, type, sizep);\n+\tif (*buffer && S_ISREG(mode)) {\n \t\tstruct strbuf buf = STRBUF_INIT;\n \t\tsize_t size = 0;\n \n-\t\tstrbuf_attach(&buf, buffer, *sizep, *sizep + 1);\n+\t\tstrbuf_attach(&buf, *buffer, *sizep, *sizep + 1);\n \t\tconvert_to_working_tree(path, buf.buf, buf.len, &buf);\n \t\tif (commit)\n \t\t\tformat_subst(commit, buf.buf, buf.len, &buf);\n-\t\tbuffer = strbuf_detach(&buf, &size);\n+\t\t*buffer = strbuf_detach(&buf, &size);\n \t\t*sizep = size;\n \t}\n-\n-\treturn buffer;\n }\n \n static void setup_archive_check(struct git_attr_check *check)\n@@ -97,6 +107,7 @@ static void setup_archive_check(struct git_attr_check *check)\n struct archiver_context {\n \tstruct archiver_args *args;\n \twrite_archive_entry_fn_t write_entry;\n+\tint stream_ok;\n };\n \n static int write_archive_entry(const unsigned char *sha1, const char *base,\n@@ -109,6 +120,7 @@ static int write_archive_entry(const unsigned char *sha1, const char *base,\n \twrite_archive_entry_fn_t write_entry = c->write_entry;\n \tstruct git_attr_check check[2];\n \tconst char *path_without_prefix;\n+\tstruct git_istream *stream = NULL;\n \tint convert = 0;\n \tint err;\n \tenum object_type type;\n@@ -133,25 +145,29 @@ static int write_archive_entry(const unsigned char *sha1, const char *base,\n \t\tstrbuf_addch(&path, '/');\n \t\tif (args->verbose)\n \t\t\tfprintf(stderr, \"%.*s\\n\", (int)path.len, path.buf);\n-\t\terr = write_entry(args, sha1, path.buf, path.len, mode, NULL, 0);\n+\t\terr = write_entry(args, sha1, path.buf, path.len, mode, NULL, NULL, 0);\n \t\tif (err)\n \t\t\treturn err;\n \t\treturn (S_ISDIR(mode) ? READ_TREE_RECURSIVE : 0);\n \t}\n \n-\tbuffer = sha1_file_to_archive(path_without_prefix, sha1, mode,\n-\t\t\t&type, &size, convert ? args->commit : NULL);\n-\tif (!buffer)\n+\tsha1_file_to_archive(&buffer, c->stream_ok ? &stream : NULL,\n+\t\t\t     path_without_prefix, sha1, mode,\n+\t\t\t     &type, &size, convert ? args->commit : NULL);\n+\tif (!buffer && !stream)\n \t\treturn error(\"cannot read %s\", sha1_to_hex(sha1));\n \tif (args->verbose)\n \t\tfprintf(stderr, \"%.*s\\n\", (int)path.len, path.buf);\n-\terr = write_entry(args, sha1, path.buf, path.len, mode, buffer, size);\n+\terr = write_entry(args, sha1, path.buf, path.len, mode, buffer, stream, size);\n+\tif (stream)\n+\t\tclose_istream(stream);\n \tfree(buffer);\n \treturn err;\n }\n \n int write_archive_entries(struct archiver_args *args,\n-\t\twrite_archive_entry_fn_t write_entry)\n+\t\t\t  write_archive_entry_fn_t write_entry,\n+\t\t\t  int stream_ok)\n {\n \tstruct archiver_context context;\n \tstruct unpack_trees_options opts;\n@@ -167,13 +183,14 @@ int write_archive_entries(struct archiver_args *args,\n \t\tif (args->verbose)\n \t\t\tfprintf(stderr, \"%.*s\\n\", (int)len, args->base);\n \t\terr = write_entry(args, args->tree->object.sha1, args->base,\n-\t\t\t\tlen, 040777, NULL, 0);\n+\t\t\t\t  len, 040777, NULL, NULL, 0);\n \t\tif (err)\n \t\t\treturn err;\n \t}\n \n \tcontext.args = args;\n \tcontext.write_entry = write_entry;\n+\tcontext.stream_ok = stream_ok;\n \n \t/*\n \t * Setup index and instruct attr to read index only\ndiff --git a/archive.h b/archive.h\nindex 2b0884f..370cca9 100644\n--- a/archive.h\n+++ b/archive.h\n@@ -27,9 +27,16 @@ extern void register_archiver(struct archiver *);\n extern void init_tar_archiver(void);\n extern void init_zip_archiver(void);\n \n-typedef int (*write_archive_entry_fn_t)(struct archiver_args *args, const unsigned char *sha1, const char *path, size_t pathlen, unsigned int mode, void *buffer, unsigned long size);\n+struct git_istream;\n+typedef int (*write_archive_entry_fn_t)(struct archiver_args *args,\n+\t\t\t\t\tconst unsigned char *sha1,\n+\t\t\t\t\tconst char *path, size_t pathlen,\n+\t\t\t\t\tunsigned int mode,\n+\t\t\t\t\tvoid *buffer,\n+\t\t\t\t\tstruct git_istream *stream,\n+\t\t\t\t\tunsigned long size);\n \n-extern int write_archive_entries(struct archiver_args *args, write_archive_entry_fn_t write_entry);\n+extern int write_archive_entries(struct archiver_args *args, write_archive_entry_fn_t write_entry, int stream_ok);\n extern int write_archive(int argc, const char **argv, const char *prefix, int setup_prefix, const char *name_hint, int remote);\n \n const char *archive_format_from_filename(const char *filename);\ndiff --git a/t/t1050-large.sh b/t/t1050-large.sh\nindex 52acae5..5336eb8 100755\n--- a/t/t1050-large.sh\n+++ b/t/t1050-large.sh\n@@ -151,7 +151,7 @@ test_expect_failure 'repack' '\n \tgit repack -ad\n '\n \n-test_expect_failure 'tar achiving' '\n+test_expect_success 'tar achiving' '\n \tgit archive --format=tar HEAD >/dev/null\n '\n \n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"186079","messageId":"1330919028-6611-11-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330865996-2069-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v3 10/11] fsck: use streaming interface for writing lost-found blobs","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-05T03:43:47Z","receivedAt":"2012-03-05T03:43:47Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/fsck.c |    8 ++------\n 1 files changed, 2 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/fsck.c b/builtin/fsck.c\nindex 8c479a7..7fcb33e 100644\n--- a/builtin/fsck.c\n+++ b/builtin/fsck.c\n@@ -12,6 +12,7 @@\n #include \"parse-options.h\"\n #include \"dir.h\"\n #include \"progress.h\"\n+#include \"streaming.h\"\n \n #define REACHABLE 0x0001\n #define SEEN      0x0002\n@@ -236,13 +237,8 @@ static void check_unreachable_object(struct object *obj)\n \t\t\tif (!(f = fopen(filename, \"w\")))\n \t\t\t\tdie_errno(\"Could not open '%s'\", filename);\n \t\t\tif (obj->type == OBJ_BLOB) {\n-\t\t\t\tenum object_type type;\n-\t\t\t\tunsigned long size;\n-\t\t\t\tchar *buf = read_sha1_file(obj->sha1,\n-\t\t\t\t\t\t&type, &size);\n-\t\t\t\tif (buf && fwrite(buf, 1, size, f) != size)\n+\t\t\t\tif (stream_blob_to_fd(fileno(f), obj->sha1, NULL, 1))\n \t\t\t\t\tdie_errno(\"Could not write '%s'\", filename);\n-\t\t\t\tfree(buf);\n \t\t\t} else\n \t\t\t\tfprintf(f, \"%s\\n\", sha1_to_hex(obj->sha1));\n \t\t\tif (fclose(f))\n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"186080","messageId":"1330919028-6611-12-git-send-email-pclouds@gmail.com","threadId":"29753","inReplyTo":"1330865996-2069-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v3 11/11] update-server-info: respect core.bigfilethreshold","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-05T03:43:48Z","receivedAt":"2012-03-05T03:43:48Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"This command indirectly calls check_sha1_signature() (add_info_ref ->\nderef_tag -> parse_object -> ..) , which may put whole blob in memory\nif the blob's size is under core.bigfilethreshold. As config is not\nread, the threshold is always 512MB. Respect user settings here.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n builtin/update-server-info.c |    1 +\n t/t1050-large.sh             |    2 +-\n 2 files changed, 2 insertions(+), 1 deletions(-)\n\ndiff --git a/builtin/update-server-info.c b/builtin/update-server-info.c\nindex b90dce6..0d63c44 100644\n--- a/builtin/update-server-info.c\n+++ b/builtin/update-server-info.c\n@@ -15,6 +15,7 @@ int cmd_update_server_info(int argc, const char **argv, const char *prefix)\n \t\tOPT_END()\n \t};\n \n+\tgit_config(git_default_config, NULL);\n \targc = parse_options(argc, argv, prefix, options,\n \t\t\t     update_server_info_usage, 0);\n \tif (argc > 0)\ndiff --git a/t/t1050-large.sh b/t/t1050-large.sh\nindex 5336eb8..9197b89 100755\n--- a/t/t1050-large.sh\n+++ b/t/t1050-large.sh\n@@ -147,7 +147,7 @@ test_expect_success 'fsck' '\n \tgit fsck --full\n '\n \n-test_expect_failure 'repack' '\n+test_expect_success 'repack' '\n \tgit repack -ad\n '\n \n-- \n1.7.3.1.256.g2539c.dirty\n"},{"id":"186161","messageId":"7vty22wnpe.fsf@alter.siamese.dyndns.org","threadId":"29753","inReplyTo":"1330919028-6611-5-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH v3 04/11] parse_object: special code path for blobs to avoid putting whole object in memory","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-03-06T00:57:33Z","receivedAt":"2012-03-06T00:57:33Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n\nThe code looks OK but the updated API into check_sha1_signature()\nneeds to be explained both in-code comment and the log message.\n\nI'll push out an updated version later on 'pu'.\n"},{"id":"186162","messageId":"7vmx7uwnp0.fsf@alter.siamese.dyndns.org","threadId":"29753","inReplyTo":"1330919028-6611-10-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH v3 09/11] archive: support streaming large files to a tar archive","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-03-06T00:57:47Z","receivedAt":"2012-03-06T00:57:47Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n\nThis is *way* *too* underdocumented.\n\nFor example, it is totally unclear from the patch what determines\nthe last parameter to write_archive_entries(), OK_TO_STREAM.  Does\nit depend on the nature of the payload?  Does the backend decide it,\nin other words, if it is prepared to read from a streaming API or not?\n\nI wanted to first take all the \"do not slurp things in core, and\ninstead read from streaming API\" patches from this series, but I\nhad to stop at this one.\n\n> ---\n>  archive-tar.c    |   35 ++++++++++++++++++++++++++++-------\n>  archive-zip.c    |    9 +++++----\n>  archive.c        |   51 ++++++++++++++++++++++++++++++++++-----------------\n>  archive.h        |   11 +++++++++--\n>  t/t1050-large.sh |    2 +-\n>  5 files changed, 77 insertions(+), 31 deletions(-)\n>\n> diff --git a/archive-tar.c b/archive-tar.c\n> index 20af005..5bffe49 100644\n> --- a/archive-tar.c\n> +++ b/archive-tar.c\n> @@ -5,6 +5,7 @@\n>  #include \"tar.h\"\n>  #include \"archive.h\"\n>  #include \"run-command.h\"\n> +#include \"streaming.h\"\n>  \n>  #define RECORDSIZE\t(512)\n>  #define BLOCKSIZE\t(RECORDSIZE * 20)\n> @@ -123,9 +124,29 @@ static size_t get_path_prefix(const char *path, size_t pathlen, size_t maxlen)\n>  \treturn i;\n>  }\n>  \n> +static void write_file(struct git_istream *stream, const void *buffer,\n> +\t\t       unsigned long size)\n> +{\n> +\tif (!stream) {\n> +\t\twrite_blocked(buffer, size);\n> +\t\treturn;\n> +\t}\n> +\tfor (;;) {\n> +\t\tchar buf[1024 * 16];\n> +\t\tssize_t readlen;\n> +\n> +\t\treadlen = read_istream(stream, buf, sizeof(buf));\n> +\n> +\t\tif (!readlen)\n> +\t\t\tbreak;\n> +\t\twrite_blocked(buf, readlen);\n> +\t}\n> +}\n> +\n>  static int write_tar_entry(struct archiver_args *args,\n> -\t\tconst unsigned char *sha1, const char *path, size_t pathlen,\n> -\t\tunsigned int mode, void *buffer, unsigned long size)\n> +\t\t\t   const unsigned char *sha1, const char *path,\n> +\t\t\t   size_t pathlen, unsigned int mode, void *buffer,\n> +\t\t\t   struct git_istream *stream, unsigned long size)\n>  {\n>  \tstruct ustar_header header;\n>  \tstruct strbuf ext_header = STRBUF_INIT;\n> @@ -200,14 +221,14 @@ static int write_tar_entry(struct archiver_args *args,\n>  \n>  \tif (ext_header.len > 0) {\n>  \t\terr = write_tar_entry(args, sha1, NULL, 0, 0, ext_header.buf,\n> -\t\t\t\text_header.len);\n> +\t\t\t\t      NULL, ext_header.len);\n>  \t\tif (err)\n>  \t\t\treturn err;\n>  \t}\n>  \tstrbuf_release(&ext_header);\n>  \twrite_blocked(&header, sizeof(header));\n> -\tif (S_ISREG(mode) && buffer && size > 0)\n> -\t\twrite_blocked(buffer, size);\n> +\tif (S_ISREG(mode) && size > 0)\n> +\t\twrite_file(stream, buffer, size);\n>  \treturn err;\n>  }\n>  \n> @@ -219,7 +240,7 @@ static int write_global_extended_header(struct archiver_args *args)\n>  \n>  \tstrbuf_append_ext_header(&ext_header, \"comment\", sha1_to_hex(sha1), 40);\n>  \terr = write_tar_entry(args, NULL, NULL, 0, 0, ext_header.buf,\n> -\t\t\text_header.len);\n> +\t\t\t      NULL, ext_header.len);\n>  \tstrbuf_release(&ext_header);\n>  \treturn err;\n>  }\n> @@ -308,7 +329,7 @@ static int write_tar_archive(const struct archiver *ar,\n>  \tif (args->commit_sha1)\n>  \t\terr = write_global_extended_header(args);\n>  \tif (!err)\n> -\t\terr = write_archive_entries(args, write_tar_entry);\n> +\t\terr = write_archive_entries(args, write_tar_entry, 1);\n>  \tif (!err)\n>  \t\twrite_trailer();\n>  \treturn err;\n> diff --git a/archive-zip.c b/archive-zip.c\n> index 02d1f37..4a1e917 100644\n> --- a/archive-zip.c\n> +++ b/archive-zip.c\n> @@ -120,9 +120,10 @@ static void *zlib_deflate(void *data, unsigned long size,\n>  \treturn buffer;\n>  }\n>  \n> -static int write_zip_entry(struct archiver_args *args,\n> -\t\tconst unsigned char *sha1, const char *path, size_t pathlen,\n> -\t\tunsigned int mode, void *buffer, unsigned long size)\n> +int write_zip_entry(struct archiver_args *args,\n> +\t\t\t   const unsigned char *sha1, const char *path,\n> +\t\t\t   size_t pathlen, unsigned int mode, void *buffer,\n> +\t\t\t   struct git_istream *stream, unsigned long size)\n>  {\n>  \tstruct zip_local_header header;\n>  \tstruct zip_dir_header dirent;\n> @@ -271,7 +272,7 @@ static int write_zip_archive(const struct archiver *ar,\n>  \tzip_dir = xmalloc(ZIP_DIRECTORY_MIN_SIZE);\n>  \tzip_dir_size = ZIP_DIRECTORY_MIN_SIZE;\n>  \n> -\terr = write_archive_entries(args, write_zip_entry);\n> +\terr = write_archive_entries(args, write_zip_entry, 0);\n>  \tif (!err)\n>  \t\twrite_zip_trailer(args->commit_sha1);\n>  \n> diff --git a/archive.c b/archive.c\n> index 1ee837d..257eadf 100644\n> --- a/archive.c\n> +++ b/archive.c\n> @@ -5,6 +5,7 @@\n>  #include \"archive.h\"\n>  #include \"parse-options.h\"\n>  #include \"unpack-trees.h\"\n> +#include \"streaming.h\"\n>  \n>  static char const * const archive_usage[] = {\n>  \t\"git archive [options] <tree-ish> [<path>...]\",\n> @@ -59,26 +60,35 @@ static void format_subst(const struct commit *commit,\n>  \tfree(to_free);\n>  }\n>  \n> -static void *sha1_file_to_archive(const char *path, const unsigned char *sha1,\n> -\t\tunsigned int mode, enum object_type *type,\n> -\t\tunsigned long *sizep, const struct commit *commit)\n> +void sha1_file_to_archive(void **buffer, struct git_istream **stream,\n> +\t\t\t  const char *path, const unsigned char *sha1,\n> +\t\t\t  unsigned int mode, enum object_type *type,\n> +\t\t\t  unsigned long *sizep,\n> +\t\t\t  const struct commit *commit)\n>  {\n> -\tvoid *buffer;\n> +\tif (stream) {\n> +\t\tstruct stream_filter *filter;\n> +\t\tfilter = get_stream_filter(path, sha1);\n> +\t\tif (!commit && S_ISREG(mode) && is_null_stream_filter(filter)) {\n> +\t\t\t*buffer = NULL;\n> +\t\t\t*stream = open_istream(sha1, type, sizep, NULL);\n> +\t\t\treturn;\n> +\t\t}\n> +\t\t*stream = NULL;\n> +\t}\n>  \n> -\tbuffer = read_sha1_file(sha1, type, sizep);\n> -\tif (buffer && S_ISREG(mode)) {\n> +\t*buffer = read_sha1_file(sha1, type, sizep);\n> +\tif (*buffer && S_ISREG(mode)) {\n>  \t\tstruct strbuf buf = STRBUF_INIT;\n>  \t\tsize_t size = 0;\n>  \n> -\t\tstrbuf_attach(&buf, buffer, *sizep, *sizep + 1);\n> +\t\tstrbuf_attach(&buf, *buffer, *sizep, *sizep + 1);\n>  \t\tconvert_to_working_tree(path, buf.buf, buf.len, &buf);\n>  \t\tif (commit)\n>  \t\t\tformat_subst(commit, buf.buf, buf.len, &buf);\n> -\t\tbuffer = strbuf_detach(&buf, &size);\n> +\t\t*buffer = strbuf_detach(&buf, &size);\n>  \t\t*sizep = size;\n>  \t}\n> -\n> -\treturn buffer;\n>  }\n>  \n>  static void setup_archive_check(struct git_attr_check *check)\n> @@ -97,6 +107,7 @@ static void setup_archive_check(struct git_attr_check *check)\n>  struct archiver_context {\n>  \tstruct archiver_args *args;\n>  \twrite_archive_entry_fn_t write_entry;\n> +\tint stream_ok;\n>  };\n>  \n>  static int write_archive_entry(const unsigned char *sha1, const char *base,\n> @@ -109,6 +120,7 @@ static int write_archive_entry(const unsigned char *sha1, const char *base,\n>  \twrite_archive_entry_fn_t write_entry = c->write_entry;\n>  \tstruct git_attr_check check[2];\n>  \tconst char *path_without_prefix;\n> +\tstruct git_istream *stream = NULL;\n>  \tint convert = 0;\n>  \tint err;\n>  \tenum object_type type;\n> @@ -133,25 +145,29 @@ static int write_archive_entry(const unsigned char *sha1, const char *base,\n>  \t\tstrbuf_addch(&path, '/');\n>  \t\tif (args->verbose)\n>  \t\t\tfprintf(stderr, \"%.*s\\n\", (int)path.len, path.buf);\n> -\t\terr = write_entry(args, sha1, path.buf, path.len, mode, NULL, 0);\n> +\t\terr = write_entry(args, sha1, path.buf, path.len, mode, NULL, NULL, 0);\n>  \t\tif (err)\n>  \t\t\treturn err;\n>  \t\treturn (S_ISDIR(mode) ? READ_TREE_RECURSIVE : 0);\n>  \t}\n>  \n> -\tbuffer = sha1_file_to_archive(path_without_prefix, sha1, mode,\n> -\t\t\t&type, &size, convert ? args->commit : NULL);\n> -\tif (!buffer)\n> +\tsha1_file_to_archive(&buffer, c->stream_ok ? &stream : NULL,\n> +\t\t\t     path_without_prefix, sha1, mode,\n> +\t\t\t     &type, &size, convert ? args->commit : NULL);\n> +\tif (!buffer && !stream)\n>  \t\treturn error(\"cannot read %s\", sha1_to_hex(sha1));\n>  \tif (args->verbose)\n>  \t\tfprintf(stderr, \"%.*s\\n\", (int)path.len, path.buf);\n> -\terr = write_entry(args, sha1, path.buf, path.len, mode, buffer, size);\n> +\terr = write_entry(args, sha1, path.buf, path.len, mode, buffer, stream, size);\n> +\tif (stream)\n> +\t\tclose_istream(stream);\n>  \tfree(buffer);\n>  \treturn err;\n>  }\n>  \n>  int write_archive_entries(struct archiver_args *args,\n> -\t\twrite_archive_entry_fn_t write_entry)\n> +\t\t\t  write_archive_entry_fn_t write_entry,\n> +\t\t\t  int stream_ok)\n>  {\n>  \tstruct archiver_context context;\n>  \tstruct unpack_trees_options opts;\n> @@ -167,13 +183,14 @@ int write_archive_entries(struct archiver_args *args,\n>  \t\tif (args->verbose)\n>  \t\t\tfprintf(stderr, \"%.*s\\n\", (int)len, args->base);\n>  \t\terr = write_entry(args, args->tree->object.sha1, args->base,\n> -\t\t\t\tlen, 040777, NULL, 0);\n> +\t\t\t\t  len, 040777, NULL, NULL, 0);\n>  \t\tif (err)\n>  \t\t\treturn err;\n>  \t}\n>  \n>  \tcontext.args = args;\n>  \tcontext.write_entry = write_entry;\n> +\tcontext.stream_ok = stream_ok;\n>  \n>  \t/*\n>  \t * Setup index and instruct attr to read index only\n> diff --git a/archive.h b/archive.h\n> index 2b0884f..370cca9 100644\n> --- a/archive.h\n> +++ b/archive.h\n> @@ -27,9 +27,16 @@ extern void register_archiver(struct archiver *);\n>  extern void init_tar_archiver(void);\n>  extern void init_zip_archiver(void);\n>  \n> -typedef int (*write_archive_entry_fn_t)(struct archiver_args *args, const unsigned char *sha1, const char *path, size_t pathlen, unsigned int mode, void *buffer, unsigned long size);\n> +struct git_istream;\n> +typedef int (*write_archive_entry_fn_t)(struct archiver_args *args,\n> +\t\t\t\t\tconst unsigned char *sha1,\n> +\t\t\t\t\tconst char *path, size_t pathlen,\n> +\t\t\t\t\tunsigned int mode,\n> +\t\t\t\t\tvoid *buffer,\n> +\t\t\t\t\tstruct git_istream *stream,\n> +\t\t\t\t\tunsigned long size);\n>  \n> -extern int write_archive_entries(struct archiver_args *args, write_archive_entry_fn_t write_entry);\n> +extern int write_archive_entries(struct archiver_args *args, write_archive_entry_fn_t write_entry, int stream_ok);\n>  extern int write_archive(int argc, const char **argv, const char *prefix, int setup_prefix, const char *name_hint, int remote);\n>  \n>  const char *archive_format_from_filename(const char *filename);\n> diff --git a/t/t1050-large.sh b/t/t1050-large.sh\n> index 52acae5..5336eb8 100755\n> --- a/t/t1050-large.sh\n> +++ b/t/t1050-large.sh\n> @@ -151,7 +151,7 @@ test_expect_failure 'repack' '\n>  \tgit repack -ad\n>  '\n>  \n> -test_expect_failure 'tar achiving' '\n> +test_expect_success 'tar achiving' '\n>  \tgit archive --format=tar HEAD >/dev/null\n>  '\n"},{"id":"186163","messageId":"7vipiiwnmm.fsf@alter.siamese.dyndns.org","threadId":"29753","inReplyTo":"1330865996-2069-2-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH v2 01/10] Add more large blob test cases","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-03-06T00:59:13Z","receivedAt":"2012-03-06T00:59:13Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> diff --git a/wrapper.c b/wrapper.c\n> index 85f09df..d4c0972 100644\n> --- a/wrapper.c\n> +++ b/wrapper.c\n> @@ -9,6 +9,18 @@ static void do_nothing(size_t size)\n>  \n>  static void (*try_to_free_routine)(size_t size) = do_nothing;\n>  \n> +static void memory_limit_check(size_t size)\n> +{\n> +\tstatic int limit = -1;\n> +\tif (limit == -1) {\n> +\t\tconst char *env = getenv(\"GIT_ALLOC_LIMIT\");\n> +\t\tlimit = env ? atoi(env) * 1024 : 0;\n> +\t}\n> +\tif (limit && size > limit)\n> +\t\tdie(\"attempting to allocate %d over limit %d\",\n> +\t\t    size, limit);\n\nsize is size_t and %d calls for an int.\n\nI'll push out a fixed-up version later to 'pu'.\n"}]}