{"thread":{"id":"56672","subject":"[PATCH] unpack-objects: unpack large object in stream","startedAt":"2021-10-09T08:21:18Z","lastAt":"2022-07-01T02:01:11Z","messageCount":211,"participants":["Han Xin","Philip Oakley","Jiang Xin","Junio C Hamano","Derrick Stolee","Jeff King","Ævar Arnfjörð Bjarmason","René Scharfe","Neeraj Singh","Johannes Schindelin"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"438396","messageId":"20211009082058.41138-1-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":null,"subject":"[PATCH] unpack-objects: unpack large object in stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-10-09T08:20:58Z","receivedAt":"2021-10-09T08:21:18Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWhen calling \"unpack_non_delta_entry()\", will allocate full memory for\nthe whole size of the unpacked object and write the buffer to loose file\non disk. This may lead to OOM for the git-unpack-objects process when\nunpacking a very large object.\n\nIn function \"unpack_delta_entry()\", will also allocate full memory to\nbuffer the whole delta, but since there will be no delta for an object\nlarger than \"core.bigFileThreshold\", this issue is moderate.\n\nTo resolve the OOM issue in \"git-unpack-objects\", we can unpack large\nobject to file in stream, and use the setting of \"core.bigFileThreshold\" as\nthe threshold for large object.\n\nReviewed-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n builtin/unpack-objects.c          |  41 +++++++-\n object-file.c                     | 149 +++++++++++++++++++++++++++---\n object-store.h                    |   9 ++\n t/t5590-receive-unpack-objects.sh |  92 ++++++++++++++++++\n 4 files changed, 279 insertions(+), 12 deletions(-)\n create mode 100755 t/t5590-receive-unpack-objects.sh\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 4a9466295b..8ac77e60a8 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -320,11 +320,50 @@ static void added_object(unsigned nr, enum object_type type,\n \t}\n }\n \n+static void fill_stream(struct git_zstream *stream)\n+{\n+\tstream->next_in = fill(1);\n+\tstream->avail_in = len;\n+}\n+\n+static void use_stream(struct git_zstream *stream)\n+{\n+\tuse(len - stream->avail_in);\n+}\n+\n+static void write_stream_blob(unsigned nr, unsigned long size)\n+{\n+\tstruct git_zstream_reader reader;\n+\tstruct object_id *oid = &obj_list[nr].oid;\n+\n+\treader.fill = &fill_stream;\n+\treader.use = &use_stream;\n+\n+\tif (write_stream_object_file(&reader, size, type_name(OBJ_BLOB),\n+\t\t\t\t     oid, dry_run))\n+\t\tdie(\"failed to write object in stream\");\n+\tif (strict && !dry_run) {\n+\t\tstruct blob *blob = lookup_blob(the_repository, oid);\n+\t\tif (blob)\n+\t\t\tblob->object.flags |= FLAG_WRITTEN;\n+\t\telse\n+\t\t\tdie(\"invalid blob object from stream\");\n+\t}\n+\tobj_list[nr].obj = NULL;\n+}\n+\n static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n \t\t\t\t   unsigned nr)\n {\n-\tvoid *buf = get_data(size);\n+\tvoid *buf;\n+\n+\t/* Write large blob in stream without allocating full buffer. */\n+\tif (type == OBJ_BLOB && size > big_file_threshold) {\n+\t\twrite_stream_blob(nr, size);\n+\t\treturn;\n+\t}\n \n+\tbuf = get_data(size);\n \tif (!dry_run && buf)\n \t\twrite_object(nr, type, buf, size);\n \telse\ndiff --git a/object-file.c b/object-file.c\nindex a8be899481..06c1693675 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1913,6 +1913,28 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n \treturn fd;\n }\n \n+static int write_object_buffer(struct git_zstream *stream, git_hash_ctx *c,\n+\t\t\t       int fd, unsigned char *compressed,\n+\t\t\t       int compressed_len, const void *buf,\n+\t\t\t       size_t len, int flush)\n+{\n+\tint ret;\n+\n+\tstream->next_in = (void *)buf;\n+\tstream->avail_in = len;\n+\tdo {\n+\t\tunsigned char *in0 = stream->next_in;\n+\t\tret = git_deflate(stream, flush);\n+\t\tthe_hash_algo->update_fn(c, in0, stream->next_in - in0);\n+\t\tif (write_buffer(fd, compressed, stream->next_out - compressed) < 0)\n+\t\t\tdie(_(\"unable to write loose object file\"));\n+\t\tstream->next_out = compressed;\n+\t\tstream->avail_out = compressed_len;\n+\t} while (ret == Z_OK);\n+\n+\treturn ret;\n+}\n+\n static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\t\t      int hdrlen, const void *buf, unsigned long len,\n \t\t\t      time_t mtime)\n@@ -1949,17 +1971,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n \n \t/* Then the data itself.. */\n-\tstream.next_in = (void *)buf;\n-\tstream.avail_in = len;\n-\tdo {\n-\t\tunsigned char *in0 = stream.next_in;\n-\t\tret = git_deflate(&stream, Z_FINISH);\n-\t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n-\t\tif (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n-\t\t\tdie(_(\"unable to write loose object file\"));\n-\t\tstream.next_out = compressed;\n-\t\tstream.avail_out = sizeof(compressed);\n-\t} while (ret == Z_OK);\n+\tret = write_object_buffer(&stream, &c, fd, compressed,\n+\t\t\t\t  sizeof(compressed), buf, len,\n+\t\t\t\t  Z_FINISH);\n \n \tif (ret != Z_STREAM_END)\n \t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n@@ -2020,6 +2034,119 @@ int write_object_file(const void *buf, unsigned long len, const char *type,\n \treturn write_loose_object(oid, hdr, hdrlen, buf, len, 0);\n }\n \n+int write_stream_object_file(struct git_zstream_reader *reader,\n+\t\t\t     unsigned long len, const char *type,\n+\t\t\t     struct object_id *oid,\n+\t\t\t     int dry_run)\n+{\n+\tgit_zstream istream, ostream;\n+\tunsigned char buf[8192], compressed[4096];\n+\tchar hdr[MAX_HEADER_LEN];\n+\tint istatus, ostatus, fd = 0, hdrlen, dirlen, flush = 0;\n+\tint ret = 0;\n+\tgit_hash_ctx c;\n+\tstruct strbuf tmp_file = STRBUF_INIT;\n+\tstruct strbuf filename = STRBUF_INIT;\n+\n+\t/* Write tmpfile in objects dir, because oid is unknown */\n+\tif (!dry_run) {\n+\t\tstrbuf_addstr(&filename, the_repository->objects->odb->path);\n+\t\tstrbuf_addch(&filename, '/');\n+\t\tfd = create_tmpfile(&tmp_file, filename.buf);\n+\t\tif (fd < 0) {\n+\t\t\tif (errno == EACCES)\n+\t\t\t\tret = error(_(\"insufficient permission for adding an object to repository database %s\"),\n+\t\t\t\t\tget_object_directory());\n+\t\t\telse\n+\t\t\t\tret = error_errno(_(\"unable to create temporary file\"));\n+\t\t\tgoto cleanup;\n+\t\t}\n+\t}\n+\n+\tmemset(&istream, 0, sizeof(istream));\n+\tistream.next_out = buf;\n+\tistream.avail_out = sizeof(buf);\n+\tgit_inflate_init(&istream);\n+\n+\tif (!dry_run) {\n+\t\t/* Set it up */\n+\t\tgit_deflate_init(&ostream, zlib_compression_level);\n+\t\tostream.next_out = compressed;\n+\t\tostream.avail_out = sizeof(compressed);\n+\t\tthe_hash_algo->init_fn(&c);\n+\n+\t\t/* First header */\n+\t\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\" PRIuMAX, type,\n+\t\t\t\t(uintmax_t)len) + 1;\n+\t\tostream.next_in = (unsigned char *)hdr;\n+\t\tostream.avail_in = hdrlen;\n+\t\twhile (git_deflate(&ostream, 0) == Z_OK)\n+\t\t\t; /* nothing */\n+\t\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n+\t}\n+\n+\t/* Then the data itself */\n+\tdo {\n+\t\tunsigned char *last_out = istream.next_out;\n+\t\treader->fill(&istream);\n+\t\tistatus = git_inflate(&istream, 0);\n+\t\tif (istatus == Z_STREAM_END)\n+\t\t\tflush = Z_FINISH;\n+\t\treader->use(&istream);\n+\t\tif (!dry_run)\n+\t\t\tostatus = write_object_buffer(&ostream, &c, fd, compressed,\n+\t\t\t\t\t\t      sizeof(compressed), last_out,\n+\t\t\t\t\t\t      istream.next_out - last_out,\n+\t\t\t\t\t\t      flush);\n+\t\tistream.next_out = buf;\n+\t\tistream.avail_out = sizeof(buf);\n+\t} while (istatus == Z_OK);\n+\n+\tif (istream.total_out != len || istatus != Z_STREAM_END)\n+\t\tdie( _(\"inflate returned %d\"), istatus);\n+\tgit_inflate_end(&istream);\n+\n+\tif (dry_run)\n+\t\tgoto cleanup;\n+\n+\tif (ostatus != Z_STREAM_END)\n+\t\tdie(_(\"unable to deflate new object (%d)\"), ostatus);\n+\tostatus = git_deflate_end_gently(&ostream);\n+\tif (ostatus != Z_OK)\n+\t\tdie(_(\"deflateEnd on object failed (%d)\"), ostatus);\n+\tthe_hash_algo->final_fn(oid->hash, &c);\n+\tclose_loose_object(fd);\n+\n+\t/* We get the oid now */\n+\tloose_object_path(the_repository, &filename, oid);\n+\n+\tdirlen = directory_size(filename.buf);\n+\tif (dirlen) {\n+\t\tstruct strbuf dir = STRBUF_INIT;\n+\t\t/*\n+\t\t * Make sure the directory exists; note that the contents\n+\t\t * of the buffer are undefined after mkstemp returns an\n+\t\t * error, so we have to rewrite the whole buffer from\n+\t\t * scratch.\n+\t\t */\n+\t\tstrbuf_add(&dir, filename.buf, dirlen - 1);\n+\t\tif (mkdir(dir.buf, 0777) && errno != EEXIST) {\n+\t\t\tunlink_or_warn(tmp_file.buf);\n+\t\t\tstrbuf_release(&dir);\n+\t\t\tret = -1;\n+\t\t\tgoto cleanup;\n+\t\t}\n+\t\tstrbuf_release(&dir);\n+\t}\n+\n+\tret = finalize_object_file(tmp_file.buf, filename.buf);\n+\n+cleanup:\n+\tstrbuf_release(&tmp_file);\n+\tstrbuf_release(&filename);\n+\treturn ret;\n+}\n+\n int hash_object_file_literally(const void *buf, unsigned long len,\n \t\t\t       const char *type, struct object_id *oid,\n \t\t\t       unsigned flags)\ndiff --git a/object-store.h b/object-store.h\nindex d24915ced1..12b113ef93 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -33,6 +33,11 @@ struct object_directory {\n \tchar *path;\n };\n \n+struct git_zstream_reader {\n+\tvoid (*fill)(struct git_zstream *);\n+\tvoid (*use)(struct git_zstream *);\n+};\n+\n KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n \tstruct object_directory *, 1, fspathhash, fspatheq)\n \n@@ -225,6 +230,10 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n int write_object_file(const void *buf, unsigned long len,\n \t\t      const char *type, struct object_id *oid);\n \n+int write_stream_object_file(struct git_zstream_reader *reader,\n+\t\t\t     unsigned long len, const char *type,\n+\t\t\t     struct object_id *oid, int dry_run);\n+\n int hash_object_file_literally(const void *buf, unsigned long len,\n \t\t\t       const char *type, struct object_id *oid,\n \t\t\t       unsigned flags);\ndiff --git a/t/t5590-receive-unpack-objects.sh b/t/t5590-receive-unpack-objects.sh\nnew file mode 100755\nindex 0000000000..7e63dfc0db\n--- /dev/null\n+++ b/t/t5590-receive-unpack-objects.sh\n@@ -0,0 +1,92 @@\n+#!/bin/sh\n+#\n+# Copyright (c) 2021 Han Xin\n+#\n+\n+test_description='Test unpack-objects when receive pack'\n+\n+GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n+export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n+\n+. ./test-lib.sh\n+\n+test_expect_success \"create commit with big blobs (1.5 MB)\" '\n+\ttest-tool genrandom foo 1500000 >big-blob &&\n+\ttest_commit --append foo big-blob &&\n+\ttest-tool genrandom bar 1500000 >big-blob &&\n+\ttest_commit --append bar big-blob &&\n+\t(\n+\t\tcd .git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >expect &&\n+\tgit repack -ad\n+'\n+\n+test_expect_success 'setup GIT_ALLOC_LIMIT to 1MB' '\n+\tGIT_ALLOC_LIMIT=1m &&\n+\texport GIT_ALLOC_LIMIT\n+'\n+\n+test_expect_success 'prepare dest repository' '\n+\tgit init --bare dest.git &&\n+\tgit -C dest.git config core.bigFileThreshold 2m &&\n+\tgit -C dest.git config receive.unpacklimit 100\n+'\n+\n+test_expect_success 'fail to push: cannot allocate' '\n+\ttest_must_fail git push dest.git HEAD 2>err &&\n+\ttest_i18ngrep \"remote: fatal: attempting to allocate\" err &&\n+\t(\n+\t\tcd dest.git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >actual &&\n+\t! test_cmp expect actual\n+'\n+\n+test_expect_success 'set a lower bigfile threshold' '\n+\tgit -C dest.git config core.bigFileThreshold 1m\n+'\n+\n+test_expect_success 'unpack big object in stream' '\n+\tgit push dest.git HEAD &&\n+\tgit -C dest.git fsck &&\n+\t(\n+\t\tcd dest.git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'setup for unpack-objects dry-run test' '\n+\tPACK=$(echo main | git pack-objects --progress --revs test) &&\n+\tunset GIT_ALLOC_LIMIT &&\n+\tgit init --bare unpack-test.git\n+'\n+\n+test_expect_success 'unpack-objects dry-run with large threshold' '\n+\t(\n+\t\tcd unpack-test.git &&\n+\t\tgit config core.bigFileThreshold 2m &&\n+\t\tgit unpack-objects -n <../test-$PACK.pack\n+\t) &&\n+\t(\n+\t\tcd unpack-test.git &&\n+\t\tfind objects/ -type f\n+\t) >actual &&\n+\ttest_must_be_empty actual\n+'\n+\n+test_expect_success 'unpack-objects dry-run with small threshold' '\n+\t(\n+\t\tcd unpack-test.git &&\n+\t\tgit config core.bigFileThreshold 1m &&\n+\t\tgit unpack-objects -n <../test-$PACK.pack\n+\t) &&\n+\t(\n+\t\tcd unpack-test.git &&\n+\t\tfind objects/ -type f\n+\t) >actual &&\n+\ttest_must_be_empty actual\n+'\n+\n+test_done\n-- \n2.33.0.1.g09a6bb964f.dirty\n\n"},{"id":"439017","messageId":"CAO0brD2J1zEP5=A1uuwvLsFMGZ9X6WaZuhEYZt7AHhzzKoJX5Q@mail.gmail.com","threadId":"56672","inReplyTo":"20211009082058.41138-1-chiyutianyi@gmail.com","subject":"Re: [PATCH] unpack-objects: unpack large object in stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-10-19T07:37:30Z","receivedAt":"2021-10-19T07:37:45Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"Any suggestions?\n\nHan Xin <chiyutianyi@gmail.com> 于2021年10月9日周六 下午4:21写道：\n>\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> When calling \"unpack_non_delta_entry()\", will allocate full memory for\n> the whole size of the unpacked object and write the buffer to loose file\n> on disk. This may lead to OOM for the git-unpack-objects process when\n> unpacking a very large object.\n>\n> In function \"unpack_delta_entry()\", will also allocate full memory to\n> buffer the whole delta, but since there will be no delta for an object\n> larger than \"core.bigFileThreshold\", this issue is moderate.\n>\n> To resolve the OOM issue in \"git-unpack-objects\", we can unpack large\n> object to file in stream, and use the setting of \"core.bigFileThreshold\" as\n> the threshold for large object.\n>\n> Reviewed-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  builtin/unpack-objects.c          |  41 +++++++-\n>  object-file.c                     | 149 +++++++++++++++++++++++++++---\n>  object-store.h                    |   9 ++\n>  t/t5590-receive-unpack-objects.sh |  92 ++++++++++++++++++\n>  4 files changed, 279 insertions(+), 12 deletions(-)\n>  create mode 100755 t/t5590-receive-unpack-objects.sh\n>\n> diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n> index 4a9466295b..8ac77e60a8 100644\n> --- a/builtin/unpack-objects.c\n> +++ b/builtin/unpack-objects.c\n> @@ -320,11 +320,50 @@ static void added_object(unsigned nr, enum object_type type,\n>         }\n>  }\n>\n> +static void fill_stream(struct git_zstream *stream)\n> +{\n> +       stream->next_in = fill(1);\n> +       stream->avail_in = len;\n> +}\n> +\n> +static void use_stream(struct git_zstream *stream)\n> +{\n> +       use(len - stream->avail_in);\n> +}\n> +\n> +static void write_stream_blob(unsigned nr, unsigned long size)\n> +{\n> +       struct git_zstream_reader reader;\n> +       struct object_id *oid = &obj_list[nr].oid;\n> +\n> +       reader.fill = &fill_stream;\n> +       reader.use = &use_stream;\n> +\n> +       if (write_stream_object_file(&reader, size, type_name(OBJ_BLOB),\n> +                                    oid, dry_run))\n> +               die(\"failed to write object in stream\");\n> +       if (strict && !dry_run) {\n> +               struct blob *blob = lookup_blob(the_repository, oid);\n> +               if (blob)\n> +                       blob->object.flags |= FLAG_WRITTEN;\n> +               else\n> +                       die(\"invalid blob object from stream\");\n> +       }\n> +       obj_list[nr].obj = NULL;\n> +}\n> +\n>  static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n>                                    unsigned nr)\n>  {\n> -       void *buf = get_data(size);\n> +       void *buf;\n> +\n> +       /* Write large blob in stream without allocating full buffer. */\n> +       if (type == OBJ_BLOB && size > big_file_threshold) {\n> +               write_stream_blob(nr, size);\n> +               return;\n> +       }\n>\n> +       buf = get_data(size);\n>         if (!dry_run && buf)\n>                 write_object(nr, type, buf, size);\n>         else\n> diff --git a/object-file.c b/object-file.c\n> index a8be899481..06c1693675 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1913,6 +1913,28 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n>         return fd;\n>  }\n>\n> +static int write_object_buffer(struct git_zstream *stream, git_hash_ctx *c,\n> +                              int fd, unsigned char *compressed,\n> +                              int compressed_len, const void *buf,\n> +                              size_t len, int flush)\n> +{\n> +       int ret;\n> +\n> +       stream->next_in = (void *)buf;\n> +       stream->avail_in = len;\n> +       do {\n> +               unsigned char *in0 = stream->next_in;\n> +               ret = git_deflate(stream, flush);\n> +               the_hash_algo->update_fn(c, in0, stream->next_in - in0);\n> +               if (write_buffer(fd, compressed, stream->next_out - compressed) < 0)\n> +                       die(_(\"unable to write loose object file\"));\n> +               stream->next_out = compressed;\n> +               stream->avail_out = compressed_len;\n> +       } while (ret == Z_OK);\n> +\n> +       return ret;\n> +}\n> +\n>  static int write_loose_object(const struct object_id *oid, char *hdr,\n>                               int hdrlen, const void *buf, unsigned long len,\n>                               time_t mtime)\n> @@ -1949,17 +1971,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>         the_hash_algo->update_fn(&c, hdr, hdrlen);\n>\n>         /* Then the data itself.. */\n> -       stream.next_in = (void *)buf;\n> -       stream.avail_in = len;\n> -       do {\n> -               unsigned char *in0 = stream.next_in;\n> -               ret = git_deflate(&stream, Z_FINISH);\n> -               the_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n> -               if (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n> -                       die(_(\"unable to write loose object file\"));\n> -               stream.next_out = compressed;\n> -               stream.avail_out = sizeof(compressed);\n> -       } while (ret == Z_OK);\n> +       ret = write_object_buffer(&stream, &c, fd, compressed,\n> +                                 sizeof(compressed), buf, len,\n> +                                 Z_FINISH);\n>\n>         if (ret != Z_STREAM_END)\n>                 die(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n> @@ -2020,6 +2034,119 @@ int write_object_file(const void *buf, unsigned long len, const char *type,\n>         return write_loose_object(oid, hdr, hdrlen, buf, len, 0);\n>  }\n>\n> +int write_stream_object_file(struct git_zstream_reader *reader,\n> +                            unsigned long len, const char *type,\n> +                            struct object_id *oid,\n> +                            int dry_run)\n> +{\n> +       git_zstream istream, ostream;\n> +       unsigned char buf[8192], compressed[4096];\n> +       char hdr[MAX_HEADER_LEN];\n> +       int istatus, ostatus, fd = 0, hdrlen, dirlen, flush = 0;\n> +       int ret = 0;\n> +       git_hash_ctx c;\n> +       struct strbuf tmp_file = STRBUF_INIT;\n> +       struct strbuf filename = STRBUF_INIT;\n> +\n> +       /* Write tmpfile in objects dir, because oid is unknown */\n> +       if (!dry_run) {\n> +               strbuf_addstr(&filename, the_repository->objects->odb->path);\n> +               strbuf_addch(&filename, '/');\n> +               fd = create_tmpfile(&tmp_file, filename.buf);\n> +               if (fd < 0) {\n> +                       if (errno == EACCES)\n> +                               ret = error(_(\"insufficient permission for adding an object to repository database %s\"),\n> +                                       get_object_directory());\n> +                       else\n> +                               ret = error_errno(_(\"unable to create temporary file\"));\n> +                       goto cleanup;\n> +               }\n> +       }\n> +\n> +       memset(&istream, 0, sizeof(istream));\n> +       istream.next_out = buf;\n> +       istream.avail_out = sizeof(buf);\n> +       git_inflate_init(&istream);\n> +\n> +       if (!dry_run) {\n> +               /* Set it up */\n> +               git_deflate_init(&ostream, zlib_compression_level);\n> +               ostream.next_out = compressed;\n> +               ostream.avail_out = sizeof(compressed);\n> +               the_hash_algo->init_fn(&c);\n> +\n> +               /* First header */\n> +               hdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\" PRIuMAX, type,\n> +                               (uintmax_t)len) + 1;\n> +               ostream.next_in = (unsigned char *)hdr;\n> +               ostream.avail_in = hdrlen;\n> +               while (git_deflate(&ostream, 0) == Z_OK)\n> +                       ; /* nothing */\n> +               the_hash_algo->update_fn(&c, hdr, hdrlen);\n> +       }\n> +\n> +       /* Then the data itself */\n> +       do {\n> +               unsigned char *last_out = istream.next_out;\n> +               reader->fill(&istream);\n> +               istatus = git_inflate(&istream, 0);\n> +               if (istatus == Z_STREAM_END)\n> +                       flush = Z_FINISH;\n> +               reader->use(&istream);\n> +               if (!dry_run)\n> +                       ostatus = write_object_buffer(&ostream, &c, fd, compressed,\n> +                                                     sizeof(compressed), last_out,\n> +                                                     istream.next_out - last_out,\n> +                                                     flush);\n> +               istream.next_out = buf;\n> +               istream.avail_out = sizeof(buf);\n> +       } while (istatus == Z_OK);\n> +\n> +       if (istream.total_out != len || istatus != Z_STREAM_END)\n> +               die( _(\"inflate returned %d\"), istatus);\n> +       git_inflate_end(&istream);\n> +\n> +       if (dry_run)\n> +               goto cleanup;\n> +\n> +       if (ostatus != Z_STREAM_END)\n> +               die(_(\"unable to deflate new object (%d)\"), ostatus);\n> +       ostatus = git_deflate_end_gently(&ostream);\n> +       if (ostatus != Z_OK)\n> +               die(_(\"deflateEnd on object failed (%d)\"), ostatus);\n> +       the_hash_algo->final_fn(oid->hash, &c);\n> +       close_loose_object(fd);\n> +\n> +       /* We get the oid now */\n> +       loose_object_path(the_repository, &filename, oid);\n> +\n> +       dirlen = directory_size(filename.buf);\n> +       if (dirlen) {\n> +               struct strbuf dir = STRBUF_INIT;\n> +               /*\n> +                * Make sure the directory exists; note that the contents\n> +                * of the buffer are undefined after mkstemp returns an\n> +                * error, so we have to rewrite the whole buffer from\n> +                * scratch.\n> +                */\n> +               strbuf_add(&dir, filename.buf, dirlen - 1);\n> +               if (mkdir(dir.buf, 0777) && errno != EEXIST) {\n> +                       unlink_or_warn(tmp_file.buf);\n> +                       strbuf_release(&dir);\n> +                       ret = -1;\n> +                       goto cleanup;\n> +               }\n> +               strbuf_release(&dir);\n> +       }\n> +\n> +       ret = finalize_object_file(tmp_file.buf, filename.buf);\n> +\n> +cleanup:\n> +       strbuf_release(&tmp_file);\n> +       strbuf_release(&filename);\n> +       return ret;\n> +}\n> +\n>  int hash_object_file_literally(const void *buf, unsigned long len,\n>                                const char *type, struct object_id *oid,\n>                                unsigned flags)\n> diff --git a/object-store.h b/object-store.h\n> index d24915ced1..12b113ef93 100644\n> --- a/object-store.h\n> +++ b/object-store.h\n> @@ -33,6 +33,11 @@ struct object_directory {\n>         char *path;\n>  };\n>\n> +struct git_zstream_reader {\n> +       void (*fill)(struct git_zstream *);\n> +       void (*use)(struct git_zstream *);\n> +};\n> +\n>  KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n>         struct object_directory *, 1, fspathhash, fspatheq)\n>\n> @@ -225,6 +230,10 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n>  int write_object_file(const void *buf, unsigned long len,\n>                       const char *type, struct object_id *oid);\n>\n> +int write_stream_object_file(struct git_zstream_reader *reader,\n> +                            unsigned long len, const char *type,\n> +                            struct object_id *oid, int dry_run);\n> +\n>  int hash_object_file_literally(const void *buf, unsigned long len,\n>                                const char *type, struct object_id *oid,\n>                                unsigned flags);\n> diff --git a/t/t5590-receive-unpack-objects.sh b/t/t5590-receive-unpack-objects.sh\n> new file mode 100755\n> index 0000000000..7e63dfc0db\n> --- /dev/null\n> +++ b/t/t5590-receive-unpack-objects.sh\n> @@ -0,0 +1,92 @@\n> +#!/bin/sh\n> +#\n> +# Copyright (c) 2021 Han Xin\n> +#\n> +\n> +test_description='Test unpack-objects when receive pack'\n> +\n> +GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n> +export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n> +\n> +. ./test-lib.sh\n> +\n> +test_expect_success \"create commit with big blobs (1.5 MB)\" '\n> +       test-tool genrandom foo 1500000 >big-blob &&\n> +       test_commit --append foo big-blob &&\n> +       test-tool genrandom bar 1500000 >big-blob &&\n> +       test_commit --append bar big-blob &&\n> +       (\n> +               cd .git &&\n> +               find objects/?? -type f | sort\n> +       ) >expect &&\n> +       git repack -ad\n> +'\n> +\n> +test_expect_success 'setup GIT_ALLOC_LIMIT to 1MB' '\n> +       GIT_ALLOC_LIMIT=1m &&\n> +       export GIT_ALLOC_LIMIT\n> +'\n> +\n> +test_expect_success 'prepare dest repository' '\n> +       git init --bare dest.git &&\n> +       git -C dest.git config core.bigFileThreshold 2m &&\n> +       git -C dest.git config receive.unpacklimit 100\n> +'\n> +\n> +test_expect_success 'fail to push: cannot allocate' '\n> +       test_must_fail git push dest.git HEAD 2>err &&\n> +       test_i18ngrep \"remote: fatal: attempting to allocate\" err &&\n> +       (\n> +               cd dest.git &&\n> +               find objects/?? -type f | sort\n> +       ) >actual &&\n> +       ! test_cmp expect actual\n> +'\n> +\n> +test_expect_success 'set a lower bigfile threshold' '\n> +       git -C dest.git config core.bigFileThreshold 1m\n> +'\n> +\n> +test_expect_success 'unpack big object in stream' '\n> +       git push dest.git HEAD &&\n> +       git -C dest.git fsck &&\n> +       (\n> +               cd dest.git &&\n> +               find objects/?? -type f | sort\n> +       ) >actual &&\n> +       test_cmp expect actual\n> +'\n> +\n> +test_expect_success 'setup for unpack-objects dry-run test' '\n> +       PACK=$(echo main | git pack-objects --progress --revs test) &&\n> +       unset GIT_ALLOC_LIMIT &&\n> +       git init --bare unpack-test.git\n> +'\n> +\n> +test_expect_success 'unpack-objects dry-run with large threshold' '\n> +       (\n> +               cd unpack-test.git &&\n> +               git config core.bigFileThreshold 2m &&\n> +               git unpack-objects -n <../test-$PACK.pack\n> +       ) &&\n> +       (\n> +               cd unpack-test.git &&\n> +               find objects/ -type f\n> +       ) >actual &&\n> +       test_must_be_empty actual\n> +'\n> +\n> +test_expect_success 'unpack-objects dry-run with small threshold' '\n> +       (\n> +               cd unpack-test.git &&\n> +               git config core.bigFileThreshold 1m &&\n> +               git unpack-objects -n <../test-$PACK.pack\n> +       ) &&\n> +       (\n> +               cd unpack-test.git &&\n> +               find objects/ -type f\n> +       ) >actual &&\n> +       test_must_be_empty actual\n> +'\n> +\n> +test_done\n> --\n> 2.33.0.1.g09a6bb964f.dirty\n>\n"},{"id":"439111","messageId":"aeb1c7cb-7e01-8a21-5f7a-d27bd43d7604@iee.email","threadId":"56672","inReplyTo":"20211009082058.41138-1-chiyutianyi@gmail.com","subject":"Re: [PATCH] unpack-objects: unpack large object in stream","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2021-10-20T14:42:58Z","receivedAt":"2021-10-20T14:43:03Z","isPatch":true,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"On 09/10/2021 09:20, Han Xin wrote:\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> When calling \"unpack_non_delta_entry()\", will allocate full memory for\n> the whole size of the unpacked object and write the buffer to loose file\n> on disk. This may lead to OOM for the git-unpack-objects process when\n> unpacking a very large object.\n>\n> In function \"unpack_delta_entry()\", will also allocate full memory to\n> buffer the whole delta, but since there will be no delta for an object\n> larger than \"core.bigFileThreshold\", this issue is moderate.\n>\n> To resolve the OOM issue in \"git-unpack-objects\", we can unpack large\n> object to file in stream, and use the setting of \"core.bigFileThreshold\" as\n> the threshold for large object.\n>\n> Reviewed-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  builtin/unpack-objects.c          |  41 +++++++-\n>  object-file.c                     | 149 +++++++++++++++++++++++++++---\n>  object-store.h                    |   9 ++\n>  t/t5590-receive-unpack-objects.sh |  92 ++++++++++++++++++\n>  4 files changed, 279 insertions(+), 12 deletions(-)\n>  create mode 100755 t/t5590-receive-unpack-objects.sh\n>\n> diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n> index 4a9466295b..8ac77e60a8 100644\n> --- a/builtin/unpack-objects.c\n> +++ b/builtin/unpack-objects.c\n> @@ -320,11 +320,50 @@ static void added_object(unsigned nr, enum object_type type,\n>  \t}\n>  }\n>  \n> +static void fill_stream(struct git_zstream *stream)\n> +{\n> +\tstream->next_in = fill(1);\n> +\tstream->avail_in = len;\n> +}\n> +\n> +static void use_stream(struct git_zstream *stream)\n> +{\n> +\tuse(len - stream->avail_in);\n> +}\n> +\n> +static void write_stream_blob(unsigned nr, unsigned long size)\n\nCan we use size_t for the `size`, and possibly `nr`, to improve\ncompatibility with Windows systems where unsigned long is only 32 bits?\n\nThere has been some work in the past on providing large file support on\nWindows, which requires numerous long -> size_t changes.\n\nPhilip\n> +{\n> +\tstruct git_zstream_reader reader;\n> +\tstruct object_id *oid = &obj_list[nr].oid;\n> +\n> +\treader.fill = &fill_stream;\n> +\treader.use = &use_stream;\n> +\n> +\tif (write_stream_object_file(&reader, size, type_name(OBJ_BLOB),\n> +\t\t\t\t     oid, dry_run))\n> +\t\tdie(\"failed to write object in stream\");\n> +\tif (strict && !dry_run) {\n> +\t\tstruct blob *blob = lookup_blob(the_repository, oid);\n> +\t\tif (blob)\n> +\t\t\tblob->object.flags |= FLAG_WRITTEN;\n> +\t\telse\n> +\t\t\tdie(\"invalid blob object from stream\");\n> +\t}\n> +\tobj_list[nr].obj = NULL;\n> +}\n> +\n>  static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n>  \t\t\t\t   unsigned nr)\n>  {\n> -\tvoid *buf = get_data(size);\n> +\tvoid *buf;\n> +\n> +\t/* Write large blob in stream without allocating full buffer. */\n> +\tif (type == OBJ_BLOB && size > big_file_threshold) {\n> +\t\twrite_stream_blob(nr, size);\n> +\t\treturn;\n> +\t}\n>  \n> +\tbuf = get_data(size);\n>  \tif (!dry_run && buf)\n>  \t\twrite_object(nr, type, buf, size);\n>  \telse\n> diff --git a/object-file.c b/object-file.c\n> index a8be899481..06c1693675 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1913,6 +1913,28 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n>  \treturn fd;\n>  }\n>  \n> +static int write_object_buffer(struct git_zstream *stream, git_hash_ctx *c,\n> +\t\t\t       int fd, unsigned char *compressed,\n> +\t\t\t       int compressed_len, const void *buf,\n> +\t\t\t       size_t len, int flush)\n> +{\n> +\tint ret;\n> +\n> +\tstream->next_in = (void *)buf;\n> +\tstream->avail_in = len;\n> +\tdo {\n> +\t\tunsigned char *in0 = stream->next_in;\n> +\t\tret = git_deflate(stream, flush);\n> +\t\tthe_hash_algo->update_fn(c, in0, stream->next_in - in0);\n> +\t\tif (write_buffer(fd, compressed, stream->next_out - compressed) < 0)\n> +\t\t\tdie(_(\"unable to write loose object file\"));\n> +\t\tstream->next_out = compressed;\n> +\t\tstream->avail_out = compressed_len;\n> +\t} while (ret == Z_OK);\n> +\n> +\treturn ret;\n> +}\n> +\n>  static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \t\t\t      int hdrlen, const void *buf, unsigned long len,\n>  \t\t\t      time_t mtime)\n> @@ -1949,17 +1971,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n>  \n>  \t/* Then the data itself.. */\n> -\tstream.next_in = (void *)buf;\n> -\tstream.avail_in = len;\n> -\tdo {\n> -\t\tunsigned char *in0 = stream.next_in;\n> -\t\tret = git_deflate(&stream, Z_FINISH);\n> -\t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n> -\t\tif (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n> -\t\t\tdie(_(\"unable to write loose object file\"));\n> -\t\tstream.next_out = compressed;\n> -\t\tstream.avail_out = sizeof(compressed);\n> -\t} while (ret == Z_OK);\n> +\tret = write_object_buffer(&stream, &c, fd, compressed,\n> +\t\t\t\t  sizeof(compressed), buf, len,\n> +\t\t\t\t  Z_FINISH);\n>  \n>  \tif (ret != Z_STREAM_END)\n>  \t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n> @@ -2020,6 +2034,119 @@ int write_object_file(const void *buf, unsigned long len, const char *type,\n>  \treturn write_loose_object(oid, hdr, hdrlen, buf, len, 0);\n>  }\n>  \n> +int write_stream_object_file(struct git_zstream_reader *reader,\n> +\t\t\t     unsigned long len, const char *type,\n> +\t\t\t     struct object_id *oid,\n> +\t\t\t     int dry_run)\n> +{\n> +\tgit_zstream istream, ostream;\n> +\tunsigned char buf[8192], compressed[4096];\n> +\tchar hdr[MAX_HEADER_LEN];\n> +\tint istatus, ostatus, fd = 0, hdrlen, dirlen, flush = 0;\n> +\tint ret = 0;\n> +\tgit_hash_ctx c;\n> +\tstruct strbuf tmp_file = STRBUF_INIT;\n> +\tstruct strbuf filename = STRBUF_INIT;\n> +\n> +\t/* Write tmpfile in objects dir, because oid is unknown */\n> +\tif (!dry_run) {\n> +\t\tstrbuf_addstr(&filename, the_repository->objects->odb->path);\n> +\t\tstrbuf_addch(&filename, '/');\n> +\t\tfd = create_tmpfile(&tmp_file, filename.buf);\n> +\t\tif (fd < 0) {\n> +\t\t\tif (errno == EACCES)\n> +\t\t\t\tret = error(_(\"insufficient permission for adding an object to repository database %s\"),\n> +\t\t\t\t\tget_object_directory());\n> +\t\t\telse\n> +\t\t\t\tret = error_errno(_(\"unable to create temporary file\"));\n> +\t\t\tgoto cleanup;\n> +\t\t}\n> +\t}\n> +\n> +\tmemset(&istream, 0, sizeof(istream));\n> +\tistream.next_out = buf;\n> +\tistream.avail_out = sizeof(buf);\n> +\tgit_inflate_init(&istream);\n> +\n> +\tif (!dry_run) {\n> +\t\t/* Set it up */\n> +\t\tgit_deflate_init(&ostream, zlib_compression_level);\n> +\t\tostream.next_out = compressed;\n> +\t\tostream.avail_out = sizeof(compressed);\n> +\t\tthe_hash_algo->init_fn(&c);\n> +\n> +\t\t/* First header */\n> +\t\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\" PRIuMAX, type,\n> +\t\t\t\t(uintmax_t)len) + 1;\n> +\t\tostream.next_in = (unsigned char *)hdr;\n> +\t\tostream.avail_in = hdrlen;\n> +\t\twhile (git_deflate(&ostream, 0) == Z_OK)\n> +\t\t\t; /* nothing */\n> +\t\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n> +\t}\n> +\n> +\t/* Then the data itself */\n> +\tdo {\n> +\t\tunsigned char *last_out = istream.next_out;\n> +\t\treader->fill(&istream);\n> +\t\tistatus = git_inflate(&istream, 0);\n> +\t\tif (istatus == Z_STREAM_END)\n> +\t\t\tflush = Z_FINISH;\n> +\t\treader->use(&istream);\n> +\t\tif (!dry_run)\n> +\t\t\tostatus = write_object_buffer(&ostream, &c, fd, compressed,\n> +\t\t\t\t\t\t      sizeof(compressed), last_out,\n> +\t\t\t\t\t\t      istream.next_out - last_out,\n> +\t\t\t\t\t\t      flush);\n> +\t\tistream.next_out = buf;\n> +\t\tistream.avail_out = sizeof(buf);\n> +\t} while (istatus == Z_OK);\n> +\n> +\tif (istream.total_out != len || istatus != Z_STREAM_END)\n> +\t\tdie( _(\"inflate returned %d\"), istatus);\n> +\tgit_inflate_end(&istream);\n> +\n> +\tif (dry_run)\n> +\t\tgoto cleanup;\n> +\n> +\tif (ostatus != Z_STREAM_END)\n> +\t\tdie(_(\"unable to deflate new object (%d)\"), ostatus);\n> +\tostatus = git_deflate_end_gently(&ostream);\n> +\tif (ostatus != Z_OK)\n> +\t\tdie(_(\"deflateEnd on object failed (%d)\"), ostatus);\n> +\tthe_hash_algo->final_fn(oid->hash, &c);\n> +\tclose_loose_object(fd);\n> +\n> +\t/* We get the oid now */\n> +\tloose_object_path(the_repository, &filename, oid);\n> +\n> +\tdirlen = directory_size(filename.buf);\n> +\tif (dirlen) {\n> +\t\tstruct strbuf dir = STRBUF_INIT;\n> +\t\t/*\n> +\t\t * Make sure the directory exists; note that the contents\n> +\t\t * of the buffer are undefined after mkstemp returns an\n> +\t\t * error, so we have to rewrite the whole buffer from\n> +\t\t * scratch.\n> +\t\t */\n> +\t\tstrbuf_add(&dir, filename.buf, dirlen - 1);\n> +\t\tif (mkdir(dir.buf, 0777) && errno != EEXIST) {\n> +\t\t\tunlink_or_warn(tmp_file.buf);\n> +\t\t\tstrbuf_release(&dir);\n> +\t\t\tret = -1;\n> +\t\t\tgoto cleanup;\n> +\t\t}\n> +\t\tstrbuf_release(&dir);\n> +\t}\n> +\n> +\tret = finalize_object_file(tmp_file.buf, filename.buf);\n> +\n> +cleanup:\n> +\tstrbuf_release(&tmp_file);\n> +\tstrbuf_release(&filename);\n> +\treturn ret;\n> +}\n> +\n>  int hash_object_file_literally(const void *buf, unsigned long len,\n>  \t\t\t       const char *type, struct object_id *oid,\n>  \t\t\t       unsigned flags)\n> diff --git a/object-store.h b/object-store.h\n> index d24915ced1..12b113ef93 100644\n> --- a/object-store.h\n> +++ b/object-store.h\n> @@ -33,6 +33,11 @@ struct object_directory {\n>  \tchar *path;\n>  };\n>  \n> +struct git_zstream_reader {\n> +\tvoid (*fill)(struct git_zstream *);\n> +\tvoid (*use)(struct git_zstream *);\n> +};\n> +\n>  KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n>  \tstruct object_directory *, 1, fspathhash, fspatheq)\n>  \n> @@ -225,6 +230,10 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n>  int write_object_file(const void *buf, unsigned long len,\n>  \t\t      const char *type, struct object_id *oid);\n>  \n> +int write_stream_object_file(struct git_zstream_reader *reader,\n> +\t\t\t     unsigned long len, const char *type,\n> +\t\t\t     struct object_id *oid, int dry_run);\n> +\n>  int hash_object_file_literally(const void *buf, unsigned long len,\n>  \t\t\t       const char *type, struct object_id *oid,\n>  \t\t\t       unsigned flags);\n> diff --git a/t/t5590-receive-unpack-objects.sh b/t/t5590-receive-unpack-objects.sh\n> new file mode 100755\n> index 0000000000..7e63dfc0db\n> --- /dev/null\n> +++ b/t/t5590-receive-unpack-objects.sh\n> @@ -0,0 +1,92 @@\n> +#!/bin/sh\n> +#\n> +# Copyright (c) 2021 Han Xin\n> +#\n> +\n> +test_description='Test unpack-objects when receive pack'\n> +\n> +GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n> +export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n> +\n> +. ./test-lib.sh\n> +\n> +test_expect_success \"create commit with big blobs (1.5 MB)\" '\n> +\ttest-tool genrandom foo 1500000 >big-blob &&\n> +\ttest_commit --append foo big-blob &&\n> +\ttest-tool genrandom bar 1500000 >big-blob &&\n> +\ttest_commit --append bar big-blob &&\n> +\t(\n> +\t\tcd .git &&\n> +\t\tfind objects/?? -type f | sort\n> +\t) >expect &&\n> +\tgit repack -ad\n> +'\n> +\n> +test_expect_success 'setup GIT_ALLOC_LIMIT to 1MB' '\n> +\tGIT_ALLOC_LIMIT=1m &&\n> +\texport GIT_ALLOC_LIMIT\n> +'\n> +\n> +test_expect_success 'prepare dest repository' '\n> +\tgit init --bare dest.git &&\n> +\tgit -C dest.git config core.bigFileThreshold 2m &&\n> +\tgit -C dest.git config receive.unpacklimit 100\n> +'\n> +\n> +test_expect_success 'fail to push: cannot allocate' '\n> +\ttest_must_fail git push dest.git HEAD 2>err &&\n> +\ttest_i18ngrep \"remote: fatal: attempting to allocate\" err &&\n> +\t(\n> +\t\tcd dest.git &&\n> +\t\tfind objects/?? -type f | sort\n> +\t) >actual &&\n> +\t! test_cmp expect actual\n> +'\n> +\n> +test_expect_success 'set a lower bigfile threshold' '\n> +\tgit -C dest.git config core.bigFileThreshold 1m\n> +'\n> +\n> +test_expect_success 'unpack big object in stream' '\n> +\tgit push dest.git HEAD &&\n> +\tgit -C dest.git fsck &&\n> +\t(\n> +\t\tcd dest.git &&\n> +\t\tfind objects/?? -type f | sort\n> +\t) >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_expect_success 'setup for unpack-objects dry-run test' '\n> +\tPACK=$(echo main | git pack-objects --progress --revs test) &&\n> +\tunset GIT_ALLOC_LIMIT &&\n> +\tgit init --bare unpack-test.git\n> +'\n> +\n> +test_expect_success 'unpack-objects dry-run with large threshold' '\n> +\t(\n> +\t\tcd unpack-test.git &&\n> +\t\tgit config core.bigFileThreshold 2m &&\n> +\t\tgit unpack-objects -n <../test-$PACK.pack\n> +\t) &&\n> +\t(\n> +\t\tcd unpack-test.git &&\n> +\t\tfind objects/ -type f\n> +\t) >actual &&\n> +\ttest_must_be_empty actual\n> +'\n> +\n> +test_expect_success 'unpack-objects dry-run with small threshold' '\n> +\t(\n> +\t\tcd unpack-test.git &&\n> +\t\tgit config core.bigFileThreshold 1m &&\n> +\t\tgit unpack-objects -n <../test-$PACK.pack\n> +\t) &&\n> +\t(\n> +\t\tcd unpack-test.git &&\n> +\t\tfind objects/ -type f\n> +\t) >actual &&\n> +\ttest_must_be_empty actual\n> +'\n> +\n> +test_done\n\n"},{"id":"439172","messageId":"CAO0brD1CXhuAwQe5mmjWcD1dF-cdRdk4o1rm0wyygbhEjRR84g@mail.gmail.com","threadId":"56672","inReplyTo":"aeb1c7cb-7e01-8a21-5f7a-d27bd43d7604@iee.email","subject":"Re: [PATCH] unpack-objects: unpack large object in stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-10-21T03:42:46Z","receivedAt":"2021-10-21T03:43:27Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"Philip Oakley <philipoakley@iee.email> 于2021年10月20日周三 下午10:43写道：\n>\n> On 09/10/2021 09:20, Han Xin wrote:\n> > From: Han Xin <hanxin.hx@alibaba-inc.com>\n> >\n> > When calling \"unpack_non_delta_entry()\", will allocate full memory for\n> > the whole size of the unpacked object and write the buffer to loose file\n> > on disk. This may lead to OOM for the git-unpack-objects process when\n> > unpacking a very large object.\n> >\n> > In function \"unpack_delta_entry()\", will also allocate full memory to\n> > buffer the whole delta, but since there will be no delta for an object\n> > larger than \"core.bigFileThreshold\", this issue is moderate.\n> >\n> > To resolve the OOM issue in \"git-unpack-objects\", we can unpack large\n> > object to file in stream, and use the setting of \"core.bigFileThreshold\" as\n> > the threshold for large object.\n> >\n> > Reviewed-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> > Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> > ---\n> >  builtin/unpack-objects.c          |  41 +++++++-\n> >  object-file.c                     | 149 +++++++++++++++++++++++++++---\n> >  object-store.h                    |   9 ++\n> >  t/t5590-receive-unpack-objects.sh |  92 ++++++++++++++++++\n> >  4 files changed, 279 insertions(+), 12 deletions(-)\n> >  create mode 100755 t/t5590-receive-unpack-objects.sh\n> >\n> > diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n> > index 4a9466295b..8ac77e60a8 100644\n> > --- a/builtin/unpack-objects.c\n> > +++ b/builtin/unpack-objects.c\n> > @@ -320,11 +320,50 @@ static void added_object(unsigned nr, enum object_type type,\n> >       }\n> >  }\n> >\n> > +static void fill_stream(struct git_zstream *stream)\n> > +{\n> > +     stream->next_in = fill(1);\n> > +     stream->avail_in = len;\n> > +}\n> > +\n> > +static void use_stream(struct git_zstream *stream)\n> > +{\n> > +     use(len - stream->avail_in);\n> > +}\n> > +\n> > +static void write_stream_blob(unsigned nr, unsigned long size)\n>\n> Can we use size_t for the `size`, and possibly `nr`, to improve\n> compatibility with Windows systems where unsigned long is only 32 bits?\n>\n> There has been some work in the past on providing large file support on\n> Windows, which requires numerous long -> size_t changes.\n>\n> Philip\n\nThanks for your review. I'm not sure if I should do this change in this patch,\nit will also change the type defined in `unpack_one()`,`unpack_non_delta_entry`,\n`write_object()` and many others.\n\n> > +{\n> > +     struct git_zstream_reader reader;\n> > +     struct object_id *oid = &obj_list[nr].oid;\n> > +\n> > +     reader.fill = &fill_stream;\n> > +     reader.use = &use_stream;\n> > +\n> > +     if (write_stream_object_file(&reader, size, type_name(OBJ_BLOB),\n> > +                                  oid, dry_run))\n> > +             die(\"failed to write object in stream\");\n> > +     if (strict && !dry_run) {\n> > +             struct blob *blob = lookup_blob(the_repository, oid);\n> > +             if (blob)\n> > +                     blob->object.flags |= FLAG_WRITTEN;\n> > +             else\n> > +                     die(\"invalid blob object from stream\");\n> > +     }\n> > +     obj_list[nr].obj = NULL;\n> > +}\n> > +\n> >  static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n> >                                  unsigned nr)\n> >  {\n> > -     void *buf = get_data(size);\n> > +     void *buf;\n> > +\n> > +     /* Write large blob in stream without allocating full buffer. */\n> > +     if (type == OBJ_BLOB && size > big_file_threshold) {\n> > +             write_stream_blob(nr, size);\n> > +             return;\n> > +     }\n> >\n> > +     buf = get_data(size);\n> >       if (!dry_run && buf)\n> >               write_object(nr, type, buf, size);\n> >       else\n> > diff --git a/object-file.c b/object-file.c\n> > index a8be899481..06c1693675 100644\n> > --- a/object-file.c\n> > +++ b/object-file.c\n> > @@ -1913,6 +1913,28 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n> >       return fd;\n> >  }\n> >\n> > +static int write_object_buffer(struct git_zstream *stream, git_hash_ctx *c,\n> > +                            int fd, unsigned char *compressed,\n> > +                            int compressed_len, const void *buf,\n> > +                            size_t len, int flush)\n> > +{\n> > +     int ret;\n> > +\n> > +     stream->next_in = (void *)buf;\n> > +     stream->avail_in = len;\n> > +     do {\n> > +             unsigned char *in0 = stream->next_in;\n> > +             ret = git_deflate(stream, flush);\n> > +             the_hash_algo->update_fn(c, in0, stream->next_in - in0);\n> > +             if (write_buffer(fd, compressed, stream->next_out - compressed) < 0)\n> > +                     die(_(\"unable to write loose object file\"));\n> > +             stream->next_out = compressed;\n> > +             stream->avail_out = compressed_len;\n> > +     } while (ret == Z_OK);\n> > +\n> > +     return ret;\n> > +}\n> > +\n> >  static int write_loose_object(const struct object_id *oid, char *hdr,\n> >                             int hdrlen, const void *buf, unsigned long len,\n> >                             time_t mtime)\n> > @@ -1949,17 +1971,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n> >       the_hash_algo->update_fn(&c, hdr, hdrlen);\n> >\n> >       /* Then the data itself.. */\n> > -     stream.next_in = (void *)buf;\n> > -     stream.avail_in = len;\n> > -     do {\n> > -             unsigned char *in0 = stream.next_in;\n> > -             ret = git_deflate(&stream, Z_FINISH);\n> > -             the_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n> > -             if (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n> > -                     die(_(\"unable to write loose object file\"));\n> > -             stream.next_out = compressed;\n> > -             stream.avail_out = sizeof(compressed);\n> > -     } while (ret == Z_OK);\n> > +     ret = write_object_buffer(&stream, &c, fd, compressed,\n> > +                               sizeof(compressed), buf, len,\n> > +                               Z_FINISH);\n> >\n> >       if (ret != Z_STREAM_END)\n> >               die(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n> > @@ -2020,6 +2034,119 @@ int write_object_file(const void *buf, unsigned long len, const char *type,\n> >       return write_loose_object(oid, hdr, hdrlen, buf, len, 0);\n> >  }\n> >\n> > +int write_stream_object_file(struct git_zstream_reader *reader,\n> > +                          unsigned long len, const char *type,\n> > +                          struct object_id *oid,\n> > +                          int dry_run)\n> > +{\n> > +     git_zstream istream, ostream;\n> > +     unsigned char buf[8192], compressed[4096];\n> > +     char hdr[MAX_HEADER_LEN];\n> > +     int istatus, ostatus, fd = 0, hdrlen, dirlen, flush = 0;\n> > +     int ret = 0;\n> > +     git_hash_ctx c;\n> > +     struct strbuf tmp_file = STRBUF_INIT;\n> > +     struct strbuf filename = STRBUF_INIT;\n> > +\n> > +     /* Write tmpfile in objects dir, because oid is unknown */\n> > +     if (!dry_run) {\n> > +             strbuf_addstr(&filename, the_repository->objects->odb->path);\n> > +             strbuf_addch(&filename, '/');\n> > +             fd = create_tmpfile(&tmp_file, filename.buf);\n> > +             if (fd < 0) {\n> > +                     if (errno == EACCES)\n> > +                             ret = error(_(\"insufficient permission for adding an object to repository database %s\"),\n> > +                                     get_object_directory());\n> > +                     else\n> > +                             ret = error_errno(_(\"unable to create temporary file\"));\n> > +                     goto cleanup;\n> > +             }\n> > +     }\n> > +\n> > +     memset(&istream, 0, sizeof(istream));\n> > +     istream.next_out = buf;\n> > +     istream.avail_out = sizeof(buf);\n> > +     git_inflate_init(&istream);\n> > +\n> > +     if (!dry_run) {\n> > +             /* Set it up */\n> > +             git_deflate_init(&ostream, zlib_compression_level);\n> > +             ostream.next_out = compressed;\n> > +             ostream.avail_out = sizeof(compressed);\n> > +             the_hash_algo->init_fn(&c);\n> > +\n> > +             /* First header */\n> > +             hdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\" PRIuMAX, type,\n> > +                             (uintmax_t)len) + 1;\n> > +             ostream.next_in = (unsigned char *)hdr;\n> > +             ostream.avail_in = hdrlen;\n> > +             while (git_deflate(&ostream, 0) == Z_OK)\n> > +                     ; /* nothing */\n> > +             the_hash_algo->update_fn(&c, hdr, hdrlen);\n> > +     }\n> > +\n> > +     /* Then the data itself */\n> > +     do {\n> > +             unsigned char *last_out = istream.next_out;\n> > +             reader->fill(&istream);\n> > +             istatus = git_inflate(&istream, 0);\n> > +             if (istatus == Z_STREAM_END)\n> > +                     flush = Z_FINISH;\n> > +             reader->use(&istream);\n> > +             if (!dry_run)\n> > +                     ostatus = write_object_buffer(&ostream, &c, fd, compressed,\n> > +                                                   sizeof(compressed), last_out,\n> > +                                                   istream.next_out - last_out,\n> > +                                                   flush);\n> > +             istream.next_out = buf;\n> > +             istream.avail_out = sizeof(buf);\n> > +     } while (istatus == Z_OK);\n> > +\n> > +     if (istream.total_out != len || istatus != Z_STREAM_END)\n> > +             die( _(\"inflate returned %d\"), istatus);\n> > +     git_inflate_end(&istream);\n> > +\n> > +     if (dry_run)\n> > +             goto cleanup;\n> > +\n> > +     if (ostatus != Z_STREAM_END)\n> > +             die(_(\"unable to deflate new object (%d)\"), ostatus);\n> > +     ostatus = git_deflate_end_gently(&ostream);\n> > +     if (ostatus != Z_OK)\n> > +             die(_(\"deflateEnd on object failed (%d)\"), ostatus);\n> > +     the_hash_algo->final_fn(oid->hash, &c);\n> > +     close_loose_object(fd);\n> > +\n> > +     /* We get the oid now */\n> > +     loose_object_path(the_repository, &filename, oid);\n> > +\n> > +     dirlen = directory_size(filename.buf);\n> > +     if (dirlen) {\n> > +             struct strbuf dir = STRBUF_INIT;\n> > +             /*\n> > +              * Make sure the directory exists; note that the contents\n> > +              * of the buffer are undefined after mkstemp returns an\n> > +              * error, so we have to rewrite the whole buffer from\n> > +              * scratch.\n> > +              */\n> > +             strbuf_add(&dir, filename.buf, dirlen - 1);\n> > +             if (mkdir(dir.buf, 0777) && errno != EEXIST) {\n> > +                     unlink_or_warn(tmp_file.buf);\n> > +                     strbuf_release(&dir);\n> > +                     ret = -1;\n> > +                     goto cleanup;\n> > +             }\n> > +             strbuf_release(&dir);\n> > +     }\n> > +\n> > +     ret = finalize_object_file(tmp_file.buf, filename.buf);\n> > +\n> > +cleanup:\n> > +     strbuf_release(&tmp_file);\n> > +     strbuf_release(&filename);\n> > +     return ret;\n> > +}\n> > +\n> >  int hash_object_file_literally(const void *buf, unsigned long len,\n> >                              const char *type, struct object_id *oid,\n> >                              unsigned flags)\n> > diff --git a/object-store.h b/object-store.h\n> > index d24915ced1..12b113ef93 100644\n> > --- a/object-store.h\n> > +++ b/object-store.h\n> > @@ -33,6 +33,11 @@ struct object_directory {\n> >       char *path;\n> >  };\n> >\n> > +struct git_zstream_reader {\n> > +     void (*fill)(struct git_zstream *);\n> > +     void (*use)(struct git_zstream *);\n> > +};\n> > +\n> >  KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n> >       struct object_directory *, 1, fspathhash, fspatheq)\n> >\n> > @@ -225,6 +230,10 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n> >  int write_object_file(const void *buf, unsigned long len,\n> >                     const char *type, struct object_id *oid);\n> >\n> > +int write_stream_object_file(struct git_zstream_reader *reader,\n> > +                          unsigned long len, const char *type,\n> > +                          struct object_id *oid, int dry_run);\n> > +\n> >  int hash_object_file_literally(const void *buf, unsigned long len,\n> >                              const char *type, struct object_id *oid,\n> >                              unsigned flags);\n> > diff --git a/t/t5590-receive-unpack-objects.sh b/t/t5590-receive-unpack-objects.sh\n> > new file mode 100755\n> > index 0000000000..7e63dfc0db\n> > --- /dev/null\n> > +++ b/t/t5590-receive-unpack-objects.sh\n> > @@ -0,0 +1,92 @@\n> > +#!/bin/sh\n> > +#\n> > +# Copyright (c) 2021 Han Xin\n> > +#\n> > +\n> > +test_description='Test unpack-objects when receive pack'\n> > +\n> > +GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n> > +export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n> > +\n> > +. ./test-lib.sh\n> > +\n> > +test_expect_success \"create commit with big blobs (1.5 MB)\" '\n> > +     test-tool genrandom foo 1500000 >big-blob &&\n> > +     test_commit --append foo big-blob &&\n> > +     test-tool genrandom bar 1500000 >big-blob &&\n> > +     test_commit --append bar big-blob &&\n> > +     (\n> > +             cd .git &&\n> > +             find objects/?? -type f | sort\n> > +     ) >expect &&\n> > +     git repack -ad\n> > +'\n> > +\n> > +test_expect_success 'setup GIT_ALLOC_LIMIT to 1MB' '\n> > +     GIT_ALLOC_LIMIT=1m &&\n> > +     export GIT_ALLOC_LIMIT\n> > +'\n> > +\n> > +test_expect_success 'prepare dest repository' '\n> > +     git init --bare dest.git &&\n> > +     git -C dest.git config core.bigFileThreshold 2m &&\n> > +     git -C dest.git config receive.unpacklimit 100\n> > +'\n> > +\n> > +test_expect_success 'fail to push: cannot allocate' '\n> > +     test_must_fail git push dest.git HEAD 2>err &&\n> > +     test_i18ngrep \"remote: fatal: attempting to allocate\" err &&\n> > +     (\n> > +             cd dest.git &&\n> > +             find objects/?? -type f | sort\n> > +     ) >actual &&\n> > +     ! test_cmp expect actual\n> > +'\n> > +\n> > +test_expect_success 'set a lower bigfile threshold' '\n> > +     git -C dest.git config core.bigFileThreshold 1m\n> > +'\n> > +\n> > +test_expect_success 'unpack big object in stream' '\n> > +     git push dest.git HEAD &&\n> > +     git -C dest.git fsck &&\n> > +     (\n> > +             cd dest.git &&\n> > +             find objects/?? -type f | sort\n> > +     ) >actual &&\n> > +     test_cmp expect actual\n> > +'\n> > +\n> > +test_expect_success 'setup for unpack-objects dry-run test' '\n> > +     PACK=$(echo main | git pack-objects --progress --revs test) &&\n> > +     unset GIT_ALLOC_LIMIT &&\n> > +     git init --bare unpack-test.git\n> > +'\n> > +\n> > +test_expect_success 'unpack-objects dry-run with large threshold' '\n> > +     (\n> > +             cd unpack-test.git &&\n> > +             git config core.bigFileThreshold 2m &&\n> > +             git unpack-objects -n <../test-$PACK.pack\n> > +     ) &&\n> > +     (\n> > +             cd unpack-test.git &&\n> > +             find objects/ -type f\n> > +     ) >actual &&\n> > +     test_must_be_empty actual\n> > +'\n> > +\n> > +test_expect_success 'unpack-objects dry-run with small threshold' '\n> > +     (\n> > +             cd unpack-test.git &&\n> > +             git config core.bigFileThreshold 1m &&\n> > +             git unpack-objects -n <../test-$PACK.pack\n> > +     ) &&\n> > +     (\n> > +             cd unpack-test.git &&\n> > +             find objects/ -type f\n> > +     ) >actual &&\n> > +     test_must_be_empty actual\n> > +'\n> > +\n> > +test_done\n>\n"},{"id":"439299","messageId":"07e7e056-7f50-4a1b-5ca4-b9994b846e7b@iee.email","threadId":"56672","inReplyTo":"CAO0brD1CXhuAwQe5mmjWcD1dF-cdRdk4o1rm0wyygbhEjRR84g@mail.gmail.com","subject":"Re: [PATCH] unpack-objects: unpack large object in stream","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2021-10-21T22:47:14Z","receivedAt":"2021-10-21T22:47:17Z","isPatch":true,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"On 21/10/2021 04:42, Han Xin wrote:\n>>> +static void write_stream_blob(unsigned nr, unsigned long size)\n>> Can we use size_t for the `size`, and possibly `nr`, to improve\n>> compatibility with Windows systems where unsigned long is only 32 bits?\n>>\n>> There has been some work in the past on providing large file support on\n>> Windows, which requires numerous long -> size_t changes.\n>>\n>> Philip\n> Thanks for your review. I'm not sure if I should do this change in this patch,\n> it will also change the type defined in `unpack_one()`,`unpack_non_delta_entry`,\n> `write_object()` and many others.\n>\nI was mainly raising the issue regarding the 4GB (sometime 2GB)\nlimitations on Windows which has been a problem for many years.\n\nI had been thinking of not changing the `nr` (number of objects limit)\nas 2G objects is hopefully already sufficient, even for thargest of\nrepos (though IIUC their index file size did break the 32bit size limit).\n\nStaying with the existing types won't make the situation any worse, so\nfrom that perspective the change isn't needed.\n--\nPhilip\n"},{"id":"440358","messageId":"CAO0brD3uT6y0ytPvjMzi9LdNRUR9bWXf3-o+D7RbdSLAJxCfAw@mail.gmail.com","threadId":"56672","inReplyTo":"20211009082058.41138-1-chiyutianyi@gmail.com","subject":"Re: [PATCH] unpack-objects: unpack large object in stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-11-03T01:48:35Z","receivedAt":"2021-11-03T01:49:00Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"Any more suggestions?\n\nHan Xin <chiyutianyi@gmail.com> 于2021年10月9日周六 下午4:21写道：\n>\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> When calling \"unpack_non_delta_entry()\", will allocate full memory for\n> the whole size of the unpacked object and write the buffer to loose file\n> on disk. This may lead to OOM for the git-unpack-objects process when\n> unpacking a very large object.\n>\n> In function \"unpack_delta_entry()\", will also allocate full memory to\n> buffer the whole delta, but since there will be no delta for an object\n> larger than \"core.bigFileThreshold\", this issue is moderate.\n>\n> To resolve the OOM issue in \"git-unpack-objects\", we can unpack large\n> object to file in stream, and use the setting of \"core.bigFileThreshold\" as\n> the threshold for large object.\n>\n> Reviewed-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  builtin/unpack-objects.c          |  41 +++++++-\n>  object-file.c                     | 149 +++++++++++++++++++++++++++---\n>  object-store.h                    |   9 ++\n>  t/t5590-receive-unpack-objects.sh |  92 ++++++++++++++++++\n>  4 files changed, 279 insertions(+), 12 deletions(-)\n>  create mode 100755 t/t5590-receive-unpack-objects.sh\n>\n> diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n> index 4a9466295b..8ac77e60a8 100644\n> --- a/builtin/unpack-objects.c\n> +++ b/builtin/unpack-objects.c\n> @@ -320,11 +320,50 @@ static void added_object(unsigned nr, enum object_type type,\n>         }\n>  }\n>\n> +static void fill_stream(struct git_zstream *stream)\n> +{\n> +       stream->next_in = fill(1);\n> +       stream->avail_in = len;\n> +}\n> +\n> +static void use_stream(struct git_zstream *stream)\n> +{\n> +       use(len - stream->avail_in);\n> +}\n> +\n> +static void write_stream_blob(unsigned nr, unsigned long size)\n> +{\n> +       struct git_zstream_reader reader;\n> +       struct object_id *oid = &obj_list[nr].oid;\n> +\n> +       reader.fill = &fill_stream;\n> +       reader.use = &use_stream;\n> +\n> +       if (write_stream_object_file(&reader, size, type_name(OBJ_BLOB),\n> +                                    oid, dry_run))\n> +               die(\"failed to write object in stream\");\n> +       if (strict && !dry_run) {\n> +               struct blob *blob = lookup_blob(the_repository, oid);\n> +               if (blob)\n> +                       blob->object.flags |= FLAG_WRITTEN;\n> +               else\n> +                       die(\"invalid blob object from stream\");\n> +       }\n> +       obj_list[nr].obj = NULL;\n> +}\n> +\n>  static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n>                                    unsigned nr)\n>  {\n> -       void *buf = get_data(size);\n> +       void *buf;\n> +\n> +       /* Write large blob in stream without allocating full buffer. */\n> +       if (type == OBJ_BLOB && size > big_file_threshold) {\n> +               write_stream_blob(nr, size);\n> +               return;\n> +       }\n>\n> +       buf = get_data(size);\n>         if (!dry_run && buf)\n>                 write_object(nr, type, buf, size);\n>         else\n> diff --git a/object-file.c b/object-file.c\n> index a8be899481..06c1693675 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1913,6 +1913,28 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n>         return fd;\n>  }\n>\n> +static int write_object_buffer(struct git_zstream *stream, git_hash_ctx *c,\n> +                              int fd, unsigned char *compressed,\n> +                              int compressed_len, const void *buf,\n> +                              size_t len, int flush)\n> +{\n> +       int ret;\n> +\n> +       stream->next_in = (void *)buf;\n> +       stream->avail_in = len;\n> +       do {\n> +               unsigned char *in0 = stream->next_in;\n> +               ret = git_deflate(stream, flush);\n> +               the_hash_algo->update_fn(c, in0, stream->next_in - in0);\n> +               if (write_buffer(fd, compressed, stream->next_out - compressed) < 0)\n> +                       die(_(\"unable to write loose object file\"));\n> +               stream->next_out = compressed;\n> +               stream->avail_out = compressed_len;\n> +       } while (ret == Z_OK);\n> +\n> +       return ret;\n> +}\n> +\n>  static int write_loose_object(const struct object_id *oid, char *hdr,\n>                               int hdrlen, const void *buf, unsigned long len,\n>                               time_t mtime)\n> @@ -1949,17 +1971,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>         the_hash_algo->update_fn(&c, hdr, hdrlen);\n>\n>         /* Then the data itself.. */\n> -       stream.next_in = (void *)buf;\n> -       stream.avail_in = len;\n> -       do {\n> -               unsigned char *in0 = stream.next_in;\n> -               ret = git_deflate(&stream, Z_FINISH);\n> -               the_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n> -               if (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n> -                       die(_(\"unable to write loose object file\"));\n> -               stream.next_out = compressed;\n> -               stream.avail_out = sizeof(compressed);\n> -       } while (ret == Z_OK);\n> +       ret = write_object_buffer(&stream, &c, fd, compressed,\n> +                                 sizeof(compressed), buf, len,\n> +                                 Z_FINISH);\n>\n>         if (ret != Z_STREAM_END)\n>                 die(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n> @@ -2020,6 +2034,119 @@ int write_object_file(const void *buf, unsigned long len, const char *type,\n>         return write_loose_object(oid, hdr, hdrlen, buf, len, 0);\n>  }\n>\n> +int write_stream_object_file(struct git_zstream_reader *reader,\n> +                            unsigned long len, const char *type,\n> +                            struct object_id *oid,\n> +                            int dry_run)\n> +{\n> +       git_zstream istream, ostream;\n> +       unsigned char buf[8192], compressed[4096];\n> +       char hdr[MAX_HEADER_LEN];\n> +       int istatus, ostatus, fd = 0, hdrlen, dirlen, flush = 0;\n> +       int ret = 0;\n> +       git_hash_ctx c;\n> +       struct strbuf tmp_file = STRBUF_INIT;\n> +       struct strbuf filename = STRBUF_INIT;\n> +\n> +       /* Write tmpfile in objects dir, because oid is unknown */\n> +       if (!dry_run) {\n> +               strbuf_addstr(&filename, the_repository->objects->odb->path);\n> +               strbuf_addch(&filename, '/');\n> +               fd = create_tmpfile(&tmp_file, filename.buf);\n> +               if (fd < 0) {\n> +                       if (errno == EACCES)\n> +                               ret = error(_(\"insufficient permission for adding an object to repository database %s\"),\n> +                                       get_object_directory());\n> +                       else\n> +                               ret = error_errno(_(\"unable to create temporary file\"));\n> +                       goto cleanup;\n> +               }\n> +       }\n> +\n> +       memset(&istream, 0, sizeof(istream));\n> +       istream.next_out = buf;\n> +       istream.avail_out = sizeof(buf);\n> +       git_inflate_init(&istream);\n> +\n> +       if (!dry_run) {\n> +               /* Set it up */\n> +               git_deflate_init(&ostream, zlib_compression_level);\n> +               ostream.next_out = compressed;\n> +               ostream.avail_out = sizeof(compressed);\n> +               the_hash_algo->init_fn(&c);\n> +\n> +               /* First header */\n> +               hdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\" PRIuMAX, type,\n> +                               (uintmax_t)len) + 1;\n> +               ostream.next_in = (unsigned char *)hdr;\n> +               ostream.avail_in = hdrlen;\n> +               while (git_deflate(&ostream, 0) == Z_OK)\n> +                       ; /* nothing */\n> +               the_hash_algo->update_fn(&c, hdr, hdrlen);\n> +       }\n> +\n> +       /* Then the data itself */\n> +       do {\n> +               unsigned char *last_out = istream.next_out;\n> +               reader->fill(&istream);\n> +               istatus = git_inflate(&istream, 0);\n> +               if (istatus == Z_STREAM_END)\n> +                       flush = Z_FINISH;\n> +               reader->use(&istream);\n> +               if (!dry_run)\n> +                       ostatus = write_object_buffer(&ostream, &c, fd, compressed,\n> +                                                     sizeof(compressed), last_out,\n> +                                                     istream.next_out - last_out,\n> +                                                     flush);\n> +               istream.next_out = buf;\n> +               istream.avail_out = sizeof(buf);\n> +       } while (istatus == Z_OK);\n> +\n> +       if (istream.total_out != len || istatus != Z_STREAM_END)\n> +               die( _(\"inflate returned %d\"), istatus);\n> +       git_inflate_end(&istream);\n> +\n> +       if (dry_run)\n> +               goto cleanup;\n> +\n> +       if (ostatus != Z_STREAM_END)\n> +               die(_(\"unable to deflate new object (%d)\"), ostatus);\n> +       ostatus = git_deflate_end_gently(&ostream);\n> +       if (ostatus != Z_OK)\n> +               die(_(\"deflateEnd on object failed (%d)\"), ostatus);\n> +       the_hash_algo->final_fn(oid->hash, &c);\n> +       close_loose_object(fd);\n> +\n> +       /* We get the oid now */\n> +       loose_object_path(the_repository, &filename, oid);\n> +\n> +       dirlen = directory_size(filename.buf);\n> +       if (dirlen) {\n> +               struct strbuf dir = STRBUF_INIT;\n> +               /*\n> +                * Make sure the directory exists; note that the contents\n> +                * of the buffer are undefined after mkstemp returns an\n> +                * error, so we have to rewrite the whole buffer from\n> +                * scratch.\n> +                */\n> +               strbuf_add(&dir, filename.buf, dirlen - 1);\n> +               if (mkdir(dir.buf, 0777) && errno != EEXIST) {\n> +                       unlink_or_warn(tmp_file.buf);\n> +                       strbuf_release(&dir);\n> +                       ret = -1;\n> +                       goto cleanup;\n> +               }\n> +               strbuf_release(&dir);\n> +       }\n> +\n> +       ret = finalize_object_file(tmp_file.buf, filename.buf);\n> +\n> +cleanup:\n> +       strbuf_release(&tmp_file);\n> +       strbuf_release(&filename);\n> +       return ret;\n> +}\n> +\n>  int hash_object_file_literally(const void *buf, unsigned long len,\n>                                const char *type, struct object_id *oid,\n>                                unsigned flags)\n> diff --git a/object-store.h b/object-store.h\n> index d24915ced1..12b113ef93 100644\n> --- a/object-store.h\n> +++ b/object-store.h\n> @@ -33,6 +33,11 @@ struct object_directory {\n>         char *path;\n>  };\n>\n> +struct git_zstream_reader {\n> +       void (*fill)(struct git_zstream *);\n> +       void (*use)(struct git_zstream *);\n> +};\n> +\n>  KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n>         struct object_directory *, 1, fspathhash, fspatheq)\n>\n> @@ -225,6 +230,10 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n>  int write_object_file(const void *buf, unsigned long len,\n>                       const char *type, struct object_id *oid);\n>\n> +int write_stream_object_file(struct git_zstream_reader *reader,\n> +                            unsigned long len, const char *type,\n> +                            struct object_id *oid, int dry_run);\n> +\n>  int hash_object_file_literally(const void *buf, unsigned long len,\n>                                const char *type, struct object_id *oid,\n>                                unsigned flags);\n> diff --git a/t/t5590-receive-unpack-objects.sh b/t/t5590-receive-unpack-objects.sh\n> new file mode 100755\n> index 0000000000..7e63dfc0db\n> --- /dev/null\n> +++ b/t/t5590-receive-unpack-objects.sh\n> @@ -0,0 +1,92 @@\n> +#!/bin/sh\n> +#\n> +# Copyright (c) 2021 Han Xin\n> +#\n> +\n> +test_description='Test unpack-objects when receive pack'\n> +\n> +GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n> +export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n> +\n> +. ./test-lib.sh\n> +\n> +test_expect_success \"create commit with big blobs (1.5 MB)\" '\n> +       test-tool genrandom foo 1500000 >big-blob &&\n> +       test_commit --append foo big-blob &&\n> +       test-tool genrandom bar 1500000 >big-blob &&\n> +       test_commit --append bar big-blob &&\n> +       (\n> +               cd .git &&\n> +               find objects/?? -type f | sort\n> +       ) >expect &&\n> +       git repack -ad\n> +'\n> +\n> +test_expect_success 'setup GIT_ALLOC_LIMIT to 1MB' '\n> +       GIT_ALLOC_LIMIT=1m &&\n> +       export GIT_ALLOC_LIMIT\n> +'\n> +\n> +test_expect_success 'prepare dest repository' '\n> +       git init --bare dest.git &&\n> +       git -C dest.git config core.bigFileThreshold 2m &&\n> +       git -C dest.git config receive.unpacklimit 100\n> +'\n> +\n> +test_expect_success 'fail to push: cannot allocate' '\n> +       test_must_fail git push dest.git HEAD 2>err &&\n> +       test_i18ngrep \"remote: fatal: attempting to allocate\" err &&\n> +       (\n> +               cd dest.git &&\n> +               find objects/?? -type f | sort\n> +       ) >actual &&\n> +       ! test_cmp expect actual\n> +'\n> +\n> +test_expect_success 'set a lower bigfile threshold' '\n> +       git -C dest.git config core.bigFileThreshold 1m\n> +'\n> +\n> +test_expect_success 'unpack big object in stream' '\n> +       git push dest.git HEAD &&\n> +       git -C dest.git fsck &&\n> +       (\n> +               cd dest.git &&\n> +               find objects/?? -type f | sort\n> +       ) >actual &&\n> +       test_cmp expect actual\n> +'\n> +\n> +test_expect_success 'setup for unpack-objects dry-run test' '\n> +       PACK=$(echo main | git pack-objects --progress --revs test) &&\n> +       unset GIT_ALLOC_LIMIT &&\n> +       git init --bare unpack-test.git\n> +'\n> +\n> +test_expect_success 'unpack-objects dry-run with large threshold' '\n> +       (\n> +               cd unpack-test.git &&\n> +               git config core.bigFileThreshold 2m &&\n> +               git unpack-objects -n <../test-$PACK.pack\n> +       ) &&\n> +       (\n> +               cd unpack-test.git &&\n> +               find objects/ -type f\n> +       ) >actual &&\n> +       test_must_be_empty actual\n> +'\n> +\n> +test_expect_success 'unpack-objects dry-run with small threshold' '\n> +       (\n> +               cd unpack-test.git &&\n> +               git config core.bigFileThreshold 1m &&\n> +               git unpack-objects -n <../test-$PACK.pack\n> +       ) &&\n> +       (\n> +               cd unpack-test.git &&\n> +               find objects/ -type f\n> +       ) >actual &&\n> +       test_must_be_empty actual\n> +'\n> +\n> +test_done\n> --\n> 2.33.0.1.g09a6bb964f.dirty\n>\n"},{"id":"440364","messageId":"828d1d4c-43da-7cbb-bc39-18f6892e1562@iee.email","threadId":"56672","inReplyTo":"CAO0brD3uT6y0ytPvjMzi9LdNRUR9bWXf3-o+D7RbdSLAJxCfAw@mail.gmail.com","subject":"Re: [PATCH] unpack-objects: unpack large object in stream","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2021-11-03T10:07:33Z","receivedAt":"2021-11-03T10:07:43Z","isPatch":true,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"(replies to the alibaba-inc.com aren't getting through for me)\n\nOn 03/11/2021 01:48, Han Xin wrote:\n> Any more suggestions?\n>\n> Han Xin <chiyutianyi@gmail.com> 于2021年10月9日周六 下午4:21写道：\n>> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>>\n>> When calling \"unpack_non_delta_entry()\", will allocate full memory for\n>> the whole size of the unpacked object and write the buffer to loose file\n>> on disk. This may lead to OOM for the git-unpack-objects process when\n>> unpacking a very large object.\n\nIs it possible to split the patch into smaller pieces, taking each item\nseparately?\n\nFor large files (as above), it should be possible to stream the\nunpacking direct to disk, in the same way that the zlib reading is\nchunked. However having the same 'code' in two places would need to be\naddressed (the DRY principle).\n\nAt the moment on LLP64 systems (Windows) there is already a long (32bit)\nvs size_t (64bit) problem there (zlib stream), and the size_t problem\nthen permeates the wider codebase.\n\nThe normal Git file operations does tend to memory map whole files, but\nhere it looks like you can bypass that.\n>>\n>> In function \"unpack_delta_entry()\", will also allocate full memory to\n>> buffer the whole delta, but since there will be no delta for an object\n>> larger than \"core.bigFileThreshold\", this issue is moderate.\n\nWhat does 'moderate' mean here? Does it mean there is a simple test that\nallows you to side step the whole problem?\n\n>>\n>> To resolve the OOM issue in \"git-unpack-objects\", we can unpack large\n>> object to file in stream, and use the setting of \"core.bigFileThreshold\" as\n>> the threshold for large object.\n\nIs this \"core.bigFileThreshold\" the core element? If so, it is too far\ndown the commit message. The readers have already (potentially) misread\nthe message and reacted too soon.  Perhaps: \"use `core.bigFileThreshold`\nto avoid mmap OOM limits when unpacking\".\n\n--\nPhilip\n>>\n>> Reviewed-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n>> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n>> ---\n>>  builtin/unpack-objects.c          |  41 +++++++-\n>>  object-file.c                     | 149 +++++++++++++++++++++++++++---\n>>  object-store.h                    |   9 ++\n>>  t/t5590-receive-unpack-objects.sh |  92 ++++++++++++++++++\n>>  4 files changed, 279 insertions(+), 12 deletions(-)\n>>  create mode 100755 t/t5590-receive-unpack-objects.sh\n>>\n>> diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n>> index 4a9466295b..8ac77e60a8 100644\n>> --- a/builtin/unpack-objects.c\n>> +++ b/builtin/unpack-objects.c\n>> @@ -320,11 +320,50 @@ static void added_object(unsigned nr, enum object_type type,\n>>         }\n>>  }\n>>\n>> +static void fill_stream(struct git_zstream *stream)\n>> +{\n>> +       stream->next_in = fill(1);\n>> +       stream->avail_in = len;\n>> +}\n>> +\n>> +static void use_stream(struct git_zstream *stream)\n>> +{\n>> +       use(len - stream->avail_in);\n>> +}\n>> +\n>> +static void write_stream_blob(unsigned nr, unsigned long size)\n>> +{\n>> +       struct git_zstream_reader reader;\n>> +       struct object_id *oid = &obj_list[nr].oid;\n>> +\n>> +       reader.fill = &fill_stream;\n>> +       reader.use = &use_stream;\n>> +\n>> +       if (write_stream_object_file(&reader, size, type_name(OBJ_BLOB),\n>> +                                    oid, dry_run))\n>> +               die(\"failed to write object in stream\");\n>> +       if (strict && !dry_run) {\n>> +               struct blob *blob = lookup_blob(the_repository, oid);\n>> +               if (blob)\n>> +                       blob->object.flags |= FLAG_WRITTEN;\n>> +               else\n>> +                       die(\"invalid blob object from stream\");\n>> +       }\n>> +       obj_list[nr].obj = NULL;\n>> +}\n>> +\n>>  static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n>>                                    unsigned nr)\n>>  {\n>> -       void *buf = get_data(size);\n>> +       void *buf;\n>> +\n>> +       /* Write large blob in stream without allocating full buffer. */\n>> +       if (type == OBJ_BLOB && size > big_file_threshold) {\n>> +               write_stream_blob(nr, size);\n>> +               return;\n>> +       }\n>>\n>> +       buf = get_data(size);\n>>         if (!dry_run && buf)\n>>                 write_object(nr, type, buf, size);\n>>         else\n>> diff --git a/object-file.c b/object-file.c\n>> index a8be899481..06c1693675 100644\n>> --- a/object-file.c\n>> +++ b/object-file.c\n>> @@ -1913,6 +1913,28 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n>>         return fd;\n>>  }\n>>\n>> +static int write_object_buffer(struct git_zstream *stream, git_hash_ctx *c,\n>> +                              int fd, unsigned char *compressed,\n>> +                              int compressed_len, const void *buf,\n>> +                              size_t len, int flush)\n>> +{\n>> +       int ret;\n>> +\n>> +       stream->next_in = (void *)buf;\n>> +       stream->avail_in = len;\n>> +       do {\n>> +               unsigned char *in0 = stream->next_in;\n>> +               ret = git_deflate(stream, flush);\n>> +               the_hash_algo->update_fn(c, in0, stream->next_in - in0);\n>> +               if (write_buffer(fd, compressed, stream->next_out - compressed) < 0)\n>> +                       die(_(\"unable to write loose object file\"));\n>> +               stream->next_out = compressed;\n>> +               stream->avail_out = compressed_len;\n>> +       } while (ret == Z_OK);\n>> +\n>> +       return ret;\n>> +}\n>> +\n>>  static int write_loose_object(const struct object_id *oid, char *hdr,\n>>                               int hdrlen, const void *buf, unsigned long len,\n>>                               time_t mtime)\n>> @@ -1949,17 +1971,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>>         the_hash_algo->update_fn(&c, hdr, hdrlen);\n>>\n>>         /* Then the data itself.. */\n>> -       stream.next_in = (void *)buf;\n>> -       stream.avail_in = len;\n>> -       do {\n>> -               unsigned char *in0 = stream.next_in;\n>> -               ret = git_deflate(&stream, Z_FINISH);\n>> -               the_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n>> -               if (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n>> -                       die(_(\"unable to write loose object file\"));\n>> -               stream.next_out = compressed;\n>> -               stream.avail_out = sizeof(compressed);\n>> -       } while (ret == Z_OK);\n>> +       ret = write_object_buffer(&stream, &c, fd, compressed,\n>> +                                 sizeof(compressed), buf, len,\n>> +                                 Z_FINISH);\n>>\n>>         if (ret != Z_STREAM_END)\n>>                 die(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n>> @@ -2020,6 +2034,119 @@ int write_object_file(const void *buf, unsigned long len, const char *type,\n>>         return write_loose_object(oid, hdr, hdrlen, buf, len, 0);\n>>  }\n>>\n>> +int write_stream_object_file(struct git_zstream_reader *reader,\n>> +                            unsigned long len, const char *type,\n>> +                            struct object_id *oid,\n>> +                            int dry_run)\n>> +{\n>> +       git_zstream istream, ostream;\n>> +       unsigned char buf[8192], compressed[4096];\n>> +       char hdr[MAX_HEADER_LEN];\n>> +       int istatus, ostatus, fd = 0, hdrlen, dirlen, flush = 0;\n>> +       int ret = 0;\n>> +       git_hash_ctx c;\n>> +       struct strbuf tmp_file = STRBUF_INIT;\n>> +       struct strbuf filename = STRBUF_INIT;\n>> +\n>> +       /* Write tmpfile in objects dir, because oid is unknown */\n>> +       if (!dry_run) {\n>> +               strbuf_addstr(&filename, the_repository->objects->odb->path);\n>> +               strbuf_addch(&filename, '/');\n>> +               fd = create_tmpfile(&tmp_file, filename.buf);\n>> +               if (fd < 0) {\n>> +                       if (errno == EACCES)\n>> +                               ret = error(_(\"insufficient permission for adding an object to repository database %s\"),\n>> +                                       get_object_directory());\n>> +                       else\n>> +                               ret = error_errno(_(\"unable to create temporary file\"));\n>> +                       goto cleanup;\n>> +               }\n>> +       }\n>> +\n>> +       memset(&istream, 0, sizeof(istream));\n>> +       istream.next_out = buf;\n>> +       istream.avail_out = sizeof(buf);\n>> +       git_inflate_init(&istream);\n>> +\n>> +       if (!dry_run) {\n>> +               /* Set it up */\n>> +               git_deflate_init(&ostream, zlib_compression_level);\n>> +               ostream.next_out = compressed;\n>> +               ostream.avail_out = sizeof(compressed);\n>> +               the_hash_algo->init_fn(&c);\n>> +\n>> +               /* First header */\n>> +               hdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\" PRIuMAX, type,\n>> +                               (uintmax_t)len) + 1;\n>> +               ostream.next_in = (unsigned char *)hdr;\n>> +               ostream.avail_in = hdrlen;\n>> +               while (git_deflate(&ostream, 0) == Z_OK)\n>> +                       ; /* nothing */\n>> +               the_hash_algo->update_fn(&c, hdr, hdrlen);\n>> +       }\n>> +\n>> +       /* Then the data itself */\n>> +       do {\n>> +               unsigned char *last_out = istream.next_out;\n>> +               reader->fill(&istream);\n>> +               istatus = git_inflate(&istream, 0);\n>> +               if (istatus == Z_STREAM_END)\n>> +                       flush = Z_FINISH;\n>> +               reader->use(&istream);\n>> +               if (!dry_run)\n>> +                       ostatus = write_object_buffer(&ostream, &c, fd, compressed,\n>> +                                                     sizeof(compressed), last_out,\n>> +                                                     istream.next_out - last_out,\n>> +                                                     flush);\n>> +               istream.next_out = buf;\n>> +               istream.avail_out = sizeof(buf);\n>> +       } while (istatus == Z_OK);\n>> +\n>> +       if (istream.total_out != len || istatus != Z_STREAM_END)\n>> +               die( _(\"inflate returned %d\"), istatus);\n>> +       git_inflate_end(&istream);\n>> +\n>> +       if (dry_run)\n>> +               goto cleanup;\n>> +\n>> +       if (ostatus != Z_STREAM_END)\n>> +               die(_(\"unable to deflate new object (%d)\"), ostatus);\n>> +       ostatus = git_deflate_end_gently(&ostream);\n>> +       if (ostatus != Z_OK)\n>> +               die(_(\"deflateEnd on object failed (%d)\"), ostatus);\n>> +       the_hash_algo->final_fn(oid->hash, &c);\n>> +       close_loose_object(fd);\n>> +\n>> +       /* We get the oid now */\n>> +       loose_object_path(the_repository, &filename, oid);\n>> +\n>> +       dirlen = directory_size(filename.buf);\n>> +       if (dirlen) {\n>> +               struct strbuf dir = STRBUF_INIT;\n>> +               /*\n>> +                * Make sure the directory exists; note that the contents\n>> +                * of the buffer are undefined after mkstemp returns an\n>> +                * error, so we have to rewrite the whole buffer from\n>> +                * scratch.\n>> +                */\n>> +               strbuf_add(&dir, filename.buf, dirlen - 1);\n>> +               if (mkdir(dir.buf, 0777) && errno != EEXIST) {\n>> +                       unlink_or_warn(tmp_file.buf);\n>> +                       strbuf_release(&dir);\n>> +                       ret = -1;\n>> +                       goto cleanup;\n>> +               }\n>> +               strbuf_release(&dir);\n>> +       }\n>> +\n>> +       ret = finalize_object_file(tmp_file.buf, filename.buf);\n>> +\n>> +cleanup:\n>> +       strbuf_release(&tmp_file);\n>> +       strbuf_release(&filename);\n>> +       return ret;\n>> +}\n>> +\n>>  int hash_object_file_literally(const void *buf, unsigned long len,\n>>                                const char *type, struct object_id *oid,\n>>                                unsigned flags)\n>> diff --git a/object-store.h b/object-store.h\n>> index d24915ced1..12b113ef93 100644\n>> --- a/object-store.h\n>> +++ b/object-store.h\n>> @@ -33,6 +33,11 @@ struct object_directory {\n>>         char *path;\n>>  };\n>>\n>> +struct git_zstream_reader {\n>> +       void (*fill)(struct git_zstream *);\n>> +       void (*use)(struct git_zstream *);\n>> +};\n>> +\n>>  KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n>>         struct object_directory *, 1, fspathhash, fspatheq)\n>>\n>> @@ -225,6 +230,10 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n>>  int write_object_file(const void *buf, unsigned long len,\n>>                       const char *type, struct object_id *oid);\n>>\n>> +int write_stream_object_file(struct git_zstream_reader *reader,\n>> +                            unsigned long len, const char *type,\n>> +                            struct object_id *oid, int dry_run);\n>> +\n>>  int hash_object_file_literally(const void *buf, unsigned long len,\n>>                                const char *type, struct object_id *oid,\n>>                                unsigned flags);\n>> diff --git a/t/t5590-receive-unpack-objects.sh b/t/t5590-receive-unpack-objects.sh\n>> new file mode 100755\n>> index 0000000000..7e63dfc0db\n>> --- /dev/null\n>> +++ b/t/t5590-receive-unpack-objects.sh\n>> @@ -0,0 +1,92 @@\n>> +#!/bin/sh\n>> +#\n>> +# Copyright (c) 2021 Han Xin\n>> +#\n>> +\n>> +test_description='Test unpack-objects when receive pack'\n>> +\n>> +GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n>> +export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n>> +\n>> +. ./test-lib.sh\n>> +\n>> +test_expect_success \"create commit with big blobs (1.5 MB)\" '\n>> +       test-tool genrandom foo 1500000 >big-blob &&\n>> +       test_commit --append foo big-blob &&\n>> +       test-tool genrandom bar 1500000 >big-blob &&\n>> +       test_commit --append bar big-blob &&\n>> +       (\n>> +               cd .git &&\n>> +               find objects/?? -type f | sort\n>> +       ) >expect &&\n>> +       git repack -ad\n>> +'\n>> +\n>> +test_expect_success 'setup GIT_ALLOC_LIMIT to 1MB' '\n>> +       GIT_ALLOC_LIMIT=1m &&\n>> +       export GIT_ALLOC_LIMIT\n>> +'\n>> +\n>> +test_expect_success 'prepare dest repository' '\n>> +       git init --bare dest.git &&\n>> +       git -C dest.git config core.bigFileThreshold 2m &&\n>> +       git -C dest.git config receive.unpacklimit 100\n>> +'\n>> +\n>> +test_expect_success 'fail to push: cannot allocate' '\n>> +       test_must_fail git push dest.git HEAD 2>err &&\n>> +       test_i18ngrep \"remote: fatal: attempting to allocate\" err &&\n>> +       (\n>> +               cd dest.git &&\n>> +               find objects/?? -type f | sort\n>> +       ) >actual &&\n>> +       ! test_cmp expect actual\n>> +'\n>> +\n>> +test_expect_success 'set a lower bigfile threshold' '\n>> +       git -C dest.git config core.bigFileThreshold 1m\n>> +'\n>> +\n>> +test_expect_success 'unpack big object in stream' '\n>> +       git push dest.git HEAD &&\n>> +       git -C dest.git fsck &&\n>> +       (\n>> +               cd dest.git &&\n>> +               find objects/?? -type f | sort\n>> +       ) >actual &&\n>> +       test_cmp expect actual\n>> +'\n>> +\n>> +test_expect_success 'setup for unpack-objects dry-run test' '\n>> +       PACK=$(echo main | git pack-objects --progress --revs test) &&\n>> +       unset GIT_ALLOC_LIMIT &&\n>> +       git init --bare unpack-test.git\n>> +'\n>> +\n>> +test_expect_success 'unpack-objects dry-run with large threshold' '\n>> +       (\n>> +               cd unpack-test.git &&\n>> +               git config core.bigFileThreshold 2m &&\n>> +               git unpack-objects -n <../test-$PACK.pack\n>> +       ) &&\n>> +       (\n>> +               cd unpack-test.git &&\n>> +               find objects/ -type f\n>> +       ) >actual &&\n>> +       test_must_be_empty actual\n>> +'\n>> +\n>> +test_expect_success 'unpack-objects dry-run with small threshold' '\n>> +       (\n>> +               cd unpack-test.git &&\n>> +               git config core.bigFileThreshold 1m &&\n>> +               git unpack-objects -n <../test-$PACK.pack\n>> +       ) &&\n>> +       (\n>> +               cd unpack-test.git &&\n>> +               find objects/ -type f\n>> +       ) >actual &&\n>> +       test_must_be_empty actual\n>> +'\n>> +\n>> +test_done\n>> --\n>> 2.33.0.1.g09a6bb964f.dirty\n>>\n\n"},{"id":"440975","messageId":"20211112094010.73468-1-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211009082058.41138-1-chiyutianyi@gmail.com","subject":"[PATCH v2 1/6] object-file: refactor write_loose_object() to support inputstream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-11-12T09:40:05Z","receivedAt":"2021-11-12T09:41:44Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nRefactor write_loose_object() to support inputstream, in the same way\nthat zlib reading is chunked.\n\nUsing \"in_stream\" instead of \"void *buf\", we needn't to allocate enough\nmemory in advance, and only part of the contents will be read when\ncalled \"in_stream.read()\".\n\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c  | 50 ++++++++++++++++++++++++++++++++++++++++++++++----\n object-store.h |  5 +++++\n 2 files changed, 51 insertions(+), 4 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 02b7970274..1ad2cb579c 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1860,8 +1860,26 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n \treturn fd;\n }\n \n+struct input_data_from_buffer {\n+\tconst char *buf;\n+\tunsigned long len;\n+};\n+\n+static const char *read_input_stream_from_buffer(void *data, unsigned long *len)\n+{\n+\tstruct input_data_from_buffer *input = (struct input_data_from_buffer *)data;\n+\n+\tif (input->len == 0) {\n+\t\t*len = 0;\n+\t\treturn NULL;\n+\t}\n+\t*len = input->len;\n+\tinput->len = 0;\n+\treturn input->buf;\n+}\n+\n static int write_loose_object(const struct object_id *oid, char *hdr,\n-\t\t\t      int hdrlen, const void *buf, unsigned long len,\n+\t\t\t      int hdrlen, struct input_stream *in_stream,\n \t\t\t      time_t mtime, unsigned flags)\n {\n \tint fd, ret;\n@@ -1871,6 +1889,8 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tstruct object_id parano_oid;\n \tstatic struct strbuf tmp_file = STRBUF_INIT;\n \tstatic struct strbuf filename = STRBUF_INIT;\n+\tconst char *buf;\n+\tunsigned long len;\n \n \tloose_object_path(the_repository, &filename, oid);\n \n@@ -1898,6 +1918,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n \n \t/* Then the data itself.. */\n+\tbuf = in_stream->read(in_stream->data, &len);\n \tstream.next_in = (void *)buf;\n \tstream.avail_in = len;\n \tdo {\n@@ -1960,6 +1981,13 @@ int write_object_file_flags(const void *buf, unsigned long len,\n {\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen = sizeof(hdr);\n+\tstruct input_stream in_stream = {\n+\t\t.read = read_input_stream_from_buffer,\n+\t\t.data = (void *)&(struct input_data_from_buffer) {\n+\t\t\t.buf = buf,\n+\t\t\t.len = len,\n+\t\t},\n+\t};\n \n \t/* Normally if we have it in the pack then we do not bother writing\n \t * it out into .git/objects/??/?{38} file.\n@@ -1968,7 +1996,7 @@ int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t\t  &hdrlen);\n \tif (freshen_packed_object(oid) || freshen_loose_object(oid))\n \t\treturn 0;\n-\treturn write_loose_object(oid, hdr, hdrlen, buf, len, 0, flags);\n+\treturn write_loose_object(oid, hdr, hdrlen, &in_stream, 0, flags);\n }\n \n int hash_object_file_literally(const void *buf, unsigned long len,\n@@ -1977,6 +2005,13 @@ int hash_object_file_literally(const void *buf, unsigned long len,\n {\n \tchar *header;\n \tint hdrlen, status = 0;\n+\tstruct input_stream in_stream = {\n+\t\t.read = read_input_stream_from_buffer,\n+\t\t.data = (void *)&(struct input_data_from_buffer) {\n+\t\t\t.buf = buf,\n+\t\t\t.len = len,\n+\t\t},\n+\t};\n \n \t/* type string, SP, %lu of the length plus NUL must fit this */\n \thdrlen = strlen(type) + MAX_HEADER_LEN;\n@@ -1988,7 +2023,7 @@ int hash_object_file_literally(const void *buf, unsigned long len,\n \t\tgoto cleanup;\n \tif (freshen_packed_object(oid) || freshen_loose_object(oid))\n \t\tgoto cleanup;\n-\tstatus = write_loose_object(oid, header, hdrlen, buf, len, 0, 0);\n+\tstatus = write_loose_object(oid, header, hdrlen, &in_stream, 0, 0);\n \n cleanup:\n \tfree(header);\n@@ -2003,14 +2038,21 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen;\n \tint ret;\n+\tstruct input_data_from_buffer data;\n+\tstruct input_stream in_stream = {\n+\t\t.read = read_input_stream_from_buffer,\n+\t\t.data = &data,\n+\t};\n \n \tif (has_loose_object(oid))\n \t\treturn 0;\n \tbuf = read_object(the_repository, oid, &type, &len);\n \tif (!buf)\n \t\treturn error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n+\tdata.buf = buf;\n+\tdata.len = len;\n \thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n-\tret = write_loose_object(oid, hdr, hdrlen, buf, len, mtime, 0);\n+\tret = write_loose_object(oid, hdr, hdrlen, &in_stream, mtime, 0);\n \tfree(buf);\n \n \treturn ret;\ndiff --git a/object-store.h b/object-store.h\nindex 952efb6a4b..f1b67e9100 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -34,6 +34,11 @@ struct object_directory {\n \tchar *path;\n };\n \n+struct input_stream {\n+\tconst char *(*read)(void* data, unsigned long *len);\n+\tvoid *data;\n+};\n+\n KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n \tstruct object_directory *, 1, fspathhash, fspatheq)\n \n-- \n2.33.1.44.g9344627884.agit.6.5.4\n\n"},{"id":"440976","messageId":"20211112094010.73468-2-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211009082058.41138-1-chiyutianyi@gmail.com","subject":"[PATCH v2 2/6] object-file.c: add dry_run mode for write_loose_object()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-11-12T09:40:06Z","receivedAt":"2021-11-12T09:42:02Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWe will use \"write_loose_object()\" later to handle large blob object,\nwhich needs to work in dry_run mode.\n\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c | 32 +++++++++++++++++++-------------\n 1 file changed, 19 insertions(+), 13 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 1ad2cb579c..b0838c847e 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1880,9 +1880,10 @@ static const char *read_input_stream_from_buffer(void *data, unsigned long *len)\n \n static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\t\t      int hdrlen, struct input_stream *in_stream,\n+\t\t\t      int dry_run,\n \t\t\t      time_t mtime, unsigned flags)\n {\n-\tint fd, ret;\n+\tint fd, ret = 0;\n \tunsigned char compressed[4096];\n \tgit_zstream stream;\n \tgit_hash_ctx c;\n@@ -1894,14 +1895,16 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \n \tloose_object_path(the_repository, &filename, oid);\n \n-\tfd = create_tmpfile(&tmp_file, filename.buf);\n-\tif (fd < 0) {\n-\t\tif (flags & HASH_SILENT)\n-\t\t\treturn -1;\n-\t\telse if (errno == EACCES)\n-\t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n-\t\telse\n-\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n+\tif (!dry_run) {\n+\t\tfd = create_tmpfile(&tmp_file, filename.buf);\n+\t\tif (fd < 0) {\n+\t\t\tif (flags & HASH_SILENT)\n+\t\t\t\treturn -1;\n+\t\t\telse if (errno == EACCES)\n+\t\t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n+\t\t\telse\n+\t\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n+\t\t}\n \t}\n \n \t/* Set it up */\n@@ -1925,7 +1928,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tunsigned char *in0 = stream.next_in;\n \t\tret = git_deflate(&stream, Z_FINISH);\n \t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n-\t\tif (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n+\t\tif (!dry_run && write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n \t\t\tdie(_(\"unable to write loose object file\"));\n \t\tstream.next_out = compressed;\n \t\tstream.avail_out = sizeof(compressed);\n@@ -1943,6 +1946,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n \n+\tif (dry_run)\n+\t\treturn 0;\n+\n \tclose_loose_object(fd);\n \n \tif (mtime) {\n@@ -1996,7 +2002,7 @@ int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t\t  &hdrlen);\n \tif (freshen_packed_object(oid) || freshen_loose_object(oid))\n \t\treturn 0;\n-\treturn write_loose_object(oid, hdr, hdrlen, &in_stream, 0, flags);\n+\treturn write_loose_object(oid, hdr, hdrlen, &in_stream, 0, 0, flags);\n }\n \n int hash_object_file_literally(const void *buf, unsigned long len,\n@@ -2023,7 +2029,7 @@ int hash_object_file_literally(const void *buf, unsigned long len,\n \t\tgoto cleanup;\n \tif (freshen_packed_object(oid) || freshen_loose_object(oid))\n \t\tgoto cleanup;\n-\tstatus = write_loose_object(oid, header, hdrlen, &in_stream, 0, 0);\n+\tstatus = write_loose_object(oid, header, hdrlen, &in_stream, 0, 0, 0);\n \n cleanup:\n \tfree(header);\n@@ -2052,7 +2058,7 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n \tdata.buf = buf;\n \tdata.len = len;\n \thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n-\tret = write_loose_object(oid, hdr, hdrlen, &in_stream, mtime, 0);\n+\tret = write_loose_object(oid, hdr, hdrlen, &in_stream, 0, mtime, 0);\n \tfree(buf);\n \n \treturn ret;\n-- \n2.33.1.44.g9344627884.agit.6.5.4\n\n"},{"id":"440977","messageId":"20211112094010.73468-3-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211009082058.41138-1-chiyutianyi@gmail.com","subject":"[PATCH v2 3/6] object-file.c: handle nil oid in write_loose_object()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-11-12T09:40:07Z","receivedAt":"2021-11-12T09:42:05Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWhen read input stream, oid can't get before reading all, and it will be\nfilled after reading.\n\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c | 34 ++++++++++++++++++++++++++++++++--\n 1 file changed, 32 insertions(+), 2 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex b0838c847e..8393659f0d 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1893,7 +1893,13 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tconst char *buf;\n \tunsigned long len;\n \n-\tloose_object_path(the_repository, &filename, oid);\n+\tif (is_null_oid(oid)) {\n+\t\t/* When oid is not determined, save tmp file to odb path. */\n+\t\tstrbuf_reset(&filename);\n+\t\tstrbuf_addstr(&filename, the_repository->objects->odb->path);\n+\t\tstrbuf_addch(&filename, '/');\n+\t} else\n+\t\tloose_object_path(the_repository, &filename, oid);\n \n \tif (!dry_run) {\n \t\tfd = create_tmpfile(&tmp_file, filename.buf);\n@@ -1942,7 +1948,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n \t\t    ret);\n \tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n-\tif (!oideq(oid, &parano_oid))\n+\tif (!is_null_oid(oid) && !oideq(oid, &parano_oid))\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n \n@@ -1951,6 +1957,30 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \n \tclose_loose_object(fd);\n \n+\tif (is_null_oid(oid)) {\n+\t\tint dirlen;\n+\n+\t\t/* copy oid */\n+\t\toidcpy((struct object_id *)oid, &parano_oid);\n+\t\t/* We get the oid now */\n+\t\tloose_object_path(the_repository, &filename, oid);\n+\n+\t\tdirlen = directory_size(filename.buf);\n+\t\tif (dirlen) {\n+\t\t\tstruct strbuf dir = STRBUF_INIT;\n+\t\t\t/*\n+\t\t\t * Make sure the directory exists; note that the\n+\t\t\t * contents of the buffer are undefined after mkstemp\n+\t\t\t * returns an error, so we have to rewrite the whole\n+\t\t\t * buffer from scratch.\n+\t\t\t */\n+\t\t\tstrbuf_reset(&dir);\n+\t\t\tstrbuf_add(&dir, filename.buf, dirlen - 1);\n+\t\t\tif (mkdir(dir.buf, 0777) && errno != EEXIST)\n+\t\t\t\treturn -1;\n+\t\t}\n+\t}\n+\n \tif (mtime) {\n \t\tstruct utimbuf utb;\n \t\tutb.actime = mtime;\n-- \n2.33.1.44.g9344627884.agit.6.5.4\n\n"},{"id":"440978","messageId":"20211112094010.73468-4-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211009082058.41138-1-chiyutianyi@gmail.com","subject":"[PATCH v2 4/6] object-file.c: read input stream repeatedly in write_loose_object()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-11-12T09:40:08Z","receivedAt":"2021-11-12T09:42:11Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nRead input stream repeatedly in write_loose_object() unless reach the\nend, so that we can divide the large blob write into many small blocks.\n\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c | 14 +++++++++-----\n 1 file changed, 9 insertions(+), 5 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 8393659f0d..e333448c54 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1891,7 +1891,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tstatic struct strbuf tmp_file = STRBUF_INIT;\n \tstatic struct strbuf filename = STRBUF_INIT;\n \tconst char *buf;\n-\tunsigned long len;\n+\tint flush = 0;\n \n \tif (is_null_oid(oid)) {\n \t\t/* When oid is not determined, save tmp file to odb path. */\n@@ -1927,12 +1927,16 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n \n \t/* Then the data itself.. */\n-\tbuf = in_stream->read(in_stream->data, &len);\n-\tstream.next_in = (void *)buf;\n-\tstream.avail_in = len;\n \tdo {\n \t\tunsigned char *in0 = stream.next_in;\n-\t\tret = git_deflate(&stream, Z_FINISH);\n+\t\tif (!stream.avail_in) {\n+\t\t\tif ((buf = in_stream->read(in_stream->data, &stream.avail_in))) {\n+\t\t\t\tstream.next_in = (void *)buf;\n+\t\t\t\tin0 = (unsigned char *)buf;\n+\t\t\t} else\n+\t\t\t\tflush = Z_FINISH;\n+\t\t}\n+\t\tret = git_deflate(&stream, flush);\n \t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n \t\tif (!dry_run && write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n \t\t\tdie(_(\"unable to write loose object file\"));\n-- \n2.33.1.44.g9344627884.agit.6.5.4\n\n"},{"id":"440979","messageId":"20211112094010.73468-5-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211009082058.41138-1-chiyutianyi@gmail.com","subject":"[PATCH v2 5/6] object-store.h: add write_loose_object()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-11-12T09:40:09Z","receivedAt":"2021-11-12T09:42:14Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nFor large loose object files, that should be possible to stream it\ndirect to disk with \"write_loose_object()\".\nUnlike \"write_object_file()\", you need to implement an \"input_stream\"\ninstead of giving void *buf.\n\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c  | 8 ++++----\n object-store.h | 5 +++++\n 2 files changed, 9 insertions(+), 4 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex e333448c54..60eb29db97 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1878,10 +1878,10 @@ static const char *read_input_stream_from_buffer(void *data, unsigned long *len)\n \treturn input->buf;\n }\n \n-static int write_loose_object(const struct object_id *oid, char *hdr,\n-\t\t\t      int hdrlen, struct input_stream *in_stream,\n-\t\t\t      int dry_run,\n-\t\t\t      time_t mtime, unsigned flags)\n+int write_loose_object(const struct object_id *oid, char *hdr,\n+\t\t       int hdrlen, struct input_stream *in_stream,\n+\t\t       int dry_run,\n+\t\t       time_t mtime, unsigned flags)\n {\n \tint fd, ret = 0;\n \tunsigned char compressed[4096];\ndiff --git a/object-store.h b/object-store.h\nindex f1b67e9100..f6faa8d6d3 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -228,6 +228,11 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n \t\t     unsigned long len, const char *type,\n \t\t     struct object_id *oid);\n \n+int write_loose_object(const struct object_id *oid, char *hdr,\n+\t\t       int hdrlen, struct input_stream *in_stream,\n+\t\t       int dry_run,\n+\t\t       time_t mtime, unsigned flags);\n+\n int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t    const char *type, struct object_id *oid,\n \t\t\t    unsigned flags);\n-- \n2.33.1.44.g9344627884.agit.6.5.4\n\n"},{"id":"440980","messageId":"20211112094010.73468-6-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211009082058.41138-1-chiyutianyi@gmail.com","subject":"[PATCH v2 6/6] unpack-objects: unpack large object in stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-11-12T09:40:10Z","receivedAt":"2021-11-12T09:42:15Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWhen calling \"unpack_non_delta_entry()\", will allocate full memory for\nthe whole size of the unpacked object and write the buffer to loose file\non disk. This may lead to OOM for the git-unpack-objects process when\nunpacking a very large object.\n\nIn function \"unpack_delta_entry()\", will also allocate full memory to\nbuffer the whole delta, but since there will be no delta for an object\nlarger than \"core.bigFileThreshold\", this issue is moderate.\n\nTo resolve the OOM issue in \"git-unpack-objects\", we can unpack large\nobject to file in stream, and use \"core.bigFileThreshold\" to avoid OOM\nlimits when called \"get_data()\".\n\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n builtin/unpack-objects.c          | 76 ++++++++++++++++++++++++-\n t/t5590-receive-unpack-objects.sh | 92 +++++++++++++++++++++++++++++++\n 2 files changed, 167 insertions(+), 1 deletion(-)\n create mode 100755 t/t5590-receive-unpack-objects.sh\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 4a9466295b..6c757d823b 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -320,11 +320,85 @@ static void added_object(unsigned nr, enum object_type type,\n \t}\n }\n \n+struct input_data_from_zstream {\n+\tgit_zstream *zstream;\n+\tunsigned char buf[4096];\n+\tint status;\n+};\n+\n+static const char *read_inflate_in_stream(void *data, unsigned long *readlen)\n+{\n+\tstruct input_data_from_zstream *input = data;\n+\tgit_zstream *zstream = input->zstream;\n+\tvoid *in = fill(1);\n+\n+\tif (!len || input->status == Z_STREAM_END) {\n+\t\t*readlen = 0;\n+\t\treturn NULL;\n+\t}\n+\n+\tzstream->next_out = input->buf;\n+\tzstream->avail_out = sizeof(input->buf);\n+\tzstream->next_in = in;\n+\tzstream->avail_in = len;\n+\n+\tinput->status = git_inflate(zstream, 0);\n+\tuse(len - zstream->avail_in);\n+\t*readlen = sizeof(input->buf) - zstream->avail_out;\n+\n+\treturn (const char *)input->buf;\n+}\n+\n+static void write_stream_blob(unsigned nr, unsigned long size)\n+{\n+\tchar hdr[32];\n+\tint hdrlen;\n+\tgit_zstream zstream;\n+\tstruct input_data_from_zstream data;\n+\tstruct input_stream in_stream = {\n+\t\t.read = read_inflate_in_stream,\n+\t\t.data = &data,\n+\t};\n+\tstruct object_id *oid = &obj_list[nr].oid;\n+\tint ret;\n+\n+\tmemset(&zstream, 0, sizeof(zstream));\n+\tmemset(&data, 0, sizeof(data));\n+\tdata.zstream = &zstream;\n+\tgit_inflate_init(&zstream);\n+\n+\t/* Generate the header */\n+\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX, type_name(OBJ_BLOB), (uintmax_t)size) + 1;\n+\n+\tif ((ret = write_loose_object(oid, hdr, hdrlen, &in_stream, dry_run, 0, 0)))\n+\t\tdie(_(\"failed to write object in stream %d\"), ret);\n+\n+\tif (zstream.total_out != size || data.status != Z_STREAM_END)\n+\t\tdie(_(\"inflate returned %d\"), data.status);\n+\tgit_inflate_end(&zstream);\n+\n+\tif (strict && !dry_run) {\n+\t\tstruct blob *blob = lookup_blob(the_repository, oid);\n+\t\tif (blob)\n+\t\t\tblob->object.flags |= FLAG_WRITTEN;\n+\t\telse\n+\t\t\tdie(\"invalid blob object from stream\");\n+\t}\n+\tobj_list[nr].obj = NULL;\n+}\n+\n static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n \t\t\t\t   unsigned nr)\n {\n-\tvoid *buf = get_data(size);\n+\tvoid *buf;\n+\n+\t/* Write large blob in stream without allocating full buffer. */\n+\tif (type == OBJ_BLOB && size > big_file_threshold) {\n+\t\twrite_stream_blob(nr, size);\n+\t\treturn;\n+\t}\n \n+\tbuf = get_data(size);\n \tif (!dry_run && buf)\n \t\twrite_object(nr, type, buf, size);\n \telse\ndiff --git a/t/t5590-receive-unpack-objects.sh b/t/t5590-receive-unpack-objects.sh\nnew file mode 100755\nindex 0000000000..7e63dfc0db\n--- /dev/null\n+++ b/t/t5590-receive-unpack-objects.sh\n@@ -0,0 +1,92 @@\n+#!/bin/sh\n+#\n+# Copyright (c) 2021 Han Xin\n+#\n+\n+test_description='Test unpack-objects when receive pack'\n+\n+GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n+export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n+\n+. ./test-lib.sh\n+\n+test_expect_success \"create commit with big blobs (1.5 MB)\" '\n+\ttest-tool genrandom foo 1500000 >big-blob &&\n+\ttest_commit --append foo big-blob &&\n+\ttest-tool genrandom bar 1500000 >big-blob &&\n+\ttest_commit --append bar big-blob &&\n+\t(\n+\t\tcd .git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >expect &&\n+\tgit repack -ad\n+'\n+\n+test_expect_success 'setup GIT_ALLOC_LIMIT to 1MB' '\n+\tGIT_ALLOC_LIMIT=1m &&\n+\texport GIT_ALLOC_LIMIT\n+'\n+\n+test_expect_success 'prepare dest repository' '\n+\tgit init --bare dest.git &&\n+\tgit -C dest.git config core.bigFileThreshold 2m &&\n+\tgit -C dest.git config receive.unpacklimit 100\n+'\n+\n+test_expect_success 'fail to push: cannot allocate' '\n+\ttest_must_fail git push dest.git HEAD 2>err &&\n+\ttest_i18ngrep \"remote: fatal: attempting to allocate\" err &&\n+\t(\n+\t\tcd dest.git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >actual &&\n+\t! test_cmp expect actual\n+'\n+\n+test_expect_success 'set a lower bigfile threshold' '\n+\tgit -C dest.git config core.bigFileThreshold 1m\n+'\n+\n+test_expect_success 'unpack big object in stream' '\n+\tgit push dest.git HEAD &&\n+\tgit -C dest.git fsck &&\n+\t(\n+\t\tcd dest.git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'setup for unpack-objects dry-run test' '\n+\tPACK=$(echo main | git pack-objects --progress --revs test) &&\n+\tunset GIT_ALLOC_LIMIT &&\n+\tgit init --bare unpack-test.git\n+'\n+\n+test_expect_success 'unpack-objects dry-run with large threshold' '\n+\t(\n+\t\tcd unpack-test.git &&\n+\t\tgit config core.bigFileThreshold 2m &&\n+\t\tgit unpack-objects -n <../test-$PACK.pack\n+\t) &&\n+\t(\n+\t\tcd unpack-test.git &&\n+\t\tfind objects/ -type f\n+\t) >actual &&\n+\ttest_must_be_empty actual\n+'\n+\n+test_expect_success 'unpack-objects dry-run with small threshold' '\n+\t(\n+\t\tcd unpack-test.git &&\n+\t\tgit config core.bigFileThreshold 1m &&\n+\t\tgit unpack-objects -n <../test-$PACK.pack\n+\t) &&\n+\t(\n+\t\tcd unpack-test.git &&\n+\t\tfind objects/ -type f\n+\t) >actual &&\n+\ttest_must_be_empty actual\n+'\n+\n+test_done\n-- \n2.33.1.44.g9344627884.agit.6.5.4\n\n"},{"id":"441565","messageId":"CANYiYbHiTnhpzyWLakVZ6tfmG0pKO=qHdZsgaocX6eJ=PN_06g@mail.gmail.com","threadId":"56672","inReplyTo":"20211112094010.73468-1-chiyutianyi@gmail.com","subject":"Re: [PATCH v2 1/6] object-file: refactor write_loose_object() to support inputstream","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2021-11-18T04:59:04Z","receivedAt":"2021-11-18T04:59:18Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"On Fri, Nov 12, 2021 at 5:43 PM Han Xin <chiyutianyi@gmail.com> wrote:\n>\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nIt would be better to provide a cover letter describing changes in v2, such as:\n\n* Make \"write_loose_object()\" a public method, so we can\n   reuse it in \"unpack_non_delta_entry()\".\n   (But I doubt we can use \"write_object_file_flags()\" public\n     function, without make this change.)\n\n* Add an new interface \"input_stream\" as an argument for\n   \"write_loose_object()\", so that we can feed data to\n   \"write_loose_object()\" from buffer or from zlib stream.\n\n> Refactor write_loose_object() to support inputstream, in the same way\n> that zlib reading is chunked.\n\nIn the beginning of your commit log, you should describe the problem, such as:\n\nWe used to read the full content of a blob into buffer in\n\"unpack_non_delta_entry()\" by calling:\n\n    void *buf = get_data(size);\n\nThis will consume lots of memory for a very big blob object.\n\n> Using \"in_stream\" instead of \"void *buf\", we needn't to allocate enough\n> memory in advance, and only part of the contents will be read when\n> called \"in_stream.read()\".\n>\n> Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  object-file.c  | 50 ++++++++++++++++++++++++++++++++++++++++++++++----\n>  object-store.h |  5 +++++\n>  2 files changed, 51 insertions(+), 4 deletions(-)\n>\n> diff --git a/object-file.c b/object-file.c\n> index 02b7970274..1ad2cb579c 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1860,8 +1860,26 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n>         return fd;\n>  }\n>\n> +struct input_data_from_buffer {\n> +       const char *buf;\n> +       unsigned long len;\n> +};\n> +\n> +static const char *read_input_stream_from_buffer(void *data, unsigned long *len)\n\nUse \"const void *\" for the type of return variable, just like input\nargument for write_loose_object()?\n\n> +{\n> +       struct input_data_from_buffer *input = (struct input_data_from_buffer *)data;\n> +\n> +       if (input->len == 0) {\n> +               *len = 0;\n> +               return NULL;\n> +       }\n> +       *len = input->len;\n> +       input->len = 0;\n> +       return input->buf;\n> +}\n> +\n>  static int write_loose_object(const struct object_id *oid, char *hdr,\n> -                             int hdrlen, const void *buf, unsigned long len,\n> +                             int hdrlen, struct input_stream *in_stream,\n>                               time_t mtime, unsigned flags)\n>  {\n>         int fd, ret;\n> @@ -1871,6 +1889,8 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>         struct object_id parano_oid;\n>         static struct strbuf tmp_file = STRBUF_INIT;\n>         static struct strbuf filename = STRBUF_INIT;\n> +       const char *buf;\n\nCan we use the same prototype as the original:  \"const void *buf\" ?\n\n> +       unsigned long len;\n>\n>         loose_object_path(the_repository, &filename, oid);\n>\n> @@ -1898,6 +1918,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>         the_hash_algo->update_fn(&c, hdr, hdrlen);\n>\n>         /* Then the data itself.. */\n> +       buf = in_stream->read(in_stream->data, &len);\n>         stream.next_in = (void *)buf;\n>         stream.avail_in = len;\n>         do {\n> @@ -1960,6 +1981,13 @@ int write_object_file_flags(const void *buf, unsigned long len,\n>  {\n>         char hdr[MAX_HEADER_LEN];\n>         int hdrlen = sizeof(hdr);\n> +       struct input_stream in_stream = {\n> +               .read = read_input_stream_from_buffer,\n> +               .data = (void *)&(struct input_data_from_buffer) {\n> +                       .buf = buf,\n> +                       .len = len,\n> +               },\n> +       };\n>\n>         /* Normally if we have it in the pack then we do not bother writing\n>          * it out into .git/objects/??/?{38} file.\n> @@ -1968,7 +1996,7 @@ int write_object_file_flags(const void *buf, unsigned long len,\n>                                   &hdrlen);\n>         if (freshen_packed_object(oid) || freshen_loose_object(oid))\n>                 return 0;\n> -       return write_loose_object(oid, hdr, hdrlen, buf, len, 0, flags);\n> +       return write_loose_object(oid, hdr, hdrlen, &in_stream, 0, flags);\n>  }\n>\n>  int hash_object_file_literally(const void *buf, unsigned long len,\n> @@ -1977,6 +2005,13 @@ int hash_object_file_literally(const void *buf, unsigned long len,\n>  {\n>         char *header;\n>         int hdrlen, status = 0;\n> +       struct input_stream in_stream = {\n> +               .read = read_input_stream_from_buffer,\n> +               .data = (void *)&(struct input_data_from_buffer) {\n> +                       .buf = buf,\n> +                       .len = len,\n> +               },\n> +       };\n>\n>         /* type string, SP, %lu of the length plus NUL must fit this */\n>         hdrlen = strlen(type) + MAX_HEADER_LEN;\n> @@ -1988,7 +2023,7 @@ int hash_object_file_literally(const void *buf, unsigned long len,\n>                 goto cleanup;\n>         if (freshen_packed_object(oid) || freshen_loose_object(oid))\n>                 goto cleanup;\n> -       status = write_loose_object(oid, header, hdrlen, buf, len, 0, 0);\n> +       status = write_loose_object(oid, header, hdrlen, &in_stream, 0, 0);\n>\n>  cleanup:\n>         free(header);\n> @@ -2003,14 +2038,21 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n>         char hdr[MAX_HEADER_LEN];\n>         int hdrlen;\n>         int ret;\n> +       struct input_data_from_buffer data;\n> +       struct input_stream in_stream = {\n> +               .read = read_input_stream_from_buffer,\n> +               .data = &data,\n> +       };\n>\n>         if (has_loose_object(oid))\n>                 return 0;\n>         buf = read_object(the_repository, oid, &type, &len);\n>         if (!buf)\n>                 return error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n> +       data.buf = buf;\n> +       data.len = len;\n>         hdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n> -       ret = write_loose_object(oid, hdr, hdrlen, buf, len, mtime, 0);\n> +       ret = write_loose_object(oid, hdr, hdrlen, &in_stream, mtime, 0);\n>         free(buf);\n>\n>         return ret;\n> diff --git a/object-store.h b/object-store.h\n> index 952efb6a4b..f1b67e9100 100644\n> --- a/object-store.h\n> +++ b/object-store.h\n> @@ -34,6 +34,11 @@ struct object_directory {\n>         char *path;\n>  };\n>\n> +struct input_stream {\n> +       const char *(*read)(void* data, unsigned long *len);\n> +       void *data;\n> +};\n> +\n>  KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n>         struct object_directory *, 1, fspathhash, fspatheq)\n>\n> --\n> 2.33.1.44.g9344627884.agit.6.5.4\n>\n"},{"id":"441568","messageId":"CANYiYbHvpVH7CvhGOUZojVkXd7eztC+wQ_R8=Ta9hPwjjAvWoQ@mail.gmail.com","threadId":"56672","inReplyTo":"20211112094010.73468-2-chiyutianyi@gmail.com","subject":"Re: [PATCH v2 2/6] object-file.c: add dry_run mode for write_loose_object()","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2021-11-18T05:42:36Z","receivedAt":"2021-11-18T05:42:51Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"On Fri, Nov 12, 2021 at 5:42 PM Han Xin <chiyutianyi@gmail.com> wrote:\n>\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> We will use \"write_loose_object()\" later to handle large blob object,\n> which needs to work in dry_run mode.\n\nThe dry_run mode comes from \"builtin/unpack-object.c\", throw the\nbuffer read from \"get_data()\".\nSo why not add \"dry_run\" to \"get_data()\" instead?\n\nIf we have a dry_run version of get_data, such as \"get_data(size,\ndry_run)\", we do not have to add dry_run mode for ”\nwrite_loose_object()\".\n\nSee: git grep -A5 get_data builtin/unpack-objects.c\nbuiltin/unpack-objects.c:       void *buf = get_data(size);\nbuiltin/unpack-objects.c-\nbuiltin/unpack-objects.c-       if (!dry_run && buf)\nbuiltin/unpack-objects.c-               write_object(nr, type, buf, size);\nbuiltin/unpack-objects.c-       else\nbuiltin/unpack-objects.c-               free(buf);\n--\nbuiltin/unpack-objects.c:               delta_data = get_data(delta_size);\nbuiltin/unpack-objects.c-               if (dry_run || !delta_data) {\nbuiltin/unpack-objects.c-                       free(delta_data);\nbuiltin/unpack-objects.c-                       return;\nbuiltin/unpack-objects.c-               }\n--\nbuiltin/unpack-objects.c:               delta_data = get_data(delta_size);\nbuiltin/unpack-objects.c-               if (dry_run || !delta_data) {\nbuiltin/unpack-objects.c-                       free(delta_data);\nbuiltin/unpack-objects.c-                       return;\nbuiltin/unpack-objects.c-               }\n\n\n> Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  object-file.c | 32 +++++++++++++++++++-------------\n>  1 file changed, 19 insertions(+), 13 deletions(-)\n>\n> diff --git a/object-file.c b/object-file.c\n> index 1ad2cb579c..b0838c847e 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1880,9 +1880,10 @@ static const char *read_input_stream_from_buffer(void *data, unsigned long *len)\n>\n>  static int write_loose_object(const struct object_id *oid, char *hdr,\n>                               int hdrlen, struct input_stream *in_stream,\n> +                             int dry_run,\n>                               time_t mtime, unsigned flags)\n>  {\n> -       int fd, ret;\n> +       int fd, ret = 0;\n>         unsigned char compressed[4096];\n>         git_zstream stream;\n>         git_hash_ctx c;\n> @@ -1894,14 +1895,16 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>\n>         loose_object_path(the_repository, &filename, oid);\n>\n> -       fd = create_tmpfile(&tmp_file, filename.buf);\n> -       if (fd < 0) {\n> -               if (flags & HASH_SILENT)\n> -                       return -1;\n> -               else if (errno == EACCES)\n> -                       return error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n> -               else\n> -                       return error_errno(_(\"unable to create temporary file\"));\n> +       if (!dry_run) {\n> +               fd = create_tmpfile(&tmp_file, filename.buf);\n> +               if (fd < 0) {\n> +                       if (flags & HASH_SILENT)\n> +                               return -1;\n> +                       else if (errno == EACCES)\n> +                               return error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n> +                       else\n> +                               return error_errno(_(\"unable to create temporary file\"));\n> +               }\n>         }\n>\n>         /* Set it up */\n> @@ -1925,7 +1928,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>                 unsigned char *in0 = stream.next_in;\n>                 ret = git_deflate(&stream, Z_FINISH);\n>                 the_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n> -               if (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n> +               if (!dry_run && write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n>                         die(_(\"unable to write loose object file\"));\n>                 stream.next_out = compressed;\n>                 stream.avail_out = sizeof(compressed);\n> @@ -1943,6 +1946,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>                 die(_(\"confused by unstable object source data for %s\"),\n>                     oid_to_hex(oid));\n>\n> +       if (dry_run)\n> +               return 0;\n> +\n>         close_loose_object(fd);\n>\n>         if (mtime) {\n> @@ -1996,7 +2002,7 @@ int write_object_file_flags(const void *buf, unsigned long len,\n>                                   &hdrlen);\n>         if (freshen_packed_object(oid) || freshen_loose_object(oid))\n>                 return 0;\n> -       return write_loose_object(oid, hdr, hdrlen, &in_stream, 0, flags);\n> +       return write_loose_object(oid, hdr, hdrlen, &in_stream, 0, 0, flags);\n>  }\n>\n>  int hash_object_file_literally(const void *buf, unsigned long len,\n> @@ -2023,7 +2029,7 @@ int hash_object_file_literally(const void *buf, unsigned long len,\n>                 goto cleanup;\n>         if (freshen_packed_object(oid) || freshen_loose_object(oid))\n>                 goto cleanup;\n> -       status = write_loose_object(oid, header, hdrlen, &in_stream, 0, 0);\n> +       status = write_loose_object(oid, header, hdrlen, &in_stream, 0, 0, 0);\n>\n>  cleanup:\n>         free(header);\n> @@ -2052,7 +2058,7 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n>         data.buf = buf;\n>         data.len = len;\n>         hdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n> -       ret = write_loose_object(oid, hdr, hdrlen, &in_stream, mtime, 0);\n> +       ret = write_loose_object(oid, hdr, hdrlen, &in_stream, 0, mtime, 0);\n>         free(buf);\n>\n>         return ret;\n> --\n> 2.33.1.44.g9344627884.agit.6.5.4\n>\n"},{"id":"441569","messageId":"CANYiYbESTJKOGQ0_70X6V6JuwaDZTL-a3fLqyH_SapOoQJw9BQ@mail.gmail.com","threadId":"56672","inReplyTo":"20211112094010.73468-3-chiyutianyi@gmail.com","subject":"Re: [PATCH v2 3/6] object-file.c: handle nil oid in write_loose_object()","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2021-11-18T05:49:45Z","receivedAt":"2021-11-18T05:50:01Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"On Fri, Nov 12, 2021 at 5:42 PM Han Xin <chiyutianyi@gmail.com> wrote:\n>\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> When read input stream, oid can't get before reading all, and it will be\n> filled after reading.\n\nUnder what circumstances is the oid a null oid?  Can we get the oid\nfrom “obj_list[nr].oid” ?\nSee unpack_non_delta_entry() of builtin/unpack-objects.c.\n\n> Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  object-file.c | 34 ++++++++++++++++++++++++++++++++--\n>  1 file changed, 32 insertions(+), 2 deletions(-)\n>\n> diff --git a/object-file.c b/object-file.c\n> index b0838c847e..8393659f0d 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1893,7 +1893,13 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>         const char *buf;\n>         unsigned long len;\n>\n> -       loose_object_path(the_repository, &filename, oid);\n> +       if (is_null_oid(oid)) {\n> +               /* When oid is not determined, save tmp file to odb path. */\n> +               strbuf_reset(&filename);\n> +               strbuf_addstr(&filename, the_repository->objects->odb->path);\n> +               strbuf_addch(&filename, '/');\n> +       } else\n> +               loose_object_path(the_repository, &filename, oid);\n>\n>         if (!dry_run) {\n>                 fd = create_tmpfile(&tmp_file, filename.buf);\n> @@ -1942,7 +1948,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>                 die(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n>                     ret);\n>         the_hash_algo->final_oid_fn(&parano_oid, &c);\n> -       if (!oideq(oid, &parano_oid))\n> +       if (!is_null_oid(oid) && !oideq(oid, &parano_oid))\n>                 die(_(\"confused by unstable object source data for %s\"),\n>                     oid_to_hex(oid));\n>\n> @@ -1951,6 +1957,30 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>\n>         close_loose_object(fd);\n>\n> +       if (is_null_oid(oid)) {\n> +               int dirlen;\n> +\n> +               /* copy oid */\n> +               oidcpy((struct object_id *)oid, &parano_oid);\n> +               /* We get the oid now */\n> +               loose_object_path(the_repository, &filename, oid);\n> +\n> +               dirlen = directory_size(filename.buf);\n> +               if (dirlen) {\n> +                       struct strbuf dir = STRBUF_INIT;\n> +                       /*\n> +                        * Make sure the directory exists; note that the\n> +                        * contents of the buffer are undefined after mkstemp\n> +                        * returns an error, so we have to rewrite the whole\n> +                        * buffer from scratch.\n> +                        */\n> +                       strbuf_reset(&dir);\n> +                       strbuf_add(&dir, filename.buf, dirlen - 1);\n> +                       if (mkdir(dir.buf, 0777) && errno != EEXIST)\n> +                               return -1;\n> +               }\n> +       }\n> +\n>         if (mtime) {\n>                 struct utimbuf utb;\n>                 utb.actime = mtime;\n> --\n> 2.33.1.44.g9344627884.agit.6.5.4\n>\n"},{"id":"441570","messageId":"CANYiYbF83iHZb=kr-yAwp8rBPx47e6O=80Avp23092f8J1m2RA@mail.gmail.com","threadId":"56672","inReplyTo":"20211112094010.73468-4-chiyutianyi@gmail.com","subject":"Re: [PATCH v2 4/6] object-file.c: read input stream repeatedly in write_loose_object()","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2021-11-18T05:56:25Z","receivedAt":"2021-11-18T05:56:44Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"On Fri, Nov 12, 2021 at 5:43 PM Han Xin <chiyutianyi@gmail.com> wrote:\n>\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> Read input stream repeatedly in write_loose_object() unless reach the\n> end, so that we can divide the large blob write into many small blocks.\n\nIn order to prepare the stream version of \"write_loose_object()\", we need ...\n\n>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  object-file.c | 14 +++++++++-----\n>  1 file changed, 9 insertions(+), 5 deletions(-)\n>\n> diff --git a/object-file.c b/object-file.c\n> index 8393659f0d..e333448c54 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1891,7 +1891,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>         static struct strbuf tmp_file = STRBUF_INIT;\n>         static struct strbuf filename = STRBUF_INIT;\n>         const char *buf;\n> -       unsigned long len;\n> +       int flush = 0;\n>\n>         if (is_null_oid(oid)) {\n>                 /* When oid is not determined, save tmp file to odb path. */\n> @@ -1927,12 +1927,16 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>         the_hash_algo->update_fn(&c, hdr, hdrlen);\n>\n>         /* Then the data itself.. */\n> -       buf = in_stream->read(in_stream->data, &len);\n> -       stream.next_in = (void *)buf;\n> -       stream.avail_in = len;\n>         do {\n>                 unsigned char *in0 = stream.next_in;\n> -               ret = git_deflate(&stream, Z_FINISH);\n> +               if (!stream.avail_in) {\n> +                       if ((buf = in_stream->read(in_stream->data, &stream.avail_in))) {\n\nif ((buf = in_stream->read(in_stream->data, &stream.avail_in)) != NULL) {\n\nOr split this long line into:\n\n    buf = in_stream->read(in_stream->data, &stream.avail_in);\n    if (buf) {\n\n> +                               stream.next_in = (void *)buf;\n> +                               in0 = (unsigned char *)buf;\n> +                       } else\n> +                               flush = Z_FINISH;\n\nAdd {} around this single line, see:\n\n  https://github.com/git/git/blob/master/Documentation/CodingGuidelines#L279-L289\n\n> +               }\n> +               ret = git_deflate(&stream, flush);\n>                 the_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n>                 if (!dry_run && write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n>                         die(_(\"unable to write loose object file\"));\n> --\n> 2.33.1.44.g9344627884.agit.6.5.4\n>\n"},{"id":"441573","messageId":"xmqq7dd6ynb9.fsf@gitster.g","threadId":"56672","inReplyTo":"CANYiYbHiTnhpzyWLakVZ6tfmG0pKO=qHdZsgaocX6eJ=PN_06g@mail.gmail.com","subject":"Re: [PATCH v2 1/6] object-file: refactor write_loose_object() to support inputstream","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-11-18T06:45:14Z","receivedAt":"2021-11-18T06:45:25Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jiang Xin <worldhello.net@gmail.com> writes:\n\n> On Fri, Nov 12, 2021 at 5:43 PM Han Xin <chiyutianyi@gmail.com> wrote:\n>>\n>> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> It would be better to provide a cover letter describing changes in v2, such as:\n>\n> * Make \"write_loose_object()\" a public method, so we can\n>    reuse it in \"unpack_non_delta_entry()\".\n>    (But I doubt we can use \"write_object_file_flags()\" public\n>      function, without make this change.)\n>\n> * Add an new interface \"input_stream\" as an argument for\n>    \"write_loose_object()\", so that we can feed data to\n>    \"write_loose_object()\" from buffer or from zlib stream.\n>\n>> Refactor write_loose_object() to support inputstream, in the same way\n>> that zlib reading is chunked.\n>\n> In the beginning of your commit log, you should describe the problem, such as:\n>\n> We used to read the full content of a blob into buffer in\n> \"unpack_non_delta_entry()\" by calling:\n>\n>     void *buf = get_data(size);\n>\n> This will consume lots of memory for a very big blob object.\n\nI was not sure where \"in_stream\" came from---\"use X insteads of Y\",\nwhen X is what these patches invent and introduce, does not make a\ngood explanation without explaining what X is, what problem X is\nattempting to solve and how.\n\nThanks for helping to clarify the proposed log message.  \n"},{"id":"441577","messageId":"CANYiYbHMoH=pEhpx36ev-KWa7AXQtXpSiyjYObP1=XEx=Y8UNQ@mail.gmail.com","threadId":"56672","inReplyTo":"20211112094010.73468-6-chiyutianyi@gmail.com","subject":"Re: [PATCH v2 6/6] unpack-objects: unpack large object in stream","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2021-11-18T07:14:50Z","receivedAt":"2021-11-18T07:15:07Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"On Fri, Nov 12, 2021 at 5:42 PM Han Xin <chiyutianyi@gmail.com> wrote:\n>\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> When calling \"unpack_non_delta_entry()\", will allocate full memory for\n> the whole size of the unpacked object and write the buffer to loose file\n> on disk. This may lead to OOM for the git-unpack-objects process when\n> unpacking a very large object.\n>\n> In function \"unpack_delta_entry()\", will also allocate full memory to\n> buffer the whole delta, but since there will be no delta for an object\n> larger than \"core.bigFileThreshold\", this issue is moderate.\n>\n> To resolve the OOM issue in \"git-unpack-objects\", we can unpack large\n> object to file in stream, and use \"core.bigFileThreshold\" to avoid OOM\n> limits when called \"get_data()\".\n>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  builtin/unpack-objects.c          | 76 ++++++++++++++++++++++++-\n>  t/t5590-receive-unpack-objects.sh | 92 +++++++++++++++++++++++++++++++\n>  2 files changed, 167 insertions(+), 1 deletion(-)\n>  create mode 100755 t/t5590-receive-unpack-objects.sh\n>\n> diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n> index 4a9466295b..6c757d823b 100644\n> --- a/builtin/unpack-objects.c\n> +++ b/builtin/unpack-objects.c\n> @@ -320,11 +320,85 @@ static void added_object(unsigned nr, enum object_type type,\n>         }\n>  }\n>\n> +struct input_data_from_zstream {\n> +       git_zstream *zstream;\n> +       unsigned char buf[4096];\n> +       int status;\n> +};\n> +\n> +static const char *read_inflate_in_stream(void *data, unsigned long *readlen)\n> +{\n> +       struct input_data_from_zstream *input = data;\n> +       git_zstream *zstream = input->zstream;\n> +       void *in = fill(1);\n> +\n> +       if (!len || input->status == Z_STREAM_END) {\n> +               *readlen = 0;\n> +               return NULL;\n> +       }\n> +\n> +       zstream->next_out = input->buf;\n> +       zstream->avail_out = sizeof(input->buf);\n> +       zstream->next_in = in;\n> +       zstream->avail_in = len;\n> +\n> +       input->status = git_inflate(zstream, 0);\n> +       use(len - zstream->avail_in);\n> +       *readlen = sizeof(input->buf) - zstream->avail_out;\n> +\n> +       return (const char *)input->buf;\n> +}\n> +\n> +static void write_stream_blob(unsigned nr, unsigned long size)\n> +{\n> +       char hdr[32];\n> +       int hdrlen;\n> +       git_zstream zstream;\n> +       struct input_data_from_zstream data;\n> +       struct input_stream in_stream = {\n> +               .read = read_inflate_in_stream,\n> +               .data = &data,\n> +       };\n> +       struct object_id *oid = &obj_list[nr].oid;\n> +       int ret;\n> +\n> +       memset(&zstream, 0, sizeof(zstream));\n> +       memset(&data, 0, sizeof(data));\n> +       data.zstream = &zstream;\n> +       git_inflate_init(&zstream);\n> +\n> +       /* Generate the header */\n> +       hdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX, type_name(OBJ_BLOB), (uintmax_t)size) + 1;\n> +\n> +       if ((ret = write_loose_object(oid, hdr, hdrlen, &in_stream, dry_run, 0, 0)))\n> +               die(_(\"failed to write object in stream %d\"), ret);\n> +\n> +       if (zstream.total_out != size || data.status != Z_STREAM_END)\n> +               die(_(\"inflate returned %d\"), data.status);\n> +       git_inflate_end(&zstream);\n> +\n> +       if (strict && !dry_run) {\n> +               struct blob *blob = lookup_blob(the_repository, oid);\n> +               if (blob)\n> +                       blob->object.flags |= FLAG_WRITTEN;\n> +               else\n> +                       die(\"invalid blob object from stream\");\n> +       }\n> +       obj_list[nr].obj = NULL;\n> +}\n> +\n>  static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n>                                    unsigned nr)\n>  {\n> -       void *buf = get_data(size);\n> +       void *buf;\n> +\n> +       /* Write large blob in stream without allocating full buffer. */\n> +       if (type == OBJ_BLOB && size > big_file_threshold) {\n\nDefault size of big_file_threshold is 512m.  Can we use\n\"write_stream_blob\" for all objects?  Can we get a more suitable\nthreshold through some benchmark data?\n\n> +               write_stream_blob(nr, size);\n> +               return;\n> +       }\n\n--\nJiang Xin\n"},{"id":"441896","messageId":"20211122033220.32883-1-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211009082058.41138-1-chiyutianyi@gmail.com","subject":"[PATCH v3 0/5] unpack large objects in stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-11-22T03:32:15Z","receivedAt":"2021-11-22T03:35:20Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nAlthough we do not recommend users push large binary files to the git repositories, \nit's difficult to prevent them from doing so. Once, we found a problem with a surge \nin memory usage on the server. The source of the problem is that a user submitted \na single object with a size of 15GB. Once someone initiates a git push, the git \nprocess will immediately allocate 15G of memory, resulting in an OOM risk.\n\nThrough further analysis, we found that when we execute git unpack-objects, in \nunpack_non_delta_entry(), \"void *buf = get_data(size);\" will directly allocate \nmemory equal to the size of the object. This is quite a scary thing, because the \npre-receive hook has not been executed at this time, and we cannot avoid this by hooks.\n\nI got inspiration from the deflate process of zlib, maybe it would be a good idea \nto change unpack-objects to stream deflate.\n\nChanges since v2:\n* Rewrite commit messages and make changes suggested by Jiang Xin.\n* Remove the commit \"object-file.c: add dry_run mode for write_loose_object()\" and\n  use a new commit \"unpack-objects.c: add dry_run mode for get_data()\" instead.\n\nHan Xin (5):\n  object-file: refactor write_loose_object() to read buffer from stream\n  object-file.c: handle undetermined oid in write_loose_object()\n  object-file.c: read stream in a loop in write_loose_object()\n  unpack-objects.c: add dry_run mode for get_data()\n  unpack-objects: unpack_non_delta_entry() read data in a stream\n\n builtin/unpack-objects.c            | 92 +++++++++++++++++++++++++--\n object-file.c                       | 98 +++++++++++++++++++++++++----\n object-store.h                      |  9 +++\n t/t5590-unpack-non-delta-objects.sh | 76 ++++++++++++++++++++++\n 4 files changed, 257 insertions(+), 18 deletions(-)\n create mode 100755 t/t5590-unpack-non-delta-objects.sh\n\nRange-diff against v2:\n1:  01672f50a0 ! 1:  8640b04f6d object-file: refactor write_loose_object() to support inputstream\n    @@ Metadata\n     Author: Han Xin <hanxin.hx@alibaba-inc.com>\n     \n      ## Commit message ##\n    -    object-file: refactor write_loose_object() to support inputstream\n    +    object-file: refactor write_loose_object() to read buffer from stream\n     \n    -    Refactor write_loose_object() to support inputstream, in the same way\n    -    that zlib reading is chunked.\n    +    We used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\n    +    entire contents of a blob object, no matter how big it is. This\n    +    implementation may consume all the memory and cause OOM.\n     \n    -    Using \"in_stream\" instead of \"void *buf\", we needn't to allocate enough\n    -    memory in advance, and only part of the contents will be read when\n    -    called \"in_stream.read()\".\n    +    This can be improved by feeding data to \"write_loose_object()\" in a\n    +    stream. The input stream is implemented as an interface. In the first\n    +    step, we make a simple implementation, feeding the entire buffer in the\n    +    \"stream\" to \"write_loose_object()\" as a refactor.\n     \n         Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n         Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n    @@ object-file.c: static int create_tmpfile(struct strbuf *tmp, const char *filenam\n      \treturn fd;\n      }\n      \n    -+struct input_data_from_buffer {\n    -+\tconst char *buf;\n    ++struct simple_input_stream_data {\n    ++\tconst void *buf;\n     +\tunsigned long len;\n     +};\n     +\n    -+static const char *read_input_stream_from_buffer(void *data, unsigned long *len)\n    ++static const void *feed_simple_input_stream(struct input_stream *in_stream, unsigned long *len)\n     +{\n    -+\tstruct input_data_from_buffer *input = (struct input_data_from_buffer *)data;\n    ++\tstruct simple_input_stream_data *data = in_stream->data;\n     +\n    -+\tif (input->len == 0) {\n    ++\tif (data->len == 0) {\n     +\t\t*len = 0;\n     +\t\treturn NULL;\n     +\t}\n    -+\t*len = input->len;\n    -+\tinput->len = 0;\n    -+\treturn input->buf;\n    ++\t*len = data->len;\n    ++\tdata->len = 0;\n    ++\treturn data->buf;\n     +}\n     +\n      static int write_loose_object(const struct object_id *oid, char *hdr,\n    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *\n      \tstruct object_id parano_oid;\n      \tstatic struct strbuf tmp_file = STRBUF_INIT;\n      \tstatic struct strbuf filename = STRBUF_INIT;\n    -+\tconst char *buf;\n    ++\tconst void *buf;\n     +\tunsigned long len;\n      \n      \tloose_object_path(the_repository, &filename, oid);\n    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *\n      \tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n      \n      \t/* Then the data itself.. */\n    -+\tbuf = in_stream->read(in_stream->data, &len);\n    ++\tbuf = in_stream->read(in_stream, &len);\n      \tstream.next_in = (void *)buf;\n      \tstream.avail_in = len;\n      \tdo {\n    @@ object-file.c: int write_object_file_flags(const void *buf, unsigned long len,\n      \tchar hdr[MAX_HEADER_LEN];\n      \tint hdrlen = sizeof(hdr);\n     +\tstruct input_stream in_stream = {\n    -+\t\t.read = read_input_stream_from_buffer,\n    -+\t\t.data = (void *)&(struct input_data_from_buffer) {\n    ++\t\t.read = feed_simple_input_stream,\n    ++\t\t.data = (void *)&(struct simple_input_stream_data) {\n     +\t\t\t.buf = buf,\n     +\t\t\t.len = len,\n     +\t\t},\n    @@ object-file.c: int hash_object_file_literally(const void *buf, unsigned long len\n      \tchar *header;\n      \tint hdrlen, status = 0;\n     +\tstruct input_stream in_stream = {\n    -+\t\t.read = read_input_stream_from_buffer,\n    -+\t\t.data = (void *)&(struct input_data_from_buffer) {\n    ++\t\t.read = feed_simple_input_stream,\n    ++\t\t.data = (void *)&(struct simple_input_stream_data) {\n     +\t\t\t.buf = buf,\n     +\t\t\t.len = len,\n     +\t\t},\n    @@ object-file.c: int force_object_loose(const struct object_id *oid, time_t mtime)\n      \tchar hdr[MAX_HEADER_LEN];\n      \tint hdrlen;\n      \tint ret;\n    -+\tstruct input_data_from_buffer data;\n    ++\tstruct simple_input_stream_data data;\n     +\tstruct input_stream in_stream = {\n    -+\t\t.read = read_input_stream_from_buffer,\n    ++\t\t.read = feed_simple_input_stream,\n     +\t\t.data = &data,\n     +\t};\n      \n    @@ object-store.h: struct object_directory {\n      };\n      \n     +struct input_stream {\n    -+\tconst char *(*read)(void* data, unsigned long *len);\n    ++\tconst void *(*read)(struct input_stream *, unsigned long *len);\n     +\tvoid *data;\n     +};\n     +\n2:  a309b7e391 < -:  ---------- object-file.c: add dry_run mode for write_loose_object()\n3:  b0a5b53710 ! 2:  d4a2caf2bd object-file.c: handle nil oid in write_loose_object()\n    @@ Metadata\n     Author: Han Xin <hanxin.hx@alibaba-inc.com>\n     \n      ## Commit message ##\n    -    object-file.c: handle nil oid in write_loose_object()\n    +    object-file.c: handle undetermined oid in write_loose_object()\n     \n    -    When read input stream, oid can't get before reading all, and it will be\n    -    filled after reading.\n    +    When streaming a large blob object to \"write_loose_object()\", we have no\n    +    chance to run \"write_object_file_prepare()\" to calculate the oid in\n    +    advance. So we need to handle undetermined oid in function\n    +    \"write_loose_object()\".\n    +\n    +    In the original implementation, we know the oid and we can write the\n    +    temporary file in the same directory as the final object, but for an\n    +    object with an undetermined oid, we don't know the exact directory for\n    +    the object, so we have to save the temporary file in \".git/objects/\"\n    +    directory instead.\n     \n         Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n         Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n     \n      ## object-file.c ##\n     @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *hdr,\n    - \tconst char *buf;\n    + \tconst void *buf;\n      \tunsigned long len;\n      \n     -\tloose_object_path(the_repository, &filename, oid);\n    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *\n     +\t\tstrbuf_reset(&filename);\n     +\t\tstrbuf_addstr(&filename, the_repository->objects->odb->path);\n     +\t\tstrbuf_addch(&filename, '/');\n    -+\t} else\n    ++\t} else {\n     +\t\tloose_object_path(the_repository, &filename, oid);\n    ++\t}\n      \n    - \tif (!dry_run) {\n    - \t\tfd = create_tmpfile(&tmp_file, filename.buf);\n    + \tfd = create_tmpfile(&tmp_file, filename.buf);\n    + \tif (fd < 0) {\n     @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *hdr,\n      \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n      \t\t    ret);\n    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *\n      \t\tdie(_(\"confused by unstable object source data for %s\"),\n      \t\t    oid_to_hex(oid));\n      \n    -@@ object-file.c: static int write_loose_object(const struct object_id *oid, char *hdr,\n    - \n      \tclose_loose_object(fd);\n      \n     +\tif (is_null_oid(oid)) {\n     +\t\tint dirlen;\n     +\n    -+\t\t/* copy oid */\n     +\t\toidcpy((struct object_id *)oid, &parano_oid);\n    -+\t\t/* We get the oid now */\n     +\t\tloose_object_path(the_repository, &filename, oid);\n     +\n    ++\t\t/* We finally know the object path, and create the missing dir. */\n     +\t\tdirlen = directory_size(filename.buf);\n     +\t\tif (dirlen) {\n     +\t\t\tstruct strbuf dir = STRBUF_INIT;\n    -+\t\t\t/*\n    -+\t\t\t * Make sure the directory exists; note that the\n    -+\t\t\t * contents of the buffer are undefined after mkstemp\n    -+\t\t\t * returns an error, so we have to rewrite the whole\n    -+\t\t\t * buffer from scratch.\n    -+\t\t\t */\n    -+\t\t\tstrbuf_reset(&dir);\n     +\t\t\tstrbuf_add(&dir, filename.buf, dirlen - 1);\n     +\t\t\tif (mkdir(dir.buf, 0777) && errno != EEXIST)\n     +\t\t\t\treturn -1;\n    ++\t\t\tif (adjust_shared_perm(dir.buf))\n    ++\t\t\t\treturn -1;\n    ++\t\t\tstrbuf_release(&dir);\n     +\t\t}\n     +\t}\n     +\n4:  09d438b692 ! 3:  2575900449 object-file.c: read input stream repeatedly in write_loose_object()\n    @@ Metadata\n     Author: Han Xin <hanxin.hx@alibaba-inc.com>\n     \n      ## Commit message ##\n    -    object-file.c: read input stream repeatedly in write_loose_object()\n    +    object-file.c: read stream in a loop in write_loose_object()\n     \n    -    Read input stream repeatedly in write_loose_object() unless reach the\n    -    end, so that we can divide the large blob write into many small blocks.\n    +    In order to prepare the stream version of \"write_loose_object()\", read\n    +    the input stream in a loop in \"write_loose_object()\", so that we can\n    +    feed the contents of large blob object to \"write_loose_object()\" using\n    +    a small fixed buffer.\n     \n    +    Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n         Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n     \n      ## object-file.c ##\n     @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *hdr,\n      \tstatic struct strbuf tmp_file = STRBUF_INIT;\n      \tstatic struct strbuf filename = STRBUF_INIT;\n    - \tconst char *buf;\n    + \tconst void *buf;\n     -\tunsigned long len;\n     +\tint flush = 0;\n      \n    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *\n      \tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n      \n      \t/* Then the data itself.. */\n    --\tbuf = in_stream->read(in_stream->data, &len);\n    +-\tbuf = in_stream->read(in_stream, &len);\n     -\tstream.next_in = (void *)buf;\n     -\tstream.avail_in = len;\n      \tdo {\n      \t\tunsigned char *in0 = stream.next_in;\n     -\t\tret = git_deflate(&stream, Z_FINISH);\n     +\t\tif (!stream.avail_in) {\n    -+\t\t\tif ((buf = in_stream->read(in_stream->data, &stream.avail_in))) {\n    ++\t\t\tbuf = in_stream->read(in_stream, &stream.avail_in);\n    ++\t\t\tif (buf) {\n     +\t\t\t\tstream.next_in = (void *)buf;\n     +\t\t\t\tin0 = (unsigned char *)buf;\n    -+\t\t\t} else\n    ++\t\t\t} else {\n     +\t\t\t\tflush = Z_FINISH;\n    ++\t\t\t}\n     +\t\t}\n     +\t\tret = git_deflate(&stream, flush);\n      \t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n    - \t\tif (!dry_run && write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n    + \t\tif (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n      \t\t\tdie(_(\"unable to write loose object file\"));\n5:  9fb188d437 < -:  ---------- object-store.h: add write_loose_object()\n-:  ---------- > 4:  ca93ecc780 unpack-objects.c: add dry_run mode for get_data()\n6:  80468a6fbc ! 5:  39a072ee2a unpack-objects: unpack large object in stream\n    @@ Metadata\n     Author: Han Xin <hanxin.hx@alibaba-inc.com>\n     \n      ## Commit message ##\n    -    unpack-objects: unpack large object in stream\n    +    unpack-objects: unpack_non_delta_entry() read data in a stream\n     \n    -    When calling \"unpack_non_delta_entry()\", will allocate full memory for\n    -    the whole size of the unpacked object and write the buffer to loose file\n    -    on disk. This may lead to OOM for the git-unpack-objects process when\n    -    unpacking a very large object.\n    +    We used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\n    +    entire contents of a blob object, no matter how big it is. This\n    +    implementation may consume all the memory and cause OOM.\n     \n    -    In function \"unpack_delta_entry()\", will also allocate full memory to\n    -    buffer the whole delta, but since there will be no delta for an object\n    -    larger than \"core.bigFileThreshold\", this issue is moderate.\n    +    By implementing a zstream version of input_stream interface, we can use\n    +    a small fixed buffer for \"unpack_non_delta_entry()\".\n     \n    -    To resolve the OOM issue in \"git-unpack-objects\", we can unpack large\n    -    object to file in stream, and use \"core.bigFileThreshold\" to avoid OOM\n    -    limits when called \"get_data()\".\n    +    However, unpack non-delta objects from a stream instead of from an entrie\n    +    buffer will have 10% performance penalty. Therefore, only unpack object\n    +    larger than the \"big_file_threshold\" in zstream. See the following\n    +    benchmarks:\n     \n    +        $ hyperfine \\\n    +        --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n    +        'git -C dest.git unpack-objects <binary_320M.pack'\n    +        Benchmark 1: git -C dest.git unpack-objects <binary_320M.pack\n    +          Time (mean ± σ):     10.029 s ±  0.270 s    [User: 8.265 s, System: 1.522 s]\n    +          Range (min … max):    9.786 s … 10.603 s    10 runs\n    +\n    +        $ hyperfine \\\n    +        --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n    +        'git -c core.bigFileThreshold=2m -C dest.git unpack-objects <binary_320M.pack'\n    +        Benchmark 1: git -c core.bigFileThreshold=2m -C dest.git unpack-objects <binary_320M.pack\n    +          Time (mean ± σ):     10.859 s ±  0.774 s    [User: 8.813 s, System: 1.898 s]\n    +          Range (min … max):    9.884 s … 12.192 s    10 runs\n    +\n    +        $ hyperfine \\\n    +        --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n    +        'git -C dest.git unpack-objects <binary_96M.pack'\n    +        Benchmark 1: git -C dest.git unpack-objects <binary_96M.pack\n    +          Time (mean ± σ):      2.678 s ±  0.037 s    [User: 2.205 s, System: 0.450 s]\n    +          Range (min … max):    2.639 s …  2.743 s    10 runs\n    +\n    +        $ hyperfine \\\n    +        --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n    +        'git -c core.bigFileThreshold=2m -C dest.git unpack-objects <binary_96M.pack'\n    +        Benchmark 1: git -c core.bigFileThreshold=2m -C dest.git unpack-objects <binary_96M.pack\n    +          Time (mean ± σ):      2.819 s ±  0.124 s    [User: 2.216 s, System: 0.564 s]\n    +          Range (min … max):    2.679 s …  3.125 s    10 runs\n    +\n    +    Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n         Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n     \n      ## builtin/unpack-objects.c ##\n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n      \t}\n      }\n      \n    -+struct input_data_from_zstream {\n    ++struct input_zstream_data {\n     +\tgit_zstream *zstream;\n     +\tunsigned char buf[4096];\n     +\tint status;\n     +};\n     +\n    -+static const char *read_inflate_in_stream(void *data, unsigned long *readlen)\n    ++static const void *feed_input_zstream(struct input_stream *in_stream, unsigned long *readlen)\n     +{\n    -+\tstruct input_data_from_zstream *input = data;\n    -+\tgit_zstream *zstream = input->zstream;\n    ++\tstruct input_zstream_data *data = in_stream->data;\n    ++\tgit_zstream *zstream = data->zstream;\n     +\tvoid *in = fill(1);\n     +\n    -+\tif (!len || input->status == Z_STREAM_END) {\n    ++\tif (!len || data->status == Z_STREAM_END) {\n     +\t\t*readlen = 0;\n     +\t\treturn NULL;\n     +\t}\n     +\n    -+\tzstream->next_out = input->buf;\n    -+\tzstream->avail_out = sizeof(input->buf);\n    ++\tzstream->next_out = data->buf;\n    ++\tzstream->avail_out = sizeof(data->buf);\n     +\tzstream->next_in = in;\n     +\tzstream->avail_in = len;\n     +\n    -+\tinput->status = git_inflate(zstream, 0);\n    ++\tdata->status = git_inflate(zstream, 0);\n     +\tuse(len - zstream->avail_in);\n    -+\t*readlen = sizeof(input->buf) - zstream->avail_out;\n    ++\t*readlen = sizeof(data->buf) - zstream->avail_out;\n     +\n    -+\treturn (const char *)input->buf;\n    ++\treturn data->buf;\n     +}\n     +\n     +static void write_stream_blob(unsigned nr, unsigned long size)\n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n     +\tchar hdr[32];\n     +\tint hdrlen;\n     +\tgit_zstream zstream;\n    -+\tstruct input_data_from_zstream data;\n    ++\tstruct input_zstream_data data;\n     +\tstruct input_stream in_stream = {\n    -+\t\t.read = read_inflate_in_stream,\n    ++\t\t.read = feed_input_zstream,\n     +\t\t.data = &data,\n     +\t};\n     +\tstruct object_id *oid = &obj_list[nr].oid;\n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n     +\t/* Generate the header */\n     +\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX, type_name(OBJ_BLOB), (uintmax_t)size) + 1;\n     +\n    -+\tif ((ret = write_loose_object(oid, hdr, hdrlen, &in_stream, dry_run, 0, 0)))\n    ++\tif ((ret = write_loose_object(oid, hdr, hdrlen, &in_stream, 0, 0)))\n     +\t\tdie(_(\"failed to write object in stream %d\"), ret);\n     +\n     +\tif (zstream.total_out != size || data.status != Z_STREAM_END)\n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n      static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n      \t\t\t\t   unsigned nr)\n      {\n    --\tvoid *buf = get_data(size);\n    +-\tvoid *buf = get_data(size, dry_run);\n     +\tvoid *buf;\n     +\n     +\t/* Write large blob in stream without allocating full buffer. */\n    -+\tif (type == OBJ_BLOB && size > big_file_threshold) {\n    ++\tif (!dry_run && type == OBJ_BLOB && size > big_file_threshold) {\n     +\t\twrite_stream_blob(nr, size);\n     +\t\treturn;\n     +\t}\n      \n    -+\tbuf = get_data(size);\n    ++\tbuf = get_data(size, dry_run);\n      \tif (!dry_run && buf)\n      \t\twrite_object(nr, type, buf, size);\n      \telse\n     \n    - ## t/t5590-receive-unpack-objects.sh (new) ##\n    + ## object-file.c ##\n    +@@ object-file.c: static const void *feed_simple_input_stream(struct input_stream *in_stream, unsi\n    + \treturn data->buf;\n    + }\n    + \n    +-static int write_loose_object(const struct object_id *oid, char *hdr,\n    +-\t\t\t      int hdrlen, struct input_stream *in_stream,\n    +-\t\t\t      time_t mtime, unsigned flags)\n    ++int write_loose_object(const struct object_id *oid, char *hdr,\n    ++\t\t       int hdrlen, struct input_stream *in_stream,\n    ++\t\t       time_t mtime, unsigned flags)\n    + {\n    + \tint fd, ret;\n    + \tunsigned char compressed[4096];\n    +\n    + ## object-store.h ##\n    +@@ object-store.h: int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n    + \t\t     unsigned long len, const char *type,\n    + \t\t     struct object_id *oid);\n    + \n    ++int write_loose_object(const struct object_id *oid, char *hdr,\n    ++\t\t       int hdrlen, struct input_stream *in_stream,\n    ++\t\t       time_t mtime, unsigned flags);\n    ++\n    + int write_object_file_flags(const void *buf, unsigned long len,\n    + \t\t\t    const char *type, struct object_id *oid,\n    + \t\t\t    unsigned flags);\n    +\n    + ## t/t5590-unpack-non-delta-objects.sh (new) ##\n     @@\n     +#!/bin/sh\n     +#\n    @@ t/t5590-receive-unpack-objects.sh (new)\n     +\t\tcd .git &&\n     +\t\tfind objects/?? -type f | sort\n     +\t) >expect &&\n    -+\tgit repack -ad\n    ++\tPACK=$(echo main | git pack-objects --progress --revs test)\n     +'\n     +\n     +test_expect_success 'setup GIT_ALLOC_LIMIT to 1MB' '\n    @@ t/t5590-receive-unpack-objects.sh (new)\n     +\tgit -C dest.git config receive.unpacklimit 100\n     +'\n     +\n    -+test_expect_success 'fail to push: cannot allocate' '\n    -+\ttest_must_fail git push dest.git HEAD 2>err &&\n    -+\ttest_i18ngrep \"remote: fatal: attempting to allocate\" err &&\n    ++test_expect_success 'fail to unpack-objects: cannot allocate' '\n    ++\ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n    ++\ttest_i18ngrep \"fatal: attempting to allocate\" err &&\n     +\t(\n     +\t\tcd dest.git &&\n     +\t\tfind objects/?? -type f | sort\n    @@ t/t5590-receive-unpack-objects.sh (new)\n     +'\n     +\n     +test_expect_success 'unpack big object in stream' '\n    -+\tgit push dest.git HEAD &&\n    ++\tgit -C dest.git unpack-objects <test-$PACK.pack &&\n     +\tgit -C dest.git fsck &&\n     +\t(\n     +\t\tcd dest.git &&\n    @@ t/t5590-receive-unpack-objects.sh (new)\n     +'\n     +\n     +test_expect_success 'setup for unpack-objects dry-run test' '\n    -+\tPACK=$(echo main | git pack-objects --progress --revs test) &&\n    -+\tunset GIT_ALLOC_LIMIT &&\n     +\tgit init --bare unpack-test.git\n     +'\n     +\n    -+test_expect_success 'unpack-objects dry-run with large threshold' '\n    -+\t(\n    -+\t\tcd unpack-test.git &&\n    -+\t\tgit config core.bigFileThreshold 2m &&\n    -+\t\tgit unpack-objects -n <../test-$PACK.pack\n    -+\t) &&\n    -+\t(\n    -+\t\tcd unpack-test.git &&\n    -+\t\tfind objects/ -type f\n    -+\t) >actual &&\n    -+\ttest_must_be_empty actual\n    -+'\n    -+\n    -+test_expect_success 'unpack-objects dry-run with small threshold' '\n    ++test_expect_success 'unpack-objects dry-run' '\n     +\t(\n     +\t\tcd unpack-test.git &&\n    -+\t\tgit config core.bigFileThreshold 1m &&\n     +\t\tgit unpack-objects -n <../test-$PACK.pack\n     +\t) &&\n     +\t(\n-- \n2.34.0.6.g676eedc724\n\n"},{"id":"441897","messageId":"20211122033220.32883-3-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211009082058.41138-1-chiyutianyi@gmail.com","subject":"[PATCH v3 2/5] object-file.c: handle undetermined oid in write_loose_object()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-11-22T03:32:17Z","receivedAt":"2021-11-22T03:35:22Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWhen streaming a large blob object to \"write_loose_object()\", we have no\nchance to run \"write_object_file_prepare()\" to calculate the oid in\nadvance. So we need to handle undetermined oid in function\n\"write_loose_object()\".\n\nIn the original implementation, we know the oid and we can write the\ntemporary file in the same directory as the final object, but for an\nobject with an undetermined oid, we don't know the exact directory for\nthe object, so we have to save the temporary file in \".git/objects/\"\ndirectory instead.\n\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c | 30 ++++++++++++++++++++++++++++--\n 1 file changed, 28 insertions(+), 2 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 227f53a0de..78fd2a5d39 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1892,7 +1892,14 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tconst void *buf;\n \tunsigned long len;\n \n-\tloose_object_path(the_repository, &filename, oid);\n+\tif (is_null_oid(oid)) {\n+\t\t/* When oid is not determined, save tmp file to odb path. */\n+\t\tstrbuf_reset(&filename);\n+\t\tstrbuf_addstr(&filename, the_repository->objects->odb->path);\n+\t\tstrbuf_addch(&filename, '/');\n+\t} else {\n+\t\tloose_object_path(the_repository, &filename, oid);\n+\t}\n \n \tfd = create_tmpfile(&tmp_file, filename.buf);\n \tif (fd < 0) {\n@@ -1939,12 +1946,31 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n \t\t    ret);\n \tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n-\tif (!oideq(oid, &parano_oid))\n+\tif (!is_null_oid(oid) && !oideq(oid, &parano_oid))\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n \n \tclose_loose_object(fd);\n \n+\tif (is_null_oid(oid)) {\n+\t\tint dirlen;\n+\n+\t\toidcpy((struct object_id *)oid, &parano_oid);\n+\t\tloose_object_path(the_repository, &filename, oid);\n+\n+\t\t/* We finally know the object path, and create the missing dir. */\n+\t\tdirlen = directory_size(filename.buf);\n+\t\tif (dirlen) {\n+\t\t\tstruct strbuf dir = STRBUF_INIT;\n+\t\t\tstrbuf_add(&dir, filename.buf, dirlen - 1);\n+\t\t\tif (mkdir(dir.buf, 0777) && errno != EEXIST)\n+\t\t\t\treturn -1;\n+\t\t\tif (adjust_shared_perm(dir.buf))\n+\t\t\t\treturn -1;\n+\t\t\tstrbuf_release(&dir);\n+\t\t}\n+\t}\n+\n \tif (mtime) {\n \t\tstruct utimbuf utb;\n \t\tutb.actime = mtime;\n-- \n2.34.0.6.g676eedc724\n\n"},{"id":"441898","messageId":"20211122033220.32883-2-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211009082058.41138-1-chiyutianyi@gmail.com","subject":"[PATCH v3 1/5] object-file: refactor write_loose_object() to read buffer from stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-11-22T03:32:16Z","receivedAt":"2021-11-22T03:35:23Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWe used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\nentire contents of a blob object, no matter how big it is. This\nimplementation may consume all the memory and cause OOM.\n\nThis can be improved by feeding data to \"write_loose_object()\" in a\nstream. The input stream is implemented as an interface. In the first\nstep, we make a simple implementation, feeding the entire buffer in the\n\"stream\" to \"write_loose_object()\" as a refactor.\n\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c  | 50 ++++++++++++++++++++++++++++++++++++++++++++++----\n object-store.h |  5 +++++\n 2 files changed, 51 insertions(+), 4 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex c3d866a287..227f53a0de 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1860,8 +1860,26 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n \treturn fd;\n }\n \n+struct simple_input_stream_data {\n+\tconst void *buf;\n+\tunsigned long len;\n+};\n+\n+static const void *feed_simple_input_stream(struct input_stream *in_stream, unsigned long *len)\n+{\n+\tstruct simple_input_stream_data *data = in_stream->data;\n+\n+\tif (data->len == 0) {\n+\t\t*len = 0;\n+\t\treturn NULL;\n+\t}\n+\t*len = data->len;\n+\tdata->len = 0;\n+\treturn data->buf;\n+}\n+\n static int write_loose_object(const struct object_id *oid, char *hdr,\n-\t\t\t      int hdrlen, const void *buf, unsigned long len,\n+\t\t\t      int hdrlen, struct input_stream *in_stream,\n \t\t\t      time_t mtime, unsigned flags)\n {\n \tint fd, ret;\n@@ -1871,6 +1889,8 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tstruct object_id parano_oid;\n \tstatic struct strbuf tmp_file = STRBUF_INIT;\n \tstatic struct strbuf filename = STRBUF_INIT;\n+\tconst void *buf;\n+\tunsigned long len;\n \n \tloose_object_path(the_repository, &filename, oid);\n \n@@ -1898,6 +1918,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n \n \t/* Then the data itself.. */\n+\tbuf = in_stream->read(in_stream, &len);\n \tstream.next_in = (void *)buf;\n \tstream.avail_in = len;\n \tdo {\n@@ -1960,6 +1981,13 @@ int write_object_file_flags(const void *buf, unsigned long len,\n {\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen = sizeof(hdr);\n+\tstruct input_stream in_stream = {\n+\t\t.read = feed_simple_input_stream,\n+\t\t.data = (void *)&(struct simple_input_stream_data) {\n+\t\t\t.buf = buf,\n+\t\t\t.len = len,\n+\t\t},\n+\t};\n \n \t/* Normally if we have it in the pack then we do not bother writing\n \t * it out into .git/objects/??/?{38} file.\n@@ -1968,7 +1996,7 @@ int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t\t  &hdrlen);\n \tif (freshen_packed_object(oid) || freshen_loose_object(oid))\n \t\treturn 0;\n-\treturn write_loose_object(oid, hdr, hdrlen, buf, len, 0, flags);\n+\treturn write_loose_object(oid, hdr, hdrlen, &in_stream, 0, flags);\n }\n \n int hash_object_file_literally(const void *buf, unsigned long len,\n@@ -1977,6 +2005,13 @@ int hash_object_file_literally(const void *buf, unsigned long len,\n {\n \tchar *header;\n \tint hdrlen, status = 0;\n+\tstruct input_stream in_stream = {\n+\t\t.read = feed_simple_input_stream,\n+\t\t.data = (void *)&(struct simple_input_stream_data) {\n+\t\t\t.buf = buf,\n+\t\t\t.len = len,\n+\t\t},\n+\t};\n \n \t/* type string, SP, %lu of the length plus NUL must fit this */\n \thdrlen = strlen(type) + MAX_HEADER_LEN;\n@@ -1988,7 +2023,7 @@ int hash_object_file_literally(const void *buf, unsigned long len,\n \t\tgoto cleanup;\n \tif (freshen_packed_object(oid) || freshen_loose_object(oid))\n \t\tgoto cleanup;\n-\tstatus = write_loose_object(oid, header, hdrlen, buf, len, 0, 0);\n+\tstatus = write_loose_object(oid, header, hdrlen, &in_stream, 0, 0);\n \n cleanup:\n \tfree(header);\n@@ -2003,14 +2038,21 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen;\n \tint ret;\n+\tstruct simple_input_stream_data data;\n+\tstruct input_stream in_stream = {\n+\t\t.read = feed_simple_input_stream,\n+\t\t.data = &data,\n+\t};\n \n \tif (has_loose_object(oid))\n \t\treturn 0;\n \tbuf = read_object(the_repository, oid, &type, &len);\n \tif (!buf)\n \t\treturn error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n+\tdata.buf = buf;\n+\tdata.len = len;\n \thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n-\tret = write_loose_object(oid, hdr, hdrlen, buf, len, mtime, 0);\n+\tret = write_loose_object(oid, hdr, hdrlen, &in_stream, mtime, 0);\n \tfree(buf);\n \n \treturn ret;\ndiff --git a/object-store.h b/object-store.h\nindex 952efb6a4b..ccc1fc9c1a 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -34,6 +34,11 @@ struct object_directory {\n \tchar *path;\n };\n \n+struct input_stream {\n+\tconst void *(*read)(struct input_stream *, unsigned long *len);\n+\tvoid *data;\n+};\n+\n KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n \tstruct object_directory *, 1, fspathhash, fspatheq)\n \n-- \n2.34.0.6.g676eedc724\n\n"},{"id":"441899","messageId":"20211122033220.32883-4-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211009082058.41138-1-chiyutianyi@gmail.com","subject":"[PATCH v3 3/5] object-file.c: read stream in a loop in write_loose_object()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-11-22T03:32:18Z","receivedAt":"2021-11-22T03:35:24Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nIn order to prepare the stream version of \"write_loose_object()\", read\nthe input stream in a loop in \"write_loose_object()\", so that we can\nfeed the contents of large blob object to \"write_loose_object()\" using\na small fixed buffer.\n\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c | 16 +++++++++++-----\n 1 file changed, 11 insertions(+), 5 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 78fd2a5d39..93bcfaca50 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1890,7 +1890,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tstatic struct strbuf tmp_file = STRBUF_INIT;\n \tstatic struct strbuf filename = STRBUF_INIT;\n \tconst void *buf;\n-\tunsigned long len;\n+\tint flush = 0;\n \n \tif (is_null_oid(oid)) {\n \t\t/* When oid is not determined, save tmp file to odb path. */\n@@ -1925,12 +1925,18 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n \n \t/* Then the data itself.. */\n-\tbuf = in_stream->read(in_stream, &len);\n-\tstream.next_in = (void *)buf;\n-\tstream.avail_in = len;\n \tdo {\n \t\tunsigned char *in0 = stream.next_in;\n-\t\tret = git_deflate(&stream, Z_FINISH);\n+\t\tif (!stream.avail_in) {\n+\t\t\tbuf = in_stream->read(in_stream, &stream.avail_in);\n+\t\t\tif (buf) {\n+\t\t\t\tstream.next_in = (void *)buf;\n+\t\t\t\tin0 = (unsigned char *)buf;\n+\t\t\t} else {\n+\t\t\t\tflush = Z_FINISH;\n+\t\t\t}\n+\t\t}\n+\t\tret = git_deflate(&stream, flush);\n \t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n \t\tif (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n \t\t\tdie(_(\"unable to write loose object file\"));\n-- \n2.34.0.6.g676eedc724\n\n"},{"id":"441900","messageId":"20211122033220.32883-5-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211009082058.41138-1-chiyutianyi@gmail.com","subject":"[PATCH v3 4/5] unpack-objects.c: add dry_run mode for get_data()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-11-22T03:32:19Z","receivedAt":"2021-11-22T03:35:26Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nIn dry_run mode, \"get_data()\" is used to verify the inflation of data,\nand the returned buffer will not be used at all and will be freed\nimmediately. Even in dry_run mode, it is dangerous to allocate a\nfull-size buffer for a large blob object. Therefore, only allocate a\nlow memory footprint when calling \"get_data()\" in dry_run mode.\n\nSuggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n builtin/unpack-objects.c | 18 ++++++++++++------\n 1 file changed, 12 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 4a9466295b..8d68acd662 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -96,15 +96,16 @@ static void use(int bytes)\n \tdisplay_throughput(progress, consumed_bytes);\n }\n \n-static void *get_data(unsigned long size)\n+static void *get_data(unsigned long size, int dry_run)\n {\n \tgit_zstream stream;\n-\tvoid *buf = xmallocz(size);\n+\tunsigned long bufsize = dry_run ? 4096 : size;\n+\tvoid *buf = xmallocz(bufsize);\n \n \tmemset(&stream, 0, sizeof(stream));\n \n \tstream.next_out = buf;\n-\tstream.avail_out = size;\n+\tstream.avail_out = bufsize;\n \tstream.next_in = fill(1);\n \tstream.avail_in = len;\n \tgit_inflate_init(&stream);\n@@ -124,6 +125,11 @@ static void *get_data(unsigned long size)\n \t\t}\n \t\tstream.next_in = fill(1);\n \t\tstream.avail_in = len;\n+\t\tif (dry_run) {\n+\t\t\t/* reuse the buffer in dry_run mode */\n+\t\t\tstream.next_out = buf;\n+\t\t\tstream.avail_out = bufsize;\n+\t\t}\n \t}\n \tgit_inflate_end(&stream);\n \treturn buf;\n@@ -323,7 +329,7 @@ static void added_object(unsigned nr, enum object_type type,\n static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n \t\t\t\t   unsigned nr)\n {\n-\tvoid *buf = get_data(size);\n+\tvoid *buf = get_data(size, dry_run);\n \n \tif (!dry_run && buf)\n \t\twrite_object(nr, type, buf, size);\n@@ -357,7 +363,7 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \tif (type == OBJ_REF_DELTA) {\n \t\toidread(&base_oid, fill(the_hash_algo->rawsz));\n \t\tuse(the_hash_algo->rawsz);\n-\t\tdelta_data = get_data(delta_size);\n+\t\tdelta_data = get_data(delta_size, dry_run);\n \t\tif (dry_run || !delta_data) {\n \t\t\tfree(delta_data);\n \t\t\treturn;\n@@ -396,7 +402,7 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\tif (base_offset <= 0 || base_offset >= obj_list[nr].offset)\n \t\t\tdie(\"offset value out of bound for delta base object\");\n \n-\t\tdelta_data = get_data(delta_size);\n+\t\tdelta_data = get_data(delta_size, dry_run);\n \t\tif (dry_run || !delta_data) {\n \t\t\tfree(delta_data);\n \t\t\treturn;\n-- \n2.34.0.6.g676eedc724\n\n"},{"id":"441901","messageId":"20211122033220.32883-6-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211009082058.41138-1-chiyutianyi@gmail.com","subject":"[PATCH v3 5/5] unpack-objects: unpack_non_delta_entry() read data in a stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-11-22T03:32:20Z","receivedAt":"2021-11-22T03:35:28Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWe used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\nentire contents of a blob object, no matter how big it is. This\nimplementation may consume all the memory and cause OOM.\n\nBy implementing a zstream version of input_stream interface, we can use\na small fixed buffer for \"unpack_non_delta_entry()\".\n\nHowever, unpack non-delta objects from a stream instead of from an entrie\nbuffer will have 10% performance penalty. Therefore, only unpack object\nlarger than the \"big_file_threshold\" in zstream. See the following\nbenchmarks:\n\n    $ hyperfine \\\n    --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n    'git -C dest.git unpack-objects <binary_320M.pack'\n    Benchmark 1: git -C dest.git unpack-objects <binary_320M.pack\n      Time (mean ± σ):     10.029 s ±  0.270 s    [User: 8.265 s, System: 1.522 s]\n      Range (min … max):    9.786 s … 10.603 s    10 runs\n\n    $ hyperfine \\\n    --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n    'git -c core.bigFileThreshold=2m -C dest.git unpack-objects <binary_320M.pack'\n    Benchmark 1: git -c core.bigFileThreshold=2m -C dest.git unpack-objects <binary_320M.pack\n      Time (mean ± σ):     10.859 s ±  0.774 s    [User: 8.813 s, System: 1.898 s]\n      Range (min … max):    9.884 s … 12.192 s    10 runs\n\n    $ hyperfine \\\n    --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n    'git -C dest.git unpack-objects <binary_96M.pack'\n    Benchmark 1: git -C dest.git unpack-objects <binary_96M.pack\n      Time (mean ± σ):      2.678 s ±  0.037 s    [User: 2.205 s, System: 0.450 s]\n      Range (min … max):    2.639 s …  2.743 s    10 runs\n\n    $ hyperfine \\\n    --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n    'git -c core.bigFileThreshold=2m -C dest.git unpack-objects <binary_96M.pack'\n    Benchmark 1: git -c core.bigFileThreshold=2m -C dest.git unpack-objects <binary_96M.pack\n      Time (mean ± σ):      2.819 s ±  0.124 s    [User: 2.216 s, System: 0.564 s]\n      Range (min … max):    2.679 s …  3.125 s    10 runs\n\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n builtin/unpack-objects.c            | 76 ++++++++++++++++++++++++++++-\n object-file.c                       |  6 +--\n object-store.h                      |  4 ++\n t/t5590-unpack-non-delta-objects.sh | 76 +++++++++++++++++++++++++++++\n 4 files changed, 158 insertions(+), 4 deletions(-)\n create mode 100755 t/t5590-unpack-non-delta-objects.sh\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 8d68acd662..bfc254a236 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -326,11 +326,85 @@ static void added_object(unsigned nr, enum object_type type,\n \t}\n }\n \n+struct input_zstream_data {\n+\tgit_zstream *zstream;\n+\tunsigned char buf[4096];\n+\tint status;\n+};\n+\n+static const void *feed_input_zstream(struct input_stream *in_stream, unsigned long *readlen)\n+{\n+\tstruct input_zstream_data *data = in_stream->data;\n+\tgit_zstream *zstream = data->zstream;\n+\tvoid *in = fill(1);\n+\n+\tif (!len || data->status == Z_STREAM_END) {\n+\t\t*readlen = 0;\n+\t\treturn NULL;\n+\t}\n+\n+\tzstream->next_out = data->buf;\n+\tzstream->avail_out = sizeof(data->buf);\n+\tzstream->next_in = in;\n+\tzstream->avail_in = len;\n+\n+\tdata->status = git_inflate(zstream, 0);\n+\tuse(len - zstream->avail_in);\n+\t*readlen = sizeof(data->buf) - zstream->avail_out;\n+\n+\treturn data->buf;\n+}\n+\n+static void write_stream_blob(unsigned nr, unsigned long size)\n+{\n+\tchar hdr[32];\n+\tint hdrlen;\n+\tgit_zstream zstream;\n+\tstruct input_zstream_data data;\n+\tstruct input_stream in_stream = {\n+\t\t.read = feed_input_zstream,\n+\t\t.data = &data,\n+\t};\n+\tstruct object_id *oid = &obj_list[nr].oid;\n+\tint ret;\n+\n+\tmemset(&zstream, 0, sizeof(zstream));\n+\tmemset(&data, 0, sizeof(data));\n+\tdata.zstream = &zstream;\n+\tgit_inflate_init(&zstream);\n+\n+\t/* Generate the header */\n+\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX, type_name(OBJ_BLOB), (uintmax_t)size) + 1;\n+\n+\tif ((ret = write_loose_object(oid, hdr, hdrlen, &in_stream, 0, 0)))\n+\t\tdie(_(\"failed to write object in stream %d\"), ret);\n+\n+\tif (zstream.total_out != size || data.status != Z_STREAM_END)\n+\t\tdie(_(\"inflate returned %d\"), data.status);\n+\tgit_inflate_end(&zstream);\n+\n+\tif (strict && !dry_run) {\n+\t\tstruct blob *blob = lookup_blob(the_repository, oid);\n+\t\tif (blob)\n+\t\t\tblob->object.flags |= FLAG_WRITTEN;\n+\t\telse\n+\t\t\tdie(\"invalid blob object from stream\");\n+\t}\n+\tobj_list[nr].obj = NULL;\n+}\n+\n static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n \t\t\t\t   unsigned nr)\n {\n-\tvoid *buf = get_data(size, dry_run);\n+\tvoid *buf;\n+\n+\t/* Write large blob in stream without allocating full buffer. */\n+\tif (!dry_run && type == OBJ_BLOB && size > big_file_threshold) {\n+\t\twrite_stream_blob(nr, size);\n+\t\treturn;\n+\t}\n \n+\tbuf = get_data(size, dry_run);\n \tif (!dry_run && buf)\n \t\twrite_object(nr, type, buf, size);\n \telse\ndiff --git a/object-file.c b/object-file.c\nindex 93bcfaca50..bd7631f7ef 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1878,9 +1878,9 @@ static const void *feed_simple_input_stream(struct input_stream *in_stream, unsi\n \treturn data->buf;\n }\n \n-static int write_loose_object(const struct object_id *oid, char *hdr,\n-\t\t\t      int hdrlen, struct input_stream *in_stream,\n-\t\t\t      time_t mtime, unsigned flags)\n+int write_loose_object(const struct object_id *oid, char *hdr,\n+\t\t       int hdrlen, struct input_stream *in_stream,\n+\t\t       time_t mtime, unsigned flags)\n {\n \tint fd, ret;\n \tunsigned char compressed[4096];\ndiff --git a/object-store.h b/object-store.h\nindex ccc1fc9c1a..cbd95c47e2 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -228,6 +228,10 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n \t\t     unsigned long len, const char *type,\n \t\t     struct object_id *oid);\n \n+int write_loose_object(const struct object_id *oid, char *hdr,\n+\t\t       int hdrlen, struct input_stream *in_stream,\n+\t\t       time_t mtime, unsigned flags);\n+\n int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t    const char *type, struct object_id *oid,\n \t\t\t    unsigned flags);\ndiff --git a/t/t5590-unpack-non-delta-objects.sh b/t/t5590-unpack-non-delta-objects.sh\nnew file mode 100755\nindex 0000000000..01d950d119\n--- /dev/null\n+++ b/t/t5590-unpack-non-delta-objects.sh\n@@ -0,0 +1,76 @@\n+#!/bin/sh\n+#\n+# Copyright (c) 2021 Han Xin\n+#\n+\n+test_description='Test unpack-objects when receive pack'\n+\n+GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n+export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n+\n+. ./test-lib.sh\n+\n+test_expect_success \"create commit with big blobs (1.5 MB)\" '\n+\ttest-tool genrandom foo 1500000 >big-blob &&\n+\ttest_commit --append foo big-blob &&\n+\ttest-tool genrandom bar 1500000 >big-blob &&\n+\ttest_commit --append bar big-blob &&\n+\t(\n+\t\tcd .git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >expect &&\n+\tPACK=$(echo main | git pack-objects --progress --revs test)\n+'\n+\n+test_expect_success 'setup GIT_ALLOC_LIMIT to 1MB' '\n+\tGIT_ALLOC_LIMIT=1m &&\n+\texport GIT_ALLOC_LIMIT\n+'\n+\n+test_expect_success 'prepare dest repository' '\n+\tgit init --bare dest.git &&\n+\tgit -C dest.git config core.bigFileThreshold 2m &&\n+\tgit -C dest.git config receive.unpacklimit 100\n+'\n+\n+test_expect_success 'fail to unpack-objects: cannot allocate' '\n+\ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n+\ttest_i18ngrep \"fatal: attempting to allocate\" err &&\n+\t(\n+\t\tcd dest.git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >actual &&\n+\t! test_cmp expect actual\n+'\n+\n+test_expect_success 'set a lower bigfile threshold' '\n+\tgit -C dest.git config core.bigFileThreshold 1m\n+'\n+\n+test_expect_success 'unpack big object in stream' '\n+\tgit -C dest.git unpack-objects <test-$PACK.pack &&\n+\tgit -C dest.git fsck &&\n+\t(\n+\t\tcd dest.git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'setup for unpack-objects dry-run test' '\n+\tgit init --bare unpack-test.git\n+'\n+\n+test_expect_success 'unpack-objects dry-run' '\n+\t(\n+\t\tcd unpack-test.git &&\n+\t\tgit unpack-objects -n <../test-$PACK.pack\n+\t) &&\n+\t(\n+\t\tcd unpack-test.git &&\n+\t\tfind objects/ -type f\n+\t) >actual &&\n+\ttest_must_be_empty actual\n+'\n+\n+test_done\n-- \n2.34.0.6.g676eedc724\n\n"},{"id":"442209","messageId":"xmqqczmq78x7.fsf@gitster.g","threadId":"56672","inReplyTo":"20211122033220.32883-2-chiyutianyi@gmail.com","subject":"Re: [PATCH v3 1/5] object-file: refactor write_loose_object() to read buffer from stream","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-11-23T23:24:04Z","receivedAt":"2021-11-23T23:24:10Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Han Xin <chiyutianyi@gmail.com> writes:\n\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> We used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\n> entire contents of a blob object, no matter how big it is. This\n> implementation may consume all the memory and cause OOM.\n>\n> This can be improved by feeding data to \"write_loose_object()\" in a\n> stream. The input stream is implemented as an interface. In the first\n> step, we make a simple implementation, feeding the entire buffer in the\n> \"stream\" to \"write_loose_object()\" as a refactor.\n\nPossibly a stupid question (not a review).\n\nHow does this compare with \"struct git_istream\" implemented for a\nfew existing codepaths?  It seems that the existing users are\npack-objects, index-pack and archive and all of them use the\ninterface to obtain data given an object name without having to grab\neverything in core at once.\n\nIf we are adding a new streaming interface to go in the opposite\ndirection, i.e. from the working tree data to object store, I would\nunderstand it as a complementary interface (but then I suspect there\nis a half of it already in bulk-checkin API), but I am not sure how\nthis new thing fits in the larger picture.\n\n\n\n> Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  object-file.c  | 50 ++++++++++++++++++++++++++++++++++++++++++++++----\n>  object-store.h |  5 +++++\n>  2 files changed, 51 insertions(+), 4 deletions(-)\n>\n> diff --git a/object-file.c b/object-file.c\n> index c3d866a287..227f53a0de 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1860,8 +1860,26 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n>  \treturn fd;\n>  }\n>  \n> +struct simple_input_stream_data {\n> +\tconst void *buf;\n> +\tunsigned long len;\n> +};\n> +\n> +static const void *feed_simple_input_stream(struct input_stream *in_stream, unsigned long *len)\n> +{\n> +\tstruct simple_input_stream_data *data = in_stream->data;\n> +\n> +\tif (data->len == 0) {\n> +\t\t*len = 0;\n> +\t\treturn NULL;\n> +\t}\n> +\t*len = data->len;\n> +\tdata->len = 0;\n> +\treturn data->buf;\n> +}\n> +\n>  static int write_loose_object(const struct object_id *oid, char *hdr,\n> -\t\t\t      int hdrlen, const void *buf, unsigned long len,\n> +\t\t\t      int hdrlen, struct input_stream *in_stream,\n>  \t\t\t      time_t mtime, unsigned flags)\n>  {\n>  \tint fd, ret;\n> @@ -1871,6 +1889,8 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \tstruct object_id parano_oid;\n>  \tstatic struct strbuf tmp_file = STRBUF_INIT;\n>  \tstatic struct strbuf filename = STRBUF_INIT;\n> +\tconst void *buf;\n> +\tunsigned long len;\n>  \n>  \tloose_object_path(the_repository, &filename, oid);\n>  \n> @@ -1898,6 +1918,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n>  \n>  \t/* Then the data itself.. */\n> +\tbuf = in_stream->read(in_stream, &len);\n>  \tstream.next_in = (void *)buf;\n>  \tstream.avail_in = len;\n>  \tdo {\n> @@ -1960,6 +1981,13 @@ int write_object_file_flags(const void *buf, unsigned long len,\n>  {\n>  \tchar hdr[MAX_HEADER_LEN];\n>  \tint hdrlen = sizeof(hdr);\n> +\tstruct input_stream in_stream = {\n> +\t\t.read = feed_simple_input_stream,\n> +\t\t.data = (void *)&(struct simple_input_stream_data) {\n> +\t\t\t.buf = buf,\n> +\t\t\t.len = len,\n> +\t\t},\n> +\t};\n>  \n>  \t/* Normally if we have it in the pack then we do not bother writing\n>  \t * it out into .git/objects/??/?{38} file.\n> @@ -1968,7 +1996,7 @@ int write_object_file_flags(const void *buf, unsigned long len,\n>  \t\t\t\t  &hdrlen);\n>  \tif (freshen_packed_object(oid) || freshen_loose_object(oid))\n>  \t\treturn 0;\n> -\treturn write_loose_object(oid, hdr, hdrlen, buf, len, 0, flags);\n> +\treturn write_loose_object(oid, hdr, hdrlen, &in_stream, 0, flags);\n>  }\n>  \n>  int hash_object_file_literally(const void *buf, unsigned long len,\n> @@ -1977,6 +2005,13 @@ int hash_object_file_literally(const void *buf, unsigned long len,\n>  {\n>  \tchar *header;\n>  \tint hdrlen, status = 0;\n> +\tstruct input_stream in_stream = {\n> +\t\t.read = feed_simple_input_stream,\n> +\t\t.data = (void *)&(struct simple_input_stream_data) {\n> +\t\t\t.buf = buf,\n> +\t\t\t.len = len,\n> +\t\t},\n> +\t};\n>  \n>  \t/* type string, SP, %lu of the length plus NUL must fit this */\n>  \thdrlen = strlen(type) + MAX_HEADER_LEN;\n> @@ -1988,7 +2023,7 @@ int hash_object_file_literally(const void *buf, unsigned long len,\n>  \t\tgoto cleanup;\n>  \tif (freshen_packed_object(oid) || freshen_loose_object(oid))\n>  \t\tgoto cleanup;\n> -\tstatus = write_loose_object(oid, header, hdrlen, buf, len, 0, 0);\n> +\tstatus = write_loose_object(oid, header, hdrlen, &in_stream, 0, 0);\n>  \n>  cleanup:\n>  \tfree(header);\n> @@ -2003,14 +2038,21 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n>  \tchar hdr[MAX_HEADER_LEN];\n>  \tint hdrlen;\n>  \tint ret;\n> +\tstruct simple_input_stream_data data;\n> +\tstruct input_stream in_stream = {\n> +\t\t.read = feed_simple_input_stream,\n> +\t\t.data = &data,\n> +\t};\n>  \n>  \tif (has_loose_object(oid))\n>  \t\treturn 0;\n>  \tbuf = read_object(the_repository, oid, &type, &len);\n>  \tif (!buf)\n>  \t\treturn error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n> +\tdata.buf = buf;\n> +\tdata.len = len;\n>  \thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n> -\tret = write_loose_object(oid, hdr, hdrlen, buf, len, mtime, 0);\n> +\tret = write_loose_object(oid, hdr, hdrlen, &in_stream, mtime, 0);\n>  \tfree(buf);\n>  \n>  \treturn ret;\n> diff --git a/object-store.h b/object-store.h\n> index 952efb6a4b..ccc1fc9c1a 100644\n> --- a/object-store.h\n> +++ b/object-store.h\n> @@ -34,6 +34,11 @@ struct object_directory {\n>  \tchar *path;\n>  };\n>  \n> +struct input_stream {\n> +\tconst void *(*read)(struct input_stream *, unsigned long *len);\n> +\tvoid *data;\n> +};\n> +\n>  KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n>  \tstruct object_directory *, 1, fspathhash, fspatheq)\n"},{"id":"442235","messageId":"CAO0brD3tQuzyHXbdndJrgqYd21pANYntmg5h7YKasne7QQ6Now@mail.gmail.com","threadId":"56672","inReplyTo":"xmqqczmq78x7.fsf@gitster.g","subject":"Re: [PATCH v3 1/5] object-file: refactor write_loose_object() to read buffer from stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-11-24T09:00:00Z","receivedAt":"2021-11-24T09:00:16Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n>\n> Han Xin <chiyutianyi@gmail.com> writes:\n>\n> > From: Han Xin <hanxin.hx@alibaba-inc.com>\n> >\n> > We used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\n> > entire contents of a blob object, no matter how big it is. This\n> > implementation may consume all the memory and cause OOM.\n> >\n> > This can be improved by feeding data to \"write_loose_object()\" in a\n> > stream. The input stream is implemented as an interface. In the first\n> > step, we make a simple implementation, feeding the entire buffer in the\n> > \"stream\" to \"write_loose_object()\" as a refactor.\n>\n> Possibly a stupid question (not a review).\n>\n> How does this compare with \"struct git_istream\" implemented for a\n> few existing codepaths?  It seems that the existing users are\n> pack-objects, index-pack and archive and all of them use the\n> interface to obtain data given an object name without having to grab\n> everything in core at once.\n>\n> If we are adding a new streaming interface to go in the opposite\n> direction, i.e. from the working tree data to object store, I would\n> understand it as a complementary interface (but then I suspect there\n> is a half of it already in bulk-checkin API), but I am not sure how\n> this new thing fits in the larger picture.\n>\n\nThank you for your reply.\n\nBefore starting to make this patch, I did consider whether I should\nreuse \"struct  git_istream\" to solve the problem, but I found that in the\nprocess of git unpack-objects, the data comes from stdin, and we\ncannot get an oid in advance until the whole object data is read.\nAlso, we can't do \"lseek()“ on stdin to change the data reading position.\n\nI compared the implementation of \"bulk-checkin\", and they do have\nsome similarities.\nI think the difference in the reverse implementation is that we do not\nalways clearly know where the boundary of the target data is. For\nexample, in the process of \"unpack-objects\", the \"buffer\" has been\npartially read after calling \"fill()\". And the \"buffer\" remaining after\nreading cannot be discarded because it is the beginning of the next\nobject.\nPerhaps \"struct input_stream\" can make some improvements to\n\"index_bulk_checkin()\", so that it can read from an inner buffer in\naddition to reading from \"fd\" if necessary.\n\n>\n>\n> > Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> > Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> > ---\n> >  object-file.c  | 50 ++++++++++++++++++++++++++++++++++++++++++++++----\n> >  object-store.h |  5 +++++\n> >  2 files changed, 51 insertions(+), 4 deletions(-)\n> >\n> > diff --git a/object-file.c b/object-file.c\n> > index c3d866a287..227f53a0de 100644\n> > --- a/object-file.c\n> > +++ b/object-file.c\n> > @@ -1860,8 +1860,26 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n> >       return fd;\n> >  }\n> >\n> > +struct simple_input_stream_data {\n> > +     const void *buf;\n> > +     unsigned long len;\n> > +};\n> > +\n> > +static const void *feed_simple_input_stream(struct input_stream *in_stream, unsigned long *len)\n> > +{\n> > +     struct simple_input_stream_data *data = in_stream->data;\n> > +\n> > +     if (data->len == 0) {\n> > +             *len = 0;\n> > +             return NULL;\n> > +     }\n> > +     *len = data->len;\n> > +     data->len = 0;\n> > +     return data->buf;\n> > +}\n> > +\n> >  static int write_loose_object(const struct object_id *oid, char *hdr,\n> > -                           int hdrlen, const void *buf, unsigned long len,\n> > +                           int hdrlen, struct input_stream *in_stream,\n> >                             time_t mtime, unsigned flags)\n> >  {\n> >       int fd, ret;\n> > @@ -1871,6 +1889,8 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n> >       struct object_id parano_oid;\n> >       static struct strbuf tmp_file = STRBUF_INIT;\n> >       static struct strbuf filename = STRBUF_INIT;\n> > +     const void *buf;\n> > +     unsigned long len;\n> >\n> >       loose_object_path(the_repository, &filename, oid);\n> >\n> > @@ -1898,6 +1918,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n> >       the_hash_algo->update_fn(&c, hdr, hdrlen);\n> >\n> >       /* Then the data itself.. */\n> > +     buf = in_stream->read(in_stream, &len);\n> >       stream.next_in = (void *)buf;\n> >       stream.avail_in = len;\n> >       do {\n> > @@ -1960,6 +1981,13 @@ int write_object_file_flags(const void *buf, unsigned long len,\n> >  {\n> >       char hdr[MAX_HEADER_LEN];\n> >       int hdrlen = sizeof(hdr);\n> > +     struct input_stream in_stream = {\n> > +             .read = feed_simple_input_stream,\n> > +             .data = (void *)&(struct simple_input_stream_data) {\n> > +                     .buf = buf,\n> > +                     .len = len,\n> > +             },\n> > +     };\n> >\n> >       /* Normally if we have it in the pack then we do not bother writing\n> >        * it out into .git/objects/??/?{38} file.\n> > @@ -1968,7 +1996,7 @@ int write_object_file_flags(const void *buf, unsigned long len,\n> >                                 &hdrlen);\n> >       if (freshen_packed_object(oid) || freshen_loose_object(oid))\n> >               return 0;\n> > -     return write_loose_object(oid, hdr, hdrlen, buf, len, 0, flags);\n> > +     return write_loose_object(oid, hdr, hdrlen, &in_stream, 0, flags);\n> >  }\n> >\n> >  int hash_object_file_literally(const void *buf, unsigned long len,\n> > @@ -1977,6 +2005,13 @@ int hash_object_file_literally(const void *buf, unsigned long len,\n> >  {\n> >       char *header;\n> >       int hdrlen, status = 0;\n> > +     struct input_stream in_stream = {\n> > +             .read = feed_simple_input_stream,\n> > +             .data = (void *)&(struct simple_input_stream_data) {\n> > +                     .buf = buf,\n> > +                     .len = len,\n> > +             },\n> > +     };\n> >\n> >       /* type string, SP, %lu of the length plus NUL must fit this */\n> >       hdrlen = strlen(type) + MAX_HEADER_LEN;\n> > @@ -1988,7 +2023,7 @@ int hash_object_file_literally(const void *buf, unsigned long len,\n> >               goto cleanup;\n> >       if (freshen_packed_object(oid) || freshen_loose_object(oid))\n> >               goto cleanup;\n> > -     status = write_loose_object(oid, header, hdrlen, buf, len, 0, 0);\n> > +     status = write_loose_object(oid, header, hdrlen, &in_stream, 0, 0);\n> >\n> >  cleanup:\n> >       free(header);\n> > @@ -2003,14 +2038,21 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n> >       char hdr[MAX_HEADER_LEN];\n> >       int hdrlen;\n> >       int ret;\n> > +     struct simple_input_stream_data data;\n> > +     struct input_stream in_stream = {\n> > +             .read = feed_simple_input_stream,\n> > +             .data = &data,\n> > +     };\n> >\n> >       if (has_loose_object(oid))\n> >               return 0;\n> >       buf = read_object(the_repository, oid, &type, &len);\n> >       if (!buf)\n> >               return error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n> > +     data.buf = buf;\n> > +     data.len = len;\n> >       hdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n> > -     ret = write_loose_object(oid, hdr, hdrlen, buf, len, mtime, 0);\n> > +     ret = write_loose_object(oid, hdr, hdrlen, &in_stream, mtime, 0);\n> >       free(buf);\n> >\n> >       return ret;\n> > diff --git a/object-store.h b/object-store.h\n> > index 952efb6a4b..ccc1fc9c1a 100644\n> > --- a/object-store.h\n> > +++ b/object-store.h\n> > @@ -34,6 +34,11 @@ struct object_directory {\n> >       char *path;\n> >  };\n> >\n> > +struct input_stream {\n> > +     const void *(*read)(struct input_stream *, unsigned long *len);\n> > +     void *data;\n> > +};\n> > +\n> >  KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n> >       struct object_directory *, 1, fspathhash, fspatheq)\n"},{"id":"442431","messageId":"CAO0brD3VPtUrpCE2kCJDram=bLMN=89++=bgf1TddriTYo-nsA@mail.gmail.com","threadId":"56672","inReplyTo":"20211122033220.32883-1-chiyutianyi@gmail.com","subject":"Re: [PATCH v3 0/5] unpack large objects in stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-11-29T07:01:47Z","receivedAt":"2021-11-29T07:06:54Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"Han Xin <chiyutianyi@gmail.com> writes:\n>\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> Although we do not recommend users push large binary files to the git repositories,\n> it's difficult to prevent them from doing so. Once, we found a problem with a surge\n> in memory usage on the server. The source of the problem is that a user submitted\n> a single object with a size of 15GB. Once someone initiates a git push, the git\n> process will immediately allocate 15G of memory, resulting in an OOM risk.\n>\n> Through further analysis, we found that when we execute git unpack-objects, in\n> unpack_non_delta_entry(), \"void *buf = get_data(size);\" will directly allocate\n> memory equal to the size of the object. This is quite a scary thing, because the\n> pre-receive hook has not been executed at this time, and we cannot avoid this by hooks.\n>\n> I got inspiration from the deflate process of zlib, maybe it would be a good idea\n> to change unpack-objects to stream deflate.\n>\n\nHi, Jeff.\n\nI hope you can share with me how Github solves this problem.\n\nAs you said in your reply at：\nhttps://lore.kernel.org/git/YVaw6agcPNclhws8@coredump.intra.peff.net/\n\"we don't have a match in unpack-objects, but we always run index-pack\non incoming packs\".\n\nIn the original implementation of \"index-pack\", for objects larger than\nbig_file_threshold, \"fixed_buf\" with a size of 8192 will be used to\ncomplete the calculation of \"oid\".\n\nI tried the implementation in jk/no-more-unpack-objects, as you noted:\n  /* XXX This will expand too-large objects! */\n  if (!data)\n  data = new_data = get_data_from_pack(obj_entry);\nIf the conditions of --unpack are given, there will be risks here.\nWhen I create an object larger than 1GB and execute index-pack, the\nresult is as follows:\n  $GIT_ALLOC_LIMIT=1024m git index-pack --unpack --stdin <large.pack\n  fatal: attempting to allocate 1228800001 over limit 1073741824\n\nLooking forward to your reply.\n"},{"id":"442469","messageId":"47b3e2ad-4fa1-040a-24c1-6da0445bd1a5@gmail.com","threadId":"56672","inReplyTo":"20211122033220.32883-3-chiyutianyi@gmail.com","subject":"Re: [PATCH v3 2/5] object-file.c: handle undetermined oid in write_loose_object()","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-11-29T15:10:39Z","receivedAt":"2021-11-29T18:50:55Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 11/21/2021 10:32 PM, Han Xin wrote:\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n> \n> When streaming a large blob object to \"write_loose_object()\", we have no\n> chance to run \"write_object_file_prepare()\" to calculate the oid in\n> advance. So we need to handle undetermined oid in function\n> \"write_loose_object()\".\n> \n> In the original implementation, we know the oid and we can write the\n> temporary file in the same directory as the final object, but for an\n> object with an undetermined oid, we don't know the exact directory for\n> the object, so we have to save the temporary file in \".git/objects/\"\n> directory instead.\n\nMy first reaction is to not write into .git/objects/ directly, but\ninstead make a .git/objects/tmp/ directory and write within that\ndirectory. The idea is to prevent leaving stale files in the\n.git/objects/ directory if the process terminates strangely (say,\na power outage or segfault).\n\nIf this was an interesting idea to pursue, it does leave a question:\nshould we clean up the tmp/ directory when it is empty? That would\nrequire adding a check in finalize_object_file() that is probably\nbest left unchecked (the lstat() would add a cost per loose object\nwrite that is probably too costly). I would rather leave an empty\ntmp/ directory than add that cost per loose object write.\n\nI suppose another way to do it would be to register the check as\nan event at the end of the process, so we only check once, and\nthat only happens if we created a loose object with this streaming\nmethod.\n\nWith all of these complications in mind, I think cleaning up the\nstale tmp/ directory could (at the very least) be delayed to another\ncommit or patch series. Hopefully adding the directory is not too\nmuch complication to add here.\n\n> -\tloose_object_path(the_repository, &filename, oid);\n> +\tif (is_null_oid(oid)) {\n> +\t\t/* When oid is not determined, save tmp file to odb path. */\n> +\t\tstrbuf_reset(&filename);\n> +\t\tstrbuf_addstr(&filename, the_repository->objects->odb->path);\n> +\t\tstrbuf_addch(&filename, '/');\n\nHere, you could instead of the strbuf_addch() do\n\n\tstrbuf_add(&filename, \"/tmp/\", 5);\n\tif (safe_create_leading_directories(filename.buf)) {\n\t\terror(_(\"failed to create '%s'\"));\n\t\tstrbuf_release(&filename);\n\t\treturn -1;\n\t}\t\t\n\n> +\t} else {\n> +\t\tloose_object_path(the_repository, &filename, oid);\n> +\t}\n>  \n>  \tfd = create_tmpfile(&tmp_file, filename.buf);\n>  \tif (fd < 0) {\n> @@ -1939,12 +1946,31 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n>  \t\t    ret);\n>  \tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n> -\tif (!oideq(oid, &parano_oid))\n> +\tif (!is_null_oid(oid) && !oideq(oid, &parano_oid))\n>  \t\tdie(_(\"confused by unstable object source data for %s\"),\n>  \t\t    oid_to_hex(oid));\n>  \n>  \tclose_loose_object(fd);\n>  \n> +\tif (is_null_oid(oid)) {\n> +\t\tint dirlen;\n> +\n> +\t\toidcpy((struct object_id *)oid, &parano_oid);\n> +\t\tloose_object_path(the_repository, &filename, oid);\n> +\n> +\t\t/* We finally know the object path, and create the missing dir. */\n> +\t\tdirlen = directory_size(filename.buf);\n> +\t\tif (dirlen) {\n> +\t\t\tstruct strbuf dir = STRBUF_INIT;\n> +\t\t\tstrbuf_add(&dir, filename.buf, dirlen - 1);\n> +\t\t\tif (mkdir(dir.buf, 0777) && errno != EEXIST)\n> +\t\t\t\treturn -1;\n> +\t\t\tif (adjust_shared_perm(dir.buf))\n> +\t\t\t\treturn -1;\n> +\t\t\tstrbuf_release(&dir);\n> +\t\t}\n> +\t}\n> +\n\nUpon first reading I was asking \"where is the file rename?\" but\nit is part of finalize_object_file() which is called further down.\n\nThanks,\n-Stolee\n"},{"id":"442470","messageId":"YaUmFpIeCvHdKixj@coredump.intra.peff.net","threadId":"56672","inReplyTo":"CAO0brD3VPtUrpCE2kCJDram=bLMN=89++=bgf1TddriTYo-nsA@mail.gmail.com","subject":"Re: [PATCH v3 0/5] unpack large objects in stream","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-11-29T19:12:22Z","receivedAt":"2021-11-29T19:14:25Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Nov 29, 2021 at 03:01:47PM +0800, Han Xin wrote:\n\n> Han Xin <chiyutianyi@gmail.com> writes:\n> >\n> > From: Han Xin <hanxin.hx@alibaba-inc.com>\n> >\n> > Although we do not recommend users push large binary files to the git repositories,\n> > it's difficult to prevent them from doing so. Once, we found a problem with a surge\n> > in memory usage on the server. The source of the problem is that a user submitted\n> > a single object with a size of 15GB. Once someone initiates a git push, the git\n> > process will immediately allocate 15G of memory, resulting in an OOM risk.\n> >\n> > Through further analysis, we found that when we execute git unpack-objects, in\n> > unpack_non_delta_entry(), \"void *buf = get_data(size);\" will directly allocate\n> > memory equal to the size of the object. This is quite a scary thing, because the\n> > pre-receive hook has not been executed at this time, and we cannot avoid this by hooks.\n> >\n> > I got inspiration from the deflate process of zlib, maybe it would be a good idea\n> > to change unpack-objects to stream deflate.\n> >\n> \n> Hi, Jeff.\n> \n> I hope you can share with me how Github solves this problem.\n> \n> As you said in your reply at：\n> https://lore.kernel.org/git/YVaw6agcPNclhws8@coredump.intra.peff.net/\n> \"we don't have a match in unpack-objects, but we always run index-pack\n> on incoming packs\".\n> \n> In the original implementation of \"index-pack\", for objects larger than\n> big_file_threshold, \"fixed_buf\" with a size of 8192 will be used to\n> complete the calculation of \"oid\".\n\nWe set transfer.unpackLimit to \"1\", so we never run unpack-objects at\nall. We always run index-pack, and every push, no matter how small,\nresults in a pack.\n\nWe also set GIT_ALLOC_LIMIT to limit any single allocation. We also have\ncustom code in index-pack to detect large objects (where our definition\nof \"large\" is 100MB by default):\n\n  - for large blobs, we do index it as normal, writing the oid out to a\n    file which is then processed by a pre-receive hook (since people\n    often push up large files accidentally, the hook generates a nice\n    error message, including finding the path at which the blob is\n    referenced)\n\n  - for other large objects, we die immediately (with an error message).\n    100MB commit messages aren't a common user error, and it closes off\n    a whole set of possible integer-overflow parsing attacks (e.g.,\n    index-pack in strict-mode will run every tree through fsck_tree(),\n    so there's otherwise nothing stopping you from having a 4GB filename\n    in a tree).\n\n> I tried the implementation in jk/no-more-unpack-objects, as you noted:\n>   /* XXX This will expand too-large objects! */\n>   if (!data)\n>   data = new_data = get_data_from_pack(obj_entry);\n> If the conditions of --unpack are given, there will be risks here.\n> When I create an object larger than 1GB and execute index-pack, the\n> result is as follows:\n>   $GIT_ALLOC_LIMIT=1024m git index-pack --unpack --stdin <large.pack\n>   fatal: attempting to allocate 1228800001 over limit 1073741824\n\nYeah, that issue was one of the reasons I never sent the \"index-pack\n--unpack\" code to the list. We don't actually use those patches at\nGitHub. It was something I was working on for upstream but never\nfinished.\n\n-Peff\n"},{"id":"442488","messageId":"xmqqsfve669w.fsf@gitster.g","threadId":"56672","inReplyTo":"47b3e2ad-4fa1-040a-24c1-6da0445bd1a5@gmail.com","subject":"Re: [PATCH v3 2/5] object-file.c: handle undetermined oid in write_loose_object()","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-11-29T20:44:43Z","receivedAt":"2021-11-29T20:46:49Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Derrick Stolee <stolee@gmail.com> writes:\n\n> My first reaction is to not write into .git/objects/ directly, but\n> instead make a .git/objects/tmp/ directory and write within that\n> directory. The idea is to prevent leaving stale files in the\n> .git/objects/ directory if the process terminates strangely (say,\n> a power outage or segfault).\n\nEven if we know the name of the object we are writing beforehand, I\ndo not think it is a good idea to open-write-close the final object\nfile.  The approach we already use everywhere is to write into a\ntmpfile/lockfile and rename it to the final name \n\nobject-file.c::write_loose_object() uses create_tmpfile() to prepare\na temporary file whose name begins with \"tmp_obj_\", so that \"gc\" can\nrecognize stale ones and remove them.\n\n> If this was an interesting idea to pursue, it does leave a question:\n> should we clean up the tmp/ directory when it is empty? That would\n> require adding a check in finalize_object_file() that is probably\n> best left unchecked (the lstat() would add a cost per loose object\n> write that is probably too costly). I would rather leave an empty\n> tmp/ directory than add that cost per loose object write.\n\nI am not sure why we want a new tmp/ directory.\n"},{"id":"442533","messageId":"2271576b-79d3-7983-d3df-5548e0a12e85@gmail.com","threadId":"56672","inReplyTo":"xmqqsfve669w.fsf@gitster.g","subject":"Re: [PATCH v3 2/5] object-file.c: handle undetermined oid in write_loose_object()","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-11-29T22:18:21Z","receivedAt":"2021-11-29T22:29:23Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 11/29/2021 3:44 PM, Junio C Hamano wrote:\n> Derrick Stolee <stolee@gmail.com> writes:\n> \n>> My first reaction is to not write into .git/objects/ directly, but\n>> instead make a .git/objects/tmp/ directory and write within that\n>> directory. The idea is to prevent leaving stale files in the\n>> .git/objects/ directory if the process terminates strangely (say,\n>> a power outage or segfault).\n> \n> Even if we know the name of the object we are writing beforehand, I\n> do not think it is a good idea to open-write-close the final object\n> file.  The approach we already use everywhere is to write into a\n> tmpfile/lockfile and rename it to the final name \n> \n> object-file.c::write_loose_object() uses create_tmpfile() to prepare\n> a temporary file whose name begins with \"tmp_obj_\", so that \"gc\" can\n> recognize stale ones and remove them.\n\nThe only difference is that the tmp_obj_* file would go into the\nloose object directory corresponding to the first two hex characters\nof the OID, but that no longer happens now.\n \n>> If this was an interesting idea to pursue, it does leave a question:\n>> should we clean up the tmp/ directory when it is empty? That would\n>> require adding a check in finalize_object_file() that is probably\n>> best left unchecked (the lstat() would add a cost per loose object\n>> write that is probably too costly). I would rather leave an empty\n>> tmp/ directory than add that cost per loose object write.\n> \n> I am not sure why we want a new tmp/ directory.\n\nI'm just thinking of a case where this fails repeatedly I would\nrather have those failed tmp_obj_* files isolated in their own\ndirectory. It's an extremely minor point, so I'm fine to drop\nthe recommendation.\n\nThanks,\n-Stolee\n"},{"id":"442555","messageId":"8ff89e50-1b80-7932-f0e2-af401ee04bb1@gmail.com","threadId":"56672","inReplyTo":"20211122033220.32883-6-chiyutianyi@gmail.com","subject":"Re: [PATCH v3 5/5] unpack-objects: unpack_non_delta_entry() read data in a stream","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-11-29T17:37:13Z","receivedAt":"2021-11-29T23:02:25Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 11/21/2021 10:32 PM, Han Xin wrote:\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n> \n> We used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\n> entire contents of a blob object, no matter how big it is. This\n> implementation may consume all the memory and cause OOM.\n> \n> By implementing a zstream version of input_stream interface, we can use\n> a small fixed buffer for \"unpack_non_delta_entry()\".\n> \n> However, unpack non-delta objects from a stream instead of from an entrie\n> buffer will have 10% performance penalty. Therefore, only unpack object\n> larger than the \"big_file_threshold\" in zstream. See the following\n> benchmarks:\n> \n>     $ hyperfine \\\n>     --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n>     'git -C dest.git unpack-objects <binary_320M.pack'\n>     Benchmark 1: git -C dest.git unpack-objects <binary_320M.pack\n>       Time (mean ± σ):     10.029 s ±  0.270 s    [User: 8.265 s, System: 1.522 s]\n>       Range (min … max):    9.786 s … 10.603 s    10 runs\n> \n>     $ hyperfine \\\n>     --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n>     'git -c core.bigFileThreshold=2m -C dest.git unpack-objects <binary_320M.pack'\n>     Benchmark 1: git -c core.bigFileThreshold=2m -C dest.git unpack-objects <binary_320M.pack\n>       Time (mean ± σ):     10.859 s ±  0.774 s    [User: 8.813 s, System: 1.898 s]\n>       Range (min … max):    9.884 s … 12.192 s    10 runs\n\nIt seems that you want us to compare this pair of results, and\nhyperfine can assist with that by including multiple benchmarks\n(with labels, using '-n') as follows:\n\n$ hyperfine \\\n        --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n        -n 'old' '~/_git/git-upstream/git -C dest.git unpack-objects <big.pack' \\\n        -n 'new' '~/_git/git/git -C dest.git unpack-objects <big.pack' \\\n        -n 'new (small threshold)' '~/_git/git/git -c core.bigfilethreshold=64k -C dest.git unpack-objects <big.pack'\n\nBenchmark 1: old\n  Time (mean ± σ):     20.835 s ±  0.058 s    [User: 14.510 s, System: 6.284 s]\n  Range (min … max):   20.741 s … 20.909 s    10 runs\n \nBenchmark 2: new\n  Time (mean ± σ):     26.515 s ±  0.072 s    [User: 19.783 s, System: 6.696 s]\n  Range (min … max):   26.419 s … 26.611 s    10 runs\n \nBenchmark 3: new (small threshold)\n  Time (mean ± σ):     26.523 s ±  0.101 s    [User: 19.805 s, System: 6.680 s]\n  Range (min … max):   26.416 s … 26.739 s    10 runs\n \nSummary\n  'old' ran\n    1.27 ± 0.00 times faster than 'new'\n    1.27 ± 0.01 times faster than 'new (small threshold)'\n\n(Here, 'old' is testing a compiled version of the latest 'master'\nbranch, while 'new' has your patches applied on top.)\n\nNotice from this example I had a pack with many small objects (mostly\ncommits and trees) and I see that this change introduces significant\noverhead to this case.\n\nIt would be nice to understand this overhead and fix it before taking\nthis change any further.\n\nThanks,\n-Stolee\n"},{"id":"442594","messageId":"CAO0brD3bSaYabKgdDZjqJd97CJ2hN1XuxWGc+Ww1GZ2wpKA4ZQ@mail.gmail.com","threadId":"56672","inReplyTo":"YaUmFpIeCvHdKixj@coredump.intra.peff.net","subject":"Re: [PATCH v3 0/5] unpack large objects in stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-11-30T02:57:40Z","receivedAt":"2021-11-30T02:57:55Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Tue, Nov 30, 2021 at 3:12 AM Jeff King <peff@peff.net> wrote:\n> We set transfer.unpackLimit to \"1\", so we never run unpack-objects at\n> all. We always run index-pack, and every push, no matter how small,\n> results in a pack.\n>\n> We also set GIT_ALLOC_LIMIT to limit any single allocation. We also have\n> custom code in index-pack to detect large objects (where our definition\n> of \"large\" is 100MB by default):\n>\n>   - for large blobs, we do index it as normal, writing the oid out to a\n>     file which is then processed by a pre-receive hook (since people\n>     often push up large files accidentally, the hook generates a nice\n>     error message, including finding the path at which the blob is\n>     referenced)\n>\n>   - for other large objects, we die immediately (with an error message).\n>     100MB commit messages aren't a common user error, and it closes off\n>     a whole set of possible integer-overflow parsing attacks (e.g.,\n>     index-pack in strict-mode will run every tree through fsck_tree(),\n>     so there's otherwise nothing stopping you from having a 4GB filename\n>     in a tree).\n\nThank you very much for sharing.\n\nThe way Github handles it reminds me of what Shawn Pearce introduced in\n\"Scaling up JGit\". I guess \"mulit-pack-index\" and \"bitmap\" must play an\nimportant role in this.\n\nI will seriously consider this solution, thanks a lot.\n"},{"id":"442595","messageId":"CAO0brD1U+zbqBGjOV+Vc5+AQAdrG5ULp_0=PMYeUQSxqvWQ65w@mail.gmail.com","threadId":"56672","inReplyTo":"2271576b-79d3-7983-d3df-5548e0a12e85@gmail.com","subject":"Re: [PATCH v3 2/5] object-file.c: handle undetermined oid in write_loose_object()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-11-30T03:23:29Z","receivedAt":"2021-11-30T03:23:44Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Tue, Nov 30, 2021 at 6:18 AM Derrick Stolee <stolee@gmail.com> wrote:\n>\n> On 11/29/2021 3:44 PM, Junio C Hamano wrote:\n> > Derrick Stolee <stolee@gmail.com> writes:\n> >\n> >> My first reaction is to not write into .git/objects/ directly, but\n> >> instead make a .git/objects/tmp/ directory and write within that\n> >> directory. The idea is to prevent leaving stale files in the\n> >> .git/objects/ directory if the process terminates strangely (say,\n> >> a power outage or segfault).\n> >\n> > Even if we know the name of the object we are writing beforehand, I\n> > do not think it is a good idea to open-write-close the final object\n> > file.  The approach we already use everywhere is to write into a\n> > tmpfile/lockfile and rename it to the final name\n> >\n> > object-file.c::write_loose_object() uses create_tmpfile() to prepare\n> > a temporary file whose name begins with \"tmp_obj_\", so that \"gc\" can\n> > recognize stale ones and remove them.\n>\n> The only difference is that the tmp_obj_* file would go into the\n> loose object directory corresponding to the first two hex characters\n> of the OID, but that no longer happens now.\n>\n\nAt the beginning of this patch, I did save the temporary object in a\ntwo hex characters directory of \"null_oid\", but this is also a very\nstrange behavior. \"Gc\" will indeed clean up these tmp_obj_* files, no\nmatter if they are in .git/objects/ or .git/objects/xx.\n\nThanks,\n-Han Xin\n"},{"id":"442650","messageId":"CAO0brD0oPHMwGNQXpC2XVhU=fY7XrrtBeu-x8GmJndeVptJaBg@mail.gmail.com","threadId":"56672","inReplyTo":"8ff89e50-1b80-7932-f0e2-af401ee04bb1@gmail.com","subject":"Re: [PATCH v3 5/5] unpack-objects: unpack_non_delta_entry() read data in a stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-11-30T13:49:53Z","receivedAt":"2021-11-30T13:50:10Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Tue, Nov 30, 2021 at 1:37 AM Derrick Stolee <stolee@gmail.com> wrote:\n>\n> On 11/21/2021 10:32 PM, Han Xin wrote:\n> > From: Han Xin <hanxin.hx@alibaba-inc.com>\n> >\n> > We used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\n> > entire contents of a blob object, no matter how big it is. This\n> > implementation may consume all the memory and cause OOM.\n> >\n> > By implementing a zstream version of input_stream interface, we can use\n> > a small fixed buffer for \"unpack_non_delta_entry()\".\n> >\n> > However, unpack non-delta objects from a stream instead of from an entrie\n> > buffer will have 10% performance penalty. Therefore, only unpack object\n> > larger than the \"big_file_threshold\" in zstream. See the following\n> > benchmarks:\n> >\n> >     $ hyperfine \\\n> >     --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n> >     'git -C dest.git unpack-objects <binary_320M.pack'\n> >     Benchmark 1: git -C dest.git unpack-objects <binary_320M.pack\n> >       Time (mean ± σ):     10.029 s ±  0.270 s    [User: 8.265 s, System: 1.522 s]\n> >       Range (min … max):    9.786 s … 10.603 s    10 runs\n> >\n> >     $ hyperfine \\\n> >     --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n> >     'git -c core.bigFileThreshold=2m -C dest.git unpack-objects <binary_320M.pack'\n> >     Benchmark 1: git -c core.bigFileThreshold=2m -C dest.git unpack-objects <binary_320M.pack\n> >       Time (mean ± σ):     10.859 s ±  0.774 s    [User: 8.813 s, System: 1.898 s]\n> >       Range (min … max):    9.884 s … 12.192 s    10 runs\n>\n> It seems that you want us to compare this pair of results, and\n> hyperfine can assist with that by including multiple benchmarks\n> (with labels, using '-n') as follows:\n>\n> $ hyperfine \\\n>         --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n>         -n 'old' '~/_git/git-upstream/git -C dest.git unpack-objects <big.pack' \\\n>         -n 'new' '~/_git/git/git -C dest.git unpack-objects <big.pack' \\\n>         -n 'new (small threshold)' '~/_git/git/git -c core.bigfilethreshold=64k -C dest.git unpack-objects <big.pack'\n>\n> Benchmark 1: old\n>   Time (mean ± σ):     20.835 s ±  0.058 s    [User: 14.510 s, System: 6.284 s]\n>   Range (min … max):   20.741 s … 20.909 s    10 runs\n>\n> Benchmark 2: new\n>   Time (mean ± σ):     26.515 s ±  0.072 s    [User: 19.783 s, System: 6.696 s]\n>   Range (min … max):   26.419 s … 26.611 s    10 runs\n>\n> Benchmark 3: new (small threshold)\n>   Time (mean ± σ):     26.523 s ±  0.101 s    [User: 19.805 s, System: 6.680 s]\n>   Range (min … max):   26.416 s … 26.739 s    10 runs\n>\n> Summary\n>   'old' ran\n>     1.27 ± 0.00 times faster than 'new'\n>     1.27 ± 0.01 times faster than 'new (small threshold)'\n>\n> (Here, 'old' is testing a compiled version of the latest 'master'\n> branch, while 'new' has your patches applied on top.)\n>\n> Notice from this example I had a pack with many small objects (mostly\n> commits and trees) and I see that this change introduces significant\n> overhead to this case.\n>\n> It would be nice to understand this overhead and fix it before taking\n> this change any further.\n>\n> Thanks,\n> -Stolee\n\nCan you show me the specific information of the repository you\ntested, so that I can analyze it further.\n\nI test this repository, but did not meet the problem:\n\n Unpacking objects: 100% (18345/18345), 43.15 MiB\n\nhyperfine \\\n        --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n        -n 'old' 'git -C dest.git unpack-objects <big.pack' \\\n        -n 'new' 'new/git -C dest.git unpack-objects <big.pack' \\\n        -n 'new (small threshold)' 'new/git -c\ncore.bigfilethreshold=64k -C dest.git unpack-objects <big.pack'\nBenchmark 1: old\n  Time (mean ± σ):     17.403 s ±  0.880 s    [User: 4.996 s, System: 11.803 s]\n  Range (min … max):   15.911 s … 19.368 s    10 runs\n\nBenchmark 2: new\n  Time (mean ± σ):     17.788 s ±  0.199 s    [User: 5.054 s, System: 12.257 s]\n  Range (min … max):   17.420 s … 18.195 s    10 runs\n\nBenchmark 3: new (small threshold)\n  Time (mean ± σ):     18.433 s ±  0.711 s    [User: 4.982 s, System: 12.338 s]\n  Range (min … max):   17.518 s … 19.775 s    10 runs\n\nSummary\n  'old' ran\n    1.02 ± 0.05 times faster than 'new'\n    1.06 ± 0.07 times faster than 'new (small threshold)'\n\nThanks,\n- Han Xin\n"},{"id":"442673","messageId":"446c3677-140f-3033-138f-1ef9b1f546a5@gmail.com","threadId":"56672","inReplyTo":"CAO0brD0oPHMwGNQXpC2XVhU=fY7XrrtBeu-x8GmJndeVptJaBg@mail.gmail.com","subject":"Re: [PATCH v3 5/5] unpack-objects: unpack_non_delta_entry() read data in a stream","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-11-30T18:38:07Z","receivedAt":"2021-11-30T18:38:11Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 11/30/2021 8:49 AM, Han Xin wrote:\n> On Tue, Nov 30, 2021 at 1:37 AM Derrick Stolee <stolee@gmail.com> wrote:\n>> $ hyperfine \\\n>>         --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n>>         -n 'old' '~/_git/git-upstream/git -C dest.git unpack-objects <big.pack' \\\n>>         -n 'new' '~/_git/git/git -C dest.git unpack-objects <big.pack' \\\n>>         -n 'new (small threshold)' '~/_git/git/git -c core.bigfilethreshold=64k -C dest.git unpack-objects <big.pack'\n>>\n>> Benchmark 1: old\n>>   Time (mean ± σ):     20.835 s ±  0.058 s    [User: 14.510 s, System: 6.284 s]\n>>   Range (min … max):   20.741 s … 20.909 s    10 runs\n>>\n>> Benchmark 2: new\n>>   Time (mean ± σ):     26.515 s ±  0.072 s    [User: 19.783 s, System: 6.696 s]\n>>   Range (min … max):   26.419 s … 26.611 s    10 runs\n>>\n>> Benchmark 3: new (small threshold)\n>>   Time (mean ± σ):     26.523 s ±  0.101 s    [User: 19.805 s, System: 6.680 s]\n>>   Range (min … max):   26.416 s … 26.739 s    10 runs\n>>\n>> Summary\n>>   'old' ran\n>>     1.27 ± 0.00 times faster than 'new'\n>>     1.27 ± 0.01 times faster than 'new (small threshold)'\n>>\n>> (Here, 'old' is testing a compiled version of the latest 'master'\n>> branch, while 'new' has your patches applied on top.)\n>>\n>> Notice from this example I had a pack with many small objects (mostly\n>> commits and trees) and I see that this change introduces significant\n>> overhead to this case.\n>>\n>> It would be nice to understand this overhead and fix it before taking\n>> this change any further.\n>>\n>> Thanks,\n>> -Stolee\n> \n> Can you show me the specific information of the repository you\n> tested, so that I can analyze it further.\n\nI used a pack-file from an internal repo. It happened to be using\npartial clone, so here is a repro with the git/git repository\nafter cloning this way:\n\n$ git clone --no-checkout --filter=blob:none https://github.com/git/git\n\n(copy the large .pack from git/.git/objects/pack/ to big.pack)\n\n$ hyperfine \\\n\t--prepare 'rm -rf dest.git && git init --bare dest.git' \\\n\t-n 'old' '~/_git/git-upstream/git -C dest.git unpack-objects <big.pack' \\\n\t-n 'new' '~/_git/git/git -C dest.git unpack-objects <big.pack' \\\n\t-n 'new (small threshold)' '~/_git/git/git -c core.bigfilethreshold=64k -C dest.git unpack-objects <big.pack'\n\nBenchmark 1: old\n  Time (mean ± σ):     82.748 s ±  0.445 s    [User: 50.512 s, System: 32.049 s]\n  Range (min … max):   82.042 s … 83.587 s    10 runs\n \nBenchmark 2: new\n  Time (mean ± σ):     101.644 s ±  0.524 s    [User: 67.470 s, System: 34.047 s]\n  Range (min … max):   100.866 s … 102.633 s    10 runs\n \nBenchmark 3: new (small threshold)\n  Time (mean ± σ):     101.093 s ±  0.269 s    [User: 67.404 s, System: 33.559 s]\n  Range (min … max):   100.639 s … 101.375 s    10 runs\n \nSummary\n  'old' ran\n    1.22 ± 0.01 times faster than 'new (small threshold)'\n    1.23 ± 0.01 times faster than 'new'\n\nI'm also able to repro this with a smaller repo (microsoft/scalar)\nso the tests complete much faster:\n\n$ hyperfine \\\n        --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n        -n 'old' '~/_git/git-upstream/git -C dest.git unpack-objects <small.pack' \\\n        -n 'new' '~/_git/git/git -C dest.git unpack-objects <small.pack' \\\n        -n 'new (small threshold)' '~/_git/git/git -c core.bigfilethreshold=64k -C dest.git unpack-objects <small.pack'\n\nBenchmark 1: old\n  Time (mean ± σ):      3.295 s ±  0.023 s    [User: 1.063 s, System: 2.228 s]\n  Range (min … max):    3.269 s …  3.351 s    10 runs\n \nBenchmark 2: new\n  Time (mean ± σ):      3.592 s ±  0.105 s    [User: 1.261 s, System: 2.328 s]\n  Range (min … max):    3.378 s …  3.679 s    10 runs\n \nBenchmark 3: new (small threshold)\n  Time (mean ± σ):      3.584 s ±  0.144 s    [User: 1.241 s, System: 2.339 s]\n  Range (min … max):    3.359 s …  3.747 s    10 runs\n \nSummary\n  'old' ran\n    1.09 ± 0.04 times faster than 'new (small threshold)'\n    1.09 ± 0.03 times faster than 'new'\n\nIt's not the same relative overhead, but still significant.\n\nThese pack-files contain (mostly) small objects, no large blobs.\nI know that's not the target of your efforts, but it would be\ngood to avoid a regression here.\n\nThanks,\n-Stolee\n"},{"id":"442803","messageId":"211201.86r1aw9gbd.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"446c3677-140f-3033-138f-1ef9b1f546a5@gmail.com","subject":"\"git hyperfine\" (was: [PATCH v3 5/5] unpack-objects[...])","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-01T20:37:00Z","receivedAt":"2021-12-01T21:16:28Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nI hadn't sent a shameless plug for my \"git hyperfine\" script to the\nlist, perhaps this is a good time. It's just a thin shellscript wrapper\naround \"hyperfine\" that I wrote the other day, which...\n\nOn Tue, Nov 30 2021, Derrick Stolee wrote:\n\n> [...]\n> I used a pack-file from an internal repo. It happened to be using\n> partial clone, so here is a repro with the git/git repository\n> after cloning this way:\n>\n> $ git clone --no-checkout --filter=blob:none https://github.com/git/git\n>\n> (copy the large .pack from git/.git/objects/pack/ to big.pack)\n>\n> $ hyperfine \\\n> \t--prepare 'rm -rf dest.git && git init --bare dest.git' \\\n> \t-n 'old' '~/_git/git-upstream/git -C dest.git unpack-objects <big.pack' \\\n> \t-n 'new' '~/_git/git/git -C dest.git unpack-objects <big.pack' \\\n> \t-n 'new (small threshold)' '~/_git/git/git -c core.bigfilethreshold=64k -C dest.git unpack-objects <big.pack'\n>\n> Benchmark 1: old\n>   Time (mean ± σ):     82.748 s ±  0.445 s    [User: 50.512 s, System: 32.049 s]\n>   Range (min … max):   82.042 s … 83.587 s    10 runs\n>  \n> Benchmark 2: new\n>   Time (mean ± σ):     101.644 s ±  0.524 s    [User: 67.470 s, System: 34.047 s]\n>   Range (min … max):   100.866 s … 102.633 s    10 runs\n>  \n> Benchmark 3: new (small threshold)\n>   Time (mean ± σ):     101.093 s ±  0.269 s    [User: 67.404 s, System: 33.559 s]\n>   Range (min … max):   100.639 s … 101.375 s    10 runs\n>  \n> Summary\n>   'old' ran\n>     1.22 ± 0.01 times faster than 'new (small threshold)'\n>     1.23 ± 0.01 times faster than 'new'\n\n...adds enough sugar around \"hyperfine\" itself to do this as e.g. (the\n\"-s\" is a feature I submitted to hyperfine itself, it's not in a release\nyet[1], but in this case you could also use \"-p\"):\n\n    git hyperfine -L rev v2.20.0,origin/master \\\n        -s 'if ! test -d redis.git; then git clone --bare --filter=blob:none https://github.com/redis/redis; fi && make' \\\n        -p 'rm -rf dest.git; git init --bare dest.git' \\\n        './git -C dest.git unpack-objects <$(echo redis.git/objects/pack/*.pack)'\n\nThe sugar being that for each named \"rev\" parameter it'll set up \"git\nworktree\" for you, so under the hood each of those is chdir-ing to the\nrespective revision of:\n    \n    $ git worktree list\n    [...]\n    /run/user/1001/git-hyperfine/origin/master  abe6bb39053 (detached HEAD)\n    /run/user/1001/git-hyperfine/v2.33.0        225bc32a989 (detached HEAD)\n\nThat they're named revisions and not git-rev-parse'd is intentional,\nsince you'll benefit from faster incremental \"make\" (even if using\n\"ccache\"). I'm typically benchmarking HEAD~1,HEAD~0.\n\nThe output will then use those \"rev\" parameters, and be e.g.:\n    \n    Benchmark 1: ./git -C dest.git unpack-objects <$(echo redis.git/objects/pack/*.pack)' in 'v2.20.0\n      Time (mean ± σ):      6.678 s ±  0.046 s    [User: 4.525 s, System: 2.117 s]\n      Range (min … max):    6.619 s …  6.765 s    10 runs\n     \n    Benchmark 2: ./git -C dest.git unpack-objects <$(echo redis.git/objects/pack/*.pack)' in 'origin/master\n      Time (mean ± σ):      6.756 s ±  0.074 s    [User: 4.586 s, System: 2.134 s]\n      Range (min … max):    6.691 s …  6.941 s    10 runs\n     \n    Summary\n      './git -C dest.git unpack-objects <$(echo redis.git/objects/pack/*.pack)' in 'v2.20.0' ran\n        1.01 ± 0.01 times faster than './git -C dest.git unpack-objects <$(echo redis.git/objects/pack/*.pack)' in 'origin/master'\n\nI think if you're routinely benchmarking N different git versions you'll\nfind it handy, it also has configurable hook support (using git config),\nso e.g. it's easy to copy your config.mak in-place in the\nworktrees. E.g. my config is:\n\n    $ git -P config --get-regexp '^hyperfine'\n    hyperfine.run-dir $XDG_RUNTIME_DIR/git-hyperfine\n    hyperfine.xargs-options -r\n    hyperfine.hook.setup ~/g/git.meta/config.mak.sh\n\nIt's hosted at https://github.com/avar/git-hyperfine/ and\nhttps://gitlab.com/avar/git-hyperfine/; It's implemented in (portable)\nPOSIX shell script.\n\nThere's surely some bugs in it, one known one is that unlike hyperfine\nit doesn't accept there being spaces in the parameters to -L, because\nI'm screwing up some quoting-within-quoting in the (shellscript)\nimplementation (suggestions for that particular one most welcome).\n\nI hacked it up after this suggestion from Jeff King[2] of moving t/perf\nover to it.\n\nI haven't done any of that legwork, but I think a wrapper like\n\"git-hyperfine\" that prepares worktrees for the N revisions we're\nbenchmarking is a good direction to go in.\n\nWe don't use git-worktrees in t/perf, but probably could for most/all\ntests. In any case it would be easy to have the script setup the revs to\nbe benchmarked in some hookable custom manner to have it do exactly what\nt/perf/run is doing now.\n\n1. https://github.com/sharkdp/hyperfine/commit/017d55a\n2. https://lore.kernel.org/git/YV+zFqi4VmBVJYex@coredump.intra.peff.net/\n"},{"id":"442856","messageId":"CAO0brD2bGXwAJVrPSybbVNeSmQ1S85a_ykmceVKrg=pE3MfsnA@mail.gmail.com","threadId":"56672","inReplyTo":"446c3677-140f-3033-138f-1ef9b1f546a5@gmail.com","subject":"Re: [PATCH v3 5/5] unpack-objects: unpack_non_delta_entry() read data in a stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-02T07:33:04Z","receivedAt":"2021-12-02T07:33:18Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Wed, Dec 1, 2021 at 2:38 AM Derrick Stolee <stolee@gmail.com> wrote:\n>\n> I used a pack-file from an internal repo. It happened to be using\n> partial clone, so here is a repro with the git/git repository\n> after cloning this way:\n>\n> $ git clone --no-checkout --filter=blob:none https://github.com/git/git\n>\n> (copy the large .pack from git/.git/objects/pack/ to big.pack)\n>\n> $ hyperfine \\\n>         --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n>         -n 'old' '~/_git/git-upstream/git -C dest.git unpack-objects <big.pack' \\\n>         -n 'new' '~/_git/git/git -C dest.git unpack-objects <big.pack' \\\n>         -n 'new (small threshold)' '~/_git/git/git -c core.bigfilethreshold=64k -C dest.git unpack-objects <big.pack'\n>\n> Benchmark 1: old\n>   Time (mean ± σ):     82.748 s ±  0.445 s    [User: 50.512 s, System: 32.049 s]\n>   Range (min … max):   82.042 s … 83.587 s    10 runs\n>\n> Benchmark 2: new\n>   Time (mean ± σ):     101.644 s ±  0.524 s    [User: 67.470 s, System: 34.047 s]\n>   Range (min … max):   100.866 s … 102.633 s    10 runs\n>\n> Benchmark 3: new (small threshold)\n>   Time (mean ± σ):     101.093 s ±  0.269 s    [User: 67.404 s, System: 33.559 s]\n>   Range (min … max):   100.639 s … 101.375 s    10 runs\n>\n> Summary\n>   'old' ran\n>     1.22 ± 0.01 times faster than 'new (small threshold)'\n>     1.23 ± 0.01 times faster than 'new'\n>\n> I'm also able to repro this with a smaller repo (microsoft/scalar)\n> so the tests complete much faster:\n>\n> $ hyperfine \\\n>         --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n>         -n 'old' '~/_git/git-upstream/git -C dest.git unpack-objects <small.pack' \\\n>         -n 'new' '~/_git/git/git -C dest.git unpack-objects <small.pack' \\\n>         -n 'new (small threshold)' '~/_git/git/git -c core.bigfilethreshold=64k -C dest.git unpack-objects <small.pack'\n>\n> Benchmark 1: old\n>   Time (mean ± σ):      3.295 s ±  0.023 s    [User: 1.063 s, System: 2.228 s]\n>   Range (min … max):    3.269 s …  3.351 s    10 runs\n>\n> Benchmark 2: new\n>   Time (mean ± σ):      3.592 s ±  0.105 s    [User: 1.261 s, System: 2.328 s]\n>   Range (min … max):    3.378 s …  3.679 s    10 runs\n>\n> Benchmark 3: new (small threshold)\n>   Time (mean ± σ):      3.584 s ±  0.144 s    [User: 1.241 s, System: 2.339 s]\n>   Range (min … max):    3.359 s …  3.747 s    10 runs\n>\n> Summary\n>   'old' ran\n>     1.09 ± 0.04 times faster than 'new (small threshold)'\n>     1.09 ± 0.03 times faster than 'new'\n>\n> It's not the same relative overhead, but still significant.\n>\n> These pack-files contain (mostly) small objects, no large blobs.\n> I know that's not the target of your efforts, but it would be\n> good to avoid a regression here.\n>\n> Thanks,\n> -Stolee\n\nWith your help, I did catch this performance problem, which was\nintroduced in this patch:\nhttps://lore.kernel.org/git/20211122033220.32883-4-chiyutianyi@gmail.com/\n\nThis patch changes the original data reading ino to stream reading, but\nits problem is that even for the original reading of the whole object data,\nit still generates an additional git_deflate() and subsequent transfer.\n\nI will fix it in a follow-up patch.\n\nThanks,\n-Han Xin\n"},{"id":"442865","messageId":"993f83e8-cb53-cad2-8457-e8e94ac56ca2@gmail.com","threadId":"56672","inReplyTo":"CAO0brD2bGXwAJVrPSybbVNeSmQ1S85a_ykmceVKrg=pE3MfsnA@mail.gmail.com","subject":"Re: [PATCH v3 5/5] unpack-objects: unpack_non_delta_entry() read data in a stream","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-12-02T13:53:12Z","receivedAt":"2021-12-02T13:53:19Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 12/2/2021 2:33 AM, Han Xin wrote:\n> On Wed, Dec 1, 2021 at 2:38 AM Derrick Stolee <stolee@gmail.com> wrote:\n>> These pack-files contain (mostly) small objects, no large blobs.\n>> I know that's not the target of your efforts, but it would be\n>> good to avoid a regression here.\n>>\n>> Thanks,\n>> -Stolee\n> \n> With your help, I did catch this performance problem, which was\n> introduced in this patch:\n> https://lore.kernel.org/git/20211122033220.32883-4-chiyutianyi@gmail.com/\n> \n> This patch changes the original data reading ino to stream reading, but\n> its problem is that even for the original reading of the whole object data,\n> it still generates an additional git_deflate() and subsequent transfer.\n\nI'm glad you found it!\n\n> I will fix it in a follow-up patch.\n\nLooking forward to it.\n\nThanks,\n-Stolee\n\n"},{"id":"442952","messageId":"20211203093530.93589-1-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211122033220.32883-1-chiyutianyi@gmail.com","subject":"[PATCH v4 0/5] unpack large objects in stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-03T09:35:25Z","receivedAt":"2021-12-03T09:36:07Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nChanges since v3:\n* Add \"size\" to \"struct input_stream\" which used by following commits.\n\n* Increase the buffer size of \"struct input_zstream_data\" from 4096 to\n  8192, which is consistent with the \"fixed_buf\" in the \"index-pack.c\".\n\n* Refactor \"read stream in a loop in write_loose_object()\" which\n  introduced a performance problem reported by Derrick Stolee[1].\n\n* Rewrite benchmarks in \"unpack-objects: unpack_non_delta_entry() read\n  data in a stream\" with sugguestions by Derrick Stolee[1] and\n  Ævar Arnfjörð Bjarmason[2]. \n  Now use \"scalar.git\" to benchmark, which contains more than 28000\n  objects and 96 objects larger than 16kB.\n\n1. https://lore.kernel.org/git/8ff89e50-1b80-7932-f0e2-af401ee04bb1@gmail.com/\n2. https://lore.kernel.org/git/211201.86r1aw9gbd.gmgdl@evledraar.gmail.com/\n\nHan Xin (5):\n  object-file: refactor write_loose_object() to read buffer from stream\n  object-file.c: handle undetermined oid in write_loose_object()\n  object-file.c: read stream in a loop in write_loose_object()\n  unpack-objects.c: add dry_run mode for get_data()\n  unpack-objects: unpack_non_delta_entry() read data in a stream\n\n builtin/unpack-objects.c            |  93 +++++++++++++++++++++++--\n object-file.c                       | 102 ++++++++++++++++++++++++----\n object-store.h                      |  10 +++\n t/t5590-unpack-non-delta-objects.sh |  76 +++++++++++++++++++++\n 4 files changed, 262 insertions(+), 19 deletions(-)\n create mode 100755 t/t5590-unpack-non-delta-objects.sh\n\nRange-diff against v3:\n1:  8640b04f6d ! 1:  af707ef304 object-file: refactor write_loose_object() to read buffer from stream\n    @@ object-file.c: int write_object_file_flags(const void *buf, unsigned long len,\n     +\t\t\t.buf = buf,\n     +\t\t\t.len = len,\n     +\t\t},\n    ++\t\t.size = len,\n     +\t};\n      \n      \t/* Normally if we have it in the pack then we do not bother writing\n    @@ object-file.c: int hash_object_file_literally(const void *buf, unsigned long len\n     +\t\t\t.buf = buf,\n     +\t\t\t.len = len,\n     +\t\t},\n    ++\t\t.size = len,\n     +\t};\n      \n      \t/* type string, SP, %lu of the length plus NUL must fit this */\n    @@ object-file.c: int force_object_loose(const struct object_id *oid, time_t mtime)\n      \tif (has_loose_object(oid))\n      \t\treturn 0;\n      \tbuf = read_object(the_repository, oid, &type, &len);\n    ++\tin_stream.size = len;\n      \tif (!buf)\n      \t\treturn error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n     +\tdata.buf = buf;\n    @@ object-store.h: struct object_directory {\n     +struct input_stream {\n     +\tconst void *(*read)(struct input_stream *, unsigned long *len);\n     +\tvoid *data;\n    ++\tsize_t size;\n     +};\n     +\n      KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n2:  d4a2caf2bd = 2:  321ad90d8e object-file.c: handle undetermined oid in write_loose_object()\n3:  2575900449 ! 3:  1992ac39af object-file.c: read stream in a loop in write_loose_object()\n    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *\n     -\t\tret = git_deflate(&stream, Z_FINISH);\n     +\t\tif (!stream.avail_in) {\n     +\t\t\tbuf = in_stream->read(in_stream, &stream.avail_in);\n    -+\t\t\tif (buf) {\n    -+\t\t\t\tstream.next_in = (void *)buf;\n    -+\t\t\t\tin0 = (unsigned char *)buf;\n    -+\t\t\t} else {\n    ++\t\t\tstream.next_in = (void *)buf;\n    ++\t\t\tin0 = (unsigned char *)buf;\n    ++\t\t\t/* All data has been read. */\n    ++\t\t\tif (in_stream->size + hdrlen == stream.total_in + stream.avail_in)\n     +\t\t\t\tflush = Z_FINISH;\n    -+\t\t\t}\n     +\t\t}\n     +\t\tret = git_deflate(&stream, flush);\n      \t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n      \t\tif (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n      \t\t\tdie(_(\"unable to write loose object file\"));\n    + \t\tstream.next_out = compressed;\n    + \t\tstream.avail_out = sizeof(compressed);\n    +-\t} while (ret == Z_OK);\n    ++\t} while (ret == Z_OK || ret == Z_BUF_ERROR);\n    + \n    + \tif (ret != Z_STREAM_END)\n    + \t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n4:  ca93ecc780 = 4:  c41eb06533 unpack-objects.c: add dry_run mode for get_data()\n5:  39a072ee2a ! 5:  9427775bdc unpack-objects: unpack_non_delta_entry() read data in a stream\n    @@ Commit message\n         larger than the \"big_file_threshold\" in zstream. See the following\n         benchmarks:\n     \n    -        $ hyperfine \\\n    -        --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n    -        'git -C dest.git unpack-objects <binary_320M.pack'\n    -        Benchmark 1: git -C dest.git unpack-objects <binary_320M.pack\n    -          Time (mean ± σ):     10.029 s ±  0.270 s    [User: 8.265 s, System: 1.522 s]\n    -          Range (min … max):    9.786 s … 10.603 s    10 runs\n    +        hyperfine \\\n    +          --setup \\\n    +          'if ! test -d scalar.git; then git clone --bare https://github.com/microsoft/scalar.git; cp scalar.git/objects/pack/*.pack small.pack; fi' \\\n    +          --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n    +          -n 'old' 'git -C dest.git unpack-objects <small.pack' \\\n    +          -n 'new' 'new/git -C dest.git unpack-objects <small.pack' \\\n    +          -n 'new (small threshold)' \\\n    +          'new/git -c core.bigfilethreshold=16k -C dest.git unpack-objects <small.pack'\n    +        Benchmark 1: old\n    +          Time (mean ± σ):      6.075 s ±  0.069 s    [User: 5.047 s, System: 0.991 s]\n    +          Range (min … max):    6.018 s …  6.189 s    10 runs\n     \n    -        $ hyperfine \\\n    -        --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n    -        'git -c core.bigFileThreshold=2m -C dest.git unpack-objects <binary_320M.pack'\n    -        Benchmark 1: git -c core.bigFileThreshold=2m -C dest.git unpack-objects <binary_320M.pack\n    -          Time (mean ± σ):     10.859 s ±  0.774 s    [User: 8.813 s, System: 1.898 s]\n    -          Range (min … max):    9.884 s … 12.192 s    10 runs\n    +        Benchmark 2: new\n    +          Time (mean ± σ):      6.090 s ±  0.033 s    [User: 5.075 s, System: 0.976 s]\n    +          Range (min … max):    6.030 s …  6.142 s    10 runs\n     \n    -        $ hyperfine \\\n    -        --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n    -        'git -C dest.git unpack-objects <binary_96M.pack'\n    -        Benchmark 1: git -C dest.git unpack-objects <binary_96M.pack\n    -          Time (mean ± σ):      2.678 s ±  0.037 s    [User: 2.205 s, System: 0.450 s]\n    -          Range (min … max):    2.639 s …  2.743 s    10 runs\n    +        Benchmark 3: new (small threshold)\n    +          Time (mean ± σ):      6.755 s ±  0.029 s    [User: 5.150 s, System: 1.560 s]\n    +          Range (min … max):    6.711 s …  6.809 s    10 runs\n     \n    -        $ hyperfine \\\n    -        --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n    -        'git -c core.bigFileThreshold=2m -C dest.git unpack-objects <binary_96M.pack'\n    -        Benchmark 1: git -c core.bigFileThreshold=2m -C dest.git unpack-objects <binary_96M.pack\n    -          Time (mean ± σ):      2.819 s ±  0.124 s    [User: 2.216 s, System: 0.564 s]\n    -          Range (min … max):    2.679 s …  3.125 s    10 runs\n    +        Summary\n    +          'old' ran\n    +            1.00 ± 0.01 times faster than 'new'\n    +            1.11 ± 0.01 times faster than 'new (small threshold)'\n     \n    +    Helped-by: Derrick Stolee <stolee@gmail.com>\n         Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n         Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n     \n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n      \n     +struct input_zstream_data {\n     +\tgit_zstream *zstream;\n    -+\tunsigned char buf[4096];\n    ++\tunsigned char buf[8192];\n     +\tint status;\n     +};\n     +\n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n     +\tstruct input_stream in_stream = {\n     +\t\t.read = feed_input_zstream,\n     +\t\t.data = &data,\n    ++\t\t.size = size,\n     +\t};\n     +\tstruct object_id *oid = &obj_list[nr].oid;\n     +\tint ret;\n-- \n2.34.0\n\n"},{"id":"442953","messageId":"20211203093530.93589-2-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211122033220.32883-1-chiyutianyi@gmail.com","subject":"[PATCH v4 1/5] object-file: refactor write_loose_object() to read buffer from stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-03T09:35:26Z","receivedAt":"2021-12-03T09:36:11Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWe used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\nentire contents of a blob object, no matter how big it is. This\nimplementation may consume all the memory and cause OOM.\n\nThis can be improved by feeding data to \"write_loose_object()\" in a\nstream. The input stream is implemented as an interface. In the first\nstep, we make a simple implementation, feeding the entire buffer in the\n\"stream\" to \"write_loose_object()\" as a refactor.\n\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c  | 53 ++++++++++++++++++++++++++++++++++++++++++++++----\n object-store.h |  6 ++++++\n 2 files changed, 55 insertions(+), 4 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex eb972cdccd..82656f7428 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1860,8 +1860,26 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n \treturn fd;\n }\n \n+struct simple_input_stream_data {\n+\tconst void *buf;\n+\tunsigned long len;\n+};\n+\n+static const void *feed_simple_input_stream(struct input_stream *in_stream, unsigned long *len)\n+{\n+\tstruct simple_input_stream_data *data = in_stream->data;\n+\n+\tif (data->len == 0) {\n+\t\t*len = 0;\n+\t\treturn NULL;\n+\t}\n+\t*len = data->len;\n+\tdata->len = 0;\n+\treturn data->buf;\n+}\n+\n static int write_loose_object(const struct object_id *oid, char *hdr,\n-\t\t\t      int hdrlen, const void *buf, unsigned long len,\n+\t\t\t      int hdrlen, struct input_stream *in_stream,\n \t\t\t      time_t mtime, unsigned flags)\n {\n \tint fd, ret;\n@@ -1871,6 +1889,8 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tstruct object_id parano_oid;\n \tstatic struct strbuf tmp_file = STRBUF_INIT;\n \tstatic struct strbuf filename = STRBUF_INIT;\n+\tconst void *buf;\n+\tunsigned long len;\n \n \tloose_object_path(the_repository, &filename, oid);\n \n@@ -1898,6 +1918,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n \n \t/* Then the data itself.. */\n+\tbuf = in_stream->read(in_stream, &len);\n \tstream.next_in = (void *)buf;\n \tstream.avail_in = len;\n \tdo {\n@@ -1960,6 +1981,14 @@ int write_object_file_flags(const void *buf, unsigned long len,\n {\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen = sizeof(hdr);\n+\tstruct input_stream in_stream = {\n+\t\t.read = feed_simple_input_stream,\n+\t\t.data = (void *)&(struct simple_input_stream_data) {\n+\t\t\t.buf = buf,\n+\t\t\t.len = len,\n+\t\t},\n+\t\t.size = len,\n+\t};\n \n \t/* Normally if we have it in the pack then we do not bother writing\n \t * it out into .git/objects/??/?{38} file.\n@@ -1968,7 +1997,7 @@ int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t\t  &hdrlen);\n \tif (freshen_packed_object(oid) || freshen_loose_object(oid))\n \t\treturn 0;\n-\treturn write_loose_object(oid, hdr, hdrlen, buf, len, 0, flags);\n+\treturn write_loose_object(oid, hdr, hdrlen, &in_stream, 0, flags);\n }\n \n int hash_object_file_literally(const void *buf, unsigned long len,\n@@ -1977,6 +2006,14 @@ int hash_object_file_literally(const void *buf, unsigned long len,\n {\n \tchar *header;\n \tint hdrlen, status = 0;\n+\tstruct input_stream in_stream = {\n+\t\t.read = feed_simple_input_stream,\n+\t\t.data = (void *)&(struct simple_input_stream_data) {\n+\t\t\t.buf = buf,\n+\t\t\t.len = len,\n+\t\t},\n+\t\t.size = len,\n+\t};\n \n \t/* type string, SP, %lu of the length plus NUL must fit this */\n \thdrlen = strlen(type) + MAX_HEADER_LEN;\n@@ -1988,7 +2025,7 @@ int hash_object_file_literally(const void *buf, unsigned long len,\n \t\tgoto cleanup;\n \tif (freshen_packed_object(oid) || freshen_loose_object(oid))\n \t\tgoto cleanup;\n-\tstatus = write_loose_object(oid, header, hdrlen, buf, len, 0, 0);\n+\tstatus = write_loose_object(oid, header, hdrlen, &in_stream, 0, 0);\n \n cleanup:\n \tfree(header);\n@@ -2003,14 +2040,22 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen;\n \tint ret;\n+\tstruct simple_input_stream_data data;\n+\tstruct input_stream in_stream = {\n+\t\t.read = feed_simple_input_stream,\n+\t\t.data = &data,\n+\t};\n \n \tif (has_loose_object(oid))\n \t\treturn 0;\n \tbuf = read_object(the_repository, oid, &type, &len);\n+\tin_stream.size = len;\n \tif (!buf)\n \t\treturn error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n+\tdata.buf = buf;\n+\tdata.len = len;\n \thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n-\tret = write_loose_object(oid, hdr, hdrlen, buf, len, mtime, 0);\n+\tret = write_loose_object(oid, hdr, hdrlen, &in_stream, mtime, 0);\n \tfree(buf);\n \n \treturn ret;\ndiff --git a/object-store.h b/object-store.h\nindex 952efb6a4b..a84d891d60 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -34,6 +34,12 @@ struct object_directory {\n \tchar *path;\n };\n \n+struct input_stream {\n+\tconst void *(*read)(struct input_stream *, unsigned long *len);\n+\tvoid *data;\n+\tsize_t size;\n+};\n+\n KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n \tstruct object_directory *, 1, fspathhash, fspatheq)\n \n-- \n2.34.0\n\n"},{"id":"442954","messageId":"20211203093530.93589-3-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211122033220.32883-1-chiyutianyi@gmail.com","subject":"[PATCH v4 2/5] object-file.c: handle undetermined oid in write_loose_object()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-03T09:35:27Z","receivedAt":"2021-12-03T09:36:13Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWhen streaming a large blob object to \"write_loose_object()\", we have no\nchance to run \"write_object_file_prepare()\" to calculate the oid in\nadvance. So we need to handle undetermined oid in function\n\"write_loose_object()\".\n\nIn the original implementation, we know the oid and we can write the\ntemporary file in the same directory as the final object, but for an\nobject with an undetermined oid, we don't know the exact directory for\nthe object, so we have to save the temporary file in \".git/objects/\"\ndirectory instead.\n\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c | 30 ++++++++++++++++++++++++++++--\n 1 file changed, 28 insertions(+), 2 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 82656f7428..1c41587bfb 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1892,7 +1892,14 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tconst void *buf;\n \tunsigned long len;\n \n-\tloose_object_path(the_repository, &filename, oid);\n+\tif (is_null_oid(oid)) {\n+\t\t/* When oid is not determined, save tmp file to odb path. */\n+\t\tstrbuf_reset(&filename);\n+\t\tstrbuf_addstr(&filename, the_repository->objects->odb->path);\n+\t\tstrbuf_addch(&filename, '/');\n+\t} else {\n+\t\tloose_object_path(the_repository, &filename, oid);\n+\t}\n \n \tfd = create_tmpfile(&tmp_file, filename.buf);\n \tif (fd < 0) {\n@@ -1939,12 +1946,31 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n \t\t    ret);\n \tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n-\tif (!oideq(oid, &parano_oid))\n+\tif (!is_null_oid(oid) && !oideq(oid, &parano_oid))\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n \n \tclose_loose_object(fd);\n \n+\tif (is_null_oid(oid)) {\n+\t\tint dirlen;\n+\n+\t\toidcpy((struct object_id *)oid, &parano_oid);\n+\t\tloose_object_path(the_repository, &filename, oid);\n+\n+\t\t/* We finally know the object path, and create the missing dir. */\n+\t\tdirlen = directory_size(filename.buf);\n+\t\tif (dirlen) {\n+\t\t\tstruct strbuf dir = STRBUF_INIT;\n+\t\t\tstrbuf_add(&dir, filename.buf, dirlen - 1);\n+\t\t\tif (mkdir(dir.buf, 0777) && errno != EEXIST)\n+\t\t\t\treturn -1;\n+\t\t\tif (adjust_shared_perm(dir.buf))\n+\t\t\t\treturn -1;\n+\t\t\tstrbuf_release(&dir);\n+\t\t}\n+\t}\n+\n \tif (mtime) {\n \t\tstruct utimbuf utb;\n \t\tutb.actime = mtime;\n-- \n2.34.0\n\n"},{"id":"442955","messageId":"20211203093530.93589-4-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211122033220.32883-1-chiyutianyi@gmail.com","subject":"[PATCH v4 3/5] object-file.c: read stream in a loop in write_loose_object()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-03T09:35:28Z","receivedAt":"2021-12-03T09:36:15Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nIn order to prepare the stream version of \"write_loose_object()\", read\nthe input stream in a loop in \"write_loose_object()\", so that we can\nfeed the contents of large blob object to \"write_loose_object()\" using\na small fixed buffer.\n\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c | 17 +++++++++++------\n 1 file changed, 11 insertions(+), 6 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 1c41587bfb..fa54e39c2c 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1890,7 +1890,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tstatic struct strbuf tmp_file = STRBUF_INIT;\n \tstatic struct strbuf filename = STRBUF_INIT;\n \tconst void *buf;\n-\tunsigned long len;\n+\tint flush = 0;\n \n \tif (is_null_oid(oid)) {\n \t\t/* When oid is not determined, save tmp file to odb path. */\n@@ -1925,18 +1925,23 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n \n \t/* Then the data itself.. */\n-\tbuf = in_stream->read(in_stream, &len);\n-\tstream.next_in = (void *)buf;\n-\tstream.avail_in = len;\n \tdo {\n \t\tunsigned char *in0 = stream.next_in;\n-\t\tret = git_deflate(&stream, Z_FINISH);\n+\t\tif (!stream.avail_in) {\n+\t\t\tbuf = in_stream->read(in_stream, &stream.avail_in);\n+\t\t\tstream.next_in = (void *)buf;\n+\t\t\tin0 = (unsigned char *)buf;\n+\t\t\t/* All data has been read. */\n+\t\t\tif (in_stream->size + hdrlen == stream.total_in + stream.avail_in)\n+\t\t\t\tflush = Z_FINISH;\n+\t\t}\n+\t\tret = git_deflate(&stream, flush);\n \t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n \t\tif (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n \t\t\tdie(_(\"unable to write loose object file\"));\n \t\tstream.next_out = compressed;\n \t\tstream.avail_out = sizeof(compressed);\n-\t} while (ret == Z_OK);\n+\t} while (ret == Z_OK || ret == Z_BUF_ERROR);\n \n \tif (ret != Z_STREAM_END)\n \t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n-- \n2.34.0\n\n"},{"id":"442956","messageId":"20211203093530.93589-5-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211122033220.32883-1-chiyutianyi@gmail.com","subject":"[PATCH v4 4/5] unpack-objects.c: add dry_run mode for get_data()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-03T09:35:29Z","receivedAt":"2021-12-03T09:36:18Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nIn dry_run mode, \"get_data()\" is used to verify the inflation of data,\nand the returned buffer will not be used at all and will be freed\nimmediately. Even in dry_run mode, it is dangerous to allocate a\nfull-size buffer for a large blob object. Therefore, only allocate a\nlow memory footprint when calling \"get_data()\" in dry_run mode.\n\nSuggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n builtin/unpack-objects.c | 18 ++++++++++++------\n 1 file changed, 12 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 4a9466295b..8d68acd662 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -96,15 +96,16 @@ static void use(int bytes)\n \tdisplay_throughput(progress, consumed_bytes);\n }\n \n-static void *get_data(unsigned long size)\n+static void *get_data(unsigned long size, int dry_run)\n {\n \tgit_zstream stream;\n-\tvoid *buf = xmallocz(size);\n+\tunsigned long bufsize = dry_run ? 4096 : size;\n+\tvoid *buf = xmallocz(bufsize);\n \n \tmemset(&stream, 0, sizeof(stream));\n \n \tstream.next_out = buf;\n-\tstream.avail_out = size;\n+\tstream.avail_out = bufsize;\n \tstream.next_in = fill(1);\n \tstream.avail_in = len;\n \tgit_inflate_init(&stream);\n@@ -124,6 +125,11 @@ static void *get_data(unsigned long size)\n \t\t}\n \t\tstream.next_in = fill(1);\n \t\tstream.avail_in = len;\n+\t\tif (dry_run) {\n+\t\t\t/* reuse the buffer in dry_run mode */\n+\t\t\tstream.next_out = buf;\n+\t\t\tstream.avail_out = bufsize;\n+\t\t}\n \t}\n \tgit_inflate_end(&stream);\n \treturn buf;\n@@ -323,7 +329,7 @@ static void added_object(unsigned nr, enum object_type type,\n static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n \t\t\t\t   unsigned nr)\n {\n-\tvoid *buf = get_data(size);\n+\tvoid *buf = get_data(size, dry_run);\n \n \tif (!dry_run && buf)\n \t\twrite_object(nr, type, buf, size);\n@@ -357,7 +363,7 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \tif (type == OBJ_REF_DELTA) {\n \t\toidread(&base_oid, fill(the_hash_algo->rawsz));\n \t\tuse(the_hash_algo->rawsz);\n-\t\tdelta_data = get_data(delta_size);\n+\t\tdelta_data = get_data(delta_size, dry_run);\n \t\tif (dry_run || !delta_data) {\n \t\t\tfree(delta_data);\n \t\t\treturn;\n@@ -396,7 +402,7 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\tif (base_offset <= 0 || base_offset >= obj_list[nr].offset)\n \t\t\tdie(\"offset value out of bound for delta base object\");\n \n-\t\tdelta_data = get_data(delta_size);\n+\t\tdelta_data = get_data(delta_size, dry_run);\n \t\tif (dry_run || !delta_data) {\n \t\t\tfree(delta_data);\n \t\t\treturn;\n-- \n2.34.0\n\n"},{"id":"442957","messageId":"20211203093530.93589-6-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211122033220.32883-1-chiyutianyi@gmail.com","subject":"[PATCH v4 5/5] unpack-objects: unpack_non_delta_entry() read data in a stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-03T09:35:30Z","receivedAt":"2021-12-03T09:36:22Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWe used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\nentire contents of a blob object, no matter how big it is. This\nimplementation may consume all the memory and cause OOM.\n\nBy implementing a zstream version of input_stream interface, we can use\na small fixed buffer for \"unpack_non_delta_entry()\".\n\nHowever, unpack non-delta objects from a stream instead of from an entrie\nbuffer will have 10% performance penalty. Therefore, only unpack object\nlarger than the \"big_file_threshold\" in zstream. See the following\nbenchmarks:\n\n    hyperfine \\\n      --setup \\\n      'if ! test -d scalar.git; then git clone --bare https://github.com/microsoft/scalar.git; cp scalar.git/objects/pack/*.pack small.pack; fi' \\\n      --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n      -n 'old' 'git -C dest.git unpack-objects <small.pack' \\\n      -n 'new' 'new/git -C dest.git unpack-objects <small.pack' \\\n      -n 'new (small threshold)' \\\n      'new/git -c core.bigfilethreshold=16k -C dest.git unpack-objects <small.pack'\n    Benchmark 1: old\n      Time (mean ± σ):      6.075 s ±  0.069 s    [User: 5.047 s, System: 0.991 s]\n      Range (min … max):    6.018 s …  6.189 s    10 runs\n\n    Benchmark 2: new\n      Time (mean ± σ):      6.090 s ±  0.033 s    [User: 5.075 s, System: 0.976 s]\n      Range (min … max):    6.030 s …  6.142 s    10 runs\n\n    Benchmark 3: new (small threshold)\n      Time (mean ± σ):      6.755 s ±  0.029 s    [User: 5.150 s, System: 1.560 s]\n      Range (min … max):    6.711 s …  6.809 s    10 runs\n\n    Summary\n      'old' ran\n        1.00 ± 0.01 times faster than 'new'\n        1.11 ± 0.01 times faster than 'new (small threshold)'\n\nHelped-by: Derrick Stolee <stolee@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n builtin/unpack-objects.c            | 77 ++++++++++++++++++++++++++++-\n object-file.c                       |  6 +--\n object-store.h                      |  4 ++\n t/t5590-unpack-non-delta-objects.sh | 76 ++++++++++++++++++++++++++++\n 4 files changed, 159 insertions(+), 4 deletions(-)\n create mode 100755 t/t5590-unpack-non-delta-objects.sh\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 8d68acd662..bedc494e2d 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -326,11 +326,86 @@ static void added_object(unsigned nr, enum object_type type,\n \t}\n }\n \n+struct input_zstream_data {\n+\tgit_zstream *zstream;\n+\tunsigned char buf[8192];\n+\tint status;\n+};\n+\n+static const void *feed_input_zstream(struct input_stream *in_stream, unsigned long *readlen)\n+{\n+\tstruct input_zstream_data *data = in_stream->data;\n+\tgit_zstream *zstream = data->zstream;\n+\tvoid *in = fill(1);\n+\n+\tif (!len || data->status == Z_STREAM_END) {\n+\t\t*readlen = 0;\n+\t\treturn NULL;\n+\t}\n+\n+\tzstream->next_out = data->buf;\n+\tzstream->avail_out = sizeof(data->buf);\n+\tzstream->next_in = in;\n+\tzstream->avail_in = len;\n+\n+\tdata->status = git_inflate(zstream, 0);\n+\tuse(len - zstream->avail_in);\n+\t*readlen = sizeof(data->buf) - zstream->avail_out;\n+\n+\treturn data->buf;\n+}\n+\n+static void write_stream_blob(unsigned nr, unsigned long size)\n+{\n+\tchar hdr[32];\n+\tint hdrlen;\n+\tgit_zstream zstream;\n+\tstruct input_zstream_data data;\n+\tstruct input_stream in_stream = {\n+\t\t.read = feed_input_zstream,\n+\t\t.data = &data,\n+\t\t.size = size,\n+\t};\n+\tstruct object_id *oid = &obj_list[nr].oid;\n+\tint ret;\n+\n+\tmemset(&zstream, 0, sizeof(zstream));\n+\tmemset(&data, 0, sizeof(data));\n+\tdata.zstream = &zstream;\n+\tgit_inflate_init(&zstream);\n+\n+\t/* Generate the header */\n+\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX, type_name(OBJ_BLOB), (uintmax_t)size) + 1;\n+\n+\tif ((ret = write_loose_object(oid, hdr, hdrlen, &in_stream, 0, 0)))\n+\t\tdie(_(\"failed to write object in stream %d\"), ret);\n+\n+\tif (zstream.total_out != size || data.status != Z_STREAM_END)\n+\t\tdie(_(\"inflate returned %d\"), data.status);\n+\tgit_inflate_end(&zstream);\n+\n+\tif (strict && !dry_run) {\n+\t\tstruct blob *blob = lookup_blob(the_repository, oid);\n+\t\tif (blob)\n+\t\t\tblob->object.flags |= FLAG_WRITTEN;\n+\t\telse\n+\t\t\tdie(\"invalid blob object from stream\");\n+\t}\n+\tobj_list[nr].obj = NULL;\n+}\n+\n static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n \t\t\t\t   unsigned nr)\n {\n-\tvoid *buf = get_data(size, dry_run);\n+\tvoid *buf;\n+\n+\t/* Write large blob in stream without allocating full buffer. */\n+\tif (!dry_run && type == OBJ_BLOB && size > big_file_threshold) {\n+\t\twrite_stream_blob(nr, size);\n+\t\treturn;\n+\t}\n \n+\tbuf = get_data(size, dry_run);\n \tif (!dry_run && buf)\n \t\twrite_object(nr, type, buf, size);\n \telse\ndiff --git a/object-file.c b/object-file.c\nindex fa54e39c2c..71d510614b 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1878,9 +1878,9 @@ static const void *feed_simple_input_stream(struct input_stream *in_stream, unsi\n \treturn data->buf;\n }\n \n-static int write_loose_object(const struct object_id *oid, char *hdr,\n-\t\t\t      int hdrlen, struct input_stream *in_stream,\n-\t\t\t      time_t mtime, unsigned flags)\n+int write_loose_object(const struct object_id *oid, char *hdr,\n+\t\t       int hdrlen, struct input_stream *in_stream,\n+\t\t       time_t mtime, unsigned flags)\n {\n \tint fd, ret;\n \tunsigned char compressed[4096];\ndiff --git a/object-store.h b/object-store.h\nindex a84d891d60..ac5b11ec16 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -229,6 +229,10 @@ int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n \t\t     unsigned long len, const char *type,\n \t\t     struct object_id *oid);\n \n+int write_loose_object(const struct object_id *oid, char *hdr,\n+\t\t       int hdrlen, struct input_stream *in_stream,\n+\t\t       time_t mtime, unsigned flags);\n+\n int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t    const char *type, struct object_id *oid,\n \t\t\t    unsigned flags);\ndiff --git a/t/t5590-unpack-non-delta-objects.sh b/t/t5590-unpack-non-delta-objects.sh\nnew file mode 100755\nindex 0000000000..01d950d119\n--- /dev/null\n+++ b/t/t5590-unpack-non-delta-objects.sh\n@@ -0,0 +1,76 @@\n+#!/bin/sh\n+#\n+# Copyright (c) 2021 Han Xin\n+#\n+\n+test_description='Test unpack-objects when receive pack'\n+\n+GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n+export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n+\n+. ./test-lib.sh\n+\n+test_expect_success \"create commit with big blobs (1.5 MB)\" '\n+\ttest-tool genrandom foo 1500000 >big-blob &&\n+\ttest_commit --append foo big-blob &&\n+\ttest-tool genrandom bar 1500000 >big-blob &&\n+\ttest_commit --append bar big-blob &&\n+\t(\n+\t\tcd .git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >expect &&\n+\tPACK=$(echo main | git pack-objects --progress --revs test)\n+'\n+\n+test_expect_success 'setup GIT_ALLOC_LIMIT to 1MB' '\n+\tGIT_ALLOC_LIMIT=1m &&\n+\texport GIT_ALLOC_LIMIT\n+'\n+\n+test_expect_success 'prepare dest repository' '\n+\tgit init --bare dest.git &&\n+\tgit -C dest.git config core.bigFileThreshold 2m &&\n+\tgit -C dest.git config receive.unpacklimit 100\n+'\n+\n+test_expect_success 'fail to unpack-objects: cannot allocate' '\n+\ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n+\ttest_i18ngrep \"fatal: attempting to allocate\" err &&\n+\t(\n+\t\tcd dest.git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >actual &&\n+\t! test_cmp expect actual\n+'\n+\n+test_expect_success 'set a lower bigfile threshold' '\n+\tgit -C dest.git config core.bigFileThreshold 1m\n+'\n+\n+test_expect_success 'unpack big object in stream' '\n+\tgit -C dest.git unpack-objects <test-$PACK.pack &&\n+\tgit -C dest.git fsck &&\n+\t(\n+\t\tcd dest.git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'setup for unpack-objects dry-run test' '\n+\tgit init --bare unpack-test.git\n+'\n+\n+test_expect_success 'unpack-objects dry-run' '\n+\t(\n+\t\tcd unpack-test.git &&\n+\t\tgit unpack-objects -n <../test-$PACK.pack\n+\t) &&\n+\t(\n+\t\tcd unpack-test.git &&\n+\t\tfind objects/ -type f\n+\t) >actual &&\n+\ttest_must_be_empty actual\n+'\n+\n+test_done\n-- \n2.34.0\n\n"},{"id":"442973","messageId":"211203.86zgphsu5a.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"20211203093530.93589-6-chiyutianyi@gmail.com","subject":"Re: [PATCH v4 5/5] unpack-objects: unpack_non_delta_entry() read data in a stream","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-03T13:07:58Z","receivedAt":"2021-12-03T13:19:38Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, Dec 03 2021, Han Xin wrote:\n\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> We used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\n> entire contents of a blob object, no matter how big it is. This\n> implementation may consume all the memory and cause OOM.\n>\n> By implementing a zstream version of input_stream interface, we can use\n> a small fixed buffer for \"unpack_non_delta_entry()\".\n>\n> However, unpack non-delta objects from a stream instead of from an entrie\n> buffer will have 10% performance penalty. Therefore, only unpack object\n> larger than the \"big_file_threshold\" in zstream. See the following\n> benchmarks:\n>\n>     hyperfine \\\n>       --setup \\\n>       'if ! test -d scalar.git; then git clone --bare https://github.com/microsoft/scalar.git; cp scalar.git/objects/pack/*.pack small.pack; fi' \\\n>       --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n>       -n 'old' 'git -C dest.git unpack-objects <small.pack' \\\n>       -n 'new' 'new/git -C dest.git unpack-objects <small.pack' \\\n>       -n 'new (small threshold)' \\\n>       'new/git -c core.bigfilethreshold=16k -C dest.git unpack-objects <small.pack'\n>     Benchmark 1: old\n>       Time (mean ± σ):      6.075 s ±  0.069 s    [User: 5.047 s, System: 0.991 s]\n>       Range (min … max):    6.018 s …  6.189 s    10 runs\n>\n>     Benchmark 2: new\n>       Time (mean ± σ):      6.090 s ±  0.033 s    [User: 5.075 s, System: 0.976 s]\n>       Range (min … max):    6.030 s …  6.142 s    10 runs\n>\n>     Benchmark 3: new (small threshold)\n>       Time (mean ± σ):      6.755 s ±  0.029 s    [User: 5.150 s, System: 1.560 s]\n>       Range (min … max):    6.711 s …  6.809 s    10 runs\n>\n>     Summary\n>       'old' ran\n>         1.00 ± 0.01 times faster than 'new'\n>         1.11 ± 0.01 times faster than 'new (small threshold)'\n\nSo before we wrote used core.bigfilethreshold for two things (or more?):\nWhether we show a diff for it (we mark it \"binary\") and whether it's\nsplit into a loose object.\n\nNow it's three things, we've added a \"this is a threshold when we'll\nstream the object\" to that.\n\nMight it make sense to squash something like this in, so we can have our\ncake & eat it too?\n\nWith this I get, where HEAD~0 is this change:\n    \n    Summary\n      './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'HEAD~0' ran\n        1.00 ± 0.01 times faster than './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'HEAD~1'\n        1.00 ± 0.01 times faster than './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'origin/master'\n        1.01 ± 0.01 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'HEAD~0'\n        1.06 ± 0.14 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'origin/master'\n        1.20 ± 0.01 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'HEAD~1'\n\nI.e. it's 5% slower, not 20% (haven't looked into why), but we'll not\nstream out 16k..128MB objects (maybe the repo has even bigger ones?)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex c04f62a54a1..601b7a2418f 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -424,6 +424,17 @@ be delta compressed, but larger binary media files won't be.\n +\n Common unit suffixes of 'k', 'm', or 'g' are supported.\n \n+core.bigFileStreamingThreshold::\n+\tFiles larger than this will be streamed out to a temporary\n+\tobject file while being hashed, which will when be renamed\n+\tin-place to a loose object, particularly if the\n+\t`core.bigFileThreshold' setting dictates that they're always\n+\twritten out as loose objects.\n++\n+Default is 128 MiB on all platforms.\n++\n+Common unit suffixes of 'k', 'm', or 'g' are supported.\n+\n core.excludesFile::\n \tSpecifies the pathname to the file that contains patterns to\n \tdescribe paths that are not meant to be tracked, in addition\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex bedc494e2db..94ce275c807 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -400,7 +400,7 @@ static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n \tvoid *buf;\n \n \t/* Write large blob in stream without allocating full buffer. */\n-\tif (!dry_run && type == OBJ_BLOB && size > big_file_threshold) {\n+\tif (!dry_run && type == OBJ_BLOB && size > big_file_streaming_threshold) {\n \t\twrite_stream_blob(nr, size);\n \t\treturn;\n \t}\ndiff --git a/cache.h b/cache.h\nindex eba12487b99..4037c7fd849 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -964,6 +964,7 @@ extern size_t packed_git_window_size;\n extern size_t packed_git_limit;\n extern size_t delta_base_cache_limit;\n extern unsigned long big_file_threshold;\n+extern unsigned long big_file_streaming_threshold;\n extern unsigned long pack_size_limit_cfg;\n \n /*\ndiff --git a/config.c b/config.c\nindex c5873f3a706..7b122a142a8 100644\n--- a/config.c\n+++ b/config.c\n@@ -1408,6 +1408,11 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"core.bigfilestreamingthreshold\")) {\n+\t\tbig_file_streaming_threshold = git_config_ulong(var, value);\n+\t\treturn 0;\n+\t}\n+\n \tif (!strcmp(var, \"core.packedgitlimit\")) {\n \t\tpacked_git_limit = git_config_ulong(var, value);\n \t\treturn 0;\ndiff --git a/environment.c b/environment.c\nindex 9da7f3c1a19..4fcc3de7417 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -46,6 +46,7 @@ size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\n unsigned long big_file_threshold = 512 * 1024 * 1024;\n+unsigned long big_file_streaming_threshold = 128 * 1024 * 1024;\n int pager_use_color = 1;\n const char *editor_program;\n const char *askpass_program;\n"},{"id":"442974","messageId":"211203.86v905stru.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"20211203093530.93589-3-chiyutianyi@gmail.com","subject":"Re: [PATCH v4 2/5] object-file.c: handle undetermined oid in write_loose_object()","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-03T13:21:08Z","receivedAt":"2021-12-03T13:27:38Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, Dec 03 2021, Han Xin wrote:\n\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> When streaming a large blob object to \"write_loose_object()\", we have no\n> chance to run \"write_object_file_prepare()\" to calculate the oid in\n> advance. So we need to handle undetermined oid in function\n> \"write_loose_object()\".\n>\n> In the original implementation, we know the oid and we can write the\n> temporary file in the same directory as the final object, but for an\n> object with an undetermined oid, we don't know the exact directory for\n> the object, so we have to save the temporary file in \".git/objects/\"\n> directory instead.\n>\n> Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  object-file.c | 30 ++++++++++++++++++++++++++++--\n>  1 file changed, 28 insertions(+), 2 deletions(-)\n>\n> diff --git a/object-file.c b/object-file.c\n> index 82656f7428..1c41587bfb 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1892,7 +1892,14 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \tconst void *buf;\n>  \tunsigned long len;\n>  \n> -\tloose_object_path(the_repository, &filename, oid);\n> +\tif (is_null_oid(oid)) {\n> +\t\t/* When oid is not determined, save tmp file to odb path. */\n> +\t\tstrbuf_reset(&filename);\n> +\t\tstrbuf_addstr(&filename, the_repository->objects->odb->path);\n> +\t\tstrbuf_addch(&filename, '/');\n> +\t} else {\n> +\t\tloose_object_path(the_repository, &filename, oid);\n> +\t}\n>  \n>  \tfd = create_tmpfile(&tmp_file, filename.buf);\n>  \tif (fd < 0) {\n> @@ -1939,12 +1946,31 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n>  \t\t    ret);\n>  \tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n> -\tif (!oideq(oid, &parano_oid))\n> +\tif (!is_null_oid(oid) && !oideq(oid, &parano_oid))\n>  \t\tdie(_(\"confused by unstable object source data for %s\"),\n>  \t\t    oid_to_hex(oid));\n>  \n>  \tclose_loose_object(fd);\n>  \n> +\tif (is_null_oid(oid)) {\n> +\t\tint dirlen;\n> +\n> +\t\toidcpy((struct object_id *)oid, &parano_oid);\n> +\t\tloose_object_path(the_repository, &filename, oid);\n\nWhy are we breaking the promise that \"oid\" is constant here? I tested\nlocally with the below on top, and it seems to work (at least no tests\nbroke). Isn't it preferrable to the cast & the caller having its \"oid\"\nchanged?\n\ndiff --git a/object-file.c b/object-file.c\nindex 71d510614b9..d014e6942ea 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1958,10 +1958,11 @@ int write_loose_object(const struct object_id *oid, char *hdr,\n \tclose_loose_object(fd);\n \n \tif (is_null_oid(oid)) {\n+\t\tstruct object_id oid2;\n \t\tint dirlen;\n \n-\t\toidcpy((struct object_id *)oid, &parano_oid);\n-\t\tloose_object_path(the_repository, &filename, oid);\n+\t\toidcpy(&oid2, &parano_oid);\n+\t\tloose_object_path(the_repository, &filename, &oid2);\n \n \t\t/* We finally know the object path, and create the missing dir. */\n \t\tdirlen = directory_size(filename.buf);\n"},{"id":"442993","messageId":"211203.86r1atst50.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"20211203093530.93589-2-chiyutianyi@gmail.com","subject":"Re: [PATCH v4 1/5] object-file: refactor write_loose_object() to read buffer from stream","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-03T13:28:24Z","receivedAt":"2021-12-03T13:41:32Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, Dec 03 2021, Han Xin wrote:\n\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> We used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\n> entire contents of a blob object, no matter how big it is. This\n> implementation may consume all the memory and cause OOM.\n>\n> This can be improved by feeding data to \"write_loose_object()\" in a\n> stream. The input stream is implemented as an interface. In the first\n> step, we make a simple implementation, feeding the entire buffer in the\n> \"stream\" to \"write_loose_object()\" as a refactor.\n>\n> Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  object-file.c  | 53 ++++++++++++++++++++++++++++++++++++++++++++++----\n>  object-store.h |  6 ++++++\n>  2 files changed, 55 insertions(+), 4 deletions(-)\n>\n> diff --git a/object-file.c b/object-file.c\n> index eb972cdccd..82656f7428 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1860,8 +1860,26 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n>  \treturn fd;\n>  }\n>  \n> +struct simple_input_stream_data {\n> +\tconst void *buf;\n> +\tunsigned long len;\n> +};\n\nI see why you picked \"const void *buf\" here, over say const char *, it's\nwhat \"struct input_stream\" uses.\n\nBut why not use size_t for the length, as input_stream does?\n\n> +static const void *feed_simple_input_stream(struct input_stream *in_stream, unsigned long *len)\n> +{\n> +\tstruct simple_input_stream_data *data = in_stream->data;\n> +\n> +\tif (data->len == 0) {\n\nnit: if (!data->len)...\n\n> +\t\t*len = 0;\n> +\t\treturn NULL;\n> +\t}\n> +\t*len = data->len;\n> +\tdata->len = 0;\n> +\treturn data->buf;\n\nBut isn't the body of this functin the same as:\n\n        *len = data->len;\n        if (!len)\n                return NULL;\n        data->len = 0;\n        return data->buf;\n\nI.e. you don't need the condition for setting \"*len\" if it's 0, then\ndata->len is also 0. You just want to return NULL afterwards, and not\nset (harmless, but no need) data->len to 0)< or return data->buf.\n> +\tstruct input_stream in_stream = {\n> +\t\t.read = feed_simple_input_stream,\n> +\t\t.data = (void *)&(struct simple_input_stream_data) {\n> +\t\t\t.buf = buf,\n> +\t\t\t.len = len,\n> +\t\t},\n> +\t\t.size = len,\n> +\t};\n\nMaybe it's that I'm unused to it, but I find this a bit more readable:\n\t\n\t@@ -2013,12 +2011,13 @@ int write_object_file_flags(const void *buf, unsigned long len,\n\t {\n\t \tchar hdr[MAX_HEADER_LEN];\n\t \tint hdrlen = sizeof(hdr);\n\t+\tstruct simple_input_stream_data tmp = {\n\t+\t\t.buf = buf,\n\t+\t\t.len = len,\n\t+\t};\n\t \tstruct input_stream in_stream = {\n\t \t\t.read = feed_simple_input_stream,\n\t-\t\t.data = (void *)&(struct simple_input_stream_data) {\n\t-\t\t\t.buf = buf,\n\t-\t\t\t.len = len,\n\t-\t\t},\n\t+\t\t.data = (void *)&tmp,\n\t \t\t.size = len,\n\t \t};\n\t\nYes there's a temporary variable, but no denser inline casting. Also\neasier to strep through in a debugger (which will have the type\ninformation on \"tmp\".\n\n>  int hash_object_file_literally(const void *buf, unsigned long len,\n> @@ -1977,6 +2006,14 @@ int hash_object_file_literally(const void *buf, unsigned long len,\n>  {\n>  \tchar *header;\n>  \tint hdrlen, status = 0;\n> +\tstruct input_stream in_stream = {\n> +\t\t.read = feed_simple_input_stream,\n> +\t\t.data = (void *)&(struct simple_input_stream_data) {\n> +\t\t\t.buf = buf,\n> +\t\t\t.len = len,\n> +\t\t},\n> +\t\t.size = len,\n> +\t};\n\nditto..\n\n>  \t/* type string, SP, %lu of the length plus NUL must fit this */\n>  \thdrlen = strlen(type) + MAX_HEADER_LEN;\n> @@ -1988,7 +2025,7 @@ int hash_object_file_literally(const void *buf, unsigned long len,\n>  \t\tgoto cleanup;\n>  \tif (freshen_packed_object(oid) || freshen_loose_object(oid))\n>  \t\tgoto cleanup;\n> -\tstatus = write_loose_object(oid, header, hdrlen, buf, len, 0, 0);\n> +\tstatus = write_loose_object(oid, header, hdrlen, &in_stream, 0, 0);\n>  \n>  cleanup:\n>  \tfree(header);\n> @@ -2003,14 +2040,22 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n>  \tchar hdr[MAX_HEADER_LEN];\n>  \tint hdrlen;\n>  \tint ret;\n> +\tstruct simple_input_stream_data data;\n> +\tstruct input_stream in_stream = {\n> +\t\t.read = feed_simple_input_stream,\n> +\t\t.data = &data,\n> +\t};\n>  \n>  \tif (has_loose_object(oid))\n>  \t\treturn 0;\n>  \tbuf = read_object(the_repository, oid, &type, &len);\n> +\tin_stream.size = len;\n\nWhy are we setting this here?...\n\n>  \tif (!buf)\n>  \t\treturn error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n\n...Insted of after this point, as we may error and never use it?\n\n> +\tdata.buf = buf;\n> +\tdata.len = len;\n\nProbably won't matter,  just a nit...\n\n> +struct input_stream {\n> +\tconst void *(*read)(struct input_stream *, unsigned long *len);\n> +\tvoid *data;\n> +\tsize_t size;\n> +};\n> +\n\nAh, and here's the size_t... :)\n"},{"id":"442994","messageId":"211203.86mtlhssj4.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"20211203093530.93589-3-chiyutianyi@gmail.com","subject":"Re: [PATCH v4 2/5] object-file.c: handle undetermined oid in write_loose_object()","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-03T13:41:28Z","receivedAt":"2021-12-03T13:54:36Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, Dec 03 2021, Han Xin wrote:\n\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> When streaming a large blob object to \"write_loose_object()\", we have no\n> chance to run \"write_object_file_prepare()\" to calculate the oid in\n> advance. So we need to handle undetermined oid in function\n> \"write_loose_object()\".\n>\n> In the original implementation, we know the oid and we can write the\n> temporary file in the same directory as the final object, but for an\n> object with an undetermined oid, we don't know the exact directory for\n> the object, so we have to save the temporary file in \".git/objects/\"\n> directory instead.\n>\n> Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  object-file.c | 30 ++++++++++++++++++++++++++++--\n>  1 file changed, 28 insertions(+), 2 deletions(-)\n>\n> diff --git a/object-file.c b/object-file.c\n> index 82656f7428..1c41587bfb 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1892,7 +1892,14 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \tconst void *buf;\n>  \tunsigned long len;\n>  \n> -\tloose_object_path(the_repository, &filename, oid);\n> +\tif (is_null_oid(oid)) {\n> +\t\t/* When oid is not determined, save tmp file to odb path. */\n> +\t\tstrbuf_reset(&filename);\n\nWhy re-use this & leak memory? An existing strbuf use in this function\ndoesn't leak in the same way. Just release it as in the below patch on\ntop (the ret v.s. err variable naming is a bit confused, maybe could do\nwith a prep cleanup step.).\n\n> +\t\tstrbuf_addstr(&filename, the_repository->objects->odb->path);\n> +\t\tstrbuf_addch(&filename, '/');\n\nAnd once we do that this could just become:\n\n\tstrbuf_addf($filename, \"%s/\", ...)\n\nIs there's existing uses of this pattern, so mayb e not worth it, but it\nallows you to remove the braces on the if/else.\n\ndiff --git a/object-file.c b/object-file.c\nindex 8bd89e7b7ba..2b52f3fc1cc 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1880,7 +1880,7 @@ int write_loose_object(const struct object_id *oid, char *hdr,\n \t\t       int hdrlen, struct input_stream *in_stream,\n \t\t       time_t mtime, unsigned flags)\n {\n-\tint fd, ret;\n+\tint fd, ret, err = 0;\n \tunsigned char compressed[4096];\n \tgit_zstream stream;\n \tgit_hash_ctx c;\n@@ -1892,7 +1892,6 @@ int write_loose_object(const struct object_id *oid, char *hdr,\n \n \tif (is_null_oid(oid)) {\n \t\t/* When oid is not determined, save tmp file to odb path. */\n-\t\tstrbuf_reset(&filename);\n \t\tstrbuf_addstr(&filename, the_repository->objects->odb->path);\n \t\tstrbuf_addch(&filename, '/');\n \t} else {\n@@ -1902,11 +1901,12 @@ int write_loose_object(const struct object_id *oid, char *hdr,\n \tfd = create_tmpfile(&tmp_file, filename.buf);\n \tif (fd < 0) {\n \t\tif (flags & HASH_SILENT)\n-\t\t\treturn -1;\n+\t\t\terr = -1;\n \t\telse if (errno == EACCES)\n-\t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n+\t\t\terr = error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n \t\telse\n-\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n+\t\t\terr = error_errno(_(\"unable to create temporary file\"));\n+\t\tgoto cleanup;\n \t}\n \n \t/* Set it up */\n@@ -1968,10 +1968,13 @@ int write_loose_object(const struct object_id *oid, char *hdr,\n \t\t\tstruct strbuf dir = STRBUF_INIT;\n \t\t\tstrbuf_add(&dir, filename.buf, dirlen - 1);\n \t\t\tif (mkdir(dir.buf, 0777) && errno != EEXIST)\n-\t\t\t\treturn -1;\n-\t\t\tif (adjust_shared_perm(dir.buf))\n-\t\t\t\treturn -1;\n-\t\t\tstrbuf_release(&dir);\n+\t\t\t\terr = -1;\n+\t\t\telse if (adjust_shared_perm(dir.buf))\n+\t\t\t\terr = -1;\n+\t\t\telse\n+\t\t\t\tstrbuf_release(&dir);\n+\t\t\tif (err < 0)\n+\t\t\t\tgoto cleanup;\n \t\t}\n \t}\n \n@@ -1984,7 +1987,10 @@ int write_loose_object(const struct object_id *oid, char *hdr,\n \t\t\twarning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n \t}\n \n-\treturn finalize_object_file(tmp_file.buf, filename.buf);\n+\terr = finalize_object_file(tmp_file.buf, filename.buf);\n+cleanup:\n+\tstrbuf_release(&filename);\n+\treturn err;\n }\n \n static int freshen_loose_object(const struct object_id *oid)\n"},{"id":"442996","messageId":"211203.86ilw5ssar.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"20211203093530.93589-6-chiyutianyi@gmail.com","subject":"Re: [PATCH v4 5/5] unpack-objects: unpack_non_delta_entry() read data in a stream","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-03T13:54:40Z","receivedAt":"2021-12-03T13:59:34Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, Dec 03 2021, Han Xin wrote:\n\n> diff --git a/t/t5590-unpack-non-delta-objects.sh b/t/t5590-unpack-non-delta-objects.sh\n> new file mode 100755\n> index 0000000000..01d950d119\n> --- /dev/null\n> +++ b/t/t5590-unpack-non-delta-objects.sh\n> @@ -0,0 +1,76 @@\n> +#!/bin/sh\n> +#\n> +# Copyright (c) 2021 Han Xin\n> +#\n> +\n> +test_description='Test unpack-objects when receive pack'\n> +\n> +GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n> +export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n> +\n> +. ./test-lib.sh\n> +\n> +test_expect_success \"create commit with big blobs (1.5 MB)\" '\n> +\ttest-tool genrandom foo 1500000 >big-blob &&\n> +\ttest_commit --append foo big-blob &&\n> +\ttest-tool genrandom bar 1500000 >big-blob &&\n> +\ttest_commit --append bar big-blob &&\n> +\t(\n> +\t\tcd .git &&\n> +\t\tfind objects/?? -type f | sort\n\n...are thse...\n\n> +\t) >expect &&\n> +\tPACK=$(echo main | git pack-objects --progress --revs test)\n\nIs --progress needed?\n\n> +'\n> +\n> +test_expect_success 'setup GIT_ALLOC_LIMIT to 1MB' '\n> +\tGIT_ALLOC_LIMIT=1m &&\n> +\texport GIT_ALLOC_LIMIT\n> +'\n> +\n> +test_expect_success 'prepare dest repository' '\n> +\tgit init --bare dest.git &&\n> +\tgit -C dest.git config core.bigFileThreshold 2m &&\n> +\tgit -C dest.git config receive.unpacklimit 100\n\nI think it would be better to just (could roll this into a function):\n\n\ttest_when_finished \"rm -rf dest.git\" &&\n\tgit init dest.git &&\n\tgit -C dest.git config ...\n\nThen you can use it with e.g. --run=3-4 and not have it error out\nbecause of skipped setup.\n\nA lot of our tests fail like that, but in this case fixing it seems\ntrivial.\n\n\n\n> +'\n> +\n> +test_expect_success 'fail to unpack-objects: cannot allocate' '\n> +\ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n> +\ttest_i18ngrep \"fatal: attempting to allocate\" err &&\n\nnit: just \"grep\", not \"test_i18ngrep\"\n\n> +\t(\n> +\t\tcd dest.git &&\n> +\t\tfind objects/?? -type f | sort\n\n...\"find\" needed over just globbing?:\n\n    obj=$(echo objects/*/*)\n\n?\n"},{"id":"442997","messageId":"211203.86ee6tss9p.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"20211203093530.93589-5-chiyutianyi@gmail.com","subject":"Re: [PATCH v4 4/5] unpack-objects.c: add dry_run mode for get_data()","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-03T13:59:29Z","receivedAt":"2021-12-03T14:00:09Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, Dec 03 2021, Han Xin wrote:\n\n> +\tunsigned long bufsize = dry_run ? 4096 : size;\n> +\tvoid *buf = xmallocz(bufsize);\n\nIt's probably nothing, but in your CL you note that you changed another\nhardcoding from 4k to 8k, should this one still be 4k?\n\nIt's probably fine, just wondering...\n"},{"id":"442999","messageId":"211203.86a6hhsqwf.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"20211203093530.93589-6-chiyutianyi@gmail.com","subject":"Re: [PATCH v4 5/5] unpack-objects: unpack_non_delta_entry() read data in a stream","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-03T14:05:44Z","receivedAt":"2021-12-03T14:29:40Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, Dec 03 2021, Han Xin wrote:\n\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n> [..]\n> +static void write_stream_blob(unsigned nr, unsigned long size)\n> +{\n> +\tchar hdr[32];\n> +\tint hdrlen;\n> +\tgit_zstream zstream;\n> +\tstruct input_zstream_data data;\n> +\tstruct input_stream in_stream = {\n> +\t\t.read = feed_input_zstream,\n> +\t\t.data = &data,\n> +\t\t.size = size,\n> +\t};\n> +\tstruct object_id *oid = &obj_list[nr].oid;\n> +\tint ret;\n> +\n> +\tmemset(&zstream, 0, sizeof(zstream));\n> +\tmemset(&data, 0, sizeof(data));\n> +\tdata.zstream = &zstream;\n> +\tgit_inflate_init(&zstream);\n> +\n> +\t/* Generate the header */\n> +\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX, type_name(OBJ_BLOB), (uintmax_t)size) + 1;\n> +\n> +\tif ((ret = write_loose_object(oid, hdr, hdrlen, &in_stream, 0, 0)))\n> +\t\tdie(_(\"failed to write object in stream %d\"), ret);\n> +\n> +\tif (zstream.total_out != size || data.status != Z_STREAM_END)\n> +\t\tdie(_(\"inflate returned %d\"), data.status);\n> +\tgit_inflate_end(&zstream);\n> +\n> +\tif (strict && !dry_run) {\n> +\t\tstruct blob *blob = lookup_blob(the_repository, oid);\n> +\t\tif (blob)\n> +\t\t\tblob->object.flags |= FLAG_WRITTEN;\n> +\t\telse\n> +\t\t\tdie(\"invalid blob object from stream\");\n> +\t}\n> +\tobj_list[nr].obj = NULL;\n> +}\n\nJust a side-note, I think (but am not 100% sure) that these existing\noccurances aren't needed due to our use of CALLOC_ARRAY():\n    \n    diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n    index 4a9466295ba..00b349412c5 100644\n    --- a/builtin/unpack-objects.c\n    +++ b/builtin/unpack-objects.c\n    @@ -248,7 +248,6 @@ static void write_object(unsigned nr, enum object_type type,\n                            die(\"failed to write object\");\n                    added_object(nr, type, buf, size);\n                    free(buf);\n    -               obj_list[nr].obj = NULL;\n            } else if (type == OBJ_BLOB) {\n                    struct blob *blob;\n                    if (write_object_file(buf, size, type_name(type),\n    @@ -262,7 +261,6 @@ static void write_object(unsigned nr, enum object_type type,\n                            blob->object.flags |= FLAG_WRITTEN;\n                    else\n                            die(\"invalid blob object\");\n    -               obj_list[nr].obj = NULL;\n            } else {\n                    struct object *obj;\n                    int eaten;\n\nThe reason I'm noting it is that the same seems to be true of your new\naddition here. I.e. are these assignments to NULL needed?\n\nAnyway, the reason I started poking at this it tha this\nwrite_stream_blob() seems to duplicate much of write_object(). AFAICT\nonly the writing part is really different, the part where we\nlookup_blob() after, set FLAG_WRITTEN etc. is all the same.\n\nWhy can't we call write_object() here?\n\nThe obvious answer seems to be that the call to write_object_file()\nisn't prepared to do the sort of streaming that you want, so instead\nyou're bypassing it and calling write_loose_object() directly.\n\nI haven't tried this myself, but isn't a better and cleaner approach\nhere to not add another meaning to what is_null_oid() means, but to just\nadd a HASH_STREAM flag that'll get passed down as \"unsigned flags\" to\nwrite_loose_object()? See FLAG_BITS in object.h.\n\nThen the \"obj_list[nr].obj\" here could also become\n\"obj_list[nr].obj.flags |= (1u<<12)\" or whatever (but that wouldn't\nstrictly be needed I think.\n\nBut by adding the \"HASH_STREAM\" flag you could I think stop duplicating\nthe \"Generate the header\" etc. here and call write_object_file_flags().\n\nI don't so much care about how it's done within unpack-objects.c, but\nnot having another meaning to is_null_oid() in play would be really\nnice, and it this case it seems entirely avoidable.\n"},{"id":"443129","messageId":"CAO0brD2s9ZiVc+1BvE8r_86YvywJOrdyth1R+O1r5KwRtnuCrQ@mail.gmail.com","threadId":"56672","inReplyTo":"211203.86r1atst50.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v4 1/5] object-file: refactor write_loose_object() to read buffer from stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-06T02:07:25Z","receivedAt":"2021-12-06T02:07:40Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Fri, Dec 3, 2021 at 9:41 PM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>\n>\n> On Fri, Dec 03 2021, Han Xin wrote:\n>\n> > From: Han Xin <hanxin.hx@alibaba-inc.com>\n> >\n> > We used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\n> > entire contents of a blob object, no matter how big it is. This\n> > implementation may consume all the memory and cause OOM.\n> >\n> > This can be improved by feeding data to \"write_loose_object()\" in a\n> > stream. The input stream is implemented as an interface. In the first\n> > step, we make a simple implementation, feeding the entire buffer in the\n> > \"stream\" to \"write_loose_object()\" as a refactor.\n> >\n> > Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> > Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> > ---\n> >  object-file.c  | 53 ++++++++++++++++++++++++++++++++++++++++++++++----\n> >  object-store.h |  6 ++++++\n> >  2 files changed, 55 insertions(+), 4 deletions(-)\n> >\n> > diff --git a/object-file.c b/object-file.c\n> > index eb972cdccd..82656f7428 100644\n> > --- a/object-file.c\n> > +++ b/object-file.c\n> > @@ -1860,8 +1860,26 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n> >       return fd;\n> >  }\n> >\n> > +struct simple_input_stream_data {\n> > +     const void *buf;\n> > +     unsigned long len;\n> > +};\n>\n> I see why you picked \"const void *buf\" here, over say const char *, it's\n> what \"struct input_stream\" uses.\n>\n> But why not use size_t for the length, as input_stream does?\n>\n\nYes, \"size_t\" will be better here.\n\n> > +static const void *feed_simple_input_stream(struct input_stream *in_stream, unsigned long *len)\n> > +{\n> > +     struct simple_input_stream_data *data = in_stream->data;\n> > +\n> > +     if (data->len == 0) {\n>\n> nit: if (!data->len)...\n>\n\nWill apply.\n\n> > +             *len = 0;\n> > +             return NULL;\n> > +     }\n> > +     *len = data->len;\n> > +     data->len = 0;\n> > +     return data->buf;\n>\n> But isn't the body of this functin the same as:\n>\n>         *len = data->len;\n>         if (!len)\n>                 return NULL;\n>         data->len = 0;\n>         return data->buf;\n>\n> I.e. you don't need the condition for setting \"*len\" if it's 0, then\n> data->len is also 0. You just want to return NULL afterwards, and not\n> set (harmless, but no need) data->len to 0)< or return data->buf.\n\nWill apply.\n\n> > +     struct input_stream in_stream = {\n> > +             .read = feed_simple_input_stream,\n> > +             .data = (void *)&(struct simple_input_stream_data) {\n> > +                     .buf = buf,\n> > +                     .len = len,\n> > +             },\n> > +             .size = len,\n> > +     };\n>\n> Maybe it's that I'm unused to it, but I find this a bit more readable:\n>\n>         @@ -2013,12 +2011,13 @@ int write_object_file_flags(const void *buf, unsigned long len,\n>          {\n>                 char hdr[MAX_HEADER_LEN];\n>                 int hdrlen = sizeof(hdr);\n>         +       struct simple_input_stream_data tmp = {\n>         +               .buf = buf,\n>         +               .len = len,\n>         +       };\n>                 struct input_stream in_stream = {\n>                         .read = feed_simple_input_stream,\n>         -               .data = (void *)&(struct simple_input_stream_data) {\n>         -                       .buf = buf,\n>         -                       .len = len,\n>         -               },\n>         +               .data = (void *)&tmp,\n>                         .size = len,\n>                 };\n>\n> Yes there's a temporary variable, but no denser inline casting. Also\n> easier to strep through in a debugger (which will have the type\n> information on \"tmp\".\n>\n\nWill apply.\n\n> >  int hash_object_file_literally(const void *buf, unsigned long len,\n> > @@ -1977,6 +2006,14 @@ int hash_object_file_literally(const void *buf, unsigned long len,\n> >  {\n> >       char *header;\n> >       int hdrlen, status = 0;\n> > +     struct input_stream in_stream = {\n> > +             .read = feed_simple_input_stream,\n> > +             .data = (void *)&(struct simple_input_stream_data) {\n> > +                     .buf = buf,\n> > +                     .len = len,\n> > +             },\n> > +             .size = len,\n> > +     };\n>\n> ditto..\n>\n> >       /* type string, SP, %lu of the length plus NUL must fit this */\n> >       hdrlen = strlen(type) + MAX_HEADER_LEN;\n> > @@ -1988,7 +2025,7 @@ int hash_object_file_literally(const void *buf, unsigned long len,\n> >               goto cleanup;\n> >       if (freshen_packed_object(oid) || freshen_loose_object(oid))\n> >               goto cleanup;\n> > -     status = write_loose_object(oid, header, hdrlen, buf, len, 0, 0);\n> > +     status = write_loose_object(oid, header, hdrlen, &in_stream, 0, 0);\n> >\n> >  cleanup:\n> >       free(header);\n> > @@ -2003,14 +2040,22 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n> >       char hdr[MAX_HEADER_LEN];\n> >       int hdrlen;\n> >       int ret;\n> > +     struct simple_input_stream_data data;\n> > +     struct input_stream in_stream = {\n> > +             .read = feed_simple_input_stream,\n> > +             .data = &data,\n> > +     };\n> >\n> >       if (has_loose_object(oid))\n> >               return 0;\n> >       buf = read_object(the_repository, oid, &type, &len);\n> > +     in_stream.size = len;\n>\n> Why are we setting this here?...\n>\n\nYes, putting \"in_stream.size=len;\" here was a stupid decision.\n\n> >       if (!buf)\n> >               return error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n>\n> ...Insted of after this point, as we may error and never use it?\n>\n> > +     data.buf = buf;\n> > +     data.len = len;\n>\n> Probably won't matter,  just a nit...\n>\n> > +struct input_stream {\n> > +     const void *(*read)(struct input_stream *, unsigned long *len);\n> > +     void *data;\n> > +     size_t size;\n> > +};\n> > +\n>\n> Ah, and here's the size_t... :)\n"},{"id":"443131","messageId":"CAO0brD2_TvxCMhGiyZC5ex-73dk+3CafWFT43K9CPfE8WAXKXQ@mail.gmail.com","threadId":"56672","inReplyTo":"211203.86v905stru.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v4 2/5] object-file.c: handle undetermined oid in write_loose_object()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-06T02:51:44Z","receivedAt":"2021-12-06T02:51:58Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Fri, Dec 3, 2021 at 9:27 PM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>\n>\n> On Fri, Dec 03 2021, Han Xin wrote:\n>\n> > From: Han Xin <hanxin.hx@alibaba-inc.com>\n> >\n> > When streaming a large blob object to \"write_loose_object()\", we have no\n> > chance to run \"write_object_file_prepare()\" to calculate the oid in\n> > advance. So we need to handle undetermined oid in function\n> > \"write_loose_object()\".\n> >\n> > In the original implementation, we know the oid and we can write the\n> > temporary file in the same directory as the final object, but for an\n> > object with an undetermined oid, we don't know the exact directory for\n> > the object, so we have to save the temporary file in \".git/objects/\"\n> > directory instead.\n> >\n> > Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> > Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> > ---\n> >  object-file.c | 30 ++++++++++++++++++++++++++++--\n> >  1 file changed, 28 insertions(+), 2 deletions(-)\n> >\n> > diff --git a/object-file.c b/object-file.c\n> > index 82656f7428..1c41587bfb 100644\n> > --- a/object-file.c\n> > +++ b/object-file.c\n> > @@ -1892,7 +1892,14 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n> >       const void *buf;\n> >       unsigned long len;\n> >\n> > -     loose_object_path(the_repository, &filename, oid);\n> > +     if (is_null_oid(oid)) {\n> > +             /* When oid is not determined, save tmp file to odb path. */\n> > +             strbuf_reset(&filename);\n> > +             strbuf_addstr(&filename, the_repository->objects->odb->path);\n> > +             strbuf_addch(&filename, '/');\n> > +     } else {\n> > +             loose_object_path(the_repository, &filename, oid);\n> > +     }\n> >\n> >       fd = create_tmpfile(&tmp_file, filename.buf);\n> >       if (fd < 0) {\n> > @@ -1939,12 +1946,31 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n> >               die(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n> >                   ret);\n> >       the_hash_algo->final_oid_fn(&parano_oid, &c);\n> > -     if (!oideq(oid, &parano_oid))\n> > +     if (!is_null_oid(oid) && !oideq(oid, &parano_oid))\n> >               die(_(\"confused by unstable object source data for %s\"),\n> >                   oid_to_hex(oid));\n> >\n> >       close_loose_object(fd);\n> >\n> > +     if (is_null_oid(oid)) {\n> > +             int dirlen;\n> > +\n> > +             oidcpy((struct object_id *)oid, &parano_oid);\n> > +             loose_object_path(the_repository, &filename, oid);\n>\n> Why are we breaking the promise that \"oid\" is constant here? I tested\n> locally with the below on top, and it seems to work (at least no tests\n> broke). Isn't it preferrable to the cast & the caller having its \"oid\"\n> changed?\n>\n> diff --git a/object-file.c b/object-file.c\n> index 71d510614b9..d014e6942ea 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1958,10 +1958,11 @@ int write_loose_object(const struct object_id *oid, char *hdr,\n>         close_loose_object(fd);\n>\n>         if (is_null_oid(oid)) {\n> +               struct object_id oid2;\n>                 int dirlen;\n>\n> -               oidcpy((struct object_id *)oid, &parano_oid);\n> -               loose_object_path(the_repository, &filename, oid);\n> +               oidcpy(&oid2, &parano_oid);\n> +               loose_object_path(the_repository, &filename, &oid2);\n>\n>                 /* We finally know the object path, and create the missing dir. */\n>                 dirlen = directory_size(filename.buf);\n\nMaybe I should change the promise that \"oid\" is constant in\n\"write_loose_object()\".\n\nThe original write_object_file_flags() defines a variable \"oid\", and\ncompletes the calculation of the \"oid\" in\n\"write_object_file_prepare()\" which will be passed to\n\"write_loose_object()\".\n\nIf a null oid is maintained after calling \"write_loose_object()\",\n\"--strict\" will become meaningless, although it does not break existing\ntest cases.\n"},{"id":"443134","messageId":"CAO0brD3=J1hQPypqqmyYqL9wP+xxThAO6womU_h_s0wDBqJWbg@mail.gmail.com","threadId":"56672","inReplyTo":"211203.86mtlhssj4.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v4 2/5] object-file.c: handle undetermined oid in write_loose_object()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-06T03:12:11Z","receivedAt":"2021-12-06T03:12:26Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Fri, Dec 3, 2021 at 9:54 PM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>\n>\n> On Fri, Dec 03 2021, Han Xin wrote:\n>\n> > From: Han Xin <hanxin.hx@alibaba-inc.com>\n> >\n> > When streaming a large blob object to \"write_loose_object()\", we have no\n> > chance to run \"write_object_file_prepare()\" to calculate the oid in\n> > advance. So we need to handle undetermined oid in function\n> > \"write_loose_object()\".\n> >\n> > In the original implementation, we know the oid and we can write the\n> > temporary file in the same directory as the final object, but for an\n> > object with an undetermined oid, we don't know the exact directory for\n> > the object, so we have to save the temporary file in \".git/objects/\"\n> > directory instead.\n> >\n> > Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> > Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> > ---\n> >  object-file.c | 30 ++++++++++++++++++++++++++++--\n> >  1 file changed, 28 insertions(+), 2 deletions(-)\n> >\n> > diff --git a/object-file.c b/object-file.c\n> > index 82656f7428..1c41587bfb 100644\n> > --- a/object-file.c\n> > +++ b/object-file.c\n> > @@ -1892,7 +1892,14 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n> >       const void *buf;\n> >       unsigned long len;\n> >\n> > -     loose_object_path(the_repository, &filename, oid);\n> > +     if (is_null_oid(oid)) {\n> > +             /* When oid is not determined, save tmp file to odb path. */\n> > +             strbuf_reset(&filename);\n>\n> Why re-use this & leak memory? An existing strbuf use in this function\n> doesn't leak in the same way. Just release it as in the below patch on\n> top (the ret v.s. err variable naming is a bit confused, maybe could do\n> with a prep cleanup step.).\n>\n> > +             strbuf_addstr(&filename, the_repository->objects->odb->path);\n> > +             strbuf_addch(&filename, '/');\n>\n> And once we do that this could just become:\n>\n>         strbuf_addf($filename, \"%s/\", ...)\n>\n> Is there's existing uses of this pattern, so mayb e not worth it, but it\n> allows you to remove the braces on the if/else.\n>\n> diff --git a/object-file.c b/object-file.c\n> index 8bd89e7b7ba..2b52f3fc1cc 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1880,7 +1880,7 @@ int write_loose_object(const struct object_id *oid, char *hdr,\n>                        int hdrlen, struct input_stream *in_stream,\n>                        time_t mtime, unsigned flags)\n>  {\n> -       int fd, ret;\n> +       int fd, ret, err = 0;\n>         unsigned char compressed[4096];\n>         git_zstream stream;\n>         git_hash_ctx c;\n> @@ -1892,7 +1892,6 @@ int write_loose_object(const struct object_id *oid, char *hdr,\n>\n>         if (is_null_oid(oid)) {\n>                 /* When oid is not determined, save tmp file to odb path. */\n> -               strbuf_reset(&filename);\n>                 strbuf_addstr(&filename, the_repository->objects->odb->path);\n>                 strbuf_addch(&filename, '/');\n>         } else {\n> @@ -1902,11 +1901,12 @@ int write_loose_object(const struct object_id *oid, char *hdr,\n>         fd = create_tmpfile(&tmp_file, filename.buf);\n>         if (fd < 0) {\n>                 if (flags & HASH_SILENT)\n> -                       return -1;\n> +                       err = -1;\n>                 else if (errno == EACCES)\n> -                       return error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n> +                       err = error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n>                 else\n> -                       return error_errno(_(\"unable to create temporary file\"));\n> +                       err = error_errno(_(\"unable to create temporary file\"));\n> +               goto cleanup;\n>         }\n>\n>         /* Set it up */\n> @@ -1968,10 +1968,13 @@ int write_loose_object(const struct object_id *oid, char *hdr,\n>                         struct strbuf dir = STRBUF_INIT;\n>                         strbuf_add(&dir, filename.buf, dirlen - 1);\n>                         if (mkdir(dir.buf, 0777) && errno != EEXIST)\n> -                               return -1;\n> -                       if (adjust_shared_perm(dir.buf))\n> -                               return -1;\n> -                       strbuf_release(&dir);\n> +                               err = -1;\n> +                       else if (adjust_shared_perm(dir.buf))\n> +                               err = -1;\n> +                       else\n> +                               strbuf_release(&dir);\n> +                       if (err < 0)\n> +                               goto cleanup;\n>                 }\n>         }\n>\n> @@ -1984,7 +1987,10 @@ int write_loose_object(const struct object_id *oid, char *hdr,\n>                         warning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n>         }\n>\n> -       return finalize_object_file(tmp_file.buf, filename.buf);\n> +       err = finalize_object_file(tmp_file.buf, filename.buf);\n> +cleanup:\n> +       strbuf_release(&filename);\n> +       return err;\n>  }\n>\n>  static int freshen_loose_object(const struct object_id *oid)\n\nYes, this will be much better. Will apply.\n"},{"id":"443135","messageId":"CAO0brD2zHETRWz-0Bcy1+tjkFS0XF94+SdBj-sDfZcq5iG=6ZQ@mail.gmail.com","threadId":"56672","inReplyTo":"211203.86ee6tss9p.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v4 4/5] unpack-objects.c: add dry_run mode for get_data()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-06T03:20:23Z","receivedAt":"2021-12-06T03:20:37Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Fri, Dec 3, 2021 at 10:00 PM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Fri, Dec 03 2021, Han Xin wrote:\n>\n> > +     unsigned long bufsize = dry_run ? 4096 : size;\n> > +     void *buf = xmallocz(bufsize);\n>\n> It's probably nothing, but in your CL you note that you changed another\n> hardcoding from 4k to 8k, should this one still be 4k?\n>\n> It's probably fine, just wondering...\n\nYes, I think this is an omission from my work.\n"},{"id":"443236","messageId":"CAO0brD32=0Gc4_5-aJS5MYicV3uuuMkqpo3adVAELq21V8EvdA@mail.gmail.com","threadId":"56672","inReplyTo":"211203.86ilw5ssar.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v4 5/5] unpack-objects: unpack_non_delta_entry() read data in a stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-07T06:17:39Z","receivedAt":"2021-12-07T06:17:54Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Fri, Dec 3, 2021 at 9:59 PM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>\n>\n> On Fri, Dec 03 2021, Han Xin wrote:\n>\n> > diff --git a/t/t5590-unpack-non-delta-objects.sh b/t/t5590-unpack-non-delta-objects.sh\n> > new file mode 100755\n> > index 0000000000..01d950d119\n> > --- /dev/null\n> > +++ b/t/t5590-unpack-non-delta-objects.sh\n> > @@ -0,0 +1,76 @@\n> > +#!/bin/sh\n> > +#\n> > +# Copyright (c) 2021 Han Xin\n> > +#\n> > +\n> > +test_description='Test unpack-objects when receive pack'\n> > +\n> > +GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n> > +export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n> > +\n> > +. ./test-lib.sh\n> > +\n> > +test_expect_success \"create commit with big blobs (1.5 MB)\" '\n> > +     test-tool genrandom foo 1500000 >big-blob &&\n> > +     test_commit --append foo big-blob &&\n> > +     test-tool genrandom bar 1500000 >big-blob &&\n> > +     test_commit --append bar big-blob &&\n> > +     (\n> > +             cd .git &&\n> > +             find objects/?? -type f | sort\n>\n> ...are thse...\n>\n> > +     ) >expect &&\n> > +     PACK=$(echo main | git pack-objects --progress --revs test)\n>\n> Is --progress needed?\n>\n\n\"--progress\" is not necessary.\n\n> > +'\n> > +\n> > +test_expect_success 'setup GIT_ALLOC_LIMIT to 1MB' '\n> > +     GIT_ALLOC_LIMIT=1m &&\n> > +     export GIT_ALLOC_LIMIT\n> > +'\n> > +\n> > +test_expect_success 'prepare dest repository' '\n> > +     git init --bare dest.git &&\n> > +     git -C dest.git config core.bigFileThreshold 2m &&\n> > +     git -C dest.git config receive.unpacklimit 100\n>\n> I think it would be better to just (could roll this into a function):\n>\n>         test_when_finished \"rm -rf dest.git\" &&\n>         git init dest.git &&\n>         git -C dest.git config ...\n>\n> Then you can use it with e.g. --run=3-4 and not have it error out\n> because of skipped setup.\n>\n> A lot of our tests fail like that, but in this case fixing it seems\n> trivial.\n>\n>\n\nOK, I will take it.\n\n>\n> > +'\n> > +\n> > +test_expect_success 'fail to unpack-objects: cannot allocate' '\n> > +     test_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n> > +     test_i18ngrep \"fatal: attempting to allocate\" err &&\n>\n> nit: just \"grep\", not \"test_i18ngrep\"\n>\n> > +     (\n> > +             cd dest.git &&\n> > +             find objects/?? -type f | sort\n>\n> ...\"find\" needed over just globbing?:\n>\n>     obj=$(echo objects/*/*)\n>\n> ?\n\nI tried to use \"echo\" instead of \"find\". It works well on my personal\ncomputer, but fails due to the \"info/commit-graph\" generated when CI on\nGithub.\nSo it seems that \".git/objects/??\" will be more rigorous?\n"},{"id":"443237","messageId":"CAO0brD2AYAw9KaKLdMgQURh0RkdcvuGZJTNrhF6ZnpUvhk3d=g@mail.gmail.com","threadId":"56672","inReplyTo":"211203.86zgphsu5a.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v4 5/5] unpack-objects: unpack_non_delta_entry() read data in a stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-07T06:42:23Z","receivedAt":"2021-12-07T06:42:38Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Fri, Dec 3, 2021 at 9:19 PM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>\n>\n> On Fri, Dec 03 2021, Han Xin wrote:\n>\n> > From: Han Xin <hanxin.hx@alibaba-inc.com>\n> >\n> > We used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\n> > entire contents of a blob object, no matter how big it is. This\n> > implementation may consume all the memory and cause OOM.\n> >\n> > By implementing a zstream version of input_stream interface, we can use\n> > a small fixed buffer for \"unpack_non_delta_entry()\".\n> >\n> > However, unpack non-delta objects from a stream instead of from an entrie\n> > buffer will have 10% performance penalty. Therefore, only unpack object\n> > larger than the \"big_file_threshold\" in zstream. See the following\n> > benchmarks:\n> >\n> >     hyperfine \\\n> >       --setup \\\n> >       'if ! test -d scalar.git; then git clone --bare https://github.com/microsoft/scalar.git; cp scalar.git/objects/pack/*.pack small.pack; fi' \\\n> >       --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n> >       -n 'old' 'git -C dest.git unpack-objects <small.pack' \\\n> >       -n 'new' 'new/git -C dest.git unpack-objects <small.pack' \\\n> >       -n 'new (small threshold)' \\\n> >       'new/git -c core.bigfilethreshold=16k -C dest.git unpack-objects <small.pack'\n> >     Benchmark 1: old\n> >       Time (mean ± σ):      6.075 s ±  0.069 s    [User: 5.047 s, System: 0.991 s]\n> >       Range (min … max):    6.018 s …  6.189 s    10 runs\n> >\n> >     Benchmark 2: new\n> >       Time (mean ± σ):      6.090 s ±  0.033 s    [User: 5.075 s, System: 0.976 s]\n> >       Range (min … max):    6.030 s …  6.142 s    10 runs\n> >\n> >     Benchmark 3: new (small threshold)\n> >       Time (mean ± σ):      6.755 s ±  0.029 s    [User: 5.150 s, System: 1.560 s]\n> >       Range (min … max):    6.711 s …  6.809 s    10 runs\n> >\n> >     Summary\n> >       'old' ran\n> >         1.00 ± 0.01 times faster than 'new'\n> >         1.11 ± 0.01 times faster than 'new (small threshold)'\n>\n> So before we wrote used core.bigfilethreshold for two things (or more?):\n> Whether we show a diff for it (we mark it \"binary\") and whether it's\n> split into a loose object.\n>\n> Now it's three things, we've added a \"this is a threshold when we'll\n> stream the object\" to that.\n>\n> Might it make sense to squash something like this in, so we can have our\n> cake & eat it too?\n>\n> With this I get, where HEAD~0 is this change:\n>\n>     Summary\n>       './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'HEAD~0' ran\n>         1.00 ± 0.01 times faster than './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'HEAD~1'\n>         1.00 ± 0.01 times faster than './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'origin/master'\n>         1.01 ± 0.01 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'HEAD~0'\n>         1.06 ± 0.14 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'origin/master'\n>         1.20 ± 0.01 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'HEAD~1'\n>\n> I.e. it's 5% slower, not 20% (haven't looked into why), but we'll not\n> stream out 16k..128MB objects (maybe the repo has even bigger ones?)\n>\n> diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> index c04f62a54a1..601b7a2418f 100644\n> --- a/Documentation/config/core.txt\n> +++ b/Documentation/config/core.txt\n> @@ -424,6 +424,17 @@ be delta compressed, but larger binary media files won't be.\n>  +\n>  Common unit suffixes of 'k', 'm', or 'g' are supported.\n>\n> +core.bigFileStreamingThreshold::\n> +       Files larger than this will be streamed out to a temporary\n> +       object file while being hashed, which will when be renamed\n> +       in-place to a loose object, particularly if the\n> +       `core.bigFileThreshold' setting dictates that they're always\n> +       written out as loose objects.\n> ++\n> +Default is 128 MiB on all platforms.\n> ++\n> +Common unit suffixes of 'k', 'm', or 'g' are supported.\n> +\n>  core.excludesFile::\n>         Specifies the pathname to the file that contains patterns to\n>         describe paths that are not meant to be tracked, in addition\n> diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n> index bedc494e2db..94ce275c807 100644\n> --- a/builtin/unpack-objects.c\n> +++ b/builtin/unpack-objects.c\n> @@ -400,7 +400,7 @@ static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n>         void *buf;\n>\n>         /* Write large blob in stream without allocating full buffer. */\n> -       if (!dry_run && type == OBJ_BLOB && size > big_file_threshold) {\n> +       if (!dry_run && type == OBJ_BLOB && size > big_file_streaming_threshold) {\n>                 write_stream_blob(nr, size);\n>                 return;\n>         }\n> diff --git a/cache.h b/cache.h\n> index eba12487b99..4037c7fd849 100644\n> --- a/cache.h\n> +++ b/cache.h\n> @@ -964,6 +964,7 @@ extern size_t packed_git_window_size;\n>  extern size_t packed_git_limit;\n>  extern size_t delta_base_cache_limit;\n>  extern unsigned long big_file_threshold;\n> +extern unsigned long big_file_streaming_threshold;\n>  extern unsigned long pack_size_limit_cfg;\n>\n>  /*\n> diff --git a/config.c b/config.c\n> index c5873f3a706..7b122a142a8 100644\n> --- a/config.c\n> +++ b/config.c\n> @@ -1408,6 +1408,11 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n>                 return 0;\n>         }\n>\n> +       if (!strcmp(var, \"core.bigfilestreamingthreshold\")) {\n> +               big_file_streaming_threshold = git_config_ulong(var, value);\n> +               return 0;\n> +       }\n> +\n>         if (!strcmp(var, \"core.packedgitlimit\")) {\n>                 packed_git_limit = git_config_ulong(var, value);\n>                 return 0;\n> diff --git a/environment.c b/environment.c\n> index 9da7f3c1a19..4fcc3de7417 100644\n> --- a/environment.c\n> +++ b/environment.c\n> @@ -46,6 +46,7 @@ size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n>  size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n>  size_t delta_base_cache_limit = 96 * 1024 * 1024;\n>  unsigned long big_file_threshold = 512 * 1024 * 1024;\n> +unsigned long big_file_streaming_threshold = 128 * 1024 * 1024;\n>  int pager_use_color = 1;\n>  const char *editor_program;\n>  const char *askpass_program;\n\nI'm not sure if we need an additional \"core.bigFileStreamingThreshold\"\nhere, because \"core.bigFileThreshold\" has been widely used in\n\"index-pack\", \"read_object\" and so on.\n\nIn the test case which uses \"core.bigFileStreamingThreshold\" instead of\n\"core.bigFileThreshold\", I found the test case execution failed because\nof \"fsck\", who tried to allocate 15MB of memory.\nIn the process of \"fsck_loose()\", \"read_loose_object()\" will be called,\nwhich contains the following content:\n\n  if (*oi->typep == OBJ_BLOB && *size> big_file_threshold) {\n    if (check_stream_oid(&stream, hdr, *size, path, expected_oid) <0)\n    goto out;\n  } else {\n    /* this will allocate 15MB of memory */\n    *contents = unpack_loose_rest(&stream, hdr, *size, expected_oid);\n    ...\n  }\n\nThe same case can be found in \"unpack_entry_data()\":\n\n  static char fixed_buf[8192];\n  ...\n  if (type == OBJ_BLOB && size > big_file_threshold)\n    buf = fixed_buf;\n  else\n    buf = xmallocz(size);\n ...\n\nAlthough I know that setting a \"core.bigfilethreshold\" smaller than the\ndefault value on the server side does not help me prevent users from\ncreating large delta objects on the client side, it can still\neffectively help me reduce the Memory allocation in \"receive-pack\".\n\nIf this is not the correct way to use \"core.bigfilethreshold\", maybe\nyou can share some better solutions to me, if you want.\n\nThanks.\n-Han Xin\n"},{"id":"443238","messageId":"CAO0brD1_h0qC=Qk2K1c1aZ=0u73BQnE50xZ0W6py0=m4TgB3XA@mail.gmail.com","threadId":"56672","inReplyTo":"211203.86a6hhsqwf.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v4 5/5] unpack-objects: unpack_non_delta_entry() read data in a stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-07T06:48:51Z","receivedAt":"2021-12-07T06:49:08Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Fri, Dec 3, 2021 at 10:29 PM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Fri, Dec 03 2021, Han Xin wrote:\n>\n> > From: Han Xin <hanxin.hx@alibaba-inc.com>\n> > [..]\n> > +static void write_stream_blob(unsigned nr, unsigned long size)\n> > +{\n> > +     char hdr[32];\n> > +     int hdrlen;\n> > +     git_zstream zstream;\n> > +     struct input_zstream_data data;\n> > +     struct input_stream in_stream = {\n> > +             .read = feed_input_zstream,\n> > +             .data = &data,\n> > +             .size = size,\n> > +     };\n> > +     struct object_id *oid = &obj_list[nr].oid;\n> > +     int ret;\n> > +\n> > +     memset(&zstream, 0, sizeof(zstream));\n> > +     memset(&data, 0, sizeof(data));\n> > +     data.zstream = &zstream;\n> > +     git_inflate_init(&zstream);\n> > +\n> > +     /* Generate the header */\n> > +     hdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX, type_name(OBJ_BLOB), (uintmax_t)size) + 1;\n> > +\n> > +     if ((ret = write_loose_object(oid, hdr, hdrlen, &in_stream, 0, 0)))\n> > +             die(_(\"failed to write object in stream %d\"), ret);\n> > +\n> > +     if (zstream.total_out != size || data.status != Z_STREAM_END)\n> > +             die(_(\"inflate returned %d\"), data.status);\n> > +     git_inflate_end(&zstream);\n> > +\n> > +     if (strict && !dry_run) {\n> > +             struct blob *blob = lookup_blob(the_repository, oid);\n> > +             if (blob)\n> > +                     blob->object.flags |= FLAG_WRITTEN;\n> > +             else\n> > +                     die(\"invalid blob object from stream\");\n> > +     }\n> > +     obj_list[nr].obj = NULL;\n> > +}\n>\n> Just a side-note, I think (but am not 100% sure) that these existing\n> occurances aren't needed due to our use of CALLOC_ARRAY():\n>\n>     diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n>     index 4a9466295ba..00b349412c5 100644\n>     --- a/builtin/unpack-objects.c\n>     +++ b/builtin/unpack-objects.c\n>     @@ -248,7 +248,6 @@ static void write_object(unsigned nr, enum object_type type,\n>                             die(\"failed to write object\");\n>                     added_object(nr, type, buf, size);\n>                     free(buf);\n>     -               obj_list[nr].obj = NULL;\n>             } else if (type == OBJ_BLOB) {\n>                     struct blob *blob;\n>                     if (write_object_file(buf, size, type_name(type),\n>     @@ -262,7 +261,6 @@ static void write_object(unsigned nr, enum object_type type,\n>                             blob->object.flags |= FLAG_WRITTEN;\n>                     else\n>                             die(\"invalid blob object\");\n>     -               obj_list[nr].obj = NULL;\n>             } else {\n>                     struct object *obj;\n>                     int eaten;\n>\n> The reason I'm noting it is that the same seems to be true of your new\n> addition here. I.e. are these assignments to NULL needed?\n>\n> Anyway, the reason I started poking at this it tha this\n> write_stream_blob() seems to duplicate much of write_object(). AFAICT\n> only the writing part is really different, the part where we\n> lookup_blob() after, set FLAG_WRITTEN etc. is all the same.\n>\n> Why can't we call write_object() here?\n>\n> The obvious answer seems to be that the call to write_object_file()\n> isn't prepared to do the sort of streaming that you want, so instead\n> you're bypassing it and calling write_loose_object() directly.\n>\n> I haven't tried this myself, but isn't a better and cleaner approach\n> here to not add another meaning to what is_null_oid() means, but to just\n> add a HASH_STREAM flag that'll get passed down as \"unsigned flags\" to\n> write_loose_object()? See FLAG_BITS in object.h.\n>\n> Then the \"obj_list[nr].obj\" here could also become\n> \"obj_list[nr].obj.flags |= (1u<<12)\" or whatever (but that wouldn't\n> strictly be needed I think.\n>\n> But by adding the \"HASH_STREAM\" flag you could I think stop duplicating\n> the \"Generate the header\" etc. here and call write_object_file_flags().\n>\n> I don't so much care about how it's done within unpack-objects.c, but\n> not having another meaning to is_null_oid() in play would be really\n> nice, and it this case it seems entirely avoidable.\n\nI did refactor it according to your suggestions in my next patch version.\nUsing a HASH_STREAM tag is indeed a better way to deal with it, and it\ncan also reduce my refactor to the original contents.\n\nThanks.\n-Han Xin\n"},{"id":"443292","messageId":"8594ac5f-8c09-6959-2bc7-208f6d888b4b@gmail.com","threadId":"56672","inReplyTo":"20211203093530.93589-1-chiyutianyi@gmail.com","subject":"Re: [PATCH v4 0/5] unpack large objects in stream","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2021-12-07T16:18:02Z","receivedAt":"2021-12-07T16:18:11Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 12/3/2021 4:35 AM, Han Xin wrote:\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n> \n> Changes since v3:\n> * Add \"size\" to \"struct input_stream\" which used by following commits.\n> \n> * Increase the buffer size of \"struct input_zstream_data\" from 4096 to\n>   8192, which is consistent with the \"fixed_buf\" in the \"index-pack.c\".\n> \n> * Refactor \"read stream in a loop in write_loose_object()\" which\n>   introduced a performance problem reported by Derrick Stolee[1].\n\nThank you for finding the issue. It seems simple enough to add that size\ninformation and regain the performance back to nearly no overhead. Your\nhyperfine statistics are within noise, which is great. Thanks!\n\n-Stolee\n"},{"id":"443777","messageId":"20211210103435.83656-1-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211203093530.93589-1-chiyutianyi@gmail.com","subject":"[PATCH v5 0/6] unpack large blobs in stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-10T10:34:29Z","receivedAt":"2021-12-10T10:35:01Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nChanges since v4:\n* Refactor to \"struct input_stream\" implementations so that we can\n  reduce the changes to \"write_loose_object()\" sugguest by\n  Ævar Arnfjörð Bjarmason.\n\n* Add a new flag called \"HASH_STREAM\" to support this feature.\n\n* Add a new config \"core.bigFileStreamingThreshold\" instread of\n  \"core.bigFileThreshold\" sugguest by Ævar Arnfjörð Bjarmason[1].\n\n* Roll destination repository preparement into a function in \n  \"t5590-unpack-non-delta-objects.sh\", so that we can run testcases\n  with --run=setup,3,4.\n\n1. https://lore.kernel.org/git/211203.86zgphsu5a.gmgdl@evledraar.gmail.com/\n\nHan Xin (6):\n  object-file: refactor write_loose_object() to support read from stream\n  object-file.c: handle undetermined oid in write_loose_object()\n  object-file.c: read stream in a loop in write_loose_object()\n  unpack-objects.c: add dry_run mode for get_data()\n  object-file.c: make \"write_object_file_flags()\" to support \"HASH_STREAM\"\n  unpack-objects: unpack_non_delta_entry() read data in a stream\n\n Documentation/config/core.txt       | 11 ++++\n builtin/unpack-objects.c            | 86 +++++++++++++++++++++++++++--\n cache.h                             |  2 +\n config.c                            |  5 ++\n environment.c                       |  1 +\n object-file.c                       | 73 +++++++++++++++++++-----\n object-store.h                      |  5 ++\n t/t5590-unpack-non-delta-objects.sh | 70 +++++++++++++++++++++++\n 8 files changed, 234 insertions(+), 19 deletions(-)\n create mode 100755 t/t5590-unpack-non-delta-objects.sh\n\nRange-diff against v4:\n1:  af707ef304 < -:  ---------- object-file: refactor write_loose_object() to read buffer from stream\n2:  321ad90d8e < -:  ---------- object-file.c: handle undetermined oid in write_loose_object()\n3:  1992ac39af < -:  ---------- object-file.c: read stream in a loop in write_loose_object()\n-:  ---------- > 1:  f3595e68cc object-file: refactor write_loose_object() to support read from stream\n-:  ---------- > 2:  c25fdd1fe5 object-file.c: handle undetermined oid in write_loose_object()\n-:  ---------- > 3:  ed226f2f9f object-file.c: read stream in a loop in write_loose_object()\n4:  c41eb06533 ! 4:  2f91e540f6 unpack-objects.c: add dry_run mode for get_data()\n    @@ builtin/unpack-objects.c: static void use(int bytes)\n      {\n      \tgit_zstream stream;\n     -\tvoid *buf = xmallocz(size);\n    -+\tunsigned long bufsize = dry_run ? 4096 : size;\n    ++\tunsigned long bufsize = dry_run ? 8192 : size;\n     +\tvoid *buf = xmallocz(bufsize);\n      \n      \tmemset(&stream, 0, sizeof(stream));\n-:  ---------- > 5:  7698938eac object-file.c: make \"write_object_file_flags()\" to support \"HASH_STREAM\"\n5:  9427775bdc ! 6:  103bb1db06 unpack-objects: unpack_non_delta_entry() read data in a stream\n    @@ Commit message\n     \n         However, unpack non-delta objects from a stream instead of from an entrie\n         buffer will have 10% performance penalty. Therefore, only unpack object\n    -    larger than the \"big_file_threshold\" in zstream. See the following\n    +    larger than the \"core.BigFileStreamingThreshold\" in zstream. See the following\n         benchmarks:\n     \n             hyperfine \\\n               --setup \\\n               'if ! test -d scalar.git; then git clone --bare https://github.com/microsoft/scalar.git; cp scalar.git/objects/pack/*.pack small.pack; fi' \\\n    -          --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n    -          -n 'old' 'git -C dest.git unpack-objects <small.pack' \\\n    -          -n 'new' 'new/git -C dest.git unpack-objects <small.pack' \\\n    -          -n 'new (small threshold)' \\\n    -          'new/git -c core.bigfilethreshold=16k -C dest.git unpack-objects <small.pack'\n    -        Benchmark 1: old\n    -          Time (mean ± σ):      6.075 s ±  0.069 s    [User: 5.047 s, System: 0.991 s]\n    -          Range (min … max):    6.018 s …  6.189 s    10 runs\n    -\n    -        Benchmark 2: new\n    -          Time (mean ± σ):      6.090 s ±  0.033 s    [User: 5.075 s, System: 0.976 s]\n    -          Range (min … max):    6.030 s …  6.142 s    10 runs\n    -\n    -        Benchmark 3: new (small threshold)\n    -          Time (mean ± σ):      6.755 s ±  0.029 s    [User: 5.150 s, System: 1.560 s]\n    -          Range (min … max):    6.711 s …  6.809 s    10 runs\n    +          --prepare 'rm -rf dest.git && git init --bare dest.git'\n     \n             Summary\n    -          'old' ran\n    -            1.00 ± 0.01 times faster than 'new'\n    -            1.11 ± 0.01 times faster than 'new (small threshold)'\n    +          './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'origin/master'\n    +            1.01 ± 0.04 times faster than './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'HEAD~1'\n    +            1.01 ± 0.04 times faster than './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'HEAD~0'\n    +            1.03 ± 0.10 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'origin/master'\n    +            1.02 ± 0.07 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'HEAD~0'\n    +            1.10 ± 0.04 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'HEAD~1'\n     \n    +    Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n         Helped-by: Derrick Stolee <stolee@gmail.com>\n         Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n         Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n     \n    + ## Documentation/config/core.txt ##\n    +@@ Documentation/config/core.txt: be delta compressed, but larger binary media files won't be.\n    + +\n    + Common unit suffixes of 'k', 'm', or 'g' are supported.\n    + \n    ++core.bigFileStreamingThreshold::\n    ++\tFiles larger than this will be streamed out to a temporary\n    ++\tobject file while being hashed, which will when be renamed\n    ++\tin-place to a loose object, particularly if the\n    ++\t`core.bigFileThreshold' setting dictates that they're always\n    ++\twritten out as loose objects.\n    +++\n    ++Default is 128 MiB on all platforms.\n    +++\n    ++Common unit suffixes of 'k', 'm', or 'g' are supported.\n    ++\n    + core.excludesFile::\n    + \tSpecifies the pathname to the file that contains patterns to\n    + \tdescribe paths that are not meant to be tracked, in addition\n    +\n      ## builtin/unpack-objects.c ##\n     @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type type,\n      \t}\n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n     +\n     +static void write_stream_blob(unsigned nr, unsigned long size)\n     +{\n    -+\tchar hdr[32];\n    -+\tint hdrlen;\n     +\tgit_zstream zstream;\n     +\tstruct input_zstream_data data;\n     +\tstruct input_stream in_stream = {\n     +\t\t.read = feed_input_zstream,\n     +\t\t.data = &data,\n    -+\t\t.size = size,\n     +\t};\n    -+\tstruct object_id *oid = &obj_list[nr].oid;\n     +\tint ret;\n     +\n     +\tmemset(&zstream, 0, sizeof(zstream));\n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n     +\tdata.zstream = &zstream;\n     +\tgit_inflate_init(&zstream);\n     +\n    -+\t/* Generate the header */\n    -+\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX, type_name(OBJ_BLOB), (uintmax_t)size) + 1;\n    -+\n    -+\tif ((ret = write_loose_object(oid, hdr, hdrlen, &in_stream, 0, 0)))\n    ++\tif ((ret = write_object_file_flags(&in_stream, size, type_name(OBJ_BLOB) ,&obj_list[nr].oid, HASH_STREAM)))\n     +\t\tdie(_(\"failed to write object in stream %d\"), ret);\n     +\n     +\tif (zstream.total_out != size || data.status != Z_STREAM_END)\n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n     +\tgit_inflate_end(&zstream);\n     +\n     +\tif (strict && !dry_run) {\n    -+\t\tstruct blob *blob = lookup_blob(the_repository, oid);\n    ++\t\tstruct blob *blob = lookup_blob(the_repository, &obj_list[nr].oid);\n     +\t\tif (blob)\n     +\t\t\tblob->object.flags |= FLAG_WRITTEN;\n     +\t\telse\n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n     +\tvoid *buf;\n     +\n     +\t/* Write large blob in stream without allocating full buffer. */\n    -+\tif (!dry_run && type == OBJ_BLOB && size > big_file_threshold) {\n    ++\tif (!dry_run && type == OBJ_BLOB && size > big_file_streaming_threshold) {\n     +\t\twrite_stream_blob(nr, size);\n     +\t\treturn;\n     +\t}\n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n      \t\twrite_object(nr, type, buf, size);\n      \telse\n     \n    - ## object-file.c ##\n    -@@ object-file.c: static const void *feed_simple_input_stream(struct input_stream *in_stream, unsi\n    - \treturn data->buf;\n    - }\n    + ## cache.h ##\n    +@@ cache.h: extern size_t packed_git_window_size;\n    + extern size_t packed_git_limit;\n    + extern size_t delta_base_cache_limit;\n    + extern unsigned long big_file_threshold;\n    ++extern unsigned long big_file_streaming_threshold;\n    + extern unsigned long pack_size_limit_cfg;\n      \n    --static int write_loose_object(const struct object_id *oid, char *hdr,\n    --\t\t\t      int hdrlen, struct input_stream *in_stream,\n    --\t\t\t      time_t mtime, unsigned flags)\n    -+int write_loose_object(const struct object_id *oid, char *hdr,\n    -+\t\t       int hdrlen, struct input_stream *in_stream,\n    -+\t\t       time_t mtime, unsigned flags)\n    - {\n    - \tint fd, ret;\n    - \tunsigned char compressed[4096];\n    + /*\n     \n    - ## object-store.h ##\n    -@@ object-store.h: int hash_object_file(const struct git_hash_algo *algo, const void *buf,\n    - \t\t     unsigned long len, const char *type,\n    - \t\t     struct object_id *oid);\n    + ## config.c ##\n    +@@ config.c: static int git_default_core_config(const char *var, const char *value, void *cb)\n    + \t\treturn 0;\n    + \t}\n      \n    -+int write_loose_object(const struct object_id *oid, char *hdr,\n    -+\t\t       int hdrlen, struct input_stream *in_stream,\n    -+\t\t       time_t mtime, unsigned flags);\n    ++\tif (!strcmp(var, \"core.bigfilestreamingthreshold\")) {\n    ++\t\tbig_file_streaming_threshold = git_config_ulong(var, value);\n    ++\t\treturn 0;\n    ++\t}\n     +\n    - int write_object_file_flags(const void *buf, unsigned long len,\n    - \t\t\t    const char *type, struct object_id *oid,\n    - \t\t\t    unsigned flags);\n    + \tif (!strcmp(var, \"core.packedgitlimit\")) {\n    + \t\tpacked_git_limit = git_config_ulong(var, value);\n    + \t\treturn 0;\n    +\n    + ## environment.c ##\n    +@@ environment.c: size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n    + size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n    + size_t delta_base_cache_limit = 96 * 1024 * 1024;\n    + unsigned long big_file_threshold = 512 * 1024 * 1024;\n    ++unsigned long big_file_streaming_threshold = 128 * 1024 * 1024;\n    + int pager_use_color = 1;\n    + const char *editor_program;\n    + const char *askpass_program;\n     \n      ## t/t5590-unpack-non-delta-objects.sh (new) ##\n     @@\n    @@ t/t5590-unpack-non-delta-objects.sh (new)\n     +\n     +. ./test-lib.sh\n     +\n    -+test_expect_success \"create commit with big blobs (1.5 MB)\" '\n    ++prepare_dest () {\n    ++\ttest_when_finished \"rm -rf dest.git\" &&\n    ++\tgit init --bare dest.git &&\n    ++\tgit -C dest.git config core.bigFileStreamingThreshold $1\n    ++\tgit -C dest.git config core.bigFileThreshold $1\n    ++}\n    ++\n    ++test_expect_success \"setup repo with big blobs (1.5 MB)\" '\n     +\ttest-tool genrandom foo 1500000 >big-blob &&\n     +\ttest_commit --append foo big-blob &&\n     +\ttest-tool genrandom bar 1500000 >big-blob &&\n    @@ t/t5590-unpack-non-delta-objects.sh (new)\n     +\t\tcd .git &&\n     +\t\tfind objects/?? -type f | sort\n     +\t) >expect &&\n    -+\tPACK=$(echo main | git pack-objects --progress --revs test)\n    ++\tPACK=$(echo main | git pack-objects --revs test)\n     +'\n     +\n    -+test_expect_success 'setup GIT_ALLOC_LIMIT to 1MB' '\n    ++test_expect_success 'setup env: GIT_ALLOC_LIMIT to 1MB' '\n     +\tGIT_ALLOC_LIMIT=1m &&\n     +\texport GIT_ALLOC_LIMIT\n     +'\n     +\n    -+test_expect_success 'prepare dest repository' '\n    -+\tgit init --bare dest.git &&\n    -+\tgit -C dest.git config core.bigFileThreshold 2m &&\n    -+\tgit -C dest.git config receive.unpacklimit 100\n    -+'\n    -+\n     +test_expect_success 'fail to unpack-objects: cannot allocate' '\n    ++\tprepare_dest 2m &&\n     +\ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n    -+\ttest_i18ngrep \"fatal: attempting to allocate\" err &&\n    ++\tgrep \"fatal: attempting to allocate\" err &&\n     +\t(\n     +\t\tcd dest.git &&\n     +\t\tfind objects/?? -type f | sort\n     +\t) >actual &&\n    ++\ttest_file_not_empty actual &&\n     +\t! test_cmp expect actual\n     +'\n     +\n    -+test_expect_success 'set a lower bigfile threshold' '\n    -+\tgit -C dest.git config core.bigFileThreshold 1m\n    -+'\n    -+\n     +test_expect_success 'unpack big object in stream' '\n    ++\tprepare_dest 1m &&\n     +\tgit -C dest.git unpack-objects <test-$PACK.pack &&\n     +\tgit -C dest.git fsck &&\n     +\t(\n    @@ t/t5590-unpack-non-delta-objects.sh (new)\n     +\ttest_cmp expect actual\n     +'\n     +\n    -+test_expect_success 'setup for unpack-objects dry-run test' '\n    -+\tgit init --bare unpack-test.git\n    -+'\n    -+\n     +test_expect_success 'unpack-objects dry-run' '\n    ++\tprepare_dest 1m &&\n    ++\tgit -C dest.git unpack-objects -n <test-$PACK.pack &&\n     +\t(\n    -+\t\tcd unpack-test.git &&\n    -+\t\tgit unpack-objects -n <../test-$PACK.pack\n    -+\t) &&\n    -+\t(\n    -+\t\tcd unpack-test.git &&\n    ++\t\tcd dest.git &&\n     +\t\tfind objects/ -type f\n     +\t) >actual &&\n     +\ttest_must_be_empty actual\n-- \n2.34.0\n\n"},{"id":"443778","messageId":"20211210103435.83656-2-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211203093530.93589-1-chiyutianyi@gmail.com","subject":"[PATCH v5 1/6] object-file: refactor write_loose_object() to support read from stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-10T10:34:30Z","receivedAt":"2021-12-10T10:35:06Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWe used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\nentire contents of a blob object, no matter how big it is. This\nimplementation may consume all the memory and cause OOM.\n\nThis can be improved by feeding data to \"write_loose_object()\" in a\nstream. The input stream is implemented as an interface.\n\nIn the first step, we add a new flag called \"HASH_STREAM\" and make a\nsimple implementation, feeding the entire buffer in the stream to\n\"write_loose_object()\" as a refactor.\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n cache.h        | 1 +\n object-file.c  | 7 ++++++-\n object-store.h | 5 +++++\n 3 files changed, 12 insertions(+), 1 deletion(-)\n\ndiff --git a/cache.h b/cache.h\nindex eba12487b9..51bd435dea 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -888,6 +888,7 @@ int ie_modified(struct index_state *, const struct cache_entry *, struct stat *,\n #define HASH_FORMAT_CHECK 2\n #define HASH_RENORMALIZE  4\n #define HASH_SILENT 8\n+#define HASH_STREAM 16\n int index_fd(struct index_state *istate, struct object_id *oid, int fd, struct stat *st, enum object_type type, const char *path, unsigned flags);\n int index_path(struct index_state *istate, struct object_id *oid, const char *path, struct stat *st, unsigned flags);\n \ndiff --git a/object-file.c b/object-file.c\nindex eb972cdccd..06375a90d6 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1898,7 +1898,12 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n \n \t/* Then the data itself.. */\n-\tstream.next_in = (void *)buf;\n+\tif (flags & HASH_STREAM) {\n+\t\tstruct input_stream *in_stream = (struct input_stream *)buf;\n+\t\tstream.next_in = (void *)in_stream->read(in_stream, &len);\n+\t} else {\n+\t\tstream.next_in = (void *)buf;\n+\t}\n \tstream.avail_in = len;\n \tdo {\n \t\tunsigned char *in0 = stream.next_in;\ndiff --git a/object-store.h b/object-store.h\nindex 952efb6a4b..ccc1fc9c1a 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -34,6 +34,11 @@ struct object_directory {\n \tchar *path;\n };\n \n+struct input_stream {\n+\tconst void *(*read)(struct input_stream *, unsigned long *len);\n+\tvoid *data;\n+};\n+\n KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n \tstruct object_directory *, 1, fspathhash, fspatheq)\n \n-- \n2.34.0\n\n"},{"id":"443779","messageId":"20211210103435.83656-3-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211203093530.93589-1-chiyutianyi@gmail.com","subject":"[PATCH v5 2/6] object-file.c: handle undetermined oid in write_loose_object()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-10T10:34:31Z","receivedAt":"2021-12-10T10:35:07Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWhen streaming a large blob object to \"write_loose_object()\", we have no\nchance to run \"write_object_file_prepare()\" to calculate the oid in\nadvance. So we need to handle undetermined oid in function\n\"write_loose_object()\".\n\nIn the original implementation, we know the oid and we can write the\ntemporary file in the same directory as the final object, but for an\nobject with an undetermined oid, we don't know the exact directory for\nthe object, so we have to save the temporary file in \".git/objects/\"\ndirectory instead.\n\nThe promise that \"oid\" is constant in \"write_loose_object()\" has been\nremoved because it will be filled after reading all stream data.\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c | 48 +++++++++++++++++++++++++++++++++++++++---------\n 1 file changed, 39 insertions(+), 9 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 06375a90d6..41099b137f 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1860,11 +1860,11 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n \treturn fd;\n }\n \n-static int write_loose_object(const struct object_id *oid, char *hdr,\n+static int write_loose_object(struct object_id *oid, char *hdr,\n \t\t\t      int hdrlen, const void *buf, unsigned long len,\n \t\t\t      time_t mtime, unsigned flags)\n {\n-\tint fd, ret;\n+\tint fd, ret, err = 0;\n \tunsigned char compressed[4096];\n \tgit_zstream stream;\n \tgit_hash_ctx c;\n@@ -1872,16 +1872,21 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tstatic struct strbuf tmp_file = STRBUF_INIT;\n \tstatic struct strbuf filename = STRBUF_INIT;\n \n-\tloose_object_path(the_repository, &filename, oid);\n+\tif (flags & HASH_STREAM)\n+\t\t/* When oid is not determined, save tmp file to odb path. */\n+\t\tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n+\telse\n+\t\tloose_object_path(the_repository, &filename, oid);\n \n \tfd = create_tmpfile(&tmp_file, filename.buf);\n \tif (fd < 0) {\n \t\tif (flags & HASH_SILENT)\n-\t\t\treturn -1;\n+\t\t\terr = -1;\n \t\telse if (errno == EACCES)\n-\t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n+\t\t\terr = error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n \t\telse\n-\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n+\t\t\terr = error_errno(_(\"unable to create temporary file\"));\n+\t\tgoto cleanup;\n \t}\n \n \t/* Set it up */\n@@ -1923,12 +1928,34 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n \t\t    ret);\n \tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n-\tif (!oideq(oid, &parano_oid))\n+\tif (!(flags & HASH_STREAM) && !oideq(oid, &parano_oid))\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n \n \tclose_loose_object(fd);\n \n+\tif (flags & HASH_STREAM) {\n+\t\tint dirlen;\n+\n+\t\toidcpy((struct object_id *)oid, &parano_oid);\n+\t\tloose_object_path(the_repository, &filename, oid);\n+\n+\t\t/* We finally know the object path, and create the missing dir. */\n+\t\tdirlen = directory_size(filename.buf);\n+\t\tif (dirlen) {\n+\t\t\tstruct strbuf dir = STRBUF_INIT;\n+\t\t\tstrbuf_add(&dir, filename.buf, dirlen - 1);\n+\t\t\tif (mkdir(dir.buf, 0777) && errno != EEXIST)\n+\t\t\t\terr = -1;\n+\t\t\telse if (adjust_shared_perm(dir.buf))\n+\t\t\t\terr = -1;\n+\t\t\telse\n+\t\t\t\tstrbuf_release(&dir);\n+\t\t\tif (err < 0)\n+\t\t\t\tgoto cleanup;\n+\t\t}\n+\t}\n+\n \tif (mtime) {\n \t\tstruct utimbuf utb;\n \t\tutb.actime = mtime;\n@@ -1938,7 +1965,10 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\t\twarning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n \t}\n \n-\treturn finalize_object_file(tmp_file.buf, filename.buf);\n+\terr = finalize_object_file(tmp_file.buf, filename.buf);\n+cleanup:\n+\tstrbuf_release(&filename);\n+\treturn err;\n }\n \n static int freshen_loose_object(const struct object_id *oid)\n@@ -2015,7 +2045,7 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n \tif (!buf)\n \t\treturn error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n \thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n-\tret = write_loose_object(oid, hdr, hdrlen, buf, len, mtime, 0);\n+\tret = write_loose_object((struct object_id*) oid, hdr, hdrlen, buf, len, mtime, 0);\n \tfree(buf);\n \n \treturn ret;\n-- \n2.34.0\n\n"},{"id":"443780","messageId":"20211210103435.83656-4-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211203093530.93589-1-chiyutianyi@gmail.com","subject":"[PATCH v5 3/6] object-file.c: read stream in a loop in write_loose_object()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-10T10:34:32Z","receivedAt":"2021-12-10T10:35:08Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nIn order to prepare the stream version of \"write_loose_object()\", read\nthe input stream in a loop in \"write_loose_object()\", so that we can\nfeed the contents of large blob object to \"write_loose_object()\" using\na small fixed buffer.\n\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c | 23 +++++++++++++++--------\n 1 file changed, 15 insertions(+), 8 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 41099b137f..455ab3c06e 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1864,7 +1864,7 @@ static int write_loose_object(struct object_id *oid, char *hdr,\n \t\t\t      int hdrlen, const void *buf, unsigned long len,\n \t\t\t      time_t mtime, unsigned flags)\n {\n-\tint fd, ret, err = 0;\n+\tint fd, ret, err = 0, flush = 0;\n \tunsigned char compressed[4096];\n \tgit_zstream stream;\n \tgit_hash_ctx c;\n@@ -1903,22 +1903,29 @@ static int write_loose_object(struct object_id *oid, char *hdr,\n \tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n \n \t/* Then the data itself.. */\n-\tif (flags & HASH_STREAM) {\n-\t\tstruct input_stream *in_stream = (struct input_stream *)buf;\n-\t\tstream.next_in = (void *)in_stream->read(in_stream, &len);\n-\t} else {\n+\tif (!(flags & HASH_STREAM)) {\n \t\tstream.next_in = (void *)buf;\n+\t\tstream.avail_in = len;\n+\t\tflush = Z_FINISH;\n \t}\n-\tstream.avail_in = len;\n \tdo {\n \t\tunsigned char *in0 = stream.next_in;\n-\t\tret = git_deflate(&stream, Z_FINISH);\n+\t\tif (flags & HASH_STREAM && !stream.avail_in) {\n+\t\t\tstruct input_stream *in_stream = (struct input_stream *)buf;\n+\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n+\t\t\tstream.next_in = (void *)in;\n+\t\t\tin0 = (unsigned char *)in;\n+\t\t\t/* All data has been read. */\n+\t\t\tif (len + hdrlen == stream.total_in + stream.avail_in)\n+\t\t\t\tflush = Z_FINISH;\n+\t\t}\n+\t\tret = git_deflate(&stream, flush);\n \t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n \t\tif (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n \t\t\tdie(_(\"unable to write loose object file\"));\n \t\tstream.next_out = compressed;\n \t\tstream.avail_out = sizeof(compressed);\n-\t} while (ret == Z_OK);\n+\t} while (ret == Z_OK || ret == Z_BUF_ERROR);\n \n \tif (ret != Z_STREAM_END)\n \t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n-- \n2.34.0\n\n"},{"id":"443781","messageId":"20211210103435.83656-5-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211203093530.93589-1-chiyutianyi@gmail.com","subject":"[PATCH v5 4/6] unpack-objects.c: add dry_run mode for get_data()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-10T10:34:33Z","receivedAt":"2021-12-10T10:35:11Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nIn dry_run mode, \"get_data()\" is used to verify the inflation of data,\nand the returned buffer will not be used at all and will be freed\nimmediately. Even in dry_run mode, it is dangerous to allocate a\nfull-size buffer for a large blob object. Therefore, only allocate a\nlow memory footprint when calling \"get_data()\" in dry_run mode.\n\nSuggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n builtin/unpack-objects.c | 18 ++++++++++++------\n 1 file changed, 12 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 4a9466295b..d878e2f8b4 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -96,15 +96,16 @@ static void use(int bytes)\n \tdisplay_throughput(progress, consumed_bytes);\n }\n \n-static void *get_data(unsigned long size)\n+static void *get_data(unsigned long size, int dry_run)\n {\n \tgit_zstream stream;\n-\tvoid *buf = xmallocz(size);\n+\tunsigned long bufsize = dry_run ? 8192 : size;\n+\tvoid *buf = xmallocz(bufsize);\n \n \tmemset(&stream, 0, sizeof(stream));\n \n \tstream.next_out = buf;\n-\tstream.avail_out = size;\n+\tstream.avail_out = bufsize;\n \tstream.next_in = fill(1);\n \tstream.avail_in = len;\n \tgit_inflate_init(&stream);\n@@ -124,6 +125,11 @@ static void *get_data(unsigned long size)\n \t\t}\n \t\tstream.next_in = fill(1);\n \t\tstream.avail_in = len;\n+\t\tif (dry_run) {\n+\t\t\t/* reuse the buffer in dry_run mode */\n+\t\t\tstream.next_out = buf;\n+\t\t\tstream.avail_out = bufsize;\n+\t\t}\n \t}\n \tgit_inflate_end(&stream);\n \treturn buf;\n@@ -323,7 +329,7 @@ static void added_object(unsigned nr, enum object_type type,\n static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n \t\t\t\t   unsigned nr)\n {\n-\tvoid *buf = get_data(size);\n+\tvoid *buf = get_data(size, dry_run);\n \n \tif (!dry_run && buf)\n \t\twrite_object(nr, type, buf, size);\n@@ -357,7 +363,7 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \tif (type == OBJ_REF_DELTA) {\n \t\toidread(&base_oid, fill(the_hash_algo->rawsz));\n \t\tuse(the_hash_algo->rawsz);\n-\t\tdelta_data = get_data(delta_size);\n+\t\tdelta_data = get_data(delta_size, dry_run);\n \t\tif (dry_run || !delta_data) {\n \t\t\tfree(delta_data);\n \t\t\treturn;\n@@ -396,7 +402,7 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\tif (base_offset <= 0 || base_offset >= obj_list[nr].offset)\n \t\t\tdie(\"offset value out of bound for delta base object\");\n \n-\t\tdelta_data = get_data(delta_size);\n+\t\tdelta_data = get_data(delta_size, dry_run);\n \t\tif (dry_run || !delta_data) {\n \t\t\tfree(delta_data);\n \t\t\treturn;\n-- \n2.34.0\n\n"},{"id":"443782","messageId":"20211210103435.83656-6-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211203093530.93589-1-chiyutianyi@gmail.com","subject":"[PATCH v5 5/6] object-file.c: make \"write_object_file_flags()\" to support \"HASH_STREAM\"","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-10T10:34:34Z","receivedAt":"2021-12-10T10:35:14Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWe will use \"write_object_file_flags()\" in \"unpack_non_delta_entry()\" to\nread the entire data contents in stream. When read in stream, we needn't\nprepare \"oid\" before \"write_loose_object()\", only generate the header.\n\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c | 5 +++++\n 1 file changed, 5 insertions(+)\n\ndiff --git a/object-file.c b/object-file.c\nindex 455ab3c06e..906590dae5 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2002,6 +2002,11 @@ int write_object_file_flags(const void *buf, unsigned long len,\n {\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen = sizeof(hdr);\n+\tif (flags & HASH_STREAM) {\n+\t\t/* Generate the header */\n+\t\thdrlen = xsnprintf(hdr, hdrlen, \"%s %\"PRIuMAX , type, (uintmax_t)len)+1;\n+\t\treturn write_loose_object(oid, hdr, hdrlen, buf, len, 0, flags);\n+\t}\n \n \t/* Normally if we have it in the pack then we do not bother writing\n \t * it out into .git/objects/??/?{38} file.\n-- \n2.34.0\n\n"},{"id":"443783","messageId":"20211210103435.83656-7-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211203093530.93589-1-chiyutianyi@gmail.com","subject":"[PATCH v5 6/6] unpack-objects: unpack_non_delta_entry() read data in a stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-10T10:34:35Z","receivedAt":"2021-12-10T10:35:18Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWe used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\nentire contents of a blob object, no matter how big it is. This\nimplementation may consume all the memory and cause OOM.\n\nBy implementing a zstream version of input_stream interface, we can use\na small fixed buffer for \"unpack_non_delta_entry()\".\n\nHowever, unpack non-delta objects from a stream instead of from an entrie\nbuffer will have 10% performance penalty. Therefore, only unpack object\nlarger than the \"core.BigFileStreamingThreshold\" in zstream. See the following\nbenchmarks:\n\n    hyperfine \\\n      --setup \\\n      'if ! test -d scalar.git; then git clone --bare https://github.com/microsoft/scalar.git; cp scalar.git/objects/pack/*.pack small.pack; fi' \\\n      --prepare 'rm -rf dest.git && git init --bare dest.git'\n\n    Summary\n      './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'origin/master'\n        1.01 ± 0.04 times faster than './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'HEAD~1'\n        1.01 ± 0.04 times faster than './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'HEAD~0'\n        1.03 ± 0.10 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'origin/master'\n        1.02 ± 0.07 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'HEAD~0'\n        1.10 ± 0.04 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'HEAD~1'\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Derrick Stolee <stolee@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n Documentation/config/core.txt       | 11 +++++\n builtin/unpack-objects.c            | 70 ++++++++++++++++++++++++++++-\n cache.h                             |  1 +\n config.c                            |  5 +++\n environment.c                       |  1 +\n t/t5590-unpack-non-delta-objects.sh | 70 +++++++++++++++++++++++++++++\n 6 files changed, 157 insertions(+), 1 deletion(-)\n create mode 100755 t/t5590-unpack-non-delta-objects.sh\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex c04f62a54a..601b7a2418 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -424,6 +424,17 @@ be delta compressed, but larger binary media files won't be.\n +\n Common unit suffixes of 'k', 'm', or 'g' are supported.\n \n+core.bigFileStreamingThreshold::\n+\tFiles larger than this will be streamed out to a temporary\n+\tobject file while being hashed, which will when be renamed\n+\tin-place to a loose object, particularly if the\n+\t`core.bigFileThreshold' setting dictates that they're always\n+\twritten out as loose objects.\n++\n+Default is 128 MiB on all platforms.\n++\n+Common unit suffixes of 'k', 'm', or 'g' are supported.\n+\n core.excludesFile::\n \tSpecifies the pathname to the file that contains patterns to\n \tdescribe paths that are not meant to be tracked, in addition\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex d878e2f8b4..0df115ab0d 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -326,11 +326,79 @@ static void added_object(unsigned nr, enum object_type type,\n \t}\n }\n \n+struct input_zstream_data {\n+\tgit_zstream *zstream;\n+\tunsigned char buf[8192];\n+\tint status;\n+};\n+\n+static const void *feed_input_zstream(struct input_stream *in_stream, unsigned long *readlen)\n+{\n+\tstruct input_zstream_data *data = in_stream->data;\n+\tgit_zstream *zstream = data->zstream;\n+\tvoid *in = fill(1);\n+\n+\tif (!len || data->status == Z_STREAM_END) {\n+\t\t*readlen = 0;\n+\t\treturn NULL;\n+\t}\n+\n+\tzstream->next_out = data->buf;\n+\tzstream->avail_out = sizeof(data->buf);\n+\tzstream->next_in = in;\n+\tzstream->avail_in = len;\n+\n+\tdata->status = git_inflate(zstream, 0);\n+\tuse(len - zstream->avail_in);\n+\t*readlen = sizeof(data->buf) - zstream->avail_out;\n+\n+\treturn data->buf;\n+}\n+\n+static void write_stream_blob(unsigned nr, unsigned long size)\n+{\n+\tgit_zstream zstream;\n+\tstruct input_zstream_data data;\n+\tstruct input_stream in_stream = {\n+\t\t.read = feed_input_zstream,\n+\t\t.data = &data,\n+\t};\n+\tint ret;\n+\n+\tmemset(&zstream, 0, sizeof(zstream));\n+\tmemset(&data, 0, sizeof(data));\n+\tdata.zstream = &zstream;\n+\tgit_inflate_init(&zstream);\n+\n+\tif ((ret = write_object_file_flags(&in_stream, size, type_name(OBJ_BLOB) ,&obj_list[nr].oid, HASH_STREAM)))\n+\t\tdie(_(\"failed to write object in stream %d\"), ret);\n+\n+\tif (zstream.total_out != size || data.status != Z_STREAM_END)\n+\t\tdie(_(\"inflate returned %d\"), data.status);\n+\tgit_inflate_end(&zstream);\n+\n+\tif (strict && !dry_run) {\n+\t\tstruct blob *blob = lookup_blob(the_repository, &obj_list[nr].oid);\n+\t\tif (blob)\n+\t\t\tblob->object.flags |= FLAG_WRITTEN;\n+\t\telse\n+\t\t\tdie(\"invalid blob object from stream\");\n+\t}\n+\tobj_list[nr].obj = NULL;\n+}\n+\n static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n \t\t\t\t   unsigned nr)\n {\n-\tvoid *buf = get_data(size, dry_run);\n+\tvoid *buf;\n+\n+\t/* Write large blob in stream without allocating full buffer. */\n+\tif (!dry_run && type == OBJ_BLOB && size > big_file_streaming_threshold) {\n+\t\twrite_stream_blob(nr, size);\n+\t\treturn;\n+\t}\n \n+\tbuf = get_data(size, dry_run);\n \tif (!dry_run && buf)\n \t\twrite_object(nr, type, buf, size);\n \telse\ndiff --git a/cache.h b/cache.h\nindex 51bd435dea..78548cd67a 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -965,6 +965,7 @@ extern size_t packed_git_window_size;\n extern size_t packed_git_limit;\n extern size_t delta_base_cache_limit;\n extern unsigned long big_file_threshold;\n+extern unsigned long big_file_streaming_threshold;\n extern unsigned long pack_size_limit_cfg;\n \n /*\ndiff --git a/config.c b/config.c\nindex c5873f3a70..7b122a142a 100644\n--- a/config.c\n+++ b/config.c\n@@ -1408,6 +1408,11 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"core.bigfilestreamingthreshold\")) {\n+\t\tbig_file_streaming_threshold = git_config_ulong(var, value);\n+\t\treturn 0;\n+\t}\n+\n \tif (!strcmp(var, \"core.packedgitlimit\")) {\n \t\tpacked_git_limit = git_config_ulong(var, value);\n \t\treturn 0;\ndiff --git a/environment.c b/environment.c\nindex 9da7f3c1a1..4fcc3de741 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -46,6 +46,7 @@ size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\n unsigned long big_file_threshold = 512 * 1024 * 1024;\n+unsigned long big_file_streaming_threshold = 128 * 1024 * 1024;\n int pager_use_color = 1;\n const char *editor_program;\n const char *askpass_program;\ndiff --git a/t/t5590-unpack-non-delta-objects.sh b/t/t5590-unpack-non-delta-objects.sh\nnew file mode 100755\nindex 0000000000..ff4c78900b\n--- /dev/null\n+++ b/t/t5590-unpack-non-delta-objects.sh\n@@ -0,0 +1,70 @@\n+#!/bin/sh\n+#\n+# Copyright (c) 2021 Han Xin\n+#\n+\n+test_description='Test unpack-objects when receive pack'\n+\n+GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n+export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n+\n+. ./test-lib.sh\n+\n+prepare_dest () {\n+\ttest_when_finished \"rm -rf dest.git\" &&\n+\tgit init --bare dest.git &&\n+\tgit -C dest.git config core.bigFileStreamingThreshold $1\n+\tgit -C dest.git config core.bigFileThreshold $1\n+}\n+\n+test_expect_success \"setup repo with big blobs (1.5 MB)\" '\n+\ttest-tool genrandom foo 1500000 >big-blob &&\n+\ttest_commit --append foo big-blob &&\n+\ttest-tool genrandom bar 1500000 >big-blob &&\n+\ttest_commit --append bar big-blob &&\n+\t(\n+\t\tcd .git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >expect &&\n+\tPACK=$(echo main | git pack-objects --revs test)\n+'\n+\n+test_expect_success 'setup env: GIT_ALLOC_LIMIT to 1MB' '\n+\tGIT_ALLOC_LIMIT=1m &&\n+\texport GIT_ALLOC_LIMIT\n+'\n+\n+test_expect_success 'fail to unpack-objects: cannot allocate' '\n+\tprepare_dest 2m &&\n+\ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n+\tgrep \"fatal: attempting to allocate\" err &&\n+\t(\n+\t\tcd dest.git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >actual &&\n+\ttest_file_not_empty actual &&\n+\t! test_cmp expect actual\n+'\n+\n+test_expect_success 'unpack big object in stream' '\n+\tprepare_dest 1m &&\n+\tgit -C dest.git unpack-objects <test-$PACK.pack &&\n+\tgit -C dest.git fsck &&\n+\t(\n+\t\tcd dest.git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'unpack-objects dry-run' '\n+\tprepare_dest 1m &&\n+\tgit -C dest.git unpack-objects -n <test-$PACK.pack &&\n+\t(\n+\t\tcd dest.git &&\n+\t\tfind objects/ -type f\n+\t) >actual &&\n+\ttest_must_be_empty actual\n+'\n+\n+test_done\n-- \n2.34.0\n\n"},{"id":"443963","messageId":"211213.86bl1l9bfz.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"20211210103435.83656-3-chiyutianyi@gmail.com","subject":"Re: [PATCH v5 2/6] object-file.c: handle undetermined oid in write_loose_object()","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-13T07:32:58Z","receivedAt":"2021-12-13T08:05:40Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, Dec 10 2021, Han Xin wrote:\n\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> When streaming a large blob object to \"write_loose_object()\", we have no\n> chance to run \"write_object_file_prepare()\" to calculate the oid in\n> advance. So we need to handle undetermined oid in function\n> \"write_loose_object()\".\n>\n> In the original implementation, we know the oid and we can write the\n> temporary file in the same directory as the final object, but for an\n> object with an undetermined oid, we don't know the exact directory for\n> the object, so we have to save the temporary file in \".git/objects/\"\n> directory instead.\n>\n> The promise that \"oid\" is constant in \"write_loose_object()\" has been\n> removed because it will be filled after reading all stream data.\n>\n> Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  object-file.c | 48 +++++++++++++++++++++++++++++++++++++++---------\n>  1 file changed, 39 insertions(+), 9 deletions(-)\n>\n> diff --git a/object-file.c b/object-file.c\n> index 06375a90d6..41099b137f 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1860,11 +1860,11 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n>  \treturn fd;\n>  }\n>  \n> -static int write_loose_object(const struct object_id *oid, char *hdr,\n> +static int write_loose_object(struct object_id *oid, char *hdr,\n>  \t\t\t      int hdrlen, const void *buf, unsigned long len,\n>  \t\t\t      time_t mtime, unsigned flags)\n>  {\n> -\tint fd, ret;\n> +\tint fd, ret, err = 0;\n>  \tunsigned char compressed[4096];\n>  \tgit_zstream stream;\n>  \tgit_hash_ctx c;\n> @@ -1872,16 +1872,21 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \tstatic struct strbuf tmp_file = STRBUF_INIT;\n>  \tstatic struct strbuf filename = STRBUF_INIT;\n>  \n> -\tloose_object_path(the_repository, &filename, oid);\n> +\tif (flags & HASH_STREAM)\n> +\t\t/* When oid is not determined, save tmp file to odb path. */\n> +\t\tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n> +\telse\n> +\t\tloose_object_path(the_repository, &filename, oid);\n>  \n>  \tfd = create_tmpfile(&tmp_file, filename.buf);\n>  \tif (fd < 0) {\n>  \t\tif (flags & HASH_SILENT)\n> -\t\t\treturn -1;\n> +\t\t\terr = -1;\n>  \t\telse if (errno == EACCES)\n> -\t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n> +\t\t\terr = error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n>  \t\telse\n> -\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n> +\t\t\terr = error_errno(_(\"unable to create temporary file\"));\n> +\t\tgoto cleanup;\n>  \t}\n>  \n>  \t/* Set it up */\n> @@ -1923,12 +1928,34 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n>  \t\t    ret);\n>  \tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n> -\tif (!oideq(oid, &parano_oid))\n> +\tif (!(flags & HASH_STREAM) && !oideq(oid, &parano_oid))\n>  \t\tdie(_(\"confused by unstable object source data for %s\"),\n>  \t\t    oid_to_hex(oid));\n\nHere we don't have a meaningful \"const\" OID anymore, but still if we die\nwe use the \"oid\". \n\n>  \tclose_loose_object(fd);\n>  \n> +\tif (flags & HASH_STREAM) {\n> +\t\tint dirlen;\n> +\n> +\t\toidcpy((struct object_id *)oid, &parano_oid);\n\nThis cast isn't needed anymore now that you stripped the \"const\" off,\nbut more on that later...\n\n> +\t\tloose_object_path(the_repository, &filename, oid);\n> +\n> +\t\t/* We finally know the object path, and create the missing dir. */\n> +\t\tdirlen = directory_size(filename.buf);\n> +\t\tif (dirlen) {\n> +\t\t\tstruct strbuf dir = STRBUF_INIT;\n> +\t\t\tstrbuf_add(&dir, filename.buf, dirlen - 1);\n> +\t\t\tif (mkdir(dir.buf, 0777) && errno != EEXIST)\n> +\t\t\t\terr = -1;\n> +\t\t\telse if (adjust_shared_perm(dir.buf))\n> +\t\t\t\terr = -1;\n> +\t\t\telse\n> +\t\t\t\tstrbuf_release(&dir);\n> +\t\t\tif (err < 0)\n> +\t\t\t\tgoto cleanup;\n\nCan't we use one of the existing utility functions for this? Testing\nlocally I could replace this with:\n\t\n\tdiff --git a/object-file.c b/object-file.c\n\tindex 7c93db11b2d..05e1fae893d 100644\n\t--- a/object-file.c\n\t+++ b/object-file.c\n\t@@ -1952,14 +1952,11 @@ static int write_loose_object(struct object_id *oid, char *hdr,\n\t \t\tif (dirlen) {\n\t \t\t\tstruct strbuf dir = STRBUF_INIT;\n\t \t\t\tstrbuf_add(&dir, filename.buf, dirlen - 1);\n\t-\t\t\tif (mkdir(dir.buf, 0777) && errno != EEXIST)\n\t+\t\t\t\n\t+\t\t\tif (mkdir_in_gitdir(dir.buf) < 0) {\n\t \t\t\t\terr = -1;\n\t-\t\t\telse if (adjust_shared_perm(dir.buf))\n\t-\t\t\t\terr = -1;\n\t-\t\t\telse\n\t-\t\t\t\tstrbuf_release(&dir);\n\t-\t\t\tif (err < 0)\n\t \t\t\t\tgoto cleanup;\n\t+\t\t\t}\n\t \t\t}\n\t \t}\n\nAnd your tests still pass. Maybe they have a blind spot, or maybe we can\njust use the existing function.\n\t \n> +\t\t}\n> +\t}\n> +\n>  \tif (mtime) {\n>  \t\tstruct utimbuf utb;\n>  \t\tutb.actime = mtime;\n> @@ -1938,7 +1965,10 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \t\t\twarning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n>  \t}\n>  \n> -\treturn finalize_object_file(tmp_file.buf, filename.buf);\n> +\terr = finalize_object_file(tmp_file.buf, filename.buf);\n> +cleanup:\n> +\tstrbuf_release(&filename);\n> +\treturn err;\n>  }\n\nReading this series is an odd mixture of of things that would really be\nmuch easier to understand if they were combined, e.g. 1/6 adding APIs\nthat aren't used by anything, but then adding one codepath (also\nunused), that we then use later. Could just add it at the same time as\nthe use and the patch would be easier to read....\n\n...and then this, which *is* something that could be split up into an\nearlier cleanup step, i.e. the strbuf leak here exists before this\nseries, fixing it is good, but splitting that up into its own patch\nwould make this diff smaller & the actual behavior changes easier to\nreason about.\n\n>  static int freshen_loose_object(const struct object_id *oid)\n> @@ -2015,7 +2045,7 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n>  \tif (!buf)\n>  \t\treturn error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n>  \thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n> -\tret = write_loose_object(oid, hdr, hdrlen, buf, len, mtime, 0);\n> +\tret = write_loose_object((struct object_id*) oid, hdr, hdrlen, buf, len, mtime, 0);\n>  \tfree(buf);\n>  \n>  \treturn ret;\n\n ...on the \"more on that later\", here we're casting the \"oid\" from const\n for a function that's never going to be involved in the streaming\n codepath.\n\nI know I suggested the HASH_STREAM flag, but what I was really going for\nwas \"let's share more of the code?\", looking at this v5 (which is\nalready much better than v4) I think a better approach is to split up\nwrite_loose_object().\n\nI.e. it already calls close_loose_object() and finalize_object_file() to\ndo some of its work, but around that we have:\n\n 1. Figuring out a path for the (temp) object file\n 2. Creating the tempfile\n 3. Setting up zlib\n 4. Once zlib is set up inspect its state, die with a message\n    about oid_to_hex(oid) if we failed\n 5. Optionally, do HASH_STREAM stuff\n    Maybe force a loose object if \"mtime\".\n\nI think if that's split up so that each of those is its own little\nfunction what's now write_loose_object() can call those in sequence, and\na new stream_loose_object() can just do #1 differentl, followed by the\nsame #2 and #4, but do #4 differently etc.\n\nYou'll still be able to re-use the write_object_file_prepare()\netc. logic.\n\nAs an example your 5/6 copy/pastes the xsnprintf() formatting of the\nobject header. It's just one line, but it's also code that's very\ncentral to git, so I think instead of just copy/pasting it a prep step\nof factoring it out would make sense, and that would be a prep cleanup\nthat would help later readability. E.g.:\n\t\n\tdiff --git a/object-file.c b/object-file.c\n\tindex eac67f6f5f9..a7dcbd929e9 100644\n\t--- a/object-file.c\n\t+++ b/object-file.c\n\t@@ -1009,6 +1009,13 @@ void *xmmap(void *start, size_t length,\n\t \treturn ret;\n\t }\n\t \n\t+static int generate_object_header(char *buf, int bufsz, const char *type_name,\n\t+\t\t\t\t  unsigned long size)\n\t+{\n\t+\treturn xsnprintf(buf, bufsz, \"%s %\"PRIuMAX , type_name,\n\t+\t\t\t (uintmax_t)size) + 1;\n\t+}\n\t+\n\t /*\n\t  * With an in-core object data in \"map\", rehash it to make sure the\n\t  * object name actually matches \"oid\" to detect object corruption.\n\t@@ -1037,7 +1044,7 @@ int check_object_signature(struct repository *r, const struct object_id *oid,\n\t \t\treturn -1;\n\t \n\t \t/* Generate the header */\n\t-\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(obj_type), (uintmax_t)size) + 1;\n\t+\thdrlen = generate_object_header(hdr, sizeof(hdr), type_name(obj_type), size);\n\t \n\t \t/* Sha1.. */\n\t \tr->hash_algo->init_fn(&c);\n\t@@ -1737,7 +1744,7 @@ static void write_object_file_prepare(const struct git_hash_algo *algo,\n\t \tgit_hash_ctx c;\n\t \n\t \t/* Generate the header */\n\t-\t*hdrlen = xsnprintf(hdr, *hdrlen, \"%s %\"PRIuMAX , type, (uintmax_t)len)+1;\n\t+\t*hdrlen = generate_object_header(hdr, *hdrlen, type, len);\n\t \n\t \t/* Sha1.. */\n\t \talgo->init_fn(&c);\n\t@@ -2009,7 +2016,7 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n\t \tbuf = read_object(the_repository, oid, &type, &len);\n\t \tif (!buf)\n\t \t\treturn error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n\t-\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n\t+\thdrlen = generate_object_header(hdr, sizeof(hdr), type_name(type), len);\n\t \tret = write_loose_object(oid, hdr, hdrlen, buf, len, mtime, 0);\n\t \tfree(buf);\n\nThen in your change on top you just call that generate_object_header(),\nor better yet your amended write_object_file_flags() can just call a\nsimilarly amended write_object_file_prepare() directly.\n"},{"id":"443965","messageId":"211213.867dc8ansq.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"20211210103435.83656-7-chiyutianyi@gmail.com","subject":"Re: [PATCH v5 6/6] unpack-objects: unpack_non_delta_entry() read data in a stream","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-13T08:05:55Z","receivedAt":"2021-12-13T08:53:29Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, Dec 10 2021, Han Xin wrote:\n\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n> [...]\n> +\tif ((ret = write_object_file_flags(&in_stream, size, type_name(OBJ_BLOB) ,&obj_list[nr].oid, HASH_STREAM)))\n\nThere's some odd code formatting here, i.e.. \") ,&\" not \"), &\". Could\nalso use line-wrapping at 79 characters.\n"},{"id":"444342","messageId":"20211217112629.12334-1-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211210103435.83656-1-chiyutianyi@gmail.com","subject":"[PATCH v5 0/6] unpack large blobs in stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-17T11:26:23Z","receivedAt":"2021-12-17T11:28:39Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nChanges since v5:\n* Refactor write_loose_object() to reuse in stream version sugguest by\n  Ævar Arnfjörð Bjarmason [1].\n\n* Add a new testcase into t5590-unpack-non-delta-objects to cover the case of\n  unpacking existing objects.\n\n* Fix code formatting in unpack-objects.c sugguest by\n  Ævar Arnfjörð Bjarmason [2].\n\n1. https://lore.kernel.org/git/211213.86bl1l9bfz.gmgdl@evledraar.gmail.com/\n2. https://lore.kernel.org/git/211213.867dc8ansq.gmgdl@evledraar.gmail.com/\n\nHan Xin (6):\n  object-file.c: release strbuf in write_loose_object()\n  object-file.c: refactor object header generation into a function\n  object-file.c: refactor write_loose_object() to reuse in stream\n    version\n  object-file.c: make \"write_object_file_flags()\" to support read in\n    stream\n  unpack-objects.c: add dry_run mode for get_data()\n  unpack-objects: unpack_non_delta_entry() read data in a stream\n\n Documentation/config/core.txt       |  11 ++\n builtin/unpack-objects.c            |  94 ++++++++++++-\n cache.h                             |   2 +\n config.c                            |   5 +\n environment.c                       |   1 +\n object-file.c                       | 207 +++++++++++++++++++++++-----\n object-store.h                      |   5 +\n t/t5590-unpack-non-delta-objects.sh |  87 ++++++++++++\n 8 files changed, 370 insertions(+), 42 deletions(-)\n create mode 100755 t/t5590-unpack-non-delta-objects.sh\n\nRange-diff against v5:\n1:  f3595e68cc < -:  ---------- object-file: refactor write_loose_object() to support read from stream\n2:  c25fdd1fe5 < -:  ---------- object-file.c: handle undetermined oid in write_loose_object()\n3:  ed226f2f9f < -:  ---------- object-file.c: read stream in a loop in write_loose_object()\n-:  ---------- > 1:  59d35dac5f object-file.c: release strbuf in write_loose_object()\n-:  ---------- > 2:  2174a6cbad object-file.c: refactor object header generation into a function\n-:  ---------- > 3:  8a704ecc59 object-file.c: refactor write_loose_object() to reuse in stream version\n-:  ---------- > 4:  96f05632a2 object-file.c: make \"write_object_file_flags()\" to support read in stream\n4:  2f91e540f6 ! 5:  1acbb6e849 unpack-objects.c: add dry_run mode for get_data()\n    @@ builtin/unpack-objects.c: static void use(int bytes)\n      {\n      \tgit_zstream stream;\n     -\tvoid *buf = xmallocz(size);\n    -+\tunsigned long bufsize = dry_run ? 8192 : size;\n    -+\tvoid *buf = xmallocz(bufsize);\n    ++\tunsigned long bufsize;\n    ++\tvoid *buf;\n      \n      \tmemset(&stream, 0, sizeof(stream));\n    ++\tif (dry_run && size > 8192)\n    ++\t\tbufsize = 8192;\n    ++\telse\n    ++\t\tbufsize = size;\n    ++\tbuf = xmallocz(bufsize);\n      \n      \tstream.next_out = buf;\n     -\tstream.avail_out = size;\n5:  7698938eac < -:  ---------- object-file.c: make \"write_object_file_flags()\" to support \"HASH_STREAM\"\n6:  92d69cb84a ! 6:  476aaba527 unpack-objects: unpack_non_delta_entry() read data in a stream\n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n     +\tint status;\n     +};\n     +\n    -+static const void *feed_input_zstream(struct input_stream *in_stream, unsigned long *readlen)\n    ++static const void *feed_input_zstream(const struct input_stream *in_stream,\n    ++\t\t\t\t      unsigned long *readlen)\n     +{\n     +\tstruct input_zstream_data *data = in_stream->data;\n     +\tgit_zstream *zstream = data->zstream;\n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n     +\t\t.read = feed_input_zstream,\n     +\t\t.data = &data,\n     +\t};\n    -+\tint ret;\n     +\n     +\tmemset(&zstream, 0, sizeof(zstream));\n     +\tmemset(&data, 0, sizeof(data));\n     +\tdata.zstream = &zstream;\n     +\tgit_inflate_init(&zstream);\n     +\n    -+\tif ((ret = write_object_file_flags(&in_stream, size, type_name(OBJ_BLOB) ,&obj_list[nr].oid, HASH_STREAM)))\n    -+\t\tdie(_(\"failed to write object in stream %d\"), ret);\n    ++\tif (write_object_file_flags(&in_stream, size,\n    ++\t\t\t\t    type_name(OBJ_BLOB),\n    ++\t\t\t\t    &obj_list[nr].oid,\n    ++\t\t\t\t    HASH_STREAM))\n    ++\t\tdie(_(\"failed to write object in stream\"));\n     +\n     +\tif (zstream.total_out != size || data.status != Z_STREAM_END)\n     +\t\tdie(_(\"inflate returned %d\"), data.status);\n     +\tgit_inflate_end(&zstream);\n     +\n    -+\tif (strict && !dry_run) {\n    ++\tif (strict) {\n     +\t\tstruct blob *blob = lookup_blob(the_repository, &obj_list[nr].oid);\n     +\t\tif (blob)\n     +\t\t\tblob->object.flags |= FLAG_WRITTEN;\n     +\t\telse\n    -+\t\t\tdie(\"invalid blob object from stream\");\n    ++\t\t\tdie(_(\"invalid blob object from stream\"));\n     +\t}\n     +\tobj_list[nr].obj = NULL;\n     +}\n    @@ t/t5590-unpack-non-delta-objects.sh (new)\n     +prepare_dest () {\n     +\ttest_when_finished \"rm -rf dest.git\" &&\n     +\tgit init --bare dest.git &&\n    -+\tgit -C dest.git config core.bigFileStreamingThreshold $1\n    ++\tgit -C dest.git config core.bigFileStreamingThreshold $1 &&\n     +\tgit -C dest.git config core.bigFileThreshold $1\n     +}\n     +\n    @@ t/t5590-unpack-non-delta-objects.sh (new)\n     +\ttest_cmp expect actual\n     +'\n     +\n    ++test_expect_success 'unpack big object in stream with existing oids' '\n    ++\tprepare_dest 1m &&\n    ++\tgit -C dest.git index-pack --stdin <test-$PACK.pack &&\n    ++\t(\n    ++\t\tcd dest.git &&\n    ++\t\tfind objects/?? -type f | sort\n    ++\t) >actual &&\n    ++\ttest_must_be_empty actual &&\n    ++\tgit -C dest.git unpack-objects <test-$PACK.pack &&\n    ++\tgit -C dest.git fsck &&\n    ++\t(\n    ++\t\tcd dest.git &&\n    ++\t\tfind objects/?? -type f | sort\n    ++\t) >actual &&\n    ++\ttest_must_be_empty actual\n    ++'\n    ++\n     +test_expect_success 'unpack-objects dry-run' '\n     +\tprepare_dest 1m &&\n     +\tgit -C dest.git unpack-objects -n <test-$PACK.pack &&\n-- \n2.34.1.52.gfcc2252aea.agit.6.5.6\n\n"},{"id":"444343","messageId":"20211217112629.12334-2-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211210103435.83656-1-chiyutianyi@gmail.com","subject":"[PATCH v6 1/6] object-file.c: release strbuf in write_loose_object()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-17T11:26:24Z","receivedAt":"2021-12-17T11:28:43Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nFix a strbuf leak in \"write_loose_object()\" sugguested by\nÆvar Arnfjörð Bjarmason.\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c | 15 +++++++++++----\n 1 file changed, 11 insertions(+), 4 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex eb1426f98c..32acf1dad6 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1874,11 +1874,14 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tfd = create_tmpfile(&tmp_file, filename.buf);\n \tif (fd < 0) {\n \t\tif (flags & HASH_SILENT)\n-\t\t\treturn -1;\n+\t\t\tret = -1;\n \t\telse if (errno == EACCES)\n-\t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n+\t\t\tret = error(_(\"insufficient permission for adding an \"\n+\t\t\t\t      \"object to repository database %s\"),\n+\t\t\t\t    get_object_directory());\n \t\telse\n-\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n+\t\t\tret = error_errno(_(\"unable to create temporary file\"));\n+\t\tgoto cleanup;\n \t}\n \n \t/* Set it up */\n@@ -1930,7 +1933,11 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\t\twarning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n \t}\n \n-\treturn finalize_object_file(tmp_file.buf, filename.buf);\n+\tret = finalize_object_file(tmp_file.buf, filename.buf);\n+cleanup:\n+\tstrbuf_release(&filename);\n+\tstrbuf_release(&tmp_file);\n+\treturn ret;\n }\n \n static int freshen_loose_object(const struct object_id *oid)\n-- \n2.34.1.52.gfcc2252aea.agit.6.5.6\n\n"},{"id":"444344","messageId":"20211217112629.12334-3-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211210103435.83656-1-chiyutianyi@gmail.com","subject":"[PATCH v6 2/6] object-file.c: refactor object header generation into a function","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-17T11:26:25Z","receivedAt":"2021-12-17T11:28:44Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nThere are 3 places where \"xsnprintf\" is used to generate the object\nheader, and I originally planned to add a fourth in the latter patch.\n\nAccording to Ævar Arnfjörð Bjarmason’s suggestion, although it's just\none line, it's also code that's very central to git, so reafactor them\ninto a function which will help later readability.\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c | 14 +++++++++++---\n 1 file changed, 11 insertions(+), 3 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 32acf1dad6..95fcd5435d 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1006,6 +1006,14 @@ void *xmmap(void *start, size_t length,\n \treturn ret;\n }\n \n+static inline int generate_object_header(char *buf, int bufsz,\n+\t\t\t\t\t const char *type_name,\n+\t\t\t\t\t unsigned long size)\n+{\n+\treturn xsnprintf(buf, bufsz, \"%s %\"PRIuMAX, type_name,\n+\t\t\t (uintmax_t)size) + 1;\n+}\n+\n /*\n  * With an in-core object data in \"map\", rehash it to make sure the\n  * object name actually matches \"oid\" to detect object corruption.\n@@ -1034,7 +1042,7 @@ int check_object_signature(struct repository *r, const struct object_id *oid,\n \t\treturn -1;\n \n \t/* Generate the header */\n-\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(obj_type), (uintmax_t)size) + 1;\n+\thdrlen = generate_object_header(hdr, sizeof(hdr), type_name(obj_type), size);\n \n \t/* Sha1.. */\n \tr->hash_algo->init_fn(&c);\n@@ -1734,7 +1742,7 @@ static void write_object_file_prepare(const struct git_hash_algo *algo,\n \tgit_hash_ctx c;\n \n \t/* Generate the header */\n-\t*hdrlen = xsnprintf(hdr, *hdrlen, \"%s %\"PRIuMAX , type, (uintmax_t)len)+1;\n+\t*hdrlen = generate_object_header(hdr, *hdrlen, type, len);\n \n \t/* Sha1.. */\n \talgo->init_fn(&c);\n@@ -2013,7 +2021,7 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n \tbuf = read_object(the_repository, oid, &type, &len);\n \tif (!buf)\n \t\treturn error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n-\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n+\thdrlen = generate_object_header(hdr, sizeof(hdr), type_name(type), len);\n \tret = write_loose_object(oid, hdr, hdrlen, buf, len, mtime, 0);\n \tfree(buf);\n \n-- \n2.34.1.52.gfcc2252aea.agit.6.5.6\n\n"},{"id":"444345","messageId":"20211217112629.12334-4-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211210103435.83656-1-chiyutianyi@gmail.com","subject":"[PATCH v6 3/6] object-file.c: refactor write_loose_object() to reuse in stream version","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-17T11:26:26Z","receivedAt":"2021-12-17T11:28:47Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWe used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\nentire contents of a blob object, no matter how big it is. This\nimplementation may consume all the memory and cause OOM.\n\nThis can be improved by feeding data to \"stream_loose_object()\" in\nstream instead of read into the whole buf.\n\nAs this new method \"stream_loose_object()\" has many similarities with\n\"write_loose_object()\", we split up \"write_loose_object()\" into some\nsteps:\n 1. Figuring out a path for the (temp) object file.\n 2. Creating the tempfile.\n 3. Setting up zlib and write header.\n 4. Write object data and handle errors.\n 5. Optionally, do someting after write, maybe force a loose object if\n\"mtime\".\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c | 98 +++++++++++++++++++++++++++++++++------------------\n 1 file changed, 63 insertions(+), 35 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 95fcd5435d..dd29e5372e 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1751,6 +1751,25 @@ static void write_object_file_prepare(const struct git_hash_algo *algo,\n \talgo->final_oid_fn(oid, &c);\n }\n \n+/*\n+ * Move the just written object with proper mtime into its final resting place.\n+ */\n+static int finalize_object_file_with_mtime(const char *tmpfile,\n+\t\t\t\t\t   const char *filename,\n+\t\t\t\t\t   time_t mtime,\n+\t\t\t\t\t   unsigned flags)\n+{\n+\tstruct utimbuf utb;\n+\n+\tif (mtime) {\n+\t\tutb.actime = mtime;\n+\t\tutb.modtime = mtime;\n+\t\tif (utime(tmpfile, &utb) < 0 && !(flags & HASH_SILENT))\n+\t\t\twarning_errno(_(\"failed utime() on %s\"), tmpfile);\n+\t}\n+\treturn finalize_object_file(tmpfile, filename);\n+}\n+\n /*\n  * Move the just written object into its final resting place.\n  */\n@@ -1836,7 +1855,8 @@ static inline int directory_size(const char *filename)\n  * We want to avoid cross-directory filename renames, because those\n  * can have problems on various filesystems (FAT, NFS, Coda).\n  */\n-static int create_tmpfile(struct strbuf *tmp, const char *filename)\n+static int create_tmpfile(struct strbuf *tmp, const char *filename,\n+\t\t\t  unsigned flags)\n {\n \tint fd, dirlen = directory_size(filename);\n \n@@ -1844,7 +1864,9 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n \tstrbuf_add(tmp, filename, dirlen);\n \tstrbuf_addstr(tmp, \"tmp_obj_XXXXXX\");\n \tfd = git_mkstemp_mode(tmp->buf, 0444);\n-\tif (fd < 0 && dirlen && errno == ENOENT) {\n+\tdo {\n+\t\tif (fd >= 0 || !dirlen || errno != ENOENT)\n+\t\t\tbreak;\n \t\t/*\n \t\t * Make sure the directory exists; note that the contents\n \t\t * of the buffer are undefined after mkstemp returns an\n@@ -1854,17 +1876,48 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n \t\tstrbuf_reset(tmp);\n \t\tstrbuf_add(tmp, filename, dirlen - 1);\n \t\tif (mkdir(tmp->buf, 0777) && errno != EEXIST)\n-\t\t\treturn -1;\n+\t\t\tbreak;\n \t\tif (adjust_shared_perm(tmp->buf))\n-\t\t\treturn -1;\n+\t\t\tbreak;\n \n \t\t/* Try again */\n \t\tstrbuf_addstr(tmp, \"/tmp_obj_XXXXXX\");\n \t\tfd = git_mkstemp_mode(tmp->buf, 0444);\n+\t} while (0);\n+\n+\tif (fd < 0 && !(flags & HASH_SILENT)) {\n+\t\tif (errno == EACCES)\n+\t\t\treturn error(_(\"insufficient permission for adding an \"\n+\t\t\t\t       \"object to repository database %s\"),\n+\t\t\t\t     get_object_directory());\n+\t\telse\n+\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n \t}\n+\n \treturn fd;\n }\n \n+static void setup_stream_and_header(git_zstream *stream,\n+\t\t\t\t    unsigned char *compressed,\n+\t\t\t\t    unsigned long compressed_size,\n+\t\t\t\t    git_hash_ctx *c,\n+\t\t\t\t    char *hdr,\n+\t\t\t\t    int hdrlen)\n+{\n+\t/* Set it up */\n+\tgit_deflate_init(stream, zlib_compression_level);\n+\tstream->next_out = compressed;\n+\tstream->avail_out = compressed_size;\n+\tthe_hash_algo->init_fn(c);\n+\n+\t/* First header.. */\n+\tstream->next_in = (unsigned char *)hdr;\n+\tstream->avail_in = hdrlen;\n+\twhile (git_deflate(stream, 0) == Z_OK)\n+\t\t; /* nothing */\n+\tthe_hash_algo->update_fn(c, hdr, hdrlen);\n+}\n+\n static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\t\t      int hdrlen, const void *buf, unsigned long len,\n \t\t\t      time_t mtime, unsigned flags)\n@@ -1879,31 +1932,15 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \n \tloose_object_path(the_repository, &filename, oid);\n \n-\tfd = create_tmpfile(&tmp_file, filename.buf);\n+\tfd = create_tmpfile(&tmp_file, filename.buf, flags);\n \tif (fd < 0) {\n-\t\tif (flags & HASH_SILENT)\n-\t\t\tret = -1;\n-\t\telse if (errno == EACCES)\n-\t\t\tret = error(_(\"insufficient permission for adding an \"\n-\t\t\t\t      \"object to repository database %s\"),\n-\t\t\t\t    get_object_directory());\n-\t\telse\n-\t\t\tret = error_errno(_(\"unable to create temporary file\"));\n+\t\tret = -1;\n \t\tgoto cleanup;\n \t}\n \n-\t/* Set it up */\n-\tgit_deflate_init(&stream, zlib_compression_level);\n-\tstream.next_out = compressed;\n-\tstream.avail_out = sizeof(compressed);\n-\tthe_hash_algo->init_fn(&c);\n-\n-\t/* First header.. */\n-\tstream.next_in = (unsigned char *)hdr;\n-\tstream.avail_in = hdrlen;\n-\twhile (git_deflate(&stream, 0) == Z_OK)\n-\t\t; /* nothing */\n-\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n+\t/* Set it up and write header */\n+\tsetup_stream_and_header(&stream, compressed, sizeof(compressed),\n+\t\t\t\t&c, hdr, hdrlen);\n \n \t/* Then the data itself.. */\n \tstream.next_in = (void *)buf;\n@@ -1932,16 +1969,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \n \tclose_loose_object(fd);\n \n-\tif (mtime) {\n-\t\tstruct utimbuf utb;\n-\t\tutb.actime = mtime;\n-\t\tutb.modtime = mtime;\n-\t\tif (utime(tmp_file.buf, &utb) < 0 &&\n-\t\t    !(flags & HASH_SILENT))\n-\t\t\twarning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n-\t}\n-\n-\tret = finalize_object_file(tmp_file.buf, filename.buf);\n+\tret = finalize_object_file_with_mtime(tmp_file.buf, filename.buf, mtime, flags);\n cleanup:\n \tstrbuf_release(&filename);\n \tstrbuf_release(&tmp_file);\n-- \n2.34.1.52.gfcc2252aea.agit.6.5.6\n\n"},{"id":"444346","messageId":"20211217112629.12334-5-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211210103435.83656-1-chiyutianyi@gmail.com","subject":"[PATCH v6 4/6] object-file.c: make \"write_object_file_flags()\" to support read in stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-17T11:26:27Z","receivedAt":"2021-12-17T11:28:52Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWe used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\nentire contents of a blob object, no matter how big it is. This\nimplementation may consume all the memory and cause OOM.\n\nThis can be improved by feeding data to \"stream_loose_object()\" in a\nstream. The input stream is implemented as an interface.\n\nWhen streaming a large blob object to \"write_loose_object()\", we have no\nchance to run \"write_object_file_prepare()\" to calculate the oid in\nadvance. So we need to handle undetermined oid in a new function called\n\"stream_loose_object()\".\n\nIn \"write_loose_object()\", we know the oid and we can write the\ntemporary file in the same directory as the final object, but for an\nobject with an undetermined oid, we don't know the exact directory for\nthe object, so we have to save the temporary file in \".git/objects/\"\ndirectory instead.\n\nWe will reuse \"write_object_file_flags()\" in \"unpack_non_delta_entry()\" to\nread the entire data contents in stream, so a new flag \"HASH_STREAM\" is\nadded. When read in stream, we needn't prepare the \"oid\" before\n\"write_loose_object()\", only generate the header.\n\"freshen_packed_object()\" or \"freshen_loose_object()\" will be called\ninside \"stream_loose_object()\" after obtaining the \"oid\".\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n cache.h        |  1 +\n object-file.c  | 92 ++++++++++++++++++++++++++++++++++++++++++++++++++\n object-store.h |  5 +++\n 3 files changed, 98 insertions(+)\n\ndiff --git a/cache.h b/cache.h\nindex cfba463aa9..6d68fd10a3 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -898,6 +898,7 @@ int ie_modified(struct index_state *, const struct cache_entry *, struct stat *,\n #define HASH_FORMAT_CHECK 2\n #define HASH_RENORMALIZE  4\n #define HASH_SILENT 8\n+#define HASH_STREAM 16\n int index_fd(struct index_state *istate, struct object_id *oid, int fd, struct stat *st, enum object_type type, const char *path, unsigned flags);\n int index_path(struct index_state *istate, struct object_id *oid, const char *path, struct stat *st, unsigned flags);\n \ndiff --git a/object-file.c b/object-file.c\nindex dd29e5372e..2ef1d4fb00 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1994,6 +1994,88 @@ static int freshen_packed_object(const struct object_id *oid)\n \treturn 1;\n }\n \n+static int stream_loose_object(struct object_id *oid, char *hdr, int hdrlen,\n+\t\t\t       const struct input_stream *in_stream,\n+\t\t\t       unsigned long len, time_t mtime, unsigned flags)\n+{\n+\tint fd, ret, err = 0, flush = 0;\n+\tunsigned char compressed[4096];\n+\tgit_zstream stream;\n+\tgit_hash_ctx c;\n+\tstruct object_id parano_oid;\n+\tstatic struct strbuf tmp_file = STRBUF_INIT;\n+\tstatic struct strbuf filename = STRBUF_INIT;\n+\tint dirlen;\n+\n+\t/* When oid is not determined, save tmp file to odb path. */\n+\tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n+\n+\tfd = create_tmpfile(&tmp_file, filename.buf, flags);\n+\tif (fd < 0) {\n+\t\terr = -1;\n+\t\tgoto cleanup;\n+\t}\n+\n+\t/* Set it up and write header */\n+\tsetup_stream_and_header(&stream, compressed, sizeof(compressed),\n+\t\t\t\t&c, hdr, hdrlen);\n+\n+\t/* Then the data itself.. */\n+\tdo {\n+\t\tunsigned char *in0 = stream.next_in;\n+\t\tif (!stream.avail_in) {\n+\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n+\t\t\tstream.next_in = (void *)in;\n+\t\t\tin0 = (unsigned char *)in;\n+\t\t\t/* All data has been read. */\n+\t\t\tif (len + hdrlen == stream.total_in + stream.avail_in)\n+\t\t\t\tflush = Z_FINISH;\n+\t\t}\n+\t\tret = git_deflate(&stream, flush);\n+\t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n+\t\tif (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n+\t\t\tdie(_(\"unable to write loose object file\"));\n+\t\tstream.next_out = compressed;\n+\t\tstream.avail_out = sizeof(compressed);\n+\t} while (ret == Z_OK || ret == Z_BUF_ERROR);\n+\n+\tif (ret != Z_STREAM_END)\n+\t\tdie(_(\"unable to deflate new object streamingly (%d)\"), ret);\n+\tret = git_deflate_end_gently(&stream);\n+\tif (ret != Z_OK)\n+\t\tdie(_(\"deflateEnd on object streamingly failed (%d)\"), ret);\n+\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n+\n+\tclose_loose_object(fd);\n+\n+\toidcpy(oid, &parano_oid);\n+\n+\tif (freshen_packed_object(oid) || freshen_loose_object(oid)) {\n+\t\tunlink_or_warn(tmp_file.buf);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tloose_object_path(the_repository, &filename, oid);\n+\n+\t/* We finally know the object path, and create the missing dir. */\n+\tdirlen = directory_size(filename.buf);\n+\tif (dirlen) {\n+\t\tstruct strbuf dir = STRBUF_INIT;\n+\t\tstrbuf_add(&dir, filename.buf, dirlen - 1);\n+\n+\t\tif (mkdir_in_gitdir(dir.buf) < 0) {\n+\t\t\terr = -1;\n+\t\t\tgoto cleanup;\n+\t\t}\n+\t}\n+\n+\terr = finalize_object_file_with_mtime(tmp_file.buf, filename.buf, mtime, flags);\n+cleanup:\n+\tstrbuf_release(&tmp_file);\n+\tstrbuf_release(&filename);\n+\treturn err;\n+}\n+\n int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t    const char *type, struct object_id *oid,\n \t\t\t    unsigned flags)\n@@ -2001,6 +2083,16 @@ int write_object_file_flags(const void *buf, unsigned long len,\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen = sizeof(hdr);\n \n+\t/* When streaming a large blob object (marked as HASH_STREAM),\n+\t * we have no chance to run \"write_object_file_prepare()\" to\n+\t * calculate the \"oid\" in advance.  Call \"stream_loose_object()\"\n+\t * to write loose object in stream.\n+\t */\n+\tif (flags & HASH_STREAM) {\n+\t\thdrlen = generate_object_header(hdr, hdrlen, type, len);\n+\t\treturn stream_loose_object(oid, hdr, hdrlen, buf, len, 0, flags);\n+\t}\n+\n \t/* Normally if we have it in the pack then we do not bother writing\n \t * it out into .git/objects/??/?{38} file.\n \t */\ndiff --git a/object-store.h b/object-store.h\nindex 952efb6a4b..4040e2c40a 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -34,6 +34,11 @@ struct object_directory {\n \tchar *path;\n };\n \n+struct input_stream {\n+\tconst void *(*read)(const struct input_stream *, unsigned long *len);\n+\tvoid *data;\n+};\n+\n KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n \tstruct object_directory *, 1, fspathhash, fspatheq)\n \n-- \n2.34.1.52.gfcc2252aea.agit.6.5.6\n\n"},{"id":"444347","messageId":"20211217112629.12334-6-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211210103435.83656-1-chiyutianyi@gmail.com","subject":"[PATCH v6 5/6] unpack-objects.c: add dry_run mode for get_data()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-17T11:26:28Z","receivedAt":"2021-12-17T11:28:54Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nIn dry_run mode, \"get_data()\" is used to verify the inflation of data,\nand the returned buffer will not be used at all and will be freed\nimmediately. Even in dry_run mode, it is dangerous to allocate a\nfull-size buffer for a large blob object. Therefore, only allocate a\nlow memory footprint when calling \"get_data()\" in dry_run mode.\n\nSuggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n builtin/unpack-objects.c | 23 +++++++++++++++++------\n 1 file changed, 17 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 4a9466295b..c4a17bdb44 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -96,15 +96,21 @@ static void use(int bytes)\n \tdisplay_throughput(progress, consumed_bytes);\n }\n \n-static void *get_data(unsigned long size)\n+static void *get_data(unsigned long size, int dry_run)\n {\n \tgit_zstream stream;\n-\tvoid *buf = xmallocz(size);\n+\tunsigned long bufsize;\n+\tvoid *buf;\n \n \tmemset(&stream, 0, sizeof(stream));\n+\tif (dry_run && size > 8192)\n+\t\tbufsize = 8192;\n+\telse\n+\t\tbufsize = size;\n+\tbuf = xmallocz(bufsize);\n \n \tstream.next_out = buf;\n-\tstream.avail_out = size;\n+\tstream.avail_out = bufsize;\n \tstream.next_in = fill(1);\n \tstream.avail_in = len;\n \tgit_inflate_init(&stream);\n@@ -124,6 +130,11 @@ static void *get_data(unsigned long size)\n \t\t}\n \t\tstream.next_in = fill(1);\n \t\tstream.avail_in = len;\n+\t\tif (dry_run) {\n+\t\t\t/* reuse the buffer in dry_run mode */\n+\t\t\tstream.next_out = buf;\n+\t\t\tstream.avail_out = bufsize;\n+\t\t}\n \t}\n \tgit_inflate_end(&stream);\n \treturn buf;\n@@ -323,7 +334,7 @@ static void added_object(unsigned nr, enum object_type type,\n static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n \t\t\t\t   unsigned nr)\n {\n-\tvoid *buf = get_data(size);\n+\tvoid *buf = get_data(size, dry_run);\n \n \tif (!dry_run && buf)\n \t\twrite_object(nr, type, buf, size);\n@@ -357,7 +368,7 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \tif (type == OBJ_REF_DELTA) {\n \t\toidread(&base_oid, fill(the_hash_algo->rawsz));\n \t\tuse(the_hash_algo->rawsz);\n-\t\tdelta_data = get_data(delta_size);\n+\t\tdelta_data = get_data(delta_size, dry_run);\n \t\tif (dry_run || !delta_data) {\n \t\t\tfree(delta_data);\n \t\t\treturn;\n@@ -396,7 +407,7 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\tif (base_offset <= 0 || base_offset >= obj_list[nr].offset)\n \t\t\tdie(\"offset value out of bound for delta base object\");\n \n-\t\tdelta_data = get_data(delta_size);\n+\t\tdelta_data = get_data(delta_size, dry_run);\n \t\tif (dry_run || !delta_data) {\n \t\t\tfree(delta_data);\n \t\t\treturn;\n-- \n2.34.1.52.gfcc2252aea.agit.6.5.6\n\n"},{"id":"444348","messageId":"20211217112629.12334-7-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211210103435.83656-1-chiyutianyi@gmail.com","subject":"[PATCH v6 6/6] unpack-objects: unpack_non_delta_entry() read data in a stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-17T11:26:29Z","receivedAt":"2021-12-17T11:28:58Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWe used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\nentire contents of a blob object, no matter how big it is. This\nimplementation may consume all the memory and cause OOM.\n\nBy implementing a zstream version of input_stream interface, we can use\na small fixed buffer for \"unpack_non_delta_entry()\".\n\nHowever, unpack non-delta objects from a stream instead of from an\nentrie buffer will have 10% performance penalty. Therefore, only unpack\nobject larger than the \"core.BigFileStreamingThreshold\" in zstream. See\nthe following benchmarks:\n\n    hyperfine \\\n      --setup \\\n      'if ! test -d scalar.git; then git clone --bare https://github.com/microsoft/scalar.git; cp scalar.git/objects/pack/*.pack small.pack; fi' \\\n      --prepare 'rm -rf dest.git && git init --bare dest.git'\n\n    Summary\n      './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'origin/master'\n        1.01 ± 0.04 times faster than './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'HEAD~1'\n        1.01 ± 0.04 times faster than './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'HEAD~0'\n        1.03 ± 0.10 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'origin/master'\n        1.02 ± 0.07 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'HEAD~0'\n        1.10 ± 0.04 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'HEAD~1'\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Derrick Stolee <stolee@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n Documentation/config/core.txt       | 11 ++++\n builtin/unpack-objects.c            | 73 +++++++++++++++++++++++-\n cache.h                             |  1 +\n config.c                            |  5 ++\n environment.c                       |  1 +\n t/t5590-unpack-non-delta-objects.sh | 87 +++++++++++++++++++++++++++++\n 6 files changed, 177 insertions(+), 1 deletion(-)\n create mode 100755 t/t5590-unpack-non-delta-objects.sh\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex c04f62a54a..601b7a2418 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -424,6 +424,17 @@ be delta compressed, but larger binary media files won't be.\n +\n Common unit suffixes of 'k', 'm', or 'g' are supported.\n \n+core.bigFileStreamingThreshold::\n+\tFiles larger than this will be streamed out to a temporary\n+\tobject file while being hashed, which will when be renamed\n+\tin-place to a loose object, particularly if the\n+\t`core.bigFileThreshold' setting dictates that they're always\n+\twritten out as loose objects.\n++\n+Default is 128 MiB on all platforms.\n++\n+Common unit suffixes of 'k', 'm', or 'g' are supported.\n+\n core.excludesFile::\n \tSpecifies the pathname to the file that contains patterns to\n \tdescribe paths that are not meant to be tracked, in addition\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex c4a17bdb44..42e1033d85 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -331,11 +331,82 @@ static void added_object(unsigned nr, enum object_type type,\n \t}\n }\n \n+struct input_zstream_data {\n+\tgit_zstream *zstream;\n+\tunsigned char buf[8192];\n+\tint status;\n+};\n+\n+static const void *feed_input_zstream(const struct input_stream *in_stream,\n+\t\t\t\t      unsigned long *readlen)\n+{\n+\tstruct input_zstream_data *data = in_stream->data;\n+\tgit_zstream *zstream = data->zstream;\n+\tvoid *in = fill(1);\n+\n+\tif (!len || data->status == Z_STREAM_END) {\n+\t\t*readlen = 0;\n+\t\treturn NULL;\n+\t}\n+\n+\tzstream->next_out = data->buf;\n+\tzstream->avail_out = sizeof(data->buf);\n+\tzstream->next_in = in;\n+\tzstream->avail_in = len;\n+\n+\tdata->status = git_inflate(zstream, 0);\n+\tuse(len - zstream->avail_in);\n+\t*readlen = sizeof(data->buf) - zstream->avail_out;\n+\n+\treturn data->buf;\n+}\n+\n+static void write_stream_blob(unsigned nr, unsigned long size)\n+{\n+\tgit_zstream zstream;\n+\tstruct input_zstream_data data;\n+\tstruct input_stream in_stream = {\n+\t\t.read = feed_input_zstream,\n+\t\t.data = &data,\n+\t};\n+\n+\tmemset(&zstream, 0, sizeof(zstream));\n+\tmemset(&data, 0, sizeof(data));\n+\tdata.zstream = &zstream;\n+\tgit_inflate_init(&zstream);\n+\n+\tif (write_object_file_flags(&in_stream, size,\n+\t\t\t\t    type_name(OBJ_BLOB),\n+\t\t\t\t    &obj_list[nr].oid,\n+\t\t\t\t    HASH_STREAM))\n+\t\tdie(_(\"failed to write object in stream\"));\n+\n+\tif (zstream.total_out != size || data.status != Z_STREAM_END)\n+\t\tdie(_(\"inflate returned %d\"), data.status);\n+\tgit_inflate_end(&zstream);\n+\n+\tif (strict) {\n+\t\tstruct blob *blob = lookup_blob(the_repository, &obj_list[nr].oid);\n+\t\tif (blob)\n+\t\t\tblob->object.flags |= FLAG_WRITTEN;\n+\t\telse\n+\t\t\tdie(_(\"invalid blob object from stream\"));\n+\t}\n+\tobj_list[nr].obj = NULL;\n+}\n+\n static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n \t\t\t\t   unsigned nr)\n {\n-\tvoid *buf = get_data(size, dry_run);\n+\tvoid *buf;\n+\n+\t/* Write large blob in stream without allocating full buffer. */\n+\tif (!dry_run && type == OBJ_BLOB && size > big_file_streaming_threshold) {\n+\t\twrite_stream_blob(nr, size);\n+\t\treturn;\n+\t}\n \n+\tbuf = get_data(size, dry_run);\n \tif (!dry_run && buf)\n \t\twrite_object(nr, type, buf, size);\n \telse\ndiff --git a/cache.h b/cache.h\nindex 6d68fd10a3..976f9cf656 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -975,6 +975,7 @@ extern size_t packed_git_window_size;\n extern size_t packed_git_limit;\n extern size_t delta_base_cache_limit;\n extern unsigned long big_file_threshold;\n+extern unsigned long big_file_streaming_threshold;\n extern unsigned long pack_size_limit_cfg;\n \n /*\ndiff --git a/config.c b/config.c\nindex c5873f3a70..7b122a142a 100644\n--- a/config.c\n+++ b/config.c\n@@ -1408,6 +1408,11 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"core.bigfilestreamingthreshold\")) {\n+\t\tbig_file_streaming_threshold = git_config_ulong(var, value);\n+\t\treturn 0;\n+\t}\n+\n \tif (!strcmp(var, \"core.packedgitlimit\")) {\n \t\tpacked_git_limit = git_config_ulong(var, value);\n \t\treturn 0;\ndiff --git a/environment.c b/environment.c\nindex 0d06a31024..04bba593de 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -47,6 +47,7 @@ size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\n unsigned long big_file_threshold = 512 * 1024 * 1024;\n+unsigned long big_file_streaming_threshold = 128 * 1024 * 1024;\n int pager_use_color = 1;\n const char *editor_program;\n const char *askpass_program;\ndiff --git a/t/t5590-unpack-non-delta-objects.sh b/t/t5590-unpack-non-delta-objects.sh\nnew file mode 100755\nindex 0000000000..11c70e192c\n--- /dev/null\n+++ b/t/t5590-unpack-non-delta-objects.sh\n@@ -0,0 +1,87 @@\n+#!/bin/sh\n+#\n+# Copyright (c) 2021 Han Xin\n+#\n+\n+test_description='Test unpack-objects when receive pack'\n+\n+GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n+export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n+\n+. ./test-lib.sh\n+\n+prepare_dest () {\n+\ttest_when_finished \"rm -rf dest.git\" &&\n+\tgit init --bare dest.git &&\n+\tgit -C dest.git config core.bigFileStreamingThreshold $1 &&\n+\tgit -C dest.git config core.bigFileThreshold $1\n+}\n+\n+test_expect_success \"setup repo with big blobs (1.5 MB)\" '\n+\ttest-tool genrandom foo 1500000 >big-blob &&\n+\ttest_commit --append foo big-blob &&\n+\ttest-tool genrandom bar 1500000 >big-blob &&\n+\ttest_commit --append bar big-blob &&\n+\t(\n+\t\tcd .git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >expect &&\n+\tPACK=$(echo main | git pack-objects --revs test)\n+'\n+\n+test_expect_success 'setup env: GIT_ALLOC_LIMIT to 1MB' '\n+\tGIT_ALLOC_LIMIT=1m &&\n+\texport GIT_ALLOC_LIMIT\n+'\n+\n+test_expect_success 'fail to unpack-objects: cannot allocate' '\n+\tprepare_dest 2m &&\n+\ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n+\tgrep \"fatal: attempting to allocate\" err &&\n+\t(\n+\t\tcd dest.git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >actual &&\n+\ttest_file_not_empty actual &&\n+\t! test_cmp expect actual\n+'\n+\n+test_expect_success 'unpack big object in stream' '\n+\tprepare_dest 1m &&\n+\tgit -C dest.git unpack-objects <test-$PACK.pack &&\n+\tgit -C dest.git fsck &&\n+\t(\n+\t\tcd dest.git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'unpack big object in stream with existing oids' '\n+\tprepare_dest 1m &&\n+\tgit -C dest.git index-pack --stdin <test-$PACK.pack &&\n+\t(\n+\t\tcd dest.git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >actual &&\n+\ttest_must_be_empty actual &&\n+\tgit -C dest.git unpack-objects <test-$PACK.pack &&\n+\tgit -C dest.git fsck &&\n+\t(\n+\t\tcd dest.git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >actual &&\n+\ttest_must_be_empty actual\n+'\n+\n+test_expect_success 'unpack-objects dry-run' '\n+\tprepare_dest 1m &&\n+\tgit -C dest.git unpack-objects -n <test-$PACK.pack &&\n+\t(\n+\t\tcd dest.git &&\n+\t\tfind objects/ -type f\n+\t) >actual &&\n+\ttest_must_be_empty actual\n+'\n+\n+test_done\n-- \n2.34.1.52.gfcc2252aea.agit.6.5.6\n\n"},{"id":"444386","messageId":"c860c56f-ce25-4391-7f65-50c9d5d80c2c@web.de","threadId":"56672","inReplyTo":"20211217112629.12334-2-chiyutianyi@gmail.com","subject":"Re: [PATCH v6 1/6] object-file.c: release strbuf in write_loose_object()","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2021-12-17T19:28:55Z","receivedAt":"2021-12-17T19:30:28Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 17.12.21 um 12:26 schrieb Han Xin:\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> Fix a strbuf leak in \"write_loose_object()\" sugguested by\n> Ævar Arnfjörð Bjarmason.\n>\n> Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  object-file.c | 15 +++++++++++----\n>  1 file changed, 11 insertions(+), 4 deletions(-)\n>\n> diff --git a/object-file.c b/object-file.c\n> index eb1426f98c..32acf1dad6 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1874,11 +1874,14 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n\n\nRelevant context lines:\n\n\tstatic struct strbuf tmp_file = STRBUF_INIT;\n\tstatic struct strbuf filename = STRBUF_INIT;\n\n\tloose_object_path(the_repository, &filename, oid);\n\n>  \tfd = create_tmpfile(&tmp_file, filename.buf);\n>  \tif (fd < 0) {\n>  \t\tif (flags & HASH_SILENT)\n> -\t\t\treturn -1;\n> +\t\t\tret = -1;\n>  \t\telse if (errno == EACCES)\n> -\t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n> +\t\t\tret = error(_(\"insufficient permission for adding an \"\n> +\t\t\t\t      \"object to repository database %s\"),\n> +\t\t\t\t    get_object_directory());\n>  \t\telse\n> -\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n> +\t\t\tret = error_errno(_(\"unable to create temporary file\"));\n> +\t\tgoto cleanup;\n>  \t}\n>\n>  \t/* Set it up */\n> @@ -1930,7 +1933,11 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \t\t\twarning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n>  \t}\n>\n> -\treturn finalize_object_file(tmp_file.buf, filename.buf);\n> +\tret = finalize_object_file(tmp_file.buf, filename.buf);\n> +cleanup:\n> +\tstrbuf_release(&filename);\n> +\tstrbuf_release(&tmp_file);\n\nThere was no leak before.  Both strbufs are static and both functions\nthey are passed to (loose_object_path() and create_tmpfile()) reset\nthem first.  So while the allocated memory was not released before,\nit was reused.\n\nNot sure if making write_loose_object() allocate and release these\nbuffers on every call has much of a performance impact.  The only\nreason I can think of for wanting such a change is to get rid of the\nstatic buffers, to allow the function to be used by concurrent\nthreads.\n\nSo I think either keeping the code as-is or also making the strbufs\nnon-static would be better (but then discussing a possible\nperformance impact in the commit message would be nice).\n\n> +\treturn ret;\n>  }\n>\n>  static int freshen_loose_object(const struct object_id *oid)\n\n"},{"id":"444394","messageId":"39ac2d4c-e153-3369-f93f-4d8124f35b87@web.de","threadId":"56672","inReplyTo":"20211217112629.12334-6-chiyutianyi@gmail.com","subject":"Re: [PATCH v6 5/6] unpack-objects.c: add dry_run mode for get_data()","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2021-12-17T21:22:01Z","receivedAt":"2021-12-17T21:22:25Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 17.12.21 um 12:26 schrieb Han Xin:\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> In dry_run mode, \"get_data()\" is used to verify the inflation of data,\n> and the returned buffer will not be used at all and will be freed\n> immediately. Even in dry_run mode, it is dangerous to allocate a\n> full-size buffer for a large blob object. Therefore, only allocate a\n> low memory footprint when calling \"get_data()\" in dry_run mode.\n\nClever.  Looks good to me.\n\nFor some reason I was expecting this patch to have some connection to\none of the earlier ones (perhaps because get_data() was mentioned),\nbut it is technically independent.\n\n>\n> Suggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  builtin/unpack-objects.c | 23 +++++++++++++++++------\n>  1 file changed, 17 insertions(+), 6 deletions(-)\n>\n> diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n> index 4a9466295b..c4a17bdb44 100644\n> --- a/builtin/unpack-objects.c\n> +++ b/builtin/unpack-objects.c\n> @@ -96,15 +96,21 @@ static void use(int bytes)\n>  \tdisplay_throughput(progress, consumed_bytes);\n>  }\n>\n> -static void *get_data(unsigned long size)\n> +static void *get_data(unsigned long size, int dry_run)\n>  {\n>  \tgit_zstream stream;\n> -\tvoid *buf = xmallocz(size);\n> +\tunsigned long bufsize;\n> +\tvoid *buf;\n>\n>  \tmemset(&stream, 0, sizeof(stream));\n> +\tif (dry_run && size > 8192)\n> +\t\tbufsize = 8192;\n> +\telse\n> +\t\tbufsize = size;\n> +\tbuf = xmallocz(bufsize);\n>\n>  \tstream.next_out = buf;\n> -\tstream.avail_out = size;\n> +\tstream.avail_out = bufsize;\n>  \tstream.next_in = fill(1);\n>  \tstream.avail_in = len;\n>  \tgit_inflate_init(&stream);\n> @@ -124,6 +130,11 @@ static void *get_data(unsigned long size)\n>  \t\t}\n>  \t\tstream.next_in = fill(1);\n>  \t\tstream.avail_in = len;\n> +\t\tif (dry_run) {\n> +\t\t\t/* reuse the buffer in dry_run mode */\n> +\t\t\tstream.next_out = buf;\n> +\t\t\tstream.avail_out = bufsize;\n> +\t\t}\n>  \t}\n>  \tgit_inflate_end(&stream);\n>  \treturn buf;\n> @@ -323,7 +334,7 @@ static void added_object(unsigned nr, enum object_type type,\n>  static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n>  \t\t\t\t   unsigned nr)\n>  {\n> -\tvoid *buf = get_data(size);\n> +\tvoid *buf = get_data(size, dry_run);\n>\n>  \tif (!dry_run && buf)\n>  \t\twrite_object(nr, type, buf, size);\n> @@ -357,7 +368,7 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n>  \tif (type == OBJ_REF_DELTA) {\n>  \t\toidread(&base_oid, fill(the_hash_algo->rawsz));\n>  \t\tuse(the_hash_algo->rawsz);\n> -\t\tdelta_data = get_data(delta_size);\n> +\t\tdelta_data = get_data(delta_size, dry_run);\n>  \t\tif (dry_run || !delta_data) {\n>  \t\t\tfree(delta_data);\n>  \t\t\treturn;\n> @@ -396,7 +407,7 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n>  \t\tif (base_offset <= 0 || base_offset >= obj_list[nr].offset)\n>  \t\t\tdie(\"offset value out of bound for delta base object\");\n>\n> -\t\tdelta_data = get_data(delta_size);\n> +\t\tdelta_data = get_data(delta_size, dry_run);\n>  \t\tif (dry_run || !delta_data) {\n>  \t\t\tfree(delta_data);\n>  \t\t\treturn;\n\n"},{"id":"444400","messageId":"e959e4f1-7500-5f6b-5bd2-2f060287eeff@web.de","threadId":"56672","inReplyTo":"20211217112629.12334-5-chiyutianyi@gmail.com","subject":"Re: [PATCH v6 4/6] object-file.c: make \"write_object_file_flags()\" to support read in stream","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2021-12-17T22:52:21Z","receivedAt":"2021-12-17T22:52:43Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 17.12.21 um 12:26 schrieb Han Xin:\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> We used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\n> entire contents of a blob object, no matter how big it is. This\n> implementation may consume all the memory and cause OOM.\n>\n> This can be improved by feeding data to \"stream_loose_object()\" in a\n> stream. The input stream is implemented as an interface.\n>\n> When streaming a large blob object to \"write_loose_object()\", we have no\n> chance to run \"write_object_file_prepare()\" to calculate the oid in\n> advance. So we need to handle undetermined oid in a new function called\n> \"stream_loose_object()\".\n>\n> In \"write_loose_object()\", we know the oid and we can write the\n> temporary file in the same directory as the final object, but for an\n> object with an undetermined oid, we don't know the exact directory for\n> the object, so we have to save the temporary file in \".git/objects/\"\n> directory instead.\n>\n> We will reuse \"write_object_file_flags()\" in \"unpack_non_delta_entry()\" to\n> read the entire data contents in stream, so a new flag \"HASH_STREAM\" is\n> added. When read in stream, we needn't prepare the \"oid\" before\n> \"write_loose_object()\", only generate the header.\n> \"freshen_packed_object()\" or \"freshen_loose_object()\" will be called\n> inside \"stream_loose_object()\" after obtaining the \"oid\".\n>\n> Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  cache.h        |  1 +\n>  object-file.c  | 92 ++++++++++++++++++++++++++++++++++++++++++++++++++\n>  object-store.h |  5 +++\n>  3 files changed, 98 insertions(+)\n>\n> diff --git a/cache.h b/cache.h\n> index cfba463aa9..6d68fd10a3 100644\n> --- a/cache.h\n> +++ b/cache.h\n> @@ -898,6 +898,7 @@ int ie_modified(struct index_state *, const struct cache_entry *, struct stat *,\n>  #define HASH_FORMAT_CHECK 2\n>  #define HASH_RENORMALIZE  4\n>  #define HASH_SILENT 8\n> +#define HASH_STREAM 16\n>  int index_fd(struct index_state *istate, struct object_id *oid, int fd, struct stat *st, enum object_type type, const char *path, unsigned flags);\n>  int index_path(struct index_state *istate, struct object_id *oid, const char *path, struct stat *st, unsigned flags);\n>\n> diff --git a/object-file.c b/object-file.c\n> index dd29e5372e..2ef1d4fb00 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1994,6 +1994,88 @@ static int freshen_packed_object(const struct object_id *oid)\n>  \treturn 1;\n>  }\n>\n> +static int stream_loose_object(struct object_id *oid, char *hdr, int hdrlen,\n> +\t\t\t       const struct input_stream *in_stream,\n> +\t\t\t       unsigned long len, time_t mtime, unsigned flags)\n> +{\n> +\tint fd, ret, err = 0, flush = 0;\n> +\tunsigned char compressed[4096];\n> +\tgit_zstream stream;\n> +\tgit_hash_ctx c;\n> +\tstruct object_id parano_oid;\n> +\tstatic struct strbuf tmp_file = STRBUF_INIT;\n> +\tstatic struct strbuf filename = STRBUF_INIT;\n\nNote these static strbufs.\n\n> +\tint dirlen;\n> +\n> +\t/* When oid is not determined, save tmp file to odb path. */\n> +\tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n> +\n> +\tfd = create_tmpfile(&tmp_file, filename.buf, flags);\n> +\tif (fd < 0) {\n> +\t\terr = -1;\n> +\t\tgoto cleanup;\n> +\t}\n> +\n> +\t/* Set it up and write header */\n> +\tsetup_stream_and_header(&stream, compressed, sizeof(compressed),\n> +\t\t\t\t&c, hdr, hdrlen);\n> +\n> +\t/* Then the data itself.. */\n> +\tdo {\n> +\t\tunsigned char *in0 = stream.next_in;\n> +\t\tif (!stream.avail_in) {\n> +\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n> +\t\t\tstream.next_in = (void *)in;\n> +\t\t\tin0 = (unsigned char *)in;\n> +\t\t\t/* All data has been read. */\n> +\t\t\tif (len + hdrlen == stream.total_in + stream.avail_in)\n> +\t\t\t\tflush = Z_FINISH;\n> +\t\t}\n> +\t\tret = git_deflate(&stream, flush);\n> +\t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n> +\t\tif (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n> +\t\t\tdie(_(\"unable to write loose object file\"));\n> +\t\tstream.next_out = compressed;\n> +\t\tstream.avail_out = sizeof(compressed);\n> +\t} while (ret == Z_OK || ret == Z_BUF_ERROR);\n> +\n> +\tif (ret != Z_STREAM_END)\n> +\t\tdie(_(\"unable to deflate new object streamingly (%d)\"), ret);\n> +\tret = git_deflate_end_gently(&stream);\n> +\tif (ret != Z_OK)\n> +\t\tdie(_(\"deflateEnd on object streamingly failed (%d)\"), ret);\n> +\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n> +\n> +\tclose_loose_object(fd);\n> +\n> +\toidcpy(oid, &parano_oid);\n> +\n> +\tif (freshen_packed_object(oid) || freshen_loose_object(oid)) {\n> +\t\tunlink_or_warn(tmp_file.buf);\n> +\t\tgoto cleanup;\n> +\t}\n> +\n> +\tloose_object_path(the_repository, &filename, oid);\n> +\n> +\t/* We finally know the object path, and create the missing dir. */\n> +\tdirlen = directory_size(filename.buf);\n> +\tif (dirlen) {\n> +\t\tstruct strbuf dir = STRBUF_INIT;\n> +\t\tstrbuf_add(&dir, filename.buf, dirlen - 1);\n> +\n> +\t\tif (mkdir_in_gitdir(dir.buf) < 0) {\n> +\t\t\terr = -1;\n> +\t\t\tgoto cleanup;\n> +\t\t}\n> +\t}\n> +\n> +\terr = finalize_object_file_with_mtime(tmp_file.buf, filename.buf, mtime, flags);\n> +cleanup:\n> +\tstrbuf_release(&tmp_file);\n> +\tstrbuf_release(&filename);\n\nThe static strbufs are released here.  That combination is strange --\nwhy keep the variable values between calls by making them static, but\nthrow away the allocated buffers instead of reusing them?\n\nGiven that this function is only used for huge objects I think making\nthe strbufs non-static and releasing them is the best choice here.\n\n> +\treturn err;\n> +}\n> +\n>  int write_object_file_flags(const void *buf, unsigned long len,\n>  \t\t\t    const char *type, struct object_id *oid,\n>  \t\t\t    unsigned flags)\n> @@ -2001,6 +2083,16 @@ int write_object_file_flags(const void *buf, unsigned long len,\n>  \tchar hdr[MAX_HEADER_LEN];\n>  \tint hdrlen = sizeof(hdr);\n>\n> +\t/* When streaming a large blob object (marked as HASH_STREAM),\n> +\t * we have no chance to run \"write_object_file_prepare()\" to\n> +\t * calculate the \"oid\" in advance.  Call \"stream_loose_object()\"\n> +\t * to write loose object in stream.\n> +\t */\n> +\tif (flags & HASH_STREAM) {\n> +\t\thdrlen = generate_object_header(hdr, hdrlen, type, len);\n> +\t\treturn stream_loose_object(oid, hdr, hdrlen, buf, len, 0, flags);\n> +\t}\n\nSo stream_loose_object() is called by passing the flag HASH_STREAM to\nwrite_object_file_flags() and passing a struct input_stream via its\nbuf pointer.  That's ... unconventional.  Certainly scary.  Why not\nexport stream_loose_object() and call it directly?  Demo patch below.\n\n> +\n>  \t/* Normally if we have it in the pack then we do not bother writing\n>  \t * it out into .git/objects/??/?{38} file.\n>  \t */\n> diff --git a/object-store.h b/object-store.h\n> index 952efb6a4b..4040e2c40a 100644\n> --- a/object-store.h\n> +++ b/object-store.h\n> @@ -34,6 +34,11 @@ struct object_directory {\n>  \tchar *path;\n>  };\n>\n> +struct input_stream {\n> +\tconst void *(*read)(const struct input_stream *, unsigned long *len);\n> +\tvoid *data;\n> +};\n> +\n>  KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n>  \tstruct object_directory *, 1, fspathhash, fspatheq)\n>\n\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 42e1033d85..07d186bd20 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -375,10 +375,8 @@ static void write_stream_blob(unsigned nr, unsigned long size)\n \tdata.zstream = &zstream;\n \tgit_inflate_init(&zstream);\n\n-\tif (write_object_file_flags(&in_stream, size,\n-\t\t\t\t    type_name(OBJ_BLOB),\n-\t\t\t\t    &obj_list[nr].oid,\n-\t\t\t\t    HASH_STREAM))\n+\tif (stream_loose_object(&in_stream, size, type_name(OBJ_BLOB), 0, 0,\n+\t\t\t\t&obj_list[nr].oid))\n \t\tdie(_(\"failed to write object in stream\"));\n\n \tif (zstream.total_out != size || data.status != Z_STREAM_END)\ndiff --git a/object-file.c b/object-file.c\nindex 2ef1d4fb00..0a6b65ab26 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1994,9 +1994,9 @@ static int freshen_packed_object(const struct object_id *oid)\n \treturn 1;\n }\n\n-static int stream_loose_object(struct object_id *oid, char *hdr, int hdrlen,\n-\t\t\t       const struct input_stream *in_stream,\n-\t\t\t       unsigned long len, time_t mtime, unsigned flags)\n+int stream_loose_object(struct input_stream *in_stream, unsigned long len,\n+\t\t\tconst char *type, time_t mtime, unsigned flags,\n+\t\t\tstruct object_id *oid)\n {\n \tint fd, ret, err = 0, flush = 0;\n \tunsigned char compressed[4096];\n@@ -2006,6 +2006,10 @@ static int stream_loose_object(struct object_id *oid, char *hdr, int hdrlen,\n \tstatic struct strbuf tmp_file = STRBUF_INIT;\n \tstatic struct strbuf filename = STRBUF_INIT;\n \tint dirlen;\n+\tchar hdr[MAX_HEADER_LEN];\n+\tint hdrlen = sizeof(hdr);\n+\n+\thdrlen = generate_object_header(hdr, hdrlen, type, len);\n\n \t/* When oid is not determined, save tmp file to odb path. */\n \tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n@@ -2083,16 +2087,6 @@ int write_object_file_flags(const void *buf, unsigned long len,\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen = sizeof(hdr);\n\n-\t/* When streaming a large blob object (marked as HASH_STREAM),\n-\t * we have no chance to run \"write_object_file_prepare()\" to\n-\t * calculate the \"oid\" in advance.  Call \"stream_loose_object()\"\n-\t * to write loose object in stream.\n-\t */\n-\tif (flags & HASH_STREAM) {\n-\t\thdrlen = generate_object_header(hdr, hdrlen, type, len);\n-\t\treturn stream_loose_object(oid, hdr, hdrlen, buf, len, 0, flags);\n-\t}\n-\n \t/* Normally if we have it in the pack then we do not bother writing\n \t * it out into .git/objects/??/?{38} file.\n \t */\ndiff --git a/object-store.h b/object-store.h\nindex 4040e2c40a..786b6435b1 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -237,6 +237,10 @@ static inline int write_object_file(const void *buf, unsigned long len,\n \treturn write_object_file_flags(buf, len, type, oid, 0);\n }\n\n+int stream_loose_object(struct input_stream *in_stream, unsigned long len,\n+\t\t\tconst char *type, time_t mtime, unsigned flags,\n+\t\t\tstruct object_id *oid);\n+\n int hash_object_file_literally(const void *buf, unsigned long len,\n \t\t\t       const char *type, struct object_id *oid,\n \t\t\t       unsigned flags);\n\n"},{"id":"444404","messageId":"xmqqa6gy22py.fsf@gitster.g","threadId":"56672","inReplyTo":"c860c56f-ce25-4391-7f65-50c9d5d80c2c@web.de","subject":"Re: [PATCH v6 1/6] object-file.c: release strbuf in write_loose_object()","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-12-18T00:09:29Z","receivedAt":"2021-12-18T00:09:35Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"René Scharfe <l.s.r@web.de> writes:\n\n> There was no leak before.  Both strbufs are static and both functions\n> they are passed to (loose_object_path() and create_tmpfile()) reset\n> them first.  So while the allocated memory was not released before,\n> it was reused.\n>\n> Not sure if making write_loose_object() allocate and release these\n> buffers on every call has much of a performance impact.  The only\n> reason I can think of for wanting such a change is to get rid of the\n> static buffers, to allow the function to be used by concurrent\n> threads.\n>\n> So I think either keeping the code as-is or also making the strbufs\n> non-static would be better (but then discussing a possible\n> performance impact in the commit message would be nice).\n\nMakes sense.\n"},{"id":"444453","messageId":"RFC-patch-1.1-bda62567f6b-20211220T120740Z-avarab@gmail.com","threadId":"56672","inReplyTo":"20211217112629.12334-3-chiyutianyi@gmail.com","subject":"[RFC PATCH] object-file API: add a format_loose_header() function","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-20T12:10:02Z","receivedAt":"2021-12-20T12:10:08Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Add a convenience function to wrap the xsnprintf() command that\ngenerates loose object headers. This code was copy/pasted in various\nparts of the codebase, let's define it in one place and re-use it from\nthere.\n\nAll except one caller of it had a valid \"enum object_type\" for us,\nit's only write_object_file_prepare() which might need to deal with\n\"git hash-object --literally\" and a potential garbage type. Let's have\nthe primary API use an \"enum object_type\", and define an *_extended()\nfunction that can take an arbitrary \"const char *\" for the type.\n\nSee [1] for the discussion that prompted this patch, i.e. new code in\nobject-file.c that wanted to copy/paste the xsnprintf() invocation.\n\n1. https://lore.kernel.org/git/211213.86bl1l9bfz.gmgdl@evledraar.gmail.com/\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n\nOn Fri, Dec 17 2021, Han Xin wrote:\n\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> There are 3 places where \"xsnprintf\" is used to generate the object\n> header, and I originally planned to add a fourth in the latter patch.\n>\n> According to Ævar Arnfjörð Bjarmason’s suggestion, although it's just\n> one line, it's also code that's very central to git, so reafactor them\n> into a function which will help later readability.\n>\n> Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n\nI came up with this after my comment on the earlier round suggesting\nto factor out that header formatting. I don't know if this more\nthorough approach is worth it or if you'd like to replace your change\nwith this one, but just posting it here as an RFC.\n\n builtin/index-pack.c |  3 +--\n bulk-checkin.c       |  4 ++--\n cache.h              | 21 +++++++++++++++++++++\n http-push.c          |  2 +-\n object-file.c        | 14 +++++++++++---\n 5 files changed, 36 insertions(+), 8 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex c23d01de7dc..900c6539f68 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -449,8 +449,7 @@ static void *unpack_entry_data(off_t offset, unsigned long size,\n \tint hdrlen;\n \n \tif (!is_delta_type(type)) {\n-\t\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX,\n-\t\t\t\t   type_name(type),(uintmax_t)size) + 1;\n+\t\thdrlen = format_loose_header(hdr, sizeof(hdr), type, (uintmax_t)size);\n \t\tthe_hash_algo->init_fn(&c);\n \t\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n \t} else\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 8785b2ac806..446dea7c516 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -220,8 +220,8 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n \tif (seekback == (off_t) -1)\n \t\treturn error(\"cannot find the current offset\");\n \n-\theader_len = xsnprintf((char *)obuf, sizeof(obuf), \"%s %\" PRIuMAX,\n-\t\t\t       type_name(type), (uintmax_t)size) + 1;\n+\theader_len = format_loose_header((char *)obuf, sizeof(obuf),\n+\t\t\t\t\t type, (uintmax_t)size);\n \tthe_hash_algo->init_fn(&ctx);\n \tthe_hash_algo->update_fn(&ctx, obuf, header_len);\n \ndiff --git a/cache.h b/cache.h\nindex d5cafba17d4..ccece21a4a2 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1309,6 +1309,27 @@ enum unpack_loose_header_result unpack_loose_header(git_zstream *stream,\n \t\t\t\t\t\t    unsigned long bufsiz,\n \t\t\t\t\t\t    struct strbuf *hdrbuf);\n \n+/**\n+ * format_loose_header() is a thin wrapper around s xsnprintf() that\n+ * writes the initial \"<type> <obj-len>\" part of the loose object\n+ * header. It returns the size that snprintf() returns + 1.\n+ *\n+ * The format_loose_header_extended() function allows for writing a\n+ * type_name that's not one of the \"enum object_type\" types. This is\n+ * used for \"git hash-object --literally\". Pass in a OBJ_NONE as the\n+ * type, and a non-NULL \"type_str\" to do that.\n+ *\n+ * format_loose_header() is a convenience wrapper for\n+ * format_loose_header_extended().\n+ */\n+int format_loose_header_extended(char *str, size_t size, enum object_type type,\n+\t\t\t\t const char *type_str, size_t objsize);\n+static inline int format_loose_header(char *str, size_t size,\n+\t\t\t\t      enum object_type type, size_t objsize)\n+{\n+\treturn format_loose_header_extended(str, size, type, NULL, objsize);\n+}\n+\n /**\n  * parse_loose_header() parses the starting \"<type> <len>\\0\" of an\n  * object. If it doesn't follow that format -1 is returned. To check\ndiff --git a/http-push.c b/http-push.c\nindex 3309aaf004a..d1a8619e0af 100644\n--- a/http-push.c\n+++ b/http-push.c\n@@ -363,7 +363,7 @@ static void start_put(struct transfer_request *request)\n \tgit_zstream stream;\n \n \tunpacked = read_object_file(&request->obj->oid, &type, &len);\n-\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n+\thdrlen = format_loose_header(hdr, sizeof(hdr), type, (uintmax_t)len);\n \n \t/* Set it up */\n \tgit_deflate_init(&stream, zlib_compression_level);\ndiff --git a/object-file.c b/object-file.c\nindex eac67f6f5f9..d94609ee48d 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1009,6 +1009,14 @@ void *xmmap(void *start, size_t length,\n \treturn ret;\n }\n \n+int format_loose_header_extended(char *str, size_t size, enum object_type type,\n+\t\t\t\t const char *typestr, size_t objsize)\n+{\n+\tconst char *s = type == OBJ_NONE ? typestr : type_name(type);\n+\n+\treturn xsnprintf(str, size, \"%s %\"PRIuMAX, s, (uintmax_t)objsize) + 1;\n+}\n+\n /*\n  * With an in-core object data in \"map\", rehash it to make sure the\n  * object name actually matches \"oid\" to detect object corruption.\n@@ -1037,7 +1045,7 @@ int check_object_signature(struct repository *r, const struct object_id *oid,\n \t\treturn -1;\n \n \t/* Generate the header */\n-\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(obj_type), (uintmax_t)size) + 1;\n+\thdrlen = format_loose_header(hdr, sizeof(hdr), obj_type, size);\n \n \t/* Sha1.. */\n \tr->hash_algo->init_fn(&c);\n@@ -1737,7 +1745,7 @@ static void write_object_file_prepare(const struct git_hash_algo *algo,\n \tgit_hash_ctx c;\n \n \t/* Generate the header */\n-\t*hdrlen = xsnprintf(hdr, *hdrlen, \"%s %\"PRIuMAX , type, (uintmax_t)len)+1;\n+\t*hdrlen = format_loose_header_extended(hdr, *hdrlen, OBJ_NONE, type, len);\n \n \t/* Sha1.. */\n \talgo->init_fn(&c);\n@@ -2009,7 +2017,7 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n \tbuf = read_object(the_repository, oid, &type, &len);\n \tif (!buf)\n \t\treturn error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n-\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n+\thdrlen = format_loose_header(hdr, sizeof(hdr), type, len);\n \tret = write_loose_object(oid, hdr, hdrlen, buf, len, mtime, 0);\n \tfree(buf);\n \n-- \n2.34.1.1119.g606023410ba\n\n"},{"id":"444456","messageId":"41bd6eed-0e29-3b52-cfd4-9b9f9ac99c72@iee.email","threadId":"56672","inReplyTo":"RFC-patch-1.1-bda62567f6b-20211220T120740Z-avarab@gmail.com","subject":"Re: [RFC PATCH] object-file API: add a format_loose_header() function","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2021-12-20T12:48:55Z","receivedAt":"2021-12-20T12:48:59Z","isPatch":true,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"Hi Ævar,\n(catching up after a week away, and noticed your patch today..)\n\nOn 20/12/2021 12:10, Ævar Arnfjörð Bjarmason wrote:\n> Add a convenience function to wrap the xsnprintf() command that\n> generates loose object headers. This code was copy/pasted in various\n> parts of the codebase, let's define it in one place and re-use it from\n> there.\n>\n> All except one caller of it had a valid \"enum object_type\" for us,\n> it's only write_object_file_prepare() which might need to deal with\n> \"git hash-object --literally\" and a potential garbage type. Let's have\n> the primary API use an \"enum object_type\", and define an *_extended()\n> function that can take an arbitrary \"const char *\" for the type.\n\nI recently completed a PR in the Git for Windows build that is focused on\n\"git hash-object --literally\" as a starter for LLP64 large file (>4GB)\ncompatibility.\n(https://github.com/git-for-windows/git/pull/3533), which Dscho has\nmerged (cc'd).\n\nI'm not sure that the `extended` version will work as expected across\nthe test suite\nas multiple fake object types are tried, though I only skimmed the patch.\n\nI'd support the general thrust, but just wanted to synchronise any changes.\n\nPhilip\n>\n> See [1] for the discussion that prompted this patch, i.e. new code in\n> object-file.c that wanted to copy/paste the xsnprintf() invocation.\n>\n> 1. https://lore.kernel.org/git/211213.86bl1l9bfz.gmgdl@evledraar.gmail.com/\n>\n> Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> ---\n>\n> On Fri, Dec 17 2021, Han Xin wrote:\n>\n>> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>>\n>> There are 3 places where \"xsnprintf\" is used to generate the object\n>> header, and I originally planned to add a fourth in the latter patch.\n>>\n>> According to Ævar Arnfjörð Bjarmason’s suggestion, although it's just\n>> one line, it's also code that's very central to git, so reafactor them\n>> into a function which will help later readability.\n>>\n>> Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n>> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> I came up with this after my comment on the earlier round suggesting\n> to factor out that header formatting. I don't know if this more\n> thorough approach is worth it or if you'd like to replace your change\n> with this one, but just posting it here as an RFC.\n>\n>  builtin/index-pack.c |  3 +--\n>  bulk-checkin.c       |  4 ++--\n>  cache.h              | 21 +++++++++++++++++++++\n>  http-push.c          |  2 +-\n>  object-file.c        | 14 +++++++++++---\n>  5 files changed, 36 insertions(+), 8 deletions(-)\n>\n> diff --git a/builtin/index-pack.c b/builtin/index-pack.c\n> index c23d01de7dc..900c6539f68 100644\n> --- a/builtin/index-pack.c\n> +++ b/builtin/index-pack.c\n> @@ -449,8 +449,7 @@ static void *unpack_entry_data(off_t offset, unsigned long size,\n>  \tint hdrlen;\n>  \n>  \tif (!is_delta_type(type)) {\n> -\t\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX,\n> -\t\t\t\t   type_name(type),(uintmax_t)size) + 1;\n> +\t\thdrlen = format_loose_header(hdr, sizeof(hdr), type, (uintmax_t)size);\n>  \t\tthe_hash_algo->init_fn(&c);\n>  \t\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n>  \t} else\n> diff --git a/bulk-checkin.c b/bulk-checkin.c\n> index 8785b2ac806..446dea7c516 100644\n> --- a/bulk-checkin.c\n> +++ b/bulk-checkin.c\n> @@ -220,8 +220,8 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n>  \tif (seekback == (off_t) -1)\n>  \t\treturn error(\"cannot find the current offset\");\n>  \n> -\theader_len = xsnprintf((char *)obuf, sizeof(obuf), \"%s %\" PRIuMAX,\n> -\t\t\t       type_name(type), (uintmax_t)size) + 1;\n> +\theader_len = format_loose_header((char *)obuf, sizeof(obuf),\n> +\t\t\t\t\t type, (uintmax_t)size);\n>  \tthe_hash_algo->init_fn(&ctx);\n>  \tthe_hash_algo->update_fn(&ctx, obuf, header_len);\n>  \n> diff --git a/cache.h b/cache.h\n> index d5cafba17d4..ccece21a4a2 100644\n> --- a/cache.h\n> +++ b/cache.h\n> @@ -1309,6 +1309,27 @@ enum unpack_loose_header_result unpack_loose_header(git_zstream *stream,\n>  \t\t\t\t\t\t    unsigned long bufsiz,\n>  \t\t\t\t\t\t    struct strbuf *hdrbuf);\n>  \n> +/**\n> + * format_loose_header() is a thin wrapper around s xsnprintf() that\n> + * writes the initial \"<type> <obj-len>\" part of the loose object\n> + * header. It returns the size that snprintf() returns + 1.\n> + *\n> + * The format_loose_header_extended() function allows for writing a\n> + * type_name that's not one of the \"enum object_type\" types. This is\n> + * used for \"git hash-object --literally\". Pass in a OBJ_NONE as the\n> + * type, and a non-NULL \"type_str\" to do that.\n> + *\n> + * format_loose_header() is a convenience wrapper for\n> + * format_loose_header_extended().\n> + */\n> +int format_loose_header_extended(char *str, size_t size, enum object_type type,\n> +\t\t\t\t const char *type_str, size_t objsize);\n> +static inline int format_loose_header(char *str, size_t size,\n> +\t\t\t\t      enum object_type type, size_t objsize)\n> +{\n> +\treturn format_loose_header_extended(str, size, type, NULL, objsize);\n> +}\n> +\n>  /**\n>   * parse_loose_header() parses the starting \"<type> <len>\\0\" of an\n>   * object. If it doesn't follow that format -1 is returned. To check\n> diff --git a/http-push.c b/http-push.c\n> index 3309aaf004a..d1a8619e0af 100644\n> --- a/http-push.c\n> +++ b/http-push.c\n> @@ -363,7 +363,7 @@ static void start_put(struct transfer_request *request)\n>  \tgit_zstream stream;\n>  \n>  \tunpacked = read_object_file(&request->obj->oid, &type, &len);\n> -\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n> +\thdrlen = format_loose_header(hdr, sizeof(hdr), type, (uintmax_t)len);\n>  \n>  \t/* Set it up */\n>  \tgit_deflate_init(&stream, zlib_compression_level);\n> diff --git a/object-file.c b/object-file.c\n> index eac67f6f5f9..d94609ee48d 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1009,6 +1009,14 @@ void *xmmap(void *start, size_t length,\n>  \treturn ret;\n>  }\n>  \n> +int format_loose_header_extended(char *str, size_t size, enum object_type type,\n> +\t\t\t\t const char *typestr, size_t objsize)\n> +{\n> +\tconst char *s = type == OBJ_NONE ? typestr : type_name(type);\n> +\n> +\treturn xsnprintf(str, size, \"%s %\"PRIuMAX, s, (uintmax_t)objsize) + 1;\n> +}\n> +\n>  /*\n>   * With an in-core object data in \"map\", rehash it to make sure the\n>   * object name actually matches \"oid\" to detect object corruption.\n> @@ -1037,7 +1045,7 @@ int check_object_signature(struct repository *r, const struct object_id *oid,\n>  \t\treturn -1;\n>  \n>  \t/* Generate the header */\n> -\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(obj_type), (uintmax_t)size) + 1;\n> +\thdrlen = format_loose_header(hdr, sizeof(hdr), obj_type, size);\n>  \n>  \t/* Sha1.. */\n>  \tr->hash_algo->init_fn(&c);\n> @@ -1737,7 +1745,7 @@ static void write_object_file_prepare(const struct git_hash_algo *algo,\n>  \tgit_hash_ctx c;\n>  \n>  \t/* Generate the header */\n> -\t*hdrlen = xsnprintf(hdr, *hdrlen, \"%s %\"PRIuMAX , type, (uintmax_t)len)+1;\n> +\t*hdrlen = format_loose_header_extended(hdr, *hdrlen, OBJ_NONE, type, len);\n>  \n>  \t/* Sha1.. */\n>  \talgo->init_fn(&c);\n> @@ -2009,7 +2017,7 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n>  \tbuf = read_object(the_repository, oid, &type, &len);\n>  \tif (!buf)\n>  \t\treturn error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n> -\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n> +\thdrlen = format_loose_header(hdr, sizeof(hdr), type, len);\n>  \tret = write_loose_object(oid, hdr, hdrlen, buf, len, mtime, 0);\n>  \tfree(buf);\n>  \n\n"},{"id":"444552","messageId":"xmqqilviud6e.fsf@gitster.g","threadId":"56672","inReplyTo":"RFC-patch-1.1-bda62567f6b-20211220T120740Z-avarab@gmail.com","subject":"Re: [RFC PATCH] object-file API: add a format_loose_header() function","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-12-20T22:25:13Z","receivedAt":"2021-12-20T22:25:20Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason  <avarab@gmail.com> writes:\n\n> Add a convenience function to wrap the xsnprintf() command that\n> generates loose object headers. This code was copy/pasted in various\n> parts of the codebase, let's define it in one place and re-use it from\n> there.\n> ...\n> +/**\n> + * format_loose_header() is a thin wrapper around s xsnprintf() that\n\nThe name should have \"object\" somewhere in it.  Not all readers can\nbe expected to know that you meant \"loose\" to be an acceptable short\nhand for \"loose object\".\n\nThat nit aside, I think it is a good idea to give people a common\nhelper function to call.  I am undecided if it is a good idea to\nmake it take enum or \"const char *\"; most everybody should be able\nto say\n\n\tformat_object_header(type_name(OBJ_COMMIT), ...)\n\njust fine, so two variants might be overkill, just to allow \n\n\tformat_object_header(OBJ_COMMIT, ...)\n\nand to forbid\n\n\tformat_object_header(\"connit\", ...)\n\nI dunno.\n"},{"id":"444569","messageId":"211221.86wnjyspd3.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"xmqqilviud6e.fsf@gitster.g","subject":"Re: [RFC PATCH] object-file API: add a format_loose_header() function","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-21T01:42:44Z","receivedAt":"2021-12-21T01:45:00Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Mon, Dec 20 2021, Junio C Hamano wrote:\n\n> Ævar Arnfjörð Bjarmason  <avarab@gmail.com> writes:\n>\n>> Add a convenience function to wrap the xsnprintf() command that\n>> generates loose object headers. This code was copy/pasted in various\n>> parts of the codebase, let's define it in one place and re-use it from\n>> there.\n>> ...\n>> +/**\n>> + * format_loose_header() is a thin wrapper around s xsnprintf() that\n>\n> The name should have \"object\" somewhere in it.  Not all readers can\n> be expected to know that you meant \"loose\" to be an acceptable short\n> hand for \"loose object\".\n\n*nod*\n\n> That nit aside, I think it is a good idea to give people a common\n> helper function to call.  I am undecided if it is a good idea to\n> make it take enum or \"const char *\"; most everybody should be able\n> to say\n>\n> \tformat_object_header(type_name(OBJ_COMMIT), ...)\n>\n> just fine, so two variants might be overkill, just to allow \n>\n> \tformat_object_header(OBJ_COMMIT, ...)\n>\n> and to forbid\n>\n> \tformat_object_header(\"connit\", ...)\n>\n> I dunno.\n\nUltimately only a single API caller in hash-object.c really cares about\nsomething else than the enum.\n\nI've got some patches locally to convert e.g. write_object_file() to use\nthe enum, and it removes the need for some callers to convert enum to\nchar *, only to have other things convert it back.\n\nSo I think for any new APIs it makes sense to work towards sidelining\nthe hash-object.c --literally caller.\n"},{"id":"444572","messageId":"xmqqmtkuso5f.fsf@gitster.g","threadId":"56672","inReplyTo":"211221.86wnjyspd3.gmgdl@evledraar.gmail.com","subject":"Re: [RFC PATCH] object-file API: add a format_loose_header() function","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-12-21T02:11:08Z","receivedAt":"2021-12-21T02:11:17Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n\n> I've got some patches locally to convert e.g. write_object_file() to use\n> the enum, and it removes the need for some callers to convert enum to\n> char *, only to have other things convert it back.\n>\n> So I think for any new APIs it makes sense to work towards sidelining\n> the hash-object.c --literally caller.\n\nYour logic is backwards to argue \"because I did something this way,\nit makes sense to do it this way\"?\n"},{"id":"444573","messageId":"211221.86sfumsn2d.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"xmqqmtkuso5f.fsf@gitster.g","subject":"Re: [RFC PATCH] object-file API: add a format_loose_header() function","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-21T02:27:44Z","receivedAt":"2021-12-21T02:34:39Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Mon, Dec 20 2021, Junio C Hamano wrote:\n\n> Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n>\n>> I've got some patches locally to convert e.g. write_object_file() to use\n>> the enum, and it removes the need for some callers to convert enum to\n>> char *, only to have other things convert it back.\n>>\n>> So I think for any new APIs it makes sense to work towards sidelining\n>> the hash-object.c --literally caller.\n>\n> Your logic is backwards to argue \"because I did something this way,\n> it makes sense to do it this way\"?\n\nNo, it's that if you look at the write_object_file() and\nhash_object_file() callers in-tree now many, including in object-file.c\nitself are taking an \"enum object_type\" only to convert it to a string,\nand then we'll in turn sometimes convert that to the \"enum object_type\"\nagain at some lower level.\n\nThat API inconsistency dates back to at least Linus's a733cb606fe\n(Change pack file format. Hopefully for the last time., 2005-06-28).\n\nI'm just pointing out that I have local patches that prove that a lot of\nback & forth is done for no good reason, and that this is one of the\ncodepaths that's tangentally involved. So it makes sense in this case to\nmake any new API take \"enum object_type\" as the primary interface.\n"},{"id":"444597","messageId":"CAO0brD0KLM7wiZJUSLTxieov1BM347iitkMSF4_+XG5C-aBAyQ@mail.gmail.com","threadId":"56672","inReplyTo":"RFC-patch-1.1-bda62567f6b-20211220T120740Z-avarab@gmail.com","subject":"Re: [RFC PATCH] object-file API: add a format_loose_header() function","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-21T11:43:30Z","receivedAt":"2021-12-21T11:43:45Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Mon, Dec 20, 2021 at 8:10 PM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n> I came up with this after my comment on the earlier round suggesting\n> to factor out that header formatting. I don't know if this more\n> thorough approach is worth it or if you'd like to replace your change\n> with this one, but just posting it here as an RFC.\n>\n\nI will take this patch and rename the function name from\n\"format_loose_header()\" to \"format_object_header()\".\n\nThanks\n-Han Xin\n"},{"id":"444598","messageId":"20211221115201.12120-1-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211217112629.12334-1-chiyutianyi@gmail.com","subject":"[PATCH v7 0/5] unpack large blobs in stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-21T11:51:56Z","receivedAt":"2021-12-21T11:54:18Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nChanges since v6:\n* Remove \"object-file.c: release strbuf in write_loose_object()\" which is not\n  needed anymore. Thanks to René Scharfe[1] for reporting this.\n\n* Reorder the patch series and put \"unpack-objects.c: add dry_run mode for get_data()\"\n  and its testcases to the front.\n\n* Replace \"refactor object header generation into a function\" with\n  \"object-file API: add a format_object_header() function\" sugguested by\n  Ævar Arnfjörð Bjarmason[2].\n\n* Export \"write_stream_object_file()\" instead of \"reusing write_object_file_flags()\"\n  sugguested by René Scharfe[3]. The new flag \"HASH_STREAM\" has been removed.\n\n* Fix the directory creation error and the \"strbuf dir\" leak in\n  \"write_stream_object_file()\".\n\n* Change \"unsigned long size\" to \"size_t size\" in \"write_stream_blob()\" and\n  \"get_data()\" in \"unpack-objects.c\".\n\n1. https://lore.kernel.org/git/c860c56f-ce25-4391-7f65-50c9d5d80c2c@web.de/\n2. https://lore.kernel.org/git/RFC-patch-1.1-bda62567f6b-20211220T120740Z-avarab@gmail.com/\n3. https://lore.kernel.org/git/e959e4f1-7500-5f6b-5bd2-2f060287eeff@web.de/\n\nHan Xin (4):\n  unpack-objects.c: add dry_run mode for get_data()\n  object-file.c: refactor write_loose_object() to reuse in stream\n    version\n  object-file.c: add \"write_stream_object_file()\" to support read in\n    stream\n  unpack-objects: unpack_non_delta_entry() read data in a stream\n\nÆvar Arnfjörð Bjarmason (1):\n  object-file API: add a format_object_header() function\n\n Documentation/config/core.txt       |  11 ++\n builtin/index-pack.c                |   3 +-\n builtin/unpack-objects.c            |  94 ++++++++++++-\n bulk-checkin.c                      |   4 +-\n cache.h                             |  22 +++\n config.c                            |   5 +\n environment.c                       |   1 +\n http-push.c                         |   2 +-\n object-file.c                       | 199 ++++++++++++++++++++++------\n object-store.h                      |   9 ++\n t/t5590-unpack-non-delta-objects.sh |  91 +++++++++++++\n 11 files changed, 392 insertions(+), 49 deletions(-)\n create mode 100755 t/t5590-unpack-non-delta-objects.sh\n\nRange-diff against v6:\n1:  59d35dac5f < -:  ---------- object-file.c: release strbuf in write_loose_object()\n2:  2174a6cbad < -:  ---------- object-file.c: refactor object header generation into a function\n5:  1acbb6e849 ! 1:  a8f232f553 unpack-objects.c: add dry_run mode for get_data()\n    @@ builtin/unpack-objects.c: static void use(int bytes)\n      }\n      \n     -static void *get_data(unsigned long size)\n    -+static void *get_data(unsigned long size, int dry_run)\n    ++static void *get_data(size_t size, int dry_run)\n      {\n      \tgit_zstream stream;\n     -\tvoid *buf = xmallocz(size);\n    -+\tunsigned long bufsize;\n    ++\tsize_t bufsize;\n     +\tvoid *buf;\n      \n      \tmemset(&stream, 0, sizeof(stream));\n    @@ builtin/unpack-objects.c: static void unpack_delta_entry(enum object_type type,\n      \t\tif (dry_run || !delta_data) {\n      \t\t\tfree(delta_data);\n      \t\t\treturn;\n    +\n    + ## t/t5590-unpack-non-delta-objects.sh (new) ##\n    +@@\n    ++#!/bin/sh\n    ++#\n    ++# Copyright (c) 2021 Han Xin\n    ++#\n    ++\n    ++test_description='Test unpack-objects with non-delta objects'\n    ++\n    ++GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n    ++export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n    ++\n    ++. ./test-lib.sh\n    ++\n    ++prepare_dest () {\n    ++\ttest_when_finished \"rm -rf dest.git\" &&\n    ++\tgit init --bare dest.git\n    ++}\n    ++\n    ++test_expect_success \"setup repo with big blobs (1.5 MB)\" '\n    ++\ttest-tool genrandom foo 1500000 >big-blob &&\n    ++\ttest_commit --append foo big-blob &&\n    ++\ttest-tool genrandom bar 1500000 >big-blob &&\n    ++\ttest_commit --append bar big-blob &&\n    ++\t(\n    ++\t\tcd .git &&\n    ++\t\tfind objects/?? -type f | sort\n    ++\t) >expect &&\n    ++\tPACK=$(echo main | git pack-objects --revs test)\n    ++'\n    ++\n    ++test_expect_success 'setup env: GIT_ALLOC_LIMIT to 1MB' '\n    ++\tGIT_ALLOC_LIMIT=1m &&\n    ++\texport GIT_ALLOC_LIMIT\n    ++'\n    ++\n    ++test_expect_success 'fail to unpack-objects: cannot allocate' '\n    ++\tprepare_dest &&\n    ++\ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n    ++\tgrep \"fatal: attempting to allocate\" err &&\n    ++\t(\n    ++\t\tcd dest.git &&\n    ++\t\tfind objects/?? -type f | sort\n    ++\t) >actual &&\n    ++\ttest_file_not_empty actual &&\n    ++\t! test_cmp expect actual\n    ++'\n    ++\n    ++test_expect_success 'unpack-objects dry-run' '\n    ++\tprepare_dest &&\n    ++\tgit -C dest.git unpack-objects -n <test-$PACK.pack &&\n    ++\t(\n    ++\t\tcd dest.git &&\n    ++\t\tfind objects/ -type f\n    ++\t) >actual &&\n    ++\ttest_must_be_empty actual\n    ++'\n    ++\n    ++test_done\n-:  ---------- > 2:  0d2e0f3a00 object-file API: add a format_object_header() function\n3:  8a704ecc59 ! 3:  a571b8f16c object-file.c: refactor write_loose_object() to reuse in stream version\n    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *\n      \tloose_object_path(the_repository, &filename, oid);\n      \n     -\tfd = create_tmpfile(&tmp_file, filename.buf);\n    -+\tfd = create_tmpfile(&tmp_file, filename.buf, flags);\n    - \tif (fd < 0) {\n    +-\tif (fd < 0) {\n     -\t\tif (flags & HASH_SILENT)\n    --\t\t\tret = -1;\n    +-\t\t\treturn -1;\n     -\t\telse if (errno == EACCES)\n    --\t\t\tret = error(_(\"insufficient permission for adding an \"\n    --\t\t\t\t      \"object to repository database %s\"),\n    --\t\t\t\t    get_object_directory());\n    +-\t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n     -\t\telse\n    --\t\t\tret = error_errno(_(\"unable to create temporary file\"));\n    -+\t\tret = -1;\n    - \t\tgoto cleanup;\n    - \t}\n    - \n    +-\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n    +-\t}\n    +-\n     -\t/* Set it up */\n     -\tgit_deflate_init(&stream, zlib_compression_level);\n     -\tstream.next_out = compressed;\n     -\tstream.avail_out = sizeof(compressed);\n     -\tthe_hash_algo->init_fn(&c);\n    --\n    ++\tfd = create_tmpfile(&tmp_file, filename.buf, flags);\n    ++\tif (fd < 0)\n    ++\t\treturn -1;\n    + \n     -\t/* First header.. */\n     -\tstream.next_in = (unsigned char *)hdr;\n     -\tstream.avail_in = hdrlen;\n    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *\n     -\t\t\twarning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n     -\t}\n     -\n    --\tret = finalize_object_file(tmp_file.buf, filename.buf);\n    -+\tret = finalize_object_file_with_mtime(tmp_file.buf, filename.buf, mtime, flags);\n    - cleanup:\n    - \tstrbuf_release(&filename);\n    - \tstrbuf_release(&tmp_file);\n    +-\treturn finalize_object_file(tmp_file.buf, filename.buf);\n    ++\treturn finalize_object_file_with_mtime(tmp_file.buf, filename.buf,\n    ++\t\t\t\t\t       mtime, flags);\n    + }\n    + \n    + static int freshen_loose_object(const struct object_id *oid)\n4:  96f05632a2 ! 4:  1de06a8f5c object-file.c: make \"write_object_file_flags()\" to support read in stream\n    @@ Metadata\n     Author: Han Xin <hanxin.hx@alibaba-inc.com>\n     \n      ## Commit message ##\n    -    object-file.c: make \"write_object_file_flags()\" to support read in stream\n    +    object-file.c: add \"write_stream_object_file()\" to support read in stream\n     \n         We used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\n         entire contents of a blob object, no matter how big it is. This\n         implementation may consume all the memory and cause OOM.\n     \n    -    This can be improved by feeding data to \"stream_loose_object()\" in a\n    -    stream. The input stream is implemented as an interface.\n    -\n    -    When streaming a large blob object to \"write_loose_object()\", we have no\n    -    chance to run \"write_object_file_prepare()\" to calculate the oid in\n    -    advance. So we need to handle undetermined oid in a new function called\n    -    \"stream_loose_object()\".\n    +    This can be improved by feeding data to \"write_stream_object_file()\"\n    +    in a stream. The input stream is implemented as an interface.\n     \n    +    The difference with \"write_loose_object()\" is that we have no chance\n    +    to run \"write_object_file_prepare()\" to calculate the oid in advance.\n         In \"write_loose_object()\", we know the oid and we can write the\n         temporary file in the same directory as the final object, but for an\n         object with an undetermined oid, we don't know the exact directory for\n         the object, so we have to save the temporary file in \".git/objects/\"\n         directory instead.\n     \n    -    We will reuse \"write_object_file_flags()\" in \"unpack_non_delta_entry()\" to\n    -    read the entire data contents in stream, so a new flag \"HASH_STREAM\" is\n    -    added. When read in stream, we needn't prepare the \"oid\" before\n    -    \"write_loose_object()\", only generate the header.\n         \"freshen_packed_object()\" or \"freshen_loose_object()\" will be called\n    -    inside \"stream_loose_object()\" after obtaining the \"oid\".\n    +    inside \"write_stream_object_file()\" after obtaining the \"oid\".\n     \n    +    Helped-by: René Scharfe <l.s.r@web.de>\n         Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n         Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n         Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n     \n    - ## cache.h ##\n    -@@ cache.h: int ie_modified(struct index_state *, const struct cache_entry *, struct stat *,\n    - #define HASH_FORMAT_CHECK 2\n    - #define HASH_RENORMALIZE  4\n    - #define HASH_SILENT 8\n    -+#define HASH_STREAM 16\n    - int index_fd(struct index_state *istate, struct object_id *oid, int fd, struct stat *st, enum object_type type, const char *path, unsigned flags);\n    - int index_path(struct index_state *istate, struct object_id *oid, const char *path, struct stat *st, unsigned flags);\n    - \n    -\n      ## object-file.c ##\n     @@ object-file.c: static int freshen_packed_object(const struct object_id *oid)\n      \treturn 1;\n      }\n      \n    -+static int stream_loose_object(struct object_id *oid, char *hdr, int hdrlen,\n    -+\t\t\t       const struct input_stream *in_stream,\n    -+\t\t\t       unsigned long len, time_t mtime, unsigned flags)\n    ++int write_stream_object_file(struct input_stream *in_stream, size_t len,\n    ++\t\t\t     enum object_type type, time_t mtime,\n    ++\t\t\t     unsigned flags, struct object_id *oid)\n     +{\n    -+\tint fd, ret, err = 0, flush = 0;\n    ++\tint fd, ret, flush = 0;\n     +\tunsigned char compressed[4096];\n     +\tgit_zstream stream;\n     +\tgit_hash_ctx c;\n    @@ object-file.c: static int freshen_packed_object(const struct object_id *oid)\n     +\tstatic struct strbuf tmp_file = STRBUF_INIT;\n     +\tstatic struct strbuf filename = STRBUF_INIT;\n     +\tint dirlen;\n    ++\tchar hdr[MAX_HEADER_LEN];\n    ++\tint hdrlen = sizeof(hdr);\n     +\n    ++\t/* Since \"filename\" is defined as static, it will be reused. So reset it\n    ++\t * first before using it. */\n    ++\tstrbuf_reset(&filename);\n     +\t/* When oid is not determined, save tmp file to odb path. */\n     +\tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n     +\n     +\tfd = create_tmpfile(&tmp_file, filename.buf, flags);\n    -+\tif (fd < 0) {\n    -+\t\terr = -1;\n    -+\t\tgoto cleanup;\n    -+\t}\n    ++\tif (fd < 0)\n    ++\t\treturn -1;\n    ++\n    ++\thdrlen = format_object_header(hdr, hdrlen, type, len);\n     +\n     +\t/* Set it up and write header */\n     +\tsetup_stream_and_header(&stream, compressed, sizeof(compressed),\n    @@ object-file.c: static int freshen_packed_object(const struct object_id *oid)\n     +\n     +\tif (freshen_packed_object(oid) || freshen_loose_object(oid)) {\n     +\t\tunlink_or_warn(tmp_file.buf);\n    -+\t\tgoto cleanup;\n    ++\t\treturn 0;\n     +\t}\n     +\n     +\tloose_object_path(the_repository, &filename, oid);\n    @@ object-file.c: static int freshen_packed_object(const struct object_id *oid)\n     +\t\tstruct strbuf dir = STRBUF_INIT;\n     +\t\tstrbuf_add(&dir, filename.buf, dirlen - 1);\n     +\n    -+\t\tif (mkdir_in_gitdir(dir.buf) < 0) {\n    -+\t\t\terr = -1;\n    -+\t\t\tgoto cleanup;\n    ++\t\tif (mkdir_in_gitdir(dir.buf) && errno != EEXIST) {\n    ++\t\t\tret = error_errno(_(\"unable to create directory %s\"), dir.buf);\n    ++\t\t\tstrbuf_release(&dir);\n    ++\t\t\treturn ret;\n     +\t\t}\n    ++\t\tstrbuf_release(&dir);\n     +\t}\n     +\n    -+\terr = finalize_object_file_with_mtime(tmp_file.buf, filename.buf, mtime, flags);\n    -+cleanup:\n    -+\tstrbuf_release(&tmp_file);\n    -+\tstrbuf_release(&filename);\n    -+\treturn err;\n    ++\treturn finalize_object_file_with_mtime(tmp_file.buf, filename.buf, mtime, flags);\n     +}\n     +\n      int write_object_file_flags(const void *buf, unsigned long len,\n      \t\t\t    const char *type, struct object_id *oid,\n      \t\t\t    unsigned flags)\n    -@@ object-file.c: int write_object_file_flags(const void *buf, unsigned long len,\n    - \tchar hdr[MAX_HEADER_LEN];\n    - \tint hdrlen = sizeof(hdr);\n    - \n    -+\t/* When streaming a large blob object (marked as HASH_STREAM),\n    -+\t * we have no chance to run \"write_object_file_prepare()\" to\n    -+\t * calculate the \"oid\" in advance.  Call \"stream_loose_object()\"\n    -+\t * to write loose object in stream.\n    -+\t */\n    -+\tif (flags & HASH_STREAM) {\n    -+\t\thdrlen = generate_object_header(hdr, hdrlen, type, len);\n    -+\t\treturn stream_loose_object(oid, hdr, hdrlen, buf, len, 0, flags);\n    -+\t}\n    -+\n    - \t/* Normally if we have it in the pack then we do not bother writing\n    - \t * it out into .git/objects/??/?{38} file.\n    - \t */\n     \n      ## object-store.h ##\n     @@ object-store.h: struct object_directory {\n    @@ object-store.h: struct object_directory {\n      };\n      \n     +struct input_stream {\n    -+\tconst void *(*read)(const struct input_stream *, unsigned long *len);\n    ++\tconst void *(*read)(struct input_stream *, unsigned long *len);\n     +\tvoid *data;\n     +};\n     +\n      KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n      \tstruct object_directory *, 1, fspathhash, fspatheq)\n      \n    +@@ object-store.h: static inline int write_object_file(const void *buf, unsigned long len,\n    + \treturn write_object_file_flags(buf, len, type, oid, 0);\n    + }\n    + \n    ++int write_stream_object_file(struct input_stream *in_stream, size_t len,\n    ++\t\t\t     enum object_type type, time_t mtime,\n    ++\t\t\t     unsigned flags, struct object_id *oid);\n    ++\n    + int hash_object_file_literally(const void *buf, unsigned long len,\n    + \t\t\t       const char *type, struct object_id *oid,\n    + \t\t\t       unsigned flags);\n6:  476aaba527 ! 5:  e7b4e426ef unpack-objects: unpack_non_delta_entry() read data in a stream\n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n     +\tint status;\n     +};\n     +\n    -+static const void *feed_input_zstream(const struct input_stream *in_stream,\n    ++static const void *feed_input_zstream(struct input_stream *in_stream,\n     +\t\t\t\t      unsigned long *readlen)\n     +{\n     +\tstruct input_zstream_data *data = in_stream->data;\n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n     +\treturn data->buf;\n     +}\n     +\n    -+static void write_stream_blob(unsigned nr, unsigned long size)\n    ++static void write_stream_blob(unsigned nr, size_t size)\n     +{\n     +\tgit_zstream zstream;\n     +\tstruct input_zstream_data data;\n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n     +\tdata.zstream = &zstream;\n     +\tgit_inflate_init(&zstream);\n     +\n    -+\tif (write_object_file_flags(&in_stream, size,\n    -+\t\t\t\t    type_name(OBJ_BLOB),\n    -+\t\t\t\t    &obj_list[nr].oid,\n    -+\t\t\t\t    HASH_STREAM))\n    ++\tif (write_stream_object_file(&in_stream, size, OBJ_BLOB, 0, 0,\n    ++\t\t\t\t     &obj_list[nr].oid))\n     +\t\tdie(_(\"failed to write object in stream\"));\n     +\n     +\tif (zstream.total_out != size || data.status != Z_STREAM_END)\n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n     +\tgit_inflate_end(&zstream);\n     +\n     +\tif (strict) {\n    -+\t\tstruct blob *blob = lookup_blob(the_repository, &obj_list[nr].oid);\n    ++\t\tstruct blob *blob =\n    ++\t\t\tlookup_blob(the_repository, &obj_list[nr].oid);\n     +\t\tif (blob)\n     +\t\t\tblob->object.flags |= FLAG_WRITTEN;\n     +\t\telse\n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n     +\tvoid *buf;\n     +\n     +\t/* Write large blob in stream without allocating full buffer. */\n    -+\tif (!dry_run && type == OBJ_BLOB && size > big_file_streaming_threshold) {\n    ++\tif (!dry_run && type == OBJ_BLOB &&\n    ++\t    size > big_file_streaming_threshold) {\n     +\t\twrite_stream_blob(nr, size);\n     +\t\treturn;\n     +\t}\n    @@ environment.c: size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n      const char *editor_program;\n      const char *askpass_program;\n     \n    - ## t/t5590-unpack-non-delta-objects.sh (new) ##\n    -@@\n    -+#!/bin/sh\n    -+#\n    -+# Copyright (c) 2021 Han Xin\n    -+#\n    -+\n    -+test_description='Test unpack-objects when receive pack'\n    -+\n    -+GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n    -+export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n    -+\n    -+. ./test-lib.sh\n    -+\n    -+prepare_dest () {\n    -+\ttest_when_finished \"rm -rf dest.git\" &&\n    -+\tgit init --bare dest.git &&\n    -+\tgit -C dest.git config core.bigFileStreamingThreshold $1 &&\n    -+\tgit -C dest.git config core.bigFileThreshold $1\n    -+}\n    -+\n    -+test_expect_success \"setup repo with big blobs (1.5 MB)\" '\n    -+\ttest-tool genrandom foo 1500000 >big-blob &&\n    -+\ttest_commit --append foo big-blob &&\n    -+\ttest-tool genrandom bar 1500000 >big-blob &&\n    -+\ttest_commit --append bar big-blob &&\n    -+\t(\n    -+\t\tcd .git &&\n    -+\t\tfind objects/?? -type f | sort\n    -+\t) >expect &&\n    -+\tPACK=$(echo main | git pack-objects --revs test)\n    -+'\n    -+\n    -+test_expect_success 'setup env: GIT_ALLOC_LIMIT to 1MB' '\n    -+\tGIT_ALLOC_LIMIT=1m &&\n    -+\texport GIT_ALLOC_LIMIT\n    -+'\n    -+\n    -+test_expect_success 'fail to unpack-objects: cannot allocate' '\n    + ## t/t5590-unpack-non-delta-objects.sh ##\n    +@@ t/t5590-unpack-non-delta-objects.sh: export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n    + prepare_dest () {\n    + \ttest_when_finished \"rm -rf dest.git\" &&\n    + \tgit init --bare dest.git\n    ++\tif test -n \"$1\"\n    ++\tthen\n    ++\t\tgit -C dest.git config core.bigFileStreamingThreshold $1\n    ++\t\tgit -C dest.git config core.bigFileThreshold $1\n    ++\tfi\n    + }\n    + \n    + test_expect_success \"setup repo with big blobs (1.5 MB)\" '\n    +@@ t/t5590-unpack-non-delta-objects.sh: test_expect_success 'setup env: GIT_ALLOC_LIMIT to 1MB' '\n    + '\n    + \n    + test_expect_success 'fail to unpack-objects: cannot allocate' '\n    +-\tprepare_dest &&\n     +\tprepare_dest 2m &&\n    -+\ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n    -+\tgrep \"fatal: attempting to allocate\" err &&\n    -+\t(\n    -+\t\tcd dest.git &&\n    -+\t\tfind objects/?? -type f | sort\n    -+\t) >actual &&\n    -+\ttest_file_not_empty actual &&\n    -+\t! test_cmp expect actual\n    -+'\n    -+\n    + \ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n    + \tgrep \"fatal: attempting to allocate\" err &&\n    + \t(\n    +@@ t/t5590-unpack-non-delta-objects.sh: test_expect_success 'fail to unpack-objects: cannot allocate' '\n    + \t! test_cmp expect actual\n    + '\n    + \n     +test_expect_success 'unpack big object in stream' '\n     +\tprepare_dest 1m &&\n    ++\tmkdir -p dest.git/objects/05 &&\n     +\tgit -C dest.git unpack-objects <test-$PACK.pack &&\n     +\tgit -C dest.git fsck &&\n     +\t(\n    @@ t/t5590-unpack-non-delta-objects.sh (new)\n     +\ttest_must_be_empty actual\n     +'\n     +\n    -+test_expect_success 'unpack-objects dry-run' '\n    -+\tprepare_dest 1m &&\n    -+\tgit -C dest.git unpack-objects -n <test-$PACK.pack &&\n    -+\t(\n    -+\t\tcd dest.git &&\n    -+\t\tfind objects/ -type f\n    -+\t) >actual &&\n    -+\ttest_must_be_empty actual\n    -+'\n    -+\n    -+test_done\n    + test_expect_success 'unpack-objects dry-run' '\n    + \tprepare_dest &&\n    + \tgit -C dest.git unpack-objects -n <test-$PACK.pack &&\n-- \n2.34.1.52.g80008efde6.agit.6.5.6\n\n"},{"id":"444599","messageId":"20211221115201.12120-2-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211217112629.12334-1-chiyutianyi@gmail.com","subject":"[PATCH v7 1/5] unpack-objects.c: add dry_run mode for get_data()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-21T11:51:57Z","receivedAt":"2021-12-21T11:54:21Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nIn dry_run mode, \"get_data()\" is used to verify the inflation of data,\nand the returned buffer will not be used at all and will be freed\nimmediately. Even in dry_run mode, it is dangerous to allocate a\nfull-size buffer for a large blob object. Therefore, only allocate a\nlow memory footprint when calling \"get_data()\" in dry_run mode.\n\nSuggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n builtin/unpack-objects.c            | 23 +++++++++---\n t/t5590-unpack-non-delta-objects.sh | 57 +++++++++++++++++++++++++++++\n 2 files changed, 74 insertions(+), 6 deletions(-)\n create mode 100755 t/t5590-unpack-non-delta-objects.sh\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 4a9466295b..9104eb48da 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -96,15 +96,21 @@ static void use(int bytes)\n \tdisplay_throughput(progress, consumed_bytes);\n }\n \n-static void *get_data(unsigned long size)\n+static void *get_data(size_t size, int dry_run)\n {\n \tgit_zstream stream;\n-\tvoid *buf = xmallocz(size);\n+\tsize_t bufsize;\n+\tvoid *buf;\n \n \tmemset(&stream, 0, sizeof(stream));\n+\tif (dry_run && size > 8192)\n+\t\tbufsize = 8192;\n+\telse\n+\t\tbufsize = size;\n+\tbuf = xmallocz(bufsize);\n \n \tstream.next_out = buf;\n-\tstream.avail_out = size;\n+\tstream.avail_out = bufsize;\n \tstream.next_in = fill(1);\n \tstream.avail_in = len;\n \tgit_inflate_init(&stream);\n@@ -124,6 +130,11 @@ static void *get_data(unsigned long size)\n \t\t}\n \t\tstream.next_in = fill(1);\n \t\tstream.avail_in = len;\n+\t\tif (dry_run) {\n+\t\t\t/* reuse the buffer in dry_run mode */\n+\t\t\tstream.next_out = buf;\n+\t\t\tstream.avail_out = bufsize;\n+\t\t}\n \t}\n \tgit_inflate_end(&stream);\n \treturn buf;\n@@ -323,7 +334,7 @@ static void added_object(unsigned nr, enum object_type type,\n static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n \t\t\t\t   unsigned nr)\n {\n-\tvoid *buf = get_data(size);\n+\tvoid *buf = get_data(size, dry_run);\n \n \tif (!dry_run && buf)\n \t\twrite_object(nr, type, buf, size);\n@@ -357,7 +368,7 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \tif (type == OBJ_REF_DELTA) {\n \t\toidread(&base_oid, fill(the_hash_algo->rawsz));\n \t\tuse(the_hash_algo->rawsz);\n-\t\tdelta_data = get_data(delta_size);\n+\t\tdelta_data = get_data(delta_size, dry_run);\n \t\tif (dry_run || !delta_data) {\n \t\t\tfree(delta_data);\n \t\t\treturn;\n@@ -396,7 +407,7 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\tif (base_offset <= 0 || base_offset >= obj_list[nr].offset)\n \t\t\tdie(\"offset value out of bound for delta base object\");\n \n-\t\tdelta_data = get_data(delta_size);\n+\t\tdelta_data = get_data(delta_size, dry_run);\n \t\tif (dry_run || !delta_data) {\n \t\t\tfree(delta_data);\n \t\t\treturn;\ndiff --git a/t/t5590-unpack-non-delta-objects.sh b/t/t5590-unpack-non-delta-objects.sh\nnew file mode 100755\nindex 0000000000..48c4fb1ba3\n--- /dev/null\n+++ b/t/t5590-unpack-non-delta-objects.sh\n@@ -0,0 +1,57 @@\n+#!/bin/sh\n+#\n+# Copyright (c) 2021 Han Xin\n+#\n+\n+test_description='Test unpack-objects with non-delta objects'\n+\n+GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n+export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n+\n+. ./test-lib.sh\n+\n+prepare_dest () {\n+\ttest_when_finished \"rm -rf dest.git\" &&\n+\tgit init --bare dest.git\n+}\n+\n+test_expect_success \"setup repo with big blobs (1.5 MB)\" '\n+\ttest-tool genrandom foo 1500000 >big-blob &&\n+\ttest_commit --append foo big-blob &&\n+\ttest-tool genrandom bar 1500000 >big-blob &&\n+\ttest_commit --append bar big-blob &&\n+\t(\n+\t\tcd .git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >expect &&\n+\tPACK=$(echo main | git pack-objects --revs test)\n+'\n+\n+test_expect_success 'setup env: GIT_ALLOC_LIMIT to 1MB' '\n+\tGIT_ALLOC_LIMIT=1m &&\n+\texport GIT_ALLOC_LIMIT\n+'\n+\n+test_expect_success 'fail to unpack-objects: cannot allocate' '\n+\tprepare_dest &&\n+\ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n+\tgrep \"fatal: attempting to allocate\" err &&\n+\t(\n+\t\tcd dest.git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >actual &&\n+\ttest_file_not_empty actual &&\n+\t! test_cmp expect actual\n+'\n+\n+test_expect_success 'unpack-objects dry-run' '\n+\tprepare_dest &&\n+\tgit -C dest.git unpack-objects -n <test-$PACK.pack &&\n+\t(\n+\t\tcd dest.git &&\n+\t\tfind objects/ -type f\n+\t) >actual &&\n+\ttest_must_be_empty actual\n+'\n+\n+test_done\n-- \n2.34.1.52.g80008efde6.agit.6.5.6\n\n"},{"id":"444600","messageId":"20211221115201.12120-3-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211217112629.12334-1-chiyutianyi@gmail.com","subject":"[PATCH v7 2/5] object-file API: add a format_object_header() function","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-21T11:51:58Z","receivedAt":"2021-12-21T11:54:24Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n\nAdd a convenience function to wrap the xsnprintf() command that\ngenerates loose object headers. This code was copy/pasted in various\nparts of the codebase, let's define it in one place and re-use it from\nthere.\n\nAll except one caller of it had a valid \"enum object_type\" for us,\nit's only write_object_file_prepare() which might need to deal with\n\"git hash-object --literally\" and a potential garbage type. Let's have\nthe primary API use an \"enum object_type\", and define an *_extended()\nfunction that can take an arbitrary \"const char *\" for the type.\n\nSee [1] for the discussion that prompted this patch, i.e. new code in\nobject-file.c that wanted to copy/paste the xsnprintf() invocation.\n\n1. https://lore.kernel.org/git/211213.86bl1l9bfz.gmgdl@evledraar.gmail.com/\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n builtin/index-pack.c |  3 +--\n bulk-checkin.c       |  4 ++--\n cache.h              | 21 +++++++++++++++++++++\n http-push.c          |  2 +-\n object-file.c        | 14 +++++++++++---\n 5 files changed, 36 insertions(+), 8 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex c23d01de7d..4a765ddae6 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -449,8 +449,7 @@ static void *unpack_entry_data(off_t offset, unsigned long size,\n \tint hdrlen;\n \n \tif (!is_delta_type(type)) {\n-\t\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX,\n-\t\t\t\t   type_name(type),(uintmax_t)size) + 1;\n+\t\thdrlen = format_object_header(hdr, sizeof(hdr), type, (uintmax_t)size);\n \t\tthe_hash_algo->init_fn(&c);\n \t\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n \t} else\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 8785b2ac80..1733a1de4f 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -220,8 +220,8 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n \tif (seekback == (off_t) -1)\n \t\treturn error(\"cannot find the current offset\");\n \n-\theader_len = xsnprintf((char *)obuf, sizeof(obuf), \"%s %\" PRIuMAX,\n-\t\t\t       type_name(type), (uintmax_t)size) + 1;\n+\theader_len = format_object_header((char *)obuf, sizeof(obuf),\n+\t\t\t\t\t type, (uintmax_t)size);\n \tthe_hash_algo->init_fn(&ctx);\n \tthe_hash_algo->update_fn(&ctx, obuf, header_len);\n \ndiff --git a/cache.h b/cache.h\nindex cfba463aa9..64071a8d80 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1310,6 +1310,27 @@ enum unpack_loose_header_result unpack_loose_header(git_zstream *stream,\n \t\t\t\t\t\t    unsigned long bufsiz,\n \t\t\t\t\t\t    struct strbuf *hdrbuf);\n \n+/**\n+ * format_object_header() is a thin wrapper around s xsnprintf() that\n+ * writes the initial \"<type> <obj-len>\" part of the loose object\n+ * header. It returns the size that snprintf() returns + 1.\n+ *\n+ * The format_object_header_extended() function allows for writing a\n+ * type_name that's not one of the \"enum object_type\" types. This is\n+ * used for \"git hash-object --literally\". Pass in a OBJ_NONE as the\n+ * type, and a non-NULL \"type_str\" to do that.\n+ *\n+ * format_object_header() is a convenience wrapper for\n+ * format_object_header_extended().\n+ */\n+int format_object_header_extended(char *str, size_t size, enum object_type type,\n+\t\t\t\t const char *type_str, size_t objsize);\n+static inline int format_object_header(char *str, size_t size,\n+\t\t\t\t      enum object_type type, size_t objsize)\n+{\n+\treturn format_object_header_extended(str, size, type, NULL, objsize);\n+}\n+\n /**\n  * parse_loose_header() parses the starting \"<type> <len>\\0\" of an\n  * object. If it doesn't follow that format -1 is returned. To check\ndiff --git a/http-push.c b/http-push.c\nindex 3309aaf004..f55e316ff4 100644\n--- a/http-push.c\n+++ b/http-push.c\n@@ -363,7 +363,7 @@ static void start_put(struct transfer_request *request)\n \tgit_zstream stream;\n \n \tunpacked = read_object_file(&request->obj->oid, &type, &len);\n-\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n+\thdrlen = format_object_header(hdr, sizeof(hdr), type, (uintmax_t)len);\n \n \t/* Set it up */\n \tgit_deflate_init(&stream, zlib_compression_level);\ndiff --git a/object-file.c b/object-file.c\nindex eb1426f98c..6bba4766f9 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1006,6 +1006,14 @@ void *xmmap(void *start, size_t length,\n \treturn ret;\n }\n \n+int format_object_header_extended(char *str, size_t size, enum object_type type,\n+\t\t\t\t const char *typestr, size_t objsize)\n+{\n+\tconst char *s = type == OBJ_NONE ? typestr : type_name(type);\n+\n+\treturn xsnprintf(str, size, \"%s %\"PRIuMAX, s, (uintmax_t)objsize) + 1;\n+}\n+\n /*\n  * With an in-core object data in \"map\", rehash it to make sure the\n  * object name actually matches \"oid\" to detect object corruption.\n@@ -1034,7 +1042,7 @@ int check_object_signature(struct repository *r, const struct object_id *oid,\n \t\treturn -1;\n \n \t/* Generate the header */\n-\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(obj_type), (uintmax_t)size) + 1;\n+\thdrlen = format_object_header(hdr, sizeof(hdr), obj_type, size);\n \n \t/* Sha1.. */\n \tr->hash_algo->init_fn(&c);\n@@ -1734,7 +1742,7 @@ static void write_object_file_prepare(const struct git_hash_algo *algo,\n \tgit_hash_ctx c;\n \n \t/* Generate the header */\n-\t*hdrlen = xsnprintf(hdr, *hdrlen, \"%s %\"PRIuMAX , type, (uintmax_t)len)+1;\n+\t*hdrlen = format_object_header_extended(hdr, *hdrlen, OBJ_NONE, type, len);\n \n \t/* Sha1.. */\n \talgo->init_fn(&c);\n@@ -2006,7 +2014,7 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n \tbuf = read_object(the_repository, oid, &type, &len);\n \tif (!buf)\n \t\treturn error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n-\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n+\thdrlen = format_object_header(hdr, sizeof(hdr), type, len);\n \tret = write_loose_object(oid, hdr, hdrlen, buf, len, mtime, 0);\n \tfree(buf);\n \n-- \n2.34.1.52.g80008efde6.agit.6.5.6\n\n"},{"id":"444601","messageId":"20211221115201.12120-4-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211217112629.12334-1-chiyutianyi@gmail.com","subject":"[PATCH v7 3/5] object-file.c: refactor write_loose_object() to reuse in stream version","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-21T11:51:59Z","receivedAt":"2021-12-21T11:54:27Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWe used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\nentire contents of a blob object, no matter how big it is. This\nimplementation may consume all the memory and cause OOM.\n\nThis can be improved by feeding data to \"stream_loose_object()\" in\nstream instead of read into the whole buf.\n\nAs this new method \"stream_loose_object()\" has many similarities with\n\"write_loose_object()\", we split up \"write_loose_object()\" into some\nsteps:\n 1. Figuring out a path for the (temp) object file.\n 2. Creating the tempfile.\n 3. Setting up zlib and write header.\n 4. Write object data and handle errors.\n 5. Optionally, do someting after write, maybe force a loose object if\n\"mtime\".\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c | 100 ++++++++++++++++++++++++++++++++------------------\n 1 file changed, 65 insertions(+), 35 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 6bba4766f9..e048f3d39e 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1751,6 +1751,25 @@ static void write_object_file_prepare(const struct git_hash_algo *algo,\n \talgo->final_oid_fn(oid, &c);\n }\n \n+/*\n+ * Move the just written object with proper mtime into its final resting place.\n+ */\n+static int finalize_object_file_with_mtime(const char *tmpfile,\n+\t\t\t\t\t   const char *filename,\n+\t\t\t\t\t   time_t mtime,\n+\t\t\t\t\t   unsigned flags)\n+{\n+\tstruct utimbuf utb;\n+\n+\tif (mtime) {\n+\t\tutb.actime = mtime;\n+\t\tutb.modtime = mtime;\n+\t\tif (utime(tmpfile, &utb) < 0 && !(flags & HASH_SILENT))\n+\t\t\twarning_errno(_(\"failed utime() on %s\"), tmpfile);\n+\t}\n+\treturn finalize_object_file(tmpfile, filename);\n+}\n+\n /*\n  * Move the just written object into its final resting place.\n  */\n@@ -1836,7 +1855,8 @@ static inline int directory_size(const char *filename)\n  * We want to avoid cross-directory filename renames, because those\n  * can have problems on various filesystems (FAT, NFS, Coda).\n  */\n-static int create_tmpfile(struct strbuf *tmp, const char *filename)\n+static int create_tmpfile(struct strbuf *tmp, const char *filename,\n+\t\t\t  unsigned flags)\n {\n \tint fd, dirlen = directory_size(filename);\n \n@@ -1844,7 +1864,9 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n \tstrbuf_add(tmp, filename, dirlen);\n \tstrbuf_addstr(tmp, \"tmp_obj_XXXXXX\");\n \tfd = git_mkstemp_mode(tmp->buf, 0444);\n-\tif (fd < 0 && dirlen && errno == ENOENT) {\n+\tdo {\n+\t\tif (fd >= 0 || !dirlen || errno != ENOENT)\n+\t\t\tbreak;\n \t\t/*\n \t\t * Make sure the directory exists; note that the contents\n \t\t * of the buffer are undefined after mkstemp returns an\n@@ -1854,17 +1876,48 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n \t\tstrbuf_reset(tmp);\n \t\tstrbuf_add(tmp, filename, dirlen - 1);\n \t\tif (mkdir(tmp->buf, 0777) && errno != EEXIST)\n-\t\t\treturn -1;\n+\t\t\tbreak;\n \t\tif (adjust_shared_perm(tmp->buf))\n-\t\t\treturn -1;\n+\t\t\tbreak;\n \n \t\t/* Try again */\n \t\tstrbuf_addstr(tmp, \"/tmp_obj_XXXXXX\");\n \t\tfd = git_mkstemp_mode(tmp->buf, 0444);\n+\t} while (0);\n+\n+\tif (fd < 0 && !(flags & HASH_SILENT)) {\n+\t\tif (errno == EACCES)\n+\t\t\treturn error(_(\"insufficient permission for adding an \"\n+\t\t\t\t       \"object to repository database %s\"),\n+\t\t\t\t     get_object_directory());\n+\t\telse\n+\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n \t}\n+\n \treturn fd;\n }\n \n+static void setup_stream_and_header(git_zstream *stream,\n+\t\t\t\t    unsigned char *compressed,\n+\t\t\t\t    unsigned long compressed_size,\n+\t\t\t\t    git_hash_ctx *c,\n+\t\t\t\t    char *hdr,\n+\t\t\t\t    int hdrlen)\n+{\n+\t/* Set it up */\n+\tgit_deflate_init(stream, zlib_compression_level);\n+\tstream->next_out = compressed;\n+\tstream->avail_out = compressed_size;\n+\tthe_hash_algo->init_fn(c);\n+\n+\t/* First header.. */\n+\tstream->next_in = (unsigned char *)hdr;\n+\tstream->avail_in = hdrlen;\n+\twhile (git_deflate(stream, 0) == Z_OK)\n+\t\t; /* nothing */\n+\tthe_hash_algo->update_fn(c, hdr, hdrlen);\n+}\n+\n static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\t\t      int hdrlen, const void *buf, unsigned long len,\n \t\t\t      time_t mtime, unsigned flags)\n@@ -1879,28 +1932,13 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \n \tloose_object_path(the_repository, &filename, oid);\n \n-\tfd = create_tmpfile(&tmp_file, filename.buf);\n-\tif (fd < 0) {\n-\t\tif (flags & HASH_SILENT)\n-\t\t\treturn -1;\n-\t\telse if (errno == EACCES)\n-\t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n-\t\telse\n-\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n-\t}\n-\n-\t/* Set it up */\n-\tgit_deflate_init(&stream, zlib_compression_level);\n-\tstream.next_out = compressed;\n-\tstream.avail_out = sizeof(compressed);\n-\tthe_hash_algo->init_fn(&c);\n+\tfd = create_tmpfile(&tmp_file, filename.buf, flags);\n+\tif (fd < 0)\n+\t\treturn -1;\n \n-\t/* First header.. */\n-\tstream.next_in = (unsigned char *)hdr;\n-\tstream.avail_in = hdrlen;\n-\twhile (git_deflate(&stream, 0) == Z_OK)\n-\t\t; /* nothing */\n-\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n+\t/* Set it up and write header */\n+\tsetup_stream_and_header(&stream, compressed, sizeof(compressed),\n+\t\t\t\t&c, hdr, hdrlen);\n \n \t/* Then the data itself.. */\n \tstream.next_in = (void *)buf;\n@@ -1929,16 +1967,8 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \n \tclose_loose_object(fd);\n \n-\tif (mtime) {\n-\t\tstruct utimbuf utb;\n-\t\tutb.actime = mtime;\n-\t\tutb.modtime = mtime;\n-\t\tif (utime(tmp_file.buf, &utb) < 0 &&\n-\t\t    !(flags & HASH_SILENT))\n-\t\t\twarning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n-\t}\n-\n-\treturn finalize_object_file(tmp_file.buf, filename.buf);\n+\treturn finalize_object_file_with_mtime(tmp_file.buf, filename.buf,\n+\t\t\t\t\t       mtime, flags);\n }\n \n static int freshen_loose_object(const struct object_id *oid)\n-- \n2.34.1.52.g80008efde6.agit.6.5.6\n\n"},{"id":"444602","messageId":"20211221115201.12120-5-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211217112629.12334-1-chiyutianyi@gmail.com","subject":"[PATCH v7 4/5] object-file.c: add \"write_stream_object_file()\" to support read in stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-21T11:52:00Z","receivedAt":"2021-12-21T11:54:33Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWe used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\nentire contents of a blob object, no matter how big it is. This\nimplementation may consume all the memory and cause OOM.\n\nThis can be improved by feeding data to \"write_stream_object_file()\"\nin a stream. The input stream is implemented as an interface.\n\nThe difference with \"write_loose_object()\" is that we have no chance\nto run \"write_object_file_prepare()\" to calculate the oid in advance.\nIn \"write_loose_object()\", we know the oid and we can write the\ntemporary file in the same directory as the final object, but for an\nobject with an undetermined oid, we don't know the exact directory for\nthe object, so we have to save the temporary file in \".git/objects/\"\ndirectory instead.\n\n\"freshen_packed_object()\" or \"freshen_loose_object()\" will be called\ninside \"write_stream_object_file()\" after obtaining the \"oid\".\n\nHelped-by: René Scharfe <l.s.r@web.de>\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c  | 85 ++++++++++++++++++++++++++++++++++++++++++++++++++\n object-store.h |  9 ++++++\n 2 files changed, 94 insertions(+)\n\ndiff --git a/object-file.c b/object-file.c\nindex e048f3d39e..d0573e2a61 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1989,6 +1989,91 @@ static int freshen_packed_object(const struct object_id *oid)\n \treturn 1;\n }\n \n+int write_stream_object_file(struct input_stream *in_stream, size_t len,\n+\t\t\t     enum object_type type, time_t mtime,\n+\t\t\t     unsigned flags, struct object_id *oid)\n+{\n+\tint fd, ret, flush = 0;\n+\tunsigned char compressed[4096];\n+\tgit_zstream stream;\n+\tgit_hash_ctx c;\n+\tstruct object_id parano_oid;\n+\tstatic struct strbuf tmp_file = STRBUF_INIT;\n+\tstatic struct strbuf filename = STRBUF_INIT;\n+\tint dirlen;\n+\tchar hdr[MAX_HEADER_LEN];\n+\tint hdrlen = sizeof(hdr);\n+\n+\t/* Since \"filename\" is defined as static, it will be reused. So reset it\n+\t * first before using it. */\n+\tstrbuf_reset(&filename);\n+\t/* When oid is not determined, save tmp file to odb path. */\n+\tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n+\n+\tfd = create_tmpfile(&tmp_file, filename.buf, flags);\n+\tif (fd < 0)\n+\t\treturn -1;\n+\n+\thdrlen = format_object_header(hdr, hdrlen, type, len);\n+\n+\t/* Set it up and write header */\n+\tsetup_stream_and_header(&stream, compressed, sizeof(compressed),\n+\t\t\t\t&c, hdr, hdrlen);\n+\n+\t/* Then the data itself.. */\n+\tdo {\n+\t\tunsigned char *in0 = stream.next_in;\n+\t\tif (!stream.avail_in) {\n+\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n+\t\t\tstream.next_in = (void *)in;\n+\t\t\tin0 = (unsigned char *)in;\n+\t\t\t/* All data has been read. */\n+\t\t\tif (len + hdrlen == stream.total_in + stream.avail_in)\n+\t\t\t\tflush = Z_FINISH;\n+\t\t}\n+\t\tret = git_deflate(&stream, flush);\n+\t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n+\t\tif (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n+\t\t\tdie(_(\"unable to write loose object file\"));\n+\t\tstream.next_out = compressed;\n+\t\tstream.avail_out = sizeof(compressed);\n+\t} while (ret == Z_OK || ret == Z_BUF_ERROR);\n+\n+\tif (ret != Z_STREAM_END)\n+\t\tdie(_(\"unable to deflate new object streamingly (%d)\"), ret);\n+\tret = git_deflate_end_gently(&stream);\n+\tif (ret != Z_OK)\n+\t\tdie(_(\"deflateEnd on object streamingly failed (%d)\"), ret);\n+\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n+\n+\tclose_loose_object(fd);\n+\n+\toidcpy(oid, &parano_oid);\n+\n+\tif (freshen_packed_object(oid) || freshen_loose_object(oid)) {\n+\t\tunlink_or_warn(tmp_file.buf);\n+\t\treturn 0;\n+\t}\n+\n+\tloose_object_path(the_repository, &filename, oid);\n+\n+\t/* We finally know the object path, and create the missing dir. */\n+\tdirlen = directory_size(filename.buf);\n+\tif (dirlen) {\n+\t\tstruct strbuf dir = STRBUF_INIT;\n+\t\tstrbuf_add(&dir, filename.buf, dirlen - 1);\n+\n+\t\tif (mkdir_in_gitdir(dir.buf) && errno != EEXIST) {\n+\t\t\tret = error_errno(_(\"unable to create directory %s\"), dir.buf);\n+\t\t\tstrbuf_release(&dir);\n+\t\t\treturn ret;\n+\t\t}\n+\t\tstrbuf_release(&dir);\n+\t}\n+\n+\treturn finalize_object_file_with_mtime(tmp_file.buf, filename.buf, mtime, flags);\n+}\n+\n int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t    const char *type, struct object_id *oid,\n \t\t\t    unsigned flags)\ndiff --git a/object-store.h b/object-store.h\nindex 952efb6a4b..061b0cb2ba 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -34,6 +34,11 @@ struct object_directory {\n \tchar *path;\n };\n \n+struct input_stream {\n+\tconst void *(*read)(struct input_stream *, unsigned long *len);\n+\tvoid *data;\n+};\n+\n KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n \tstruct object_directory *, 1, fspathhash, fspatheq)\n \n@@ -232,6 +237,10 @@ static inline int write_object_file(const void *buf, unsigned long len,\n \treturn write_object_file_flags(buf, len, type, oid, 0);\n }\n \n+int write_stream_object_file(struct input_stream *in_stream, size_t len,\n+\t\t\t     enum object_type type, time_t mtime,\n+\t\t\t     unsigned flags, struct object_id *oid);\n+\n int hash_object_file_literally(const void *buf, unsigned long len,\n \t\t\t       const char *type, struct object_id *oid,\n \t\t\t       unsigned flags);\n-- \n2.34.1.52.g80008efde6.agit.6.5.6\n\n"},{"id":"444603","messageId":"20211221115201.12120-6-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211217112629.12334-1-chiyutianyi@gmail.com","subject":"[PATCH v7 5/5] unpack-objects: unpack_non_delta_entry() read data in a stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2021-12-21T11:52:01Z","receivedAt":"2021-12-21T11:54:34Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWe used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\nentire contents of a blob object, no matter how big it is. This\nimplementation may consume all the memory and cause OOM.\n\nBy implementing a zstream version of input_stream interface, we can use\na small fixed buffer for \"unpack_non_delta_entry()\".\n\nHowever, unpack non-delta objects from a stream instead of from an\nentrie buffer will have 10% performance penalty. Therefore, only unpack\nobject larger than the \"core.BigFileStreamingThreshold\" in zstream. See\nthe following benchmarks:\n\n    hyperfine \\\n      --setup \\\n      'if ! test -d scalar.git; then git clone --bare https://github.com/microsoft/scalar.git; cp scalar.git/objects/pack/*.pack small.pack; fi' \\\n      --prepare 'rm -rf dest.git && git init --bare dest.git'\n\n    Summary\n      './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'origin/master'\n        1.01 ± 0.04 times faster than './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'HEAD~1'\n        1.01 ± 0.04 times faster than './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'HEAD~0'\n        1.03 ± 0.10 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'origin/master'\n        1.02 ± 0.07 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'HEAD~0'\n        1.10 ± 0.04 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'HEAD~1'\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Derrick Stolee <stolee@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n Documentation/config/core.txt       | 11 +++++\n builtin/unpack-objects.c            | 73 ++++++++++++++++++++++++++++-\n cache.h                             |  1 +\n config.c                            |  5 ++\n environment.c                       |  1 +\n t/t5590-unpack-non-delta-objects.sh | 36 +++++++++++++-\n 6 files changed, 125 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex c04f62a54a..601b7a2418 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -424,6 +424,17 @@ be delta compressed, but larger binary media files won't be.\n +\n Common unit suffixes of 'k', 'm', or 'g' are supported.\n \n+core.bigFileStreamingThreshold::\n+\tFiles larger than this will be streamed out to a temporary\n+\tobject file while being hashed, which will when be renamed\n+\tin-place to a loose object, particularly if the\n+\t`core.bigFileThreshold' setting dictates that they're always\n+\twritten out as loose objects.\n++\n+Default is 128 MiB on all platforms.\n++\n+Common unit suffixes of 'k', 'm', or 'g' are supported.\n+\n core.excludesFile::\n \tSpecifies the pathname to the file that contains patterns to\n \tdescribe paths that are not meant to be tracked, in addition\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 9104eb48da..72d8616e00 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -331,11 +331,82 @@ static void added_object(unsigned nr, enum object_type type,\n \t}\n }\n \n+struct input_zstream_data {\n+\tgit_zstream *zstream;\n+\tunsigned char buf[8192];\n+\tint status;\n+};\n+\n+static const void *feed_input_zstream(struct input_stream *in_stream,\n+\t\t\t\t      unsigned long *readlen)\n+{\n+\tstruct input_zstream_data *data = in_stream->data;\n+\tgit_zstream *zstream = data->zstream;\n+\tvoid *in = fill(1);\n+\n+\tif (!len || data->status == Z_STREAM_END) {\n+\t\t*readlen = 0;\n+\t\treturn NULL;\n+\t}\n+\n+\tzstream->next_out = data->buf;\n+\tzstream->avail_out = sizeof(data->buf);\n+\tzstream->next_in = in;\n+\tzstream->avail_in = len;\n+\n+\tdata->status = git_inflate(zstream, 0);\n+\tuse(len - zstream->avail_in);\n+\t*readlen = sizeof(data->buf) - zstream->avail_out;\n+\n+\treturn data->buf;\n+}\n+\n+static void write_stream_blob(unsigned nr, size_t size)\n+{\n+\tgit_zstream zstream;\n+\tstruct input_zstream_data data;\n+\tstruct input_stream in_stream = {\n+\t\t.read = feed_input_zstream,\n+\t\t.data = &data,\n+\t};\n+\n+\tmemset(&zstream, 0, sizeof(zstream));\n+\tmemset(&data, 0, sizeof(data));\n+\tdata.zstream = &zstream;\n+\tgit_inflate_init(&zstream);\n+\n+\tif (write_stream_object_file(&in_stream, size, OBJ_BLOB, 0, 0,\n+\t\t\t\t     &obj_list[nr].oid))\n+\t\tdie(_(\"failed to write object in stream\"));\n+\n+\tif (zstream.total_out != size || data.status != Z_STREAM_END)\n+\t\tdie(_(\"inflate returned %d\"), data.status);\n+\tgit_inflate_end(&zstream);\n+\n+\tif (strict) {\n+\t\tstruct blob *blob =\n+\t\t\tlookup_blob(the_repository, &obj_list[nr].oid);\n+\t\tif (blob)\n+\t\t\tblob->object.flags |= FLAG_WRITTEN;\n+\t\telse\n+\t\t\tdie(_(\"invalid blob object from stream\"));\n+\t}\n+\tobj_list[nr].obj = NULL;\n+}\n+\n static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n \t\t\t\t   unsigned nr)\n {\n-\tvoid *buf = get_data(size, dry_run);\n+\tvoid *buf;\n+\n+\t/* Write large blob in stream without allocating full buffer. */\n+\tif (!dry_run && type == OBJ_BLOB &&\n+\t    size > big_file_streaming_threshold) {\n+\t\twrite_stream_blob(nr, size);\n+\t\treturn;\n+\t}\n \n+\tbuf = get_data(size, dry_run);\n \tif (!dry_run && buf)\n \t\twrite_object(nr, type, buf, size);\n \telse\ndiff --git a/cache.h b/cache.h\nindex 64071a8d80..8c9123cb5d 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -974,6 +974,7 @@ extern size_t packed_git_window_size;\n extern size_t packed_git_limit;\n extern size_t delta_base_cache_limit;\n extern unsigned long big_file_threshold;\n+extern unsigned long big_file_streaming_threshold;\n extern unsigned long pack_size_limit_cfg;\n \n /*\ndiff --git a/config.c b/config.c\nindex c5873f3a70..7b122a142a 100644\n--- a/config.c\n+++ b/config.c\n@@ -1408,6 +1408,11 @@ static int git_default_core_config(const char *var, const char *value, void *cb)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"core.bigfilestreamingthreshold\")) {\n+\t\tbig_file_streaming_threshold = git_config_ulong(var, value);\n+\t\treturn 0;\n+\t}\n+\n \tif (!strcmp(var, \"core.packedgitlimit\")) {\n \t\tpacked_git_limit = git_config_ulong(var, value);\n \t\treturn 0;\ndiff --git a/environment.c b/environment.c\nindex 0d06a31024..04bba593de 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -47,6 +47,7 @@ size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n size_t delta_base_cache_limit = 96 * 1024 * 1024;\n unsigned long big_file_threshold = 512 * 1024 * 1024;\n+unsigned long big_file_streaming_threshold = 128 * 1024 * 1024;\n int pager_use_color = 1;\n const char *editor_program;\n const char *askpass_program;\ndiff --git a/t/t5590-unpack-non-delta-objects.sh b/t/t5590-unpack-non-delta-objects.sh\nindex 48c4fb1ba3..8436cbf8db 100755\n--- a/t/t5590-unpack-non-delta-objects.sh\n+++ b/t/t5590-unpack-non-delta-objects.sh\n@@ -13,6 +13,11 @@ export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n prepare_dest () {\n \ttest_when_finished \"rm -rf dest.git\" &&\n \tgit init --bare dest.git\n+\tif test -n \"$1\"\n+\tthen\n+\t\tgit -C dest.git config core.bigFileStreamingThreshold $1\n+\t\tgit -C dest.git config core.bigFileThreshold $1\n+\tfi\n }\n \n test_expect_success \"setup repo with big blobs (1.5 MB)\" '\n@@ -33,7 +38,7 @@ test_expect_success 'setup env: GIT_ALLOC_LIMIT to 1MB' '\n '\n \n test_expect_success 'fail to unpack-objects: cannot allocate' '\n-\tprepare_dest &&\n+\tprepare_dest 2m &&\n \ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n \tgrep \"fatal: attempting to allocate\" err &&\n \t(\n@@ -44,6 +49,35 @@ test_expect_success 'fail to unpack-objects: cannot allocate' '\n \t! test_cmp expect actual\n '\n \n+test_expect_success 'unpack big object in stream' '\n+\tprepare_dest 1m &&\n+\tmkdir -p dest.git/objects/05 &&\n+\tgit -C dest.git unpack-objects <test-$PACK.pack &&\n+\tgit -C dest.git fsck &&\n+\t(\n+\t\tcd dest.git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'unpack big object in stream with existing oids' '\n+\tprepare_dest 1m &&\n+\tgit -C dest.git index-pack --stdin <test-$PACK.pack &&\n+\t(\n+\t\tcd dest.git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >actual &&\n+\ttest_must_be_empty actual &&\n+\tgit -C dest.git unpack-objects <test-$PACK.pack &&\n+\tgit -C dest.git fsck &&\n+\t(\n+\t\tcd dest.git &&\n+\t\tfind objects/?? -type f | sort\n+\t) >actual &&\n+\ttest_must_be_empty actual\n+'\n+\n test_expect_success 'unpack-objects dry-run' '\n \tprepare_dest &&\n \tgit -C dest.git unpack-objects -n <test-$PACK.pack &&\n-- \n2.34.1.52.g80008efde6.agit.6.5.6\n\n"},{"id":"444611","messageId":"211221.86bl1arqls.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"20211221115201.12120-2-chiyutianyi@gmail.com","subject":"Re: [PATCH v7 1/5] unpack-objects.c: add dry_run mode for get_data()","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-21T14:09:43Z","receivedAt":"2021-12-21T14:15:50Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Dec 21 2021, Han Xin wrote:\n\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> In dry_run mode, \"get_data()\" is used to verify the inflation of data,\n> and the returned buffer will not be used at all and will be freed\n> immediately. Even in dry_run mode, it is dangerous to allocate a\n> full-size buffer for a large blob object. Therefore, only allocate a\n> low memory footprint when calling \"get_data()\" in dry_run mode.\n>\n> Suggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  builtin/unpack-objects.c            | 23 +++++++++---\n>  t/t5590-unpack-non-delta-objects.sh | 57 +++++++++++++++++++++++++++++\n>  2 files changed, 74 insertions(+), 6 deletions(-)\n>  create mode 100755 t/t5590-unpack-non-delta-objects.sh\n>\n> diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n> index 4a9466295b..9104eb48da 100644\n> --- a/builtin/unpack-objects.c\n> +++ b/builtin/unpack-objects.c\n> @@ -96,15 +96,21 @@ static void use(int bytes)\n>  \tdisplay_throughput(progress, consumed_bytes);\n>  }\n>  \n> -static void *get_data(unsigned long size)\n> +static void *get_data(size_t size, int dry_run)\n>  {\n>  \tgit_zstream stream;\n> -\tvoid *buf = xmallocz(size);\n> +\tsize_t bufsize;\n> +\tvoid *buf;\n>  \n>  \tmemset(&stream, 0, sizeof(stream));\n> +\tif (dry_run && size > 8192)\n> +\t\tbufsize = 8192;\n> +\telse\n> +\t\tbufsize = size;\n> +\tbuf = xmallocz(bufsize);\n\nMaybe I'm misunderstanding this, but the commit message says it would be\ndangerous to allocate a very larger buffer, but here we only limit the\nsize under \"dry_run\".\n\nRemoving that \"&& size > 8192\" makes all the tests pass still, so there\nseems to be some missing coverage here in any case.\n\n> diff --git a/t/t5590-unpack-non-delta-objects.sh b/t/t5590-unpack-non-delta-objects.sh\n> new file mode 100755\n> index 0000000000..48c4fb1ba3\n> --- /dev/null\n> +++ b/t/t5590-unpack-non-delta-objects.sh\n> @@ -0,0 +1,57 @@\n> +#!/bin/sh\n> +#\n> +# Copyright (c) 2021 Han Xin\n> +#\n> +\n> +test_description='Test unpack-objects with non-delta objects'\n> +\n> +GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n> +export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n> +\n> +. ./test-lib.sh\n> +\n> +prepare_dest () {\n> +\ttest_when_finished \"rm -rf dest.git\" &&\n> +\tgit init --bare dest.git\n> +}\n> +\n> +test_expect_success \"setup repo with big blobs (1.5 MB)\" '\n> +\ttest-tool genrandom foo 1500000 >big-blob &&\n> +\ttest_commit --append foo big-blob &&\n> +\ttest-tool genrandom bar 1500000 >big-blob &&\n> +\ttest_commit --append bar big-blob &&\n> +\t(\n> +\t\tcd .git &&\n> +\t\tfind objects/?? -type f | sort\n> +\t) >expect &&\n> +\tPACK=$(echo main | git pack-objects --revs test)\n> +'\n> +\n> +test_expect_success 'setup env: GIT_ALLOC_LIMIT to 1MB' '\n> +\tGIT_ALLOC_LIMIT=1m &&\n> +\texport GIT_ALLOC_LIMIT\n> +'\n> +\n> +test_expect_success 'fail to unpack-objects: cannot allocate' '\n> +\tprepare_dest &&\n> +\ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n> +\tgrep \"fatal: attempting to allocate\" err &&\n> +\t(\n> +\t\tcd dest.git &&\n> +\t\tfind objects/?? -type f | sort\n> +\t) >actual &&\n> +\ttest_file_not_empty actual &&\n> +\t! test_cmp expect actual\n> +'\n> +\n> +test_expect_success 'unpack-objects dry-run' '\n> +\tprepare_dest &&\n> +\tgit -C dest.git unpack-objects -n <test-$PACK.pack &&\n> +\t(\n> +\t\tcd dest.git &&\n> +\t\tfind objects/ -type f\n> +\t) >actual &&\n> +\ttest_must_be_empty actual\n> +'\n> +\n> +test_done\n\nI commented on this \"find\" usage in an earlier round, I think there's a\nmuch easier way to do this. You're really just going back and forth\nbetween checking whether or not all the objects are loose.\n\nI think that the below fix-up on top of this series is a better way to\ndo that, and more accurate. I.e. in your test here you check \"!\ntest_cmp\", which means that we could have some packed and some loose,\nbut really what you're meaning to check is a flip-flop between \"all\nloose?\" and \"no loose?.\n\nIn addition to that there was no reason to hardcode \"main\", we can just\nuse HEAD. All in all I think the below fix-up makes sense:\n\ndiff --git a/t/t5590-unpack-non-delta-objects.sh b/t/t5590-unpack-non-delta-objects.sh\nindex 8436cbf8db6..d78bb89225d 100755\n--- a/t/t5590-unpack-non-delta-objects.sh\n+++ b/t/t5590-unpack-non-delta-objects.sh\n@@ -5,9 +5,6 @@\n \n test_description='Test unpack-objects with non-delta objects'\n \n-GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n-export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n-\n . ./test-lib.sh\n \n prepare_dest () {\n@@ -20,16 +17,22 @@ prepare_dest () {\n \tfi\n }\n \n+assert_no_loose () {\n+\tglob=dest.git/objects/?? &&\n+\techo \"$glob\" >expect &&\n+\techo $glob >actual &&\n+\ttest_cmp expect actual\n+}\n+\n test_expect_success \"setup repo with big blobs (1.5 MB)\" '\n \ttest-tool genrandom foo 1500000 >big-blob &&\n \ttest_commit --append foo big-blob &&\n \ttest-tool genrandom bar 1500000 >big-blob &&\n \ttest_commit --append bar big-blob &&\n-\t(\n-\t\tcd .git &&\n-\t\tfind objects/?? -type f | sort\n-\t) >expect &&\n-\tPACK=$(echo main | git pack-objects --revs test)\n+\n+\t# Everything is loose\n+\trmdir .git/objects/pack &&\n+\tPACK=$(echo HEAD | git pack-objects --revs test)\n '\n \n test_expect_success 'setup env: GIT_ALLOC_LIMIT to 1MB' '\n@@ -41,51 +44,27 @@ test_expect_success 'fail to unpack-objects: cannot allocate' '\n \tprepare_dest 2m &&\n \ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n \tgrep \"fatal: attempting to allocate\" err &&\n-\t(\n-\t\tcd dest.git &&\n-\t\tfind objects/?? -type f | sort\n-\t) >actual &&\n-\ttest_file_not_empty actual &&\n-\t! test_cmp expect actual\n+\trmdir dest.git/objects/pack\n '\n \n test_expect_success 'unpack big object in stream' '\n \tprepare_dest 1m &&\n \tmkdir -p dest.git/objects/05 &&\n \tgit -C dest.git unpack-objects <test-$PACK.pack &&\n-\tgit -C dest.git fsck &&\n-\t(\n-\t\tcd dest.git &&\n-\t\tfind objects/?? -type f | sort\n-\t) >actual &&\n-\ttest_cmp expect actual\n+\trmdir dest.git/objects/pack\n '\n \n test_expect_success 'unpack big object in stream with existing oids' '\n \tprepare_dest 1m &&\n \tgit -C dest.git index-pack --stdin <test-$PACK.pack &&\n-\t(\n-\t\tcd dest.git &&\n-\t\tfind objects/?? -type f | sort\n-\t) >actual &&\n-\ttest_must_be_empty actual &&\n \tgit -C dest.git unpack-objects <test-$PACK.pack &&\n-\tgit -C dest.git fsck &&\n-\t(\n-\t\tcd dest.git &&\n-\t\tfind objects/?? -type f | sort\n-\t) >actual &&\n-\ttest_must_be_empty actual\n+\tassert_no_loose\n '\n \n test_expect_success 'unpack-objects dry-run' '\n \tprepare_dest &&\n \tgit -C dest.git unpack-objects -n <test-$PACK.pack &&\n-\t(\n-\t\tcd dest.git &&\n-\t\tfind objects/ -type f\n-\t) >actual &&\n-\ttest_must_be_empty actual\n+\tassert_no_loose\n '\n \n test_done\n"},{"id":"444612","messageId":"211221.867dbyrqe9.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"20211221115201.12120-4-chiyutianyi@gmail.com","subject":"Re: [PATCH v7 3/5] object-file.c: refactor write_loose_object() to reuse in stream version","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-21T14:16:57Z","receivedAt":"2021-12-21T14:20:18Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Dec 21 2021, Han Xin wrote:\n\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n> [...]\n> @@ -1854,17 +1876,48 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n>  \t\tstrbuf_reset(tmp);\n>  \t\tstrbuf_add(tmp, filename, dirlen - 1);\n>  \t\tif (mkdir(tmp->buf, 0777) && errno != EEXIST)\n> -\t\t\treturn -1;\n> +\t\t\tbreak;\n>  \t\tif (adjust_shared_perm(tmp->buf))\n> -\t\t\treturn -1;\n> +\t\t\tbreak;\n>  \n>  \t\t/* Try again */\n>  \t\tstrbuf_addstr(tmp, \"/tmp_obj_XXXXXX\");\n>  \t\tfd = git_mkstemp_mode(tmp->buf, 0444);\n> +\t} while (0);\n> +\n> +\tif (fd < 0 && !(flags & HASH_SILENT)) {\n> +\t\tif (errno == EACCES)\n> +\t\t\treturn error(_(\"insufficient permission for adding an \"\n> +\t\t\t\t       \"object to repository database %s\"),\n> +\t\t\t\t     get_object_directory());\n\nThis should be an error_errno() instead, ...\n\n> +\t\telse\n> +\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n\n...and we can just fold this whole if/else into one condition with a\nbriefer message, e.g.:\n\n    error_errno(_(\"unable to add object to '%s'\"), get_object_directory());\n\nOr whatever, unless there's another bug here where you inverted these\nconditions, and the \"else\" really should not use \"error_errno\" but\n\"error\".... (I don't know...)\n"},{"id":"444615","messageId":"b2dee243-1a38-531e-02b1-ffd66c465fa5@web.de","threadId":"56672","inReplyTo":"20211221115201.12120-3-chiyutianyi@gmail.com","subject":"Re: [PATCH v7 2/5] object-file API: add a format_object_header() function","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2021-12-21T14:30:27Z","receivedAt":"2021-12-21T14:31:07Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 21.12.21 um 12:51 schrieb Han Xin:\n> From: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n>\n> Add a convenience function to wrap the xsnprintf() command that\n> generates loose object headers. This code was copy/pasted in various\n> parts of the codebase, let's define it in one place and re-use it from\n> there.\n>\n> All except one caller of it had a valid \"enum object_type\" for us,\n> it's only write_object_file_prepare() which might need to deal with\n> \"git hash-object --literally\" and a potential garbage type. Let's have\n> the primary API use an \"enum object_type\", and define an *_extended()\n> function that can take an arbitrary \"const char *\" for the type.\n>\n> See [1] for the discussion that prompted this patch, i.e. new code in\n> object-file.c that wanted to copy/paste the xsnprintf() invocation.\n>\n> 1. https://lore.kernel.org/git/211213.86bl1l9bfz.gmgdl@evledraar.gmail.com/\n>\n> Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  builtin/index-pack.c |  3 +--\n>  bulk-checkin.c       |  4 ++--\n>  cache.h              | 21 +++++++++++++++++++++\n>  http-push.c          |  2 +-\n>  object-file.c        | 14 +++++++++++---\n>  5 files changed, 36 insertions(+), 8 deletions(-)\n>\n> diff --git a/builtin/index-pack.c b/builtin/index-pack.c\n> index c23d01de7d..4a765ddae6 100644\n> --- a/builtin/index-pack.c\n> +++ b/builtin/index-pack.c\n> @@ -449,8 +449,7 @@ static void *unpack_entry_data(off_t offset, unsigned long size,\n>  \tint hdrlen;\n>\n>  \tif (!is_delta_type(type)) {\n> -\t\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX,\n> -\t\t\t\t   type_name(type),(uintmax_t)size) + 1;\n> +\t\thdrlen = format_object_header(hdr, sizeof(hdr), type, (uintmax_t)size);\n                                                                      ^^^^^^^^^^^\nThis explicit cast is unnecessary.  It was needed with xsnprintf(), but\nthat implementation detail is handled inside the new helper function.\n\n(format_object_header() takes a size_t; even if unsigned long would be\nwider than that on some weird architecture, casting the size to\nuintmax_t will not avoid the implicit truncation happening during the\nfunction call.)\n\n>  \t\tthe_hash_algo->init_fn(&c);\n>  \t\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n>  \t} else\n> diff --git a/bulk-checkin.c b/bulk-checkin.c\n> index 8785b2ac80..1733a1de4f 100644\n> --- a/bulk-checkin.c\n> +++ b/bulk-checkin.c\n> @@ -220,8 +220,8 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n>  \tif (seekback == (off_t) -1)\n>  \t\treturn error(\"cannot find the current offset\");\n>\n> -\theader_len = xsnprintf((char *)obuf, sizeof(obuf), \"%s %\" PRIuMAX,\n> -\t\t\t       type_name(type), (uintmax_t)size) + 1;\n> +\theader_len = format_object_header((char *)obuf, sizeof(obuf),\n> +\t\t\t\t\t type, (uintmax_t)size);\n                                               ^^^^^^^^^^^\nSame here, just that size is already of type size_t, so a cast makes\neven less sense.\n\n>  \tthe_hash_algo->init_fn(&ctx);\n>  \tthe_hash_algo->update_fn(&ctx, obuf, header_len);\n>\n> diff --git a/cache.h b/cache.h\n> index cfba463aa9..64071a8d80 100644\n> --- a/cache.h\n> +++ b/cache.h\n> @@ -1310,6 +1310,27 @@ enum unpack_loose_header_result unpack_loose_header(git_zstream *stream,\n>  \t\t\t\t\t\t    unsigned long bufsiz,\n>  \t\t\t\t\t\t    struct strbuf *hdrbuf);\n>\n> +/**\n> + * format_object_header() is a thin wrapper around s xsnprintf() that\n> + * writes the initial \"<type> <obj-len>\" part of the loose object\n> + * header. It returns the size that snprintf() returns + 1.\n> + *\n> + * The format_object_header_extended() function allows for writing a\n> + * type_name that's not one of the \"enum object_type\" types. This is\n> + * used for \"git hash-object --literally\". Pass in a OBJ_NONE as the\n> + * type, and a non-NULL \"type_str\" to do that.\n> + *\n> + * format_object_header() is a convenience wrapper for\n> + * format_object_header_extended().\n> + */\n> +int format_object_header_extended(char *str, size_t size, enum object_type type,\n> +\t\t\t\t const char *type_str, size_t objsize);\n> +static inline int format_object_header(char *str, size_t size,\n> +\t\t\t\t      enum object_type type, size_t objsize)\n> +{\n> +\treturn format_object_header_extended(str, size, type, NULL, objsize);\n> +}\n> +\n>  /**\n>   * parse_loose_header() parses the starting \"<type> <len>\\0\" of an\n>   * object. If it doesn't follow that format -1 is returned. To check\n> diff --git a/http-push.c b/http-push.c\n> index 3309aaf004..f55e316ff4 100644\n> --- a/http-push.c\n> +++ b/http-push.c\n> @@ -363,7 +363,7 @@ static void start_put(struct transfer_request *request)\n>  \tgit_zstream stream;\n>\n>  \tunpacked = read_object_file(&request->obj->oid, &type, &len);\n> -\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n> +\thdrlen = format_object_header(hdr, sizeof(hdr), type, (uintmax_t)len);\n                                                              ^^^^^^^^^^^\nSame here; len is of type unsigned long.\n\n>\n>  \t/* Set it up */\n>  \tgit_deflate_init(&stream, zlib_compression_level);\n> diff --git a/object-file.c b/object-file.c\n> index eb1426f98c..6bba4766f9 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1006,6 +1006,14 @@ void *xmmap(void *start, size_t length,\n>  \treturn ret;\n>  }\n>\n> +int format_object_header_extended(char *str, size_t size, enum object_type type,\n> +\t\t\t\t const char *typestr, size_t objsize)\n> +{\n> +\tconst char *s = type == OBJ_NONE ? typestr : type_name(type);\n> +\n> +\treturn xsnprintf(str, size, \"%s %\"PRIuMAX, s, (uintmax_t)objsize) + 1;\n                                                      ^^^^^^^^^^^\nThis cast is necessary to match PRIuMAX.  And that is used because the z\nmodifier (as in e.g. printf(\"%zu\", sizeof(size_t));) was only added in\nC99 and not all platforms may have it.  (Perhaps this cautious approach\nis worth revisiting separately, now that some time has passed, but this\npatch series should still use PRIuMAX, as it does.)\n\n> +}\n> +\n>  /*\n>   * With an in-core object data in \"map\", rehash it to make sure the\n>   * object name actually matches \"oid\" to detect object corruption.\n> @@ -1034,7 +1042,7 @@ int check_object_signature(struct repository *r, const struct object_id *oid,\n>  \t\treturn -1;\n>\n>  \t/* Generate the header */\n> -\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(obj_type), (uintmax_t)size) + 1;\n> +\thdrlen = format_object_header(hdr, sizeof(hdr), obj_type, size);\n>\n>  \t/* Sha1.. */\n>  \tr->hash_algo->init_fn(&c);\n> @@ -1734,7 +1742,7 @@ static void write_object_file_prepare(const struct git_hash_algo *algo,\n>  \tgit_hash_ctx c;\n>\n>  \t/* Generate the header */\n> -\t*hdrlen = xsnprintf(hdr, *hdrlen, \"%s %\"PRIuMAX , type, (uintmax_t)len)+1;\n> +\t*hdrlen = format_object_header_extended(hdr, *hdrlen, OBJ_NONE, type, len);\n>\n>  \t/* Sha1.. */\n>  \talgo->init_fn(&c);\n> @@ -2006,7 +2014,7 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n>  \tbuf = read_object(the_repository, oid, &type, &len);\n>  \tif (!buf)\n>  \t\treturn error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n> -\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n> +\thdrlen = format_object_header(hdr, sizeof(hdr), type, len);\n>  \tret = write_loose_object(oid, hdr, hdrlen, buf, len, mtime, 0);\n>  \tfree(buf);\n>\n\nNo explicit cast in these three cases -- good.  They all pass an\nunsigned long as last parameter btw.\n\nRené\n"},{"id":"444616","messageId":"211221.8635mmrpps.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"20211221115201.12120-5-chiyutianyi@gmail.com","subject":"Re: [PATCH v7 4/5] object-file.c: add \"write_stream_object_file()\" to support read in stream","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-21T14:20:22Z","receivedAt":"2021-12-21T14:34:59Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Dec 21 2021, Han Xin wrote:\n\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n> [...]\n> +int write_stream_object_file(struct input_stream *in_stream, size_t len,\n> +\t\t\t     enum object_type type, time_t mtime,\n> +\t\t\t     unsigned flags, struct object_id *oid)\n> +{\n> +\tint fd, ret, flush = 0;\n> +\tunsigned char compressed[4096];\n> +\tgit_zstream stream;\n> +\tgit_hash_ctx c;\n> +\tstruct object_id parano_oid;\n> +\tstatic struct strbuf tmp_file = STRBUF_INIT;\n> +\tstatic struct strbuf filename = STRBUF_INIT;\n> +\tint dirlen;\n> +\tchar hdr[MAX_HEADER_LEN];\n> +\tint hdrlen = sizeof(hdr);\n> +\n> +\t/* Since \"filename\" is defined as static, it will be reused. So reset it\n> +\t * first before using it. */\n> +\tstrbuf_reset(&filename);\n> +\t/* When oid is not determined, save tmp file to odb path. */\n> +\tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n\nI realize this is somewhat following the pattern of code you moved\naround earlier, but FWIW I think these sorts of comments are really\nover-doing it. I.e. we try not to comment on things that are obvious\nfrom the code itself.\n\nAlso René's comment on v6 still applies here:\n\n    Given that this function is only used for huge objects I think making\n    the strbufs non-static and releasing them is the best choice here.\n\nI thin just making them non-static and doing a strbuf_release() as he\nsuggested is best here.\n\n> +\n> +\tfd = create_tmpfile(&tmp_file, filename.buf, flags);\n> +\tif (fd < 0)\n> +\t\treturn -1;\n> +\n> +\thdrlen = format_object_header(hdr, hdrlen, type, len);\n> +\n> +\t/* Set it up and write header */\n> +\tsetup_stream_and_header(&stream, compressed, sizeof(compressed),\n> +\t\t\t\t&c, hdr, hdrlen);\n> +\n> +\t/* Then the data itself.. */\n> +\tdo {\n> +\t\tunsigned char *in0 = stream.next_in;\n> +\t\tif (!stream.avail_in) {\n> +\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n> +\t\t\tstream.next_in = (void *)in;\n> +\t\t\tin0 = (unsigned char *)in;\n> +\t\t\t/* All data has been read. */\n> +\t\t\tif (len + hdrlen == stream.total_in + stream.avail_in)\n> +\t\t\t\tflush = Z_FINISH;\n> +\t\t}\n> +\t\tret = git_deflate(&stream, flush);\n> +\t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n> +\t\tif (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n> +\t\t\tdie(_(\"unable to write loose object file\"));\n> +\t\tstream.next_out = compressed;\n> +\t\tstream.avail_out = sizeof(compressed);\n> +\t} while (ret == Z_OK || ret == Z_BUF_ERROR);\n> +\n> +\tif (ret != Z_STREAM_END)\n> +\t\tdie(_(\"unable to deflate new object streamingly (%d)\"), ret);\n> +\tret = git_deflate_end_gently(&stream);\n> +\tif (ret != Z_OK)\n> +\t\tdie(_(\"deflateEnd on object streamingly failed (%d)\"), ret);\n\nnit: let's say \"unable to stream deflate new object\" or something, and\nnot use the confusing (invented?) word \"streamingly\".\n\n> +\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n> +\n> +\tclose_loose_object(fd);\n> +\n> +\toidcpy(oid, &parano_oid);\n\nI see there's still quite a bit of duplication between this and\nwrite_loose_object(), but maybe it's not easy to factor out.\n\n> +\tif (freshen_packed_object(oid) || freshen_loose_object(oid)) {\n> +\t\tunlink_or_warn(tmp_file.buf);\n> +\t\treturn 0;\n> +\t}\n> +\n> +\tloose_object_path(the_repository, &filename, oid);\n> +\n> +\t/* We finally know the object path, and create the missing dir. */\n> +\tdirlen = directory_size(filename.buf);\n> +\tif (dirlen) {\n> +\t\tstruct strbuf dir = STRBUF_INIT;\n> +\t\tstrbuf_add(&dir, filename.buf, dirlen - 1);\n\nJust a minor nit, but I noticed we could have this on top, i.e. this\n\"remove the slash\" is now what 1/3 users of it wan:\n\t\n\t object-file.c | 10 +++++-----\n\t 1 file changed, 5 insertions(+), 5 deletions(-)\n\t\n\tdiff --git a/object-file.c b/object-file.c\n\tindex 77a3217fd0e..b0dea96906e 100644\n\t--- a/object-file.c\n\t+++ b/object-file.c\n\t@@ -1878,13 +1878,13 @@ static void close_loose_object(int fd)\n\t \t\tdie_errno(_(\"error when closing loose object file\"));\n\t }\n\t \n\t-/* Size of directory component, including the ending '/' */\n\t+/* Size of directory component, excluding the ending '/' */\n\t static inline int directory_size(const char *filename)\n\t {\n\t \tconst char *s = strrchr(filename, '/');\n\t \tif (!s)\n\t \t\treturn 0;\n\t-\treturn s - filename + 1;\n\t+\treturn s - filename;\n\t }\n\t \n\t /*\n\t@@ -1901,7 +1901,7 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename,\n\t \n\t \tstrbuf_reset(tmp);\n\t \tstrbuf_add(tmp, filename, dirlen);\n\t-\tstrbuf_addstr(tmp, \"tmp_obj_XXXXXX\");\n\t+\tstrbuf_addstr(tmp, \"/tmp_obj_XXXXXX\");\n\t \tfd = git_mkstemp_mode(tmp->buf, 0444);\n\t \tdo {\n\t \t\tif (fd >= 0 || !dirlen || errno != ENOENT)\n\t@@ -1913,7 +1913,7 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename,\n\t \t\t * scratch.\n\t \t\t */\n\t \t\tstrbuf_reset(tmp);\n\t-\t\tstrbuf_add(tmp, filename, dirlen - 1);\n\t+\t\tstrbuf_add(tmp, filename, dirlen);\n\t \t\tif (mkdir(tmp->buf, 0777) && errno != EEXIST)\n\t \t\t\tbreak;\n\t \t\tif (adjust_shared_perm(tmp->buf))\n\t@@ -2100,7 +2100,7 @@ int write_stream_object_file(struct input_stream *in_stream, size_t len,\n\t \tdirlen = directory_size(filename.buf);\n\t \tif (dirlen) {\n\t \t\tstruct strbuf dir = STRBUF_INIT;\n\t-\t\tstrbuf_add(&dir, filename.buf, dirlen - 1);\n\t+\t\tstrbuf_add(&dir, filename.buf, dirlen);\n\t \n\t \t\tif (mkdir_in_gitdir(dir.buf) && errno != EEXIST) {\n\t \t\t\tret = error_errno(_(\"unable to create directory %s\"), dir.buf);\n\nOn my platform (linux) it's not needed either way, a \"mkdir foo\" works\nas well as \"mkdir foo/\", but maybe some oS's have trouble with it.\n"},{"id":"444617","messageId":"3f1e3a59-1288-ecff-88b5-53bbd510a6a0@web.de","threadId":"56672","inReplyTo":"211221.86bl1arqls.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v7 1/5] unpack-objects.c: add dry_run mode for get_data()","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2021-12-21T14:43:20Z","receivedAt":"2021-12-21T14:43:47Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 21.12.21 um 15:09 schrieb Ævar Arnfjörð Bjarmason:\n>\n> On Tue, Dec 21 2021, Han Xin wrote:\n>\n>> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>>\n>> In dry_run mode, \"get_data()\" is used to verify the inflation of data,\n>> and the returned buffer will not be used at all and will be freed\n>> immediately. Even in dry_run mode, it is dangerous to allocate a\n>> full-size buffer for a large blob object. Therefore, only allocate a\n>> low memory footprint when calling \"get_data()\" in dry_run mode.\n>>\n>> Suggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n>> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n>> ---\n>>  builtin/unpack-objects.c            | 23 +++++++++---\n>>  t/t5590-unpack-non-delta-objects.sh | 57 +++++++++++++++++++++++++++++\n>>  2 files changed, 74 insertions(+), 6 deletions(-)\n>>  create mode 100755 t/t5590-unpack-non-delta-objects.sh\n>>\n>> diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n>> index 4a9466295b..9104eb48da 100644\n>> --- a/builtin/unpack-objects.c\n>> +++ b/builtin/unpack-objects.c\n>> @@ -96,15 +96,21 @@ static void use(int bytes)\n>>  \tdisplay_throughput(progress, consumed_bytes);\n>>  }\n>>\n>> -static void *get_data(unsigned long size)\n>> +static void *get_data(size_t size, int dry_run)\n>>  {\n>>  \tgit_zstream stream;\n>> -\tvoid *buf = xmallocz(size);\n>> +\tsize_t bufsize;\n>> +\tvoid *buf;\n>>\n>>  \tmemset(&stream, 0, sizeof(stream));\n>> +\tif (dry_run && size > 8192)\n>> +\t\tbufsize = 8192;\n>> +\telse\n>> +\t\tbufsize = size;\n>> +\tbuf = xmallocz(bufsize);\n>\n> Maybe I'm misunderstanding this, but the commit message says it would be\n> dangerous to allocate a very larger buffer, but here we only limit the\n> size under \"dry_run\".\n\nThis patch reduces the memory usage of dry runs, as its commit message\nsays.  The memory usage of one type of actual (non-dry) unpack is reduced\nby patch 5.\n\n> Removing that \"&& size > 8192\" makes all the tests pass still, so there\n> seems to be some missing coverage here in any case.\n\nHow would you test that an 8KB buffer is allocated even though a smaller\none would suffice?  And why?  Wasting a few KB shouldn't be noticeable.\n\nRené\n"},{"id":"444621","messageId":"211221.86y24eq9rn.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"3f1e3a59-1288-ecff-88b5-53bbd510a6a0@web.de","subject":"Re: [PATCH v7 1/5] unpack-objects.c: add dry_run mode for get_data()","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-21T15:04:23Z","receivedAt":"2021-12-21T15:04:49Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Dec 21 2021, René Scharfe wrote:\n\n> Am 21.12.21 um 15:09 schrieb Ævar Arnfjörð Bjarmason:\n>>\n>> On Tue, Dec 21 2021, Han Xin wrote:\n>>\n>>> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>>>\n>>> In dry_run mode, \"get_data()\" is used to verify the inflation of data,\n>>> and the returned buffer will not be used at all and will be freed\n>>> immediately. Even in dry_run mode, it is dangerous to allocate a\n>>> full-size buffer for a large blob object. Therefore, only allocate a\n>>> low memory footprint when calling \"get_data()\" in dry_run mode.\n>>>\n>>> Suggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n>>> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n>>> ---\n>>>  builtin/unpack-objects.c            | 23 +++++++++---\n>>>  t/t5590-unpack-non-delta-objects.sh | 57 +++++++++++++++++++++++++++++\n>>>  2 files changed, 74 insertions(+), 6 deletions(-)\n>>>  create mode 100755 t/t5590-unpack-non-delta-objects.sh\n>>>\n>>> diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n>>> index 4a9466295b..9104eb48da 100644\n>>> --- a/builtin/unpack-objects.c\n>>> +++ b/builtin/unpack-objects.c\n>>> @@ -96,15 +96,21 @@ static void use(int bytes)\n>>>  \tdisplay_throughput(progress, consumed_bytes);\n>>>  }\n>>>\n>>> -static void *get_data(unsigned long size)\n>>> +static void *get_data(size_t size, int dry_run)\n>>>  {\n>>>  \tgit_zstream stream;\n>>> -\tvoid *buf = xmallocz(size);\n>>> +\tsize_t bufsize;\n>>> +\tvoid *buf;\n>>>\n>>>  \tmemset(&stream, 0, sizeof(stream));\n>>> +\tif (dry_run && size > 8192)\n>>> +\t\tbufsize = 8192;\n>>> +\telse\n>>> +\t\tbufsize = size;\n>>> +\tbuf = xmallocz(bufsize);\n>>\n>> Maybe I'm misunderstanding this, but the commit message says it would be\n>> dangerous to allocate a very larger buffer, but here we only limit the\n>> size under \"dry_run\".\n>\n> This patch reduces the memory usage of dry runs, as its commit message\n> says.  The memory usage of one type of actual (non-dry) unpack is reduced\n> by patch 5.\n>\n>> Removing that \"&& size > 8192\" makes all the tests pass still, so there\n>> seems to be some missing coverage here in any case.\n>\n> How would you test that an 8KB buffer is allocated even though a smaller\n> one would suffice?  And why?  Wasting a few KB shouldn't be noticeable.\n\nThat doesn't sound like it needs to be tested. I was just trying to grok\nwhat this was all doing. Thanks!\n"},{"id":"444622","messageId":"211221.86tuf2q9os.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"211221.8635mmrpps.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v7 4/5] object-file.c: add \"write_stream_object_file()\" to support read in stream","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-21T15:05:11Z","receivedAt":"2021-12-21T15:06:31Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Dec 21 2021, Ævar Arnfjörð Bjarmason wrote:\n\n> On Tue, Dec 21 2021, Han Xin wrote:\n\n>> +\t/* Then the data itself.. */\n>> +\tdo {\n>> +\t\tunsigned char *in0 = stream.next_in;\n>> +\t\tif (!stream.avail_in) {\n>> +\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n>> +\t\t\tstream.next_in = (void *)in;\n>> +\t\t\tin0 = (unsigned char *)in;\n>> +\t\t\t/* All data has been read. */\n>> +\t\t\tif (len + hdrlen == stream.total_in + stream.avail_in)\n>> +\t\t\t\tflush = Z_FINISH;\n>> +\t\t}\n>> +\t\tret = git_deflate(&stream, flush);\n>> +\t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n>> +\t\tif (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n>> +\t\t\tdie(_(\"unable to write loose object file\"));\n>> +\t\tstream.next_out = compressed;\n>> +\t\tstream.avail_out = sizeof(compressed);\n>> +\t} while (ret == Z_OK || ret == Z_BUF_ERROR);\n>> +\n>> +\tif (ret != Z_STREAM_END)\n>> +\t\tdie(_(\"unable to deflate new object streamingly (%d)\"), ret);\n>> +\tret = git_deflate_end_gently(&stream);\n>> +\tif (ret != Z_OK)\n>> +\t\tdie(_(\"deflateEnd on object streamingly failed (%d)\"), ret);\n>\n> nit: let's say \"unable to stream deflate new object\" or something, and\n> not use the confusing (invented?) word \"streamingly\".\n>\n>> +\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n>> +\n>> +\tclose_loose_object(fd);\n>> +\n>> +\toidcpy(oid, &parano_oid);\n>\n> I see there's still quite a bit of duplication between this and\n> write_loose_object(), but maybe it's not easy to factor out.\n\nFor what it's worth I tried to do that and the result doesn't really\nseem worth it. I.e. something like the below. The inner loop of the\ndo/while looks like it could get a similar treatment, but likewise\ndoesn't seem worth the effort.\n\ndiff --git a/object-file.c b/object-file.c\nindex b0dea96906e..7fc2363cfa1 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1957,6 +1957,46 @@ static void setup_stream_and_header(git_zstream *stream,\n \tthe_hash_algo->update_fn(c, hdr, hdrlen);\n }\n \n+static int start_loose_object_common(struct strbuf *tmp_file,\n+\t\t\t\t     const char *filename, unsigned flags,\n+\t\t\t\t     git_zstream *stream,\n+\t\t\t\t     unsigned char *buf, size_t buflen,\n+\t\t\t\t     git_hash_ctx *c,\n+\t\t\t\t     enum object_type type, size_t len,\n+\t\t\t\t     char *hdr, int *hdrlen)\n+{\n+\tint fd;\n+\n+\tfd = create_tmpfile(tmp_file, filename, flags);\n+\tif (fd < 0)\n+\t\treturn -1;\n+\n+\tif (type != OBJ_NONE)\n+\t\t*hdrlen = format_object_header(hdr, *hdrlen, type, len);\n+\n+\t/* Set it up and write header */\n+\tsetup_stream_and_header(stream, buf, buflen, c, hdr, *hdrlen);\n+\n+\treturn fd;\n+\n+}\n+\n+static void end_loose_object_common(int ret, git_hash_ctx *c,\n+\t\t\t\t    git_zstream *stream,\n+\t\t\t\t    struct object_id *parano_oid,\n+\t\t\t\t    const struct object_id *expected_oid,\n+\t\t\t\t    const char *zstream_end_fmt,\n+\t\t\t\t    const char *z_ok_fmt)\n+{\n+\tif (ret != Z_STREAM_END)\n+\t\tdie(_(zstream_end_fmt), ret, expected_oid);\n+\tret = git_deflate_end_gently(stream);\n+\tif (ret != Z_OK)\n+\t\tdie(_(z_ok_fmt), ret, expected_oid);\n+\tthe_hash_algo->final_oid_fn(parano_oid, c);\n+}\n+\n+\n static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\t\t      int hdrlen, const void *buf, unsigned long len,\n \t\t\t      time_t mtime, unsigned flags)\n@@ -1970,15 +2010,12 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tstatic struct strbuf filename = STRBUF_INIT;\n \n \tloose_object_path(the_repository, &filename, oid);\n-\n-\tfd = create_tmpfile(&tmp_file, filename.buf, flags);\n+\tfd = start_loose_object_common(&tmp_file, filename.buf, flags,\n+\t\t\t\t       &stream, compressed, sizeof(compressed),\n+\t\t\t\t       &c, OBJ_NONE, 0, hdr, &hdrlen);\n \tif (fd < 0)\n \t\treturn -1;\n \n-\t/* Set it up and write header */\n-\tsetup_stream_and_header(&stream, compressed, sizeof(compressed),\n-\t\t\t\t&c, hdr, hdrlen);\n-\n \t/* Then the data itself.. */\n \tstream.next_in = (void *)buf;\n \tstream.avail_in = len;\n@@ -1992,14 +2029,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tstream.avail_out = sizeof(compressed);\n \t} while (ret == Z_OK);\n \n-\tif (ret != Z_STREAM_END)\n-\t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n-\t\t    ret);\n-\tret = git_deflate_end_gently(&stream);\n-\tif (ret != Z_OK)\n-\t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n-\t\t    ret);\n-\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n+\tend_loose_object_common(ret, &c, &stream, &parano_oid, oid,\n+\t\t\t\tN_(\"unable to deflate new object %s (%d)\"),\n+\t\t\t\tN_(\"deflateEnd on object %s failed (%d)\"));\n \tif (!oideq(oid, &parano_oid))\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n@@ -2049,16 +2081,12 @@ int write_stream_object_file(struct input_stream *in_stream, size_t len,\n \t/* When oid is not determined, save tmp file to odb path. */\n \tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n \n-\tfd = create_tmpfile(&tmp_file, filename.buf, flags);\n+\tfd = start_loose_object_common(&tmp_file, filename.buf, flags,\n+\t\t\t\t       &stream, compressed, sizeof(compressed),\n+\t\t\t\t       &c, type, len, hdr, &hdrlen);\n \tif (fd < 0)\n \t\treturn -1;\n \n-\thdrlen = format_object_header(hdr, hdrlen, type, len);\n-\n-\t/* Set it up and write header */\n-\tsetup_stream_and_header(&stream, compressed, sizeof(compressed),\n-\t\t\t\t&c, hdr, hdrlen);\n-\n \t/* Then the data itself.. */\n \tdo {\n \t\tunsigned char *in0 = stream.next_in;\n@@ -2078,12 +2106,9 @@ int write_stream_object_file(struct input_stream *in_stream, size_t len,\n \t\tstream.avail_out = sizeof(compressed);\n \t} while (ret == Z_OK || ret == Z_BUF_ERROR);\n \n-\tif (ret != Z_STREAM_END)\n-\t\tdie(_(\"unable to deflate new object streamingly (%d)\"), ret);\n-\tret = git_deflate_end_gently(&stream);\n-\tif (ret != Z_OK)\n-\t\tdie(_(\"deflateEnd on object streamingly failed (%d)\"), ret);\n-\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n+\tend_loose_object_common(ret, &c, &stream, &parano_oid, NULL,\n+\t\t\t\tN_(\"unable to deflate new object streamingly (%d)\"),\n+\t\t\t\tN_(\"deflateEnd on object streamingly failed (%d)\"));\n \n \tclose_loose_object(fd);\n \n"},{"id":"444623","messageId":"211221.86pmpqq9aj.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"20211221115201.12120-6-chiyutianyi@gmail.com","subject":"Re: [PATCH v7 5/5] unpack-objects: unpack_non_delta_entry() read data in a stream","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-12-21T15:06:37Z","receivedAt":"2021-12-21T15:15:05Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Dec 21 2021, Han Xin wrote:\n\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> We used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\n> entire contents of a blob object, no matter how big it is. This\n> implementation may consume all the memory and cause OOM.\n>\n> By implementing a zstream version of input_stream interface, we can use\n> a small fixed buffer for \"unpack_non_delta_entry()\".\n>\n> However, unpack non-delta objects from a stream instead of from an\n> entrie buffer will have 10% performance penalty. Therefore, only unpack\n> object larger than the \"core.BigFileStreamingThreshold\" in zstream. See\n> the following benchmarks:\n>\n>     hyperfine \\\n>       --setup \\\n>       'if ! test -d scalar.git; then git clone --bare https://github.com/microsoft/scalar.git; cp scalar.git/objects/pack/*.pack small.pack; fi' \\\n>       --prepare 'rm -rf dest.git && git init --bare dest.git'\n>\n>     Summary\n>       './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'origin/master'\n>         1.01 ± 0.04 times faster than './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'HEAD~1'\n>         1.01 ± 0.04 times faster than './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'HEAD~0'\n>         1.03 ± 0.10 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'origin/master'\n>         1.02 ± 0.07 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'HEAD~0'\n>         1.10 ± 0.04 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'HEAD~1'\n>\n> Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> Helped-by: Derrick Stolee <stolee@gmail.com>\n> Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  Documentation/config/core.txt       | 11 +++++\n>  builtin/unpack-objects.c            | 73 ++++++++++++++++++++++++++++-\n>  cache.h                             |  1 +\n>  config.c                            |  5 ++\n>  environment.c                       |  1 +\n>  t/t5590-unpack-non-delta-objects.sh | 36 +++++++++++++-\n>  6 files changed, 125 insertions(+), 2 deletions(-)\n>\n> diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> index c04f62a54a..601b7a2418 100644\n> --- a/Documentation/config/core.txt\n> +++ b/Documentation/config/core.txt\n> @@ -424,6 +424,17 @@ be delta compressed, but larger binary media files won't be.\n>  +\n>  Common unit suffixes of 'k', 'm', or 'g' are supported.\n>  \n> +core.bigFileStreamingThreshold::\n> +\tFiles larger than this will be streamed out to a temporary\n> +\tobject file while being hashed, which will when be renamed\n> +\tin-place to a loose object, particularly if the\n> +\t`core.bigFileThreshold' setting dictates that they're always\n> +\twritten out as loose objects.\n> ++\n> +Default is 128 MiB on all platforms.\n> ++\n> +Common unit suffixes of 'k', 'm', or 'g' are supported.\n> +\n>  core.excludesFile::\n>  \tSpecifies the pathname to the file that contains patterns to\n>  \tdescribe paths that are not meant to be tracked, in addition\n> diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n> index 9104eb48da..72d8616e00 100644\n> --- a/builtin/unpack-objects.c\n> +++ b/builtin/unpack-objects.c\n> @@ -331,11 +331,82 @@ static void added_object(unsigned nr, enum object_type type,\n>  \t}\n>  }\n>  \n> +struct input_zstream_data {\n> +\tgit_zstream *zstream;\n> +\tunsigned char buf[8192];\n> +\tint status;\n> +};\n> +\n> +static const void *feed_input_zstream(struct input_stream *in_stream,\n> +\t\t\t\t      unsigned long *readlen)\n> +{\n> +\tstruct input_zstream_data *data = in_stream->data;\n> +\tgit_zstream *zstream = data->zstream;\n> +\tvoid *in = fill(1);\n> +\n> +\tif (!len || data->status == Z_STREAM_END) {\n> +\t\t*readlen = 0;\n> +\t\treturn NULL;\n> +\t}\n> +\n> +\tzstream->next_out = data->buf;\n> +\tzstream->avail_out = sizeof(data->buf);\n> +\tzstream->next_in = in;\n> +\tzstream->avail_in = len;\n> +\n> +\tdata->status = git_inflate(zstream, 0);\n> +\tuse(len - zstream->avail_in);\n> +\t*readlen = sizeof(data->buf) - zstream->avail_out;\n> +\n> +\treturn data->buf;\n> +}\n> +\n> +static void write_stream_blob(unsigned nr, size_t size)\n> +{\n> +\tgit_zstream zstream;\n> +\tstruct input_zstream_data data;\n> +\tstruct input_stream in_stream = {\n> +\t\t.read = feed_input_zstream,\n> +\t\t.data = &data,\n> +\t};\n> +\n> +\tmemset(&zstream, 0, sizeof(zstream));\n> +\tmemset(&data, 0, sizeof(data));\n\nnit/style: both of these memset can be replaced by \"{ 0 }\", e.g. \"git_zstream zstream = { 0 }\".\n\n> +\tdata.zstream = &zstream;\n> +\tgit_inflate_init(&zstream);\n> +\n> +\tif (write_stream_object_file(&in_stream, size, OBJ_BLOB, 0, 0,\n> +\t\t\t\t     &obj_list[nr].oid))\n\nSo at the end of this series we never pass in anything but blob here,\nmtime is always 0 etc. So there was no reason to create a factored out\nfinalize_object_file_with_mtime() earlier in the series.\n\nWell, I don't mind the finalize_object_file_with_mtime() exiting, but\nlet's not pretend this is more generalized than it is. We're unlikely to\never want to do this for non-blobs.\n\nThis on top of this series (and my local WIP fixups as I'm reviewing\nthis, so it won't cleanly apply, but the idea should be clear) makes\nthis simpler:\n\t\n\tdiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n\tindex 2f8d34a2e47..a3a1d4b266f 100644\n\t--- a/builtin/unpack-objects.c\n\t+++ b/builtin/unpack-objects.c\n\t@@ -375,8 +375,7 @@ static void write_stream_blob(unsigned nr, size_t size)\n\t \tdata.zstream = &zstream;\n\t \tgit_inflate_init(&zstream);\n\t \n\t-\tif (write_stream_object_file(&in_stream, size, OBJ_BLOB, 0, 0,\n\t-\t\t\t\t     &obj_list[nr].oid))\n\t+\tif (write_stream_object_file(&in_stream, size, &obj_list[nr].oid))\n\t \t\tdie(_(\"failed to write object in stream\"));\n\t \n\t \tif (zstream.total_out != size || data.status != Z_STREAM_END)\n\tdiff --git a/object-file.c b/object-file.c\n\tindex 7fc2363cfa1..0572b34fc5a 100644\n\t--- a/object-file.c\n\t+++ b/object-file.c\n\t@@ -2061,8 +2061,7 @@ static int freshen_packed_object(const struct object_id *oid)\n\t }\n\t \n\t int write_stream_object_file(struct input_stream *in_stream, size_t len,\n\t-\t\t\t     enum object_type type, time_t mtime,\n\t-\t\t\t     unsigned flags, struct object_id *oid)\n\t+\t\t\t     struct object_id *oid)\n\t {\n\t \tint fd, ret, flush = 0;\n\t \tunsigned char compressed[4096];\n\t@@ -2081,9 +2080,9 @@ int write_stream_object_file(struct input_stream *in_stream, size_t len,\n\t \t/* When oid is not determined, save tmp file to odb path. */\n\t \tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n\t \n\t-\tfd = start_loose_object_common(&tmp_file, filename.buf, flags,\n\t+\tfd = start_loose_object_common(&tmp_file, filename.buf, 0,\n\t \t\t\t\t       &stream, compressed, sizeof(compressed),\n\t-\t\t\t\t       &c, type, len, hdr, &hdrlen);\n\t+\t\t\t\t       &c, OBJ_BLOB, len, hdr, &hdrlen);\n\t \tif (fd < 0)\n\t \t\treturn -1;\n\t \n\t@@ -2135,7 +2134,7 @@ int write_stream_object_file(struct input_stream *in_stream, size_t len,\n\t \t\tstrbuf_release(&dir);\n\t \t}\n\t \n\t-\treturn finalize_object_file_with_mtime(tmp_file.buf, filename.buf, mtime, flags);\n\t+\treturn finalize_object_file(tmp_file.buf, filename.buf);\n\t }\n\t \n\t int write_object_file_flags(const void *buf, unsigned long len,\n\tdiff --git a/object-store.h b/object-store.h\n\tindex 87d370d39ca..1362b58a4d3 100644\n\t--- a/object-store.h\n\t+++ b/object-store.h\n\t@@ -257,8 +257,7 @@ int hash_write_object_file_literally(const void *buf, unsigned long len,\n\t \t\t\t\t     unsigned flags);\n\t \n\t int write_stream_object_file(struct input_stream *in_stream, size_t len,\n\t-\t\t\t     enum object_type type, time_t mtime,\n\t-\t\t\t     unsigned flags, struct object_id *oid);\n\t+\t\t\t     struct object_id *oid);\n\t \n\t /*\n\t  * Add an object file to the in-memory object store, without writing it\n\t\n\n> +\t\tdie(_(\"failed to write object in stream\"));\n> diff --git a/environment.c b/environment.c\n> index 0d06a31024..04bba593de 100644\n> --- a/environment.c\n> +++ b/environment.c\n> @@ -47,6 +47,7 @@ size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n>  size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n>  size_t delta_base_cache_limit = 96 * 1024 * 1024;\n>  unsigned long big_file_threshold = 512 * 1024 * 1024;\n> +unsigned long big_file_streaming_threshold = 128 * 1024 * 1024;\n>  int pager_use_color = 1;\n>  const char *editor_program;\n>  const char *askpass_program;\n> diff --git a/t/t5590-unpack-non-delta-objects.sh b/t/t5590-unpack-non-delta-objects.sh\n> index 48c4fb1ba3..8436cbf8db 100755\n> --- a/t/t5590-unpack-non-delta-objects.sh\n> +++ b/t/t5590-unpack-non-delta-objects.sh\n> @@ -13,6 +13,11 @@ export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n>  prepare_dest () {\n>  \ttest_when_finished \"rm -rf dest.git\" &&\n>  \tgit init --bare dest.git\n> +\tif test -n \"$1\"\n> +\tthen\n> +\t\tgit -C dest.git config core.bigFileStreamingThreshold $1\n> +\t\tgit -C dest.git config core.bigFileThreshold $1\n> +\tfi\n\nAll of this new code is missing \"&&\" to chain & test forfailures.\n"},{"id":"444764","messageId":"CANYiYbELNLg+5AZ5PYS_FJKOXBC7Sw9jxsyt0Pu7GQue8H1eDA@mail.gmail.com","threadId":"56672","inReplyTo":"3f1e3a59-1288-ecff-88b5-53bbd510a6a0@web.de","subject":"Re: [PATCH v7 1/5] unpack-objects.c: add dry_run mode for get_data()","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2021-12-22T11:15:23Z","receivedAt":"2021-12-22T11:15:43Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"On Wed, Dec 22, 2021 at 9:53 AM René Scharfe <l.s.r@web.de> wrote:\n>\n> Am 21.12.21 um 15:09 schrieb Ævar Arnfjörð Bjarmason:\n> > Maybe I'm misunderstanding this, but the commit message says it would be\n> > dangerous to allocate a very larger buffer, but here we only limit the\n> > size under \"dry_run\".\n>\n> This patch reduces the memory usage of dry runs, as its commit message\n> says.  The memory usage of one type of actual (non-dry) unpack is reduced\n> by patch 5.\n>\n\nFor Han Xin and me, it is very challenging to write better commit log\nin English.  Since the commit is moved to the beginning, the commit\nlog should be rewritten as follows:\n\nunpack-objects.c: low memory footprint for get_data() in dry_run mode\n\nAs the name implies, \"get_data(size)\" will allocate and return a given\nsize of memory. Allocating memory for a large blob object may cause the\nsystem to run out of memory. Before preparing to replace calling of\n\"get_data()\" to resolve unpack issue of large blob objects, refactor\n\"get_data()\" to reduce memory footprint for dry_run mode. Because\nin dry_run mode, \"get_data()\" is only used to check the integrity of\ndata, and the returned buffer is not used at all.\n\nTherefore, add the flag \"dry_run\" as an additional parameter of\n\"get_data()\" and reuse a small buffer in dry_run mode. Because in\ndry_run mode, the return buffer is not the entire data that the user\nwants, for this reason, we will release the buffer and return NULL.\n\nHan Xin, I think you can try to free the allocated buffer for dry_run\nmode inside \"get_data()\".\n\n--\nJiang Xin\n"},{"id":"444765","messageId":"CANYiYbHdMTS9CqTqKfiKjBc=6uVsUSz91Qk+gEdghq_=JXbOHQ@mail.gmail.com","threadId":"56672","inReplyTo":"211221.86bl1arqls.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v7 1/5] unpack-objects.c: add dry_run mode for get_data()","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2021-12-22T11:29:40Z","receivedAt":"2021-12-22T11:29:55Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"On Wed, Dec 22, 2021 at 8:37 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n> On Tue, Dec 21 2021, Han Xin wrote:\n>\n> I commented on this \"find\" usage in an earlier round, I think there's a\n> much easier way to do this. You're really just going back and forth\n> between checking whether or not all the objects are loose.\n>\n> I think that the below fix-up on top of this series is a better way to\n> do that, and more accurate. I.e. in your test here you check \"!\n> test_cmp\", which means that we could have some packed and some loose,\n> but really what you're meaning to check is a flip-flop between \"all\n> loose?\" and \"no loose?.\n>\n> In addition to that there was no reason to hardcode \"main\", we can just\n> use HEAD. All in all I think the below fix-up makes sense:\n>\n> diff --git a/t/t5590-unpack-non-delta-objects.sh b/t/t5590-unpack-non-delta-objects.sh\n> index 8436cbf8db6..d78bb89225d 100755\n> --- a/t/t5590-unpack-non-delta-objects.sh\n> +++ b/t/t5590-unpack-non-delta-objects.sh\n> @@ -5,9 +5,6 @@\n>\n>  test_description='Test unpack-objects with non-delta objects'\n>\n> -GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME=main\n> -export GIT_TEST_DEFAULT_INITIAL_BRANCH_NAME\n> -\n>  . ./test-lib.sh\n>\n>  prepare_dest () {\n> @@ -20,16 +17,22 @@ prepare_dest () {\n>         fi\n>  }\n>\n> +assert_no_loose () {\n> +       glob=dest.git/objects/?? &&\n> +       echo \"$glob\" >expect &&\n> +       echo $glob >actual &&\n\nIncompatible for zsh. This may work:\n\n    eval \"echo $glob\" >actual &&\n\n--\nJiang Xin\n"},{"id":"444767","messageId":"CANYiYbFN4fn6hG8_=kLDENnvUtGCsXnksgTeBNG=VYzyaM9yrQ@mail.gmail.com","threadId":"56672","inReplyTo":"211221.867dbyrqe9.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v7 3/5] object-file.c: refactor write_loose_object() to reuse in stream version","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2021-12-22T12:02:08Z","receivedAt":"2021-12-22T12:02:22Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"On Wed, Dec 22, 2021 at 8:40 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Tue, Dec 21 2021, Han Xin wrote:\n>\n> > From: Han Xin <hanxin.hx@alibaba-inc.com>\n> > [...]\n> > @@ -1854,17 +1876,48 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n> >               strbuf_reset(tmp);\n> >               strbuf_add(tmp, filename, dirlen - 1);\n> >               if (mkdir(tmp->buf, 0777) && errno != EEXIST)\n> > -                     return -1;\n> > +                     break;\n> >               if (adjust_shared_perm(tmp->buf))\n> > -                     return -1;\n> > +                     break;\n> >\n> >               /* Try again */\n> >               strbuf_addstr(tmp, \"/tmp_obj_XXXXXX\");\n> >               fd = git_mkstemp_mode(tmp->buf, 0444);\n> > +     } while (0);\n> > +\n> > +     if (fd < 0 && !(flags & HASH_SILENT)) {\n> > +             if (errno == EACCES)\n> > +                     return error(_(\"insufficient permission for adding an \"\n> > +                                    \"object to repository database %s\"),\n> > +                                  get_object_directory());\n>\n> This should be an error_errno() instead, ...\n\nWe already know the errno (EACCESS) and output a decent error message,\nso using error() is OK.  BTW, it's just a refactor by copy & paste.\n\n>\n> > +             else\n> > +                     return error_errno(_(\"unable to create temporary file\"));\n>\n> ...and we can just fold this whole if/else into one condition with a\n> briefer message, e.g.:\n>\n>     error_errno(_(\"unable to add object to '%s'\"), get_object_directory());\n>\n> Or whatever, unless there's another bug here where you inverted these\n> conditions, and the \"else\" really should not use \"error_errno\" but\n> \"error\".... (I don't know...)\n"},{"id":"445265","messageId":"CANYiYbESw-hP8R+075uGb-H_uJpAcBftn1AcyCcDZ_UbD_S6-Q@mail.gmail.com","threadId":"56672","inReplyTo":"20211221115201.12120-2-chiyutianyi@gmail.com","subject":"Re: [PATCH v7 1/5] unpack-objects.c: add dry_run mode for get_data()","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2021-12-31T03:06:21Z","receivedAt":"2021-12-31T03:06:36Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"On Wed, Dec 22, 2021 at 2:33 AM Han Xin <chiyutianyi@gmail.com> wrote:\n>\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> In dry_run mode, \"get_data()\" is used to verify the inflation of data,\n> and the returned buffer will not be used at all and will be freed\n> immediately. Even in dry_run mode, it is dangerous to allocate a\n> full-size buffer for a large blob object. Therefore, only allocate a\n> low memory footprint when calling \"get_data()\" in dry_run mode.\n>\n> Suggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  builtin/unpack-objects.c            | 23 +++++++++---\n>  t/t5590-unpack-non-delta-objects.sh | 57 +++++++++++++++++++++++++++++\n>  2 files changed, 74 insertions(+), 6 deletions(-)\n>  create mode 100755 t/t5590-unpack-non-delta-objects.sh\n>\n> diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n> index 4a9466295b..9104eb48da 100644\n> --- a/builtin/unpack-objects.c\n> +++ b/builtin/unpack-objects.c\n> @@ -96,15 +96,21 @@ static void use(int bytes)\n>         display_throughput(progress, consumed_bytes);\n>  }\n>\n> -static void *get_data(unsigned long size)\n> +static void *get_data(size_t size, int dry_run)\n\nAfter a offline talk with Han Xin, we feel it is not necessary to pass\n\"dry_run\" as a argument, use the file-scope static variable directly\nin \"get_data()\".\n\n>  {\n>         git_zstream stream;\n> -       void *buf = xmallocz(size);\n> +       size_t bufsize;\n> +       void *buf;\n>\n>         memset(&stream, 0, sizeof(stream));\n> +       if (dry_run && size > 8192)\n\nUse the file-scope static variable \"dry_run\".\n"},{"id":"445267","messageId":"CANYiYbFbhYtrbXtu12p8z0X0id4FSbd68T-5XLBD-b0c8s=zzQ@mail.gmail.com","threadId":"56672","inReplyTo":"20211221115201.12120-3-chiyutianyi@gmail.com","subject":"Re: [PATCH v7 2/5] object-file API: add a format_object_header() function","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2021-12-31T03:12:10Z","receivedAt":"2021-12-31T03:12:25Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"On Wed, Dec 22, 2021 at 2:56 AM Han Xin <chiyutianyi@gmail.com> wrote:\n>\n> From: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n>\n> Add a convenience function to wrap the xsnprintf() command that\n> generates loose object headers. This code was copy/pasted in various\n> parts of the codebase, let's define it in one place and re-use it from\n> there.\n>\n> All except one caller of it had a valid \"enum object_type\" for us,\n> it's only write_object_file_prepare() which might need to deal with\n> \"git hash-object --literally\" and a potential garbage type. Let's have\n> the primary API use an \"enum object_type\", and define an *_extended()\n> function that can take an arbitrary \"const char *\" for the type.\n>\n> See [1] for the discussion that prompted this patch, i.e. new code in\n> object-file.c that wanted to copy/paste the xsnprintf() invocation.\n>\n> 1. https://lore.kernel.org/git/211213.86bl1l9bfz.gmgdl@evledraar.gmail.com/\n>\n> Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  builtin/index-pack.c |  3 +--\n>  bulk-checkin.c       |  4 ++--\n>  cache.h              | 21 +++++++++++++++++++++\n>  http-push.c          |  2 +-\n>  object-file.c        | 14 +++++++++++---\n>  5 files changed, 36 insertions(+), 8 deletions(-)\n\nAfter a offline review with Han Xin, we feel it's better to move this\nfixup commit to the end of this series, and this commit will also fix\nan additional \"xsnprintf()\" we introduced in this series.\n\n--\nJiang Xin\n"},{"id":"445268","messageId":"CANYiYbGKiSH6WA2HbTtxV6S3-aURAkKh2VCrV9wp-5us-uzoXA@mail.gmail.com","threadId":"56672","inReplyTo":"20211221115201.12120-6-chiyutianyi@gmail.com","subject":"Re: [PATCH v7 5/5] unpack-objects: unpack_non_delta_entry() read data in a stream","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2021-12-31T03:19:56Z","receivedAt":"2021-12-31T03:20:10Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"On Wed, Dec 22, 2021 at 2:56 AM Han Xin <chiyutianyi@gmail.com> wrote:\n>\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> We used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\n> entire contents of a blob object, no matter how big it is. This\n> implementation may consume all the memory and cause OOM.\n>\n> By implementing a zstream version of input_stream interface, we can use\n> a small fixed buffer for \"unpack_non_delta_entry()\".\n>\n> However, unpack non-delta objects from a stream instead of from an\n> entrie buffer will have 10% performance penalty. Therefore, only unpack\n> object larger than the \"core.BigFileStreamingThreshold\" in zstream. See\n> the following benchmarks:\n>\n>     hyperfine \\\n>       --setup \\\n>       'if ! test -d scalar.git; then git clone --bare https://github.com/microsoft/scalar.git; cp scalar.git/objects/pack/*.pack small.pack; fi' \\\n>       --prepare 'rm -rf dest.git && git init --bare dest.git'\n>\n>     Summary\n>       './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'origin/master'\n>         1.01 ± 0.04 times faster than './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'HEAD~1'\n>         1.01 ± 0.04 times faster than './git -C dest.git -c core.bigfilethreshold=512m unpack-objects <small.pack' in 'HEAD~0'\n>         1.03 ± 0.10 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'origin/master'\n>         1.02 ± 0.07 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'HEAD~0'\n>         1.10 ± 0.04 times faster than './git -C dest.git -c core.bigfilethreshold=16k unpack-objects <small.pack' in 'HEAD~1'\n>\n> Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> Helped-by: Derrick Stolee <stolee@gmail.com>\n> Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  Documentation/config/core.txt       | 11 +++++\n>  builtin/unpack-objects.c            | 73 ++++++++++++++++++++++++++++-\n>  cache.h                             |  1 +\n>  config.c                            |  5 ++\n>  environment.c                       |  1 +\n>  t/t5590-unpack-non-delta-objects.sh | 36 +++++++++++++-\n>  6 files changed, 125 insertions(+), 2 deletions(-)\n>\n> diff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\n> index c04f62a54a..601b7a2418 100644\n> --- a/Documentation/config/core.txt\n> +++ b/Documentation/config/core.txt\n> @@ -424,6 +424,17 @@ be delta compressed, but larger binary media files won't be.\n>  +\n>  Common unit suffixes of 'k', 'm', or 'g' are supported.\n>\n> +core.bigFileStreamingThreshold::\n> +       Files larger than this will be streamed out to a temporary\n> +       object file while being hashed, which will when be renamed\n> +       in-place to a loose object, particularly if the\n> +       `core.bigFileThreshold' setting dictates that they're always\n> +       written out as loose objects.\n\nHan Xin told me the reason to introduce another git config variable,\nbut I feel it not good to introduce an application specific config\nvariable as \"core.XXX\" and parsing it in \"config.c\".\n\nSo in patch v8, will still reuse the config variable\n\"core.bigFileThreshold\", and will introduce an application specific\nconfig variable, such as unpack.bigFileThreshold and parse the new\nconfig in \"builtin/unpack-objects.c\".\n\n--\nJiang Xin\n"},{"id":"445783","messageId":"20220108085419.79682-1-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211217112629.12334-1-chiyutianyi@gmail.com","subject":"[PATCH v8 0/6] unpack large blobs in stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-01-08T08:54:13Z","receivedAt":"2022-01-08T08:56:33Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nChanges since v7:\n* Use functions \"assert_no_loose()\" and \"assert_no_pack()\" to do tests instead\n  of \"find\" sugguseted by Ævar Arnfjörð Bjarmason[1].\n\n* \"get_data()\" now use the global \"dry_run\" and it will release the buf before\n  returning.\n\n* Add a new commit \"object-file.c: remove the slash for directory_size()\"\n  sugguseted by Ævar Arnfjörð Bjarmason[2].\n\n* Add \"int is_finished\" to \"struct input_stream\" who will tell us if there is \n  next buffer in the stream.\n\n* Remove the config \"core.bigFileStreamingThreshold\" introduced in v5, and keep\n  using \"core.bigFileThreshold\". Until now, the config variable has been used in\n  the cases listed in \"unpack-objects: unpack_non_delta_entry() read data in a\n  stream\", this new case belongs to the packfile category.\n\n* Remove unnecessary explicit cast in \"object-file API: add a \n  format_object_header() function\" sugguseted by René Scharfe[3].\n\n1. https://lore.kernel.org/git/211221.86bl1arqls.gmgdl@evledraar.gmail.com/\n2. https://lore.kernel.org/git/211221.8635mmrpps.gmgdl@evledraar.gmail.com/\n3. https://lore.kernel.org/git/b2dee243-1a38-531e-02b1-ffd66c465fa5@web.de/\n\nHan Xin (5):\n  unpack-objects: low memory footprint for get_data() in dry_run mode\n  object-file.c: refactor write_loose_object() to several steps\n  object-file.c: remove the slash for directory_size()\n  object-file.c: add \"stream_loose_object()\" to handle large object\n  unpack-objects: unpack_non_delta_entry() read data in a stream\n\nÆvar Arnfjörð Bjarmason (1):\n  object-file API: add a format_object_header() function\n\n builtin/index-pack.c            |   3 +-\n builtin/unpack-objects.c        | 110 +++++++++++--\n bulk-checkin.c                  |   4 +-\n cache.h                         |  21 +++\n http-push.c                     |   2 +-\n object-file.c                   | 272 ++++++++++++++++++++++++++------\n object-store.h                  |   9 ++\n t/t5329-unpack-large-objects.sh |  69 ++++++++\n 8 files changed, 422 insertions(+), 68 deletions(-)\n create mode 100755 t/t5329-unpack-large-objects.sh\n\nRange-diff against v7:\n1:  a8f232f553 < -:  ---------- unpack-objects.c: add dry_run mode for get_data()\n-:  ---------- > 1:  bd34da5816 unpack-objects: low memory footprint for get_data() in dry_run mode\n3:  a571b8f16c ! 2:  f9a4365a7d object-file.c: refactor write_loose_object() to reuse in stream version\n    @@ Metadata\n     Author: Han Xin <hanxin.hx@alibaba-inc.com>\n     \n      ## Commit message ##\n    -    object-file.c: refactor write_loose_object() to reuse in stream version\n    +    object-file.c: refactor write_loose_object() to several steps\n     \n    -    We used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\n    -    entire contents of a blob object, no matter how big it is. This\n    -    implementation may consume all the memory and cause OOM.\n    +    When writing a large blob using \"write_loose_object()\", we have to pass\n    +    a buffer with the whole content of the blob, and this behavior will\n    +    consume lots of memory and may cause OOM. We will introduce a stream\n    +    version function (\"stream_loose_object()\") in latter commit to resolve\n    +    this issue.\n     \n    -    This can be improved by feeding data to \"stream_loose_object()\" in\n    -    stream instead of read into the whole buf.\n    +    Before introducing a stream vesion function for writing loose object,\n    +    do some refactoring on \"write_loose_object()\" to reuse code for both\n    +    versions.\n     \n    -    As this new method \"stream_loose_object()\" has many similarities with\n    -    \"write_loose_object()\", we split up \"write_loose_object()\" into some\n    -    steps:\n    -     1. Figuring out a path for the (temp) object file.\n    -     2. Creating the tempfile.\n    -     3. Setting up zlib and write header.\n    -     4. Write object data and handle errors.\n    -     5. Optionally, do someting after write, maybe force a loose object if\n    -    \"mtime\".\n    +    Rewrite \"write_loose_object()\" as follows:\n    +\n    +     1. Figure out a path for the (temp) object file. This step is only\n    +        used in \"write_loose_object()\".\n    +\n    +     2. Move common steps for starting to write loose objects into a new\n    +        function \"start_loose_object_common()\".\n    +\n    +     3. Compress data.\n    +\n    +     4. Move common steps for ending zlib stream into a new funciton\n    +        \"end_loose_object_common()\".\n    +\n    +     5. Close fd and finalize the object file.\n     \n         Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n    +    Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n         Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n     \n      ## object-file.c ##\n    @@ object-file.c: static int create_tmpfile(struct strbuf *tmp, const char *filenam\n      \treturn fd;\n      }\n      \n    -+static void setup_stream_and_header(git_zstream *stream,\n    -+\t\t\t\t    unsigned char *compressed,\n    -+\t\t\t\t    unsigned long compressed_size,\n    -+\t\t\t\t    git_hash_ctx *c,\n    -+\t\t\t\t    char *hdr,\n    -+\t\t\t\t    int hdrlen)\n    ++static int start_loose_object_common(struct strbuf *tmp_file,\n    ++\t\t\t\t     const char *filename, unsigned flags,\n    ++\t\t\t\t     git_zstream *stream,\n    ++\t\t\t\t     unsigned char *buf, size_t buflen,\n    ++\t\t\t\t     git_hash_ctx *c,\n    ++\t\t\t\t     enum object_type type, size_t len,\n    ++\t\t\t\t     char *hdr, int hdrlen)\n     +{\n    -+\t/* Set it up */\n    ++\tint fd;\n    ++\n    ++\tfd = create_tmpfile(tmp_file, filename, flags);\n    ++\tif (fd < 0)\n    ++\t\treturn -1;\n    ++\n    ++\t/*  Setup zlib stream for compression */\n     +\tgit_deflate_init(stream, zlib_compression_level);\n    -+\tstream->next_out = compressed;\n    -+\tstream->avail_out = compressed_size;\n    ++\tstream->next_out = buf;\n    ++\tstream->avail_out = buflen;\n     +\tthe_hash_algo->init_fn(c);\n     +\n    -+\t/* First header.. */\n    ++\t/*  Start to feed header to zlib stream */\n     +\tstream->next_in = (unsigned char *)hdr;\n     +\tstream->avail_in = hdrlen;\n     +\twhile (git_deflate(stream, 0) == Z_OK)\n     +\t\t; /* nothing */\n     +\tthe_hash_algo->update_fn(c, hdr, hdrlen);\n    ++\n    ++\treturn fd;\n    ++}\n    ++\n    ++static void end_loose_object_common(int ret, git_hash_ctx *c,\n    ++\t\t\t\t    git_zstream *stream,\n    ++\t\t\t\t    struct object_id *parano_oid,\n    ++\t\t\t\t    const struct object_id *expected_oid,\n    ++\t\t\t\t    const char *die_msg1_fmt,\n    ++\t\t\t\t    const char *die_msg2_fmt)\n    ++{\n    ++\tif (ret != Z_STREAM_END)\n    ++\t\tdie(_(die_msg1_fmt), ret, expected_oid);\n    ++\tret = git_deflate_end_gently(stream);\n    ++\tif (ret != Z_OK)\n    ++\t\tdie(_(die_msg2_fmt), ret, expected_oid);\n    ++\tthe_hash_algo->final_oid_fn(parano_oid, c);\n     +}\n     +\n      static int write_loose_object(const struct object_id *oid, char *hdr,\n    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *\n     -\tstream.next_out = compressed;\n     -\tstream.avail_out = sizeof(compressed);\n     -\tthe_hash_algo->init_fn(&c);\n    -+\tfd = create_tmpfile(&tmp_file, filename.buf, flags);\n    -+\tif (fd < 0)\n    -+\t\treturn -1;\n    - \n    +-\n     -\t/* First header.. */\n     -\tstream.next_in = (unsigned char *)hdr;\n     -\tstream.avail_in = hdrlen;\n     -\twhile (git_deflate(&stream, 0) == Z_OK)\n     -\t\t; /* nothing */\n     -\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n    -+\t/* Set it up and write header */\n    -+\tsetup_stream_and_header(&stream, compressed, sizeof(compressed),\n    -+\t\t\t\t&c, hdr, hdrlen);\n    ++\t/* Common steps for write_loose_object and stream_loose_object to\n    ++\t * start writing loose oject:\n    ++\t *\n    ++\t *  - Create tmpfile for the loose object.\n    ++\t *  - Setup zlib stream for compression.\n    ++\t *  - Start to feed header to zlib stream.\n    ++\t */\n    ++\tfd = start_loose_object_common(&tmp_file, filename.buf, flags,\n    ++\t\t\t\t       &stream, compressed, sizeof(compressed),\n    ++\t\t\t\t       &c, OBJ_NONE, 0, hdr, hdrlen);\n    ++\tif (fd < 0)\n    ++\t\treturn -1;\n      \n      \t/* Then the data itself.. */\n      \tstream.next_in = (void *)buf;\n     @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *hdr,\n    + \t\tstream.avail_out = sizeof(compressed);\n    + \t} while (ret == Z_OK);\n    + \n    +-\tif (ret != Z_STREAM_END)\n    +-\t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n    +-\t\t    ret);\n    +-\tret = git_deflate_end_gently(&stream);\n    +-\tif (ret != Z_OK)\n    +-\t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n    +-\t\t    ret);\n    +-\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n    ++\t/* Common steps for write_loose_object and stream_loose_object to\n    ++\t * end writing loose oject:\n    ++\t *\n    ++\t *  - End the compression of zlib stream.\n    ++\t *  - Get the calculated oid to \"parano_oid\".\n    ++\t */\n    ++\tend_loose_object_common(ret, &c, &stream, &parano_oid, oid,\n    ++\t\t\t\tN_(\"unable to deflate new object %s (%d)\"),\n    ++\t\t\t\tN_(\"deflateEnd on object %s failed (%d)\"));\n    ++\n    + \tif (!oideq(oid, &parano_oid))\n    + \t\tdie(_(\"confused by unstable object source data for %s\"),\n    + \t\t    oid_to_hex(oid));\n      \n      \tclose_loose_object(fd);\n      \n-:  ---------- > 3:  18dd21122d object-file.c: remove the slash for directory_size()\n-:  ---------- > 4:  964715451b object-file.c: add \"stream_loose_object()\" to handle large object\n-:  ---------- > 5:  3f620466fe unpack-objects: unpack_non_delta_entry() read data in a stream\n2:  0d2e0f3a00 ! 6:  8073a3888d object-file API: add a format_object_header() function\n    @@ builtin/index-pack.c: static void *unpack_entry_data(off_t offset, unsigned long\n      \tif (!is_delta_type(type)) {\n     -\t\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX,\n     -\t\t\t\t   type_name(type),(uintmax_t)size) + 1;\n    -+\t\thdrlen = format_object_header(hdr, sizeof(hdr), type, (uintmax_t)size);\n    ++\t\thdrlen = format_object_header(hdr, sizeof(hdr), type, size);\n      \t\tthe_hash_algo->init_fn(&c);\n      \t\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n      \t} else\n    @@ bulk-checkin.c: static int deflate_to_pack(struct bulk_checkin_state *state,\n     -\theader_len = xsnprintf((char *)obuf, sizeof(obuf), \"%s %\" PRIuMAX,\n     -\t\t\t       type_name(type), (uintmax_t)size) + 1;\n     +\theader_len = format_object_header((char *)obuf, sizeof(obuf),\n    -+\t\t\t\t\t type, (uintmax_t)size);\n    ++\t\t\t\t\t type, size);\n      \tthe_hash_algo->init_fn(&ctx);\n      \tthe_hash_algo->update_fn(&ctx, obuf, header_len);\n      \n    @@ http-push.c: static void start_put(struct transfer_request *request)\n      \n      \tunpacked = read_object_file(&request->obj->oid, &type, &len);\n     -\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n    -+\thdrlen = format_object_header(hdr, sizeof(hdr), type, (uintmax_t)len);\n    ++\thdrlen = format_object_header(hdr, sizeof(hdr), type, len);\n      \n      \t/* Set it up */\n      \tgit_deflate_init(&stream, zlib_compression_level);\n    @@ object-file.c: static void write_object_file_prepare(const struct git_hash_algo\n      \n      \t/* Sha1.. */\n      \talgo->init_fn(&c);\n    +@@ object-file.c: int stream_loose_object(struct input_stream *in_stream, size_t len,\n    + \n    + \t/* Since oid is not determined, save tmp file to odb path. */\n    + \tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n    +-\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX, type_name(OBJ_BLOB), len) + 1;\n    ++\thdrlen = format_object_header(hdr, sizeof(hdr), OBJ_BLOB, len);\n    + \n    + \t/* Common steps for write_loose_object and stream_loose_object to\n    + \t * start writing loose oject:\n     @@ object-file.c: int force_object_loose(const struct object_id *oid, time_t mtime)\n      \tbuf = read_object(the_repository, oid, &type, &len);\n      \tif (!buf)\n4:  1de06a8f5c < -:  ---------- object-file.c: add \"write_stream_object_file()\" to support read in stream\n5:  e7b4e426ef < -:  ---------- unpack-objects: unpack_non_delta_entry() read data in a stream\n-- \n2.34.1.52.gc288e771b4.agit.6.5.6\n\n"},{"id":"445784","messageId":"20220108085419.79682-2-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211217112629.12334-1-chiyutianyi@gmail.com","subject":"[PATCH v8 1/6] unpack-objects: low memory footprint for get_data() in dry_run mode","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-01-08T08:54:14Z","receivedAt":"2022-01-08T08:56:36Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nAs the name implies, \"get_data(size)\" will allocate and return a given\nsize of memory. Allocating memory for a large blob object may cause the\nsystem to run out of memory. Before preparing to replace calling of\n\"get_data()\" to unpack large blob objects in latter commits, refactor\n\"get_data()\" to reduce memory footprint for dry_run mode.\n\nBecause in dry_run mode, \"get_data()\" is only used to check the\nintegrity of data, and the returned buffer is not used at all, we can\nallocate a smaller buffer and reuse it as zstream output. Therefore,\nin dry_run mode, \"get_data()\" will release the allocated buffer and\nreturn NULL instead of returning garbage data.\n\nSuggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n builtin/unpack-objects.c        | 39 ++++++++++++++++++-------\n t/t5329-unpack-large-objects.sh | 52 +++++++++++++++++++++++++++++++++\n 2 files changed, 80 insertions(+), 11 deletions(-)\n create mode 100755 t/t5329-unpack-large-objects.sh\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 4a9466295b..c6d6c17072 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -96,15 +96,31 @@ static void use(int bytes)\n \tdisplay_throughput(progress, consumed_bytes);\n }\n \n+/*\n+ * Decompress zstream from stdin and return specific size of data.\n+ * The caller is responsible to free the returned buffer.\n+ *\n+ * But for dry_run mode, \"get_data()\" is only used to check the\n+ * integrity of data, and the returned buffer is not used at all.\n+ * Therefore, in dry_run mode, \"get_data()\" will release the small\n+ * allocated buffer which is reused to hold temporary zstream output\n+ * and return NULL instead of returning garbage data.\n+ */\n static void *get_data(unsigned long size)\n {\n \tgit_zstream stream;\n-\tvoid *buf = xmallocz(size);\n+\tunsigned long bufsize;\n+\tvoid *buf;\n \n \tmemset(&stream, 0, sizeof(stream));\n+\tif (dry_run && size > 8192)\n+\t\tbufsize = 8192;\n+\telse\n+\t\tbufsize = size;\n+\tbuf = xmallocz(bufsize);\n \n \tstream.next_out = buf;\n-\tstream.avail_out = size;\n+\tstream.avail_out = bufsize;\n \tstream.next_in = fill(1);\n \tstream.avail_in = len;\n \tgit_inflate_init(&stream);\n@@ -124,8 +140,15 @@ static void *get_data(unsigned long size)\n \t\t}\n \t\tstream.next_in = fill(1);\n \t\tstream.avail_in = len;\n+\t\tif (dry_run) {\n+\t\t\t/* reuse the buffer in dry_run mode */\n+\t\t\tstream.next_out = buf;\n+\t\t\tstream.avail_out = bufsize;\n+\t\t}\n \t}\n \tgit_inflate_end(&stream);\n+\tif (dry_run)\n+\t\tFREE_AND_NULL(buf);\n \treturn buf;\n }\n \n@@ -325,10 +348,8 @@ static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n {\n \tvoid *buf = get_data(size);\n \n-\tif (!dry_run && buf)\n+\tif (buf)\n \t\twrite_object(nr, type, buf, size);\n-\telse\n-\t\tfree(buf);\n }\n \n static int resolve_against_held(unsigned nr, const struct object_id *base,\n@@ -358,10 +379,8 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\toidread(&base_oid, fill(the_hash_algo->rawsz));\n \t\tuse(the_hash_algo->rawsz);\n \t\tdelta_data = get_data(delta_size);\n-\t\tif (dry_run || !delta_data) {\n-\t\t\tfree(delta_data);\n+\t\tif (!delta_data)\n \t\t\treturn;\n-\t\t}\n \t\tif (has_object_file(&base_oid))\n \t\t\t; /* Ok we have this one */\n \t\telse if (resolve_against_held(nr, &base_oid,\n@@ -397,10 +416,8 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\t\tdie(\"offset value out of bound for delta base object\");\n \n \t\tdelta_data = get_data(delta_size);\n-\t\tif (dry_run || !delta_data) {\n-\t\t\tfree(delta_data);\n+\t\tif (!delta_data)\n \t\t\treturn;\n-\t\t}\n \t\tlo = 0;\n \t\thi = nr;\n \t\twhile (lo < hi) {\ndiff --git a/t/t5329-unpack-large-objects.sh b/t/t5329-unpack-large-objects.sh\nnew file mode 100755\nindex 0000000000..39c7a62d94\n--- /dev/null\n+++ b/t/t5329-unpack-large-objects.sh\n@@ -0,0 +1,52 @@\n+#!/bin/sh\n+#\n+# Copyright (c) 2021 Han Xin\n+#\n+\n+test_description='git unpack-objects with large objects'\n+\n+. ./test-lib.sh\n+\n+prepare_dest () {\n+\ttest_when_finished \"rm -rf dest.git\" &&\n+\tgit init --bare dest.git\n+}\n+\n+assert_no_loose () {\n+\tglob=dest.git/objects/?? &&\n+\techo \"$glob\" >expect &&\n+\teval \"echo $glob\" >actual &&\n+\ttest_cmp expect actual\n+}\n+\n+assert_no_pack () {\n+\trmdir dest.git/objects/pack\n+}\n+\n+test_expect_success \"create large objects (1.5 MB) and PACK\" '\n+\ttest-tool genrandom foo 1500000 >big-blob &&\n+\ttest_commit --append foo big-blob &&\n+\ttest-tool genrandom bar 1500000 >big-blob &&\n+\ttest_commit --append bar big-blob &&\n+\tPACK=$(echo HEAD | git pack-objects --revs test)\n+'\n+\n+test_expect_success 'set memory limitation to 1MB' '\n+\tGIT_ALLOC_LIMIT=1m &&\n+\texport GIT_ALLOC_LIMIT\n+'\n+\n+test_expect_success 'unpack-objects failed under memory limitation' '\n+\tprepare_dest &&\n+\ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n+\tgrep \"fatal: attempting to allocate\" err\n+'\n+\n+test_expect_success 'unpack-objects works with memory limitation in dry-run mode' '\n+\tprepare_dest &&\n+\tgit -C dest.git unpack-objects -n <test-$PACK.pack &&\n+\tassert_no_loose &&\n+\tassert_no_pack\n+'\n+\n+test_done\n-- \n2.34.1.52.gc288e771b4.agit.6.5.6\n\n"},{"id":"445785","messageId":"20220108085419.79682-3-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211217112629.12334-1-chiyutianyi@gmail.com","subject":"[PATCH v8 2/6] object-file.c: refactor write_loose_object() to several steps","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-01-08T08:54:15Z","receivedAt":"2022-01-08T08:56:37Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWhen writing a large blob using \"write_loose_object()\", we have to pass\na buffer with the whole content of the blob, and this behavior will\nconsume lots of memory and may cause OOM. We will introduce a stream\nversion function (\"stream_loose_object()\") in latter commit to resolve\nthis issue.\n\nBefore introducing a stream vesion function for writing loose object,\ndo some refactoring on \"write_loose_object()\" to reuse code for both\nversions.\n\nRewrite \"write_loose_object()\" as follows:\n\n 1. Figure out a path for the (temp) object file. This step is only\n    used in \"write_loose_object()\".\n\n 2. Move common steps for starting to write loose objects into a new\n    function \"start_loose_object_common()\".\n\n 3. Compress data.\n\n 4. Move common steps for ending zlib stream into a new funciton\n    \"end_loose_object_common()\".\n\n 5. Close fd and finalize the object file.\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c | 149 +++++++++++++++++++++++++++++++++++---------------\n 1 file changed, 105 insertions(+), 44 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex eb1426f98c..5d163081b1 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1743,6 +1743,25 @@ static void write_object_file_prepare(const struct git_hash_algo *algo,\n \talgo->final_oid_fn(oid, &c);\n }\n \n+/*\n+ * Move the just written object with proper mtime into its final resting place.\n+ */\n+static int finalize_object_file_with_mtime(const char *tmpfile,\n+\t\t\t\t\t   const char *filename,\n+\t\t\t\t\t   time_t mtime,\n+\t\t\t\t\t   unsigned flags)\n+{\n+\tstruct utimbuf utb;\n+\n+\tif (mtime) {\n+\t\tutb.actime = mtime;\n+\t\tutb.modtime = mtime;\n+\t\tif (utime(tmpfile, &utb) < 0 && !(flags & HASH_SILENT))\n+\t\t\twarning_errno(_(\"failed utime() on %s\"), tmpfile);\n+\t}\n+\treturn finalize_object_file(tmpfile, filename);\n+}\n+\n /*\n  * Move the just written object into its final resting place.\n  */\n@@ -1828,7 +1847,8 @@ static inline int directory_size(const char *filename)\n  * We want to avoid cross-directory filename renames, because those\n  * can have problems on various filesystems (FAT, NFS, Coda).\n  */\n-static int create_tmpfile(struct strbuf *tmp, const char *filename)\n+static int create_tmpfile(struct strbuf *tmp, const char *filename,\n+\t\t\t  unsigned flags)\n {\n \tint fd, dirlen = directory_size(filename);\n \n@@ -1836,7 +1856,9 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n \tstrbuf_add(tmp, filename, dirlen);\n \tstrbuf_addstr(tmp, \"tmp_obj_XXXXXX\");\n \tfd = git_mkstemp_mode(tmp->buf, 0444);\n-\tif (fd < 0 && dirlen && errno == ENOENT) {\n+\tdo {\n+\t\tif (fd >= 0 || !dirlen || errno != ENOENT)\n+\t\t\tbreak;\n \t\t/*\n \t\t * Make sure the directory exists; note that the contents\n \t\t * of the buffer are undefined after mkstemp returns an\n@@ -1846,17 +1868,72 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n \t\tstrbuf_reset(tmp);\n \t\tstrbuf_add(tmp, filename, dirlen - 1);\n \t\tif (mkdir(tmp->buf, 0777) && errno != EEXIST)\n-\t\t\treturn -1;\n+\t\t\tbreak;\n \t\tif (adjust_shared_perm(tmp->buf))\n-\t\t\treturn -1;\n+\t\t\tbreak;\n \n \t\t/* Try again */\n \t\tstrbuf_addstr(tmp, \"/tmp_obj_XXXXXX\");\n \t\tfd = git_mkstemp_mode(tmp->buf, 0444);\n+\t} while (0);\n+\n+\tif (fd < 0 && !(flags & HASH_SILENT)) {\n+\t\tif (errno == EACCES)\n+\t\t\treturn error(_(\"insufficient permission for adding an \"\n+\t\t\t\t       \"object to repository database %s\"),\n+\t\t\t\t     get_object_directory());\n+\t\telse\n+\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n \t}\n+\n \treturn fd;\n }\n \n+static int start_loose_object_common(struct strbuf *tmp_file,\n+\t\t\t\t     const char *filename, unsigned flags,\n+\t\t\t\t     git_zstream *stream,\n+\t\t\t\t     unsigned char *buf, size_t buflen,\n+\t\t\t\t     git_hash_ctx *c,\n+\t\t\t\t     enum object_type type, size_t len,\n+\t\t\t\t     char *hdr, int hdrlen)\n+{\n+\tint fd;\n+\n+\tfd = create_tmpfile(tmp_file, filename, flags);\n+\tif (fd < 0)\n+\t\treturn -1;\n+\n+\t/*  Setup zlib stream for compression */\n+\tgit_deflate_init(stream, zlib_compression_level);\n+\tstream->next_out = buf;\n+\tstream->avail_out = buflen;\n+\tthe_hash_algo->init_fn(c);\n+\n+\t/*  Start to feed header to zlib stream */\n+\tstream->next_in = (unsigned char *)hdr;\n+\tstream->avail_in = hdrlen;\n+\twhile (git_deflate(stream, 0) == Z_OK)\n+\t\t; /* nothing */\n+\tthe_hash_algo->update_fn(c, hdr, hdrlen);\n+\n+\treturn fd;\n+}\n+\n+static void end_loose_object_common(int ret, git_hash_ctx *c,\n+\t\t\t\t    git_zstream *stream,\n+\t\t\t\t    struct object_id *parano_oid,\n+\t\t\t\t    const struct object_id *expected_oid,\n+\t\t\t\t    const char *die_msg1_fmt,\n+\t\t\t\t    const char *die_msg2_fmt)\n+{\n+\tif (ret != Z_STREAM_END)\n+\t\tdie(_(die_msg1_fmt), ret, expected_oid);\n+\tret = git_deflate_end_gently(stream);\n+\tif (ret != Z_OK)\n+\t\tdie(_(die_msg2_fmt), ret, expected_oid);\n+\tthe_hash_algo->final_oid_fn(parano_oid, c);\n+}\n+\n static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\t\t      int hdrlen, const void *buf, unsigned long len,\n \t\t\t      time_t mtime, unsigned flags)\n@@ -1871,28 +1948,18 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \n \tloose_object_path(the_repository, &filename, oid);\n \n-\tfd = create_tmpfile(&tmp_file, filename.buf);\n-\tif (fd < 0) {\n-\t\tif (flags & HASH_SILENT)\n-\t\t\treturn -1;\n-\t\telse if (errno == EACCES)\n-\t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n-\t\telse\n-\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n-\t}\n-\n-\t/* Set it up */\n-\tgit_deflate_init(&stream, zlib_compression_level);\n-\tstream.next_out = compressed;\n-\tstream.avail_out = sizeof(compressed);\n-\tthe_hash_algo->init_fn(&c);\n-\n-\t/* First header.. */\n-\tstream.next_in = (unsigned char *)hdr;\n-\tstream.avail_in = hdrlen;\n-\twhile (git_deflate(&stream, 0) == Z_OK)\n-\t\t; /* nothing */\n-\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n+\t/* Common steps for write_loose_object and stream_loose_object to\n+\t * start writing loose oject:\n+\t *\n+\t *  - Create tmpfile for the loose object.\n+\t *  - Setup zlib stream for compression.\n+\t *  - Start to feed header to zlib stream.\n+\t */\n+\tfd = start_loose_object_common(&tmp_file, filename.buf, flags,\n+\t\t\t\t       &stream, compressed, sizeof(compressed),\n+\t\t\t\t       &c, OBJ_NONE, 0, hdr, hdrlen);\n+\tif (fd < 0)\n+\t\treturn -1;\n \n \t/* Then the data itself.. */\n \tstream.next_in = (void *)buf;\n@@ -1907,30 +1974,24 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tstream.avail_out = sizeof(compressed);\n \t} while (ret == Z_OK);\n \n-\tif (ret != Z_STREAM_END)\n-\t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n-\t\t    ret);\n-\tret = git_deflate_end_gently(&stream);\n-\tif (ret != Z_OK)\n-\t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n-\t\t    ret);\n-\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n+\t/* Common steps for write_loose_object and stream_loose_object to\n+\t * end writing loose oject:\n+\t *\n+\t *  - End the compression of zlib stream.\n+\t *  - Get the calculated oid to \"parano_oid\".\n+\t */\n+\tend_loose_object_common(ret, &c, &stream, &parano_oid, oid,\n+\t\t\t\tN_(\"unable to deflate new object %s (%d)\"),\n+\t\t\t\tN_(\"deflateEnd on object %s failed (%d)\"));\n+\n \tif (!oideq(oid, &parano_oid))\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n \n \tclose_loose_object(fd);\n \n-\tif (mtime) {\n-\t\tstruct utimbuf utb;\n-\t\tutb.actime = mtime;\n-\t\tutb.modtime = mtime;\n-\t\tif (utime(tmp_file.buf, &utb) < 0 &&\n-\t\t    !(flags & HASH_SILENT))\n-\t\t\twarning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n-\t}\n-\n-\treturn finalize_object_file(tmp_file.buf, filename.buf);\n+\treturn finalize_object_file_with_mtime(tmp_file.buf, filename.buf,\n+\t\t\t\t\t       mtime, flags);\n }\n \n static int freshen_loose_object(const struct object_id *oid)\n-- \n2.34.1.52.gc288e771b4.agit.6.5.6\n\n"},{"id":"445786","messageId":"20220108085419.79682-4-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211217112629.12334-1-chiyutianyi@gmail.com","subject":"[PATCH v8 3/6] object-file.c: remove the slash for directory_size()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-01-08T08:54:16Z","receivedAt":"2022-01-08T08:56:39Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nSince \"mkdir foo/\" works as well as \"mkdir foo\", let's remove the end\nslash as many users of it want.\n\nSuggested-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c | 8 ++++----\n 1 file changed, 4 insertions(+), 4 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 5d163081b1..4f0127e823 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1831,13 +1831,13 @@ static void close_loose_object(int fd)\n \t\tdie_errno(_(\"error when closing loose object file\"));\n }\n \n-/* Size of directory component, including the ending '/' */\n+/* Size of directory component, excluding the ending '/' */\n static inline int directory_size(const char *filename)\n {\n \tconst char *s = strrchr(filename, '/');\n \tif (!s)\n \t\treturn 0;\n-\treturn s - filename + 1;\n+\treturn s - filename;\n }\n \n /*\n@@ -1854,7 +1854,7 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename,\n \n \tstrbuf_reset(tmp);\n \tstrbuf_add(tmp, filename, dirlen);\n-\tstrbuf_addstr(tmp, \"tmp_obj_XXXXXX\");\n+\tstrbuf_addstr(tmp, \"/tmp_obj_XXXXXX\");\n \tfd = git_mkstemp_mode(tmp->buf, 0444);\n \tdo {\n \t\tif (fd >= 0 || !dirlen || errno != ENOENT)\n@@ -1866,7 +1866,7 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename,\n \t\t * scratch.\n \t\t */\n \t\tstrbuf_reset(tmp);\n-\t\tstrbuf_add(tmp, filename, dirlen - 1);\n+\t\tstrbuf_add(tmp, filename, dirlen);\n \t\tif (mkdir(tmp->buf, 0777) && errno != EEXIST)\n \t\t\tbreak;\n \t\tif (adjust_shared_perm(tmp->buf))\n-- \n2.34.1.52.gc288e771b4.agit.6.5.6\n\n"},{"id":"445787","messageId":"20220108085419.79682-5-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211217112629.12334-1-chiyutianyi@gmail.com","subject":"[PATCH v8 4/6] object-file.c: add \"stream_loose_object()\" to handle large object","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-01-08T08:54:17Z","receivedAt":"2022-01-08T08:56:47Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nIf we want unpack and write a loose object using \"write_loose_object\",\nwe have to feed it with a buffer with the same size of the object, which\nwill consume lots of memory and may cause OOM. This can be improved by\nfeeding data to \"stream_loose_object()\" in a stream.\n\nAdd a new function \"stream_loose_object()\", which is a stream version of\n\"write_loose_object()\" but with a low memory footprint. We will use this\nfunction to unpack large blob object in latter commit.\n\nAnother difference with \"write_loose_object()\" is that we have no chance\nto run \"write_object_file_prepare()\" to calculate the oid in advance.\nIn \"write_loose_object()\", we know the oid and we can write the\ntemporary file in the same directory as the final object, but for an\nobject with an undetermined oid, we don't know the exact directory for\nthe object, so we have to save the temporary file in \".git/objects/\"\ndirectory instead.\n\n\"freshen_packed_object()\" or \"freshen_loose_object()\" will be called\ninside \"stream_loose_object()\" after obtaining the \"oid\".\n\nHelped-by: René Scharfe <l.s.r@web.de>\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c  | 101 +++++++++++++++++++++++++++++++++++++++++++++++++\n object-store.h |   9 +++++\n 2 files changed, 110 insertions(+)\n\ndiff --git a/object-file.c b/object-file.c\nindex 4f0127e823..a462a21629 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2012,6 +2012,107 @@ static int freshen_packed_object(const struct object_id *oid)\n \treturn 1;\n }\n \n+int stream_loose_object(struct input_stream *in_stream, size_t len,\n+\t\t\tstruct object_id *oid)\n+{\n+\tint fd, ret, err = 0, flush = 0;\n+\tunsigned char compressed[4096];\n+\tgit_zstream stream;\n+\tgit_hash_ctx c;\n+\tstruct strbuf tmp_file = STRBUF_INIT;\n+\tstruct strbuf filename = STRBUF_INIT;\n+\tint dirlen;\n+\tchar hdr[MAX_HEADER_LEN];\n+\tint hdrlen;\n+\n+\t/* Since oid is not determined, save tmp file to odb path. */\n+\tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n+\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX, type_name(OBJ_BLOB), len) + 1;\n+\n+\t/* Common steps for write_loose_object and stream_loose_object to\n+\t * start writing loose oject:\n+\t *\n+\t *  - Create tmpfile for the loose object.\n+\t *  - Setup zlib stream for compression.\n+\t *  - Start to feed header to zlib stream.\n+\t */\n+\tfd = start_loose_object_common(&tmp_file, filename.buf, 0,\n+\t\t\t\t       &stream, compressed, sizeof(compressed),\n+\t\t\t\t       &c, OBJ_BLOB, len, hdr, hdrlen);\n+\tif (fd < 0) {\n+\t\terr = -1;\n+\t\tgoto cleanup;\n+\t}\n+\n+\t/* Then the data itself.. */\n+\tdo {\n+\t\tunsigned char *in0 = stream.next_in;\n+\t\tif (!stream.avail_in && !in_stream->is_finished) {\n+\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n+\t\t\tstream.next_in = (void *)in;\n+\t\t\tin0 = (unsigned char *)in;\n+\t\t\t/* All data has been read. */\n+\t\t\tif (in_stream->is_finished)\n+\t\t\t\tflush = Z_FINISH;\n+\t\t}\n+\t\tret = git_deflate(&stream, flush);\n+\t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n+\t\tif (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n+\t\t\tdie(_(\"unable to write loose object file\"));\n+\t\tstream.next_out = compressed;\n+\t\tstream.avail_out = sizeof(compressed);\n+\t\t/*\n+\t\t * Unlike write_loose_object(), we do not have the entire\n+\t\t * buffer. If we get Z_BUF_ERROR due to too few input bytes,\n+\t\t * then we'll replenish them in the next input_stream->read()\n+\t\t * call when we loop.\n+\t\t */\n+\t} while (ret == Z_OK || ret == Z_BUF_ERROR);\n+\n+\tif (stream.total_in != len + hdrlen)\n+\t\tdie(_(\"write stream object %ld != %\"PRIuMAX), stream.total_in,\n+\t\t    (uintmax_t)len + hdrlen);\n+\n+\t/* Common steps for write_loose_object and stream_loose_object to\n+\t * end writing loose oject:\n+\t *\n+\t *  - End the compression of zlib stream.\n+\t *  - Get the calculated oid.\n+\t */\n+\tend_loose_object_common(ret, &c, &stream, oid, NULL,\n+\t\t\t\tN_(\"unable to stream deflate new object (%d)\"),\n+\t\t\t\tN_(\"deflateEnd on stream object failed (%d)\"));\n+\n+\tclose_loose_object(fd);\n+\n+\tif (freshen_packed_object(oid) || freshen_loose_object(oid)) {\n+\t\tunlink_or_warn(tmp_file.buf);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tloose_object_path(the_repository, &filename, oid);\n+\n+\t/* We finally know the object path, and create the missing dir. */\n+\tdirlen = directory_size(filename.buf);\n+\tif (dirlen) {\n+\t\tstruct strbuf dir = STRBUF_INIT;\n+\t\tstrbuf_add(&dir, filename.buf, dirlen);\n+\n+\t\tif (mkdir_in_gitdir(dir.buf) && errno != EEXIST) {\n+\t\t\terr = error_errno(_(\"unable to create directory %s\"), dir.buf);\n+\t\t\tstrbuf_release(&dir);\n+\t\t\tgoto cleanup;\n+\t\t}\n+\t\tstrbuf_release(&dir);\n+\t}\n+\n+\terr = finalize_object_file(tmp_file.buf, filename.buf);\n+cleanup:\n+\tstrbuf_release(&tmp_file);\n+\tstrbuf_release(&filename);\n+\treturn err;\n+}\n+\n int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t    const char *type, struct object_id *oid,\n \t\t\t    unsigned flags)\ndiff --git a/object-store.h b/object-store.h\nindex 952efb6a4b..cc41c64d69 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -34,6 +34,12 @@ struct object_directory {\n \tchar *path;\n };\n \n+struct input_stream {\n+\tconst void *(*read)(struct input_stream *, unsigned long *len);\n+\tvoid *data;\n+\tint is_finished;\n+};\n+\n KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n \tstruct object_directory *, 1, fspathhash, fspatheq)\n \n@@ -232,6 +238,9 @@ static inline int write_object_file(const void *buf, unsigned long len,\n \treturn write_object_file_flags(buf, len, type, oid, 0);\n }\n \n+int stream_loose_object(struct input_stream *in_stream, size_t len,\n+\t\t\tstruct object_id *oid);\n+\n int hash_object_file_literally(const void *buf, unsigned long len,\n \t\t\t       const char *type, struct object_id *oid,\n \t\t\t       unsigned flags);\n-- \n2.34.1.52.gc288e771b4.agit.6.5.6\n\n"},{"id":"445788","messageId":"20220108085419.79682-7-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211217112629.12334-1-chiyutianyi@gmail.com","subject":"[PATCH v8 6/6] object-file API: add a format_object_header() function","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-01-08T08:54:19Z","receivedAt":"2022-01-08T08:56:50Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n\nAdd a convenience function to wrap the xsnprintf() command that\ngenerates loose object headers. This code was copy/pasted in various\nparts of the codebase, let's define it in one place and re-use it from\nthere.\n\nAll except one caller of it had a valid \"enum object_type\" for us,\nit's only write_object_file_prepare() which might need to deal with\n\"git hash-object --literally\" and a potential garbage type. Let's have\nthe primary API use an \"enum object_type\", and define an *_extended()\nfunction that can take an arbitrary \"const char *\" for the type.\n\nSee [1] for the discussion that prompted this patch, i.e. new code in\nobject-file.c that wanted to copy/paste the xsnprintf() invocation.\n\n1. https://lore.kernel.org/git/211213.86bl1l9bfz.gmgdl@evledraar.gmail.com/\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n builtin/index-pack.c |  3 +--\n bulk-checkin.c       |  4 ++--\n cache.h              | 21 +++++++++++++++++++++\n http-push.c          |  2 +-\n object-file.c        | 16 ++++++++++++----\n 5 files changed, 37 insertions(+), 9 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex c23d01de7d..8a6ce77940 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -449,8 +449,7 @@ static void *unpack_entry_data(off_t offset, unsigned long size,\n \tint hdrlen;\n \n \tif (!is_delta_type(type)) {\n-\t\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX,\n-\t\t\t\t   type_name(type),(uintmax_t)size) + 1;\n+\t\thdrlen = format_object_header(hdr, sizeof(hdr), type, size);\n \t\tthe_hash_algo->init_fn(&c);\n \t\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n \t} else\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 8785b2ac80..9e685f0f1a 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -220,8 +220,8 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n \tif (seekback == (off_t) -1)\n \t\treturn error(\"cannot find the current offset\");\n \n-\theader_len = xsnprintf((char *)obuf, sizeof(obuf), \"%s %\" PRIuMAX,\n-\t\t\t       type_name(type), (uintmax_t)size) + 1;\n+\theader_len = format_object_header((char *)obuf, sizeof(obuf),\n+\t\t\t\t\t type, size);\n \tthe_hash_algo->init_fn(&ctx);\n \tthe_hash_algo->update_fn(&ctx, obuf, header_len);\n \ndiff --git a/cache.h b/cache.h\nindex cfba463aa9..64071a8d80 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1310,6 +1310,27 @@ enum unpack_loose_header_result unpack_loose_header(git_zstream *stream,\n \t\t\t\t\t\t    unsigned long bufsiz,\n \t\t\t\t\t\t    struct strbuf *hdrbuf);\n \n+/**\n+ * format_object_header() is a thin wrapper around s xsnprintf() that\n+ * writes the initial \"<type> <obj-len>\" part of the loose object\n+ * header. It returns the size that snprintf() returns + 1.\n+ *\n+ * The format_object_header_extended() function allows for writing a\n+ * type_name that's not one of the \"enum object_type\" types. This is\n+ * used for \"git hash-object --literally\". Pass in a OBJ_NONE as the\n+ * type, and a non-NULL \"type_str\" to do that.\n+ *\n+ * format_object_header() is a convenience wrapper for\n+ * format_object_header_extended().\n+ */\n+int format_object_header_extended(char *str, size_t size, enum object_type type,\n+\t\t\t\t const char *type_str, size_t objsize);\n+static inline int format_object_header(char *str, size_t size,\n+\t\t\t\t      enum object_type type, size_t objsize)\n+{\n+\treturn format_object_header_extended(str, size, type, NULL, objsize);\n+}\n+\n /**\n  * parse_loose_header() parses the starting \"<type> <len>\\0\" of an\n  * object. If it doesn't follow that format -1 is returned. To check\ndiff --git a/http-push.c b/http-push.c\nindex 3309aaf004..f0c044dcf7 100644\n--- a/http-push.c\n+++ b/http-push.c\n@@ -363,7 +363,7 @@ static void start_put(struct transfer_request *request)\n \tgit_zstream stream;\n \n \tunpacked = read_object_file(&request->obj->oid, &type, &len);\n-\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n+\thdrlen = format_object_header(hdr, sizeof(hdr), type, len);\n \n \t/* Set it up */\n \tgit_deflate_init(&stream, zlib_compression_level);\ndiff --git a/object-file.c b/object-file.c\nindex a462a21629..d384ef2952 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1006,6 +1006,14 @@ void *xmmap(void *start, size_t length,\n \treturn ret;\n }\n \n+int format_object_header_extended(char *str, size_t size, enum object_type type,\n+\t\t\t\t const char *typestr, size_t objsize)\n+{\n+\tconst char *s = type == OBJ_NONE ? typestr : type_name(type);\n+\n+\treturn xsnprintf(str, size, \"%s %\"PRIuMAX, s, (uintmax_t)objsize) + 1;\n+}\n+\n /*\n  * With an in-core object data in \"map\", rehash it to make sure the\n  * object name actually matches \"oid\" to detect object corruption.\n@@ -1034,7 +1042,7 @@ int check_object_signature(struct repository *r, const struct object_id *oid,\n \t\treturn -1;\n \n \t/* Generate the header */\n-\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(obj_type), (uintmax_t)size) + 1;\n+\thdrlen = format_object_header(hdr, sizeof(hdr), obj_type, size);\n \n \t/* Sha1.. */\n \tr->hash_algo->init_fn(&c);\n@@ -1734,7 +1742,7 @@ static void write_object_file_prepare(const struct git_hash_algo *algo,\n \tgit_hash_ctx c;\n \n \t/* Generate the header */\n-\t*hdrlen = xsnprintf(hdr, *hdrlen, \"%s %\"PRIuMAX , type, (uintmax_t)len)+1;\n+\t*hdrlen = format_object_header_extended(hdr, *hdrlen, OBJ_NONE, type, len);\n \n \t/* Sha1.. */\n \talgo->init_fn(&c);\n@@ -2027,7 +2035,7 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n \n \t/* Since oid is not determined, save tmp file to odb path. */\n \tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n-\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX, type_name(OBJ_BLOB), len) + 1;\n+\thdrlen = format_object_header(hdr, sizeof(hdr), OBJ_BLOB, len);\n \n \t/* Common steps for write_loose_object and stream_loose_object to\n \t * start writing loose oject:\n@@ -2168,7 +2176,7 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n \tbuf = read_object(the_repository, oid, &type, &len);\n \tif (!buf)\n \t\treturn error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n-\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n+\thdrlen = format_object_header(hdr, sizeof(hdr), type, len);\n \tret = write_loose_object(oid, hdr, hdrlen, buf, len, mtime, 0);\n \tfree(buf);\n \n-- \n2.34.1.52.gc288e771b4.agit.6.5.6\n\n"},{"id":"445789","messageId":"20220108085419.79682-6-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20211217112629.12334-1-chiyutianyi@gmail.com","subject":"[PATCH v8 5/6] unpack-objects: unpack_non_delta_entry() read data in a stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-01-08T08:54:18Z","receivedAt":"2022-01-08T08:56:50Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWe used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\nentire contents of a blob object, no matter how big it is. This\nimplementation may consume all the memory and cause OOM.\n\nBy implementing a zstream version of input_stream interface, we can use\na small fixed buffer for \"unpack_non_delta_entry()\". However, unpack\nnon-delta objects from a stream instead of from an entrie buffer will\nhave 10% performance penalty.\n\n    $ hyperfine \\\n      --setup \\\n      'if ! test -d scalar.git; then git clone --bare\n       https://github.com/microsoft/scalar.git;\n       cp scalar.git/objects/pack/*.pack small.pack; fi' \\\n      --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n      ...\n\n    Summary\n      './git -C dest.git -c core.bigFileThreshold=512m\n      unpack-objects <small.pack' in 'origin/master'\n        1.01 ± 0.04 times faster than './git -C dest.git\n                -c core.bigFileThreshold=512m unpack-objects\n                <small.pack' in 'HEAD~1'\n        1.01 ± 0.04 times faster than './git -C dest.git\n                -c core.bigFileThreshold=512m unpack-objects\n                <small.pack' in 'HEAD~0'\n        1.03 ± 0.10 times faster than './git -C dest.git\n                -c core.bigFileThreshold=16k unpack-objects\n                <small.pack' in 'origin/master'\n        1.02 ± 0.07 times faster than './git -C dest.git\n                -c core.bigFileThreshold=16k unpack-objects\n                <small.pack' in 'HEAD~0'\n        1.10 ± 0.04 times faster than './git -C dest.git\n                -c core.bigFileThreshold=16k unpack-objects\n                <small.pack' in 'HEAD~1'\n\nTherefore, only unpack objects larger than the \"core.bigFileThreshold\"\nin zstream. Until now, the config variable has been used in the\nfollowing cases, and our new case belongs to the packfile category.\n\n * Archive:\n\n   + archive.c: write_entry(): write large blob entries to archive in\n     stream.\n\n * Loose objects:\n\n   + object-file.c: index_fd(): when hashing large files in worktree,\n     read files in a stream, and create one packfile per large blob if\n     want to save files to git object store.\n\n   + object-file.c: read_loose_object(): when checking loose objects\n     using \"git-fsck\", do not read full content of large loose objects.\n\n * Packfile:\n\n   + fast-import.c: parse_and_store_blob(): streaming large blob from\n     foreign source to packfile.\n\n   + index-pack.c: check_collison(): read and check large blob in stream.\n\n   + index-pack.c: unpack_entry_data(): do not return the entire\n     contents of the big blob from packfile, but uses a fixed buf to\n     perform some integrity checks on the object.\n\n   + pack-check.c: verify_packfile(): used by \"git-fsck\" and will call\n     check_object_signature() to check large blob in pack with the\n     streaming interface.\n\n   + pack-objects.c: get_object_details(): set \"no_try_delta\" for large\n     blobs when counting objects.\n\n   + pack-objects.c: write_no_reuse_object(): streaming large blob to\n     pack.\n\n   + unpack-objects.c: unpack_non_delta_entry(): unpack large blob in\n     stream from packfile.\n\n * Others:\n\n   + diff.c: diff_populate_filespec(): treat large blob file as binary.\n\n   + streaming.c: istream_source(): as a helper of \"open_istream()\" to\n     select proper streaming interface to read large blob from packfile.\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Derrick Stolee <stolee@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n builtin/unpack-objects.c        | 71 ++++++++++++++++++++++++++++++++-\n t/t5329-unpack-large-objects.sh | 23 +++++++++--\n 2 files changed, 90 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex c6d6c17072..e9ec2b349d 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -343,11 +343,80 @@ static void added_object(unsigned nr, enum object_type type,\n \t}\n }\n \n+struct input_zstream_data {\n+\tgit_zstream *zstream;\n+\tunsigned char buf[8192];\n+\tint status;\n+};\n+\n+static const void *feed_input_zstream(struct input_stream *in_stream,\n+\t\t\t\t      unsigned long *readlen)\n+{\n+\tstruct input_zstream_data *data = in_stream->data;\n+\tgit_zstream *zstream = data->zstream;\n+\tvoid *in = fill(1);\n+\n+\tif (in_stream->is_finished) {\n+\t\t*readlen = 0;\n+\t\treturn NULL;\n+\t}\n+\n+\tzstream->next_out = data->buf;\n+\tzstream->avail_out = sizeof(data->buf);\n+\tzstream->next_in = in;\n+\tzstream->avail_in = len;\n+\n+\tdata->status = git_inflate(zstream, 0);\n+\n+\tin_stream->is_finished = data->status != Z_OK;\n+\tuse(len - zstream->avail_in);\n+\t*readlen = sizeof(data->buf) - zstream->avail_out;\n+\n+\treturn data->buf;\n+}\n+\n+static void write_stream_blob(unsigned nr, size_t size)\n+{\n+\tgit_zstream zstream = { 0 };\n+\tstruct input_zstream_data data = { 0 };\n+\tstruct input_stream in_stream = {\n+\t\t.read = feed_input_zstream,\n+\t\t.data = &data,\n+\t};\n+\n+\tdata.zstream = &zstream;\n+\tgit_inflate_init(&zstream);\n+\n+\tif (stream_loose_object(&in_stream, size, &obj_list[nr].oid))\n+\t\tdie(_(\"failed to write object in stream\"));\n+\n+\tif (data.status != Z_STREAM_END)\n+\t\tdie(_(\"inflate returned (%d)\"), data.status);\n+\tgit_inflate_end(&zstream);\n+\n+\tif (strict) {\n+\t\tstruct blob *blob =\n+\t\t\tlookup_blob(the_repository, &obj_list[nr].oid);\n+\t\tif (blob)\n+\t\t\tblob->object.flags |= FLAG_WRITTEN;\n+\t\telse\n+\t\t\tdie(_(\"invalid blob object from stream\"));\n+\t}\n+\tobj_list[nr].obj = NULL;\n+}\n+\n static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n \t\t\t\t   unsigned nr)\n {\n-\tvoid *buf = get_data(size);\n+\tvoid *buf;\n+\n+\t/* Write large blob in stream without allocating full buffer. */\n+\tif (!dry_run && type == OBJ_BLOB && size > big_file_threshold) {\n+\t\twrite_stream_blob(nr, size);\n+\t\treturn;\n+\t}\n \n+\tbuf = get_data(size);\n \tif (buf)\n \t\twrite_object(nr, type, buf, size);\n }\ndiff --git a/t/t5329-unpack-large-objects.sh b/t/t5329-unpack-large-objects.sh\nindex 39c7a62d94..6f3bfb3df7 100755\n--- a/t/t5329-unpack-large-objects.sh\n+++ b/t/t5329-unpack-large-objects.sh\n@@ -9,7 +9,11 @@ test_description='git unpack-objects with large objects'\n \n prepare_dest () {\n \ttest_when_finished \"rm -rf dest.git\" &&\n-\tgit init --bare dest.git\n+\tgit init --bare dest.git &&\n+\tif test -n \"$1\"\n+\tthen\n+\t\tgit -C dest.git config core.bigFileThreshold $1\n+\tfi\n }\n \n assert_no_loose () {\n@@ -37,16 +41,29 @@ test_expect_success 'set memory limitation to 1MB' '\n '\n \n test_expect_success 'unpack-objects failed under memory limitation' '\n-\tprepare_dest &&\n+\tprepare_dest 2m &&\n \ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n \tgrep \"fatal: attempting to allocate\" err\n '\n \n test_expect_success 'unpack-objects works with memory limitation in dry-run mode' '\n-\tprepare_dest &&\n+\tprepare_dest 2m &&\n \tgit -C dest.git unpack-objects -n <test-$PACK.pack &&\n \tassert_no_loose &&\n \tassert_no_pack\n '\n \n+test_expect_success 'unpack big object in stream' '\n+\tprepare_dest 1m &&\n+\tgit -C dest.git unpack-objects <test-$PACK.pack &&\n+\tassert_no_pack\n+'\n+\n+test_expect_success 'do not unpack existing large objects' '\n+\tprepare_dest 1m &&\n+\tgit -C dest.git index-pack --stdin <test-$PACK.pack &&\n+\tgit -C dest.git unpack-objects <test-$PACK.pack &&\n+\tassert_no_loose\n+'\n+\n test_done\n-- \n2.34.1.52.gc288e771b4.agit.6.5.6\n\n"},{"id":"445791","messageId":"8f9dd345-56c4-9a20-151b-e0e6d1a5b3fa@web.de","threadId":"56672","inReplyTo":"20220108085419.79682-2-chiyutianyi@gmail.com","subject":"Re: [PATCH v8 1/6] unpack-objects: low memory footprint for get_data() in dry_run mode","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-01-08T12:28:31Z","receivedAt":"2022-01-08T12:29:06Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":" Am 08.01.22 um 09:54 schrieb Han Xin:\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> As the name implies, \"get_data(size)\" will allocate and return a given\n> size of memory. Allocating memory for a large blob object may cause the\n> system to run out of memory. Before preparing to replace calling of\n> \"get_data()\" to unpack large blob objects in latter commits, refactor\n> \"get_data()\" to reduce memory footprint for dry_run mode.\n>\n> Because in dry_run mode, \"get_data()\" is only used to check the\n> integrity of data, and the returned buffer is not used at all, we can\n> allocate a smaller buffer and reuse it as zstream output. Therefore,\n> in dry_run mode, \"get_data()\" will release the allocated buffer and\n> return NULL instead of returning garbage data.\n>\n> Suggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  builtin/unpack-objects.c        | 39 ++++++++++++++++++-------\n>  t/t5329-unpack-large-objects.sh | 52 +++++++++++++++++++++++++++++++++\n>  2 files changed, 80 insertions(+), 11 deletions(-)\n>  create mode 100755 t/t5329-unpack-large-objects.sh\n>\n> diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n> index 4a9466295b..c6d6c17072 100644\n> --- a/builtin/unpack-objects.c\n> +++ b/builtin/unpack-objects.c\n> @@ -96,15 +96,31 @@ static void use(int bytes)\n>  \tdisplay_throughput(progress, consumed_bytes);\n>  }\n>\n> +/*\n> + * Decompress zstream from stdin and return specific size of data.\n> + * The caller is responsible to free the returned buffer.\n> + *\n> + * But for dry_run mode, \"get_data()\" is only used to check the\n> + * integrity of data, and the returned buffer is not used at all.\n> + * Therefore, in dry_run mode, \"get_data()\" will release the small\n> + * allocated buffer which is reused to hold temporary zstream output\n> + * and return NULL instead of returning garbage data.\n> + */\n>  static void *get_data(unsigned long size)\n>  {\n>  \tgit_zstream stream;\n> -\tvoid *buf = xmallocz(size);\n> +\tunsigned long bufsize;\n> +\tvoid *buf;\n>\n>  \tmemset(&stream, 0, sizeof(stream));\n> +\tif (dry_run && size > 8192)\n> +\t\tbufsize = 8192;\n> +\telse\n> +\t\tbufsize = size;\n> +\tbuf = xmallocz(bufsize);\n>\n>  \tstream.next_out = buf;\n> -\tstream.avail_out = size;\n> +\tstream.avail_out = bufsize;\n>  \tstream.next_in = fill(1);\n>  \tstream.avail_in = len;\n>  \tgit_inflate_init(&stream);\n> @@ -124,8 +140,15 @@ static void *get_data(unsigned long size)\n>  \t\t}\n>  \t\tstream.next_in = fill(1);\n>  \t\tstream.avail_in = len;\n> +\t\tif (dry_run) {\n> +\t\t\t/* reuse the buffer in dry_run mode */\n> +\t\t\tstream.next_out = buf;\n> +\t\t\tstream.avail_out = bufsize;\n> +\t\t}\n>  \t}\n>  \tgit_inflate_end(&stream);\n> +\tif (dry_run)\n> +\t\tFREE_AND_NULL(buf);\n>  \treturn buf;\n>  }\n>\n> @@ -325,10 +348,8 @@ static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n>  {\n>  \tvoid *buf = get_data(size);\n>\n> -\tif (!dry_run && buf)\n> +\tif (buf)\n>  \t\twrite_object(nr, type, buf, size);\n> -\telse\n> -\t\tfree(buf);\n>  }\n>\n>  static int resolve_against_held(unsigned nr, const struct object_id *base,\n> @@ -358,10 +379,8 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n>  \t\toidread(&base_oid, fill(the_hash_algo->rawsz));\n>  \t\tuse(the_hash_algo->rawsz);\n>  \t\tdelta_data = get_data(delta_size);\n> -\t\tif (dry_run || !delta_data) {\n> -\t\t\tfree(delta_data);\n> +\t\tif (!delta_data)\n>  \t\t\treturn;\n> -\t\t}\n>  \t\tif (has_object_file(&base_oid))\n>  \t\t\t; /* Ok we have this one */\n>  \t\telse if (resolve_against_held(nr, &base_oid,\n> @@ -397,10 +416,8 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n>  \t\t\tdie(\"offset value out of bound for delta base object\");\n>\n>  \t\tdelta_data = get_data(delta_size);\n> -\t\tif (dry_run || !delta_data) {\n> -\t\t\tfree(delta_data);\n> +\t\tif (!delta_data)\n>  \t\t\treturn;\n> -\t\t}\n>  \t\tlo = 0;\n>  \t\thi = nr;\n>  \t\twhile (lo < hi) {\n\nNice!\n\n> diff --git a/t/t5329-unpack-large-objects.sh b/t/t5329-unpack-large-objects.sh\n> new file mode 100755\n> index 0000000000..39c7a62d94\n> --- /dev/null\n> +++ b/t/t5329-unpack-large-objects.sh\n> @@ -0,0 +1,52 @@\n> +#!/bin/sh\n> +#\n> +# Copyright (c) 2021 Han Xin\n> +#\n> +\n> +test_description='git unpack-objects with large objects'\n> +\n> +. ./test-lib.sh\n> +\n> +prepare_dest () {\n> +\ttest_when_finished \"rm -rf dest.git\" &&\n> +\tgit init --bare dest.git\n> +}\n> +\n> +assert_no_loose () {\n> +\tglob=dest.git/objects/?? &&\n> +\techo \"$glob\" >expect &&\n> +\teval \"echo $glob\" >actual &&\n> +\ttest_cmp expect actual\n> +}\n> +\n> +assert_no_pack () {\n> +\trmdir dest.git/objects/pack\n\nI would expect a function whose name starts with \"assert\" to have no\nside effects.  It doesn't matter here, because it's called only at the\nvery end, but that might change.  You can use test_dir_is_empty instead\nof rmdir.\n\n> +}\n> +\n> +test_expect_success \"create large objects (1.5 MB) and PACK\" '\n> +\ttest-tool genrandom foo 1500000 >big-blob &&\n> +\ttest_commit --append foo big-blob &&\n> +\ttest-tool genrandom bar 1500000 >big-blob &&\n> +\ttest_commit --append bar big-blob &&\n> +\tPACK=$(echo HEAD | git pack-objects --revs test)\n> +'\n> +\n> +test_expect_success 'set memory limitation to 1MB' '\n> +\tGIT_ALLOC_LIMIT=1m &&\n> +\texport GIT_ALLOC_LIMIT\n> +'\n> +\n> +test_expect_success 'unpack-objects failed under memory limitation' '\n> +\tprepare_dest &&\n> +\ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n> +\tgrep \"fatal: attempting to allocate\" err\n> +'\n> +\n> +test_expect_success 'unpack-objects works with memory limitation in dry-run mode' '\n> +\tprepare_dest &&\n> +\tgit -C dest.git unpack-objects -n <test-$PACK.pack &&\n> +\tassert_no_loose &&\n> +\tassert_no_pack\n> +'\n> +\n> +test_done\n"},{"id":"445792","messageId":"d4b89182-1b8e-3af9-ed33-e95171285ec4@web.de","threadId":"56672","inReplyTo":"20220108085419.79682-3-chiyutianyi@gmail.com","subject":"Re: [PATCH v8 2/6] object-file.c: refactor write_loose_object() to several steps","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-01-08T12:28:35Z","receivedAt":"2022-01-08T12:29:06Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 08.01.22 um 09:54 schrieb Han Xin:\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> When writing a large blob using \"write_loose_object()\", we have to pass\n> a buffer with the whole content of the blob, and this behavior will\n> consume lots of memory and may cause OOM. We will introduce a stream\n> version function (\"stream_loose_object()\") in latter commit to resolve\n> this issue.\n>\n> Before introducing a stream vesion function for writing loose object,\n> do some refactoring on \"write_loose_object()\" to reuse code for both\n> versions.\n>\n> Rewrite \"write_loose_object()\" as follows:\n>\n>  1. Figure out a path for the (temp) object file. This step is only\n>     used in \"write_loose_object()\".\n>\n>  2. Move common steps for starting to write loose objects into a new\n>     function \"start_loose_object_common()\".\n>\n>  3. Compress data.\n>\n>  4. Move common steps for ending zlib stream into a new funciton\n>     \"end_loose_object_common()\".\n>\n>  5. Close fd and finalize the object file.\n>\n> Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  object-file.c | 149 +++++++++++++++++++++++++++++++++++---------------\n>  1 file changed, 105 insertions(+), 44 deletions(-)\n>\n> diff --git a/object-file.c b/object-file.c\n> index eb1426f98c..5d163081b1 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1743,6 +1743,25 @@ static void write_object_file_prepare(const struct git_hash_algo *algo,\n>  \talgo->final_oid_fn(oid, &c);\n>  }\n>\n> +/*\n> + * Move the just written object with proper mtime into its final resting place.\n> + */\n> +static int finalize_object_file_with_mtime(const char *tmpfile,\n> +\t\t\t\t\t   const char *filename,\n> +\t\t\t\t\t   time_t mtime,\n> +\t\t\t\t\t   unsigned flags)\n\nThis function is called only once after your series.  Should it be used by\nstream_loose_object()?  Probably not -- the latter doesn't have a way to\nforce a certain modification time and its caller doesn't need one.  So\ncreating finalize_object_file_with_mtime() seems unnecessary for this\nseries.\n\n> +{\n> +\tstruct utimbuf utb;\n> +\n> +\tif (mtime) {\n> +\t\tutb.actime = mtime;\n> +\t\tutb.modtime = mtime;\n> +\t\tif (utime(tmpfile, &utb) < 0 && !(flags & HASH_SILENT))\n> +\t\t\twarning_errno(_(\"failed utime() on %s\"), tmpfile);\n> +\t}\n> +\treturn finalize_object_file(tmpfile, filename);\n> +}\n> +\n>  /*\n>   * Move the just written object into its final resting place.\n>   */\n> @@ -1828,7 +1847,8 @@ static inline int directory_size(const char *filename)\n>   * We want to avoid cross-directory filename renames, because those\n>   * can have problems on various filesystems (FAT, NFS, Coda).\n>   */\n> -static int create_tmpfile(struct strbuf *tmp, const char *filename)\n> +static int create_tmpfile(struct strbuf *tmp, const char *filename,\n> +\t\t\t  unsigned flags)\n\ncreate_tmpfile() is not mentioned in the commit message, yet it's\nchanged here.  Hrm.\n\n>  {\n>  \tint fd, dirlen = directory_size(filename);\n>\n> @@ -1836,7 +1856,9 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n>  \tstrbuf_add(tmp, filename, dirlen);\n>  \tstrbuf_addstr(tmp, \"tmp_obj_XXXXXX\");\n>  \tfd = git_mkstemp_mode(tmp->buf, 0444);\n> -\tif (fd < 0 && dirlen && errno == ENOENT) {\n> +\tdo {\n> +\t\tif (fd >= 0 || !dirlen || errno != ENOENT)\n> +\t\t\tbreak;\n\nWhy turn this branch into a loop?  Is this done to mkdir multiple\ncomponents, e.g. with filename being \"a/b/c/file\" to create \"a\", \"a/b\",\nand \"a/b/c\"?  It's only used for loose objects, so a fan-out directory\n(e.g. \".git/objects/ff\") can certainly be missing, but can their parent\nbe missing as well sometimes?  If that's the point then such a fix\nwould be worth its own patch.  (Which probably would benefit from using\nsafe_create_leading_directories()).\n\n>  \t\t/*\n>  \t\t * Make sure the directory exists; note that the contents\n>  \t\t * of the buffer are undefined after mkstemp returns an\n> @@ -1846,17 +1868,72 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n>  \t\tstrbuf_reset(tmp);\n>  \t\tstrbuf_add(tmp, filename, dirlen - 1);\n>  \t\tif (mkdir(tmp->buf, 0777) && errno != EEXIST)\n> -\t\t\treturn -1;\n> +\t\t\tbreak;\n>  \t\tif (adjust_shared_perm(tmp->buf))\n> -\t\t\treturn -1;\n> +\t\t\tbreak;\n\nOr is it just to replace these returns with a jump to the new error\nreporting section?\n\n>\n>  \t\t/* Try again */\n>  \t\tstrbuf_addstr(tmp, \"/tmp_obj_XXXXXX\");\n>  \t\tfd = git_mkstemp_mode(tmp->buf, 0444);\n\nIn that case a break would be missing here.\n\n> +\t} while (0);\n> +\n> +\tif (fd < 0 && !(flags & HASH_SILENT)) {\n> +\t\tif (errno == EACCES)\n> +\t\t\treturn error(_(\"insufficient permission for adding an \"\n> +\t\t\t\t       \"object to repository database %s\"),\n> +\t\t\t\t     get_object_directory());\n> +\t\telse\n> +\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n>  \t}\n\nWhy move this error reporting code into create_tmpfile()?  This function\nhas a single caller both before and after your series, so the code could\njust as well stay at its call-site, avoiding the need to add the flags\nparameter.\n\n> +\n>  \treturn fd;\n>  }\n>\n> +static int start_loose_object_common(struct strbuf *tmp_file,\n> +\t\t\t\t     const char *filename, unsigned flags,\n> +\t\t\t\t     git_zstream *stream,\n> +\t\t\t\t     unsigned char *buf, size_t buflen,\n> +\t\t\t\t     git_hash_ctx *c,\n> +\t\t\t\t     enum object_type type, size_t len,\n\nThe parameters type and len are not used by this function and thus can\nbe dropped.\n\n> +\t\t\t\t     char *hdr, int hdrlen)\n> +{\n> +\tint fd;\n> +\n> +\tfd = create_tmpfile(tmp_file, filename, flags);\n> +\tif (fd < 0)\n> +\t\treturn -1;\n> +\n> +\t/*  Setup zlib stream for compression */\n> +\tgit_deflate_init(stream, zlib_compression_level);\n> +\tstream->next_out = buf;\n> +\tstream->avail_out = buflen;\n> +\tthe_hash_algo->init_fn(c);\n> +\n> +\t/*  Start to feed header to zlib stream */\n> +\tstream->next_in = (unsigned char *)hdr;\n> +\tstream->avail_in = hdrlen;\n> +\twhile (git_deflate(stream, 0) == Z_OK)\n> +\t\t; /* nothing */\n> +\tthe_hash_algo->update_fn(c, hdr, hdrlen);\n> +\n> +\treturn fd;\n> +}\n> +\n> +static void end_loose_object_common(int ret, git_hash_ctx *c,\n> +\t\t\t\t    git_zstream *stream,\n> +\t\t\t\t    struct object_id *parano_oid,\n> +\t\t\t\t    const struct object_id *expected_oid,\n> +\t\t\t\t    const char *die_msg1_fmt,\n> +\t\t\t\t    const char *die_msg2_fmt)\n\nHmm, the signature needs as many lines as the function body.\n\n> +{\n> +\tif (ret != Z_STREAM_END)\n> +\t\tdie(_(die_msg1_fmt), ret, expected_oid);\n> +\tret = git_deflate_end_gently(stream);\n> +\tif (ret != Z_OK)\n> +\t\tdie(_(die_msg2_fmt), ret, expected_oid);\n\nThese format strings cannot be checked by the compiler.\n\nConsidering those two together I think I'd either unify the error\nmessages and move their strings here (losing the ability for users\nto see if streaming was used) or not extract the function and\nduplicate its few shared lines.  Just a feeling, though.\n\n> +\tthe_hash_algo->final_oid_fn(parano_oid, c);\n> +}\n> +\n>  static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \t\t\t      int hdrlen, const void *buf, unsigned long len,\n>  \t\t\t      time_t mtime, unsigned flags)\n> @@ -1871,28 +1948,18 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>\n>  \tloose_object_path(the_repository, &filename, oid);\n>\n> -\tfd = create_tmpfile(&tmp_file, filename.buf);\n> -\tif (fd < 0) {\n> -\t\tif (flags & HASH_SILENT)\n> -\t\t\treturn -1;\n> -\t\telse if (errno == EACCES)\n> -\t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n> -\t\telse\n> -\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n> -\t}\n> -\n> -\t/* Set it up */\n> -\tgit_deflate_init(&stream, zlib_compression_level);\n> -\tstream.next_out = compressed;\n> -\tstream.avail_out = sizeof(compressed);\n> -\tthe_hash_algo->init_fn(&c);\n> -\n> -\t/* First header.. */\n> -\tstream.next_in = (unsigned char *)hdr;\n> -\tstream.avail_in = hdrlen;\n> -\twhile (git_deflate(&stream, 0) == Z_OK)\n> -\t\t; /* nothing */\n> -\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n> +\t/* Common steps for write_loose_object and stream_loose_object to\n> +\t * start writing loose oject:\n> +\t *\n> +\t *  - Create tmpfile for the loose object.\n> +\t *  - Setup zlib stream for compression.\n> +\t *  - Start to feed header to zlib stream.\n> +\t */\n> +\tfd = start_loose_object_common(&tmp_file, filename.buf, flags,\n> +\t\t\t\t       &stream, compressed, sizeof(compressed),\n> +\t\t\t\t       &c, OBJ_NONE, 0, hdr, hdrlen);\n> +\tif (fd < 0)\n> +\t\treturn -1;\n>\n>  \t/* Then the data itself.. */\n>  \tstream.next_in = (void *)buf;\n> @@ -1907,30 +1974,24 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \t\tstream.avail_out = sizeof(compressed);\n>  \t} while (ret == Z_OK);\n>\n> -\tif (ret != Z_STREAM_END)\n> -\t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n> -\t\t    ret);\n> -\tret = git_deflate_end_gently(&stream);\n> -\tif (ret != Z_OK)\n> -\t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n> -\t\t    ret);\n> -\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n> +\t/* Common steps for write_loose_object and stream_loose_object to\n> +\t * end writing loose oject:\n> +\t *\n> +\t *  - End the compression of zlib stream.\n> +\t *  - Get the calculated oid to \"parano_oid\".\n> +\t */\n> +\tend_loose_object_common(ret, &c, &stream, &parano_oid, oid,\n> +\t\t\t\tN_(\"unable to deflate new object %s (%d)\"),\n> +\t\t\t\tN_(\"deflateEnd on object %s failed (%d)\"));\n> +\n>  \tif (!oideq(oid, &parano_oid))\n>  \t\tdie(_(\"confused by unstable object source data for %s\"),\n>  \t\t    oid_to_hex(oid));\n>\n>  \tclose_loose_object(fd);\n>\n> -\tif (mtime) {\n> -\t\tstruct utimbuf utb;\n> -\t\tutb.actime = mtime;\n> -\t\tutb.modtime = mtime;\n> -\t\tif (utime(tmp_file.buf, &utb) < 0 &&\n> -\t\t    !(flags & HASH_SILENT))\n> -\t\t\twarning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n> -\t}\n> -\n> -\treturn finalize_object_file(tmp_file.buf, filename.buf);\n> +\treturn finalize_object_file_with_mtime(tmp_file.buf, filename.buf,\n> +\t\t\t\t\t       mtime, flags);\n>  }\n>\n>  static int freshen_loose_object(const struct object_id *oid)\n"},{"id":"445795","messageId":"6d63d5d2-48db-40e9-8e5c-5b72c3d84414@web.de","threadId":"56672","inReplyTo":"20220108085419.79682-4-chiyutianyi@gmail.com","subject":"Re: [PATCH v8 3/6] object-file.c: remove the slash for directory_size()","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-01-08T17:24:04Z","receivedAt":"2022-01-08T17:24:26Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 08.01.22 um 09:54 schrieb Han Xin:\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> Since \"mkdir foo/\" works as well as \"mkdir foo\", let's remove the end\n> slash as many users of it want.\n>\n> Suggested-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> ---\n>  object-file.c | 8 ++++----\n>  1 file changed, 4 insertions(+), 4 deletions(-)\n>\n> diff --git a/object-file.c b/object-file.c\n> index 5d163081b1..4f0127e823 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1831,13 +1831,13 @@ static void close_loose_object(int fd)\n>  \t\tdie_errno(_(\"error when closing loose object file\"));\n>  }\n>\n> -/* Size of directory component, including the ending '/' */\n> +/* Size of directory component, excluding the ending '/' */\n>  static inline int directory_size(const char *filename)\n>  {\n>  \tconst char *s = strrchr(filename, '/');\n>  \tif (!s)\n>  \t\treturn 0;\n> -\treturn s - filename + 1;\n> +\treturn s - filename;\n\nThis will return zero both for \"filename\" and \"/filename\".  Hmm.  Since\nit's only used for loose object files we can assume that at least one\nslash is present, so this removal of functionality is not actually a\nproblem.  But I don't understand its benefit.\n\n>  }\n>\n>  /*\n> @@ -1854,7 +1854,7 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename,\n>\n>  \tstrbuf_reset(tmp);\n>  \tstrbuf_add(tmp, filename, dirlen);\n> -\tstrbuf_addstr(tmp, \"tmp_obj_XXXXXX\");\n> +\tstrbuf_addstr(tmp, \"/tmp_obj_XXXXXX\");\n>  \tfd = git_mkstemp_mode(tmp->buf, 0444);\n>  \tdo {\n>  \t\tif (fd >= 0 || !dirlen || errno != ENOENT)\n> @@ -1866,7 +1866,7 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename,\n>  \t\t * scratch.\n>  \t\t */\n>  \t\tstrbuf_reset(tmp);\n> -\t\tstrbuf_add(tmp, filename, dirlen - 1);\n> +\t\tstrbuf_add(tmp, filename, dirlen);\n>  \t\tif (mkdir(tmp->buf, 0777) && errno != EEXIST)\n\nThis code makes sure that mkdir(2) is called without the trailing slash,\nboth with or without this patch.  From the commit message above I\nsomehow expected a change in this regard -- but again I wouldn't\nunderstand its benefit.\n\nIs this change really needed?  Is streaming unpack not possible with the\noriginal directory_size() function?\n\n>  \t\t\tbreak;\n>  \t\tif (adjust_shared_perm(tmp->buf))\n"},{"id":"445907","messageId":"CAO0brD1qS_E7votGcJ26pSzadauoxA-5pX9eH8WTP6G0Yhr=nw@mail.gmail.com","threadId":"56672","inReplyTo":"6d63d5d2-48db-40e9-8e5c-5b72c3d84414@web.de","subject":"Re: [PATCH v8 3/6] object-file.c: remove the slash for directory_size()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-01-11T10:14:44Z","receivedAt":"2022-01-11T10:14:59Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Sun, Jan 9, 2022 at 1:24 AM René Scharfe <l.s.r@web.de> wrote:\n>\n> Am 08.01.22 um 09:54 schrieb Han Xin:\n> > From: Han Xin <hanxin.hx@alibaba-inc.com>\n> >\n> > Since \"mkdir foo/\" works as well as \"mkdir foo\", let's remove the end\n> > slash as many users of it want.\n> >\n> > Suggested-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> > Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> > ---\n> >  object-file.c | 8 ++++----\n> >  1 file changed, 4 insertions(+), 4 deletions(-)\n> >\n> > diff --git a/object-file.c b/object-file.c\n> > index 5d163081b1..4f0127e823 100644\n> > --- a/object-file.c\n> > +++ b/object-file.c\n> > @@ -1831,13 +1831,13 @@ static void close_loose_object(int fd)\n> >               die_errno(_(\"error when closing loose object file\"));\n> >  }\n> >\n> > -/* Size of directory component, including the ending '/' */\n> > +/* Size of directory component, excluding the ending '/' */\n> >  static inline int directory_size(const char *filename)\n> >  {\n> >       const char *s = strrchr(filename, '/');\n> >       if (!s)\n> >               return 0;\n> > -     return s - filename + 1;\n> > +     return s - filename;\n>\n> This will return zero both for \"filename\" and \"/filename\".  Hmm.  Since\n> it's only used for loose object files we can assume that at least one\n> slash is present, so this removal of functionality is not actually a\n> problem.  But I don't understand its benefit.\n>\n> >  }\n> >\n> >  /*\n> > @@ -1854,7 +1854,7 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename,\n> >\n> >       strbuf_reset(tmp);\n> >       strbuf_add(tmp, filename, dirlen);\n> > -     strbuf_addstr(tmp, \"tmp_obj_XXXXXX\");\n> > +     strbuf_addstr(tmp, \"/tmp_obj_XXXXXX\");\n> >       fd = git_mkstemp_mode(tmp->buf, 0444);\n> >       do {\n> >               if (fd >= 0 || !dirlen || errno != ENOENT)\n> > @@ -1866,7 +1866,7 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename,\n> >                * scratch.\n> >                */\n> >               strbuf_reset(tmp);\n> > -             strbuf_add(tmp, filename, dirlen - 1);\n> > +             strbuf_add(tmp, filename, dirlen);\n> >               if (mkdir(tmp->buf, 0777) && errno != EEXIST)\n>\n> This code makes sure that mkdir(2) is called without the trailing slash,\n> both with or without this patch.  From the commit message above I\n> somehow expected a change in this regard -- but again I wouldn't\n> understand its benefit.\n>\n> Is this change really needed?  Is streaming unpack not possible with the\n> original directory_size() function?\n>\n\n*nod*\nStreaming unpacking still works with the original directory_size().\n\nThis patch is more of a code cleanup that reduces the extra handling of\ndirectory size first increasing and then decreasing. I'll seriously consider\nif I should remove this patch, or move it to the tail.\n\nThanks\n-Han Xin\n\n> >                       break;\n> >               if (adjust_shared_perm(tmp->buf))\n"},{"id":"445908","messageId":"CAO0brD3drqKfTV=oRTNHncR2tg9nQnr_zycV+X4MccRagBYDSw@mail.gmail.com","threadId":"56672","inReplyTo":"d4b89182-1b8e-3af9-ed33-e95171285ec4@web.de","subject":"Re: [PATCH v8 2/6] object-file.c: refactor write_loose_object() to several steps","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-01-11T10:33:50Z","receivedAt":"2022-01-11T10:34:10Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Sat, Jan 8, 2022 at 8:28 PM René Scharfe <l.s.r@web.de> wrote:\n>\n> Am 08.01.22 um 09:54 schrieb Han Xin:\n> > From: Han Xin <hanxin.hx@alibaba-inc.com>\n> >\n> > When writing a large blob using \"write_loose_object()\", we have to pass\n> > a buffer with the whole content of the blob, and this behavior will\n> > consume lots of memory and may cause OOM. We will introduce a stream\n> > version function (\"stream_loose_object()\") in latter commit to resolve\n> > this issue.\n> >\n> > Before introducing a stream vesion function for writing loose object,\n> > do some refactoring on \"write_loose_object()\" to reuse code for both\n> > versions.\n> >\n> > Rewrite \"write_loose_object()\" as follows:\n> >\n> >  1. Figure out a path for the (temp) object file. This step is only\n> >     used in \"write_loose_object()\".\n> >\n> >  2. Move common steps for starting to write loose objects into a new\n> >     function \"start_loose_object_common()\".\n> >\n> >  3. Compress data.\n> >\n> >  4. Move common steps for ending zlib stream into a new funciton\n> >     \"end_loose_object_common()\".\n> >\n> >  5. Close fd and finalize the object file.\n> >\n> > Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> > Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> > Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> > ---\n> >  object-file.c | 149 +++++++++++++++++++++++++++++++++++---------------\n> >  1 file changed, 105 insertions(+), 44 deletions(-)\n> >\n> > diff --git a/object-file.c b/object-file.c\n> > index eb1426f98c..5d163081b1 100644\n> > --- a/object-file.c\n> > +++ b/object-file.c\n> > @@ -1743,6 +1743,25 @@ static void write_object_file_prepare(const struct git_hash_algo *algo,\n> >       algo->final_oid_fn(oid, &c);\n> >  }\n> >\n> > +/*\n> > + * Move the just written object with proper mtime into its final resting place.\n> > + */\n> > +static int finalize_object_file_with_mtime(const char *tmpfile,\n> > +                                        const char *filename,\n> > +                                        time_t mtime,\n> > +                                        unsigned flags)\n>\n> This function is called only once after your series.  Should it be used by\n> stream_loose_object()?  Probably not -- the latter doesn't have a way to\n> force a certain modification time and its caller doesn't need one.  So\n> creating finalize_object_file_with_mtime() seems unnecessary for this\n> series.\n>\n\nAfter accepting the suggestion by Ævar Arnfjörð Bjarmason[1] to remove\nfinalize_object_file_with_mtime() from stream_loose_object() , it seems to\nbe an overkill for write_loose_object() now. I'll put it back into\nwrite_loose_object() .\n\n1. https://lore.kernel.org/git/211221.86pmpqq9aj.gmgdl@evledraar.gmail.com/\n\nThanks\n-Han Xin\n\n> > +{\n> > +     struct utimbuf utb;\n> > +\n> > +     if (mtime) {\n> > +             utb.actime = mtime;\n> > +             utb.modtime = mtime;\n> > +             if (utime(tmpfile, &utb) < 0 && !(flags & HASH_SILENT))\n> > +                     warning_errno(_(\"failed utime() on %s\"), tmpfile);\n> > +     }\n> > +     return finalize_object_file(tmpfile, filename);\n> > +}\n> > +\n> >  /*\n> >   * Move the just written object into its final resting place.\n> >   */\n> > @@ -1828,7 +1847,8 @@ static inline int directory_size(const char *filename)\n> >   * We want to avoid cross-directory filename renames, because those\n> >   * can have problems on various filesystems (FAT, NFS, Coda).\n> >   */\n> > -static int create_tmpfile(struct strbuf *tmp, const char *filename)\n> > +static int create_tmpfile(struct strbuf *tmp, const char *filename,\n> > +                       unsigned flags)\n>\n> create_tmpfile() is not mentioned in the commit message, yet it's\n> changed here.  Hrm.\n>\n> >  {\n> >       int fd, dirlen = directory_size(filename);\n> >\n> > @@ -1836,7 +1856,9 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n> >       strbuf_add(tmp, filename, dirlen);\n> >       strbuf_addstr(tmp, \"tmp_obj_XXXXXX\");\n> >       fd = git_mkstemp_mode(tmp->buf, 0444);\n> > -     if (fd < 0 && dirlen && errno == ENOENT) {\n> > +     do {\n> > +             if (fd >= 0 || !dirlen || errno != ENOENT)\n> > +                     break;\n>\n> Why turn this branch into a loop?  Is this done to mkdir multiple\n> components, e.g. with filename being \"a/b/c/file\" to create \"a\", \"a/b\",\n> and \"a/b/c\"?  It's only used for loose objects, so a fan-out directory\n> (e.g. \".git/objects/ff\") can certainly be missing, but can their parent\n> be missing as well sometimes?  If that's the point then such a fix\n> would be worth its own patch.  (Which probably would benefit from using\n> safe_create_leading_directories()).\n>\n> >               /*\n> >                * Make sure the directory exists; note that the contents\n> >                * of the buffer are undefined after mkstemp returns an\n> > @@ -1846,17 +1868,72 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n> >               strbuf_reset(tmp);\n> >               strbuf_add(tmp, filename, dirlen - 1);\n> >               if (mkdir(tmp->buf, 0777) && errno != EEXIST)\n> > -                     return -1;\n> > +                     break;\n> >               if (adjust_shared_perm(tmp->buf))\n> > -                     return -1;\n> > +                     break;\n>\n> Or is it just to replace these returns with a jump to the new error\n> reporting section?\n>\n> >\n> >               /* Try again */\n> >               strbuf_addstr(tmp, \"/tmp_obj_XXXXXX\");\n> >               fd = git_mkstemp_mode(tmp->buf, 0444);\n>\n> In that case a break would be missing here.\n>\n> > +     } while (0);\n> > +\n> > +     if (fd < 0 && !(flags & HASH_SILENT)) {\n> > +             if (errno == EACCES)\n> > +                     return error(_(\"insufficient permission for adding an \"\n> > +                                    \"object to repository database %s\"),\n> > +                                  get_object_directory());\n> > +             else\n> > +                     return error_errno(_(\"unable to create temporary file\"));\n> >       }\n>\n> Why move this error reporting code into create_tmpfile()?  This function\n> has a single caller both before and after your series, so the code could\n> just as well stay at its call-site, avoiding the need to add the flags\n> parameter.\n>\n\nHere is a legacy from v7, now there is no step called \"Figuring out a path\nfor the (temp) object file.\", and it's only used in start_loose_object_common().\nI will bring it back to what it was.\n\nThanks\n-Han Xin\n> > +\n> >       return fd;\n> >  }\n> >\n> > +static int start_loose_object_common(struct strbuf *tmp_file,\n> > +                                  const char *filename, unsigned flags,\n> > +                                  git_zstream *stream,\n> > +                                  unsigned char *buf, size_t buflen,\n> > +                                  git_hash_ctx *c,\n> > +                                  enum object_type type, size_t len,\n>\n> The parameters type and len are not used by this function and thus can\n> be dropped.\n>\n\n*nod*\n\n> > +                                  char *hdr, int hdrlen)\n> > +{\n> > +     int fd;\n> > +\n> > +     fd = create_tmpfile(tmp_file, filename, flags);\n> > +     if (fd < 0)\n> > +             return -1;\n> > +\n> > +     /*  Setup zlib stream for compression */\n> > +     git_deflate_init(stream, zlib_compression_level);\n> > +     stream->next_out = buf;\n> > +     stream->avail_out = buflen;\n> > +     the_hash_algo->init_fn(c);\n> > +\n> > +     /*  Start to feed header to zlib stream */\n> > +     stream->next_in = (unsigned char *)hdr;\n> > +     stream->avail_in = hdrlen;\n> > +     while (git_deflate(stream, 0) == Z_OK)\n> > +             ; /* nothing */\n> > +     the_hash_algo->update_fn(c, hdr, hdrlen);\n> > +\n> > +     return fd;\n> > +}\n> > +\n> > +static void end_loose_object_common(int ret, git_hash_ctx *c,\n> > +                                 git_zstream *stream,\n> > +                                 struct object_id *parano_oid,\n> > +                                 const struct object_id *expected_oid,\n> > +                                 const char *die_msg1_fmt,\n> > +                                 const char *die_msg2_fmt)\n>\n> Hmm, the signature needs as many lines as the function body.\n>\n> > +{\n> > +     if (ret != Z_STREAM_END)\n> > +             die(_(die_msg1_fmt), ret, expected_oid);\n> > +     ret = git_deflate_end_gently(stream);\n> > +     if (ret != Z_OK)\n> > +             die(_(die_msg2_fmt), ret, expected_oid);\n>\n> These format strings cannot be checked by the compiler.\n>\n> Considering those two together I think I'd either unify the error\n> messages and move their strings here (losing the ability for users\n> to see if streaming was used) or not extract the function and\n> duplicate its few shared lines.  Just a feeling, though.\n>\n> > +     the_hash_algo->final_oid_fn(parano_oid, c);\n> > +}\n> > +\n> >  static int write_loose_object(const struct object_id *oid, char *hdr,\n> >                             int hdrlen, const void *buf, unsigned long len,\n> >                             time_t mtime, unsigned flags)\n> > @@ -1871,28 +1948,18 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n> >\n> >       loose_object_path(the_repository, &filename, oid);\n> >\n> > -     fd = create_tmpfile(&tmp_file, filename.buf);\n> > -     if (fd < 0) {\n> > -             if (flags & HASH_SILENT)\n> > -                     return -1;\n> > -             else if (errno == EACCES)\n> > -                     return error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n> > -             else\n> > -                     return error_errno(_(\"unable to create temporary file\"));\n> > -     }\n> > -\n> > -     /* Set it up */\n> > -     git_deflate_init(&stream, zlib_compression_level);\n> > -     stream.next_out = compressed;\n> > -     stream.avail_out = sizeof(compressed);\n> > -     the_hash_algo->init_fn(&c);\n> > -\n> > -     /* First header.. */\n> > -     stream.next_in = (unsigned char *)hdr;\n> > -     stream.avail_in = hdrlen;\n> > -     while (git_deflate(&stream, 0) == Z_OK)\n> > -             ; /* nothing */\n> > -     the_hash_algo->update_fn(&c, hdr, hdrlen);\n> > +     /* Common steps for write_loose_object and stream_loose_object to\n> > +      * start writing loose oject:\n> > +      *\n> > +      *  - Create tmpfile for the loose object.\n> > +      *  - Setup zlib stream for compression.\n> > +      *  - Start to feed header to zlib stream.\n> > +      */\n> > +     fd = start_loose_object_common(&tmp_file, filename.buf, flags,\n> > +                                    &stream, compressed, sizeof(compressed),\n> > +                                    &c, OBJ_NONE, 0, hdr, hdrlen);\n> > +     if (fd < 0)\n> > +             return -1;\n> >\n> >       /* Then the data itself.. */\n> >       stream.next_in = (void *)buf;\n> > @@ -1907,30 +1974,24 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n> >               stream.avail_out = sizeof(compressed);\n> >       } while (ret == Z_OK);\n> >\n> > -     if (ret != Z_STREAM_END)\n> > -             die(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n> > -                 ret);\n> > -     ret = git_deflate_end_gently(&stream);\n> > -     if (ret != Z_OK)\n> > -             die(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n> > -                 ret);\n> > -     the_hash_algo->final_oid_fn(&parano_oid, &c);\n> > +     /* Common steps for write_loose_object and stream_loose_object to\n> > +      * end writing loose oject:\n> > +      *\n> > +      *  - End the compression of zlib stream.\n> > +      *  - Get the calculated oid to \"parano_oid\".\n> > +      */\n> > +     end_loose_object_common(ret, &c, &stream, &parano_oid, oid,\n> > +                             N_(\"unable to deflate new object %s (%d)\"),\n> > +                             N_(\"deflateEnd on object %s failed (%d)\"));\n> > +\n> >       if (!oideq(oid, &parano_oid))\n> >               die(_(\"confused by unstable object source data for %s\"),\n> >                   oid_to_hex(oid));\n> >\n> >       close_loose_object(fd);\n> >\n> > -     if (mtime) {\n> > -             struct utimbuf utb;\n> > -             utb.actime = mtime;\n> > -             utb.modtime = mtime;\n> > -             if (utime(tmp_file.buf, &utb) < 0 &&\n> > -                 !(flags & HASH_SILENT))\n> > -                     warning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n> > -     }\n> > -\n> > -     return finalize_object_file(tmp_file.buf, filename.buf);\n> > +     return finalize_object_file_with_mtime(tmp_file.buf, filename.buf,\n> > +                                            mtime, flags);\n> >  }\n> >\n> >  static int freshen_loose_object(const struct object_id *oid)\n"},{"id":"445910","messageId":"CAO0brD17MC4THrGVNq70ey+vP-9-W28kZD4y8Fn1mVqyEbEbKA@mail.gmail.com","threadId":"56672","inReplyTo":"8f9dd345-56c4-9a20-151b-e0e6d1a5b3fa@web.de","subject":"Re: [PATCH v8 1/6] unpack-objects: low memory footprint for get_data() in dry_run mode","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-01-11T10:41:11Z","receivedAt":"2022-01-11T10:41:29Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Sat, Jan 8, 2022 at 8:28 PM René Scharfe <l.s.r@web.de> wrote:\n>\n>  Am 08.01.22 um 09:54 schrieb Han Xin:\n> > From: Han Xin <hanxin.hx@alibaba-inc.com>\n> >\n> > +assert_no_loose () {\n> > +     glob=dest.git/objects/?? &&\n> > +     echo \"$glob\" >expect &&\n> > +     eval \"echo $glob\" >actual &&\n> > +     test_cmp expect actual\n> > +}\n> > +\n> > +assert_no_pack () {\n> > +     rmdir dest.git/objects/pack\n>\n> I would expect a function whose name starts with \"assert\" to have no\n> side effects.  It doesn't matter here, because it's called only at the\n> very end, but that might change.  You can use test_dir_is_empty instead\n> of rmdir.\n>\n\n*nod*\nI think it would be better to rename \"assert_no_loose()\" to \"test_no_loose()\".\nI will remove \"assert_no_pack()\" and use \"test_dir_is_empty()\" instead.\n\nThanks\n-Han Xin\n"},{"id":"446548","messageId":"20220120112114.47618-2-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20220108085419.79682-1-chiyutianyi@gmail.com","subject":"[PATCH v9 1/5] unpack-objects: low memory footprint for get_data() in dry_run mode","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-01-20T11:21:10Z","receivedAt":"2022-01-20T11:22:49Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nAs the name implies, \"get_data(size)\" will allocate and return a given\nsize of memory. Allocating memory for a large blob object may cause the\nsystem to run out of memory. Before preparing to replace calling of\n\"get_data()\" to unpack large blob objects in latter commits, refactor\n\"get_data()\" to reduce memory footprint for dry_run mode.\n\nBecause in dry_run mode, \"get_data()\" is only used to check the\nintegrity of data, and the returned buffer is not used at all, we can\nallocate a smaller buffer and reuse it as zstream output. Therefore,\nin dry_run mode, \"get_data()\" will release the allocated buffer and\nreturn NULL instead of returning garbage data.\n\nSuggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n builtin/unpack-objects.c        | 39 +++++++++++++++++++--------\n t/t5328-unpack-large-objects.sh | 48 +++++++++++++++++++++++++++++++++\n 2 files changed, 76 insertions(+), 11 deletions(-)\n create mode 100755 t/t5328-unpack-large-objects.sh\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 4a9466295b..c6d6c17072 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -96,15 +96,31 @@ static void use(int bytes)\n \tdisplay_throughput(progress, consumed_bytes);\n }\n \n+/*\n+ * Decompress zstream from stdin and return specific size of data.\n+ * The caller is responsible to free the returned buffer.\n+ *\n+ * But for dry_run mode, \"get_data()\" is only used to check the\n+ * integrity of data, and the returned buffer is not used at all.\n+ * Therefore, in dry_run mode, \"get_data()\" will release the small\n+ * allocated buffer which is reused to hold temporary zstream output\n+ * and return NULL instead of returning garbage data.\n+ */\n static void *get_data(unsigned long size)\n {\n \tgit_zstream stream;\n-\tvoid *buf = xmallocz(size);\n+\tunsigned long bufsize;\n+\tvoid *buf;\n \n \tmemset(&stream, 0, sizeof(stream));\n+\tif (dry_run && size > 8192)\n+\t\tbufsize = 8192;\n+\telse\n+\t\tbufsize = size;\n+\tbuf = xmallocz(bufsize);\n \n \tstream.next_out = buf;\n-\tstream.avail_out = size;\n+\tstream.avail_out = bufsize;\n \tstream.next_in = fill(1);\n \tstream.avail_in = len;\n \tgit_inflate_init(&stream);\n@@ -124,8 +140,15 @@ static void *get_data(unsigned long size)\n \t\t}\n \t\tstream.next_in = fill(1);\n \t\tstream.avail_in = len;\n+\t\tif (dry_run) {\n+\t\t\t/* reuse the buffer in dry_run mode */\n+\t\t\tstream.next_out = buf;\n+\t\t\tstream.avail_out = bufsize;\n+\t\t}\n \t}\n \tgit_inflate_end(&stream);\n+\tif (dry_run)\n+\t\tFREE_AND_NULL(buf);\n \treturn buf;\n }\n \n@@ -325,10 +348,8 @@ static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n {\n \tvoid *buf = get_data(size);\n \n-\tif (!dry_run && buf)\n+\tif (buf)\n \t\twrite_object(nr, type, buf, size);\n-\telse\n-\t\tfree(buf);\n }\n \n static int resolve_against_held(unsigned nr, const struct object_id *base,\n@@ -358,10 +379,8 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\toidread(&base_oid, fill(the_hash_algo->rawsz));\n \t\tuse(the_hash_algo->rawsz);\n \t\tdelta_data = get_data(delta_size);\n-\t\tif (dry_run || !delta_data) {\n-\t\t\tfree(delta_data);\n+\t\tif (!delta_data)\n \t\t\treturn;\n-\t\t}\n \t\tif (has_object_file(&base_oid))\n \t\t\t; /* Ok we have this one */\n \t\telse if (resolve_against_held(nr, &base_oid,\n@@ -397,10 +416,8 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\t\tdie(\"offset value out of bound for delta base object\");\n \n \t\tdelta_data = get_data(delta_size);\n-\t\tif (dry_run || !delta_data) {\n-\t\t\tfree(delta_data);\n+\t\tif (!delta_data)\n \t\t\treturn;\n-\t\t}\n \t\tlo = 0;\n \t\thi = nr;\n \t\twhile (lo < hi) {\ndiff --git a/t/t5328-unpack-large-objects.sh b/t/t5328-unpack-large-objects.sh\nnew file mode 100755\nindex 0000000000..45a3316e06\n--- /dev/null\n+++ b/t/t5328-unpack-large-objects.sh\n@@ -0,0 +1,48 @@\n+#!/bin/sh\n+#\n+# Copyright (c) 2022 Han Xin\n+#\n+\n+test_description='git unpack-objects with large objects'\n+\n+. ./test-lib.sh\n+\n+prepare_dest () {\n+\ttest_when_finished \"rm -rf dest.git\" &&\n+\tgit init --bare dest.git\n+}\n+\n+test_no_loose () {\n+\tglob=dest.git/objects/?? &&\n+\techo \"$glob\" >expect &&\n+\teval \"echo $glob\" >actual &&\n+\ttest_cmp expect actual\n+}\n+\n+test_expect_success \"create large objects (1.5 MB) and PACK\" '\n+\ttest-tool genrandom foo 1500000 >big-blob &&\n+\ttest_commit --append foo big-blob &&\n+\ttest-tool genrandom bar 1500000 >big-blob &&\n+\ttest_commit --append bar big-blob &&\n+\tPACK=$(echo HEAD | git pack-objects --revs test)\n+'\n+\n+test_expect_success 'set memory limitation to 1MB' '\n+\tGIT_ALLOC_LIMIT=1m &&\n+\texport GIT_ALLOC_LIMIT\n+'\n+\n+test_expect_success 'unpack-objects failed under memory limitation' '\n+\tprepare_dest &&\n+\ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n+\tgrep \"fatal: attempting to allocate\" err\n+'\n+\n+test_expect_success 'unpack-objects works with memory limitation in dry-run mode' '\n+\tprepare_dest &&\n+\tgit -C dest.git unpack-objects -n <test-$PACK.pack &&\n+\ttest_no_loose &&\n+\ttest_dir_is_empty dest.git/objects/pack\n+'\n+\n+test_done\n-- \n2.34.1.52.gc288e771b4.agit.6.5.6\n\n"},{"id":"446549","messageId":"20220120112114.47618-1-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20220108085419.79682-1-chiyutianyi@gmail.com","subject":"[PATCH v9 0/5] unpack large blobs in stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-01-20T11:21:09Z","receivedAt":"2022-01-20T11:22:50Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nChanges since v8:\n* Rename \"assert_no_loose ()\" into \"test_no_loose ()\" in\n  \"t5329-unpack-large-objects.sh\". Remove \"assert_no_pack ()\" and use\n  \"test_dir_is_empty\" instead.\n\n* Revert changes to \"create_tmpfile()\" and error handling is now in\n  \"start_loose_object_common()\".\n\n* Remove \"finalize_object_file_with_mtime()\" which seems to be an overkill\n  for \"write_loose_object()\" now. \n\n* Remove the commit \"object-file.c: remove the slash for directory_size()\",\n  it can be in a separate patch if necessary.\n\nHan Xin (4):\n  unpack-objects: low memory footprint for get_data() in dry_run mode\n  object-file.c: refactor write_loose_object() to several steps\n  object-file.c: add \"stream_loose_object()\" to handle large object\n  unpack-objects: unpack_non_delta_entry() read data in a stream\n\nÆvar Arnfjörð Bjarmason (1):\n  object-file API: add a format_object_header() function\n\n builtin/index-pack.c            |   3 +-\n builtin/unpack-objects.c        | 110 ++++++++++++++--\n bulk-checkin.c                  |   4 +-\n cache.h                         |  21 +++\n http-push.c                     |   2 +-\n object-file.c                   | 220 +++++++++++++++++++++++++++-----\n object-store.h                  |   9 ++\n t/t5328-unpack-large-objects.sh |  65 ++++++++++\n 8 files changed, 384 insertions(+), 50 deletions(-)\n create mode 100755 t/t5328-unpack-large-objects.sh\n\nRange-diff against v8:\n1:  bd34da5816 ! 1:  6a6c11ba93 unpack-objects: low memory footprint for get_data() in dry_run mode\n    @@ builtin/unpack-objects.c: static void unpack_delta_entry(enum object_type type,\n      \t\thi = nr;\n      \t\twhile (lo < hi) {\n     \n    - ## t/t5329-unpack-large-objects.sh (new) ##\n    + ## t/t5328-unpack-large-objects.sh (new) ##\n     @@\n     +#!/bin/sh\n     +#\n    -+# Copyright (c) 2021 Han Xin\n    ++# Copyright (c) 2022 Han Xin\n     +#\n     +\n     +test_description='git unpack-objects with large objects'\n    @@ t/t5329-unpack-large-objects.sh (new)\n     +\tgit init --bare dest.git\n     +}\n     +\n    -+assert_no_loose () {\n    ++test_no_loose () {\n     +\tglob=dest.git/objects/?? &&\n     +\techo \"$glob\" >expect &&\n     +\teval \"echo $glob\" >actual &&\n     +\ttest_cmp expect actual\n     +}\n     +\n    -+assert_no_pack () {\n    -+\trmdir dest.git/objects/pack\n    -+}\n    -+\n     +test_expect_success \"create large objects (1.5 MB) and PACK\" '\n     +\ttest-tool genrandom foo 1500000 >big-blob &&\n     +\ttest_commit --append foo big-blob &&\n    @@ t/t5329-unpack-large-objects.sh (new)\n     +test_expect_success 'unpack-objects works with memory limitation in dry-run mode' '\n     +\tprepare_dest &&\n     +\tgit -C dest.git unpack-objects -n <test-$PACK.pack &&\n    -+\tassert_no_loose &&\n    -+\tassert_no_pack\n    ++\ttest_no_loose &&\n    ++\ttest_dir_is_empty dest.git/objects/pack\n     +'\n     +\n     +test_done\n2:  f9a4365a7d ! 2:  bab9e0402f object-file.c: refactor write_loose_object() to several steps\n    @@ Commit message\n         Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n     \n      ## object-file.c ##\n    -@@ object-file.c: static void write_object_file_prepare(const struct git_hash_algo *algo,\n    - \talgo->final_oid_fn(oid, &c);\n    - }\n    - \n    -+/*\n    -+ * Move the just written object with proper mtime into its final resting place.\n    -+ */\n    -+static int finalize_object_file_with_mtime(const char *tmpfile,\n    -+\t\t\t\t\t   const char *filename,\n    -+\t\t\t\t\t   time_t mtime,\n    -+\t\t\t\t\t   unsigned flags)\n    -+{\n    -+\tstruct utimbuf utb;\n    -+\n    -+\tif (mtime) {\n    -+\t\tutb.actime = mtime;\n    -+\t\tutb.modtime = mtime;\n    -+\t\tif (utime(tmpfile, &utb) < 0 && !(flags & HASH_SILENT))\n    -+\t\t\twarning_errno(_(\"failed utime() on %s\"), tmpfile);\n    -+\t}\n    -+\treturn finalize_object_file(tmpfile, filename);\n    -+}\n    -+\n    - /*\n    -  * Move the just written object into its final resting place.\n    -  */\n    -@@ object-file.c: static inline int directory_size(const char *filename)\n    -  * We want to avoid cross-directory filename renames, because those\n    -  * can have problems on various filesystems (FAT, NFS, Coda).\n    -  */\n    --static int create_tmpfile(struct strbuf *tmp, const char *filename)\n    -+static int create_tmpfile(struct strbuf *tmp, const char *filename,\n    -+\t\t\t  unsigned flags)\n    - {\n    - \tint fd, dirlen = directory_size(filename);\n    - \n    -@@ object-file.c: static int create_tmpfile(struct strbuf *tmp, const char *filename)\n    - \tstrbuf_add(tmp, filename, dirlen);\n    - \tstrbuf_addstr(tmp, \"tmp_obj_XXXXXX\");\n    - \tfd = git_mkstemp_mode(tmp->buf, 0444);\n    --\tif (fd < 0 && dirlen && errno == ENOENT) {\n    -+\tdo {\n    -+\t\tif (fd >= 0 || !dirlen || errno != ENOENT)\n    -+\t\t\tbreak;\n    - \t\t/*\n    - \t\t * Make sure the directory exists; note that the contents\n    - \t\t * of the buffer are undefined after mkstemp returns an\n     @@ object-file.c: static int create_tmpfile(struct strbuf *tmp, const char *filename)\n    - \t\tstrbuf_reset(tmp);\n    - \t\tstrbuf_add(tmp, filename, dirlen - 1);\n    - \t\tif (mkdir(tmp->buf, 0777) && errno != EEXIST)\n    --\t\t\treturn -1;\n    -+\t\t\tbreak;\n    - \t\tif (adjust_shared_perm(tmp->buf))\n    --\t\t\treturn -1;\n    -+\t\t\tbreak;\n    - \n    - \t\t/* Try again */\n    - \t\tstrbuf_addstr(tmp, \"/tmp_obj_XXXXXX\");\n    - \t\tfd = git_mkstemp_mode(tmp->buf, 0444);\n    -+\t} while (0);\n    -+\n    -+\tif (fd < 0 && !(flags & HASH_SILENT)) {\n    -+\t\tif (errno == EACCES)\n    -+\t\t\treturn error(_(\"insufficient permission for adding an \"\n    -+\t\t\t\t       \"object to repository database %s\"),\n    -+\t\t\t\t     get_object_directory());\n    -+\t\telse\n    -+\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n    - \t}\n    -+\n      \treturn fd;\n      }\n      \n    @@ object-file.c: static int create_tmpfile(struct strbuf *tmp, const char *filenam\n     +\t\t\t\t     git_zstream *stream,\n     +\t\t\t\t     unsigned char *buf, size_t buflen,\n     +\t\t\t\t     git_hash_ctx *c,\n    -+\t\t\t\t     enum object_type type, size_t len,\n     +\t\t\t\t     char *hdr, int hdrlen)\n     +{\n     +\tint fd;\n     +\n    -+\tfd = create_tmpfile(tmp_file, filename, flags);\n    -+\tif (fd < 0)\n    -+\t\treturn -1;\n    ++\tfd = create_tmpfile(tmp_file, filename);\n    ++\tif (fd < 0) {\n    ++\t\tif (flags & HASH_SILENT)\n    ++\t\t\treturn -1;\n    ++\t\telse if (errno == EACCES)\n    ++\t\t\treturn error(_(\"insufficient permission for adding \"\n    ++\t\t\t\t       \"an object to repository database %s\"),\n    ++\t\t\t\t     get_object_directory());\n    ++\t\telse\n    ++\t\t\treturn error_errno(\n    ++\t\t\t\t_(\"unable to create temporary file\"));\n    ++\t}\n     +\n     +\t/*  Setup zlib stream for compression */\n     +\tgit_deflate_init(stream, zlib_compression_level);\n    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *\n     +\t */\n     +\tfd = start_loose_object_common(&tmp_file, filename.buf, flags,\n     +\t\t\t\t       &stream, compressed, sizeof(compressed),\n    -+\t\t\t\t       &c, OBJ_NONE, 0, hdr, hdrlen);\n    ++\t\t\t\t       &c, hdr, hdrlen);\n     +\tif (fd < 0)\n     +\t\treturn -1;\n      \n    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *\n      \tif (!oideq(oid, &parano_oid))\n      \t\tdie(_(\"confused by unstable object source data for %s\"),\n      \t\t    oid_to_hex(oid));\n    - \n    - \tclose_loose_object(fd);\n    - \n    --\tif (mtime) {\n    --\t\tstruct utimbuf utb;\n    --\t\tutb.actime = mtime;\n    --\t\tutb.modtime = mtime;\n    --\t\tif (utime(tmp_file.buf, &utb) < 0 &&\n    --\t\t    !(flags & HASH_SILENT))\n    --\t\t\twarning_errno(_(\"failed utime() on %s\"), tmp_file.buf);\n    --\t}\n    --\n    --\treturn finalize_object_file(tmp_file.buf, filename.buf);\n    -+\treturn finalize_object_file_with_mtime(tmp_file.buf, filename.buf,\n    -+\t\t\t\t\t       mtime, flags);\n    - }\n    - \n    - static int freshen_loose_object(const struct object_id *oid)\n3:  18dd21122d < -:  ---------- object-file.c: remove the slash for directory_size()\n4:  964715451b ! 3:  dd13614985 object-file.c: add \"stream_loose_object()\" to handle large object\n    @@ object-file.c: static int freshen_packed_object(const struct object_id *oid)\n     +\t */\n     +\tfd = start_loose_object_common(&tmp_file, filename.buf, 0,\n     +\t\t\t\t       &stream, compressed, sizeof(compressed),\n    -+\t\t\t\t       &c, OBJ_BLOB, len, hdr, hdrlen);\n    ++\t\t\t\t       &c, hdr, hdrlen);\n     +\tif (fd < 0) {\n     +\t\terr = -1;\n     +\t\tgoto cleanup;\n5:  3f620466fe ! 4:  cd84e27b08 unpack-objects: unpack_non_delta_entry() read data in a stream\n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n      \t\twrite_object(nr, type, buf, size);\n      }\n     \n    - ## t/t5329-unpack-large-objects.sh ##\n    -@@ t/t5329-unpack-large-objects.sh: test_description='git unpack-objects with large objects'\n    + ## t/t5328-unpack-large-objects.sh ##\n    +@@ t/t5328-unpack-large-objects.sh: test_description='git unpack-objects with large objects'\n      \n      prepare_dest () {\n      \ttest_when_finished \"rm -rf dest.git\" &&\n    @@ t/t5329-unpack-large-objects.sh: test_description='git unpack-objects with large\n     +\tfi\n      }\n      \n    - assert_no_loose () {\n    -@@ t/t5329-unpack-large-objects.sh: test_expect_success 'set memory limitation to 1MB' '\n    + test_no_loose () {\n    +@@ t/t5328-unpack-large-objects.sh: test_expect_success 'set memory limitation to 1MB' '\n      '\n      \n      test_expect_success 'unpack-objects failed under memory limitation' '\n    @@ t/t5329-unpack-large-objects.sh: test_expect_success 'set memory limitation to 1\n     -\tprepare_dest &&\n     +\tprepare_dest 2m &&\n      \tgit -C dest.git unpack-objects -n <test-$PACK.pack &&\n    - \tassert_no_loose &&\n    - \tassert_no_pack\n    + \ttest_no_loose &&\n    + \ttest_dir_is_empty dest.git/objects/pack\n      '\n      \n     +test_expect_success 'unpack big object in stream' '\n     +\tprepare_dest 1m &&\n     +\tgit -C dest.git unpack-objects <test-$PACK.pack &&\n    -+\tassert_no_pack\n    ++\ttest_dir_is_empty dest.git/objects/pack\n     +'\n     +\n     +test_expect_success 'do not unpack existing large objects' '\n     +\tprepare_dest 1m &&\n     +\tgit -C dest.git index-pack --stdin <test-$PACK.pack &&\n     +\tgit -C dest.git unpack-objects <test-$PACK.pack &&\n    -+\tassert_no_loose\n    ++\ttest_no_loose\n     +'\n     +\n      test_done\n6:  8073a3888d = 5:  59f0ad95c7 object-file API: add a format_object_header() function\n-- \n2.34.1.52.gc288e771b4.agit.6.5.6\n\n"},{"id":"446550","messageId":"20220120112114.47618-3-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20220108085419.79682-1-chiyutianyi@gmail.com","subject":"[PATCH v9 2/5] object-file.c: refactor write_loose_object() to several steps","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-01-20T11:21:11Z","receivedAt":"2022-01-20T11:22:52Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWhen writing a large blob using \"write_loose_object()\", we have to pass\na buffer with the whole content of the blob, and this behavior will\nconsume lots of memory and may cause OOM. We will introduce a stream\nversion function (\"stream_loose_object()\") in latter commit to resolve\nthis issue.\n\nBefore introducing a stream vesion function for writing loose object,\ndo some refactoring on \"write_loose_object()\" to reuse code for both\nversions.\n\nRewrite \"write_loose_object()\" as follows:\n\n 1. Figure out a path for the (temp) object file. This step is only\n    used in \"write_loose_object()\".\n\n 2. Move common steps for starting to write loose objects into a new\n    function \"start_loose_object_common()\".\n\n 3. Compress data.\n\n 4. Move common steps for ending zlib stream into a new funciton\n    \"end_loose_object_common()\".\n\n 5. Close fd and finalize the object file.\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c | 105 +++++++++++++++++++++++++++++++++++---------------\n 1 file changed, 75 insertions(+), 30 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex eb1426f98c..422b43212a 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1857,6 +1857,59 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n \treturn fd;\n }\n \n+static int start_loose_object_common(struct strbuf *tmp_file,\n+\t\t\t\t     const char *filename, unsigned flags,\n+\t\t\t\t     git_zstream *stream,\n+\t\t\t\t     unsigned char *buf, size_t buflen,\n+\t\t\t\t     git_hash_ctx *c,\n+\t\t\t\t     char *hdr, int hdrlen)\n+{\n+\tint fd;\n+\n+\tfd = create_tmpfile(tmp_file, filename);\n+\tif (fd < 0) {\n+\t\tif (flags & HASH_SILENT)\n+\t\t\treturn -1;\n+\t\telse if (errno == EACCES)\n+\t\t\treturn error(_(\"insufficient permission for adding \"\n+\t\t\t\t       \"an object to repository database %s\"),\n+\t\t\t\t     get_object_directory());\n+\t\telse\n+\t\t\treturn error_errno(\n+\t\t\t\t_(\"unable to create temporary file\"));\n+\t}\n+\n+\t/*  Setup zlib stream for compression */\n+\tgit_deflate_init(stream, zlib_compression_level);\n+\tstream->next_out = buf;\n+\tstream->avail_out = buflen;\n+\tthe_hash_algo->init_fn(c);\n+\n+\t/*  Start to feed header to zlib stream */\n+\tstream->next_in = (unsigned char *)hdr;\n+\tstream->avail_in = hdrlen;\n+\twhile (git_deflate(stream, 0) == Z_OK)\n+\t\t; /* nothing */\n+\tthe_hash_algo->update_fn(c, hdr, hdrlen);\n+\n+\treturn fd;\n+}\n+\n+static void end_loose_object_common(int ret, git_hash_ctx *c,\n+\t\t\t\t    git_zstream *stream,\n+\t\t\t\t    struct object_id *parano_oid,\n+\t\t\t\t    const struct object_id *expected_oid,\n+\t\t\t\t    const char *die_msg1_fmt,\n+\t\t\t\t    const char *die_msg2_fmt)\n+{\n+\tif (ret != Z_STREAM_END)\n+\t\tdie(_(die_msg1_fmt), ret, expected_oid);\n+\tret = git_deflate_end_gently(stream);\n+\tif (ret != Z_OK)\n+\t\tdie(_(die_msg2_fmt), ret, expected_oid);\n+\tthe_hash_algo->final_oid_fn(parano_oid, c);\n+}\n+\n static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\t\t      int hdrlen, const void *buf, unsigned long len,\n \t\t\t      time_t mtime, unsigned flags)\n@@ -1871,28 +1924,18 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \n \tloose_object_path(the_repository, &filename, oid);\n \n-\tfd = create_tmpfile(&tmp_file, filename.buf);\n-\tif (fd < 0) {\n-\t\tif (flags & HASH_SILENT)\n-\t\t\treturn -1;\n-\t\telse if (errno == EACCES)\n-\t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n-\t\telse\n-\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n-\t}\n-\n-\t/* Set it up */\n-\tgit_deflate_init(&stream, zlib_compression_level);\n-\tstream.next_out = compressed;\n-\tstream.avail_out = sizeof(compressed);\n-\tthe_hash_algo->init_fn(&c);\n-\n-\t/* First header.. */\n-\tstream.next_in = (unsigned char *)hdr;\n-\tstream.avail_in = hdrlen;\n-\twhile (git_deflate(&stream, 0) == Z_OK)\n-\t\t; /* nothing */\n-\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n+\t/* Common steps for write_loose_object and stream_loose_object to\n+\t * start writing loose oject:\n+\t *\n+\t *  - Create tmpfile for the loose object.\n+\t *  - Setup zlib stream for compression.\n+\t *  - Start to feed header to zlib stream.\n+\t */\n+\tfd = start_loose_object_common(&tmp_file, filename.buf, flags,\n+\t\t\t\t       &stream, compressed, sizeof(compressed),\n+\t\t\t\t       &c, hdr, hdrlen);\n+\tif (fd < 0)\n+\t\treturn -1;\n \n \t/* Then the data itself.. */\n \tstream.next_in = (void *)buf;\n@@ -1907,14 +1950,16 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tstream.avail_out = sizeof(compressed);\n \t} while (ret == Z_OK);\n \n-\tif (ret != Z_STREAM_END)\n-\t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n-\t\t    ret);\n-\tret = git_deflate_end_gently(&stream);\n-\tif (ret != Z_OK)\n-\t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n-\t\t    ret);\n-\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n+\t/* Common steps for write_loose_object and stream_loose_object to\n+\t * end writing loose oject:\n+\t *\n+\t *  - End the compression of zlib stream.\n+\t *  - Get the calculated oid to \"parano_oid\".\n+\t */\n+\tend_loose_object_common(ret, &c, &stream, &parano_oid, oid,\n+\t\t\t\tN_(\"unable to deflate new object %s (%d)\"),\n+\t\t\t\tN_(\"deflateEnd on object %s failed (%d)\"));\n+\n \tif (!oideq(oid, &parano_oid))\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n-- \n2.34.1.52.gc288e771b4.agit.6.5.6\n\n"},{"id":"446551","messageId":"20220120112114.47618-4-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20220108085419.79682-1-chiyutianyi@gmail.com","subject":"[PATCH v9 3/5] object-file.c: add \"stream_loose_object()\" to handle large object","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-01-20T11:21:12Z","receivedAt":"2022-01-20T11:22:54Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nIf we want unpack and write a loose object using \"write_loose_object\",\nwe have to feed it with a buffer with the same size of the object, which\nwill consume lots of memory and may cause OOM. This can be improved by\nfeeding data to \"stream_loose_object()\" in a stream.\n\nAdd a new function \"stream_loose_object()\", which is a stream version of\n\"write_loose_object()\" but with a low memory footprint. We will use this\nfunction to unpack large blob object in latter commit.\n\nAnother difference with \"write_loose_object()\" is that we have no chance\nto run \"write_object_file_prepare()\" to calculate the oid in advance.\nIn \"write_loose_object()\", we know the oid and we can write the\ntemporary file in the same directory as the final object, but for an\nobject with an undetermined oid, we don't know the exact directory for\nthe object, so we have to save the temporary file in \".git/objects/\"\ndirectory instead.\n\n\"freshen_packed_object()\" or \"freshen_loose_object()\" will be called\ninside \"stream_loose_object()\" after obtaining the \"oid\".\n\nHelped-by: René Scharfe <l.s.r@web.de>\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c  | 101 +++++++++++++++++++++++++++++++++++++++++++++++++\n object-store.h |   9 +++++\n 2 files changed, 110 insertions(+)\n\ndiff --git a/object-file.c b/object-file.c\nindex 422b43212a..a738f47cb2 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1996,6 +1996,107 @@ static int freshen_packed_object(const struct object_id *oid)\n \treturn 1;\n }\n \n+int stream_loose_object(struct input_stream *in_stream, size_t len,\n+\t\t\tstruct object_id *oid)\n+{\n+\tint fd, ret, err = 0, flush = 0;\n+\tunsigned char compressed[4096];\n+\tgit_zstream stream;\n+\tgit_hash_ctx c;\n+\tstruct strbuf tmp_file = STRBUF_INIT;\n+\tstruct strbuf filename = STRBUF_INIT;\n+\tint dirlen;\n+\tchar hdr[MAX_HEADER_LEN];\n+\tint hdrlen;\n+\n+\t/* Since oid is not determined, save tmp file to odb path. */\n+\tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n+\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX, type_name(OBJ_BLOB), len) + 1;\n+\n+\t/* Common steps for write_loose_object and stream_loose_object to\n+\t * start writing loose oject:\n+\t *\n+\t *  - Create tmpfile for the loose object.\n+\t *  - Setup zlib stream for compression.\n+\t *  - Start to feed header to zlib stream.\n+\t */\n+\tfd = start_loose_object_common(&tmp_file, filename.buf, 0,\n+\t\t\t\t       &stream, compressed, sizeof(compressed),\n+\t\t\t\t       &c, hdr, hdrlen);\n+\tif (fd < 0) {\n+\t\terr = -1;\n+\t\tgoto cleanup;\n+\t}\n+\n+\t/* Then the data itself.. */\n+\tdo {\n+\t\tunsigned char *in0 = stream.next_in;\n+\t\tif (!stream.avail_in && !in_stream->is_finished) {\n+\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n+\t\t\tstream.next_in = (void *)in;\n+\t\t\tin0 = (unsigned char *)in;\n+\t\t\t/* All data has been read. */\n+\t\t\tif (in_stream->is_finished)\n+\t\t\t\tflush = Z_FINISH;\n+\t\t}\n+\t\tret = git_deflate(&stream, flush);\n+\t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n+\t\tif (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n+\t\t\tdie(_(\"unable to write loose object file\"));\n+\t\tstream.next_out = compressed;\n+\t\tstream.avail_out = sizeof(compressed);\n+\t\t/*\n+\t\t * Unlike write_loose_object(), we do not have the entire\n+\t\t * buffer. If we get Z_BUF_ERROR due to too few input bytes,\n+\t\t * then we'll replenish them in the next input_stream->read()\n+\t\t * call when we loop.\n+\t\t */\n+\t} while (ret == Z_OK || ret == Z_BUF_ERROR);\n+\n+\tif (stream.total_in != len + hdrlen)\n+\t\tdie(_(\"write stream object %ld != %\"PRIuMAX), stream.total_in,\n+\t\t    (uintmax_t)len + hdrlen);\n+\n+\t/* Common steps for write_loose_object and stream_loose_object to\n+\t * end writing loose oject:\n+\t *\n+\t *  - End the compression of zlib stream.\n+\t *  - Get the calculated oid.\n+\t */\n+\tend_loose_object_common(ret, &c, &stream, oid, NULL,\n+\t\t\t\tN_(\"unable to stream deflate new object (%d)\"),\n+\t\t\t\tN_(\"deflateEnd on stream object failed (%d)\"));\n+\n+\tclose_loose_object(fd);\n+\n+\tif (freshen_packed_object(oid) || freshen_loose_object(oid)) {\n+\t\tunlink_or_warn(tmp_file.buf);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tloose_object_path(the_repository, &filename, oid);\n+\n+\t/* We finally know the object path, and create the missing dir. */\n+\tdirlen = directory_size(filename.buf);\n+\tif (dirlen) {\n+\t\tstruct strbuf dir = STRBUF_INIT;\n+\t\tstrbuf_add(&dir, filename.buf, dirlen);\n+\n+\t\tif (mkdir_in_gitdir(dir.buf) && errno != EEXIST) {\n+\t\t\terr = error_errno(_(\"unable to create directory %s\"), dir.buf);\n+\t\t\tstrbuf_release(&dir);\n+\t\t\tgoto cleanup;\n+\t\t}\n+\t\tstrbuf_release(&dir);\n+\t}\n+\n+\terr = finalize_object_file(tmp_file.buf, filename.buf);\n+cleanup:\n+\tstrbuf_release(&tmp_file);\n+\tstrbuf_release(&filename);\n+\treturn err;\n+}\n+\n int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t    const char *type, struct object_id *oid,\n \t\t\t    unsigned flags)\ndiff --git a/object-store.h b/object-store.h\nindex 952efb6a4b..cc41c64d69 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -34,6 +34,12 @@ struct object_directory {\n \tchar *path;\n };\n \n+struct input_stream {\n+\tconst void *(*read)(struct input_stream *, unsigned long *len);\n+\tvoid *data;\n+\tint is_finished;\n+};\n+\n KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n \tstruct object_directory *, 1, fspathhash, fspatheq)\n \n@@ -232,6 +238,9 @@ static inline int write_object_file(const void *buf, unsigned long len,\n \treturn write_object_file_flags(buf, len, type, oid, 0);\n }\n \n+int stream_loose_object(struct input_stream *in_stream, size_t len,\n+\t\t\tstruct object_id *oid);\n+\n int hash_object_file_literally(const void *buf, unsigned long len,\n \t\t\t       const char *type, struct object_id *oid,\n \t\t\t       unsigned flags);\n-- \n2.34.1.52.gc288e771b4.agit.6.5.6\n\n"},{"id":"446552","messageId":"20220120112114.47618-5-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20220108085419.79682-1-chiyutianyi@gmail.com","subject":"[PATCH v9 4/5] unpack-objects: unpack_non_delta_entry() read data in a stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-01-20T11:21:13Z","receivedAt":"2022-01-20T11:22:57Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWe used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\nentire contents of a blob object, no matter how big it is. This\nimplementation may consume all the memory and cause OOM.\n\nBy implementing a zstream version of input_stream interface, we can use\na small fixed buffer for \"unpack_non_delta_entry()\". However, unpack\nnon-delta objects from a stream instead of from an entrie buffer will\nhave 10% performance penalty.\n\n    $ hyperfine \\\n      --setup \\\n      'if ! test -d scalar.git; then git clone --bare\n       https://github.com/microsoft/scalar.git;\n       cp scalar.git/objects/pack/*.pack small.pack; fi' \\\n      --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n      ...\n\n    Summary\n      './git -C dest.git -c core.bigFileThreshold=512m\n      unpack-objects <small.pack' in 'origin/master'\n        1.01 ± 0.04 times faster than './git -C dest.git\n                -c core.bigFileThreshold=512m unpack-objects\n                <small.pack' in 'HEAD~1'\n        1.01 ± 0.04 times faster than './git -C dest.git\n                -c core.bigFileThreshold=512m unpack-objects\n                <small.pack' in 'HEAD~0'\n        1.03 ± 0.10 times faster than './git -C dest.git\n                -c core.bigFileThreshold=16k unpack-objects\n                <small.pack' in 'origin/master'\n        1.02 ± 0.07 times faster than './git -C dest.git\n                -c core.bigFileThreshold=16k unpack-objects\n                <small.pack' in 'HEAD~0'\n        1.10 ± 0.04 times faster than './git -C dest.git\n                -c core.bigFileThreshold=16k unpack-objects\n                <small.pack' in 'HEAD~1'\n\nTherefore, only unpack objects larger than the \"core.bigFileThreshold\"\nin zstream. Until now, the config variable has been used in the\nfollowing cases, and our new case belongs to the packfile category.\n\n * Archive:\n\n   + archive.c: write_entry(): write large blob entries to archive in\n     stream.\n\n * Loose objects:\n\n   + object-file.c: index_fd(): when hashing large files in worktree,\n     read files in a stream, and create one packfile per large blob if\n     want to save files to git object store.\n\n   + object-file.c: read_loose_object(): when checking loose objects\n     using \"git-fsck\", do not read full content of large loose objects.\n\n * Packfile:\n\n   + fast-import.c: parse_and_store_blob(): streaming large blob from\n     foreign source to packfile.\n\n   + index-pack.c: check_collison(): read and check large blob in stream.\n\n   + index-pack.c: unpack_entry_data(): do not return the entire\n     contents of the big blob from packfile, but uses a fixed buf to\n     perform some integrity checks on the object.\n\n   + pack-check.c: verify_packfile(): used by \"git-fsck\" and will call\n     check_object_signature() to check large blob in pack with the\n     streaming interface.\n\n   + pack-objects.c: get_object_details(): set \"no_try_delta\" for large\n     blobs when counting objects.\n\n   + pack-objects.c: write_no_reuse_object(): streaming large blob to\n     pack.\n\n   + unpack-objects.c: unpack_non_delta_entry(): unpack large blob in\n     stream from packfile.\n\n * Others:\n\n   + diff.c: diff_populate_filespec(): treat large blob file as binary.\n\n   + streaming.c: istream_source(): as a helper of \"open_istream()\" to\n     select proper streaming interface to read large blob from packfile.\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Derrick Stolee <stolee@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n builtin/unpack-objects.c        | 71 ++++++++++++++++++++++++++++++++-\n t/t5328-unpack-large-objects.sh | 23 +++++++++--\n 2 files changed, 90 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex c6d6c17072..e9ec2b349d 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -343,11 +343,80 @@ static void added_object(unsigned nr, enum object_type type,\n \t}\n }\n \n+struct input_zstream_data {\n+\tgit_zstream *zstream;\n+\tunsigned char buf[8192];\n+\tint status;\n+};\n+\n+static const void *feed_input_zstream(struct input_stream *in_stream,\n+\t\t\t\t      unsigned long *readlen)\n+{\n+\tstruct input_zstream_data *data = in_stream->data;\n+\tgit_zstream *zstream = data->zstream;\n+\tvoid *in = fill(1);\n+\n+\tif (in_stream->is_finished) {\n+\t\t*readlen = 0;\n+\t\treturn NULL;\n+\t}\n+\n+\tzstream->next_out = data->buf;\n+\tzstream->avail_out = sizeof(data->buf);\n+\tzstream->next_in = in;\n+\tzstream->avail_in = len;\n+\n+\tdata->status = git_inflate(zstream, 0);\n+\n+\tin_stream->is_finished = data->status != Z_OK;\n+\tuse(len - zstream->avail_in);\n+\t*readlen = sizeof(data->buf) - zstream->avail_out;\n+\n+\treturn data->buf;\n+}\n+\n+static void write_stream_blob(unsigned nr, size_t size)\n+{\n+\tgit_zstream zstream = { 0 };\n+\tstruct input_zstream_data data = { 0 };\n+\tstruct input_stream in_stream = {\n+\t\t.read = feed_input_zstream,\n+\t\t.data = &data,\n+\t};\n+\n+\tdata.zstream = &zstream;\n+\tgit_inflate_init(&zstream);\n+\n+\tif (stream_loose_object(&in_stream, size, &obj_list[nr].oid))\n+\t\tdie(_(\"failed to write object in stream\"));\n+\n+\tif (data.status != Z_STREAM_END)\n+\t\tdie(_(\"inflate returned (%d)\"), data.status);\n+\tgit_inflate_end(&zstream);\n+\n+\tif (strict) {\n+\t\tstruct blob *blob =\n+\t\t\tlookup_blob(the_repository, &obj_list[nr].oid);\n+\t\tif (blob)\n+\t\t\tblob->object.flags |= FLAG_WRITTEN;\n+\t\telse\n+\t\t\tdie(_(\"invalid blob object from stream\"));\n+\t}\n+\tobj_list[nr].obj = NULL;\n+}\n+\n static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n \t\t\t\t   unsigned nr)\n {\n-\tvoid *buf = get_data(size);\n+\tvoid *buf;\n+\n+\t/* Write large blob in stream without allocating full buffer. */\n+\tif (!dry_run && type == OBJ_BLOB && size > big_file_threshold) {\n+\t\twrite_stream_blob(nr, size);\n+\t\treturn;\n+\t}\n \n+\tbuf = get_data(size);\n \tif (buf)\n \t\twrite_object(nr, type, buf, size);\n }\ndiff --git a/t/t5328-unpack-large-objects.sh b/t/t5328-unpack-large-objects.sh\nindex 45a3316e06..f4129979f9 100755\n--- a/t/t5328-unpack-large-objects.sh\n+++ b/t/t5328-unpack-large-objects.sh\n@@ -9,7 +9,11 @@ test_description='git unpack-objects with large objects'\n \n prepare_dest () {\n \ttest_when_finished \"rm -rf dest.git\" &&\n-\tgit init --bare dest.git\n+\tgit init --bare dest.git &&\n+\tif test -n \"$1\"\n+\tthen\n+\t\tgit -C dest.git config core.bigFileThreshold $1\n+\tfi\n }\n \n test_no_loose () {\n@@ -33,16 +37,29 @@ test_expect_success 'set memory limitation to 1MB' '\n '\n \n test_expect_success 'unpack-objects failed under memory limitation' '\n-\tprepare_dest &&\n+\tprepare_dest 2m &&\n \ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n \tgrep \"fatal: attempting to allocate\" err\n '\n \n test_expect_success 'unpack-objects works with memory limitation in dry-run mode' '\n-\tprepare_dest &&\n+\tprepare_dest 2m &&\n \tgit -C dest.git unpack-objects -n <test-$PACK.pack &&\n \ttest_no_loose &&\n \ttest_dir_is_empty dest.git/objects/pack\n '\n \n+test_expect_success 'unpack big object in stream' '\n+\tprepare_dest 1m &&\n+\tgit -C dest.git unpack-objects <test-$PACK.pack &&\n+\ttest_dir_is_empty dest.git/objects/pack\n+'\n+\n+test_expect_success 'do not unpack existing large objects' '\n+\tprepare_dest 1m &&\n+\tgit -C dest.git index-pack --stdin <test-$PACK.pack &&\n+\tgit -C dest.git unpack-objects <test-$PACK.pack &&\n+\ttest_no_loose\n+'\n+\n test_done\n-- \n2.34.1.52.gc288e771b4.agit.6.5.6\n\n"},{"id":"446553","messageId":"20220120112114.47618-6-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"20220108085419.79682-1-chiyutianyi@gmail.com","subject":"[PATCH v9 5/5] object-file API: add a format_object_header() function","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-01-20T11:21:14Z","receivedAt":"2022-01-20T11:22:58Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n\nAdd a convenience function to wrap the xsnprintf() command that\ngenerates loose object headers. This code was copy/pasted in various\nparts of the codebase, let's define it in one place and re-use it from\nthere.\n\nAll except one caller of it had a valid \"enum object_type\" for us,\nit's only write_object_file_prepare() which might need to deal with\n\"git hash-object --literally\" and a potential garbage type. Let's have\nthe primary API use an \"enum object_type\", and define an *_extended()\nfunction that can take an arbitrary \"const char *\" for the type.\n\nSee [1] for the discussion that prompted this patch, i.e. new code in\nobject-file.c that wanted to copy/paste the xsnprintf() invocation.\n\n1. https://lore.kernel.org/git/211213.86bl1l9bfz.gmgdl@evledraar.gmail.com/\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n builtin/index-pack.c |  3 +--\n bulk-checkin.c       |  4 ++--\n cache.h              | 21 +++++++++++++++++++++\n http-push.c          |  2 +-\n object-file.c        | 16 ++++++++++++----\n 5 files changed, 37 insertions(+), 9 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex c23d01de7d..8a6ce77940 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -449,8 +449,7 @@ static void *unpack_entry_data(off_t offset, unsigned long size,\n \tint hdrlen;\n \n \tif (!is_delta_type(type)) {\n-\t\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX,\n-\t\t\t\t   type_name(type),(uintmax_t)size) + 1;\n+\t\thdrlen = format_object_header(hdr, sizeof(hdr), type, size);\n \t\tthe_hash_algo->init_fn(&c);\n \t\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n \t} else\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 8785b2ac80..9e685f0f1a 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -220,8 +220,8 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n \tif (seekback == (off_t) -1)\n \t\treturn error(\"cannot find the current offset\");\n \n-\theader_len = xsnprintf((char *)obuf, sizeof(obuf), \"%s %\" PRIuMAX,\n-\t\t\t       type_name(type), (uintmax_t)size) + 1;\n+\theader_len = format_object_header((char *)obuf, sizeof(obuf),\n+\t\t\t\t\t type, size);\n \tthe_hash_algo->init_fn(&ctx);\n \tthe_hash_algo->update_fn(&ctx, obuf, header_len);\n \ndiff --git a/cache.h b/cache.h\nindex cfba463aa9..64071a8d80 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1310,6 +1310,27 @@ enum unpack_loose_header_result unpack_loose_header(git_zstream *stream,\n \t\t\t\t\t\t    unsigned long bufsiz,\n \t\t\t\t\t\t    struct strbuf *hdrbuf);\n \n+/**\n+ * format_object_header() is a thin wrapper around s xsnprintf() that\n+ * writes the initial \"<type> <obj-len>\" part of the loose object\n+ * header. It returns the size that snprintf() returns + 1.\n+ *\n+ * The format_object_header_extended() function allows for writing a\n+ * type_name that's not one of the \"enum object_type\" types. This is\n+ * used for \"git hash-object --literally\". Pass in a OBJ_NONE as the\n+ * type, and a non-NULL \"type_str\" to do that.\n+ *\n+ * format_object_header() is a convenience wrapper for\n+ * format_object_header_extended().\n+ */\n+int format_object_header_extended(char *str, size_t size, enum object_type type,\n+\t\t\t\t const char *type_str, size_t objsize);\n+static inline int format_object_header(char *str, size_t size,\n+\t\t\t\t      enum object_type type, size_t objsize)\n+{\n+\treturn format_object_header_extended(str, size, type, NULL, objsize);\n+}\n+\n /**\n  * parse_loose_header() parses the starting \"<type> <len>\\0\" of an\n  * object. If it doesn't follow that format -1 is returned. To check\ndiff --git a/http-push.c b/http-push.c\nindex 3309aaf004..f0c044dcf7 100644\n--- a/http-push.c\n+++ b/http-push.c\n@@ -363,7 +363,7 @@ static void start_put(struct transfer_request *request)\n \tgit_zstream stream;\n \n \tunpacked = read_object_file(&request->obj->oid, &type, &len);\n-\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n+\thdrlen = format_object_header(hdr, sizeof(hdr), type, len);\n \n \t/* Set it up */\n \tgit_deflate_init(&stream, zlib_compression_level);\ndiff --git a/object-file.c b/object-file.c\nindex a738f47cb2..0dce5d2fec 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1006,6 +1006,14 @@ void *xmmap(void *start, size_t length,\n \treturn ret;\n }\n \n+int format_object_header_extended(char *str, size_t size, enum object_type type,\n+\t\t\t\t const char *typestr, size_t objsize)\n+{\n+\tconst char *s = type == OBJ_NONE ? typestr : type_name(type);\n+\n+\treturn xsnprintf(str, size, \"%s %\"PRIuMAX, s, (uintmax_t)objsize) + 1;\n+}\n+\n /*\n  * With an in-core object data in \"map\", rehash it to make sure the\n  * object name actually matches \"oid\" to detect object corruption.\n@@ -1034,7 +1042,7 @@ int check_object_signature(struct repository *r, const struct object_id *oid,\n \t\treturn -1;\n \n \t/* Generate the header */\n-\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(obj_type), (uintmax_t)size) + 1;\n+\thdrlen = format_object_header(hdr, sizeof(hdr), obj_type, size);\n \n \t/* Sha1.. */\n \tr->hash_algo->init_fn(&c);\n@@ -1734,7 +1742,7 @@ static void write_object_file_prepare(const struct git_hash_algo *algo,\n \tgit_hash_ctx c;\n \n \t/* Generate the header */\n-\t*hdrlen = xsnprintf(hdr, *hdrlen, \"%s %\"PRIuMAX , type, (uintmax_t)len)+1;\n+\t*hdrlen = format_object_header_extended(hdr, *hdrlen, OBJ_NONE, type, len);\n \n \t/* Sha1.. */\n \talgo->init_fn(&c);\n@@ -2011,7 +2019,7 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n \n \t/* Since oid is not determined, save tmp file to odb path. */\n \tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n-\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX, type_name(OBJ_BLOB), len) + 1;\n+\thdrlen = format_object_header(hdr, sizeof(hdr), OBJ_BLOB, len);\n \n \t/* Common steps for write_loose_object and stream_loose_object to\n \t * start writing loose oject:\n@@ -2152,7 +2160,7 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n \tbuf = read_object(the_repository, oid, &type, &len);\n \tif (!buf)\n \t\treturn error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n-\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX , type_name(type), (uintmax_t)len) + 1;\n+\thdrlen = format_object_header(hdr, sizeof(hdr), type, len);\n \tret = write_loose_object(oid, hdr, hdrlen, buf, len, mtime, 0);\n \tfree(buf);\n \n-- \n2.34.1.52.gc288e771b4.agit.6.5.6\n\n"},{"id":"447406","messageId":"220201.86mtjaacc5.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"b2dee243-1a38-531e-02b1-ffd66c465fa5@web.de","subject":"C99 %z (was: [PATCH v7 2/5] object-file API: add a format_object_header() function)","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-02-01T14:28:30Z","receivedAt":"2022-02-01T14:30:39Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Dec 21 2021, René Scharfe wrote:\n\n> Am 21.12.21 um 12:51 schrieb Han Xin:\n>> From: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n>> [...]\n>>  \t\tthe_hash_algo->init_fn(&c);\n>>  \t\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n>>  \t} else\n>> diff --git a/bulk-checkin.c b/bulk-checkin.c\n>> index 8785b2ac80..1733a1de4f 100644\n>> --- a/bulk-checkin.c\n>> +++ b/bulk-checkin.c\n>> @@ -220,8 +220,8 @@ static int deflate_to_pack(struct bulk_checkin_state *state,\n>>  \tif (seekback == (off_t) -1)\n>>  \t\treturn error(\"cannot find the current offset\");\n>>\n>> -\theader_len = xsnprintf((char *)obuf, sizeof(obuf), \"%s %\" PRIuMAX,\n>> -\t\t\t       type_name(type), (uintmax_t)size) + 1;\n>> +\theader_len = format_object_header((char *)obuf, sizeof(obuf),\n>> +\t\t\t\t\t type, (uintmax_t)size);\n>                                                ^^^^^^^^^^^\n> Same here, just that size is already of type size_t, so a cast makes\n> even less sense.\n\nThanks, this and the below is something I made sure to include in a\nre-roll I'm about to send (to do these cleanups in object-file.c\nseparately from Han Xin's series).\n\n>> +int format_object_header_extended(char *str, size_t size, enum object_type type,\n>> +\t\t\t\t const char *typestr, size_t objsize)\n>> +{\n>> +\tconst char *s = type == OBJ_NONE ? typestr : type_name(type);\n>> +\n>> +\treturn xsnprintf(str, size, \"%s %\"PRIuMAX, s, (uintmax_t)objsize) + 1;\n>                                                       ^^^^^^^^^^^\n> This cast is necessary to match PRIuMAX.  And that is used because the z\n> modifier (as in e.g. printf(\"%zu\", sizeof(size_t));) was only added in\n> C99 and not all platforms may have it.  (Perhaps this cautious approach\n> is worth revisiting separately, now that some time has passed, but this\n> patch series should still use PRIuMAX, as it does.)\n\nI tried to use %z recently and found that the CI breaks on Windows, but\nthis was a few months ago. But I think the status of that particular C99\nfeature is that we can't use it freely, unfortunately. I may be wrong\nabout that, I haven't looked it any detail beyond running those CI\nerrors.\n"},{"id":"447466","messageId":"220201.86wnie8eg0.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"20220120112114.47618-1-chiyutianyi@gmail.com","subject":"Re: [PATCH v9 0/5] unpack large blobs in stream","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-02-01T21:24:46Z","receivedAt":"2022-02-01T21:28:05Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, Jan 20 2022, Han Xin wrote:\n\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> Changes since v8:\n> * Rename \"assert_no_loose ()\" into \"test_no_loose ()\" in\n>   \"t5329-unpack-large-objects.sh\". Remove \"assert_no_pack ()\" and use\n>   \"test_dir_is_empty\" instead.\n>\n> * Revert changes to \"create_tmpfile()\" and error handling is now in\n>   \"start_loose_object_common()\".\n>\n> * Remove \"finalize_object_file_with_mtime()\" which seems to be an overkill\n>   for \"write_loose_object()\" now. \n>\n> * Remove the commit \"object-file.c: remove the slash for directory_size()\",\n>   it can be in a separate patch if necessary.\n>\n> Han Xin (4):\n>   unpack-objects: low memory footprint for get_data() in dry_run mode\n>   object-file.c: refactor write_loose_object() to several steps\n>   object-file.c: add \"stream_loose_object()\" to handle large object\n>   unpack-objects: unpack_non_delta_entry() read data in a stream\n>\n> Ævar Arnfjörð Bjarmason (1):\n>   object-file API: add a format_object_header() function\n\nI sent\nhttps://lore.kernel.org/git/cover-00.10-00000000000-20220201T144803Z-avarab@gmail.com/\ntoday which suggests splitting out the 5/5 cleanup you'd integrated.\n\nI then rebased these patches of yours on top of that, the result is\nhere:\nhttps://github.com/avar/git/tree/han-xin-avar/unpack-loose-object-streaming-9\n\nThe range-diff to your version is below. There's a few unrelated\nfixes/nits in it.\n\nI think with/without basing this on top of my series above your patches\nhere look good with the nits pointed out in the diff below addressed\n(and some don't need to be). I.e. the dependency on it is rather\ntrivial, and the two could be split up.\n\nWhat do you think is a good way to proceed? I could just submit the\nbelow as a proposed v10 if you'd like & agree...\n\n1:  553a9377eb3 ! 1:  61fcfe7b840 unpack-objects: low memory footprint for get_data() in dry_run mode\n    @@ Commit message\n         unpack-objects: low memory footprint for get_data() in dry_run mode\n     \n         As the name implies, \"get_data(size)\" will allocate and return a given\n    -    size of memory. Allocating memory for a large blob object may cause the\n    +    amount of memory. Allocating memory for a large blob object may cause the\n         system to run out of memory. Before preparing to replace calling of\n         \"get_data()\" to unpack large blob objects in latter commits, refactor\n         \"get_data()\" to reduce memory footprint for dry_run mode.\n    @@ Commit message\n     \n         Suggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n         Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n    +    Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n      ## builtin/unpack-objects.c ##\n     @@ builtin/unpack-objects.c: static void use(int bytes)\n    @@ t/t5328-unpack-large-objects.sh (new)\n     +\n     +test_no_loose () {\n     +\tglob=dest.git/objects/?? &&\n    -+\techo \"$glob\" >expect &&\n    -+\teval \"echo $glob\" >actual &&\n    ++\techo $glob >expect &&\n    ++\techo \"$glob\" >actual &&\n     +\ttest_cmp expect actual\n     +}\n     +\n-:  ----------- > 2:  c6b0437db03 object-file.c: do fsync() and close() before post-write die()\n2:  88c91affd61 ! 3:  77bcfe3da6f object-file.c: refactor write_loose_object() to several steps\n    @@ Commit message\n         When writing a large blob using \"write_loose_object()\", we have to pass\n         a buffer with the whole content of the blob, and this behavior will\n         consume lots of memory and may cause OOM. We will introduce a stream\n    -    version function (\"stream_loose_object()\") in latter commit to resolve\n    +    version function (\"stream_loose_object()\") in later commit to resolve\n         this issue.\n     \n    -    Before introducing a stream vesion function for writing loose object,\n    -    do some refactoring on \"write_loose_object()\" to reuse code for both\n    -    versions.\n    +    Before introducing that streaming function, do some refactoring on\n    +    \"write_loose_object()\" to reuse code for both versions.\n     \n         Rewrite \"write_loose_object()\" as follows:\n     \n    @@ Commit message\n     \n          3. Compress data.\n     \n    -     4. Move common steps for ending zlib stream into a new funciton\n    +     4. Move common steps for ending zlib stream into a new function\n             \"end_loose_object_common()\".\n     \n          5. Close fd and finalize the object file.\n    @@ Commit message\n         Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n         Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n         Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n    +    Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n      ## object-file.c ##\n     @@ object-file.c: static int create_tmpfile(struct strbuf *tmp, const char *filename)\n      \treturn fd;\n      }\n      \n    ++/**\n    ++ * Common steps for loose object writers to start writing loose\n    ++ * objects:\n    ++ *\n    ++ * - Create tmpfile for the loose object.\n    ++ * - Setup zlib stream for compression.\n    ++ * - Start to feed header to zlib stream.\n    ++ *\n    ++ * Returns a \"fd\", which should later be provided to\n    ++ * end_loose_object_common().\n    ++ */\n     +static int start_loose_object_common(struct strbuf *tmp_file,\n     +\t\t\t\t     const char *filename, unsigned flags,\n     +\t\t\t\t     git_zstream *stream,\n    @@ object-file.c: static int create_tmpfile(struct strbuf *tmp, const char *filenam\n     +\treturn fd;\n     +}\n     +\n    -+static void end_loose_object_common(int ret, git_hash_ctx *c,\n    ++/**\n    ++ * Common steps for loose object writers to end writing loose objects:\n    ++ *\n    ++ * - End the compression of zlib stream.\n    ++ * - Get the calculated oid to \"parano_oid\".\n    ++ * - fsync() and close() the \"fd\"\n    ++ */\n    ++static void end_loose_object_common(int fd, int ret, git_hash_ctx *c,\n     +\t\t\t\t    git_zstream *stream,\n     +\t\t\t\t    struct object_id *parano_oid,\n     +\t\t\t\t    const struct object_id *expected_oid,\n    @@ object-file.c: static int create_tmpfile(struct strbuf *tmp, const char *filenam\n     +\tif (ret != Z_OK)\n     +\t\tdie(_(die_msg2_fmt), ret, expected_oid);\n     +\tthe_hash_algo->final_oid_fn(parano_oid, c);\n    ++\n    ++\t/*\n    ++\t * We already did a write_buffer() to the \"fd\", let's fsync()\n    ++\t * and close().\n    ++\t *\n    ++\t * We might still die() on a subsequent sanity check, but\n    ++\t * let's not add to that confusion by not flushing any\n    ++\t * outstanding writes to disk first.\n    ++\t */\n    ++\tclose_loose_object(fd);\n     +}\n     +\n      static int write_loose_object(const struct object_id *oid, char *hdr,\n    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *\n     -\twhile (git_deflate(&stream, 0) == Z_OK)\n     -\t\t; /* nothing */\n     -\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n    -+\t/* Common steps for write_loose_object and stream_loose_object to\n    -+\t * start writing loose oject:\n    -+\t *\n    -+\t *  - Create tmpfile for the loose object.\n    -+\t *  - Setup zlib stream for compression.\n    -+\t *  - Start to feed header to zlib stream.\n    -+\t */\n     +\tfd = start_loose_object_common(&tmp_file, filename.buf, flags,\n     +\t\t\t\t       &stream, compressed, sizeof(compressed),\n     +\t\t\t\t       &c, hdr, hdrlen);\n    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *\n     -\t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n     -\t\t    ret);\n     -\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n    -+\t/* Common steps for write_loose_object and stream_loose_object to\n    -+\t * end writing loose oject:\n    -+\t *\n    -+\t *  - End the compression of zlib stream.\n    -+\t *  - Get the calculated oid to \"parano_oid\".\n    -+\t */\n    -+\tend_loose_object_common(ret, &c, &stream, &parano_oid, oid,\n    +-\n    +-\t/*\n    +-\t * We already did a write_buffer() to the \"fd\", let's fsync()\n    +-\t * and close().\n    +-\t *\n    +-\t * We might still die() on a subsequent sanity check, but\n    +-\t * let's not add to that confusion by not flushing any\n    +-\t * outstanding writes to disk first.\n    +-\t */\n    +-\tclose_loose_object(fd);\n    ++\tend_loose_object_common(fd, ret, &c, &stream, &parano_oid, oid,\n     +\t\t\t\tN_(\"unable to deflate new object %s (%d)\"),\n     +\t\t\t\tN_(\"deflateEnd on object %s failed (%d)\"));\n    -+\n    + \n      \tif (!oideq(oid, &parano_oid))\n      \t\tdie(_(\"confused by unstable object source data for %s\"),\n    - \t\t    oid_to_hex(oid));\n3:  054a00ed21d ! 4:  71c10e734d1 object-file.c: add \"stream_loose_object()\" to handle large object\n    @@ Commit message\n     \n         Add a new function \"stream_loose_object()\", which is a stream version of\n         \"write_loose_object()\" but with a low memory footprint. We will use this\n    -    function to unpack large blob object in latter commit.\n    +    function to unpack large blob object in later commit.\n     \n         Another difference with \"write_loose_object()\" is that we have no chance\n         to run \"write_object_file_prepare()\" to calculate the oid in advance.\n         In \"write_loose_object()\", we know the oid and we can write the\n         temporary file in the same directory as the final object, but for an\n         object with an undetermined oid, we don't know the exact directory for\n    -    the object, so we have to save the temporary file in \".git/objects/\"\n    -    directory instead.\n    +    the object.\n    +\n    +    Still, we need to save the temporary file we're preparing\n    +    somewhere. We'll do that in the top-level \".git/objects/\"\n    +    directory (or whatever \"GIT_OBJECT_DIRECTORY\" is set to). Once we've\n    +    streamed it we'll know the OID, and will move it to its canonical\n    +    path.\n     \n         \"freshen_packed_object()\" or \"freshen_loose_object()\" will be called\n         inside \"stream_loose_object()\" after obtaining the \"oid\".\n    @@ object-file.c: static int freshen_packed_object(const struct object_id *oid)\n     +\n     +\t/* Since oid is not determined, save tmp file to odb path. */\n     +\tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n    -+\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX, type_name(OBJ_BLOB), len) + 1;\n    ++\thdrlen = format_object_header(hdr, sizeof(hdr), OBJ_BLOB, len);\n     +\n     +\t/* Common steps for write_loose_object and stream_loose_object to\n     +\t * start writing loose oject:\n    @@ object-file.c: static int freshen_packed_object(const struct object_id *oid)\n     +\t *  - End the compression of zlib stream.\n     +\t *  - Get the calculated oid.\n     +\t */\n    -+\tend_loose_object_common(ret, &c, &stream, oid, NULL,\n    ++\tend_loose_object_common(fd, ret, &c, &stream, oid, NULL,\n     +\t\t\t\tN_(\"unable to stream deflate new object (%d)\"),\n     +\t\t\t\tN_(\"deflateEnd on stream object failed (%d)\"));\n     +\n    -+\tclose_loose_object(fd);\n    -+\n     +\tif (freshen_packed_object(oid) || freshen_loose_object(oid)) {\n     +\t\tunlink_or_warn(tmp_file.buf);\n     +\t\tgoto cleanup;\n    @@ object-file.c: static int freshen_packed_object(const struct object_id *oid)\n     +}\n     +\n      int write_object_file_flags(const void *buf, unsigned long len,\n    - \t\t\t    const char *type, struct object_id *oid,\n    + \t\t\t    enum object_type type, struct object_id *oid,\n      \t\t\t    unsigned flags)\n     \n      ## object-store.h ##\n    @@ object-store.h: static inline int write_object_file(const void *buf, unsigned lo\n      \n     +int stream_loose_object(struct input_stream *in_stream, size_t len,\n     +\t\t\tstruct object_id *oid);\n    -+\n    - int hash_object_file_literally(const void *buf, unsigned long len,\n    - \t\t\t       const char *type, struct object_id *oid,\n    - \t\t\t       unsigned flags);\n    + int hash_write_object_file_literally(const void *buf, unsigned long len,\n    + \t\t\t\t     const char *type, struct object_id *oid,\n    + \t\t\t\t     unsigned flags);\n-:  ----------- > 5:  3c1d788d69d core doc: modernize core.bigFileThreshold documentation\n4:  6bcba6bce66 ! 6:  8b83f6d6b83 unpack-objects: unpack_non_delta_entry() read data in a stream\n    @@ Metadata\n     Author: Han Xin <hanxin.hx@alibaba-inc.com>\n     \n      ## Commit message ##\n    -    unpack-objects: unpack_non_delta_entry() read data in a stream\n    +    unpack-objects: use stream_loose_object() to unpack large objects\n     \n    -    We used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\n    -    entire contents of a blob object, no matter how big it is. This\n    -    implementation may consume all the memory and cause OOM.\n    +    Make use of the stream_loose_object() function introduced in the\n    +    preceding commit to unpack large objects. Before this we'd need to\n    +    malloc() the size of the blob before unpacking it, which could cause\n    +    OOM with very large blobs.\n     \n    -    By implementing a zstream version of input_stream interface, we can use\n    -    a small fixed buffer for \"unpack_non_delta_entry()\". However, unpack\n    -    non-delta objects from a stream instead of from an entrie buffer will\n    -    have 10% performance penalty.\n    +    We could use this new interface to unpack all blobs, but doing so\n    +    would result in a performance penalty of around 10%, as the below\n    +    \"hyperfine\" benchmark will show. We therefore limit this to files\n    +    larger than \"core.bigFileThreshold\":\n     \n             $ hyperfine \\\n               --setup \\\n    @@ Commit message\n                         -c core.bigFileThreshold=16k unpack-objects\n                         <small.pack' in 'HEAD~1'\n     \n    -    Therefore, only unpack objects larger than the \"core.bigFileThreshold\"\n    -    in zstream. Until now, the config variable has been used in the\n    -    following cases, and our new case belongs to the packfile category.\n    +    An earlier version of this patch introduced a new\n    +    \"core.bigFileStreamingThreshold\" instead of re-using the existing\n    +    \"core.bigFileThreshold\" variable[1]. As noted in a detailed overview\n    +    of its users in [2] using it has several different meanings.\n     \n    -     * Archive:\n    +    Still, we consider it good enough to simply re-use it. While it's\n    +    possible that someone might want to e.g. consider objects \"small\" for\n    +    the purposes of diffing but \"big\" for the purposes of writing them\n    +    such use-cases are probably too obscure to worry about. We can always\n    +    split up \"core.bigFileThreshold\" in the future if there's a need for\n    +    that.\n     \n    -       + archive.c: write_entry(): write large blob entries to archive in\n    -         stream.\n    -\n    -     * Loose objects:\n    -\n    -       + object-file.c: index_fd(): when hashing large files in worktree,\n    -         read files in a stream, and create one packfile per large blob if\n    -         want to save files to git object store.\n    -\n    -       + object-file.c: read_loose_object(): when checking loose objects\n    -         using \"git-fsck\", do not read full content of large loose objects.\n    -\n    -     * Packfile:\n    -\n    -       + fast-import.c: parse_and_store_blob(): streaming large blob from\n    -         foreign source to packfile.\n    -\n    -       + index-pack.c: check_collison(): read and check large blob in stream.\n    -\n    -       + index-pack.c: unpack_entry_data(): do not return the entire\n    -         contents of the big blob from packfile, but uses a fixed buf to\n    -         perform some integrity checks on the object.\n    -\n    -       + pack-check.c: verify_packfile(): used by \"git-fsck\" and will call\n    -         check_object_signature() to check large blob in pack with the\n    -         streaming interface.\n    -\n    -       + pack-objects.c: get_object_details(): set \"no_try_delta\" for large\n    -         blobs when counting objects.\n    -\n    -       + pack-objects.c: write_no_reuse_object(): streaming large blob to\n    -         pack.\n    -\n    -       + unpack-objects.c: unpack_non_delta_entry(): unpack large blob in\n    -         stream from packfile.\n    -\n    -     * Others:\n    -\n    -       + diff.c: diff_populate_filespec(): treat large blob file as binary.\n    -\n    -       + streaming.c: istream_source(): as a helper of \"open_istream()\" to\n    -         select proper streaming interface to read large blob from packfile.\n    +    1. https://lore.kernel.org/git/20211210103435.83656-1-chiyutianyi@gmail.com/\n    +    2. https://lore.kernel.org/git/20220120112114.47618-5-chiyutianyi@gmail.com/\n     \n         Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n         Helped-by: Derrick Stolee <stolee@gmail.com>\n         Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n         Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n     \n    + ## Documentation/config/core.txt ##\n    +@@ Documentation/config/core.txt: usage, at the slight expense of increased disk usage.\n    + * Will be generally be streamed when written, which avoids excessive\n    + memory usage, at the cost of some fixed overhead. Commands that make\n    + use of this include linkgit:git-archive[1],\n    +-linkgit:git-fast-import[1], linkgit:git-index-pack[1] and\n    +-linkgit:git-fsck[1].\n    ++linkgit:git-fast-import[1], linkgit:git-index-pack[1],\n    ++linkgit:git-unpack-objects[1] and linkgit:git-fsck[1].\n    + \n    + core.excludesFile::\n    + \tSpecifies the pathname to the file that contains patterns to\n    +\n      ## builtin/unpack-objects.c ##\n     @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type type,\n      \t}\n5:  1bfaf89ee0b < -:  ----------- object-file API: add a format_object_header() function\n"},{"id":"447523","messageId":"CAO0brD2Pe0aKSiBphZS861gC=nZk+q2GtXDN4pPjAQnPdns3TA@mail.gmail.com","threadId":"56672","inReplyTo":"220201.86wnie8eg0.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v9 0/5] unpack large blobs in stream","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-02-02T08:32:40Z","receivedAt":"2022-02-02T08:32:56Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Wed, Feb 2, 2022 at 5:28 AM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>\n>\n> On Thu, Jan 20 2022, Han Xin wrote:\n>\n> > From: Han Xin <hanxin.hx@alibaba-inc.com>\n> >\n> > Changes since v8:\n> > * Rename \"assert_no_loose ()\" into \"test_no_loose ()\" in\n> >   \"t5329-unpack-large-objects.sh\". Remove \"assert_no_pack ()\" and use\n> >   \"test_dir_is_empty\" instead.\n> >\n> > * Revert changes to \"create_tmpfile()\" and error handling is now in\n> >   \"start_loose_object_common()\".\n> >\n> > * Remove \"finalize_object_file_with_mtime()\" which seems to be an overkill\n> >   for \"write_loose_object()\" now.\n> >\n> > * Remove the commit \"object-file.c: remove the slash for directory_size()\",\n> >   it can be in a separate patch if necessary.\n> >\n> > Han Xin (4):\n> >   unpack-objects: low memory footprint for get_data() in dry_run mode\n> >   object-file.c: refactor write_loose_object() to several steps\n> >   object-file.c: add \"stream_loose_object()\" to handle large object\n> >   unpack-objects: unpack_non_delta_entry() read data in a stream\n> >\n> > Ævar Arnfjörð Bjarmason (1):\n> >   object-file API: add a format_object_header() function\n>\n> I sent\n> https://lore.kernel.org/git/cover-00.10-00000000000-20220201T144803Z-avarab@gmail.com/\n> today which suggests splitting out the 5/5 cleanup you'd integrated.\n>\n> I then rebased these patches of yours on top of that, the result is\n> here:\n> https://github.com/avar/git/tree/han-xin-avar/unpack-loose-object-streaming-9\n>\n> The range-diff to your version is below. There's a few unrelated\n> fixes/nits in it.\n>\n> I think with/without basing this on top of my series above your patches\n> here look good with the nits pointed out in the diff below addressed\n> (and some don't need to be). I.e. the dependency on it is rather\n> trivial, and the two could be split up.\n>\n> What do you think is a good way to proceed? I could just submit the\n> below as a proposed v10 if you'd like & agree...\n>\n\nYes, thanks for the suggestions, and I'm glad you're happy to do so.\n\nThanks.\n-Han Xin\n\n> 1:  553a9377eb3 ! 1:  61fcfe7b840 unpack-objects: low memory footprint for get_data() in dry_run mode\n>     @@ Commit message\n>          unpack-objects: low memory footprint for get_data() in dry_run mode\n>\n>          As the name implies, \"get_data(size)\" will allocate and return a given\n>     -    size of memory. Allocating memory for a large blob object may cause the\n>     +    amount of memory. Allocating memory for a large blob object may cause the\n>          system to run out of memory. Before preparing to replace calling of\n>          \"get_data()\" to unpack large blob objects in latter commits, refactor\n>          \"get_data()\" to reduce memory footprint for dry_run mode.\n>     @@ Commit message\n>\n>          Suggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n>          Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n>     +    Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n>\n>       ## builtin/unpack-objects.c ##\n>      @@ builtin/unpack-objects.c: static void use(int bytes)\n>     @@ t/t5328-unpack-large-objects.sh (new)\n>      +\n>      +test_no_loose () {\n>      +  glob=dest.git/objects/?? &&\n>     -+  echo \"$glob\" >expect &&\n>     -+  eval \"echo $glob\" >actual &&\n>     ++  echo $glob >expect &&\n>     ++  echo \"$glob\" >actual &&\n>      +  test_cmp expect actual\n>      +}\n>      +\n\nI have a small doubt with this, it works fine with dash, but not\nothers like zsh. Wouldn't\nit be better to do compatibility, or would it introduce other issues\nthat I don't know?\n\nThanks.\n-Han Xin\n"},{"id":"447528","messageId":"220202.86czk58rcu.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"CAO0brD2Pe0aKSiBphZS861gC=nZk+q2GtXDN4pPjAQnPdns3TA@mail.gmail.com","subject":"Re: [PATCH v9 0/5] unpack large blobs in stream","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-02-02T10:59:47Z","receivedAt":"2022-02-02T11:01:27Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Feb 02 2022, Han Xin wrote:\n\n> On Wed, Feb 2, 2022 at 5:28 AM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>>\n>>\n>> On Thu, Jan 20 2022, Han Xin wrote:\n>>\n>> > From: Han Xin <hanxin.hx@alibaba-inc.com>\n>> >\n>> > Changes since v8:\n>> > * Rename \"assert_no_loose ()\" into \"test_no_loose ()\" in\n>> >   \"t5329-unpack-large-objects.sh\". Remove \"assert_no_pack ()\" and use\n>> >   \"test_dir_is_empty\" instead.\n>> >\n>> > * Revert changes to \"create_tmpfile()\" and error handling is now in\n>> >   \"start_loose_object_common()\".\n>> >\n>> > * Remove \"finalize_object_file_with_mtime()\" which seems to be an overkill\n>> >   for \"write_loose_object()\" now.\n>> >\n>> > * Remove the commit \"object-file.c: remove the slash for directory_size()\",\n>> >   it can be in a separate patch if necessary.\n>> >\n>> > Han Xin (4):\n>> >   unpack-objects: low memory footprint for get_data() in dry_run mode\n>> >   object-file.c: refactor write_loose_object() to several steps\n>> >   object-file.c: add \"stream_loose_object()\" to handle large object\n>> >   unpack-objects: unpack_non_delta_entry() read data in a stream\n>> >\n>> > Ævar Arnfjörð Bjarmason (1):\n>> >   object-file API: add a format_object_header() function\n>>\n>> I sent\n>> https://lore.kernel.org/git/cover-00.10-00000000000-20220201T144803Z-avarab@gmail.com/\n>> today which suggests splitting out the 5/5 cleanup you'd integrated.\n>>\n>> I then rebased these patches of yours on top of that, the result is\n>> here:\n>> https://github.com/avar/git/tree/han-xin-avar/unpack-loose-object-streaming-9\n>>\n>> The range-diff to your version is below. There's a few unrelated\n>> fixes/nits in it.\n>>\n>> I think with/without basing this on top of my series above your patches\n>> here look good with the nits pointed out in the diff below addressed\n>> (and some don't need to be). I.e. the dependency on it is rather\n>> trivial, and the two could be split up.\n>>\n>> What do you think is a good way to proceed? I could just submit the\n>> below as a proposed v10 if you'd like & agree...\n>>\n>\n> Yes, thanks for the suggestions, and I'm glad you're happy to do so.\n\nWilldo.\n\n>> 1:  553a9377eb3 ! 1:  61fcfe7b840 unpack-objects: low memory footprint for get_data() in dry_run mode\n>>     @@ Commit message\n>>          unpack-objects: low memory footprint for get_data() in dry_run mode\n>>\n>>          As the name implies, \"get_data(size)\" will allocate and return a given\n>>     -    size of memory. Allocating memory for a large blob object may cause the\n>>     +    amount of memory. Allocating memory for a large blob object may cause the\n>>          system to run out of memory. Before preparing to replace calling of\n>>          \"get_data()\" to unpack large blob objects in latter commits, refactor\n>>          \"get_data()\" to reduce memory footprint for dry_run mode.\n>>     @@ Commit message\n>>\n>>          Suggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n>>          Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n>>     +    Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n>>\n>>       ## builtin/unpack-objects.c ##\n>>      @@ builtin/unpack-objects.c: static void use(int bytes)\n>>     @@ t/t5328-unpack-large-objects.sh (new)\n>>      +\n>>      +test_no_loose () {\n>>      +  glob=dest.git/objects/?? &&\n>>     -+  echo \"$glob\" >expect &&\n>>     -+  eval \"echo $glob\" >actual &&\n>>     ++  echo $glob >expect &&\n>>     ++  echo \"$glob\" >actual &&\n>>      +  test_cmp expect actual\n>>      +}\n>>      +\n>\n> I have a small doubt with this, it works fine with dash, but not\n> others like zsh. Wouldn't\n> it be better to do compatibility, or would it introduce other issues\n> that I don't know?\n\nAh, I hadn't spotted that zsh issue. I don't think the test suite will\nrun on it in general, but in any case I'll fix this.\n\nThere's a few other tests that do this just by piping \"find\" to \"wc -l\",\nit's probably better to just follow that pattern. I think the eval\nworks, but I thought it was a bit unusual/stood out.\n"},{"id":"447739","messageId":"cover-v10-0.6-00000000000-20220204T135538Z-avarab@gmail.com","threadId":"56672","inReplyTo":"20220120112114.47618-1-chiyutianyi@gmail.com","subject":"[PATCH v10 0/6] unpack-objects: support streaming large objects to disk","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-02-04T14:07:06Z","receivedAt":"2022-02-04T14:07:24Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"This is a v10 re-roll of Han Xin's series[1] to stream large objects\nto disk in \"git unpack-objects\". This v9 had integrated a proposed\ncleanup patch of mine, which is now a part of its own series, which\nthis series now depends on: [2]. This v10 is sent with Han Xin's\napproval[3].\n\nChanges since v9:\n\n * Now based on [2]\n * Small grammar/typo fixes in commit messages\n * Replaced an echo/eval pattern in a test with a $(find ... | wc -l)\n   comparison, which is a pattern we already use in another test for\n   the same (or similar) assertion.\n * I added a new 2/6 to do an fsync() before an oideq() assertion. I\n   don't think it matters in practice, but allows 3/6 to be smaller by\n   having that code-now-utility-function share more logic among its two callers.\n * Changed inline comments in 3/6 to API docs where appropriate, the\n   helper function now gets a \"fd\" per 2/6.\n * 4/6 could use the format_object_header() function in the base\n   topic, and now does so (instead of that conversion coming later in\n   v9).\n * A new 5/6 updates the core.bigFileThreshold documentation to\n   account for 12 years of behavior changes we hadn't documented.\n * The updated 6/6 now links to those docs, and I removed a very\n   detailed accounting of all in-tree uses of core.bigFileThreshold\n   from the commit message. I think linking to the summary docs should\n   suffice, and for anyone digging in the future 5/6 links to the more\n   detailed summary in the old patch.\n\nMore generally I've been heavily involved in the review for the past\niterations, and I think barring any last minute nits in this v10 this\ntopic should be ready to advance. As the above summary shows we're\ndown to typo fixes, doc and test tweaks etc. at this point.\n\nThe core functionality being added here isn't changed in any\nmeaningful way, and has had a lot of careful review already.\n\n1. https://lore.kernel.org/git/20220120112114.47618-1-chiyutianyi@gmail.com/\n2. https://lore.kernel.org/git/cover-v2-00.11-00000000000-20220204T135005Z-avarab@gmail.com/\n3. https://lore.kernel.org/git/CAO0brD2Pe0aKSiBphZS861gC=nZk+q2GtXDN4pPjAQnPdns3TA@mail.gmail.com/\n\nHan Xin (4):\n  unpack-objects: low memory footprint for get_data() in dry_run mode\n  object-file.c: refactor write_loose_object() to several steps\n  object-file.c: add \"stream_loose_object()\" to handle large object\n  unpack-objects: use stream_loose_object() to unpack large objects\n\nÆvar Arnfjörð Bjarmason (2):\n  object-file.c: do fsync() and close() before post-write die()\n  core doc: modernize core.bigFileThreshold documentation\n\n Documentation/config/core.txt   |  33 +++--\n builtin/unpack-objects.c        | 110 ++++++++++++++--\n object-file.c                   | 221 +++++++++++++++++++++++++++-----\n object-store.h                  |   8 ++\n t/t5328-unpack-large-objects.sh |  62 +++++++++\n 5 files changed, 381 insertions(+), 53 deletions(-)\n create mode 100755 t/t5328-unpack-large-objects.sh\n\nRange-diff against v9:\n1:  553a9377eb3 ! 1:  e46eb75b98f unpack-objects: low memory footprint for get_data() in dry_run mode\n    @@ Commit message\n         unpack-objects: low memory footprint for get_data() in dry_run mode\n     \n         As the name implies, \"get_data(size)\" will allocate and return a given\n    -    size of memory. Allocating memory for a large blob object may cause the\n    +    amount of memory. Allocating memory for a large blob object may cause the\n         system to run out of memory. Before preparing to replace calling of\n         \"get_data()\" to unpack large blob objects in latter commits, refactor\n         \"get_data()\" to reduce memory footprint for dry_run mode.\n    @@ Commit message\n         in dry_run mode, \"get_data()\" will release the allocated buffer and\n         return NULL instead of returning garbage data.\n     \n    +    The \"find [...]objects/?? -type f | wc -l\" test idiom being used here\n    +    is adapted from the same \"find\" use added to another test in\n    +    d9545c7f465 (fast-import: implement unpack limit, 2016-04-25).\n    +\n         Suggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n         Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n    +    Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n      ## builtin/unpack-objects.c ##\n     @@ builtin/unpack-objects.c: static void use(int bytes)\n    @@ t/t5328-unpack-large-objects.sh (new)\n     +}\n     +\n     +test_no_loose () {\n    -+\tglob=dest.git/objects/?? &&\n    -+\techo \"$glob\" >expect &&\n    -+\teval \"echo $glob\" >actual &&\n    -+\ttest_cmp expect actual\n    ++\ttest $(find dest.git/objects/?? -type f | wc -l) = 0\n     +}\n     +\n     +test_expect_success \"create large objects (1.5 MB) and PACK\" '\n-:  ----------- > 2:  48bf9090058 object-file.c: do fsync() and close() before post-write die()\n2:  88c91affd61 ! 3:  0e33d2a6e35 object-file.c: refactor write_loose_object() to several steps\n    @@ Commit message\n         When writing a large blob using \"write_loose_object()\", we have to pass\n         a buffer with the whole content of the blob, and this behavior will\n         consume lots of memory and may cause OOM. We will introduce a stream\n    -    version function (\"stream_loose_object()\") in latter commit to resolve\n    +    version function (\"stream_loose_object()\") in later commit to resolve\n         this issue.\n     \n    -    Before introducing a stream vesion function for writing loose object,\n    -    do some refactoring on \"write_loose_object()\" to reuse code for both\n    -    versions.\n    +    Before introducing that streaming function, do some refactoring on\n    +    \"write_loose_object()\" to reuse code for both versions.\n     \n         Rewrite \"write_loose_object()\" as follows:\n     \n    @@ Commit message\n     \n          3. Compress data.\n     \n    -     4. Move common steps for ending zlib stream into a new funciton\n    +     4. Move common steps for ending zlib stream into a new function\n             \"end_loose_object_common()\".\n     \n          5. Close fd and finalize the object file.\n    @@ Commit message\n         Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n         Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n         Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n    +    Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n      ## object-file.c ##\n     @@ object-file.c: static int create_tmpfile(struct strbuf *tmp, const char *filename)\n      \treturn fd;\n      }\n      \n    ++/**\n    ++ * Common steps for loose object writers to start writing loose\n    ++ * objects:\n    ++ *\n    ++ * - Create tmpfile for the loose object.\n    ++ * - Setup zlib stream for compression.\n    ++ * - Start to feed header to zlib stream.\n    ++ *\n    ++ * Returns a \"fd\", which should later be provided to\n    ++ * end_loose_object_common().\n    ++ */\n     +static int start_loose_object_common(struct strbuf *tmp_file,\n     +\t\t\t\t     const char *filename, unsigned flags,\n     +\t\t\t\t     git_zstream *stream,\n    @@ object-file.c: static int create_tmpfile(struct strbuf *tmp, const char *filenam\n     +\treturn fd;\n     +}\n     +\n    -+static void end_loose_object_common(int ret, git_hash_ctx *c,\n    ++/**\n    ++ * Common steps for loose object writers to end writing loose objects:\n    ++ *\n    ++ * - End the compression of zlib stream.\n    ++ * - Get the calculated oid to \"parano_oid\".\n    ++ * - fsync() and close() the \"fd\"\n    ++ */\n    ++static void end_loose_object_common(int fd, int ret, git_hash_ctx *c,\n     +\t\t\t\t    git_zstream *stream,\n     +\t\t\t\t    struct object_id *parano_oid,\n     +\t\t\t\t    const struct object_id *expected_oid,\n    @@ object-file.c: static int create_tmpfile(struct strbuf *tmp, const char *filenam\n     +\tif (ret != Z_OK)\n     +\t\tdie(_(die_msg2_fmt), ret, expected_oid);\n     +\tthe_hash_algo->final_oid_fn(parano_oid, c);\n    ++\n    ++\t/*\n    ++\t * We already did a write_buffer() to the \"fd\", let's fsync()\n    ++\t * and close().\n    ++\t *\n    ++\t * We might still die() on a subsequent sanity check, but\n    ++\t * let's not add to that confusion by not flushing any\n    ++\t * outstanding writes to disk first.\n    ++\t */\n    ++\tclose_loose_object(fd);\n     +}\n     +\n      static int write_loose_object(const struct object_id *oid, char *hdr,\n    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *\n     -\twhile (git_deflate(&stream, 0) == Z_OK)\n     -\t\t; /* nothing */\n     -\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n    -+\t/* Common steps for write_loose_object and stream_loose_object to\n    -+\t * start writing loose oject:\n    -+\t *\n    -+\t *  - Create tmpfile for the loose object.\n    -+\t *  - Setup zlib stream for compression.\n    -+\t *  - Start to feed header to zlib stream.\n    -+\t */\n     +\tfd = start_loose_object_common(&tmp_file, filename.buf, flags,\n     +\t\t\t\t       &stream, compressed, sizeof(compressed),\n     +\t\t\t\t       &c, hdr, hdrlen);\n    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *\n     -\t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n     -\t\t    ret);\n     -\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n    -+\t/* Common steps for write_loose_object and stream_loose_object to\n    -+\t * end writing loose oject:\n    -+\t *\n    -+\t *  - End the compression of zlib stream.\n    -+\t *  - Get the calculated oid to \"parano_oid\".\n    -+\t */\n    -+\tend_loose_object_common(ret, &c, &stream, &parano_oid, oid,\n    +-\n    +-\t/*\n    +-\t * We already did a write_buffer() to the \"fd\", let's fsync()\n    +-\t * and close().\n    +-\t *\n    +-\t * We might still die() on a subsequent sanity check, but\n    +-\t * let's not add to that confusion by not flushing any\n    +-\t * outstanding writes to disk first.\n    +-\t */\n    +-\tclose_loose_object(fd);\n    ++\tend_loose_object_common(fd, ret, &c, &stream, &parano_oid, oid,\n     +\t\t\t\tN_(\"unable to deflate new object %s (%d)\"),\n     +\t\t\t\tN_(\"deflateEnd on object %s failed (%d)\"));\n    -+\n    + \n      \tif (!oideq(oid, &parano_oid))\n      \t\tdie(_(\"confused by unstable object source data for %s\"),\n    - \t\t    oid_to_hex(oid));\n3:  054a00ed21d ! 4:  9644df5c744 object-file.c: add \"stream_loose_object()\" to handle large object\n    @@ Commit message\n     \n         Add a new function \"stream_loose_object()\", which is a stream version of\n         \"write_loose_object()\" but with a low memory footprint. We will use this\n    -    function to unpack large blob object in latter commit.\n    +    function to unpack large blob object in later commit.\n     \n         Another difference with \"write_loose_object()\" is that we have no chance\n         to run \"write_object_file_prepare()\" to calculate the oid in advance.\n         In \"write_loose_object()\", we know the oid and we can write the\n         temporary file in the same directory as the final object, but for an\n         object with an undetermined oid, we don't know the exact directory for\n    -    the object, so we have to save the temporary file in \".git/objects/\"\n    -    directory instead.\n    +    the object.\n    +\n    +    Still, we need to save the temporary file we're preparing\n    +    somewhere. We'll do that in the top-level \".git/objects/\"\n    +    directory (or whatever \"GIT_OBJECT_DIRECTORY\" is set to). Once we've\n    +    streamed it we'll know the OID, and will move it to its canonical\n    +    path.\n     \n         \"freshen_packed_object()\" or \"freshen_loose_object()\" will be called\n         inside \"stream_loose_object()\" after obtaining the \"oid\".\n    @@ object-file.c: static int freshen_packed_object(const struct object_id *oid)\n     +\n     +\t/* Since oid is not determined, save tmp file to odb path. */\n     +\tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n    -+\thdrlen = xsnprintf(hdr, sizeof(hdr), \"%s %\"PRIuMAX, type_name(OBJ_BLOB), len) + 1;\n    ++\thdrlen = format_object_header(hdr, sizeof(hdr), OBJ_BLOB, len);\n     +\n     +\t/* Common steps for write_loose_object and stream_loose_object to\n     +\t * start writing loose oject:\n    @@ object-file.c: static int freshen_packed_object(const struct object_id *oid)\n     +\t *  - End the compression of zlib stream.\n     +\t *  - Get the calculated oid.\n     +\t */\n    -+\tend_loose_object_common(ret, &c, &stream, oid, NULL,\n    ++\tend_loose_object_common(fd, ret, &c, &stream, oid, NULL,\n     +\t\t\t\tN_(\"unable to stream deflate new object (%d)\"),\n     +\t\t\t\tN_(\"deflateEnd on stream object failed (%d)\"));\n     +\n    -+\tclose_loose_object(fd);\n    -+\n     +\tif (freshen_packed_object(oid) || freshen_loose_object(oid)) {\n     +\t\tunlink_or_warn(tmp_file.buf);\n     +\t\tgoto cleanup;\n    @@ object-file.c: static int freshen_packed_object(const struct object_id *oid)\n     +}\n     +\n      int write_object_file_flags(const void *buf, unsigned long len,\n    - \t\t\t    const char *type, struct object_id *oid,\n    + \t\t\t    enum object_type type, struct object_id *oid,\n      \t\t\t    unsigned flags)\n     \n      ## object-store.h ##\n    @@ object-store.h: struct object_directory {\n      \tstruct object_directory *, 1, fspathhash, fspatheq)\n      \n     @@ object-store.h: static inline int write_object_file(const void *buf, unsigned long len,\n    - \treturn write_object_file_flags(buf, len, type, oid, 0);\n    - }\n    - \n    + int write_object_file_literally(const void *buf, unsigned long len,\n    + \t\t\t\tconst char *type, struct object_id *oid,\n    + \t\t\t\tunsigned flags);\n     +int stream_loose_object(struct input_stream *in_stream, size_t len,\n     +\t\t\tstruct object_id *oid);\n    -+\n    - int hash_object_file_literally(const void *buf, unsigned long len,\n    - \t\t\t       const char *type, struct object_id *oid,\n    - \t\t\t       unsigned flags);\n    + \n    + /*\n    +  * Add an object file to the in-memory object store, without writing it\n-:  ----------- > 5:  4550f3a2745 core doc: modernize core.bigFileThreshold documentation\n4:  6bcba6bce66 ! 6:  6a70e49a346 unpack-objects: unpack_non_delta_entry() read data in a stream\n    @@ Metadata\n     Author: Han Xin <hanxin.hx@alibaba-inc.com>\n     \n      ## Commit message ##\n    -    unpack-objects: unpack_non_delta_entry() read data in a stream\n    +    unpack-objects: use stream_loose_object() to unpack large objects\n     \n    -    We used to call \"get_data()\" in \"unpack_non_delta_entry()\" to read the\n    -    entire contents of a blob object, no matter how big it is. This\n    -    implementation may consume all the memory and cause OOM.\n    +    Make use of the stream_loose_object() function introduced in the\n    +    preceding commit to unpack large objects. Before this we'd need to\n    +    malloc() the size of the blob before unpacking it, which could cause\n    +    OOM with very large blobs.\n     \n    -    By implementing a zstream version of input_stream interface, we can use\n    -    a small fixed buffer for \"unpack_non_delta_entry()\". However, unpack\n    -    non-delta objects from a stream instead of from an entrie buffer will\n    -    have 10% performance penalty.\n    +    We could use this new interface to unpack all blobs, but doing so\n    +    would result in a performance penalty of around 10%, as the below\n    +    \"hyperfine\" benchmark will show. We therefore limit this to files\n    +    larger than \"core.bigFileThreshold\":\n     \n             $ hyperfine \\\n               --setup \\\n    @@ Commit message\n                         -c core.bigFileThreshold=16k unpack-objects\n                         <small.pack' in 'HEAD~1'\n     \n    -    Therefore, only unpack objects larger than the \"core.bigFileThreshold\"\n    -    in zstream. Until now, the config variable has been used in the\n    -    following cases, and our new case belongs to the packfile category.\n    +    An earlier version of this patch introduced a new\n    +    \"core.bigFileStreamingThreshold\" instead of re-using the existing\n    +    \"core.bigFileThreshold\" variable[1]. As noted in a detailed overview\n    +    of its users in [2] using it has several different meanings.\n     \n    -     * Archive:\n    +    Still, we consider it good enough to simply re-use it. While it's\n    +    possible that someone might want to e.g. consider objects \"small\" for\n    +    the purposes of diffing but \"big\" for the purposes of writing them\n    +    such use-cases are probably too obscure to worry about. We can always\n    +    split up \"core.bigFileThreshold\" in the future if there's a need for\n    +    that.\n     \n    -       + archive.c: write_entry(): write large blob entries to archive in\n    -         stream.\n    -\n    -     * Loose objects:\n    -\n    -       + object-file.c: index_fd(): when hashing large files in worktree,\n    -         read files in a stream, and create one packfile per large blob if\n    -         want to save files to git object store.\n    -\n    -       + object-file.c: read_loose_object(): when checking loose objects\n    -         using \"git-fsck\", do not read full content of large loose objects.\n    -\n    -     * Packfile:\n    -\n    -       + fast-import.c: parse_and_store_blob(): streaming large blob from\n    -         foreign source to packfile.\n    -\n    -       + index-pack.c: check_collison(): read and check large blob in stream.\n    -\n    -       + index-pack.c: unpack_entry_data(): do not return the entire\n    -         contents of the big blob from packfile, but uses a fixed buf to\n    -         perform some integrity checks on the object.\n    -\n    -       + pack-check.c: verify_packfile(): used by \"git-fsck\" and will call\n    -         check_object_signature() to check large blob in pack with the\n    -         streaming interface.\n    -\n    -       + pack-objects.c: get_object_details(): set \"no_try_delta\" for large\n    -         blobs when counting objects.\n    -\n    -       + pack-objects.c: write_no_reuse_object(): streaming large blob to\n    -         pack.\n    -\n    -       + unpack-objects.c: unpack_non_delta_entry(): unpack large blob in\n    -         stream from packfile.\n    -\n    -     * Others:\n    -\n    -       + diff.c: diff_populate_filespec(): treat large blob file as binary.\n    -\n    -       + streaming.c: istream_source(): as a helper of \"open_istream()\" to\n    -         select proper streaming interface to read large blob from packfile.\n    +    1. https://lore.kernel.org/git/20211210103435.83656-1-chiyutianyi@gmail.com/\n    +    2. https://lore.kernel.org/git/20220120112114.47618-5-chiyutianyi@gmail.com/\n     \n         Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n         Helped-by: Derrick Stolee <stolee@gmail.com>\n         Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n         Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n     \n    + ## Documentation/config/core.txt ##\n    +@@ Documentation/config/core.txt: usage, at the slight expense of increased disk usage.\n    + * Will be generally be streamed when written, which avoids excessive\n    + memory usage, at the cost of some fixed overhead. Commands that make\n    + use of this include linkgit:git-archive[1],\n    +-linkgit:git-fast-import[1], linkgit:git-index-pack[1] and\n    +-linkgit:git-fsck[1].\n    ++linkgit:git-fast-import[1], linkgit:git-index-pack[1],\n    ++linkgit:git-unpack-objects[1] and linkgit:git-fsck[1].\n    + \n    + core.excludesFile::\n    + \tSpecifies the pathname to the file that contains patterns to\n    +\n      ## builtin/unpack-objects.c ##\n     @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type type,\n      \t}\n5:  1bfaf89ee0b < -:  ----------- object-file API: add a format_object_header() function\n-- \n2.35.1.940.ge7a5b4b05f2\n\n"},{"id":"447740","messageId":"patch-v10-1.6-e46eb75b98f-20220204T135538Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v10-0.6-00000000000-20220204T135538Z-avarab@gmail.com","subject":"[PATCH v10 1/6] unpack-objects: low memory footprint for get_data() in dry_run mode","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-02-04T14:07:07Z","receivedAt":"2022-02-04T14:07:26Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nAs the name implies, \"get_data(size)\" will allocate and return a given\namount of memory. Allocating memory for a large blob object may cause the\nsystem to run out of memory. Before preparing to replace calling of\n\"get_data()\" to unpack large blob objects in latter commits, refactor\n\"get_data()\" to reduce memory footprint for dry_run mode.\n\nBecause in dry_run mode, \"get_data()\" is only used to check the\nintegrity of data, and the returned buffer is not used at all, we can\nallocate a smaller buffer and reuse it as zstream output. Therefore,\nin dry_run mode, \"get_data()\" will release the allocated buffer and\nreturn NULL instead of returning garbage data.\n\nThe \"find [...]objects/?? -type f | wc -l\" test idiom being used here\nis adapted from the same \"find\" use added to another test in\nd9545c7f465 (fast-import: implement unpack limit, 2016-04-25).\n\nSuggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n builtin/unpack-objects.c        | 39 ++++++++++++++++++++--------\n t/t5328-unpack-large-objects.sh | 45 +++++++++++++++++++++++++++++++++\n 2 files changed, 73 insertions(+), 11 deletions(-)\n create mode 100755 t/t5328-unpack-large-objects.sh\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex dbeb0680a58..896ea8aceb4 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -96,15 +96,31 @@ static void use(int bytes)\n \tdisplay_throughput(progress, consumed_bytes);\n }\n \n+/*\n+ * Decompress zstream from stdin and return specific size of data.\n+ * The caller is responsible to free the returned buffer.\n+ *\n+ * But for dry_run mode, \"get_data()\" is only used to check the\n+ * integrity of data, and the returned buffer is not used at all.\n+ * Therefore, in dry_run mode, \"get_data()\" will release the small\n+ * allocated buffer which is reused to hold temporary zstream output\n+ * and return NULL instead of returning garbage data.\n+ */\n static void *get_data(unsigned long size)\n {\n \tgit_zstream stream;\n-\tvoid *buf = xmallocz(size);\n+\tunsigned long bufsize;\n+\tvoid *buf;\n \n \tmemset(&stream, 0, sizeof(stream));\n+\tif (dry_run && size > 8192)\n+\t\tbufsize = 8192;\n+\telse\n+\t\tbufsize = size;\n+\tbuf = xmallocz(bufsize);\n \n \tstream.next_out = buf;\n-\tstream.avail_out = size;\n+\tstream.avail_out = bufsize;\n \tstream.next_in = fill(1);\n \tstream.avail_in = len;\n \tgit_inflate_init(&stream);\n@@ -124,8 +140,15 @@ static void *get_data(unsigned long size)\n \t\t}\n \t\tstream.next_in = fill(1);\n \t\tstream.avail_in = len;\n+\t\tif (dry_run) {\n+\t\t\t/* reuse the buffer in dry_run mode */\n+\t\t\tstream.next_out = buf;\n+\t\t\tstream.avail_out = bufsize;\n+\t\t}\n \t}\n \tgit_inflate_end(&stream);\n+\tif (dry_run)\n+\t\tFREE_AND_NULL(buf);\n \treturn buf;\n }\n \n@@ -325,10 +348,8 @@ static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n {\n \tvoid *buf = get_data(size);\n \n-\tif (!dry_run && buf)\n+\tif (buf)\n \t\twrite_object(nr, type, buf, size);\n-\telse\n-\t\tfree(buf);\n }\n \n static int resolve_against_held(unsigned nr, const struct object_id *base,\n@@ -358,10 +379,8 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\toidread(&base_oid, fill(the_hash_algo->rawsz));\n \t\tuse(the_hash_algo->rawsz);\n \t\tdelta_data = get_data(delta_size);\n-\t\tif (dry_run || !delta_data) {\n-\t\t\tfree(delta_data);\n+\t\tif (!delta_data)\n \t\t\treturn;\n-\t\t}\n \t\tif (has_object_file(&base_oid))\n \t\t\t; /* Ok we have this one */\n \t\telse if (resolve_against_held(nr, &base_oid,\n@@ -397,10 +416,8 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\t\tdie(\"offset value out of bound for delta base object\");\n \n \t\tdelta_data = get_data(delta_size);\n-\t\tif (dry_run || !delta_data) {\n-\t\t\tfree(delta_data);\n+\t\tif (!delta_data)\n \t\t\treturn;\n-\t\t}\n \t\tlo = 0;\n \t\thi = nr;\n \t\twhile (lo < hi) {\ndiff --git a/t/t5328-unpack-large-objects.sh b/t/t5328-unpack-large-objects.sh\nnew file mode 100755\nindex 00000000000..1432dfc8386\n--- /dev/null\n+++ b/t/t5328-unpack-large-objects.sh\n@@ -0,0 +1,45 @@\n+#!/bin/sh\n+#\n+# Copyright (c) 2022 Han Xin\n+#\n+\n+test_description='git unpack-objects with large objects'\n+\n+. ./test-lib.sh\n+\n+prepare_dest () {\n+\ttest_when_finished \"rm -rf dest.git\" &&\n+\tgit init --bare dest.git\n+}\n+\n+test_no_loose () {\n+\ttest $(find dest.git/objects/?? -type f | wc -l) = 0\n+}\n+\n+test_expect_success \"create large objects (1.5 MB) and PACK\" '\n+\ttest-tool genrandom foo 1500000 >big-blob &&\n+\ttest_commit --append foo big-blob &&\n+\ttest-tool genrandom bar 1500000 >big-blob &&\n+\ttest_commit --append bar big-blob &&\n+\tPACK=$(echo HEAD | git pack-objects --revs test)\n+'\n+\n+test_expect_success 'set memory limitation to 1MB' '\n+\tGIT_ALLOC_LIMIT=1m &&\n+\texport GIT_ALLOC_LIMIT\n+'\n+\n+test_expect_success 'unpack-objects failed under memory limitation' '\n+\tprepare_dest &&\n+\ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n+\tgrep \"fatal: attempting to allocate\" err\n+'\n+\n+test_expect_success 'unpack-objects works with memory limitation in dry-run mode' '\n+\tprepare_dest &&\n+\tgit -C dest.git unpack-objects -n <test-$PACK.pack &&\n+\ttest_no_loose &&\n+\ttest_dir_is_empty dest.git/objects/pack\n+'\n+\n+test_done\n-- \n2.35.1.940.ge7a5b4b05f2\n\n"},{"id":"447741","messageId":"patch-v10-2.6-48bf9090058-20220204T135538Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v10-0.6-00000000000-20220204T135538Z-avarab@gmail.com","subject":"[PATCH v10 2/6] object-file.c: do fsync() and close() before post-write die()","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-02-04T14:07:08Z","receivedAt":"2022-02-04T14:07:30Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Change write_loose_object() to do an fsync() and close() before the\noideq() sanity check at the end. This change re-joins code that was\nsplit up by the die() sanity check added in 748af44c63e (sha1_file: be\nparanoid when creating loose objects, 2010-02-21).\n\nI don't think that this change matters in itself, if we called die()\nit was possible that our data wouldn't fully make it to disk, but in\nany case we were writing data that we'd consider corrupted. It's\npossible that a subsequent \"git fsck\" will be less confused now.\n\nThe real reason to make this change is that in a subsequent commit\nwe'll split this code in write_loose_object() into a utility function,\nall its callers will want the preceding sanity checks, but not the\n\"oideq\" check. By moving the close_loose_object() earlier it'll be\neasier to reason about the introduction of the utility function.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n object-file.c | 13 +++++++++++--\n 1 file changed, 11 insertions(+), 2 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 5c9525479c2..edebdc91221 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2001,12 +2001,21 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n \t\t    ret);\n \tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n+\n+\t/*\n+\t * We already did a write_buffer() to the \"fd\", let's fsync()\n+\t * and close().\n+\t *\n+\t * We might still die() on a subsequent sanity check, but\n+\t * let's not add to that confusion by not flushing any\n+\t * outstanding writes to disk first.\n+\t */\n+\tclose_loose_object(fd);\n+\n \tif (!oideq(oid, &parano_oid))\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n \n-\tclose_loose_object(fd);\n-\n \tif (mtime) {\n \t\tstruct utimbuf utb;\n \t\tutb.actime = mtime;\n-- \n2.35.1.940.ge7a5b4b05f2\n\n"},{"id":"447742","messageId":"patch-v10-4.6-9644df5c744-20220204T135538Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v10-0.6-00000000000-20220204T135538Z-avarab@gmail.com","subject":"[PATCH v10 4/6] object-file.c: add \"stream_loose_object()\" to handle large object","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-02-04T14:07:10Z","receivedAt":"2022-02-04T14:07:32Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nIf we want unpack and write a loose object using \"write_loose_object\",\nwe have to feed it with a buffer with the same size of the object, which\nwill consume lots of memory and may cause OOM. This can be improved by\nfeeding data to \"stream_loose_object()\" in a stream.\n\nAdd a new function \"stream_loose_object()\", which is a stream version of\n\"write_loose_object()\" but with a low memory footprint. We will use this\nfunction to unpack large blob object in later commit.\n\nAnother difference with \"write_loose_object()\" is that we have no chance\nto run \"write_object_file_prepare()\" to calculate the oid in advance.\nIn \"write_loose_object()\", we know the oid and we can write the\ntemporary file in the same directory as the final object, but for an\nobject with an undetermined oid, we don't know the exact directory for\nthe object.\n\nStill, we need to save the temporary file we're preparing\nsomewhere. We'll do that in the top-level \".git/objects/\"\ndirectory (or whatever \"GIT_OBJECT_DIRECTORY\" is set to). Once we've\nstreamed it we'll know the OID, and will move it to its canonical\npath.\n\n\"freshen_packed_object()\" or \"freshen_loose_object()\" will be called\ninside \"stream_loose_object()\" after obtaining the \"oid\".\n\nHelped-by: René Scharfe <l.s.r@web.de>\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n object-file.c  | 99 ++++++++++++++++++++++++++++++++++++++++++++++++++\n object-store.h |  8 ++++\n 2 files changed, 107 insertions(+)\n\ndiff --git a/object-file.c b/object-file.c\nindex f5c579e42e3..2ef0bf0e5c3 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2095,6 +2095,105 @@ static int freshen_packed_object(const struct object_id *oid)\n \treturn 1;\n }\n \n+int stream_loose_object(struct input_stream *in_stream, size_t len,\n+\t\t\tstruct object_id *oid)\n+{\n+\tint fd, ret, err = 0, flush = 0;\n+\tunsigned char compressed[4096];\n+\tgit_zstream stream;\n+\tgit_hash_ctx c;\n+\tstruct strbuf tmp_file = STRBUF_INIT;\n+\tstruct strbuf filename = STRBUF_INIT;\n+\tint dirlen;\n+\tchar hdr[MAX_HEADER_LEN];\n+\tint hdrlen;\n+\n+\t/* Since oid is not determined, save tmp file to odb path. */\n+\tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n+\thdrlen = format_object_header(hdr, sizeof(hdr), OBJ_BLOB, len);\n+\n+\t/* Common steps for write_loose_object and stream_loose_object to\n+\t * start writing loose oject:\n+\t *\n+\t *  - Create tmpfile for the loose object.\n+\t *  - Setup zlib stream for compression.\n+\t *  - Start to feed header to zlib stream.\n+\t */\n+\tfd = start_loose_object_common(&tmp_file, filename.buf, 0,\n+\t\t\t\t       &stream, compressed, sizeof(compressed),\n+\t\t\t\t       &c, hdr, hdrlen);\n+\tif (fd < 0) {\n+\t\terr = -1;\n+\t\tgoto cleanup;\n+\t}\n+\n+\t/* Then the data itself.. */\n+\tdo {\n+\t\tunsigned char *in0 = stream.next_in;\n+\t\tif (!stream.avail_in && !in_stream->is_finished) {\n+\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n+\t\t\tstream.next_in = (void *)in;\n+\t\t\tin0 = (unsigned char *)in;\n+\t\t\t/* All data has been read. */\n+\t\t\tif (in_stream->is_finished)\n+\t\t\t\tflush = Z_FINISH;\n+\t\t}\n+\t\tret = git_deflate(&stream, flush);\n+\t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n+\t\tif (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n+\t\t\tdie(_(\"unable to write loose object file\"));\n+\t\tstream.next_out = compressed;\n+\t\tstream.avail_out = sizeof(compressed);\n+\t\t/*\n+\t\t * Unlike write_loose_object(), we do not have the entire\n+\t\t * buffer. If we get Z_BUF_ERROR due to too few input bytes,\n+\t\t * then we'll replenish them in the next input_stream->read()\n+\t\t * call when we loop.\n+\t\t */\n+\t} while (ret == Z_OK || ret == Z_BUF_ERROR);\n+\n+\tif (stream.total_in != len + hdrlen)\n+\t\tdie(_(\"write stream object %ld != %\"PRIuMAX), stream.total_in,\n+\t\t    (uintmax_t)len + hdrlen);\n+\n+\t/* Common steps for write_loose_object and stream_loose_object to\n+\t * end writing loose oject:\n+\t *\n+\t *  - End the compression of zlib stream.\n+\t *  - Get the calculated oid.\n+\t */\n+\tend_loose_object_common(fd, ret, &c, &stream, oid, NULL,\n+\t\t\t\tN_(\"unable to stream deflate new object (%d)\"),\n+\t\t\t\tN_(\"deflateEnd on stream object failed (%d)\"));\n+\n+\tif (freshen_packed_object(oid) || freshen_loose_object(oid)) {\n+\t\tunlink_or_warn(tmp_file.buf);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tloose_object_path(the_repository, &filename, oid);\n+\n+\t/* We finally know the object path, and create the missing dir. */\n+\tdirlen = directory_size(filename.buf);\n+\tif (dirlen) {\n+\t\tstruct strbuf dir = STRBUF_INIT;\n+\t\tstrbuf_add(&dir, filename.buf, dirlen);\n+\n+\t\tif (mkdir_in_gitdir(dir.buf) && errno != EEXIST) {\n+\t\t\terr = error_errno(_(\"unable to create directory %s\"), dir.buf);\n+\t\t\tstrbuf_release(&dir);\n+\t\t\tgoto cleanup;\n+\t\t}\n+\t\tstrbuf_release(&dir);\n+\t}\n+\n+\terr = finalize_object_file(tmp_file.buf, filename.buf);\n+cleanup:\n+\tstrbuf_release(&tmp_file);\n+\tstrbuf_release(&filename);\n+\treturn err;\n+}\n+\n int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t    enum object_type type, struct object_id *oid,\n \t\t\t    unsigned flags)\ndiff --git a/object-store.h b/object-store.h\nindex bd2322ed8ce..1099455bc2e 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -46,6 +46,12 @@ struct object_directory {\n \tchar *path;\n };\n \n+struct input_stream {\n+\tconst void *(*read)(struct input_stream *, unsigned long *len);\n+\tvoid *data;\n+\tint is_finished;\n+};\n+\n KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n \tstruct object_directory *, 1, fspathhash, fspatheq)\n \n@@ -261,6 +267,8 @@ static inline int write_object_file(const void *buf, unsigned long len,\n int write_object_file_literally(const void *buf, unsigned long len,\n \t\t\t\tconst char *type, struct object_id *oid,\n \t\t\t\tunsigned flags);\n+int stream_loose_object(struct input_stream *in_stream, size_t len,\n+\t\t\tstruct object_id *oid);\n \n /*\n  * Add an object file to the in-memory object store, without writing it\n-- \n2.35.1.940.ge7a5b4b05f2\n\n"},{"id":"447743","messageId":"patch-v10-3.6-0e33d2a6e35-20220204T135538Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v10-0.6-00000000000-20220204T135538Z-avarab@gmail.com","subject":"[PATCH v10 3/6] object-file.c: refactor write_loose_object() to several steps","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-02-04T14:07:09Z","receivedAt":"2022-02-04T14:07:33Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWhen writing a large blob using \"write_loose_object()\", we have to pass\na buffer with the whole content of the blob, and this behavior will\nconsume lots of memory and may cause OOM. We will introduce a stream\nversion function (\"stream_loose_object()\") in later commit to resolve\nthis issue.\n\nBefore introducing that streaming function, do some refactoring on\n\"write_loose_object()\" to reuse code for both versions.\n\nRewrite \"write_loose_object()\" as follows:\n\n 1. Figure out a path for the (temp) object file. This step is only\n    used in \"write_loose_object()\".\n\n 2. Move common steps for starting to write loose objects into a new\n    function \"start_loose_object_common()\".\n\n 3. Compress data.\n\n 4. Move common steps for ending zlib stream into a new function\n    \"end_loose_object_common()\".\n\n 5. Close fd and finalize the object file.\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n object-file.c | 129 ++++++++++++++++++++++++++++++++++----------------\n 1 file changed, 89 insertions(+), 40 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex edebdc91221..f5c579e42e3 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1943,6 +1943,87 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n \treturn fd;\n }\n \n+/**\n+ * Common steps for loose object writers to start writing loose\n+ * objects:\n+ *\n+ * - Create tmpfile for the loose object.\n+ * - Setup zlib stream for compression.\n+ * - Start to feed header to zlib stream.\n+ *\n+ * Returns a \"fd\", which should later be provided to\n+ * end_loose_object_common().\n+ */\n+static int start_loose_object_common(struct strbuf *tmp_file,\n+\t\t\t\t     const char *filename, unsigned flags,\n+\t\t\t\t     git_zstream *stream,\n+\t\t\t\t     unsigned char *buf, size_t buflen,\n+\t\t\t\t     git_hash_ctx *c,\n+\t\t\t\t     char *hdr, int hdrlen)\n+{\n+\tint fd;\n+\n+\tfd = create_tmpfile(tmp_file, filename);\n+\tif (fd < 0) {\n+\t\tif (flags & HASH_SILENT)\n+\t\t\treturn -1;\n+\t\telse if (errno == EACCES)\n+\t\t\treturn error(_(\"insufficient permission for adding \"\n+\t\t\t\t       \"an object to repository database %s\"),\n+\t\t\t\t     get_object_directory());\n+\t\telse\n+\t\t\treturn error_errno(\n+\t\t\t\t_(\"unable to create temporary file\"));\n+\t}\n+\n+\t/*  Setup zlib stream for compression */\n+\tgit_deflate_init(stream, zlib_compression_level);\n+\tstream->next_out = buf;\n+\tstream->avail_out = buflen;\n+\tthe_hash_algo->init_fn(c);\n+\n+\t/*  Start to feed header to zlib stream */\n+\tstream->next_in = (unsigned char *)hdr;\n+\tstream->avail_in = hdrlen;\n+\twhile (git_deflate(stream, 0) == Z_OK)\n+\t\t; /* nothing */\n+\tthe_hash_algo->update_fn(c, hdr, hdrlen);\n+\n+\treturn fd;\n+}\n+\n+/**\n+ * Common steps for loose object writers to end writing loose objects:\n+ *\n+ * - End the compression of zlib stream.\n+ * - Get the calculated oid to \"parano_oid\".\n+ * - fsync() and close() the \"fd\"\n+ */\n+static void end_loose_object_common(int fd, int ret, git_hash_ctx *c,\n+\t\t\t\t    git_zstream *stream,\n+\t\t\t\t    struct object_id *parano_oid,\n+\t\t\t\t    const struct object_id *expected_oid,\n+\t\t\t\t    const char *die_msg1_fmt,\n+\t\t\t\t    const char *die_msg2_fmt)\n+{\n+\tif (ret != Z_STREAM_END)\n+\t\tdie(_(die_msg1_fmt), ret, expected_oid);\n+\tret = git_deflate_end_gently(stream);\n+\tif (ret != Z_OK)\n+\t\tdie(_(die_msg2_fmt), ret, expected_oid);\n+\tthe_hash_algo->final_oid_fn(parano_oid, c);\n+\n+\t/*\n+\t * We already did a write_buffer() to the \"fd\", let's fsync()\n+\t * and close().\n+\t *\n+\t * We might still die() on a subsequent sanity check, but\n+\t * let's not add to that confusion by not flushing any\n+\t * outstanding writes to disk first.\n+\t */\n+\tclose_loose_object(fd);\n+}\n+\n static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\t\t      int hdrlen, const void *buf, unsigned long len,\n \t\t\t      time_t mtime, unsigned flags)\n@@ -1957,28 +2038,11 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \n \tloose_object_path(the_repository, &filename, oid);\n \n-\tfd = create_tmpfile(&tmp_file, filename.buf);\n-\tif (fd < 0) {\n-\t\tif (flags & HASH_SILENT)\n-\t\t\treturn -1;\n-\t\telse if (errno == EACCES)\n-\t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n-\t\telse\n-\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n-\t}\n-\n-\t/* Set it up */\n-\tgit_deflate_init(&stream, zlib_compression_level);\n-\tstream.next_out = compressed;\n-\tstream.avail_out = sizeof(compressed);\n-\tthe_hash_algo->init_fn(&c);\n-\n-\t/* First header.. */\n-\tstream.next_in = (unsigned char *)hdr;\n-\tstream.avail_in = hdrlen;\n-\twhile (git_deflate(&stream, 0) == Z_OK)\n-\t\t; /* nothing */\n-\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n+\tfd = start_loose_object_common(&tmp_file, filename.buf, flags,\n+\t\t\t\t       &stream, compressed, sizeof(compressed),\n+\t\t\t\t       &c, hdr, hdrlen);\n+\tif (fd < 0)\n+\t\treturn -1;\n \n \t/* Then the data itself.. */\n \tstream.next_in = (void *)buf;\n@@ -1993,24 +2057,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tstream.avail_out = sizeof(compressed);\n \t} while (ret == Z_OK);\n \n-\tif (ret != Z_STREAM_END)\n-\t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n-\t\t    ret);\n-\tret = git_deflate_end_gently(&stream);\n-\tif (ret != Z_OK)\n-\t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n-\t\t    ret);\n-\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n-\n-\t/*\n-\t * We already did a write_buffer() to the \"fd\", let's fsync()\n-\t * and close().\n-\t *\n-\t * We might still die() on a subsequent sanity check, but\n-\t * let's not add to that confusion by not flushing any\n-\t * outstanding writes to disk first.\n-\t */\n-\tclose_loose_object(fd);\n+\tend_loose_object_common(fd, ret, &c, &stream, &parano_oid, oid,\n+\t\t\t\tN_(\"unable to deflate new object %s (%d)\"),\n+\t\t\t\tN_(\"deflateEnd on object %s failed (%d)\"));\n \n \tif (!oideq(oid, &parano_oid))\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n-- \n2.35.1.940.ge7a5b4b05f2\n\n"},{"id":"447744","messageId":"patch-v10-5.6-4550f3a2745-20220204T135538Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v10-0.6-00000000000-20220204T135538Z-avarab@gmail.com","subject":"[PATCH v10 5/6] core doc: modernize core.bigFileThreshold documentation","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-02-04T14:07:11Z","receivedAt":"2022-02-04T14:07:35Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"The core.bigFileThreshold documentation has been largely unchanged\nsince 5eef828bc03 (fast-import: Stream very large blobs directly to\npack, 2010-02-01).\n\nBut since then this setting has been expanded to affect a lot more\nthan that description indicated. Most notably in how \"git diff\" treats\nthem, see 6bf3b813486 (diff --stat: mark any file larger than\ncore.bigfilethreshold binary, 2014-08-16).\n\nIn addition to that, numerous commands and APIs make use of a\nstreaming mode for files above this threshold.\n\nSo let's attempt to summarize 12 years of changes in behavior, which\ncan be seen with:\n\n    git log --oneline -Gbig_file_thre 5eef828bc03.. -- '*.c'\n\nTo do that turn this into a bullet-point list. The summary Han Xin\nproduced in [1] helped a lot, but is a bit too detailed for\ndocumentation aimed at users. Let's instead summarize how\nuser-observable behavior differs, and generally describe how we tend\nto stream these files in various commands.\n\n1. https://lore.kernel.org/git/20220120112114.47618-5-chiyutianyi@gmail.com/\n\nHelped-by: Han Xin <chiyutianyi@gmail.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n Documentation/config/core.txt | 33 ++++++++++++++++++++++++---------\n 1 file changed, 24 insertions(+), 9 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex c04f62a54a1..b6a12218665 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -412,17 +412,32 @@ You probably do not need to adjust this value.\n Common unit suffixes of 'k', 'm', or 'g' are supported.\n \n core.bigFileThreshold::\n-\tFiles larger than this size are stored deflated, without\n-\tattempting delta compression.  Storing large files without\n-\tdelta compression avoids excessive memory usage, at the\n-\tslight expense of increased disk usage. Additionally files\n-\tlarger than this size are always treated as binary.\n+\tThe size of files considered \"big\", which as discussed below\n+\tchanges the behavior of numerous git commands, as well as how\n+\tsuch files are stored within the repository. The default is\n+\t512 MiB. Common unit suffixes of 'k', 'm', or 'g' are\n+\tsupported.\n +\n-Default is 512 MiB on all platforms.  This should be reasonable\n-for most projects as source code and other text files can still\n-be delta compressed, but larger binary media files won't be.\n+Files above the configured limit will be:\n +\n-Common unit suffixes of 'k', 'm', or 'g' are supported.\n+* Stored deflated, without attempting delta compression.\n++\n+The default limit is primarily set with this use-case in mind. With it\n+most projects will have their source code and other text files delta\n+compressed, but not larger binary media files.\n++\n+Storing large files without delta compression avoids excessive memory\n+usage, at the slight expense of increased disk usage.\n++\n+* Will be treated as if though they were labeled \"binary\" (see\n+  linkgit:gitattributes[5]). This means that e.g. linkgit:git-log[1]\n+  and linkgit:git-diff[1] will not diffs for files above this limit.\n++\n+* Will be generally be streamed when written, which avoids excessive\n+memory usage, at the cost of some fixed overhead. Commands that make\n+use of this include linkgit:git-archive[1],\n+linkgit:git-fast-import[1], linkgit:git-index-pack[1] and\n+linkgit:git-fsck[1].\n \n core.excludesFile::\n \tSpecifies the pathname to the file that contains patterns to\n-- \n2.35.1.940.ge7a5b4b05f2\n\n"},{"id":"447745","messageId":"patch-v10-6.6-6a70e49a346-20220204T135538Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v10-0.6-00000000000-20220204T135538Z-avarab@gmail.com","subject":"[PATCH v10 6/6] unpack-objects: use stream_loose_object() to unpack large objects","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-02-04T14:07:12Z","receivedAt":"2022-02-04T14:07:36Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nMake use of the stream_loose_object() function introduced in the\npreceding commit to unpack large objects. Before this we'd need to\nmalloc() the size of the blob before unpacking it, which could cause\nOOM with very large blobs.\n\nWe could use this new interface to unpack all blobs, but doing so\nwould result in a performance penalty of around 10%, as the below\n\"hyperfine\" benchmark will show. We therefore limit this to files\nlarger than \"core.bigFileThreshold\":\n\n    $ hyperfine \\\n      --setup \\\n      'if ! test -d scalar.git; then git clone --bare\n       https://github.com/microsoft/scalar.git;\n       cp scalar.git/objects/pack/*.pack small.pack; fi' \\\n      --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n      ...\n\n    Summary\n      './git -C dest.git -c core.bigFileThreshold=512m\n      unpack-objects <small.pack' in 'origin/master'\n        1.01 ± 0.04 times faster than './git -C dest.git\n                -c core.bigFileThreshold=512m unpack-objects\n                <small.pack' in 'HEAD~1'\n        1.01 ± 0.04 times faster than './git -C dest.git\n                -c core.bigFileThreshold=512m unpack-objects\n                <small.pack' in 'HEAD~0'\n        1.03 ± 0.10 times faster than './git -C dest.git\n                -c core.bigFileThreshold=16k unpack-objects\n                <small.pack' in 'origin/master'\n        1.02 ± 0.07 times faster than './git -C dest.git\n                -c core.bigFileThreshold=16k unpack-objects\n                <small.pack' in 'HEAD~0'\n        1.10 ± 0.04 times faster than './git -C dest.git\n                -c core.bigFileThreshold=16k unpack-objects\n                <small.pack' in 'HEAD~1'\n\nAn earlier version of this patch introduced a new\n\"core.bigFileStreamingThreshold\" instead of re-using the existing\n\"core.bigFileThreshold\" variable[1]. As noted in a detailed overview\nof its users in [2] using it has several different meanings.\n\nStill, we consider it good enough to simply re-use it. While it's\npossible that someone might want to e.g. consider objects \"small\" for\nthe purposes of diffing but \"big\" for the purposes of writing them\nsuch use-cases are probably too obscure to worry about. We can always\nsplit up \"core.bigFileThreshold\" in the future if there's a need for\nthat.\n\n1. https://lore.kernel.org/git/20211210103435.83656-1-chiyutianyi@gmail.com/\n2. https://lore.kernel.org/git/20220120112114.47618-5-chiyutianyi@gmail.com/\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Derrick Stolee <stolee@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n---\n Documentation/config/core.txt   |  4 +-\n builtin/unpack-objects.c        | 71 ++++++++++++++++++++++++++++++++-\n t/t5328-unpack-large-objects.sh | 23 +++++++++--\n 3 files changed, 92 insertions(+), 6 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex b6a12218665..5aca987632c 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -436,8 +436,8 @@ usage, at the slight expense of increased disk usage.\n * Will be generally be streamed when written, which avoids excessive\n memory usage, at the cost of some fixed overhead. Commands that make\n use of this include linkgit:git-archive[1],\n-linkgit:git-fast-import[1], linkgit:git-index-pack[1] and\n-linkgit:git-fsck[1].\n+linkgit:git-fast-import[1], linkgit:git-index-pack[1],\n+linkgit:git-unpack-objects[1] and linkgit:git-fsck[1].\n \n core.excludesFile::\n \tSpecifies the pathname to the file that contains patterns to\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 896ea8aceb4..7ce3cb61086 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -343,11 +343,80 @@ static void added_object(unsigned nr, enum object_type type,\n \t}\n }\n \n+struct input_zstream_data {\n+\tgit_zstream *zstream;\n+\tunsigned char buf[8192];\n+\tint status;\n+};\n+\n+static const void *feed_input_zstream(struct input_stream *in_stream,\n+\t\t\t\t      unsigned long *readlen)\n+{\n+\tstruct input_zstream_data *data = in_stream->data;\n+\tgit_zstream *zstream = data->zstream;\n+\tvoid *in = fill(1);\n+\n+\tif (in_stream->is_finished) {\n+\t\t*readlen = 0;\n+\t\treturn NULL;\n+\t}\n+\n+\tzstream->next_out = data->buf;\n+\tzstream->avail_out = sizeof(data->buf);\n+\tzstream->next_in = in;\n+\tzstream->avail_in = len;\n+\n+\tdata->status = git_inflate(zstream, 0);\n+\n+\tin_stream->is_finished = data->status != Z_OK;\n+\tuse(len - zstream->avail_in);\n+\t*readlen = sizeof(data->buf) - zstream->avail_out;\n+\n+\treturn data->buf;\n+}\n+\n+static void write_stream_blob(unsigned nr, size_t size)\n+{\n+\tgit_zstream zstream = { 0 };\n+\tstruct input_zstream_data data = { 0 };\n+\tstruct input_stream in_stream = {\n+\t\t.read = feed_input_zstream,\n+\t\t.data = &data,\n+\t};\n+\n+\tdata.zstream = &zstream;\n+\tgit_inflate_init(&zstream);\n+\n+\tif (stream_loose_object(&in_stream, size, &obj_list[nr].oid))\n+\t\tdie(_(\"failed to write object in stream\"));\n+\n+\tif (data.status != Z_STREAM_END)\n+\t\tdie(_(\"inflate returned (%d)\"), data.status);\n+\tgit_inflate_end(&zstream);\n+\n+\tif (strict) {\n+\t\tstruct blob *blob =\n+\t\t\tlookup_blob(the_repository, &obj_list[nr].oid);\n+\t\tif (blob)\n+\t\t\tblob->object.flags |= FLAG_WRITTEN;\n+\t\telse\n+\t\t\tdie(_(\"invalid blob object from stream\"));\n+\t}\n+\tobj_list[nr].obj = NULL;\n+}\n+\n static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n \t\t\t\t   unsigned nr)\n {\n-\tvoid *buf = get_data(size);\n+\tvoid *buf;\n+\n+\t/* Write large blob in stream without allocating full buffer. */\n+\tif (!dry_run && type == OBJ_BLOB && size > big_file_threshold) {\n+\t\twrite_stream_blob(nr, size);\n+\t\treturn;\n+\t}\n \n+\tbuf = get_data(size);\n \tif (buf)\n \t\twrite_object(nr, type, buf, size);\n }\ndiff --git a/t/t5328-unpack-large-objects.sh b/t/t5328-unpack-large-objects.sh\nindex 1432dfc8386..5c1042b4d91 100755\n--- a/t/t5328-unpack-large-objects.sh\n+++ b/t/t5328-unpack-large-objects.sh\n@@ -9,7 +9,11 @@ test_description='git unpack-objects with large objects'\n \n prepare_dest () {\n \ttest_when_finished \"rm -rf dest.git\" &&\n-\tgit init --bare dest.git\n+\tgit init --bare dest.git &&\n+\tif test -n \"$1\"\n+\tthen\n+\t\tgit -C dest.git config core.bigFileThreshold $1\n+\tfi\n }\n \n test_no_loose () {\n@@ -30,16 +34,29 @@ test_expect_success 'set memory limitation to 1MB' '\n '\n \n test_expect_success 'unpack-objects failed under memory limitation' '\n-\tprepare_dest &&\n+\tprepare_dest 2m &&\n \ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n \tgrep \"fatal: attempting to allocate\" err\n '\n \n test_expect_success 'unpack-objects works with memory limitation in dry-run mode' '\n-\tprepare_dest &&\n+\tprepare_dest 2m &&\n \tgit -C dest.git unpack-objects -n <test-$PACK.pack &&\n \ttest_no_loose &&\n \ttest_dir_is_empty dest.git/objects/pack\n '\n \n+test_expect_success 'unpack big object in stream' '\n+\tprepare_dest 1m &&\n+\tgit -C dest.git unpack-objects <test-$PACK.pack &&\n+\ttest_dir_is_empty dest.git/objects/pack\n+'\n+\n+test_expect_success 'do not unpack existing large objects' '\n+\tprepare_dest 1m &&\n+\tgit -C dest.git index-pack --stdin <test-$PACK.pack &&\n+\tgit -C dest.git unpack-objects <test-$PACK.pack &&\n+\ttest_no_loose\n+'\n+\n test_done\n-- \n2.35.1.940.ge7a5b4b05f2\n\n"},{"id":"451648","messageId":"cover-v11-0.8-00000000000-20220319T001411Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v10-0.6-00000000000-20220204T135538Z-avarab@gmail.com","subject":"[PATCH v11 0/8] unpack-objects: support streaming blobs to disk","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-19T00:23:17Z","receivedAt":"2022-03-19T00:23:47Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"This series by Han Xin was waiting on some in-flight patches that\nlanded in 430883a70c7 (Merge branch 'ab/object-file-api-updates',\n2022-03-16).\n\nThis series teaches \"git unpack-objects\" to stream objects larger than\ncore.bigFileThreshold to disk. As 8/8 shows streaming e.g. a 100MB\nblob now uses ~5MB of memory instead of ~105MB. This streaming method\nis slower if you've got memory to handle the blobs in-core, but if you\ndon't it allows you to unpack objects at all, as you might otherwise\nOOM.\n\nChanges since v10:\n\n * Renamed the new test file, its number conflicted with a\n   since-landed commit-graph test.\n\n * Some minor code changes to make diffs to the pre-image smaller\n   (e.g. the top of the range-diff below)\n\n * The whole \"find dest.git\" to see if we have loose objects is now\n   either a test for \"do we have objects at all?\" (--dry-run mode), or\n   uses a simpler implementation. We could use\n   \"test_stdout_line_count\" for that.\n\n * We also test that as we use \"unpack-objects\" to stream directly to\n   a pack that the result is byte-for-byte the same as the source.\n\n * A new 4/8 that I added allows for more code sharing in\n   object-file.c, our two end-state functions now share more logic.\n\n * Minor typo/grammar/comment etc. fixes throughout.\n\n * Updated 8/8 with benchmarks, somewhere along the line we lost the\n   code to run the benchmark mentioned in the commit message...\n\n1. https://lore.kernel.org/git/cover-v10-0.6-00000000000-20220204T135538Z-avarab@gmail.com/\n\nHan Xin (4):\n  unpack-objects: low memory footprint for get_data() in dry_run mode\n  object-file.c: refactor write_loose_object() to several steps\n  object-file.c: add \"stream_loose_object()\" to handle large object\n  unpack-objects: use stream_loose_object() to unpack large objects\n\nÆvar Arnfjörð Bjarmason (4):\n  object-file.c: do fsync() and close() before post-write die()\n  object-file.c: factor out deflate part of write_loose_object()\n  core doc: modernize core.bigFileThreshold documentation\n  unpack-objects: refactor away unpack_non_delta_entry()\n\n Documentation/config/core.txt   |  33 +++--\n builtin/unpack-objects.c        | 109 +++++++++++---\n object-file.c                   | 250 +++++++++++++++++++++++++++-----\n object-store.h                  |   8 +\n t/t5351-unpack-large-objects.sh |  61 ++++++++\n 5 files changed, 397 insertions(+), 64 deletions(-)\n create mode 100755 t/t5351-unpack-large-objects.sh\n\nRange-diff against v10:\n1:  e46eb75b98f ! 1:  2103d5bfd96 unpack-objects: low memory footprint for get_data() in dry_run mode\n    @@ builtin/unpack-objects.c: static void use(int bytes)\n      {\n      \tgit_zstream stream;\n     -\tvoid *buf = xmallocz(size);\n    -+\tunsigned long bufsize;\n    -+\tvoid *buf;\n    ++\tunsigned long bufsize = dry_run && size > 8192 ? 8192 : size;\n    ++\tvoid *buf = xmallocz(bufsize);\n      \n      \tmemset(&stream, 0, sizeof(stream));\n    -+\tif (dry_run && size > 8192)\n    -+\t\tbufsize = 8192;\n    -+\telse\n    -+\t\tbufsize = size;\n    -+\tbuf = xmallocz(bufsize);\n      \n      \tstream.next_out = buf;\n     -\tstream.avail_out = size;\n    @@ builtin/unpack-objects.c: static void unpack_delta_entry(enum object_type type,\n      \t\thi = nr;\n      \t\twhile (lo < hi) {\n     \n    - ## t/t5328-unpack-large-objects.sh (new) ##\n    + ## t/t5351-unpack-large-objects.sh (new) ##\n     @@\n     +#!/bin/sh\n     +#\n    @@ t/t5328-unpack-large-objects.sh (new)\n     +\tgit init --bare dest.git\n     +}\n     +\n    -+test_no_loose () {\n    -+\ttest $(find dest.git/objects/?? -type f | wc -l) = 0\n    -+}\n    -+\n     +test_expect_success \"create large objects (1.5 MB) and PACK\" '\n     +\ttest-tool genrandom foo 1500000 >big-blob &&\n     +\ttest_commit --append foo big-blob &&\n     +\ttest-tool genrandom bar 1500000 >big-blob &&\n     +\ttest_commit --append bar big-blob &&\n    -+\tPACK=$(echo HEAD | git pack-objects --revs test)\n    ++\tPACK=$(echo HEAD | git pack-objects --revs pack)\n     +'\n     +\n     +test_expect_success 'set memory limitation to 1MB' '\n    @@ t/t5328-unpack-large-objects.sh (new)\n     +\n     +test_expect_success 'unpack-objects failed under memory limitation' '\n     +\tprepare_dest &&\n    -+\ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n    ++\ttest_must_fail git -C dest.git unpack-objects <pack-$PACK.pack 2>err &&\n     +\tgrep \"fatal: attempting to allocate\" err\n     +'\n     +\n     +test_expect_success 'unpack-objects works with memory limitation in dry-run mode' '\n     +\tprepare_dest &&\n    -+\tgit -C dest.git unpack-objects -n <test-$PACK.pack &&\n    -+\ttest_no_loose &&\n    ++\tgit -C dest.git unpack-objects -n <pack-$PACK.pack &&\n    ++\ttest_stdout_line_count = 0 find dest.git/objects -type f &&\n     +\ttest_dir_is_empty dest.git/objects/pack\n     +'\n     +\n2:  48bf9090058 = 2:  6acd8759772 object-file.c: do fsync() and close() before post-write die()\n3:  0e33d2a6e35 = 3:  f7b02c307fc object-file.c: refactor write_loose_object() to several steps\n-:  ----------- > 4:  20d97cc2605 object-file.c: factor out deflate part of write_loose_object()\n4:  9644df5c744 ! 5:  db40f4160c4 object-file.c: add \"stream_loose_object()\" to handle large object\n    @@ Commit message\n         Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n         Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n         Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n    +    Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n      ## object-file.c ##\n     @@ object-file.c: static int freshen_packed_object(const struct object_id *oid)\n    @@ object-file.c: static int freshen_packed_object(const struct object_id *oid)\n     +\tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n     +\thdrlen = format_object_header(hdr, sizeof(hdr), OBJ_BLOB, len);\n     +\n    -+\t/* Common steps for write_loose_object and stream_loose_object to\n    -+\t * start writing loose oject:\n    ++\t/*\n    ++\t * Common steps for write_loose_object and stream_loose_object to\n    ++\t * start writing loose objects:\n     +\t *\n     +\t *  - Create tmpfile for the loose object.\n     +\t *  - Setup zlib stream for compression.\n    @@ object-file.c: static int freshen_packed_object(const struct object_id *oid)\n     +\t/* Then the data itself.. */\n     +\tdo {\n     +\t\tunsigned char *in0 = stream.next_in;\n    ++\n     +\t\tif (!stream.avail_in && !in_stream->is_finished) {\n     +\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n     +\t\t\tstream.next_in = (void *)in;\n     +\t\t\tin0 = (unsigned char *)in;\n     +\t\t\t/* All data has been read. */\n     +\t\t\tif (in_stream->is_finished)\n    -+\t\t\t\tflush = Z_FINISH;\n    ++\t\t\t\tflush = 1;\n     +\t\t}\n    -+\t\tret = git_deflate(&stream, flush);\n    -+\t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n    -+\t\tif (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n    -+\t\t\tdie(_(\"unable to write loose object file\"));\n    -+\t\tstream.next_out = compressed;\n    -+\t\tstream.avail_out = sizeof(compressed);\n    ++\t\tret = write_loose_object_common(&c, &stream, flush, in0, fd,\n    ++\t\t\t\t\t\tcompressed, sizeof(compressed));\n     +\t\t/*\n     +\t\t * Unlike write_loose_object(), we do not have the entire\n     +\t\t * buffer. If we get Z_BUF_ERROR due to too few input bytes,\n5:  4550f3a2745 = 6:  d8ae2eadb98 core doc: modernize core.bigFileThreshold documentation\n-:  ----------- > 7:  2b403e7cd9c unpack-objects: refactor away unpack_non_delta_entry()\n6:  6a70e49a346 ! 8:  5eded902496 unpack-objects: use stream_loose_object() to unpack large objects\n    @@ Commit message\n         malloc() the size of the blob before unpacking it, which could cause\n         OOM with very large blobs.\n     \n    -    We could use this new interface to unpack all blobs, but doing so\n    -    would result in a performance penalty of around 10%, as the below\n    -    \"hyperfine\" benchmark will show. We therefore limit this to files\n    -    larger than \"core.bigFileThreshold\":\n    -\n    -        $ hyperfine \\\n    -          --setup \\\n    -          'if ! test -d scalar.git; then git clone --bare\n    -           https://github.com/microsoft/scalar.git;\n    -           cp scalar.git/objects/pack/*.pack small.pack; fi' \\\n    -          --prepare 'rm -rf dest.git && git init --bare dest.git' \\\n    -          ...\n    -\n    -        Summary\n    -          './git -C dest.git -c core.bigFileThreshold=512m\n    -          unpack-objects <small.pack' in 'origin/master'\n    -            1.01 ± 0.04 times faster than './git -C dest.git\n    -                    -c core.bigFileThreshold=512m unpack-objects\n    -                    <small.pack' in 'HEAD~1'\n    -            1.01 ± 0.04 times faster than './git -C dest.git\n    -                    -c core.bigFileThreshold=512m unpack-objects\n    -                    <small.pack' in 'HEAD~0'\n    -            1.03 ± 0.10 times faster than './git -C dest.git\n    -                    -c core.bigFileThreshold=16k unpack-objects\n    -                    <small.pack' in 'origin/master'\n    -            1.02 ± 0.07 times faster than './git -C dest.git\n    -                    -c core.bigFileThreshold=16k unpack-objects\n    -                    <small.pack' in 'HEAD~0'\n    -            1.10 ± 0.04 times faster than './git -C dest.git\n    -                    -c core.bigFileThreshold=16k unpack-objects\n    -                    <small.pack' in 'HEAD~1'\n    +    We could use the new streaming interface to unpack all blobs, but\n    +    doing so would be much slower, as demonstrated e.g. with this\n    +    benchmark using git-hyperfine[0]:\n    +\n    +            rm -rf /tmp/scalar.git &&\n    +            git clone --bare https://github.com/Microsoft/scalar.git /tmp/scalar.git &&\n    +            mv /tmp/scalar.git/objects/pack/*.pack /tmp/scalar.git/my.pack &&\n    +            git hyperfine \\\n    +                    -r 2 --warmup 1 \\\n    +                    -L rev origin/master,HEAD -L v \"10,512,1k,1m\" \\\n    +                    -s 'make' \\\n    +                    -p 'git init --bare dest.git' \\\n    +                    -c 'rm -rf dest.git' \\\n    +                    './git -C dest.git -c core.bigFileThreshold={v} unpack-objects </tmp/scalar.git/my.pack'\n    +\n    +    Here we'll perform worse with lower core.bigFileThreshold settings\n    +    with this change in terms of speed, but we're getting lower memory use\n    +    in return:\n    +\n    +            Summary\n    +              './git -C dest.git -c core.bigFileThreshold=10 unpack-objects </tmp/scalar.git/my.pack' in 'origin/master' ran\n    +                1.01 ± 0.01 times faster than './git -C dest.git -c core.bigFileThreshold=1k unpack-objects </tmp/scalar.git/my.pack' in 'origin/master'\n    +                1.01 ± 0.01 times faster than './git -C dest.git -c core.bigFileThreshold=1m unpack-objects </tmp/scalar.git/my.pack' in 'origin/master'\n    +                1.01 ± 0.02 times faster than './git -C dest.git -c core.bigFileThreshold=1m unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n    +                1.02 ± 0.00 times faster than './git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/scalar.git/my.pack' in 'origin/master'\n    +                1.09 ± 0.01 times faster than './git -C dest.git -c core.bigFileThreshold=1k unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n    +                1.10 ± 0.00 times faster than './git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n    +                1.11 ± 0.00 times faster than './git -C dest.git -c core.bigFileThreshold=10 unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n    +\n    +    A better benchmark to demonstrate the benefits of that this one, which\n    +    creates an artificial repo with a 1, 25, 50, 75 and 100MB blob:\n    +\n    +            rm -rf /tmp/repo &&\n    +            git init /tmp/repo &&\n    +            (\n    +                    cd /tmp/repo &&\n    +                    for i in 1 25 50 75 100\n    +                    do\n    +                            dd if=/dev/urandom of=blob.$i count=$(($i*1024)) bs=1024\n    +                    done &&\n    +                    git add blob.* &&\n    +                    git commit -mblobs &&\n    +                    git gc &&\n    +                    PACK=$(echo .git/objects/pack/pack-*.pack) &&\n    +                    cp \"$PACK\" my.pack\n    +            ) &&\n    +            git hyperfine \\\n    +                    --show-output \\\n    +                    -L rev origin/master,HEAD -L v \"512,50m,100m\" \\\n    +                    -s 'make' \\\n    +                    -p 'git init --bare dest.git' \\\n    +                    -c 'rm -rf dest.git' \\\n    +                    '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold={v} unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum'\n    +\n    +    Using this test we'll always use >100MB of memory on\n    +    origin/master (around ~105MB), but max out at e.g. ~55MB if we set\n    +    core.bigFileThreshold=50m.\n    +\n    +    The relevant \"Maximum resident set size\" lines were manually added\n    +    below the relevant benchmark:\n    +\n    +      '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=50m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'origin/master' ran\n    +            Maximum resident set size (kbytes): 107080\n    +        1.02 ± 0.78 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'origin/master'\n    +            Maximum resident set size (kbytes): 106968\n    +        1.09 ± 0.79 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=100m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'origin/master'\n    +            Maximum resident set size (kbytes): 107032\n    +        1.42 ± 1.07 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=100m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'HEAD'\n    +            Maximum resident set size (kbytes): 107072\n    +        1.83 ± 1.02 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=50m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'HEAD'\n    +            Maximum resident set size (kbytes): 55704\n    +        2.16 ± 1.19 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'HEAD'\n    +            Maximum resident set size (kbytes): 4564\n    +\n    +    This shows that if you have enough memory this new streaming method is\n    +    slower the lower you set the streaming threshold, but the benefit is\n    +    more bounded memory use.\n     \n         An earlier version of this patch introduced a new\n         \"core.bigFileStreamingThreshold\" instead of re-using the existing\n    @@ Commit message\n         split up \"core.bigFileThreshold\" in the future if there's a need for\n         that.\n     \n    +    0. https://github.com/avar/git-hyperfine/\n         1. https://lore.kernel.org/git/20211210103435.83656-1-chiyutianyi@gmail.com/\n         2. https://lore.kernel.org/git/20220120112114.47618-5-chiyutianyi@gmail.com/\n     \n    @@ Commit message\n         Helped-by: Derrick Stolee <stolee@gmail.com>\n         Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n         Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n    +    Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n      ## Documentation/config/core.txt ##\n     @@ Documentation/config/core.txt: usage, at the slight expense of increased disk usage.\n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n     +\treturn data->buf;\n     +}\n     +\n    -+static void write_stream_blob(unsigned nr, size_t size)\n    ++static void stream_blob(unsigned long size, unsigned nr)\n     +{\n     +\tgit_zstream zstream = { 0 };\n     +\tstruct input_zstream_data data = { 0 };\n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n     +\t\t.read = feed_input_zstream,\n     +\t\t.data = &data,\n     +\t};\n    ++\tstruct obj_info *info = &obj_list[nr];\n     +\n     +\tdata.zstream = &zstream;\n     +\tgit_inflate_init(&zstream);\n     +\n    -+\tif (stream_loose_object(&in_stream, size, &obj_list[nr].oid))\n    ++\tif (stream_loose_object(&in_stream, size, &info->oid))\n     +\t\tdie(_(\"failed to write object in stream\"));\n     +\n     +\tif (data.status != Z_STREAM_END)\n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n     +\tgit_inflate_end(&zstream);\n     +\n     +\tif (strict) {\n    -+\t\tstruct blob *blob =\n    -+\t\t\tlookup_blob(the_repository, &obj_list[nr].oid);\n    -+\t\tif (blob)\n    -+\t\t\tblob->object.flags |= FLAG_WRITTEN;\n    -+\t\telse\n    ++\t\tstruct blob *blob = lookup_blob(the_repository, &info->oid);\n    ++\n    ++\t\tif (!blob)\n     +\t\t\tdie(_(\"invalid blob object from stream\"));\n    ++\t\tblob->object.flags |= FLAG_WRITTEN;\n     +\t}\n    -+\tobj_list[nr].obj = NULL;\n    ++\tinfo->obj = NULL;\n     +}\n     +\n    - static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n    - \t\t\t\t   unsigned nr)\n    + static int resolve_against_held(unsigned nr, const struct object_id *base,\n    + \t\t\t\tvoid *delta_data, unsigned long delta_size)\n      {\n    --\tvoid *buf = get_data(size);\n    -+\tvoid *buf;\n    -+\n    -+\t/* Write large blob in stream without allocating full buffer. */\n    -+\tif (!dry_run && type == OBJ_BLOB && size > big_file_threshold) {\n    -+\t\twrite_stream_blob(nr, size);\n    -+\t\treturn;\n    -+\t}\n    +@@ builtin/unpack-objects.c: static void unpack_one(unsigned nr)\n      \n    -+\tbuf = get_data(size);\n    - \tif (buf)\n    - \t\twrite_object(nr, type, buf, size);\n    - }\n    + \tswitch (type) {\n    + \tcase OBJ_BLOB:\n    ++\t\tif (!dry_run && size > big_file_threshold) {\n    ++\t\t\tstream_blob(size, nr);\n    ++\t\t\treturn;\n    ++\t\t}\n    ++\t\t/* fallthrough */\n    + \tcase OBJ_COMMIT:\n    + \tcase OBJ_TREE:\n    + \tcase OBJ_TAG:\n     \n    - ## t/t5328-unpack-large-objects.sh ##\n    -@@ t/t5328-unpack-large-objects.sh: test_description='git unpack-objects with large objects'\n    + ## t/t5351-unpack-large-objects.sh ##\n    +@@ t/t5351-unpack-large-objects.sh: test_description='git unpack-objects with large objects'\n      \n      prepare_dest () {\n      \ttest_when_finished \"rm -rf dest.git\" &&\n     -\tgit init --bare dest.git\n     +\tgit init --bare dest.git &&\n    -+\tif test -n \"$1\"\n    -+\tthen\n    -+\t\tgit -C dest.git config core.bigFileThreshold $1\n    -+\tfi\n    ++\tgit -C dest.git config core.bigFileThreshold \"$1\"\n      }\n      \n    - test_no_loose () {\n    -@@ t/t5328-unpack-large-objects.sh: test_expect_success 'set memory limitation to 1MB' '\n    + test_expect_success \"create large objects (1.5 MB) and PACK\" '\n    +@@ t/t5351-unpack-large-objects.sh: test_expect_success 'set memory limitation to 1MB' '\n      '\n      \n      test_expect_success 'unpack-objects failed under memory limitation' '\n     -\tprepare_dest &&\n     +\tprepare_dest 2m &&\n    - \ttest_must_fail git -C dest.git unpack-objects <test-$PACK.pack 2>err &&\n    + \ttest_must_fail git -C dest.git unpack-objects <pack-$PACK.pack 2>err &&\n      \tgrep \"fatal: attempting to allocate\" err\n      '\n      \n      test_expect_success 'unpack-objects works with memory limitation in dry-run mode' '\n     -\tprepare_dest &&\n     +\tprepare_dest 2m &&\n    - \tgit -C dest.git unpack-objects -n <test-$PACK.pack &&\n    - \ttest_no_loose &&\n    + \tgit -C dest.git unpack-objects -n <pack-$PACK.pack &&\n    + \ttest_stdout_line_count = 0 find dest.git/objects -type f &&\n      \ttest_dir_is_empty dest.git/objects/pack\n      '\n      \n     +test_expect_success 'unpack big object in stream' '\n     +\tprepare_dest 1m &&\n    -+\tgit -C dest.git unpack-objects <test-$PACK.pack &&\n    ++\tgit -C dest.git unpack-objects <pack-$PACK.pack &&\n     +\ttest_dir_is_empty dest.git/objects/pack\n     +'\n     +\n     +test_expect_success 'do not unpack existing large objects' '\n     +\tprepare_dest 1m &&\n    -+\tgit -C dest.git index-pack --stdin <test-$PACK.pack &&\n    -+\tgit -C dest.git unpack-objects <test-$PACK.pack &&\n    -+\ttest_no_loose\n    ++\tgit -C dest.git index-pack --stdin <pack-$PACK.pack &&\n    ++\tgit -C dest.git unpack-objects <pack-$PACK.pack &&\n    ++\n    ++\t# The destination came up with the exact same pack...\n    ++\tDEST_PACK=$(echo dest.git/objects/pack/pack-*.pack) &&\n    ++\ttest_cmp pack-$PACK.pack $DEST_PACK &&\n    ++\n    ++\t# ...and wrote no loose objects\n    ++\ttest_stdout_line_count = 0 find dest.git/objects -type f ! -name \"pack-*\"\n     +'\n     +\n      test_done\n-- \n2.35.1.1438.g8874c8eeb35\n\n"},{"id":"451650","messageId":"patch-v11-1.8-2103d5bfd96-20220319T001411Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v11-0.8-00000000000-20220319T001411Z-avarab@gmail.com","subject":"[PATCH v11 1/8] unpack-objects: low memory footprint for get_data() in dry_run mode","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-19T00:23:18Z","receivedAt":"2022-03-19T00:23:47Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nAs the name implies, \"get_data(size)\" will allocate and return a given\namount of memory. Allocating memory for a large blob object may cause the\nsystem to run out of memory. Before preparing to replace calling of\n\"get_data()\" to unpack large blob objects in latter commits, refactor\n\"get_data()\" to reduce memory footprint for dry_run mode.\n\nBecause in dry_run mode, \"get_data()\" is only used to check the\nintegrity of data, and the returned buffer is not used at all, we can\nallocate a smaller buffer and reuse it as zstream output. Therefore,\nin dry_run mode, \"get_data()\" will release the allocated buffer and\nreturn NULL instead of returning garbage data.\n\nThe \"find [...]objects/?? -type f | wc -l\" test idiom being used here\nis adapted from the same \"find\" use added to another test in\nd9545c7f465 (fast-import: implement unpack limit, 2016-04-25).\n\nSuggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n builtin/unpack-objects.c        | 34 ++++++++++++++++++---------\n t/t5351-unpack-large-objects.sh | 41 +++++++++++++++++++++++++++++++++\n 2 files changed, 64 insertions(+), 11 deletions(-)\n create mode 100755 t/t5351-unpack-large-objects.sh\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex dbeb0680a58..e3d30025979 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -96,15 +96,26 @@ static void use(int bytes)\n \tdisplay_throughput(progress, consumed_bytes);\n }\n \n+/*\n+ * Decompress zstream from stdin and return specific size of data.\n+ * The caller is responsible to free the returned buffer.\n+ *\n+ * But for dry_run mode, \"get_data()\" is only used to check the\n+ * integrity of data, and the returned buffer is not used at all.\n+ * Therefore, in dry_run mode, \"get_data()\" will release the small\n+ * allocated buffer which is reused to hold temporary zstream output\n+ * and return NULL instead of returning garbage data.\n+ */\n static void *get_data(unsigned long size)\n {\n \tgit_zstream stream;\n-\tvoid *buf = xmallocz(size);\n+\tunsigned long bufsize = dry_run && size > 8192 ? 8192 : size;\n+\tvoid *buf = xmallocz(bufsize);\n \n \tmemset(&stream, 0, sizeof(stream));\n \n \tstream.next_out = buf;\n-\tstream.avail_out = size;\n+\tstream.avail_out = bufsize;\n \tstream.next_in = fill(1);\n \tstream.avail_in = len;\n \tgit_inflate_init(&stream);\n@@ -124,8 +135,15 @@ static void *get_data(unsigned long size)\n \t\t}\n \t\tstream.next_in = fill(1);\n \t\tstream.avail_in = len;\n+\t\tif (dry_run) {\n+\t\t\t/* reuse the buffer in dry_run mode */\n+\t\t\tstream.next_out = buf;\n+\t\t\tstream.avail_out = bufsize;\n+\t\t}\n \t}\n \tgit_inflate_end(&stream);\n+\tif (dry_run)\n+\t\tFREE_AND_NULL(buf);\n \treturn buf;\n }\n \n@@ -325,10 +343,8 @@ static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n {\n \tvoid *buf = get_data(size);\n \n-\tif (!dry_run && buf)\n+\tif (buf)\n \t\twrite_object(nr, type, buf, size);\n-\telse\n-\t\tfree(buf);\n }\n \n static int resolve_against_held(unsigned nr, const struct object_id *base,\n@@ -358,10 +374,8 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\toidread(&base_oid, fill(the_hash_algo->rawsz));\n \t\tuse(the_hash_algo->rawsz);\n \t\tdelta_data = get_data(delta_size);\n-\t\tif (dry_run || !delta_data) {\n-\t\t\tfree(delta_data);\n+\t\tif (!delta_data)\n \t\t\treturn;\n-\t\t}\n \t\tif (has_object_file(&base_oid))\n \t\t\t; /* Ok we have this one */\n \t\telse if (resolve_against_held(nr, &base_oid,\n@@ -397,10 +411,8 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\t\tdie(\"offset value out of bound for delta base object\");\n \n \t\tdelta_data = get_data(delta_size);\n-\t\tif (dry_run || !delta_data) {\n-\t\t\tfree(delta_data);\n+\t\tif (!delta_data)\n \t\t\treturn;\n-\t\t}\n \t\tlo = 0;\n \t\thi = nr;\n \t\twhile (lo < hi) {\ndiff --git a/t/t5351-unpack-large-objects.sh b/t/t5351-unpack-large-objects.sh\nnew file mode 100755\nindex 00000000000..8d84313221c\n--- /dev/null\n+++ b/t/t5351-unpack-large-objects.sh\n@@ -0,0 +1,41 @@\n+#!/bin/sh\n+#\n+# Copyright (c) 2022 Han Xin\n+#\n+\n+test_description='git unpack-objects with large objects'\n+\n+. ./test-lib.sh\n+\n+prepare_dest () {\n+\ttest_when_finished \"rm -rf dest.git\" &&\n+\tgit init --bare dest.git\n+}\n+\n+test_expect_success \"create large objects (1.5 MB) and PACK\" '\n+\ttest-tool genrandom foo 1500000 >big-blob &&\n+\ttest_commit --append foo big-blob &&\n+\ttest-tool genrandom bar 1500000 >big-blob &&\n+\ttest_commit --append bar big-blob &&\n+\tPACK=$(echo HEAD | git pack-objects --revs pack)\n+'\n+\n+test_expect_success 'set memory limitation to 1MB' '\n+\tGIT_ALLOC_LIMIT=1m &&\n+\texport GIT_ALLOC_LIMIT\n+'\n+\n+test_expect_success 'unpack-objects failed under memory limitation' '\n+\tprepare_dest &&\n+\ttest_must_fail git -C dest.git unpack-objects <pack-$PACK.pack 2>err &&\n+\tgrep \"fatal: attempting to allocate\" err\n+'\n+\n+test_expect_success 'unpack-objects works with memory limitation in dry-run mode' '\n+\tprepare_dest &&\n+\tgit -C dest.git unpack-objects -n <pack-$PACK.pack &&\n+\ttest_stdout_line_count = 0 find dest.git/objects -type f &&\n+\ttest_dir_is_empty dest.git/objects/pack\n+'\n+\n+test_done\n-- \n2.35.1.1438.g8874c8eeb35\n\n"},{"id":"451649","messageId":"patch-v11-2.8-6acd8759772-20220319T001411Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v11-0.8-00000000000-20220319T001411Z-avarab@gmail.com","subject":"[PATCH v11 2/8] object-file.c: do fsync() and close() before post-write die()","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-19T00:23:19Z","receivedAt":"2022-03-19T00:23:49Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Change write_loose_object() to do an fsync() and close() before the\noideq() sanity check at the end. This change re-joins code that was\nsplit up by the die() sanity check added in 748af44c63e (sha1_file: be\nparanoid when creating loose objects, 2010-02-21).\n\nI don't think that this change matters in itself, if we called die()\nit was possible that our data wouldn't fully make it to disk, but in\nany case we were writing data that we'd consider corrupted. It's\npossible that a subsequent \"git fsck\" will be less confused now.\n\nThe real reason to make this change is that in a subsequent commit\nwe'll split this code in write_loose_object() into a utility function,\nall its callers will want the preceding sanity checks, but not the\n\"oideq\" check. By moving the close_loose_object() earlier it'll be\neasier to reason about the introduction of the utility function.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n object-file.c | 13 +++++++++++--\n 1 file changed, 11 insertions(+), 2 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex bdc5cbdd386..4c140eda6bf 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2001,12 +2001,21 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n \t\t    ret);\n \tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n+\n+\t/*\n+\t * We already did a write_buffer() to the \"fd\", let's fsync()\n+\t * and close().\n+\t *\n+\t * We might still die() on a subsequent sanity check, but\n+\t * let's not add to that confusion by not flushing any\n+\t * outstanding writes to disk first.\n+\t */\n+\tclose_loose_object(fd);\n+\n \tif (!oideq(oid, &parano_oid))\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n \n-\tclose_loose_object(fd);\n-\n \tif (mtime) {\n \t\tstruct utimbuf utb;\n \t\tutb.actime = mtime;\n-- \n2.35.1.1438.g8874c8eeb35\n\n"},{"id":"451651","messageId":"patch-v11-3.8-f7b02c307fc-20220319T001411Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v11-0.8-00000000000-20220319T001411Z-avarab@gmail.com","subject":"[PATCH v11 3/8] object-file.c: refactor write_loose_object() to several steps","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-19T00:23:20Z","receivedAt":"2022-03-19T00:23:57Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWhen writing a large blob using \"write_loose_object()\", we have to pass\na buffer with the whole content of the blob, and this behavior will\nconsume lots of memory and may cause OOM. We will introduce a stream\nversion function (\"stream_loose_object()\") in later commit to resolve\nthis issue.\n\nBefore introducing that streaming function, do some refactoring on\n\"write_loose_object()\" to reuse code for both versions.\n\nRewrite \"write_loose_object()\" as follows:\n\n 1. Figure out a path for the (temp) object file. This step is only\n    used in \"write_loose_object()\".\n\n 2. Move common steps for starting to write loose objects into a new\n    function \"start_loose_object_common()\".\n\n 3. Compress data.\n\n 4. Move common steps for ending zlib stream into a new function\n    \"end_loose_object_common()\".\n\n 5. Close fd and finalize the object file.\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n object-file.c | 129 ++++++++++++++++++++++++++++++++++----------------\n 1 file changed, 89 insertions(+), 40 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 4c140eda6bf..4fcaf7a36ce 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1943,6 +1943,87 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n \treturn fd;\n }\n \n+/**\n+ * Common steps for loose object writers to start writing loose\n+ * objects:\n+ *\n+ * - Create tmpfile for the loose object.\n+ * - Setup zlib stream for compression.\n+ * - Start to feed header to zlib stream.\n+ *\n+ * Returns a \"fd\", which should later be provided to\n+ * end_loose_object_common().\n+ */\n+static int start_loose_object_common(struct strbuf *tmp_file,\n+\t\t\t\t     const char *filename, unsigned flags,\n+\t\t\t\t     git_zstream *stream,\n+\t\t\t\t     unsigned char *buf, size_t buflen,\n+\t\t\t\t     git_hash_ctx *c,\n+\t\t\t\t     char *hdr, int hdrlen)\n+{\n+\tint fd;\n+\n+\tfd = create_tmpfile(tmp_file, filename);\n+\tif (fd < 0) {\n+\t\tif (flags & HASH_SILENT)\n+\t\t\treturn -1;\n+\t\telse if (errno == EACCES)\n+\t\t\treturn error(_(\"insufficient permission for adding \"\n+\t\t\t\t       \"an object to repository database %s\"),\n+\t\t\t\t     get_object_directory());\n+\t\telse\n+\t\t\treturn error_errno(\n+\t\t\t\t_(\"unable to create temporary file\"));\n+\t}\n+\n+\t/*  Setup zlib stream for compression */\n+\tgit_deflate_init(stream, zlib_compression_level);\n+\tstream->next_out = buf;\n+\tstream->avail_out = buflen;\n+\tthe_hash_algo->init_fn(c);\n+\n+\t/*  Start to feed header to zlib stream */\n+\tstream->next_in = (unsigned char *)hdr;\n+\tstream->avail_in = hdrlen;\n+\twhile (git_deflate(stream, 0) == Z_OK)\n+\t\t; /* nothing */\n+\tthe_hash_algo->update_fn(c, hdr, hdrlen);\n+\n+\treturn fd;\n+}\n+\n+/**\n+ * Common steps for loose object writers to end writing loose objects:\n+ *\n+ * - End the compression of zlib stream.\n+ * - Get the calculated oid to \"parano_oid\".\n+ * - fsync() and close() the \"fd\"\n+ */\n+static void end_loose_object_common(int fd, int ret, git_hash_ctx *c,\n+\t\t\t\t    git_zstream *stream,\n+\t\t\t\t    struct object_id *parano_oid,\n+\t\t\t\t    const struct object_id *expected_oid,\n+\t\t\t\t    const char *die_msg1_fmt,\n+\t\t\t\t    const char *die_msg2_fmt)\n+{\n+\tif (ret != Z_STREAM_END)\n+\t\tdie(_(die_msg1_fmt), ret, expected_oid);\n+\tret = git_deflate_end_gently(stream);\n+\tif (ret != Z_OK)\n+\t\tdie(_(die_msg2_fmt), ret, expected_oid);\n+\tthe_hash_algo->final_oid_fn(parano_oid, c);\n+\n+\t/*\n+\t * We already did a write_buffer() to the \"fd\", let's fsync()\n+\t * and close().\n+\t *\n+\t * We might still die() on a subsequent sanity check, but\n+\t * let's not add to that confusion by not flushing any\n+\t * outstanding writes to disk first.\n+\t */\n+\tclose_loose_object(fd);\n+}\n+\n static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\t\t      int hdrlen, const void *buf, unsigned long len,\n \t\t\t      time_t mtime, unsigned flags)\n@@ -1957,28 +2038,11 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \n \tloose_object_path(the_repository, &filename, oid);\n \n-\tfd = create_tmpfile(&tmp_file, filename.buf);\n-\tif (fd < 0) {\n-\t\tif (flags & HASH_SILENT)\n-\t\t\treturn -1;\n-\t\telse if (errno == EACCES)\n-\t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n-\t\telse\n-\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n-\t}\n-\n-\t/* Set it up */\n-\tgit_deflate_init(&stream, zlib_compression_level);\n-\tstream.next_out = compressed;\n-\tstream.avail_out = sizeof(compressed);\n-\tthe_hash_algo->init_fn(&c);\n-\n-\t/* First header.. */\n-\tstream.next_in = (unsigned char *)hdr;\n-\tstream.avail_in = hdrlen;\n-\twhile (git_deflate(&stream, 0) == Z_OK)\n-\t\t; /* nothing */\n-\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n+\tfd = start_loose_object_common(&tmp_file, filename.buf, flags,\n+\t\t\t\t       &stream, compressed, sizeof(compressed),\n+\t\t\t\t       &c, hdr, hdrlen);\n+\tif (fd < 0)\n+\t\treturn -1;\n \n \t/* Then the data itself.. */\n \tstream.next_in = (void *)buf;\n@@ -1993,24 +2057,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tstream.avail_out = sizeof(compressed);\n \t} while (ret == Z_OK);\n \n-\tif (ret != Z_STREAM_END)\n-\t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n-\t\t    ret);\n-\tret = git_deflate_end_gently(&stream);\n-\tif (ret != Z_OK)\n-\t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n-\t\t    ret);\n-\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n-\n-\t/*\n-\t * We already did a write_buffer() to the \"fd\", let's fsync()\n-\t * and close().\n-\t *\n-\t * We might still die() on a subsequent sanity check, but\n-\t * let's not add to that confusion by not flushing any\n-\t * outstanding writes to disk first.\n-\t */\n-\tclose_loose_object(fd);\n+\tend_loose_object_common(fd, ret, &c, &stream, &parano_oid, oid,\n+\t\t\t\tN_(\"unable to deflate new object %s (%d)\"),\n+\t\t\t\tN_(\"deflateEnd on object %s failed (%d)\"));\n \n \tif (!oideq(oid, &parano_oid))\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n-- \n2.35.1.1438.g8874c8eeb35\n\n"},{"id":"451652","messageId":"patch-v11-4.8-20d97cc2605-20220319T001411Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v11-0.8-00000000000-20220319T001411Z-avarab@gmail.com","subject":"[PATCH v11 4/8] object-file.c: factor out deflate part of write_loose_object()","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-19T00:23:21Z","receivedAt":"2022-03-19T00:24:00Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Split out the part of write_loose_object() that deals with calling\ngit_deflate() into a utility function, a subsequent commit will\nintroduce another function that'll make use of it.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n object-file.c | 31 +++++++++++++++++++++++++------\n 1 file changed, 25 insertions(+), 6 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 4fcaf7a36ce..b66dc24e4b8 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1992,6 +1992,28 @@ static int start_loose_object_common(struct strbuf *tmp_file,\n \treturn fd;\n }\n \n+/**\n+ * Common steps for the inner git_deflate() loop for writing loose\n+ * objects. Returns what git_deflate() returns.\n+ */\n+static int write_loose_object_common(git_hash_ctx *c,\n+\t\t\t\t     git_zstream *stream, const int flush,\n+\t\t\t\t     unsigned char *in0, const int fd,\n+\t\t\t\t     unsigned char *compressed,\n+\t\t\t\t     const size_t compressed_len)\n+{\n+\tint ret;\n+\n+\tret = git_deflate(stream, flush ? Z_FINISH : 0);\n+\tthe_hash_algo->update_fn(c, in0, stream->next_in - in0);\n+\tif (write_buffer(fd, compressed, stream->next_out - compressed) < 0)\n+\t\tdie(_(\"unable to write loose object file\"));\n+\tstream->next_out = compressed;\n+\tstream->avail_out = compressed_len;\n+\n+\treturn ret;\n+}\n+\n /**\n  * Common steps for loose object writers to end writing loose objects:\n  *\n@@ -2049,12 +2071,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tstream.avail_in = len;\n \tdo {\n \t\tunsigned char *in0 = stream.next_in;\n-\t\tret = git_deflate(&stream, Z_FINISH);\n-\t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n-\t\tif (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n-\t\t\tdie(_(\"unable to write loose object file\"));\n-\t\tstream.next_out = compressed;\n-\t\tstream.avail_out = sizeof(compressed);\n+\n+\t\tret = write_loose_object_common(&c, &stream, 1, in0, fd,\n+\t\t\t\t\t\tcompressed, sizeof(compressed));\n \t} while (ret == Z_OK);\n \n \tend_loose_object_common(fd, ret, &c, &stream, &parano_oid, oid,\n-- \n2.35.1.1438.g8874c8eeb35\n\n"},{"id":"451653","messageId":"patch-v11-5.8-db40f4160c4-20220319T001411Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v11-0.8-00000000000-20220319T001411Z-avarab@gmail.com","subject":"[PATCH v11 5/8] object-file.c: add \"stream_loose_object()\" to handle large object","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-19T00:23:22Z","receivedAt":"2022-03-19T00:24:02Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nIf we want unpack and write a loose object using \"write_loose_object\",\nwe have to feed it with a buffer with the same size of the object, which\nwill consume lots of memory and may cause OOM. This can be improved by\nfeeding data to \"stream_loose_object()\" in a stream.\n\nAdd a new function \"stream_loose_object()\", which is a stream version of\n\"write_loose_object()\" but with a low memory footprint. We will use this\nfunction to unpack large blob object in later commit.\n\nAnother difference with \"write_loose_object()\" is that we have no chance\nto run \"write_object_file_prepare()\" to calculate the oid in advance.\nIn \"write_loose_object()\", we know the oid and we can write the\ntemporary file in the same directory as the final object, but for an\nobject with an undetermined oid, we don't know the exact directory for\nthe object.\n\nStill, we need to save the temporary file we're preparing\nsomewhere. We'll do that in the top-level \".git/objects/\"\ndirectory (or whatever \"GIT_OBJECT_DIRECTORY\" is set to). Once we've\nstreamed it we'll know the OID, and will move it to its canonical\npath.\n\n\"freshen_packed_object()\" or \"freshen_loose_object()\" will be called\ninside \"stream_loose_object()\" after obtaining the \"oid\".\n\nHelped-by: René Scharfe <l.s.r@web.de>\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n object-file.c  | 97 ++++++++++++++++++++++++++++++++++++++++++++++++++\n object-store.h |  8 +++++\n 2 files changed, 105 insertions(+)\n\ndiff --git a/object-file.c b/object-file.c\nindex b66dc24e4b8..548fef71b98 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2114,6 +2114,103 @@ static int freshen_packed_object(const struct object_id *oid)\n \treturn 1;\n }\n \n+int stream_loose_object(struct input_stream *in_stream, size_t len,\n+\t\t\tstruct object_id *oid)\n+{\n+\tint fd, ret, err = 0, flush = 0;\n+\tunsigned char compressed[4096];\n+\tgit_zstream stream;\n+\tgit_hash_ctx c;\n+\tstruct strbuf tmp_file = STRBUF_INIT;\n+\tstruct strbuf filename = STRBUF_INIT;\n+\tint dirlen;\n+\tchar hdr[MAX_HEADER_LEN];\n+\tint hdrlen;\n+\n+\t/* Since oid is not determined, save tmp file to odb path. */\n+\tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n+\thdrlen = format_object_header(hdr, sizeof(hdr), OBJ_BLOB, len);\n+\n+\t/*\n+\t * Common steps for write_loose_object and stream_loose_object to\n+\t * start writing loose objects:\n+\t *\n+\t *  - Create tmpfile for the loose object.\n+\t *  - Setup zlib stream for compression.\n+\t *  - Start to feed header to zlib stream.\n+\t */\n+\tfd = start_loose_object_common(&tmp_file, filename.buf, 0,\n+\t\t\t\t       &stream, compressed, sizeof(compressed),\n+\t\t\t\t       &c, hdr, hdrlen);\n+\tif (fd < 0) {\n+\t\terr = -1;\n+\t\tgoto cleanup;\n+\t}\n+\n+\t/* Then the data itself.. */\n+\tdo {\n+\t\tunsigned char *in0 = stream.next_in;\n+\n+\t\tif (!stream.avail_in && !in_stream->is_finished) {\n+\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n+\t\t\tstream.next_in = (void *)in;\n+\t\t\tin0 = (unsigned char *)in;\n+\t\t\t/* All data has been read. */\n+\t\t\tif (in_stream->is_finished)\n+\t\t\t\tflush = 1;\n+\t\t}\n+\t\tret = write_loose_object_common(&c, &stream, flush, in0, fd,\n+\t\t\t\t\t\tcompressed, sizeof(compressed));\n+\t\t/*\n+\t\t * Unlike write_loose_object(), we do not have the entire\n+\t\t * buffer. If we get Z_BUF_ERROR due to too few input bytes,\n+\t\t * then we'll replenish them in the next input_stream->read()\n+\t\t * call when we loop.\n+\t\t */\n+\t} while (ret == Z_OK || ret == Z_BUF_ERROR);\n+\n+\tif (stream.total_in != len + hdrlen)\n+\t\tdie(_(\"write stream object %ld != %\"PRIuMAX), stream.total_in,\n+\t\t    (uintmax_t)len + hdrlen);\n+\n+\t/* Common steps for write_loose_object and stream_loose_object to\n+\t * end writing loose oject:\n+\t *\n+\t *  - End the compression of zlib stream.\n+\t *  - Get the calculated oid.\n+\t */\n+\tend_loose_object_common(fd, ret, &c, &stream, oid, NULL,\n+\t\t\t\tN_(\"unable to stream deflate new object (%d)\"),\n+\t\t\t\tN_(\"deflateEnd on stream object failed (%d)\"));\n+\n+\tif (freshen_packed_object(oid) || freshen_loose_object(oid)) {\n+\t\tunlink_or_warn(tmp_file.buf);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tloose_object_path(the_repository, &filename, oid);\n+\n+\t/* We finally know the object path, and create the missing dir. */\n+\tdirlen = directory_size(filename.buf);\n+\tif (dirlen) {\n+\t\tstruct strbuf dir = STRBUF_INIT;\n+\t\tstrbuf_add(&dir, filename.buf, dirlen);\n+\n+\t\tif (mkdir_in_gitdir(dir.buf) && errno != EEXIST) {\n+\t\t\terr = error_errno(_(\"unable to create directory %s\"), dir.buf);\n+\t\t\tstrbuf_release(&dir);\n+\t\t\tgoto cleanup;\n+\t\t}\n+\t\tstrbuf_release(&dir);\n+\t}\n+\n+\terr = finalize_object_file(tmp_file.buf, filename.buf);\n+cleanup:\n+\tstrbuf_release(&tmp_file);\n+\tstrbuf_release(&filename);\n+\treturn err;\n+}\n+\n int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t    enum object_type type, struct object_id *oid,\n \t\t\t    unsigned flags)\ndiff --git a/object-store.h b/object-store.h\nindex bd2322ed8ce..1099455bc2e 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -46,6 +46,12 @@ struct object_directory {\n \tchar *path;\n };\n \n+struct input_stream {\n+\tconst void *(*read)(struct input_stream *, unsigned long *len);\n+\tvoid *data;\n+\tint is_finished;\n+};\n+\n KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n \tstruct object_directory *, 1, fspathhash, fspatheq)\n \n@@ -261,6 +267,8 @@ static inline int write_object_file(const void *buf, unsigned long len,\n int write_object_file_literally(const void *buf, unsigned long len,\n \t\t\t\tconst char *type, struct object_id *oid,\n \t\t\t\tunsigned flags);\n+int stream_loose_object(struct input_stream *in_stream, size_t len,\n+\t\t\tstruct object_id *oid);\n \n /*\n  * Add an object file to the in-memory object store, without writing it\n-- \n2.35.1.1438.g8874c8eeb35\n\n"},{"id":"451654","messageId":"patch-v11-6.8-d8ae2eadb98-20220319T001411Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v11-0.8-00000000000-20220319T001411Z-avarab@gmail.com","subject":"[PATCH v11 6/8] core doc: modernize core.bigFileThreshold documentation","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-19T00:23:23Z","receivedAt":"2022-03-19T00:24:03Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"The core.bigFileThreshold documentation has been largely unchanged\nsince 5eef828bc03 (fast-import: Stream very large blobs directly to\npack, 2010-02-01).\n\nBut since then this setting has been expanded to affect a lot more\nthan that description indicated. Most notably in how \"git diff\" treats\nthem, see 6bf3b813486 (diff --stat: mark any file larger than\ncore.bigfilethreshold binary, 2014-08-16).\n\nIn addition to that, numerous commands and APIs make use of a\nstreaming mode for files above this threshold.\n\nSo let's attempt to summarize 12 years of changes in behavior, which\ncan be seen with:\n\n    git log --oneline -Gbig_file_thre 5eef828bc03.. -- '*.c'\n\nTo do that turn this into a bullet-point list. The summary Han Xin\nproduced in [1] helped a lot, but is a bit too detailed for\ndocumentation aimed at users. Let's instead summarize how\nuser-observable behavior differs, and generally describe how we tend\nto stream these files in various commands.\n\n1. https://lore.kernel.org/git/20220120112114.47618-5-chiyutianyi@gmail.com/\n\nHelped-by: Han Xin <chiyutianyi@gmail.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n Documentation/config/core.txt | 33 ++++++++++++++++++++++++---------\n 1 file changed, 24 insertions(+), 9 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex c04f62a54a1..b6a12218665 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -412,17 +412,32 @@ You probably do not need to adjust this value.\n Common unit suffixes of 'k', 'm', or 'g' are supported.\n \n core.bigFileThreshold::\n-\tFiles larger than this size are stored deflated, without\n-\tattempting delta compression.  Storing large files without\n-\tdelta compression avoids excessive memory usage, at the\n-\tslight expense of increased disk usage. Additionally files\n-\tlarger than this size are always treated as binary.\n+\tThe size of files considered \"big\", which as discussed below\n+\tchanges the behavior of numerous git commands, as well as how\n+\tsuch files are stored within the repository. The default is\n+\t512 MiB. Common unit suffixes of 'k', 'm', or 'g' are\n+\tsupported.\n +\n-Default is 512 MiB on all platforms.  This should be reasonable\n-for most projects as source code and other text files can still\n-be delta compressed, but larger binary media files won't be.\n+Files above the configured limit will be:\n +\n-Common unit suffixes of 'k', 'm', or 'g' are supported.\n+* Stored deflated, without attempting delta compression.\n++\n+The default limit is primarily set with this use-case in mind. With it\n+most projects will have their source code and other text files delta\n+compressed, but not larger binary media files.\n++\n+Storing large files without delta compression avoids excessive memory\n+usage, at the slight expense of increased disk usage.\n++\n+* Will be treated as if though they were labeled \"binary\" (see\n+  linkgit:gitattributes[5]). This means that e.g. linkgit:git-log[1]\n+  and linkgit:git-diff[1] will not diffs for files above this limit.\n++\n+* Will be generally be streamed when written, which avoids excessive\n+memory usage, at the cost of some fixed overhead. Commands that make\n+use of this include linkgit:git-archive[1],\n+linkgit:git-fast-import[1], linkgit:git-index-pack[1] and\n+linkgit:git-fsck[1].\n \n core.excludesFile::\n \tSpecifies the pathname to the file that contains patterns to\n-- \n2.35.1.1438.g8874c8eeb35\n\n"},{"id":"451655","messageId":"patch-v11-7.8-2b403e7cd9c-20220319T001411Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v11-0.8-00000000000-20220319T001411Z-avarab@gmail.com","subject":"[PATCH v11 7/8] unpack-objects: refactor away unpack_non_delta_entry()","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-19T00:23:24Z","receivedAt":"2022-03-19T00:24:04Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"The unpack_one() function will call either a non-trivial\nunpack_delta_entry() or a trivial unpack_non_delta_entry(). Let's\ninline the latter in the only caller.\n\nSince 21666f1aae4 (convert object type handling from a string to a\nnumber, 2007-02-26) the unpack_non_delta_entry() function has been\nrather trivial, and in a preceding commit the \"dry_run\" condition it\nwas handling went away.\n\nThis is not done as an optimization, as the compiler will easily\ndiscover that it can do the same, rather this makes a subsequent\ncommit easier to reason about. As it'll be handling \"OBJ_BLOB\" in a\nspecial manner let's re-arrange that \"case\" in preparation for that\nchange.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n builtin/unpack-objects.c | 18 +++++++-----------\n 1 file changed, 7 insertions(+), 11 deletions(-)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex e3d30025979..d374599d544 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -338,15 +338,6 @@ static void added_object(unsigned nr, enum object_type type,\n \t}\n }\n \n-static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n-\t\t\t\t   unsigned nr)\n-{\n-\tvoid *buf = get_data(size);\n-\n-\tif (buf)\n-\t\twrite_object(nr, type, buf, size);\n-}\n-\n static int resolve_against_held(unsigned nr, const struct object_id *base,\n \t\t\t\tvoid *delta_data, unsigned long delta_size)\n {\n@@ -479,12 +470,17 @@ static void unpack_one(unsigned nr)\n \t}\n \n \tswitch (type) {\n+\tcase OBJ_BLOB:\n \tcase OBJ_COMMIT:\n \tcase OBJ_TREE:\n-\tcase OBJ_BLOB:\n \tcase OBJ_TAG:\n-\t\tunpack_non_delta_entry(type, size, nr);\n+\t{\n+\t\tvoid *buf = get_data(size);\n+\n+\t\tif (buf)\n+\t\t\twrite_object(nr, type, buf, size);\n \t\treturn;\n+\t}\n \tcase OBJ_REF_DELTA:\n \tcase OBJ_OFS_DELTA:\n \t\tunpack_delta_entry(type, size, nr);\n-- \n2.35.1.1438.g8874c8eeb35\n\n"},{"id":"451656","messageId":"patch-v11-8.8-5eded902496-20220319T001411Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v11-0.8-00000000000-20220319T001411Z-avarab@gmail.com","subject":"[PATCH v11 8/8] unpack-objects: use stream_loose_object() to unpack large objects","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-19T00:23:25Z","receivedAt":"2022-03-19T00:24:18Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nMake use of the stream_loose_object() function introduced in the\npreceding commit to unpack large objects. Before this we'd need to\nmalloc() the size of the blob before unpacking it, which could cause\nOOM with very large blobs.\n\nWe could use the new streaming interface to unpack all blobs, but\ndoing so would be much slower, as demonstrated e.g. with this\nbenchmark using git-hyperfine[0]:\n\n\trm -rf /tmp/scalar.git &&\n\tgit clone --bare https://github.com/Microsoft/scalar.git /tmp/scalar.git &&\n\tmv /tmp/scalar.git/objects/pack/*.pack /tmp/scalar.git/my.pack &&\n\tgit hyperfine \\\n\t\t-r 2 --warmup 1 \\\n\t\t-L rev origin/master,HEAD -L v \"10,512,1k,1m\" \\\n\t\t-s 'make' \\\n\t\t-p 'git init --bare dest.git' \\\n\t\t-c 'rm -rf dest.git' \\\n\t\t'./git -C dest.git -c core.bigFileThreshold={v} unpack-objects </tmp/scalar.git/my.pack'\n\nHere we'll perform worse with lower core.bigFileThreshold settings\nwith this change in terms of speed, but we're getting lower memory use\nin return:\n\n\tSummary\n\t  './git -C dest.git -c core.bigFileThreshold=10 unpack-objects </tmp/scalar.git/my.pack' in 'origin/master' ran\n\t    1.01 ± 0.01 times faster than './git -C dest.git -c core.bigFileThreshold=1k unpack-objects </tmp/scalar.git/my.pack' in 'origin/master'\n\t    1.01 ± 0.01 times faster than './git -C dest.git -c core.bigFileThreshold=1m unpack-objects </tmp/scalar.git/my.pack' in 'origin/master'\n\t    1.01 ± 0.02 times faster than './git -C dest.git -c core.bigFileThreshold=1m unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n\t    1.02 ± 0.00 times faster than './git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/scalar.git/my.pack' in 'origin/master'\n\t    1.09 ± 0.01 times faster than './git -C dest.git -c core.bigFileThreshold=1k unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n\t    1.10 ± 0.00 times faster than './git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n\t    1.11 ± 0.00 times faster than './git -C dest.git -c core.bigFileThreshold=10 unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n\nA better benchmark to demonstrate the benefits of that this one, which\ncreates an artificial repo with a 1, 25, 50, 75 and 100MB blob:\n\n\trm -rf /tmp/repo &&\n\tgit init /tmp/repo &&\n\t(\n\t\tcd /tmp/repo &&\n\t\tfor i in 1 25 50 75 100\n\t\tdo\n\t\t\tdd if=/dev/urandom of=blob.$i count=$(($i*1024)) bs=1024\n\t\tdone &&\n\t\tgit add blob.* &&\n\t\tgit commit -mblobs &&\n\t\tgit gc &&\n\t\tPACK=$(echo .git/objects/pack/pack-*.pack) &&\n\t\tcp \"$PACK\" my.pack\n\t) &&\n\tgit hyperfine \\\n\t\t--show-output \\\n\t\t-L rev origin/master,HEAD -L v \"512,50m,100m\" \\\n\t\t-s 'make' \\\n\t\t-p 'git init --bare dest.git' \\\n\t\t-c 'rm -rf dest.git' \\\n\t\t'/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold={v} unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum'\n\nUsing this test we'll always use >100MB of memory on\norigin/master (around ~105MB), but max out at e.g. ~55MB if we set\ncore.bigFileThreshold=50m.\n\nThe relevant \"Maximum resident set size\" lines were manually added\nbelow the relevant benchmark:\n\n  '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=50m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'origin/master' ran\n        Maximum resident set size (kbytes): 107080\n    1.02 ± 0.78 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'origin/master'\n        Maximum resident set size (kbytes): 106968\n    1.09 ± 0.79 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=100m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'origin/master'\n        Maximum resident set size (kbytes): 107032\n    1.42 ± 1.07 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=100m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'HEAD'\n        Maximum resident set size (kbytes): 107072\n    1.83 ± 1.02 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=50m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'HEAD'\n        Maximum resident set size (kbytes): 55704\n    2.16 ± 1.19 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'HEAD'\n        Maximum resident set size (kbytes): 4564\n\nThis shows that if you have enough memory this new streaming method is\nslower the lower you set the streaming threshold, but the benefit is\nmore bounded memory use.\n\nAn earlier version of this patch introduced a new\n\"core.bigFileStreamingThreshold\" instead of re-using the existing\n\"core.bigFileThreshold\" variable[1]. As noted in a detailed overview\nof its users in [2] using it has several different meanings.\n\nStill, we consider it good enough to simply re-use it. While it's\npossible that someone might want to e.g. consider objects \"small\" for\nthe purposes of diffing but \"big\" for the purposes of writing them\nsuch use-cases are probably too obscure to worry about. We can always\nsplit up \"core.bigFileThreshold\" in the future if there's a need for\nthat.\n\n0. https://github.com/avar/git-hyperfine/\n1. https://lore.kernel.org/git/20211210103435.83656-1-chiyutianyi@gmail.com/\n2. https://lore.kernel.org/git/20220120112114.47618-5-chiyutianyi@gmail.com/\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Derrick Stolee <stolee@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n Documentation/config/core.txt   |  4 +-\n builtin/unpack-objects.c        | 67 +++++++++++++++++++++++++++++++++\n t/t5351-unpack-large-objects.sh | 26 +++++++++++--\n 3 files changed, 92 insertions(+), 5 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex b6a12218665..5aca987632c 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -436,8 +436,8 @@ usage, at the slight expense of increased disk usage.\n * Will be generally be streamed when written, which avoids excessive\n memory usage, at the cost of some fixed overhead. Commands that make\n use of this include linkgit:git-archive[1],\n-linkgit:git-fast-import[1], linkgit:git-index-pack[1] and\n-linkgit:git-fsck[1].\n+linkgit:git-fast-import[1], linkgit:git-index-pack[1],\n+linkgit:git-unpack-objects[1] and linkgit:git-fsck[1].\n \n core.excludesFile::\n \tSpecifies the pathname to the file that contains patterns to\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex d374599d544..9d7b325c23b 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -338,6 +338,68 @@ static void added_object(unsigned nr, enum object_type type,\n \t}\n }\n \n+struct input_zstream_data {\n+\tgit_zstream *zstream;\n+\tunsigned char buf[8192];\n+\tint status;\n+};\n+\n+static const void *feed_input_zstream(struct input_stream *in_stream,\n+\t\t\t\t      unsigned long *readlen)\n+{\n+\tstruct input_zstream_data *data = in_stream->data;\n+\tgit_zstream *zstream = data->zstream;\n+\tvoid *in = fill(1);\n+\n+\tif (in_stream->is_finished) {\n+\t\t*readlen = 0;\n+\t\treturn NULL;\n+\t}\n+\n+\tzstream->next_out = data->buf;\n+\tzstream->avail_out = sizeof(data->buf);\n+\tzstream->next_in = in;\n+\tzstream->avail_in = len;\n+\n+\tdata->status = git_inflate(zstream, 0);\n+\n+\tin_stream->is_finished = data->status != Z_OK;\n+\tuse(len - zstream->avail_in);\n+\t*readlen = sizeof(data->buf) - zstream->avail_out;\n+\n+\treturn data->buf;\n+}\n+\n+static void stream_blob(unsigned long size, unsigned nr)\n+{\n+\tgit_zstream zstream = { 0 };\n+\tstruct input_zstream_data data = { 0 };\n+\tstruct input_stream in_stream = {\n+\t\t.read = feed_input_zstream,\n+\t\t.data = &data,\n+\t};\n+\tstruct obj_info *info = &obj_list[nr];\n+\n+\tdata.zstream = &zstream;\n+\tgit_inflate_init(&zstream);\n+\n+\tif (stream_loose_object(&in_stream, size, &info->oid))\n+\t\tdie(_(\"failed to write object in stream\"));\n+\n+\tif (data.status != Z_STREAM_END)\n+\t\tdie(_(\"inflate returned (%d)\"), data.status);\n+\tgit_inflate_end(&zstream);\n+\n+\tif (strict) {\n+\t\tstruct blob *blob = lookup_blob(the_repository, &info->oid);\n+\n+\t\tif (!blob)\n+\t\t\tdie(_(\"invalid blob object from stream\"));\n+\t\tblob->object.flags |= FLAG_WRITTEN;\n+\t}\n+\tinfo->obj = NULL;\n+}\n+\n static int resolve_against_held(unsigned nr, const struct object_id *base,\n \t\t\t\tvoid *delta_data, unsigned long delta_size)\n {\n@@ -471,6 +533,11 @@ static void unpack_one(unsigned nr)\n \n \tswitch (type) {\n \tcase OBJ_BLOB:\n+\t\tif (!dry_run && size > big_file_threshold) {\n+\t\t\tstream_blob(size, nr);\n+\t\t\treturn;\n+\t\t}\n+\t\t/* fallthrough */\n \tcase OBJ_COMMIT:\n \tcase OBJ_TREE:\n \tcase OBJ_TAG:\ndiff --git a/t/t5351-unpack-large-objects.sh b/t/t5351-unpack-large-objects.sh\nindex 8d84313221c..461ca060b2b 100755\n--- a/t/t5351-unpack-large-objects.sh\n+++ b/t/t5351-unpack-large-objects.sh\n@@ -9,7 +9,8 @@ test_description='git unpack-objects with large objects'\n \n prepare_dest () {\n \ttest_when_finished \"rm -rf dest.git\" &&\n-\tgit init --bare dest.git\n+\tgit init --bare dest.git &&\n+\tgit -C dest.git config core.bigFileThreshold \"$1\"\n }\n \n test_expect_success \"create large objects (1.5 MB) and PACK\" '\n@@ -26,16 +27,35 @@ test_expect_success 'set memory limitation to 1MB' '\n '\n \n test_expect_success 'unpack-objects failed under memory limitation' '\n-\tprepare_dest &&\n+\tprepare_dest 2m &&\n \ttest_must_fail git -C dest.git unpack-objects <pack-$PACK.pack 2>err &&\n \tgrep \"fatal: attempting to allocate\" err\n '\n \n test_expect_success 'unpack-objects works with memory limitation in dry-run mode' '\n-\tprepare_dest &&\n+\tprepare_dest 2m &&\n \tgit -C dest.git unpack-objects -n <pack-$PACK.pack &&\n \ttest_stdout_line_count = 0 find dest.git/objects -type f &&\n \ttest_dir_is_empty dest.git/objects/pack\n '\n \n+test_expect_success 'unpack big object in stream' '\n+\tprepare_dest 1m &&\n+\tgit -C dest.git unpack-objects <pack-$PACK.pack &&\n+\ttest_dir_is_empty dest.git/objects/pack\n+'\n+\n+test_expect_success 'do not unpack existing large objects' '\n+\tprepare_dest 1m &&\n+\tgit -C dest.git index-pack --stdin <pack-$PACK.pack &&\n+\tgit -C dest.git unpack-objects <pack-$PACK.pack &&\n+\n+\t# The destination came up with the exact same pack...\n+\tDEST_PACK=$(echo dest.git/objects/pack/pack-*.pack) &&\n+\ttest_cmp pack-$PACK.pack $DEST_PACK &&\n+\n+\t# ...and wrote no loose objects\n+\ttest_stdout_line_count = 0 find dest.git/objects -type f ! -name \"pack-*\"\n+'\n+\n test_done\n-- \n2.35.1.1438.g8874c8eeb35\n\n"},{"id":"451670","messageId":"543b06a7-f243-a690-ee81-44f33f7678af@web.de","threadId":"56672","inReplyTo":"patch-v11-3.8-f7b02c307fc-20220319T001411Z-avarab@gmail.com","subject":"Re: [PATCH v11 3/8] object-file.c: refactor write_loose_object() to several steps","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-03-19T10:11:25Z","receivedAt":"2022-03-19T10:11:41Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 19.03.22 um 01:23 schrieb Ævar Arnfjörð Bjarmason:\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> When writing a large blob using \"write_loose_object()\", we have to pass\n> a buffer with the whole content of the blob, and this behavior will\n> consume lots of memory and may cause OOM. We will introduce a stream\n> version function (\"stream_loose_object()\") in later commit to resolve\n> this issue.\n>\n> Before introducing that streaming function, do some refactoring on\n> \"write_loose_object()\" to reuse code for both versions.\n>\n> Rewrite \"write_loose_object()\" as follows:\n>\n>  1. Figure out a path for the (temp) object file. This step is only\n>     used in \"write_loose_object()\".\n>\n>  2. Move common steps for starting to write loose objects into a new\n>     function \"start_loose_object_common()\".\n>\n>  3. Compress data.\n>\n>  4. Move common steps for ending zlib stream into a new function\n>     \"end_loose_object_common()\".\n>\n>  5. Close fd and finalize the object file.\n>\n> Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n> Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n> Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> ---\n>  object-file.c | 129 ++++++++++++++++++++++++++++++++++----------------\n>  1 file changed, 89 insertions(+), 40 deletions(-)\n>\n> diff --git a/object-file.c b/object-file.c\n> index 4c140eda6bf..4fcaf7a36ce 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1943,6 +1943,87 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n>  \treturn fd;\n>  }\n>\n> +/**\n> + * Common steps for loose object writers to start writing loose\n> + * objects:\n> + *\n> + * - Create tmpfile for the loose object.\n> + * - Setup zlib stream for compression.\n> + * - Start to feed header to zlib stream.\n> + *\n> + * Returns a \"fd\", which should later be provided to\n> + * end_loose_object_common().\n> + */\n> +static int start_loose_object_common(struct strbuf *tmp_file,\n> +\t\t\t\t     const char *filename, unsigned flags,\n> +\t\t\t\t     git_zstream *stream,\n> +\t\t\t\t     unsigned char *buf, size_t buflen,\n> +\t\t\t\t     git_hash_ctx *c,\n> +\t\t\t\t     char *hdr, int hdrlen)\n> +{\n> +\tint fd;\n> +\n> +\tfd = create_tmpfile(tmp_file, filename);\n> +\tif (fd < 0) {\n> +\t\tif (flags & HASH_SILENT)\n> +\t\t\treturn -1;\n> +\t\telse if (errno == EACCES)\n> +\t\t\treturn error(_(\"insufficient permission for adding \"\n> +\t\t\t\t       \"an object to repository database %s\"),\n> +\t\t\t\t     get_object_directory());\n> +\t\telse\n> +\t\t\treturn error_errno(\n> +\t\t\t\t_(\"unable to create temporary file\"));\n> +\t}\n> +\n> +\t/*  Setup zlib stream for compression */\n> +\tgit_deflate_init(stream, zlib_compression_level);\n> +\tstream->next_out = buf;\n> +\tstream->avail_out = buflen;\n> +\tthe_hash_algo->init_fn(c);\n> +\n> +\t/*  Start to feed header to zlib stream */\n> +\tstream->next_in = (unsigned char *)hdr;\n> +\tstream->avail_in = hdrlen;\n> +\twhile (git_deflate(stream, 0) == Z_OK)\n> +\t\t; /* nothing */\n> +\tthe_hash_algo->update_fn(c, hdr, hdrlen);\n> +\n> +\treturn fd;\n> +}\n> +\n> +/**\n> + * Common steps for loose object writers to end writing loose objects:\n> + *\n> + * - End the compression of zlib stream.\n> + * - Get the calculated oid to \"parano_oid\".\n> + * - fsync() and close() the \"fd\"\n> + */\n> +static void end_loose_object_common(int fd, int ret, git_hash_ctx *c,\n> +\t\t\t\t    git_zstream *stream,\n> +\t\t\t\t    struct object_id *parano_oid,\n> +\t\t\t\t    const struct object_id *expected_oid,\n> +\t\t\t\t    const char *die_msg1_fmt,\n> +\t\t\t\t    const char *die_msg2_fmt)\n> +{\n> +\tif (ret != Z_STREAM_END)\n> +\t\tdie(_(die_msg1_fmt), ret, expected_oid);\n> +\tret = git_deflate_end_gently(stream);\n> +\tif (ret != Z_OK)\n> +\t\tdie(_(die_msg2_fmt), ret, expected_oid);\n\nstream_loose_object(), added in patch 5, passes NULL as expected_oid,\nbut these die() messages need a valid value.  end_loose_object_common()\nhas more parameters than lines of code in its body.  Inline it to allow\nfully custom messages?\n\n> +\tthe_hash_algo->final_oid_fn(parano_oid, c);\n> +\n> +\t/*\n> +\t * We already did a write_buffer() to the \"fd\", let's fsync()\n> +\t * and close().\n> +\t *\n> +\t * We might still die() on a subsequent sanity check, but\n> +\t * let's not add to that confusion by not flushing any\n> +\t * outstanding writes to disk first.\n> +\t */\n> +\tclose_loose_object(fd);\n> +}\n> +\n>  static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \t\t\t      int hdrlen, const void *buf, unsigned long len,\n>  \t\t\t      time_t mtime, unsigned flags)\n> @@ -1957,28 +2038,11 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>\n>  \tloose_object_path(the_repository, &filename, oid);\n>\n> -\tfd = create_tmpfile(&tmp_file, filename.buf);\n> -\tif (fd < 0) {\n> -\t\tif (flags & HASH_SILENT)\n> -\t\t\treturn -1;\n> -\t\telse if (errno == EACCES)\n> -\t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n> -\t\telse\n> -\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n> -\t}\n> -\n> -\t/* Set it up */\n> -\tgit_deflate_init(&stream, zlib_compression_level);\n> -\tstream.next_out = compressed;\n> -\tstream.avail_out = sizeof(compressed);\n> -\tthe_hash_algo->init_fn(&c);\n> -\n> -\t/* First header.. */\n> -\tstream.next_in = (unsigned char *)hdr;\n> -\tstream.avail_in = hdrlen;\n> -\twhile (git_deflate(&stream, 0) == Z_OK)\n> -\t\t; /* nothing */\n> -\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n> +\tfd = start_loose_object_common(&tmp_file, filename.buf, flags,\n> +\t\t\t\t       &stream, compressed, sizeof(compressed),\n> +\t\t\t\t       &c, hdr, hdrlen);\n> +\tif (fd < 0)\n> +\t\treturn -1;\n>\n>  \t/* Then the data itself.. */\n>  \tstream.next_in = (void *)buf;\n> @@ -1993,24 +2057,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \t\tstream.avail_out = sizeof(compressed);\n>  \t} while (ret == Z_OK);\n>\n> -\tif (ret != Z_STREAM_END)\n> -\t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n> -\t\t    ret);\n> -\tret = git_deflate_end_gently(&stream);\n> -\tif (ret != Z_OK)\n> -\t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n> -\t\t    ret);\n> -\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n> -\n> -\t/*\n> -\t * We already did a write_buffer() to the \"fd\", let's fsync()\n> -\t * and close().\n> -\t *\n> -\t * We might still die() on a subsequent sanity check, but\n> -\t * let's not add to that confusion by not flushing any\n> -\t * outstanding writes to disk first.\n> -\t */\n> -\tclose_loose_object(fd);\n> +\tend_loose_object_common(fd, ret, &c, &stream, &parano_oid, oid,\n> +\t\t\t\tN_(\"unable to deflate new object %s (%d)\"),\n> +\t\t\t\tN_(\"deflateEnd on object %s failed (%d)\"));\n>\n>  \tif (!oideq(oid, &parano_oid))\n>  \t\tdie(_(\"confused by unstable object source data for %s\"),\n"},{"id":"452561","messageId":"cover-v12-0.8-00000000000-20220329T135446Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v11-0.8-00000000000-20220319T001411Z-avarab@gmail.com","subject":"[PATCH v12 0/8] unpack-objects: support streaming blobs to disk","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-29T13:56:05Z","receivedAt":"2022-03-29T13:56:27Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"This series by Han Xin was waiting on some in-flight patches that\nlanded in 430883a70c7 (Merge branch 'ab/object-file-api-updates',\n2022-03-16).\n\nThis series teaches \"git unpack-objects\" to stream objects larger than\ncore.bigFileThreshold to disk. As 8/8 shows streaming e.g. a 100MB\nblob now uses ~5MB of memory instead of ~105MB. This streaming method\nis slower if you've got memory to handle the blobs in-core, but if you\ndon't it allows you to unpack objects at all, as you might otherwise\nOOM.\n\nChanges since v10[1]:\n\n * René rightly spotted that the end_loose_object_common() function\n   was feeding NULL to a format. That's now fixed, and parts of that\n   function were pulled out into the two callers to make the trade-off\n   of factoring that logic out worth it.\n\n * This topic conflicts with ns/batch-fsync in \"seen\" (see below). I\n   moved an inline comment on close_loose_object() around to make the\n   conflict easier (and it's better placed with the function anyway,\n   as we'll get two callers of it).\n\nConflict this is the --remerge-diff with \"seen\" after resolving the\nconflict. Botht textual and semantic (there's a new caller in this\ntopic) conflicts are caught by the compiler:\n\n\tdiff --git a/object-file.c b/object-file.c\n\tremerge CONFLICT (content): Merge conflict in object-file.c\n\tindex 9c640f1f39d..6068f8ec6c4 100644\n\t--- a/object-file.c\n\t+++ b/object-file.c\n\t@@ -1887,7 +1887,6 @@ void hash_object_file(const struct git_hash_algo *algo, const void *buf,\n\t \thash_object_file_literally(algo, buf, len, type_name(type), oid);\n\t }\n\n\t-<<<<<<< 34ee6a28a54 (unpack-objects: use stream_loose_object() to unpack large objects)\n\t /*\n\t  * We already did a write_buffer() to the \"fd\", let's fsync()\n\t  * and close().\n\t@@ -1896,11 +1895,7 @@ void hash_object_file(const struct git_hash_algo *algo, const void *buf,\n\t  * subsequent sanity check, but let's not add to that confusion by not\n\t  * flushing any outstanding writes to disk first.\n\t  */\n\t-static void close_loose_object(int fd)\n\t-=======\n\t-/* Finalize a file on disk, and close it. */\n\t static void close_loose_object(int fd, const char *filename)\n\t->>>>>>> b1423c89b5a (Merge branch 'ab/reftable-aix-xlc-12' into seen)\n\t {\n\t \tif (the_repository->objects->odb->will_destroy)\n\t \t\tgoto out;\n\t@@ -2093,17 +2088,12 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n\t \tret = end_loose_object_common(&c, &stream, &parano_oid);\n\t \tif (ret != Z_OK)\n\t \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid), ret);\n\t-\tclose_loose_object(fd);\n\t+\tclose_loose_object(fd, tmp_file.buf);\n\n\t \tif (!oideq(oid, &parano_oid))\n\t \t\tdie(_(\"confused by unstable object source data for %s\"),\n\t \t\t    oid_to_hex(oid));\n\n\t-<<<<<<< 34ee6a28a54 (unpack-objects: use stream_loose_object() to unpack large objects)\n\t-=======\n\t-\tclose_loose_object(fd, tmp_file.buf);\n\t-\n\t->>>>>>> b1423c89b5a (Merge branch 'ab/reftable-aix-xlc-12' into seen)\n\t \tif (mtime) {\n\t \t\tstruct utimbuf utb;\n\t \t\tutb.actime = mtime;\n\t@@ -2206,7 +2196,7 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n\t \tret = end_loose_object_common(&c, &stream, oid);\n\t \tif (ret != Z_OK)\n\t \t\tdie(_(\"deflateEnd on stream object failed (%d)\"), ret);\n\t-\tclose_loose_object(fd);\n\t+\tclose_loose_object(fd, tmp_file.buf);\n\n\t \tif (freshen_packed_object(oid) || freshen_loose_object(oid)) {\n\t \t\tunlink_or_warn(tmp_file.buf);\n\n1. https://lore.kernel.org/git/cover-v11-0.8-00000000000-20220319T001411Z-avarab@gmail.com/\n\nHan Xin (4):\n  unpack-objects: low memory footprint for get_data() in dry_run mode\n  object-file.c: refactor write_loose_object() to several steps\n  object-file.c: add \"stream_loose_object()\" to handle large object\n  unpack-objects: use stream_loose_object() to unpack large objects\n\nÆvar Arnfjörð Bjarmason (4):\n  object-file.c: do fsync() and close() before post-write die()\n  object-file.c: factor out deflate part of write_loose_object()\n  core doc: modernize core.bigFileThreshold documentation\n  unpack-objects: refactor away unpack_non_delta_entry()\n\n Documentation/config/core.txt   |  33 +++--\n builtin/unpack-objects.c        | 109 +++++++++++---\n object-file.c                   | 246 +++++++++++++++++++++++++++-----\n object-store.h                  |   8 ++\n t/t5351-unpack-large-objects.sh |  61 ++++++++\n 5 files changed, 396 insertions(+), 61 deletions(-)\n create mode 100755 t/t5351-unpack-large-objects.sh\n\nRange-diff against v11:\n1:  2103d5bfd96 = 1:  e95f6a1cfb6 unpack-objects: low memory footprint for get_data() in dry_run mode\n2:  6acd8759772 ! 2:  54060eb8c6b object-file.c: do fsync() and close() before post-write die()\n    @@ Commit message\n         Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n      ## object-file.c ##\n    +@@ object-file.c: void hash_object_file(const struct git_hash_algo *algo, const void *buf,\n    + \thash_object_file_literally(algo, buf, len, type_name(type), oid);\n    + }\n    + \n    +-/* Finalize a file on disk, and close it. */\n    ++/*\n    ++ * We already did a write_buffer() to the \"fd\", let's fsync()\n    ++ * and close().\n    ++ *\n    ++ * Finalize a file on disk, and close it. We might still die() on a\n    ++ * subsequent sanity check, but let's not add to that confusion by not\n    ++ * flushing any outstanding writes to disk first.\n    ++ */\n    + static void close_loose_object(int fd)\n    + {\n    + \tif (the_repository->objects->odb->will_destroy)\n     @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *hdr,\n      \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n      \t\t    ret);\n      \tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n    -+\n    -+\t/*\n    -+\t * We already did a write_buffer() to the \"fd\", let's fsync()\n    -+\t * and close().\n    -+\t *\n    -+\t * We might still die() on a subsequent sanity check, but\n    -+\t * let's not add to that confusion by not flushing any\n    -+\t * outstanding writes to disk first.\n    -+\t */\n     +\tclose_loose_object(fd);\n     +\n      \tif (!oideq(oid, &parano_oid))\n3:  f7b02c307fc ! 3:  3dcaa5d6589 object-file.c: refactor write_loose_object() to several steps\n    @@ object-file.c: static int create_tmpfile(struct strbuf *tmp, const char *filenam\n     + * Common steps for loose object writers to end writing loose objects:\n     + *\n     + * - End the compression of zlib stream.\n    -+ * - Get the calculated oid to \"parano_oid\".\n    ++ * - Get the calculated oid to \"oid\".\n     + * - fsync() and close() the \"fd\"\n     + */\n    -+static void end_loose_object_common(int fd, int ret, git_hash_ctx *c,\n    -+\t\t\t\t    git_zstream *stream,\n    -+\t\t\t\t    struct object_id *parano_oid,\n    -+\t\t\t\t    const struct object_id *expected_oid,\n    -+\t\t\t\t    const char *die_msg1_fmt,\n    -+\t\t\t\t    const char *die_msg2_fmt)\n    ++static int end_loose_object_common(git_hash_ctx *c, git_zstream *stream,\n    ++\t\t\t\t   struct object_id *oid)\n     +{\n    -+\tif (ret != Z_STREAM_END)\n    -+\t\tdie(_(die_msg1_fmt), ret, expected_oid);\n    ++\tint ret;\n    ++\n     +\tret = git_deflate_end_gently(stream);\n     +\tif (ret != Z_OK)\n    -+\t\tdie(_(die_msg2_fmt), ret, expected_oid);\n    -+\tthe_hash_algo->final_oid_fn(parano_oid, c);\n    ++\t\treturn ret;\n    ++\tthe_hash_algo->final_oid_fn(oid, c);\n     +\n    -+\t/*\n    -+\t * We already did a write_buffer() to the \"fd\", let's fsync()\n    -+\t * and close().\n    -+\t *\n    -+\t * We might still die() on a subsequent sanity check, but\n    -+\t * let's not add to that confusion by not flushing any\n    -+\t * outstanding writes to disk first.\n    -+\t */\n    -+\tclose_loose_object(fd);\n    ++\treturn Z_OK;\n     +}\n     +\n      static int write_loose_object(const struct object_id *oid, char *hdr,\n    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *\n      \t/* Then the data itself.. */\n      \tstream.next_in = (void *)buf;\n     @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *hdr,\n    - \t\tstream.avail_out = sizeof(compressed);\n    - \t} while (ret == Z_OK);\n    - \n    --\tif (ret != Z_STREAM_END)\n    --\t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n    --\t\t    ret);\n    + \tif (ret != Z_STREAM_END)\n    + \t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n    + \t\t    ret);\n     -\tret = git_deflate_end_gently(&stream);\n    --\tif (ret != Z_OK)\n    ++\tret = end_loose_object_common(&c, &stream, &parano_oid);\n    + \tif (ret != Z_OK)\n     -\t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n     -\t\t    ret);\n     -\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n    --\n    --\t/*\n    --\t * We already did a write_buffer() to the \"fd\", let's fsync()\n    --\t * and close().\n    --\t *\n    --\t * We might still die() on a subsequent sanity check, but\n    --\t * let's not add to that confusion by not flushing any\n    --\t * outstanding writes to disk first.\n    --\t */\n    --\tclose_loose_object(fd);\n    -+\tend_loose_object_common(fd, ret, &c, &stream, &parano_oid, oid,\n    -+\t\t\t\tN_(\"unable to deflate new object %s (%d)\"),\n    -+\t\t\t\tN_(\"deflateEnd on object %s failed (%d)\"));\n    ++\t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid), ret);\n    + \tclose_loose_object(fd);\n      \n      \tif (!oideq(oid, &parano_oid))\n    - \t\tdie(_(\"confused by unstable object source data for %s\"),\n4:  20d97cc2605 ! 4:  03f4e91ac89 object-file.c: factor out deflate part of write_loose_object()\n    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *\n     +\t\t\t\t\t\tcompressed, sizeof(compressed));\n      \t} while (ret == Z_OK);\n      \n    - \tend_loose_object_common(fd, ret, &c, &stream, &parano_oid, oid,\n    + \tif (ret != Z_STREAM_END)\n5:  db40f4160c4 ! 5:  3d64cf1cf33 object-file.c: add \"stream_loose_object()\" to handle large object\n    @@ object-file.c: static int freshen_packed_object(const struct object_id *oid)\n     +\t *  - End the compression of zlib stream.\n     +\t *  - Get the calculated oid.\n     +\t */\n    -+\tend_loose_object_common(fd, ret, &c, &stream, oid, NULL,\n    -+\t\t\t\tN_(\"unable to stream deflate new object (%d)\"),\n    -+\t\t\t\tN_(\"deflateEnd on stream object failed (%d)\"));\n    ++\tif (ret != Z_STREAM_END)\n    ++\t\tdie(_(\"unable to stream deflate new object (%d)\"), ret);\n    ++\tret = end_loose_object_common(&c, &stream, oid);\n    ++\tif (ret != Z_OK)\n    ++\t\tdie(_(\"deflateEnd on stream object failed (%d)\"), ret);\n    ++\tclose_loose_object(fd);\n     +\n     +\tif (freshen_packed_object(oid) || freshen_loose_object(oid)) {\n     +\t\tunlink_or_warn(tmp_file.buf);\n6:  d8ae2eadb98 = 6:  33ffcbbc1f0 core doc: modernize core.bigFileThreshold documentation\n7:  2b403e7cd9c = 7:  11f7aa026b4 unpack-objects: refactor away unpack_non_delta_entry()\n8:  5eded902496 = 8:  34ee6a28a54 unpack-objects: use stream_loose_object() to unpack large objects\n-- \n2.35.1.1548.g36973b18e52\n\n"},{"id":"452562","messageId":"patch-v12-2.8-54060eb8c6b-20220329T135446Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v12-0.8-00000000000-20220329T135446Z-avarab@gmail.com","subject":"[PATCH v12 2/8] object-file.c: do fsync() and close() before post-write die()","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-29T13:56:07Z","receivedAt":"2022-03-29T13:56:29Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Change write_loose_object() to do an fsync() and close() before the\noideq() sanity check at the end. This change re-joins code that was\nsplit up by the die() sanity check added in 748af44c63e (sha1_file: be\nparanoid when creating loose objects, 2010-02-21).\n\nI don't think that this change matters in itself, if we called die()\nit was possible that our data wouldn't fully make it to disk, but in\nany case we were writing data that we'd consider corrupted. It's\npossible that a subsequent \"git fsck\" will be less confused now.\n\nThe real reason to make this change is that in a subsequent commit\nwe'll split this code in write_loose_object() into a utility function,\nall its callers will want the preceding sanity checks, but not the\n\"oideq\" check. By moving the close_loose_object() earlier it'll be\neasier to reason about the introduction of the utility function.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n object-file.c | 13 ++++++++++---\n 1 file changed, 10 insertions(+), 3 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 62ebe236c90..5da458eccbf 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1886,7 +1886,14 @@ void hash_object_file(const struct git_hash_algo *algo, const void *buf,\n \thash_object_file_literally(algo, buf, len, type_name(type), oid);\n }\n \n-/* Finalize a file on disk, and close it. */\n+/*\n+ * We already did a write_buffer() to the \"fd\", let's fsync()\n+ * and close().\n+ *\n+ * Finalize a file on disk, and close it. We might still die() on a\n+ * subsequent sanity check, but let's not add to that confusion by not\n+ * flushing any outstanding writes to disk first.\n+ */\n static void close_loose_object(int fd)\n {\n \tif (the_repository->objects->odb->will_destroy)\n@@ -2006,12 +2013,12 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n \t\t    ret);\n \tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n+\tclose_loose_object(fd);\n+\n \tif (!oideq(oid, &parano_oid))\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n \n-\tclose_loose_object(fd);\n-\n \tif (mtime) {\n \t\tstruct utimbuf utb;\n \t\tutb.actime = mtime;\n-- \n2.35.1.1548.g36973b18e52\n\n"},{"id":"452563","messageId":"patch-v12-1.8-e95f6a1cfb6-20220329T135446Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v12-0.8-00000000000-20220329T135446Z-avarab@gmail.com","subject":"[PATCH v12 1/8] unpack-objects: low memory footprint for get_data() in dry_run mode","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-29T13:56:06Z","receivedAt":"2022-03-29T13:56:30Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nAs the name implies, \"get_data(size)\" will allocate and return a given\namount of memory. Allocating memory for a large blob object may cause the\nsystem to run out of memory. Before preparing to replace calling of\n\"get_data()\" to unpack large blob objects in latter commits, refactor\n\"get_data()\" to reduce memory footprint for dry_run mode.\n\nBecause in dry_run mode, \"get_data()\" is only used to check the\nintegrity of data, and the returned buffer is not used at all, we can\nallocate a smaller buffer and reuse it as zstream output. Therefore,\nin dry_run mode, \"get_data()\" will release the allocated buffer and\nreturn NULL instead of returning garbage data.\n\nThe \"find [...]objects/?? -type f | wc -l\" test idiom being used here\nis adapted from the same \"find\" use added to another test in\nd9545c7f465 (fast-import: implement unpack limit, 2016-04-25).\n\nSuggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n builtin/unpack-objects.c        | 34 ++++++++++++++++++---------\n t/t5351-unpack-large-objects.sh | 41 +++++++++++++++++++++++++++++++++\n 2 files changed, 64 insertions(+), 11 deletions(-)\n create mode 100755 t/t5351-unpack-large-objects.sh\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex dbeb0680a58..e3d30025979 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -96,15 +96,26 @@ static void use(int bytes)\n \tdisplay_throughput(progress, consumed_bytes);\n }\n \n+/*\n+ * Decompress zstream from stdin and return specific size of data.\n+ * The caller is responsible to free the returned buffer.\n+ *\n+ * But for dry_run mode, \"get_data()\" is only used to check the\n+ * integrity of data, and the returned buffer is not used at all.\n+ * Therefore, in dry_run mode, \"get_data()\" will release the small\n+ * allocated buffer which is reused to hold temporary zstream output\n+ * and return NULL instead of returning garbage data.\n+ */\n static void *get_data(unsigned long size)\n {\n \tgit_zstream stream;\n-\tvoid *buf = xmallocz(size);\n+\tunsigned long bufsize = dry_run && size > 8192 ? 8192 : size;\n+\tvoid *buf = xmallocz(bufsize);\n \n \tmemset(&stream, 0, sizeof(stream));\n \n \tstream.next_out = buf;\n-\tstream.avail_out = size;\n+\tstream.avail_out = bufsize;\n \tstream.next_in = fill(1);\n \tstream.avail_in = len;\n \tgit_inflate_init(&stream);\n@@ -124,8 +135,15 @@ static void *get_data(unsigned long size)\n \t\t}\n \t\tstream.next_in = fill(1);\n \t\tstream.avail_in = len;\n+\t\tif (dry_run) {\n+\t\t\t/* reuse the buffer in dry_run mode */\n+\t\t\tstream.next_out = buf;\n+\t\t\tstream.avail_out = bufsize;\n+\t\t}\n \t}\n \tgit_inflate_end(&stream);\n+\tif (dry_run)\n+\t\tFREE_AND_NULL(buf);\n \treturn buf;\n }\n \n@@ -325,10 +343,8 @@ static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n {\n \tvoid *buf = get_data(size);\n \n-\tif (!dry_run && buf)\n+\tif (buf)\n \t\twrite_object(nr, type, buf, size);\n-\telse\n-\t\tfree(buf);\n }\n \n static int resolve_against_held(unsigned nr, const struct object_id *base,\n@@ -358,10 +374,8 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\toidread(&base_oid, fill(the_hash_algo->rawsz));\n \t\tuse(the_hash_algo->rawsz);\n \t\tdelta_data = get_data(delta_size);\n-\t\tif (dry_run || !delta_data) {\n-\t\t\tfree(delta_data);\n+\t\tif (!delta_data)\n \t\t\treturn;\n-\t\t}\n \t\tif (has_object_file(&base_oid))\n \t\t\t; /* Ok we have this one */\n \t\telse if (resolve_against_held(nr, &base_oid,\n@@ -397,10 +411,8 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\t\tdie(\"offset value out of bound for delta base object\");\n \n \t\tdelta_data = get_data(delta_size);\n-\t\tif (dry_run || !delta_data) {\n-\t\t\tfree(delta_data);\n+\t\tif (!delta_data)\n \t\t\treturn;\n-\t\t}\n \t\tlo = 0;\n \t\thi = nr;\n \t\twhile (lo < hi) {\ndiff --git a/t/t5351-unpack-large-objects.sh b/t/t5351-unpack-large-objects.sh\nnew file mode 100755\nindex 00000000000..8d84313221c\n--- /dev/null\n+++ b/t/t5351-unpack-large-objects.sh\n@@ -0,0 +1,41 @@\n+#!/bin/sh\n+#\n+# Copyright (c) 2022 Han Xin\n+#\n+\n+test_description='git unpack-objects with large objects'\n+\n+. ./test-lib.sh\n+\n+prepare_dest () {\n+\ttest_when_finished \"rm -rf dest.git\" &&\n+\tgit init --bare dest.git\n+}\n+\n+test_expect_success \"create large objects (1.5 MB) and PACK\" '\n+\ttest-tool genrandom foo 1500000 >big-blob &&\n+\ttest_commit --append foo big-blob &&\n+\ttest-tool genrandom bar 1500000 >big-blob &&\n+\ttest_commit --append bar big-blob &&\n+\tPACK=$(echo HEAD | git pack-objects --revs pack)\n+'\n+\n+test_expect_success 'set memory limitation to 1MB' '\n+\tGIT_ALLOC_LIMIT=1m &&\n+\texport GIT_ALLOC_LIMIT\n+'\n+\n+test_expect_success 'unpack-objects failed under memory limitation' '\n+\tprepare_dest &&\n+\ttest_must_fail git -C dest.git unpack-objects <pack-$PACK.pack 2>err &&\n+\tgrep \"fatal: attempting to allocate\" err\n+'\n+\n+test_expect_success 'unpack-objects works with memory limitation in dry-run mode' '\n+\tprepare_dest &&\n+\tgit -C dest.git unpack-objects -n <pack-$PACK.pack &&\n+\ttest_stdout_line_count = 0 find dest.git/objects -type f &&\n+\ttest_dir_is_empty dest.git/objects/pack\n+'\n+\n+test_done\n-- \n2.35.1.1548.g36973b18e52\n\n"},{"id":"452564","messageId":"patch-v12-3.8-3dcaa5d6589-20220329T135446Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v12-0.8-00000000000-20220329T135446Z-avarab@gmail.com","subject":"[PATCH v12 3/8] object-file.c: refactor write_loose_object() to several steps","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-29T13:56:08Z","receivedAt":"2022-03-29T13:56:32Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWhen writing a large blob using \"write_loose_object()\", we have to pass\na buffer with the whole content of the blob, and this behavior will\nconsume lots of memory and may cause OOM. We will introduce a stream\nversion function (\"stream_loose_object()\") in later commit to resolve\nthis issue.\n\nBefore introducing that streaming function, do some refactoring on\n\"write_loose_object()\" to reuse code for both versions.\n\nRewrite \"write_loose_object()\" as follows:\n\n 1. Figure out a path for the (temp) object file. This step is only\n    used in \"write_loose_object()\".\n\n 2. Move common steps for starting to write loose objects into a new\n    function \"start_loose_object_common()\".\n\n 3. Compress data.\n\n 4. Move common steps for ending zlib stream into a new function\n    \"end_loose_object_common()\".\n\n 5. Close fd and finalize the object file.\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n object-file.c | 102 +++++++++++++++++++++++++++++++++++++-------------\n 1 file changed, 76 insertions(+), 26 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 5da458eccbf..7f160929e00 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1955,6 +1955,75 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n \treturn fd;\n }\n \n+/**\n+ * Common steps for loose object writers to start writing loose\n+ * objects:\n+ *\n+ * - Create tmpfile for the loose object.\n+ * - Setup zlib stream for compression.\n+ * - Start to feed header to zlib stream.\n+ *\n+ * Returns a \"fd\", which should later be provided to\n+ * end_loose_object_common().\n+ */\n+static int start_loose_object_common(struct strbuf *tmp_file,\n+\t\t\t\t     const char *filename, unsigned flags,\n+\t\t\t\t     git_zstream *stream,\n+\t\t\t\t     unsigned char *buf, size_t buflen,\n+\t\t\t\t     git_hash_ctx *c,\n+\t\t\t\t     char *hdr, int hdrlen)\n+{\n+\tint fd;\n+\n+\tfd = create_tmpfile(tmp_file, filename);\n+\tif (fd < 0) {\n+\t\tif (flags & HASH_SILENT)\n+\t\t\treturn -1;\n+\t\telse if (errno == EACCES)\n+\t\t\treturn error(_(\"insufficient permission for adding \"\n+\t\t\t\t       \"an object to repository database %s\"),\n+\t\t\t\t     get_object_directory());\n+\t\telse\n+\t\t\treturn error_errno(\n+\t\t\t\t_(\"unable to create temporary file\"));\n+\t}\n+\n+\t/*  Setup zlib stream for compression */\n+\tgit_deflate_init(stream, zlib_compression_level);\n+\tstream->next_out = buf;\n+\tstream->avail_out = buflen;\n+\tthe_hash_algo->init_fn(c);\n+\n+\t/*  Start to feed header to zlib stream */\n+\tstream->next_in = (unsigned char *)hdr;\n+\tstream->avail_in = hdrlen;\n+\twhile (git_deflate(stream, 0) == Z_OK)\n+\t\t; /* nothing */\n+\tthe_hash_algo->update_fn(c, hdr, hdrlen);\n+\n+\treturn fd;\n+}\n+\n+/**\n+ * Common steps for loose object writers to end writing loose objects:\n+ *\n+ * - End the compression of zlib stream.\n+ * - Get the calculated oid to \"oid\".\n+ * - fsync() and close() the \"fd\"\n+ */\n+static int end_loose_object_common(git_hash_ctx *c, git_zstream *stream,\n+\t\t\t\t   struct object_id *oid)\n+{\n+\tint ret;\n+\n+\tret = git_deflate_end_gently(stream);\n+\tif (ret != Z_OK)\n+\t\treturn ret;\n+\tthe_hash_algo->final_oid_fn(oid, c);\n+\n+\treturn Z_OK;\n+}\n+\n static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\t\t      int hdrlen, const void *buf, unsigned long len,\n \t\t\t      time_t mtime, unsigned flags)\n@@ -1969,28 +2038,11 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \n \tloose_object_path(the_repository, &filename, oid);\n \n-\tfd = create_tmpfile(&tmp_file, filename.buf);\n-\tif (fd < 0) {\n-\t\tif (flags & HASH_SILENT)\n-\t\t\treturn -1;\n-\t\telse if (errno == EACCES)\n-\t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n-\t\telse\n-\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n-\t}\n-\n-\t/* Set it up */\n-\tgit_deflate_init(&stream, zlib_compression_level);\n-\tstream.next_out = compressed;\n-\tstream.avail_out = sizeof(compressed);\n-\tthe_hash_algo->init_fn(&c);\n-\n-\t/* First header.. */\n-\tstream.next_in = (unsigned char *)hdr;\n-\tstream.avail_in = hdrlen;\n-\twhile (git_deflate(&stream, 0) == Z_OK)\n-\t\t; /* nothing */\n-\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n+\tfd = start_loose_object_common(&tmp_file, filename.buf, flags,\n+\t\t\t\t       &stream, compressed, sizeof(compressed),\n+\t\t\t\t       &c, hdr, hdrlen);\n+\tif (fd < 0)\n+\t\treturn -1;\n \n \t/* Then the data itself.. */\n \tstream.next_in = (void *)buf;\n@@ -2008,11 +2060,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tif (ret != Z_STREAM_END)\n \t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n \t\t    ret);\n-\tret = git_deflate_end_gently(&stream);\n+\tret = end_loose_object_common(&c, &stream, &parano_oid);\n \tif (ret != Z_OK)\n-\t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n-\t\t    ret);\n-\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n+\t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid), ret);\n \tclose_loose_object(fd);\n \n \tif (!oideq(oid, &parano_oid))\n-- \n2.35.1.1548.g36973b18e52\n\n"},{"id":"452565","messageId":"patch-v12-4.8-03f4e91ac89-20220329T135446Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v12-0.8-00000000000-20220329T135446Z-avarab@gmail.com","subject":"[PATCH v12 4/8] object-file.c: factor out deflate part of write_loose_object()","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-29T13:56:09Z","receivedAt":"2022-03-29T13:56:34Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Split out the part of write_loose_object() that deals with calling\ngit_deflate() into a utility function, a subsequent commit will\nintroduce another function that'll make use of it.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n object-file.c | 31 +++++++++++++++++++++++++------\n 1 file changed, 25 insertions(+), 6 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 7f160929e00..6e2f2264f8c 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2004,6 +2004,28 @@ static int start_loose_object_common(struct strbuf *tmp_file,\n \treturn fd;\n }\n \n+/**\n+ * Common steps for the inner git_deflate() loop for writing loose\n+ * objects. Returns what git_deflate() returns.\n+ */\n+static int write_loose_object_common(git_hash_ctx *c,\n+\t\t\t\t     git_zstream *stream, const int flush,\n+\t\t\t\t     unsigned char *in0, const int fd,\n+\t\t\t\t     unsigned char *compressed,\n+\t\t\t\t     const size_t compressed_len)\n+{\n+\tint ret;\n+\n+\tret = git_deflate(stream, flush ? Z_FINISH : 0);\n+\tthe_hash_algo->update_fn(c, in0, stream->next_in - in0);\n+\tif (write_buffer(fd, compressed, stream->next_out - compressed) < 0)\n+\t\tdie(_(\"unable to write loose object file\"));\n+\tstream->next_out = compressed;\n+\tstream->avail_out = compressed_len;\n+\n+\treturn ret;\n+}\n+\n /**\n  * Common steps for loose object writers to end writing loose objects:\n  *\n@@ -2049,12 +2071,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tstream.avail_in = len;\n \tdo {\n \t\tunsigned char *in0 = stream.next_in;\n-\t\tret = git_deflate(&stream, Z_FINISH);\n-\t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n-\t\tif (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n-\t\t\tdie(_(\"unable to write loose object file\"));\n-\t\tstream.next_out = compressed;\n-\t\tstream.avail_out = sizeof(compressed);\n+\n+\t\tret = write_loose_object_common(&c, &stream, 1, in0, fd,\n+\t\t\t\t\t\tcompressed, sizeof(compressed));\n \t} while (ret == Z_OK);\n \n \tif (ret != Z_STREAM_END)\n-- \n2.35.1.1548.g36973b18e52\n\n"},{"id":"452566","messageId":"patch-v12-5.8-3d64cf1cf33-20220329T135446Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v12-0.8-00000000000-20220329T135446Z-avarab@gmail.com","subject":"[PATCH v12 5/8] object-file.c: add \"stream_loose_object()\" to handle large object","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-29T13:56:10Z","receivedAt":"2022-03-29T13:56:42Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nIf we want unpack and write a loose object using \"write_loose_object\",\nwe have to feed it with a buffer with the same size of the object, which\nwill consume lots of memory and may cause OOM. This can be improved by\nfeeding data to \"stream_loose_object()\" in a stream.\n\nAdd a new function \"stream_loose_object()\", which is a stream version of\n\"write_loose_object()\" but with a low memory footprint. We will use this\nfunction to unpack large blob object in later commit.\n\nAnother difference with \"write_loose_object()\" is that we have no chance\nto run \"write_object_file_prepare()\" to calculate the oid in advance.\nIn \"write_loose_object()\", we know the oid and we can write the\ntemporary file in the same directory as the final object, but for an\nobject with an undetermined oid, we don't know the exact directory for\nthe object.\n\nStill, we need to save the temporary file we're preparing\nsomewhere. We'll do that in the top-level \".git/objects/\"\ndirectory (or whatever \"GIT_OBJECT_DIRECTORY\" is set to). Once we've\nstreamed it we'll know the OID, and will move it to its canonical\npath.\n\n\"freshen_packed_object()\" or \"freshen_loose_object()\" will be called\ninside \"stream_loose_object()\" after obtaining the \"oid\".\n\nHelped-by: René Scharfe <l.s.r@web.de>\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n object-file.c  | 100 +++++++++++++++++++++++++++++++++++++++++++++++++\n object-store.h |   8 ++++\n 2 files changed, 108 insertions(+)\n\ndiff --git a/object-file.c b/object-file.c\nindex 6e2f2264f8c..2be2bae9afa 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2118,6 +2118,106 @@ static int freshen_packed_object(const struct object_id *oid)\n \treturn 1;\n }\n \n+int stream_loose_object(struct input_stream *in_stream, size_t len,\n+\t\t\tstruct object_id *oid)\n+{\n+\tint fd, ret, err = 0, flush = 0;\n+\tunsigned char compressed[4096];\n+\tgit_zstream stream;\n+\tgit_hash_ctx c;\n+\tstruct strbuf tmp_file = STRBUF_INIT;\n+\tstruct strbuf filename = STRBUF_INIT;\n+\tint dirlen;\n+\tchar hdr[MAX_HEADER_LEN];\n+\tint hdrlen;\n+\n+\t/* Since oid is not determined, save tmp file to odb path. */\n+\tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n+\thdrlen = format_object_header(hdr, sizeof(hdr), OBJ_BLOB, len);\n+\n+\t/*\n+\t * Common steps for write_loose_object and stream_loose_object to\n+\t * start writing loose objects:\n+\t *\n+\t *  - Create tmpfile for the loose object.\n+\t *  - Setup zlib stream for compression.\n+\t *  - Start to feed header to zlib stream.\n+\t */\n+\tfd = start_loose_object_common(&tmp_file, filename.buf, 0,\n+\t\t\t\t       &stream, compressed, sizeof(compressed),\n+\t\t\t\t       &c, hdr, hdrlen);\n+\tif (fd < 0) {\n+\t\terr = -1;\n+\t\tgoto cleanup;\n+\t}\n+\n+\t/* Then the data itself.. */\n+\tdo {\n+\t\tunsigned char *in0 = stream.next_in;\n+\n+\t\tif (!stream.avail_in && !in_stream->is_finished) {\n+\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n+\t\t\tstream.next_in = (void *)in;\n+\t\t\tin0 = (unsigned char *)in;\n+\t\t\t/* All data has been read. */\n+\t\t\tif (in_stream->is_finished)\n+\t\t\t\tflush = 1;\n+\t\t}\n+\t\tret = write_loose_object_common(&c, &stream, flush, in0, fd,\n+\t\t\t\t\t\tcompressed, sizeof(compressed));\n+\t\t/*\n+\t\t * Unlike write_loose_object(), we do not have the entire\n+\t\t * buffer. If we get Z_BUF_ERROR due to too few input bytes,\n+\t\t * then we'll replenish them in the next input_stream->read()\n+\t\t * call when we loop.\n+\t\t */\n+\t} while (ret == Z_OK || ret == Z_BUF_ERROR);\n+\n+\tif (stream.total_in != len + hdrlen)\n+\t\tdie(_(\"write stream object %ld != %\"PRIuMAX), stream.total_in,\n+\t\t    (uintmax_t)len + hdrlen);\n+\n+\t/* Common steps for write_loose_object and stream_loose_object to\n+\t * end writing loose oject:\n+\t *\n+\t *  - End the compression of zlib stream.\n+\t *  - Get the calculated oid.\n+\t */\n+\tif (ret != Z_STREAM_END)\n+\t\tdie(_(\"unable to stream deflate new object (%d)\"), ret);\n+\tret = end_loose_object_common(&c, &stream, oid);\n+\tif (ret != Z_OK)\n+\t\tdie(_(\"deflateEnd on stream object failed (%d)\"), ret);\n+\tclose_loose_object(fd);\n+\n+\tif (freshen_packed_object(oid) || freshen_loose_object(oid)) {\n+\t\tunlink_or_warn(tmp_file.buf);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tloose_object_path(the_repository, &filename, oid);\n+\n+\t/* We finally know the object path, and create the missing dir. */\n+\tdirlen = directory_size(filename.buf);\n+\tif (dirlen) {\n+\t\tstruct strbuf dir = STRBUF_INIT;\n+\t\tstrbuf_add(&dir, filename.buf, dirlen);\n+\n+\t\tif (mkdir_in_gitdir(dir.buf) && errno != EEXIST) {\n+\t\t\terr = error_errno(_(\"unable to create directory %s\"), dir.buf);\n+\t\t\tstrbuf_release(&dir);\n+\t\t\tgoto cleanup;\n+\t\t}\n+\t\tstrbuf_release(&dir);\n+\t}\n+\n+\terr = finalize_object_file(tmp_file.buf, filename.buf);\n+cleanup:\n+\tstrbuf_release(&tmp_file);\n+\tstrbuf_release(&filename);\n+\treturn err;\n+}\n+\n int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t    enum object_type type, struct object_id *oid,\n \t\t\t    unsigned flags)\ndiff --git a/object-store.h b/object-store.h\nindex bd2322ed8ce..1099455bc2e 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -46,6 +46,12 @@ struct object_directory {\n \tchar *path;\n };\n \n+struct input_stream {\n+\tconst void *(*read)(struct input_stream *, unsigned long *len);\n+\tvoid *data;\n+\tint is_finished;\n+};\n+\n KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n \tstruct object_directory *, 1, fspathhash, fspatheq)\n \n@@ -261,6 +267,8 @@ static inline int write_object_file(const void *buf, unsigned long len,\n int write_object_file_literally(const void *buf, unsigned long len,\n \t\t\t\tconst char *type, struct object_id *oid,\n \t\t\t\tunsigned flags);\n+int stream_loose_object(struct input_stream *in_stream, size_t len,\n+\t\t\tstruct object_id *oid);\n \n /*\n  * Add an object file to the in-memory object store, without writing it\n-- \n2.35.1.1548.g36973b18e52\n\n"},{"id":"452567","messageId":"patch-v12-6.8-33ffcbbc1f0-20220329T135446Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v12-0.8-00000000000-20220329T135446Z-avarab@gmail.com","subject":"[PATCH v12 6/8] core doc: modernize core.bigFileThreshold documentation","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-29T13:56:11Z","receivedAt":"2022-03-29T13:56:43Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"The core.bigFileThreshold documentation has been largely unchanged\nsince 5eef828bc03 (fast-import: Stream very large blobs directly to\npack, 2010-02-01).\n\nBut since then this setting has been expanded to affect a lot more\nthan that description indicated. Most notably in how \"git diff\" treats\nthem, see 6bf3b813486 (diff --stat: mark any file larger than\ncore.bigfilethreshold binary, 2014-08-16).\n\nIn addition to that, numerous commands and APIs make use of a\nstreaming mode for files above this threshold.\n\nSo let's attempt to summarize 12 years of changes in behavior, which\ncan be seen with:\n\n    git log --oneline -Gbig_file_thre 5eef828bc03.. -- '*.c'\n\nTo do that turn this into a bullet-point list. The summary Han Xin\nproduced in [1] helped a lot, but is a bit too detailed for\ndocumentation aimed at users. Let's instead summarize how\nuser-observable behavior differs, and generally describe how we tend\nto stream these files in various commands.\n\n1. https://lore.kernel.org/git/20220120112114.47618-5-chiyutianyi@gmail.com/\n\nHelped-by: Han Xin <chiyutianyi@gmail.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n Documentation/config/core.txt | 33 ++++++++++++++++++++++++---------\n 1 file changed, 24 insertions(+), 9 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex 9da3e5d88f6..5fccbd56995 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -412,17 +412,32 @@ You probably do not need to adjust this value.\n Common unit suffixes of 'k', 'm', or 'g' are supported.\n \n core.bigFileThreshold::\n-\tFiles larger than this size are stored deflated, without\n-\tattempting delta compression.  Storing large files without\n-\tdelta compression avoids excessive memory usage, at the\n-\tslight expense of increased disk usage. Additionally files\n-\tlarger than this size are always treated as binary.\n+\tThe size of files considered \"big\", which as discussed below\n+\tchanges the behavior of numerous git commands, as well as how\n+\tsuch files are stored within the repository. The default is\n+\t512 MiB. Common unit suffixes of 'k', 'm', or 'g' are\n+\tsupported.\n +\n-Default is 512 MiB on all platforms.  This should be reasonable\n-for most projects as source code and other text files can still\n-be delta compressed, but larger binary media files won't be.\n+Files above the configured limit will be:\n +\n-Common unit suffixes of 'k', 'm', or 'g' are supported.\n+* Stored deflated, without attempting delta compression.\n++\n+The default limit is primarily set with this use-case in mind. With it\n+most projects will have their source code and other text files delta\n+compressed, but not larger binary media files.\n++\n+Storing large files without delta compression avoids excessive memory\n+usage, at the slight expense of increased disk usage.\n++\n+* Will be treated as if though they were labeled \"binary\" (see\n+  linkgit:gitattributes[5]). This means that e.g. linkgit:git-log[1]\n+  and linkgit:git-diff[1] will not diffs for files above this limit.\n++\n+* Will be generally be streamed when written, which avoids excessive\n+memory usage, at the cost of some fixed overhead. Commands that make\n+use of this include linkgit:git-archive[1],\n+linkgit:git-fast-import[1], linkgit:git-index-pack[1] and\n+linkgit:git-fsck[1].\n \n core.excludesFile::\n \tSpecifies the pathname to the file that contains patterns to\n-- \n2.35.1.1548.g36973b18e52\n\n"},{"id":"452568","messageId":"patch-v12-7.8-11f7aa026b4-20220329T135446Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v12-0.8-00000000000-20220329T135446Z-avarab@gmail.com","subject":"[PATCH v12 7/8] unpack-objects: refactor away unpack_non_delta_entry()","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-29T13:56:12Z","receivedAt":"2022-03-29T13:56:47Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"The unpack_one() function will call either a non-trivial\nunpack_delta_entry() or a trivial unpack_non_delta_entry(). Let's\ninline the latter in the only caller.\n\nSince 21666f1aae4 (convert object type handling from a string to a\nnumber, 2007-02-26) the unpack_non_delta_entry() function has been\nrather trivial, and in a preceding commit the \"dry_run\" condition it\nwas handling went away.\n\nThis is not done as an optimization, as the compiler will easily\ndiscover that it can do the same, rather this makes a subsequent\ncommit easier to reason about. As it'll be handling \"OBJ_BLOB\" in a\nspecial manner let's re-arrange that \"case\" in preparation for that\nchange.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n builtin/unpack-objects.c | 18 +++++++-----------\n 1 file changed, 7 insertions(+), 11 deletions(-)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex e3d30025979..d374599d544 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -338,15 +338,6 @@ static void added_object(unsigned nr, enum object_type type,\n \t}\n }\n \n-static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n-\t\t\t\t   unsigned nr)\n-{\n-\tvoid *buf = get_data(size);\n-\n-\tif (buf)\n-\t\twrite_object(nr, type, buf, size);\n-}\n-\n static int resolve_against_held(unsigned nr, const struct object_id *base,\n \t\t\t\tvoid *delta_data, unsigned long delta_size)\n {\n@@ -479,12 +470,17 @@ static void unpack_one(unsigned nr)\n \t}\n \n \tswitch (type) {\n+\tcase OBJ_BLOB:\n \tcase OBJ_COMMIT:\n \tcase OBJ_TREE:\n-\tcase OBJ_BLOB:\n \tcase OBJ_TAG:\n-\t\tunpack_non_delta_entry(type, size, nr);\n+\t{\n+\t\tvoid *buf = get_data(size);\n+\n+\t\tif (buf)\n+\t\t\twrite_object(nr, type, buf, size);\n \t\treturn;\n+\t}\n \tcase OBJ_REF_DELTA:\n \tcase OBJ_OFS_DELTA:\n \t\tunpack_delta_entry(type, size, nr);\n-- \n2.35.1.1548.g36973b18e52\n\n"},{"id":"452569","messageId":"patch-v12-8.8-34ee6a28a54-20220329T135446Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v12-0.8-00000000000-20220329T135446Z-avarab@gmail.com","subject":"[PATCH v12 8/8] unpack-objects: use stream_loose_object() to unpack large objects","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-29T13:56:13Z","receivedAt":"2022-03-29T13:56:51Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nMake use of the stream_loose_object() function introduced in the\npreceding commit to unpack large objects. Before this we'd need to\nmalloc() the size of the blob before unpacking it, which could cause\nOOM with very large blobs.\n\nWe could use the new streaming interface to unpack all blobs, but\ndoing so would be much slower, as demonstrated e.g. with this\nbenchmark using git-hyperfine[0]:\n\n\trm -rf /tmp/scalar.git &&\n\tgit clone --bare https://github.com/Microsoft/scalar.git /tmp/scalar.git &&\n\tmv /tmp/scalar.git/objects/pack/*.pack /tmp/scalar.git/my.pack &&\n\tgit hyperfine \\\n\t\t-r 2 --warmup 1 \\\n\t\t-L rev origin/master,HEAD -L v \"10,512,1k,1m\" \\\n\t\t-s 'make' \\\n\t\t-p 'git init --bare dest.git' \\\n\t\t-c 'rm -rf dest.git' \\\n\t\t'./git -C dest.git -c core.bigFileThreshold={v} unpack-objects </tmp/scalar.git/my.pack'\n\nHere we'll perform worse with lower core.bigFileThreshold settings\nwith this change in terms of speed, but we're getting lower memory use\nin return:\n\n\tSummary\n\t  './git -C dest.git -c core.bigFileThreshold=10 unpack-objects </tmp/scalar.git/my.pack' in 'origin/master' ran\n\t    1.01 ± 0.01 times faster than './git -C dest.git -c core.bigFileThreshold=1k unpack-objects </tmp/scalar.git/my.pack' in 'origin/master'\n\t    1.01 ± 0.01 times faster than './git -C dest.git -c core.bigFileThreshold=1m unpack-objects </tmp/scalar.git/my.pack' in 'origin/master'\n\t    1.01 ± 0.02 times faster than './git -C dest.git -c core.bigFileThreshold=1m unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n\t    1.02 ± 0.00 times faster than './git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/scalar.git/my.pack' in 'origin/master'\n\t    1.09 ± 0.01 times faster than './git -C dest.git -c core.bigFileThreshold=1k unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n\t    1.10 ± 0.00 times faster than './git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n\t    1.11 ± 0.00 times faster than './git -C dest.git -c core.bigFileThreshold=10 unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n\nA better benchmark to demonstrate the benefits of that this one, which\ncreates an artificial repo with a 1, 25, 50, 75 and 100MB blob:\n\n\trm -rf /tmp/repo &&\n\tgit init /tmp/repo &&\n\t(\n\t\tcd /tmp/repo &&\n\t\tfor i in 1 25 50 75 100\n\t\tdo\n\t\t\tdd if=/dev/urandom of=blob.$i count=$(($i*1024)) bs=1024\n\t\tdone &&\n\t\tgit add blob.* &&\n\t\tgit commit -mblobs &&\n\t\tgit gc &&\n\t\tPACK=$(echo .git/objects/pack/pack-*.pack) &&\n\t\tcp \"$PACK\" my.pack\n\t) &&\n\tgit hyperfine \\\n\t\t--show-output \\\n\t\t-L rev origin/master,HEAD -L v \"512,50m,100m\" \\\n\t\t-s 'make' \\\n\t\t-p 'git init --bare dest.git' \\\n\t\t-c 'rm -rf dest.git' \\\n\t\t'/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold={v} unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum'\n\nUsing this test we'll always use >100MB of memory on\norigin/master (around ~105MB), but max out at e.g. ~55MB if we set\ncore.bigFileThreshold=50m.\n\nThe relevant \"Maximum resident set size\" lines were manually added\nbelow the relevant benchmark:\n\n  '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=50m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'origin/master' ran\n        Maximum resident set size (kbytes): 107080\n    1.02 ± 0.78 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'origin/master'\n        Maximum resident set size (kbytes): 106968\n    1.09 ± 0.79 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=100m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'origin/master'\n        Maximum resident set size (kbytes): 107032\n    1.42 ± 1.07 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=100m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'HEAD'\n        Maximum resident set size (kbytes): 107072\n    1.83 ± 1.02 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=50m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'HEAD'\n        Maximum resident set size (kbytes): 55704\n    2.16 ± 1.19 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'HEAD'\n        Maximum resident set size (kbytes): 4564\n\nThis shows that if you have enough memory this new streaming method is\nslower the lower you set the streaming threshold, but the benefit is\nmore bounded memory use.\n\nAn earlier version of this patch introduced a new\n\"core.bigFileStreamingThreshold\" instead of re-using the existing\n\"core.bigFileThreshold\" variable[1]. As noted in a detailed overview\nof its users in [2] using it has several different meanings.\n\nStill, we consider it good enough to simply re-use it. While it's\npossible that someone might want to e.g. consider objects \"small\" for\nthe purposes of diffing but \"big\" for the purposes of writing them\nsuch use-cases are probably too obscure to worry about. We can always\nsplit up \"core.bigFileThreshold\" in the future if there's a need for\nthat.\n\n0. https://github.com/avar/git-hyperfine/\n1. https://lore.kernel.org/git/20211210103435.83656-1-chiyutianyi@gmail.com/\n2. https://lore.kernel.org/git/20220120112114.47618-5-chiyutianyi@gmail.com/\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Derrick Stolee <stolee@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n Documentation/config/core.txt   |  4 +-\n builtin/unpack-objects.c        | 67 +++++++++++++++++++++++++++++++++\n t/t5351-unpack-large-objects.sh | 26 +++++++++++--\n 3 files changed, 92 insertions(+), 5 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex 5fccbd56995..716259b6762 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -436,8 +436,8 @@ usage, at the slight expense of increased disk usage.\n * Will be generally be streamed when written, which avoids excessive\n memory usage, at the cost of some fixed overhead. Commands that make\n use of this include linkgit:git-archive[1],\n-linkgit:git-fast-import[1], linkgit:git-index-pack[1] and\n-linkgit:git-fsck[1].\n+linkgit:git-fast-import[1], linkgit:git-index-pack[1],\n+linkgit:git-unpack-objects[1] and linkgit:git-fsck[1].\n \n core.excludesFile::\n \tSpecifies the pathname to the file that contains patterns to\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex d374599d544..9d7b325c23b 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -338,6 +338,68 @@ static void added_object(unsigned nr, enum object_type type,\n \t}\n }\n \n+struct input_zstream_data {\n+\tgit_zstream *zstream;\n+\tunsigned char buf[8192];\n+\tint status;\n+};\n+\n+static const void *feed_input_zstream(struct input_stream *in_stream,\n+\t\t\t\t      unsigned long *readlen)\n+{\n+\tstruct input_zstream_data *data = in_stream->data;\n+\tgit_zstream *zstream = data->zstream;\n+\tvoid *in = fill(1);\n+\n+\tif (in_stream->is_finished) {\n+\t\t*readlen = 0;\n+\t\treturn NULL;\n+\t}\n+\n+\tzstream->next_out = data->buf;\n+\tzstream->avail_out = sizeof(data->buf);\n+\tzstream->next_in = in;\n+\tzstream->avail_in = len;\n+\n+\tdata->status = git_inflate(zstream, 0);\n+\n+\tin_stream->is_finished = data->status != Z_OK;\n+\tuse(len - zstream->avail_in);\n+\t*readlen = sizeof(data->buf) - zstream->avail_out;\n+\n+\treturn data->buf;\n+}\n+\n+static void stream_blob(unsigned long size, unsigned nr)\n+{\n+\tgit_zstream zstream = { 0 };\n+\tstruct input_zstream_data data = { 0 };\n+\tstruct input_stream in_stream = {\n+\t\t.read = feed_input_zstream,\n+\t\t.data = &data,\n+\t};\n+\tstruct obj_info *info = &obj_list[nr];\n+\n+\tdata.zstream = &zstream;\n+\tgit_inflate_init(&zstream);\n+\n+\tif (stream_loose_object(&in_stream, size, &info->oid))\n+\t\tdie(_(\"failed to write object in stream\"));\n+\n+\tif (data.status != Z_STREAM_END)\n+\t\tdie(_(\"inflate returned (%d)\"), data.status);\n+\tgit_inflate_end(&zstream);\n+\n+\tif (strict) {\n+\t\tstruct blob *blob = lookup_blob(the_repository, &info->oid);\n+\n+\t\tif (!blob)\n+\t\t\tdie(_(\"invalid blob object from stream\"));\n+\t\tblob->object.flags |= FLAG_WRITTEN;\n+\t}\n+\tinfo->obj = NULL;\n+}\n+\n static int resolve_against_held(unsigned nr, const struct object_id *base,\n \t\t\t\tvoid *delta_data, unsigned long delta_size)\n {\n@@ -471,6 +533,11 @@ static void unpack_one(unsigned nr)\n \n \tswitch (type) {\n \tcase OBJ_BLOB:\n+\t\tif (!dry_run && size > big_file_threshold) {\n+\t\t\tstream_blob(size, nr);\n+\t\t\treturn;\n+\t\t}\n+\t\t/* fallthrough */\n \tcase OBJ_COMMIT:\n \tcase OBJ_TREE:\n \tcase OBJ_TAG:\ndiff --git a/t/t5351-unpack-large-objects.sh b/t/t5351-unpack-large-objects.sh\nindex 8d84313221c..461ca060b2b 100755\n--- a/t/t5351-unpack-large-objects.sh\n+++ b/t/t5351-unpack-large-objects.sh\n@@ -9,7 +9,8 @@ test_description='git unpack-objects with large objects'\n \n prepare_dest () {\n \ttest_when_finished \"rm -rf dest.git\" &&\n-\tgit init --bare dest.git\n+\tgit init --bare dest.git &&\n+\tgit -C dest.git config core.bigFileThreshold \"$1\"\n }\n \n test_expect_success \"create large objects (1.5 MB) and PACK\" '\n@@ -26,16 +27,35 @@ test_expect_success 'set memory limitation to 1MB' '\n '\n \n test_expect_success 'unpack-objects failed under memory limitation' '\n-\tprepare_dest &&\n+\tprepare_dest 2m &&\n \ttest_must_fail git -C dest.git unpack-objects <pack-$PACK.pack 2>err &&\n \tgrep \"fatal: attempting to allocate\" err\n '\n \n test_expect_success 'unpack-objects works with memory limitation in dry-run mode' '\n-\tprepare_dest &&\n+\tprepare_dest 2m &&\n \tgit -C dest.git unpack-objects -n <pack-$PACK.pack &&\n \ttest_stdout_line_count = 0 find dest.git/objects -type f &&\n \ttest_dir_is_empty dest.git/objects/pack\n '\n \n+test_expect_success 'unpack big object in stream' '\n+\tprepare_dest 1m &&\n+\tgit -C dest.git unpack-objects <pack-$PACK.pack &&\n+\ttest_dir_is_empty dest.git/objects/pack\n+'\n+\n+test_expect_success 'do not unpack existing large objects' '\n+\tprepare_dest 1m &&\n+\tgit -C dest.git index-pack --stdin <pack-$PACK.pack &&\n+\tgit -C dest.git unpack-objects <pack-$PACK.pack &&\n+\n+\t# The destination came up with the exact same pack...\n+\tDEST_PACK=$(echo dest.git/objects/pack/pack-*.pack) &&\n+\ttest_cmp pack-$PACK.pack $DEST_PACK &&\n+\n+\t# ...and wrote no loose objects\n+\ttest_stdout_line_count = 0 find dest.git/objects -type f ! -name \"pack-*\"\n+'\n+\n test_done\n-- \n2.35.1.1548.g36973b18e52\n\n"},{"id":"452657","messageId":"20220330071344.25676-1-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"patch-v12-3.8-3dcaa5d6589-20220329T135446Z-avarab@gmail.com","subject":"Re: [PATCH v12 3/8] object-file.c: refactor write_loose_object() to several steps","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-03-30T07:13:44Z","receivedAt":"2022-03-30T07:14:11Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Tue, Mar 29, 2022 at 3:56 PM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n> \n> +/**\n> + * Common steps for loose object writers to end writing loose objects:\n> + *\n> + * - End the compression of zlib stream.\n> + * - Get the calculated oid to \"oid\".\n> + * - fsync() and close() the \"fd\"\n\nSince we removed close_loose_object() from end_loose_object_common() , I\nthink this comment should also be removed.\n\nThanks.\n-Han Xin\n\n> + */\n> +static int end_loose_object_common(git_hash_ctx *c, git_zstream *stream,\n> +\t\t\t\t   struct object_id *oid)\n> +{\n> +\tint ret;\n> +\n> +\tret = git_deflate_end_gently(stream);\n> +\tif (ret != Z_OK)\n> +\t\treturn ret;\n> +\tthe_hash_algo->final_oid_fn(oid, c);\n> +\n> +\treturn Z_OK;\n> +}\n> +\n"},{"id":"452677","messageId":"220330.86mth7ny1r.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"20220330071344.25676-1-chiyutianyi@gmail.com","subject":"Re: [PATCH v12 3/8] object-file.c: refactor write_loose_object() to several steps","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-30T17:34:35Z","receivedAt":"2022-03-30T17:35:34Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Mar 30 2022, Han Xin wrote:\n\n> On Tue, Mar 29, 2022 at 3:56 PM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>> \n>> +/**\n>> + * Common steps for loose object writers to end writing loose objects:\n>> + *\n>> + * - End the compression of zlib stream.\n>> + * - Get the calculated oid to \"oid\".\n>> + * - fsync() and close() the \"fd\"\n>\n> Since we removed close_loose_object() from end_loose_object_common() , I\n> think this comment should also be removed.\n\nYou're right. I adjusted it for the \"parano_oid\" in this v12, but\nmanaged to miss that somehow.\n\nWill submit a re-roll with those changes, but will wait a bit more to\nsee if there's any other comments on this v12 first. Thanks!\n\n"},{"id":"452698","messageId":"2f98c63d-f2c9-26fe-cfd3-9eed6b79047a@web.de","threadId":"56672","inReplyTo":"patch-v12-7.8-11f7aa026b4-20220329T135446Z-avarab@gmail.com","subject":"Re: [PATCH v12 7/8] unpack-objects: refactor away unpack_non_delta_entry()","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-03-30T19:40:43Z","receivedAt":"2022-03-30T19:40:56Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 29.03.22 um 15:56 schrieb Ævar Arnfjörð Bjarmason:\n> The unpack_one() function will call either a non-trivial\n> unpack_delta_entry() or a trivial unpack_non_delta_entry(). Let's\n> inline the latter in the only caller.\n>\n> Since 21666f1aae4 (convert object type handling from a string to a\n> number, 2007-02-26) the unpack_non_delta_entry() function has been\n> rather trivial, and in a preceding commit the \"dry_run\" condition it\n> was handling went away.\n>\n> This is not done as an optimization, as the compiler will easily\n> discover that it can do the same, rather this makes a subsequent\n> commit easier to reason about.\n\nHow exactly does inlining the function make the next patch easier to\nunderstand or discuss?  Plugging in the stream_blob() call to handle the\nbig blobs looks the same with or without the unpack_non_delta_entry()\ncall to me.\n\n> As it'll be handling \"OBJ_BLOB\" in a\n> special manner let's re-arrange that \"case\" in preparation for that\n> change.\n>\n> Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> ---\n>  builtin/unpack-objects.c | 18 +++++++-----------\n>  1 file changed, 7 insertions(+), 11 deletions(-)\n\nReducing the number of lines can be an advantage. *shrug*\n\n>\n> diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n> index e3d30025979..d374599d544 100644\n> --- a/builtin/unpack-objects.c\n> +++ b/builtin/unpack-objects.c\n> @@ -338,15 +338,6 @@ static void added_object(unsigned nr, enum object_type type,\n>  \t}\n>  }\n>\n> -static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n> -\t\t\t\t   unsigned nr)\n> -{\n> -\tvoid *buf = get_data(size);\n> -\n> -\tif (buf)\n> -\t\twrite_object(nr, type, buf, size);\n> -}\n> -\n>  static int resolve_against_held(unsigned nr, const struct object_id *base,\n>  \t\t\t\tvoid *delta_data, unsigned long delta_size)\n>  {\n> @@ -479,12 +470,17 @@ static void unpack_one(unsigned nr)\n>  \t}\n>\n>  \tswitch (type) {\n> +\tcase OBJ_BLOB:\n>  \tcase OBJ_COMMIT:\n>  \tcase OBJ_TREE:\n> -\tcase OBJ_BLOB:\n>  \tcase OBJ_TAG:\n> -\t\tunpack_non_delta_entry(type, size, nr);\n> +\t{\n> +\t\tvoid *buf = get_data(size);\n> +\n> +\t\tif (buf)\n> +\t\t\twrite_object(nr, type, buf, size);\n>  \t\treturn;\n> +\t}\n>  \tcase OBJ_REF_DELTA:\n>  \tcase OBJ_OFS_DELTA:\n>  \t\tunpack_delta_entry(type, size, nr);\n"},{"id":"452790","messageId":"220331.86wngap9vo.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"2f98c63d-f2c9-26fe-cfd3-9eed6b79047a@web.de","subject":"Re: [PATCH v12 7/8] unpack-objects: refactor away unpack_non_delta_entry()","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-31T12:42:34Z","receivedAt":"2022-03-31T12:46:56Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Mar 30 2022, René Scharfe wrote:\n\n> Am 29.03.22 um 15:56 schrieb Ævar Arnfjörð Bjarmason:\n>> The unpack_one() function will call either a non-trivial\n>> unpack_delta_entry() or a trivial unpack_non_delta_entry(). Let's\n>> inline the latter in the only caller.\n>>\n>> Since 21666f1aae4 (convert object type handling from a string to a\n>> number, 2007-02-26) the unpack_non_delta_entry() function has been\n>> rather trivial, and in a preceding commit the \"dry_run\" condition it\n>> was handling went away.\n>>\n>> This is not done as an optimization, as the compiler will easily\n>> discover that it can do the same, rather this makes a subsequent\n>> commit easier to reason about.\n>\n> How exactly does inlining the function make the next patch easier to\n> understand or discuss?  Plugging in the stream_blob() call to handle the\n> big blobs looks the same with or without the unpack_non_delta_entry()\n> call to me.\n\nThe earlier version of it without this prep cleanup can be seen at\nhttps://lore.kernel.org/git/patch-v10-6.6-6a70e49a346-20220204T135538Z-avarab@gmail.com/\n\nSo yes, this could be skipped, but I tought with this step it was easier\nto understand.\n\nWe previously had to change \"void *buf = get_data(size);\" in the\nfunction to just \"void *buf\", and do the assignment after the condition\nthat's being checked here.\n\nI think it's also more obvious in terms of control flow if we're\nchecking OBJ_BLOB here to not call a function which has a special-case\njust for OBJ_BLOB, we can just do that here instead.\n\n>> As it'll be handling \"OBJ_BLOB\" in a\n>> special manner let's re-arrange that \"case\" in preparation for that\n>> change.\n>>\n>> Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n>> ---\n>>  builtin/unpack-objects.c | 18 +++++++-----------\n>>  1 file changed, 7 insertions(+), 11 deletions(-)\n>\n> Reducing the number of lines can be an advantage. *shrug*\n\nThere was also the (admittedly rather small) knock-on-effect on\n8/8. Before this it was 8 lines added / 1 removed when it came to the\ncode impacted by this change, now it's a 5 added/0 removed in the below\n\"switch\".\n\nSo I think it's worth keeping.\n"},{"id":"452804","messageId":"b6f9518d-0b41-d607-ba77-46a83edb1d40@web.de","threadId":"56672","inReplyTo":"220331.86wngap9vo.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v12 7/8] unpack-objects: refactor away unpack_non_delta_entry()","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-03-31T16:38:21Z","receivedAt":"2022-03-31T16:38:37Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 31.03.22 um 14:42 schrieb Ævar Arnfjörð Bjarmason:\n>\n> On Wed, Mar 30 2022, René Scharfe wrote:\n>\n>> Am 29.03.22 um 15:56 schrieb Ævar Arnfjörð Bjarmason:\n>>> The unpack_one() function will call either a non-trivial\n>>> unpack_delta_entry() or a trivial unpack_non_delta_entry(). Let's\n>>> inline the latter in the only caller.\n>>>\n>>> Since 21666f1aae4 (convert object type handling from a string to a\n>>> number, 2007-02-26) the unpack_non_delta_entry() function has been\n>>> rather trivial, and in a preceding commit the \"dry_run\" condition it\n>>> was handling went away.\n>>>\n>>> This is not done as an optimization, as the compiler will easily\n>>> discover that it can do the same, rather this makes a subsequent\n>>> commit easier to reason about.\n>>\n>> How exactly does inlining the function make the next patch easier to\n>> understand or discuss?  Plugging in the stream_blob() call to handle the\n>> big blobs looks the same with or without the unpack_non_delta_entry()\n>> call to me.\n>\n> The earlier version of it without this prep cleanup can be seen at\n> https://lore.kernel.org/git/patch-v10-6.6-6a70e49a346-20220204T135538Z-avarab@gmail.com/\n\nThis plugged the special case into unpack_non_delta_entry().  The\nalternative I had in mind was to plug it into the switch statement as\nthe current patch does, just without inlining unpack_non_delta_entry().\n\n> So yes, this could be skipped, but I tought with this step it was easier\n> to understand.\n>\n> We previously had to change \"void *buf = get_data(size);\" in the\n> function to just \"void *buf\", and do the assignment after the condition\n> that's being checked here.\n>\n> I think it's also more obvious in terms of control flow if we're\n> checking OBJ_BLOB here to not call a function which has a special-case\n> just for OBJ_BLOB, we can just do that here instead.\n>\n>>> As it'll be handling \"OBJ_BLOB\" in a\n>>> special manner let's re-arrange that \"case\" in preparation for that\n>>> change.\n>>>\n>>> Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n>>> ---\n>>>  builtin/unpack-objects.c | 18 +++++++-----------\n>>>  1 file changed, 7 insertions(+), 11 deletions(-)\n>>\n>> Reducing the number of lines can be an advantage. *shrug*\n>\n> There was also the (admittedly rather small) knock-on-effect on\n> 8/8. Before this it was 8 lines added / 1 removed when it came to the\n> code impacted by this change, now it's a 5 added/0 removed in the below\n> \"switch\".\n>\n> So I think it's worth keeping.\n"},{"id":"452829","messageId":"20220331195453.GA24610@neerajsi-x1.localdomain","threadId":"56672","inReplyTo":"patch-v12-5.8-3d64cf1cf33-20220329T135446Z-avarab@gmail.com","subject":"Re: [PATCH v12 5/8] object-file.c: add \"stream_loose_object()\" to handle large object","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-03-31T19:54:53Z","receivedAt":"2022-03-31T19:55:07Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On Tue, Mar 29, 2022 at 03:56:10PM +0200, Ævar Arnfjörð Bjarmason wrote:\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n> \n> If we want unpack and write a loose object using \"write_loose_object\",\n> we have to feed it with a buffer with the same size of the object, which\n> will consume lots of memory and may cause OOM. This can be improved by\n> feeding data to \"stream_loose_object()\" in a stream.\n> \n> Add a new function \"stream_loose_object()\", which is a stream version of\n> \"write_loose_object()\" but with a low memory footprint. We will use this\n> function to unpack large blob object in later commit.\n> \n\nJust a thought for optimization which you might want to try on top of this\nseries:\ntry using mmap on both the source and target files of your stream. Use a\nbig 'window' for the mmap (multiple MB) to reduce the TLB flush costs. TLB\nflush costs should be minimal anyway if Git is single-threaded.\n\nIf you can set the source and target buffers of zlib to the source and\ndest mappings respectively, you'd eliminate two copies of data into\nGit's stack buffers.  You might need to over-allocate the dst file if\nyou don't know the size up front, but doing an over-allocate and truncate\nshould be pretty cheap if you're working with a big file.\n\nThanks,\nNeeraj\n"},{"id":"455544","messageId":"cover.1653015534.git.chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"cover-v11-0.8-00000000000-20220319T001411Z-avarab@gmail.com","subject":"[PATCH 0/1] unpack-objects: low memory footprint for get_data() in dry_run mode","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-05-20T03:05:13Z","receivedAt":"2022-05-20T03:06:16Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"This patch teaches \"git unpack-objects\" to use a lower memory footprint\nfor \"get_data()\" in dry-run mode since the returned data is not used.\n\nThis patch is separeted from \"[PATCH v12 0/8] unpack-objects: support\nstreaming blobs to disk\"[1] because it has less impact and less controversy\non existing ones.\n\n1. https://lore.kernel.org/git/cover-v12-0.8-00000000000-20220329T135446Z-avarab@gmail.com/\n\nHan Xin (1):\n  unpack-objects: low memory footprint for get_data() in dry_run mode\n\n builtin/unpack-objects.c        | 34 ++++++++++++++++++---------\n t/t5351-unpack-large-objects.sh | 41 +++++++++++++++++++++++++++++++++\n 2 files changed, 64 insertions(+), 11 deletions(-)\n create mode 100755 t/t5351-unpack-large-objects.sh\n\n-- \n2.36.1\n\n"},{"id":"455545","messageId":"354ec53826f6af0977387a99e2204c0dd4e96b20.1653015534.git.chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"cover.1653015534.git.chiyutianyi@gmail.com","subject":"[PATCH 1/1] unpack-objects: low memory footprint for get_data() in dry_run mode","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-05-20T03:05:14Z","receivedAt":"2022-05-20T03:06:20Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nAs the name implies, \"get_data(size)\" will allocate and return a given\namount of memory. Allocating memory for a large blob object may cause the\nsystem to run out of memory. Before preparing to replace calling of\n\"get_data()\" to unpack large blob objects in latter commits, refactor\n\"get_data()\" to reduce memory footprint for dry_run mode.\n\nBecause in dry_run mode, \"get_data()\" is only used to check the\nintegrity of data, and the returned buffer is not used at all, we can\nallocate a smaller buffer and reuse it as zstream output. Therefore,\nin dry_run mode, \"get_data()\" will release the allocated buffer and\nreturn NULL instead of returning garbage data.\n\nThe \"find [...]objects/?? -type f | wc -l\" test idiom being used here\nis adapted from the same \"find\" use added to another test in\nd9545c7f465 (fast-import: implement unpack limit, 2016-04-25).\n\nSuggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n builtin/unpack-objects.c        | 34 ++++++++++++++++++---------\n t/t5351-unpack-large-objects.sh | 41 +++++++++++++++++++++++++++++++++\n 2 files changed, 64 insertions(+), 11 deletions(-)\n create mode 100755 t/t5351-unpack-large-objects.sh\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex dbeb0680a5..e3d3002597 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -96,15 +96,26 @@ static void use(int bytes)\n \tdisplay_throughput(progress, consumed_bytes);\n }\n \n+/*\n+ * Decompress zstream from stdin and return specific size of data.\n+ * The caller is responsible to free the returned buffer.\n+ *\n+ * But for dry_run mode, \"get_data()\" is only used to check the\n+ * integrity of data, and the returned buffer is not used at all.\n+ * Therefore, in dry_run mode, \"get_data()\" will release the small\n+ * allocated buffer which is reused to hold temporary zstream output\n+ * and return NULL instead of returning garbage data.\n+ */\n static void *get_data(unsigned long size)\n {\n \tgit_zstream stream;\n-\tvoid *buf = xmallocz(size);\n+\tunsigned long bufsize = dry_run && size > 8192 ? 8192 : size;\n+\tvoid *buf = xmallocz(bufsize);\n \n \tmemset(&stream, 0, sizeof(stream));\n \n \tstream.next_out = buf;\n-\tstream.avail_out = size;\n+\tstream.avail_out = bufsize;\n \tstream.next_in = fill(1);\n \tstream.avail_in = len;\n \tgit_inflate_init(&stream);\n@@ -124,8 +135,15 @@ static void *get_data(unsigned long size)\n \t\t}\n \t\tstream.next_in = fill(1);\n \t\tstream.avail_in = len;\n+\t\tif (dry_run) {\n+\t\t\t/* reuse the buffer in dry_run mode */\n+\t\t\tstream.next_out = buf;\n+\t\t\tstream.avail_out = bufsize;\n+\t\t}\n \t}\n \tgit_inflate_end(&stream);\n+\tif (dry_run)\n+\t\tFREE_AND_NULL(buf);\n \treturn buf;\n }\n \n@@ -325,10 +343,8 @@ static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n {\n \tvoid *buf = get_data(size);\n \n-\tif (!dry_run && buf)\n+\tif (buf)\n \t\twrite_object(nr, type, buf, size);\n-\telse\n-\t\tfree(buf);\n }\n \n static int resolve_against_held(unsigned nr, const struct object_id *base,\n@@ -358,10 +374,8 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\toidread(&base_oid, fill(the_hash_algo->rawsz));\n \t\tuse(the_hash_algo->rawsz);\n \t\tdelta_data = get_data(delta_size);\n-\t\tif (dry_run || !delta_data) {\n-\t\t\tfree(delta_data);\n+\t\tif (!delta_data)\n \t\t\treturn;\n-\t\t}\n \t\tif (has_object_file(&base_oid))\n \t\t\t; /* Ok we have this one */\n \t\telse if (resolve_against_held(nr, &base_oid,\n@@ -397,10 +411,8 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\t\tdie(\"offset value out of bound for delta base object\");\n \n \t\tdelta_data = get_data(delta_size);\n-\t\tif (dry_run || !delta_data) {\n-\t\t\tfree(delta_data);\n+\t\tif (!delta_data)\n \t\t\treturn;\n-\t\t}\n \t\tlo = 0;\n \t\thi = nr;\n \t\twhile (lo < hi) {\ndiff --git a/t/t5351-unpack-large-objects.sh b/t/t5351-unpack-large-objects.sh\nnew file mode 100755\nindex 0000000000..8d84313221\n--- /dev/null\n+++ b/t/t5351-unpack-large-objects.sh\n@@ -0,0 +1,41 @@\n+#!/bin/sh\n+#\n+# Copyright (c) 2022 Han Xin\n+#\n+\n+test_description='git unpack-objects with large objects'\n+\n+. ./test-lib.sh\n+\n+prepare_dest () {\n+\ttest_when_finished \"rm -rf dest.git\" &&\n+\tgit init --bare dest.git\n+}\n+\n+test_expect_success \"create large objects (1.5 MB) and PACK\" '\n+\ttest-tool genrandom foo 1500000 >big-blob &&\n+\ttest_commit --append foo big-blob &&\n+\ttest-tool genrandom bar 1500000 >big-blob &&\n+\ttest_commit --append bar big-blob &&\n+\tPACK=$(echo HEAD | git pack-objects --revs pack)\n+'\n+\n+test_expect_success 'set memory limitation to 1MB' '\n+\tGIT_ALLOC_LIMIT=1m &&\n+\texport GIT_ALLOC_LIMIT\n+'\n+\n+test_expect_success 'unpack-objects failed under memory limitation' '\n+\tprepare_dest &&\n+\ttest_must_fail git -C dest.git unpack-objects <pack-$PACK.pack 2>err &&\n+\tgrep \"fatal: attempting to allocate\" err\n+'\n+\n+test_expect_success 'unpack-objects works with memory limitation in dry-run mode' '\n+\tprepare_dest &&\n+\tgit -C dest.git unpack-objects -n <pack-$PACK.pack &&\n+\ttest_stdout_line_count = 0 find dest.git/objects -type f &&\n+\ttest_dir_is_empty dest.git/objects/pack\n+'\n+\n+test_done\n-- \n2.36.1\n\n"},{"id":"456661","messageId":"patch-v13-1.7-12873fc9915-20220604T095113Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v13-0.7-00000000000-20220604T095113Z-avarab@gmail.com","subject":"[PATCH v13 1/7] unpack-objects: low memory footprint for get_data() in dry_run mode","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-06-04T10:10:22Z","receivedAt":"2022-06-04T10:10:46Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nAs the name implies, \"get_data(size)\" will allocate and return a given\namount of memory. Allocating memory for a large blob object may cause the\nsystem to run out of memory. Before preparing to replace calling of\n\"get_data()\" to unpack large blob objects in latter commits, refactor\n\"get_data()\" to reduce memory footprint for dry_run mode.\n\nBecause in dry_run mode, \"get_data()\" is only used to check the\nintegrity of data, and the returned buffer is not used at all, we can\nallocate a smaller buffer and reuse it as zstream output. Therefore,\nin dry_run mode, \"get_data()\" will release the allocated buffer and\nreturn NULL instead of returning garbage data.\n\nThe \"find [...]objects/?? -type f | wc -l\" test idiom being used here\nis adapted from the same \"find\" use added to another test in\nd9545c7f465 (fast-import: implement unpack limit, 2016-04-25).\n\nSuggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n builtin/unpack-objects.c        | 34 ++++++++++++++++++---------\n t/t5351-unpack-large-objects.sh | 41 +++++++++++++++++++++++++++++++++\n 2 files changed, 64 insertions(+), 11 deletions(-)\n create mode 100755 t/t5351-unpack-large-objects.sh\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 56d05e2725d..64abba8dbac 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -97,15 +97,26 @@ static void use(int bytes)\n \tdisplay_throughput(progress, consumed_bytes);\n }\n \n+/*\n+ * Decompress zstream from stdin and return specific size of data.\n+ * The caller is responsible to free the returned buffer.\n+ *\n+ * But for dry_run mode, \"get_data()\" is only used to check the\n+ * integrity of data, and the returned buffer is not used at all.\n+ * Therefore, in dry_run mode, \"get_data()\" will release the small\n+ * allocated buffer which is reused to hold temporary zstream output\n+ * and return NULL instead of returning garbage data.\n+ */\n static void *get_data(unsigned long size)\n {\n \tgit_zstream stream;\n-\tvoid *buf = xmallocz(size);\n+\tunsigned long bufsize = dry_run && size > 8192 ? 8192 : size;\n+\tvoid *buf = xmallocz(bufsize);\n \n \tmemset(&stream, 0, sizeof(stream));\n \n \tstream.next_out = buf;\n-\tstream.avail_out = size;\n+\tstream.avail_out = bufsize;\n \tstream.next_in = fill(1);\n \tstream.avail_in = len;\n \tgit_inflate_init(&stream);\n@@ -125,8 +136,15 @@ static void *get_data(unsigned long size)\n \t\t}\n \t\tstream.next_in = fill(1);\n \t\tstream.avail_in = len;\n+\t\tif (dry_run) {\n+\t\t\t/* reuse the buffer in dry_run mode */\n+\t\t\tstream.next_out = buf;\n+\t\t\tstream.avail_out = bufsize;\n+\t\t}\n \t}\n \tgit_inflate_end(&stream);\n+\tif (dry_run)\n+\t\tFREE_AND_NULL(buf);\n \treturn buf;\n }\n \n@@ -326,10 +344,8 @@ static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n {\n \tvoid *buf = get_data(size);\n \n-\tif (!dry_run && buf)\n+\tif (buf)\n \t\twrite_object(nr, type, buf, size);\n-\telse\n-\t\tfree(buf);\n }\n \n static int resolve_against_held(unsigned nr, const struct object_id *base,\n@@ -359,10 +375,8 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\toidread(&base_oid, fill(the_hash_algo->rawsz));\n \t\tuse(the_hash_algo->rawsz);\n \t\tdelta_data = get_data(delta_size);\n-\t\tif (dry_run || !delta_data) {\n-\t\t\tfree(delta_data);\n+\t\tif (!delta_data)\n \t\t\treturn;\n-\t\t}\n \t\tif (has_object_file(&base_oid))\n \t\t\t; /* Ok we have this one */\n \t\telse if (resolve_against_held(nr, &base_oid,\n@@ -398,10 +412,8 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\t\tdie(\"offset value out of bound for delta base object\");\n \n \t\tdelta_data = get_data(delta_size);\n-\t\tif (dry_run || !delta_data) {\n-\t\t\tfree(delta_data);\n+\t\tif (!delta_data)\n \t\t\treturn;\n-\t\t}\n \t\tlo = 0;\n \t\thi = nr;\n \t\twhile (lo < hi) {\ndiff --git a/t/t5351-unpack-large-objects.sh b/t/t5351-unpack-large-objects.sh\nnew file mode 100755\nindex 00000000000..8d84313221c\n--- /dev/null\n+++ b/t/t5351-unpack-large-objects.sh\n@@ -0,0 +1,41 @@\n+#!/bin/sh\n+#\n+# Copyright (c) 2022 Han Xin\n+#\n+\n+test_description='git unpack-objects with large objects'\n+\n+. ./test-lib.sh\n+\n+prepare_dest () {\n+\ttest_when_finished \"rm -rf dest.git\" &&\n+\tgit init --bare dest.git\n+}\n+\n+test_expect_success \"create large objects (1.5 MB) and PACK\" '\n+\ttest-tool genrandom foo 1500000 >big-blob &&\n+\ttest_commit --append foo big-blob &&\n+\ttest-tool genrandom bar 1500000 >big-blob &&\n+\ttest_commit --append bar big-blob &&\n+\tPACK=$(echo HEAD | git pack-objects --revs pack)\n+'\n+\n+test_expect_success 'set memory limitation to 1MB' '\n+\tGIT_ALLOC_LIMIT=1m &&\n+\texport GIT_ALLOC_LIMIT\n+'\n+\n+test_expect_success 'unpack-objects failed under memory limitation' '\n+\tprepare_dest &&\n+\ttest_must_fail git -C dest.git unpack-objects <pack-$PACK.pack 2>err &&\n+\tgrep \"fatal: attempting to allocate\" err\n+'\n+\n+test_expect_success 'unpack-objects works with memory limitation in dry-run mode' '\n+\tprepare_dest &&\n+\tgit -C dest.git unpack-objects -n <pack-$PACK.pack &&\n+\ttest_stdout_line_count = 0 find dest.git/objects -type f &&\n+\ttest_dir_is_empty dest.git/objects/pack\n+'\n+\n+test_done\n-- \n2.36.1.1124.g52838f02905\n\n"},{"id":"456662","messageId":"cover-v13-0.7-00000000000-20220604T095113Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v12-0.8-00000000000-20220329T135446Z-avarab@gmail.com","subject":"[PATCH v13 0/7] unpack-objects: support streaming blobs to disk","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-06-04T10:10:21Z","receivedAt":"2022-06-04T10:10:47Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"This series makes \"unpack-objects\" capable of streaming large objects\nto disk.\n\nAs 7/7 shows streaming e.g. a 100MB blob now uses ~5MB of memory\ninstead of ~105MB. This streaming method is slower if you've got\nmemory to handle the blobs in-core, but if you don't it allows you to\nunpack objects at all, as you might otherwise OOM.\n\nThis series by Han Xin was originally waiting on some in-flight\npatches that landed in 430883a70c7 (Merge branch\n'ab/object-file-api-updates', 2022-03-16), and until yesterday with\n83937e95928 (Merge branch 'ns/batch-fsync', 2022-06-03) had a textual\nand semantic conflict with \"master\".\n\nChanges since v12:\n\n * Since v12 Han Xin submitted 1/1 here as\n   https://lore.kernel.org/git/cover.1653015534.git.chiyutianyi@gmail.com/;\n   I think this is better off reviewed as a whole, and hopefully will\n   be picked up as such.\n\n * Dropped the previous 7/8, which was a refactoring to make 8/8\n   slightly smaller. Per dicsussion with René it's better to leave it\n   out.\n\n * The rest (especially 2/8) is due to rebasing on ns/batch-fsync.\n\nHan Xin (4):\n  unpack-objects: low memory footprint for get_data() in dry_run mode\n  object-file.c: refactor write_loose_object() to several steps\n  object-file.c: add \"stream_loose_object()\" to handle large object\n  unpack-objects: use stream_loose_object() to unpack large objects\n\nÆvar Arnfjörð Bjarmason (3):\n  object-file.c: do fsync() and close() before post-write die()\n  object-file.c: factor out deflate part of write_loose_object()\n  core doc: modernize core.bigFileThreshold documentation\n\n Documentation/config/core.txt   |  33 +++--\n builtin/unpack-objects.c        | 103 ++++++++++++--\n object-file.c                   | 237 +++++++++++++++++++++++++++-----\n object-store.h                  |   8 ++\n t/t5351-unpack-large-objects.sh |  61 ++++++++\n 5 files changed, 387 insertions(+), 55 deletions(-)\n create mode 100755 t/t5351-unpack-large-objects.sh\n\nRange-diff against v12:\n1:  e95f6a1cfb6 = 1:  12873fc9915 unpack-objects: low memory footprint for get_data() in dry_run mode\n2:  54060eb8c6b ! 2:  b3568f0c5c0 object-file.c: do fsync() and close() before post-write die()\n    @@ Commit message\n         Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n      ## object-file.c ##\n    -@@ object-file.c: void hash_object_file(const struct git_hash_algo *algo, const void *buf,\n    - \thash_object_file_literally(algo, buf, len, type_name(type), oid);\n    - }\n    - \n    --/* Finalize a file on disk, and close it. */\n    -+/*\n    -+ * We already did a write_buffer() to the \"fd\", let's fsync()\n    -+ * and close().\n    -+ *\n    -+ * Finalize a file on disk, and close it. We might still die() on a\n    -+ * subsequent sanity check, but let's not add to that confusion by not\n    -+ * flushing any outstanding writes to disk first.\n    -+ */\n    - static void close_loose_object(int fd)\n    - {\n    - \tif (the_repository->objects->odb->will_destroy)\n     @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *hdr,\n      \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n      \t\t    ret);\n      \tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n    -+\tclose_loose_object(fd);\n    ++\tclose_loose_object(fd, tmp_file.buf);\n     +\n      \tif (!oideq(oid, &parano_oid))\n      \t\tdie(_(\"confused by unstable object source data for %s\"),\n      \t\t    oid_to_hex(oid));\n      \n    --\tclose_loose_object(fd);\n    +-\tclose_loose_object(fd, tmp_file.buf);\n     -\n      \tif (mtime) {\n      \t\tstruct utimbuf utb;\n3:  3dcaa5d6589 ! 3:  9dc0f56878a object-file.c: refactor write_loose_object() to several steps\n    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *\n     -\t\t    ret);\n     -\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n     +\t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid), ret);\n    - \tclose_loose_object(fd);\n    + \tclose_loose_object(fd, tmp_file.buf);\n      \n      \tif (!oideq(oid, &parano_oid))\n4:  03f4e91ac89 = 4:  a0434835fe7 object-file.c: factor out deflate part of write_loose_object()\n5:  3d64cf1cf33 ! 5:  0b07b29836b object-file.c: add \"stream_loose_object()\" to handle large object\n    @@ object-file.c: static int freshen_packed_object(const struct object_id *oid)\n     +\tret = end_loose_object_common(&c, &stream, oid);\n     +\tif (ret != Z_OK)\n     +\t\tdie(_(\"deflateEnd on stream object failed (%d)\"), ret);\n    -+\tclose_loose_object(fd);\n    ++\tclose_loose_object(fd, tmp_file.buf);\n     +\n     +\tif (freshen_packed_object(oid) || freshen_loose_object(oid)) {\n     +\t\tunlink_or_warn(tmp_file.buf);\n6:  33ffcbbc1f0 = 6:  5ed79c58b18 core doc: modernize core.bigFileThreshold documentation\n7:  11f7aa026b4 < -:  ----------- unpack-objects: refactor away unpack_non_delta_entry()\n8:  34ee6a28a54 ! 7:  5bc8fa9bc8d unpack-objects: use stream_loose_object() to unpack large objects\n    @@ Documentation/config/core.txt: usage, at the slight expense of increased disk us\n      \tSpecifies the pathname to the file that contains patterns to\n     \n      ## builtin/unpack-objects.c ##\n    -@@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type type,\n    - \t}\n    +@@ builtin/unpack-objects.c: static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n    + \t\twrite_object(nr, type, buf, size);\n      }\n      \n     +struct input_zstream_data {\n    @@ builtin/unpack-objects.c: static void added_object(unsigned nr, enum object_type\n      \t\t\t\tvoid *delta_data, unsigned long delta_size)\n      {\n     @@ builtin/unpack-objects.c: static void unpack_one(unsigned nr)\n    + \t}\n      \n      \tswitch (type) {\n    - \tcase OBJ_BLOB:\n    ++\tcase OBJ_BLOB:\n     +\t\tif (!dry_run && size > big_file_threshold) {\n     +\t\t\tstream_blob(size, nr);\n     +\t\t\treturn;\n    @@ builtin/unpack-objects.c: static void unpack_one(unsigned nr)\n     +\t\t/* fallthrough */\n      \tcase OBJ_COMMIT:\n      \tcase OBJ_TREE:\n    +-\tcase OBJ_BLOB:\n      \tcase OBJ_TAG:\n    + \t\tunpack_non_delta_entry(type, size, nr);\n    + \t\treturn;\n     \n      ## t/t5351-unpack-large-objects.sh ##\n     @@ t/t5351-unpack-large-objects.sh: test_description='git unpack-objects with large objects'\n-- \n2.36.1.1124.g52838f02905\n\n"},{"id":"456663","messageId":"patch-v13-2.7-b3568f0c5c0-20220604T095113Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v13-0.7-00000000000-20220604T095113Z-avarab@gmail.com","subject":"[PATCH v13 2/7] object-file.c: do fsync() and close() before post-write die()","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-06-04T10:10:23Z","receivedAt":"2022-06-04T10:10:51Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Change write_loose_object() to do an fsync() and close() before the\noideq() sanity check at the end. This change re-joins code that was\nsplit up by the die() sanity check added in 748af44c63e (sha1_file: be\nparanoid when creating loose objects, 2010-02-21).\n\nI don't think that this change matters in itself, if we called die()\nit was possible that our data wouldn't fully make it to disk, but in\nany case we were writing data that we'd consider corrupted. It's\npossible that a subsequent \"git fsck\" will be less confused now.\n\nThe real reason to make this change is that in a subsequent commit\nwe'll split this code in write_loose_object() into a utility function,\nall its callers will want the preceding sanity checks, but not the\n\"oideq\" check. By moving the close_loose_object() earlier it'll be\neasier to reason about the introduction of the utility function.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n object-file.c | 4 ++--\n 1 file changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 79eb8339b60..e4a83012ba4 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2012,12 +2012,12 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n \t\t    ret);\n \tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n+\tclose_loose_object(fd, tmp_file.buf);\n+\n \tif (!oideq(oid, &parano_oid))\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n \n-\tclose_loose_object(fd, tmp_file.buf);\n-\n \tif (mtime) {\n \t\tstruct utimbuf utb;\n \t\tutb.actime = mtime;\n-- \n2.36.1.1124.g52838f02905\n\n"},{"id":"456664","messageId":"patch-v13-3.7-9dc0f56878a-20220604T095113Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v13-0.7-00000000000-20220604T095113Z-avarab@gmail.com","subject":"[PATCH v13 3/7] object-file.c: refactor write_loose_object() to several steps","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-06-04T10:10:24Z","receivedAt":"2022-06-04T10:10:54Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWhen writing a large blob using \"write_loose_object()\", we have to pass\na buffer with the whole content of the blob, and this behavior will\nconsume lots of memory and may cause OOM. We will introduce a stream\nversion function (\"stream_loose_object()\") in later commit to resolve\nthis issue.\n\nBefore introducing that streaming function, do some refactoring on\n\"write_loose_object()\" to reuse code for both versions.\n\nRewrite \"write_loose_object()\" as follows:\n\n 1. Figure out a path for the (temp) object file. This step is only\n    used in \"write_loose_object()\".\n\n 2. Move common steps for starting to write loose objects into a new\n    function \"start_loose_object_common()\".\n\n 3. Compress data.\n\n 4. Move common steps for ending zlib stream into a new function\n    \"end_loose_object_common()\".\n\n 5. Close fd and finalize the object file.\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n object-file.c | 102 +++++++++++++++++++++++++++++++++++++-------------\n 1 file changed, 76 insertions(+), 26 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex e4a83012ba4..ce8b52a8dc3 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1951,6 +1951,75 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n \treturn fd;\n }\n \n+/**\n+ * Common steps for loose object writers to start writing loose\n+ * objects:\n+ *\n+ * - Create tmpfile for the loose object.\n+ * - Setup zlib stream for compression.\n+ * - Start to feed header to zlib stream.\n+ *\n+ * Returns a \"fd\", which should later be provided to\n+ * end_loose_object_common().\n+ */\n+static int start_loose_object_common(struct strbuf *tmp_file,\n+\t\t\t\t     const char *filename, unsigned flags,\n+\t\t\t\t     git_zstream *stream,\n+\t\t\t\t     unsigned char *buf, size_t buflen,\n+\t\t\t\t     git_hash_ctx *c,\n+\t\t\t\t     char *hdr, int hdrlen)\n+{\n+\tint fd;\n+\n+\tfd = create_tmpfile(tmp_file, filename);\n+\tif (fd < 0) {\n+\t\tif (flags & HASH_SILENT)\n+\t\t\treturn -1;\n+\t\telse if (errno == EACCES)\n+\t\t\treturn error(_(\"insufficient permission for adding \"\n+\t\t\t\t       \"an object to repository database %s\"),\n+\t\t\t\t     get_object_directory());\n+\t\telse\n+\t\t\treturn error_errno(\n+\t\t\t\t_(\"unable to create temporary file\"));\n+\t}\n+\n+\t/*  Setup zlib stream for compression */\n+\tgit_deflate_init(stream, zlib_compression_level);\n+\tstream->next_out = buf;\n+\tstream->avail_out = buflen;\n+\tthe_hash_algo->init_fn(c);\n+\n+\t/*  Start to feed header to zlib stream */\n+\tstream->next_in = (unsigned char *)hdr;\n+\tstream->avail_in = hdrlen;\n+\twhile (git_deflate(stream, 0) == Z_OK)\n+\t\t; /* nothing */\n+\tthe_hash_algo->update_fn(c, hdr, hdrlen);\n+\n+\treturn fd;\n+}\n+\n+/**\n+ * Common steps for loose object writers to end writing loose objects:\n+ *\n+ * - End the compression of zlib stream.\n+ * - Get the calculated oid to \"oid\".\n+ * - fsync() and close() the \"fd\"\n+ */\n+static int end_loose_object_common(git_hash_ctx *c, git_zstream *stream,\n+\t\t\t\t   struct object_id *oid)\n+{\n+\tint ret;\n+\n+\tret = git_deflate_end_gently(stream);\n+\tif (ret != Z_OK)\n+\t\treturn ret;\n+\tthe_hash_algo->final_oid_fn(oid, c);\n+\n+\treturn Z_OK;\n+}\n+\n static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\t\t      int hdrlen, const void *buf, unsigned long len,\n \t\t\t      time_t mtime, unsigned flags)\n@@ -1968,28 +2037,11 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \n \tloose_object_path(the_repository, &filename, oid);\n \n-\tfd = create_tmpfile(&tmp_file, filename.buf);\n-\tif (fd < 0) {\n-\t\tif (flags & HASH_SILENT)\n-\t\t\treturn -1;\n-\t\telse if (errno == EACCES)\n-\t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n-\t\telse\n-\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n-\t}\n-\n-\t/* Set it up */\n-\tgit_deflate_init(&stream, zlib_compression_level);\n-\tstream.next_out = compressed;\n-\tstream.avail_out = sizeof(compressed);\n-\tthe_hash_algo->init_fn(&c);\n-\n-\t/* First header.. */\n-\tstream.next_in = (unsigned char *)hdr;\n-\tstream.avail_in = hdrlen;\n-\twhile (git_deflate(&stream, 0) == Z_OK)\n-\t\t; /* nothing */\n-\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n+\tfd = start_loose_object_common(&tmp_file, filename.buf, flags,\n+\t\t\t\t       &stream, compressed, sizeof(compressed),\n+\t\t\t\t       &c, hdr, hdrlen);\n+\tif (fd < 0)\n+\t\treturn -1;\n \n \t/* Then the data itself.. */\n \tstream.next_in = (void *)buf;\n@@ -2007,11 +2059,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tif (ret != Z_STREAM_END)\n \t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n \t\t    ret);\n-\tret = git_deflate_end_gently(&stream);\n+\tret = end_loose_object_common(&c, &stream, &parano_oid);\n \tif (ret != Z_OK)\n-\t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n-\t\t    ret);\n-\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n+\t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid), ret);\n \tclose_loose_object(fd, tmp_file.buf);\n \n \tif (!oideq(oid, &parano_oid))\n-- \n2.36.1.1124.g52838f02905\n\n"},{"id":"456665","messageId":"patch-v13-4.7-a0434835fe7-20220604T095113Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v13-0.7-00000000000-20220604T095113Z-avarab@gmail.com","subject":"[PATCH v13 4/7] object-file.c: factor out deflate part of write_loose_object()","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-06-04T10:10:25Z","receivedAt":"2022-06-04T10:10:56Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Split out the part of write_loose_object() that deals with calling\ngit_deflate() into a utility function, a subsequent commit will\nintroduce another function that'll make use of it.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n object-file.c | 31 +++++++++++++++++++++++++------\n 1 file changed, 25 insertions(+), 6 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex ce8b52a8dc3..7946fa5e088 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2000,6 +2000,28 @@ static int start_loose_object_common(struct strbuf *tmp_file,\n \treturn fd;\n }\n \n+/**\n+ * Common steps for the inner git_deflate() loop for writing loose\n+ * objects. Returns what git_deflate() returns.\n+ */\n+static int write_loose_object_common(git_hash_ctx *c,\n+\t\t\t\t     git_zstream *stream, const int flush,\n+\t\t\t\t     unsigned char *in0, const int fd,\n+\t\t\t\t     unsigned char *compressed,\n+\t\t\t\t     const size_t compressed_len)\n+{\n+\tint ret;\n+\n+\tret = git_deflate(stream, flush ? Z_FINISH : 0);\n+\tthe_hash_algo->update_fn(c, in0, stream->next_in - in0);\n+\tif (write_buffer(fd, compressed, stream->next_out - compressed) < 0)\n+\t\tdie(_(\"unable to write loose object file\"));\n+\tstream->next_out = compressed;\n+\tstream->avail_out = compressed_len;\n+\n+\treturn ret;\n+}\n+\n /**\n  * Common steps for loose object writers to end writing loose objects:\n  *\n@@ -2048,12 +2070,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tstream.avail_in = len;\n \tdo {\n \t\tunsigned char *in0 = stream.next_in;\n-\t\tret = git_deflate(&stream, Z_FINISH);\n-\t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n-\t\tif (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n-\t\t\tdie(_(\"unable to write loose object file\"));\n-\t\tstream.next_out = compressed;\n-\t\tstream.avail_out = sizeof(compressed);\n+\n+\t\tret = write_loose_object_common(&c, &stream, 1, in0, fd,\n+\t\t\t\t\t\tcompressed, sizeof(compressed));\n \t} while (ret == Z_OK);\n \n \tif (ret != Z_STREAM_END)\n-- \n2.36.1.1124.g52838f02905\n\n"},{"id":"456666","messageId":"patch-v13-5.7-0b07b29836b-20220604T095113Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v13-0.7-00000000000-20220604T095113Z-avarab@gmail.com","subject":"[PATCH v13 5/7] object-file.c: add \"stream_loose_object()\" to handle large object","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-06-04T10:10:26Z","receivedAt":"2022-06-04T10:10:59Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nIf we want unpack and write a loose object using \"write_loose_object\",\nwe have to feed it with a buffer with the same size of the object, which\nwill consume lots of memory and may cause OOM. This can be improved by\nfeeding data to \"stream_loose_object()\" in a stream.\n\nAdd a new function \"stream_loose_object()\", which is a stream version of\n\"write_loose_object()\" but with a low memory footprint. We will use this\nfunction to unpack large blob object in later commit.\n\nAnother difference with \"write_loose_object()\" is that we have no chance\nto run \"write_object_file_prepare()\" to calculate the oid in advance.\nIn \"write_loose_object()\", we know the oid and we can write the\ntemporary file in the same directory as the final object, but for an\nobject with an undetermined oid, we don't know the exact directory for\nthe object.\n\nStill, we need to save the temporary file we're preparing\nsomewhere. We'll do that in the top-level \".git/objects/\"\ndirectory (or whatever \"GIT_OBJECT_DIRECTORY\" is set to). Once we've\nstreamed it we'll know the OID, and will move it to its canonical\npath.\n\n\"freshen_packed_object()\" or \"freshen_loose_object()\" will be called\ninside \"stream_loose_object()\" after obtaining the \"oid\".\n\nHelped-by: René Scharfe <l.s.r@web.de>\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n object-file.c  | 100 +++++++++++++++++++++++++++++++++++++++++++++++++\n object-store.h |   8 ++++\n 2 files changed, 108 insertions(+)\n\ndiff --git a/object-file.c b/object-file.c\nindex 7946fa5e088..9fd449693c4 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2119,6 +2119,106 @@ static int freshen_packed_object(const struct object_id *oid)\n \treturn 1;\n }\n \n+int stream_loose_object(struct input_stream *in_stream, size_t len,\n+\t\t\tstruct object_id *oid)\n+{\n+\tint fd, ret, err = 0, flush = 0;\n+\tunsigned char compressed[4096];\n+\tgit_zstream stream;\n+\tgit_hash_ctx c;\n+\tstruct strbuf tmp_file = STRBUF_INIT;\n+\tstruct strbuf filename = STRBUF_INIT;\n+\tint dirlen;\n+\tchar hdr[MAX_HEADER_LEN];\n+\tint hdrlen;\n+\n+\t/* Since oid is not determined, save tmp file to odb path. */\n+\tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n+\thdrlen = format_object_header(hdr, sizeof(hdr), OBJ_BLOB, len);\n+\n+\t/*\n+\t * Common steps for write_loose_object and stream_loose_object to\n+\t * start writing loose objects:\n+\t *\n+\t *  - Create tmpfile for the loose object.\n+\t *  - Setup zlib stream for compression.\n+\t *  - Start to feed header to zlib stream.\n+\t */\n+\tfd = start_loose_object_common(&tmp_file, filename.buf, 0,\n+\t\t\t\t       &stream, compressed, sizeof(compressed),\n+\t\t\t\t       &c, hdr, hdrlen);\n+\tif (fd < 0) {\n+\t\terr = -1;\n+\t\tgoto cleanup;\n+\t}\n+\n+\t/* Then the data itself.. */\n+\tdo {\n+\t\tunsigned char *in0 = stream.next_in;\n+\n+\t\tif (!stream.avail_in && !in_stream->is_finished) {\n+\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n+\t\t\tstream.next_in = (void *)in;\n+\t\t\tin0 = (unsigned char *)in;\n+\t\t\t/* All data has been read. */\n+\t\t\tif (in_stream->is_finished)\n+\t\t\t\tflush = 1;\n+\t\t}\n+\t\tret = write_loose_object_common(&c, &stream, flush, in0, fd,\n+\t\t\t\t\t\tcompressed, sizeof(compressed));\n+\t\t/*\n+\t\t * Unlike write_loose_object(), we do not have the entire\n+\t\t * buffer. If we get Z_BUF_ERROR due to too few input bytes,\n+\t\t * then we'll replenish them in the next input_stream->read()\n+\t\t * call when we loop.\n+\t\t */\n+\t} while (ret == Z_OK || ret == Z_BUF_ERROR);\n+\n+\tif (stream.total_in != len + hdrlen)\n+\t\tdie(_(\"write stream object %ld != %\"PRIuMAX), stream.total_in,\n+\t\t    (uintmax_t)len + hdrlen);\n+\n+\t/* Common steps for write_loose_object and stream_loose_object to\n+\t * end writing loose oject:\n+\t *\n+\t *  - End the compression of zlib stream.\n+\t *  - Get the calculated oid.\n+\t */\n+\tif (ret != Z_STREAM_END)\n+\t\tdie(_(\"unable to stream deflate new object (%d)\"), ret);\n+\tret = end_loose_object_common(&c, &stream, oid);\n+\tif (ret != Z_OK)\n+\t\tdie(_(\"deflateEnd on stream object failed (%d)\"), ret);\n+\tclose_loose_object(fd, tmp_file.buf);\n+\n+\tif (freshen_packed_object(oid) || freshen_loose_object(oid)) {\n+\t\tunlink_or_warn(tmp_file.buf);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tloose_object_path(the_repository, &filename, oid);\n+\n+\t/* We finally know the object path, and create the missing dir. */\n+\tdirlen = directory_size(filename.buf);\n+\tif (dirlen) {\n+\t\tstruct strbuf dir = STRBUF_INIT;\n+\t\tstrbuf_add(&dir, filename.buf, dirlen);\n+\n+\t\tif (mkdir_in_gitdir(dir.buf) && errno != EEXIST) {\n+\t\t\terr = error_errno(_(\"unable to create directory %s\"), dir.buf);\n+\t\t\tstrbuf_release(&dir);\n+\t\t\tgoto cleanup;\n+\t\t}\n+\t\tstrbuf_release(&dir);\n+\t}\n+\n+\terr = finalize_object_file(tmp_file.buf, filename.buf);\n+cleanup:\n+\tstrbuf_release(&tmp_file);\n+\tstrbuf_release(&filename);\n+\treturn err;\n+}\n+\n int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t    enum object_type type, struct object_id *oid,\n \t\t\t    unsigned flags)\ndiff --git a/object-store.h b/object-store.h\nindex 539ea439046..5222ee54600 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -46,6 +46,12 @@ struct object_directory {\n \tchar *path;\n };\n \n+struct input_stream {\n+\tconst void *(*read)(struct input_stream *, unsigned long *len);\n+\tvoid *data;\n+\tint is_finished;\n+};\n+\n KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n \tstruct object_directory *, 1, fspathhash, fspatheq)\n \n@@ -269,6 +275,8 @@ static inline int write_object_file(const void *buf, unsigned long len,\n int write_object_file_literally(const void *buf, unsigned long len,\n \t\t\t\tconst char *type, struct object_id *oid,\n \t\t\t\tunsigned flags);\n+int stream_loose_object(struct input_stream *in_stream, size_t len,\n+\t\t\tstruct object_id *oid);\n \n /*\n  * Add an object file to the in-memory object store, without writing it\n-- \n2.36.1.1124.g52838f02905\n\n"},{"id":"456667","messageId":"patch-v13-7.7-5bc8fa9bc8d-20220604T095113Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v13-0.7-00000000000-20220604T095113Z-avarab@gmail.com","subject":"[PATCH v13 7/7] unpack-objects: use stream_loose_object() to unpack large objects","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-06-04T10:10:28Z","receivedAt":"2022-06-04T10:11:00Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nMake use of the stream_loose_object() function introduced in the\npreceding commit to unpack large objects. Before this we'd need to\nmalloc() the size of the blob before unpacking it, which could cause\nOOM with very large blobs.\n\nWe could use the new streaming interface to unpack all blobs, but\ndoing so would be much slower, as demonstrated e.g. with this\nbenchmark using git-hyperfine[0]:\n\n\trm -rf /tmp/scalar.git &&\n\tgit clone --bare https://github.com/Microsoft/scalar.git /tmp/scalar.git &&\n\tmv /tmp/scalar.git/objects/pack/*.pack /tmp/scalar.git/my.pack &&\n\tgit hyperfine \\\n\t\t-r 2 --warmup 1 \\\n\t\t-L rev origin/master,HEAD -L v \"10,512,1k,1m\" \\\n\t\t-s 'make' \\\n\t\t-p 'git init --bare dest.git' \\\n\t\t-c 'rm -rf dest.git' \\\n\t\t'./git -C dest.git -c core.bigFileThreshold={v} unpack-objects </tmp/scalar.git/my.pack'\n\nHere we'll perform worse with lower core.bigFileThreshold settings\nwith this change in terms of speed, but we're getting lower memory use\nin return:\n\n\tSummary\n\t  './git -C dest.git -c core.bigFileThreshold=10 unpack-objects </tmp/scalar.git/my.pack' in 'origin/master' ran\n\t    1.01 ± 0.01 times faster than './git -C dest.git -c core.bigFileThreshold=1k unpack-objects </tmp/scalar.git/my.pack' in 'origin/master'\n\t    1.01 ± 0.01 times faster than './git -C dest.git -c core.bigFileThreshold=1m unpack-objects </tmp/scalar.git/my.pack' in 'origin/master'\n\t    1.01 ± 0.02 times faster than './git -C dest.git -c core.bigFileThreshold=1m unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n\t    1.02 ± 0.00 times faster than './git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/scalar.git/my.pack' in 'origin/master'\n\t    1.09 ± 0.01 times faster than './git -C dest.git -c core.bigFileThreshold=1k unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n\t    1.10 ± 0.00 times faster than './git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n\t    1.11 ± 0.00 times faster than './git -C dest.git -c core.bigFileThreshold=10 unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n\nA better benchmark to demonstrate the benefits of that this one, which\ncreates an artificial repo with a 1, 25, 50, 75 and 100MB blob:\n\n\trm -rf /tmp/repo &&\n\tgit init /tmp/repo &&\n\t(\n\t\tcd /tmp/repo &&\n\t\tfor i in 1 25 50 75 100\n\t\tdo\n\t\t\tdd if=/dev/urandom of=blob.$i count=$(($i*1024)) bs=1024\n\t\tdone &&\n\t\tgit add blob.* &&\n\t\tgit commit -mblobs &&\n\t\tgit gc &&\n\t\tPACK=$(echo .git/objects/pack/pack-*.pack) &&\n\t\tcp \"$PACK\" my.pack\n\t) &&\n\tgit hyperfine \\\n\t\t--show-output \\\n\t\t-L rev origin/master,HEAD -L v \"512,50m,100m\" \\\n\t\t-s 'make' \\\n\t\t-p 'git init --bare dest.git' \\\n\t\t-c 'rm -rf dest.git' \\\n\t\t'/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold={v} unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum'\n\nUsing this test we'll always use >100MB of memory on\norigin/master (around ~105MB), but max out at e.g. ~55MB if we set\ncore.bigFileThreshold=50m.\n\nThe relevant \"Maximum resident set size\" lines were manually added\nbelow the relevant benchmark:\n\n  '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=50m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'origin/master' ran\n        Maximum resident set size (kbytes): 107080\n    1.02 ± 0.78 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'origin/master'\n        Maximum resident set size (kbytes): 106968\n    1.09 ± 0.79 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=100m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'origin/master'\n        Maximum resident set size (kbytes): 107032\n    1.42 ± 1.07 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=100m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'HEAD'\n        Maximum resident set size (kbytes): 107072\n    1.83 ± 1.02 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=50m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'HEAD'\n        Maximum resident set size (kbytes): 55704\n    2.16 ± 1.19 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'HEAD'\n        Maximum resident set size (kbytes): 4564\n\nThis shows that if you have enough memory this new streaming method is\nslower the lower you set the streaming threshold, but the benefit is\nmore bounded memory use.\n\nAn earlier version of this patch introduced a new\n\"core.bigFileStreamingThreshold\" instead of re-using the existing\n\"core.bigFileThreshold\" variable[1]. As noted in a detailed overview\nof its users in [2] using it has several different meanings.\n\nStill, we consider it good enough to simply re-use it. While it's\npossible that someone might want to e.g. consider objects \"small\" for\nthe purposes of diffing but \"big\" for the purposes of writing them\nsuch use-cases are probably too obscure to worry about. We can always\nsplit up \"core.bigFileThreshold\" in the future if there's a need for\nthat.\n\n0. https://github.com/avar/git-hyperfine/\n1. https://lore.kernel.org/git/20211210103435.83656-1-chiyutianyi@gmail.com/\n2. https://lore.kernel.org/git/20220120112114.47618-5-chiyutianyi@gmail.com/\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Derrick Stolee <stolee@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n Documentation/config/core.txt   |  4 +-\n builtin/unpack-objects.c        | 69 ++++++++++++++++++++++++++++++++-\n t/t5351-unpack-large-objects.sh | 26 +++++++++++--\n 3 files changed, 93 insertions(+), 6 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex ff6ae6bb647..b97bc7e3e55 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -468,8 +468,8 @@ usage, at the slight expense of increased disk usage.\n * Will be generally be streamed when written, which avoids excessive\n memory usage, at the cost of some fixed overhead. Commands that make\n use of this include linkgit:git-archive[1],\n-linkgit:git-fast-import[1], linkgit:git-index-pack[1] and\n-linkgit:git-fsck[1].\n+linkgit:git-fast-import[1], linkgit:git-index-pack[1],\n+linkgit:git-unpack-objects[1] and linkgit:git-fsck[1].\n \n core.excludesFile::\n \tSpecifies the pathname to the file that contains patterns to\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 64abba8dbac..d3124202f54 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -348,6 +348,68 @@ static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n \t\twrite_object(nr, type, buf, size);\n }\n \n+struct input_zstream_data {\n+\tgit_zstream *zstream;\n+\tunsigned char buf[8192];\n+\tint status;\n+};\n+\n+static const void *feed_input_zstream(struct input_stream *in_stream,\n+\t\t\t\t      unsigned long *readlen)\n+{\n+\tstruct input_zstream_data *data = in_stream->data;\n+\tgit_zstream *zstream = data->zstream;\n+\tvoid *in = fill(1);\n+\n+\tif (in_stream->is_finished) {\n+\t\t*readlen = 0;\n+\t\treturn NULL;\n+\t}\n+\n+\tzstream->next_out = data->buf;\n+\tzstream->avail_out = sizeof(data->buf);\n+\tzstream->next_in = in;\n+\tzstream->avail_in = len;\n+\n+\tdata->status = git_inflate(zstream, 0);\n+\n+\tin_stream->is_finished = data->status != Z_OK;\n+\tuse(len - zstream->avail_in);\n+\t*readlen = sizeof(data->buf) - zstream->avail_out;\n+\n+\treturn data->buf;\n+}\n+\n+static void stream_blob(unsigned long size, unsigned nr)\n+{\n+\tgit_zstream zstream = { 0 };\n+\tstruct input_zstream_data data = { 0 };\n+\tstruct input_stream in_stream = {\n+\t\t.read = feed_input_zstream,\n+\t\t.data = &data,\n+\t};\n+\tstruct obj_info *info = &obj_list[nr];\n+\n+\tdata.zstream = &zstream;\n+\tgit_inflate_init(&zstream);\n+\n+\tif (stream_loose_object(&in_stream, size, &info->oid))\n+\t\tdie(_(\"failed to write object in stream\"));\n+\n+\tif (data.status != Z_STREAM_END)\n+\t\tdie(_(\"inflate returned (%d)\"), data.status);\n+\tgit_inflate_end(&zstream);\n+\n+\tif (strict) {\n+\t\tstruct blob *blob = lookup_blob(the_repository, &info->oid);\n+\n+\t\tif (!blob)\n+\t\t\tdie(_(\"invalid blob object from stream\"));\n+\t\tblob->object.flags |= FLAG_WRITTEN;\n+\t}\n+\tinfo->obj = NULL;\n+}\n+\n static int resolve_against_held(unsigned nr, const struct object_id *base,\n \t\t\t\tvoid *delta_data, unsigned long delta_size)\n {\n@@ -480,9 +542,14 @@ static void unpack_one(unsigned nr)\n \t}\n \n \tswitch (type) {\n+\tcase OBJ_BLOB:\n+\t\tif (!dry_run && size > big_file_threshold) {\n+\t\t\tstream_blob(size, nr);\n+\t\t\treturn;\n+\t\t}\n+\t\t/* fallthrough */\n \tcase OBJ_COMMIT:\n \tcase OBJ_TREE:\n-\tcase OBJ_BLOB:\n \tcase OBJ_TAG:\n \t\tunpack_non_delta_entry(type, size, nr);\n \t\treturn;\ndiff --git a/t/t5351-unpack-large-objects.sh b/t/t5351-unpack-large-objects.sh\nindex 8d84313221c..461ca060b2b 100755\n--- a/t/t5351-unpack-large-objects.sh\n+++ b/t/t5351-unpack-large-objects.sh\n@@ -9,7 +9,8 @@ test_description='git unpack-objects with large objects'\n \n prepare_dest () {\n \ttest_when_finished \"rm -rf dest.git\" &&\n-\tgit init --bare dest.git\n+\tgit init --bare dest.git &&\n+\tgit -C dest.git config core.bigFileThreshold \"$1\"\n }\n \n test_expect_success \"create large objects (1.5 MB) and PACK\" '\n@@ -26,16 +27,35 @@ test_expect_success 'set memory limitation to 1MB' '\n '\n \n test_expect_success 'unpack-objects failed under memory limitation' '\n-\tprepare_dest &&\n+\tprepare_dest 2m &&\n \ttest_must_fail git -C dest.git unpack-objects <pack-$PACK.pack 2>err &&\n \tgrep \"fatal: attempting to allocate\" err\n '\n \n test_expect_success 'unpack-objects works with memory limitation in dry-run mode' '\n-\tprepare_dest &&\n+\tprepare_dest 2m &&\n \tgit -C dest.git unpack-objects -n <pack-$PACK.pack &&\n \ttest_stdout_line_count = 0 find dest.git/objects -type f &&\n \ttest_dir_is_empty dest.git/objects/pack\n '\n \n+test_expect_success 'unpack big object in stream' '\n+\tprepare_dest 1m &&\n+\tgit -C dest.git unpack-objects <pack-$PACK.pack &&\n+\ttest_dir_is_empty dest.git/objects/pack\n+'\n+\n+test_expect_success 'do not unpack existing large objects' '\n+\tprepare_dest 1m &&\n+\tgit -C dest.git index-pack --stdin <pack-$PACK.pack &&\n+\tgit -C dest.git unpack-objects <pack-$PACK.pack &&\n+\n+\t# The destination came up with the exact same pack...\n+\tDEST_PACK=$(echo dest.git/objects/pack/pack-*.pack) &&\n+\ttest_cmp pack-$PACK.pack $DEST_PACK &&\n+\n+\t# ...and wrote no loose objects\n+\ttest_stdout_line_count = 0 find dest.git/objects -type f ! -name \"pack-*\"\n+'\n+\n test_done\n-- \n2.36.1.1124.g52838f02905\n\n"},{"id":"456668","messageId":"patch-v13-6.7-5ed79c58b18-20220604T095113Z-avarab@gmail.com","threadId":"56672","inReplyTo":"cover-v13-0.7-00000000000-20220604T095113Z-avarab@gmail.com","subject":"[PATCH v13 6/7] core doc: modernize core.bigFileThreshold documentation","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-06-04T10:10:27Z","receivedAt":"2022-06-04T10:11:00Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"The core.bigFileThreshold documentation has been largely unchanged\nsince 5eef828bc03 (fast-import: Stream very large blobs directly to\npack, 2010-02-01).\n\nBut since then this setting has been expanded to affect a lot more\nthan that description indicated. Most notably in how \"git diff\" treats\nthem, see 6bf3b813486 (diff --stat: mark any file larger than\ncore.bigfilethreshold binary, 2014-08-16).\n\nIn addition to that, numerous commands and APIs make use of a\nstreaming mode for files above this threshold.\n\nSo let's attempt to summarize 12 years of changes in behavior, which\ncan be seen with:\n\n    git log --oneline -Gbig_file_thre 5eef828bc03.. -- '*.c'\n\nTo do that turn this into a bullet-point list. The summary Han Xin\nproduced in [1] helped a lot, but is a bit too detailed for\ndocumentation aimed at users. Let's instead summarize how\nuser-observable behavior differs, and generally describe how we tend\nto stream these files in various commands.\n\n1. https://lore.kernel.org/git/20220120112114.47618-5-chiyutianyi@gmail.com/\n\nHelped-by: Han Xin <chiyutianyi@gmail.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n Documentation/config/core.txt | 33 ++++++++++++++++++++++++---------\n 1 file changed, 24 insertions(+), 9 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex 41e330f3069..ff6ae6bb647 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -444,17 +444,32 @@ You probably do not need to adjust this value.\n Common unit suffixes of 'k', 'm', or 'g' are supported.\n \n core.bigFileThreshold::\n-\tFiles larger than this size are stored deflated, without\n-\tattempting delta compression.  Storing large files without\n-\tdelta compression avoids excessive memory usage, at the\n-\tslight expense of increased disk usage. Additionally files\n-\tlarger than this size are always treated as binary.\n+\tThe size of files considered \"big\", which as discussed below\n+\tchanges the behavior of numerous git commands, as well as how\n+\tsuch files are stored within the repository. The default is\n+\t512 MiB. Common unit suffixes of 'k', 'm', or 'g' are\n+\tsupported.\n +\n-Default is 512 MiB on all platforms.  This should be reasonable\n-for most projects as source code and other text files can still\n-be delta compressed, but larger binary media files won't be.\n+Files above the configured limit will be:\n +\n-Common unit suffixes of 'k', 'm', or 'g' are supported.\n+* Stored deflated, without attempting delta compression.\n++\n+The default limit is primarily set with this use-case in mind. With it\n+most projects will have their source code and other text files delta\n+compressed, but not larger binary media files.\n++\n+Storing large files without delta compression avoids excessive memory\n+usage, at the slight expense of increased disk usage.\n++\n+* Will be treated as if though they were labeled \"binary\" (see\n+  linkgit:gitattributes[5]). This means that e.g. linkgit:git-log[1]\n+  and linkgit:git-diff[1] will not diffs for files above this limit.\n++\n+* Will be generally be streamed when written, which avoids excessive\n+memory usage, at the cost of some fixed overhead. Commands that make\n+use of this include linkgit:git-archive[1],\n+linkgit:git-fast-import[1], linkgit:git-index-pack[1] and\n+linkgit:git-fsck[1].\n \n core.excludesFile::\n \tSpecifies the pathname to the file that contains patterns to\n-- \n2.36.1.1124.g52838f02905\n\n"},{"id":"456739","messageId":"xmqqpmjl7i7y.fsf@gitster.g","threadId":"56672","inReplyTo":"patch-v13-1.7-12873fc9915-20220604T095113Z-avarab@gmail.com","subject":"Re: [PATCH v13 1/7] unpack-objects: low memory footprint for get_data() in dry_run mode","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-06-06T18:35:45Z","receivedAt":"2022-06-06T18:36:08Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason  <avarab@gmail.com> writes:\n\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> As the name implies, \"get_data(size)\" will allocate and return a given\n> amount of memory. Allocating memory for a large blob object may cause the\n> system to run out of memory. Before preparing to replace calling of\n> \"get_data()\" to unpack large blob objects in latter commits, refactor\n> \"get_data()\" to reduce memory footprint for dry_run mode.\n>\n> Because in dry_run mode, \"get_data()\" is only used to check the\n> integrity of data, and the returned buffer is not used at all, we can\n> allocate a smaller buffer and reuse it as zstream output. Therefore,\n\n\"reuse\" -> \"use\"\n\n> in dry_run mode, \"get_data()\" will release the allocated buffer and\n> return NULL instead of returning garbage data.\n\nIt makes it sound as if we used to return garbage data, but I do not\nthink that is what happened in reality.  Perhaps rewrite the last\nsentence like\n\n\tMake the function return NULL in the dry-run mode, as no\n\tcallers use the returned buffer.\n\nor something?\n\nThe overall logic sounds quite sensible.\n\n> The \"find [...]objects/?? -type f | wc -l\" test idiom being used here\n> is adapted from the same \"find\" use added to another test in\n> d9545c7f465 (fast-import: implement unpack limit, 2016-04-25).\n\n\n> +/*\n> + * Decompress zstream from stdin and return specific size of data.\n\n\"specific size\"?  The caller specifies the size of data (because it\nknows a-priori how many bytes the zstream should inflate to), so\n\n    Decompress zstream from the standard input into a newly\n    allocated buffer of specified size and return the buffer.\n\nor something, perhaps.  In any case, it needs to say that the caller\nis responsible for giving the \"right\" size.\n\n> + * The caller is responsible to free the returned buffer.\n> + *\n> + * But for dry_run mode, \"get_data()\" is only used to check the\n> + * integrity of data, and the returned buffer is not used at all.\n> + * Therefore, in dry_run mode, \"get_data()\" will release the small\n> + * allocated buffer which is reused to hold temporary zstream output\n> + * and return NULL instead of returning garbage data.\n> + */\n>  static void *get_data(unsigned long size)\n>  {\n>  \tgit_zstream stream;\n> -\tvoid *buf = xmallocz(size);\n> +\tunsigned long bufsize = dry_run && size > 8192 ? 8192 : size;\n> +\tvoid *buf = xmallocz(bufsize);\n\nOK.\n\n>  \tmemset(&stream, 0, sizeof(stream));\n>  \n>  \tstream.next_out = buf;\n> -\tstream.avail_out = size;\n> +\tstream.avail_out = bufsize;\n>  \tstream.next_in = fill(1);\n>  \tstream.avail_in = len;\n>  \tgit_inflate_init(&stream);\n> @@ -125,8 +136,15 @@ static void *get_data(unsigned long size)\n\nWhat's hidden in the pre-context is this bit:\n\n\t\tint ret = git_inflate(&stream, 0);\n\t\tuse(len - stream.avail_in);\n\t\tif (stream.total_out == size && ret == Z_STREAM_END)\n\t\t\tbreak;\n\t\tif (ret != Z_OK) {\n\t\t\terror(\"inflate returned %d\", ret);\n\t\t\tFREE_AND_NULL(buf);\n\t\t\tif (!recover)\n\t\t\t\texit(1);\n\t\t\thas_errors = 1;\n\t\t\tbreak;\n\t\t}\n\nand it is correct to use \"size\", not \"bufsize\", for this check.\nUnless we receive exactly the caller-specified \"size\" bytes from the\ninflated zstream with Z_STREAM_END, we want to detect an error and\nbail out.\n\nI am not sure if this is not loosening the error checking in the\ndry-run case, though.  In the original code, we set the avail_out\nto the total expected size so\n\n (1) if the caller gives too small a size, git_inflate() would stop\n     at stream.total_out with ret that is not STREAM_END nor OK,\n     bypassing the \"break\", and we catch the error.\n\n (2) if the caller gives too large a size, git_inflate() would stop\n     at the true size of inflated zstream, with STREAM_END and would\n     not hit this \"break\", and we catch the error.\n\nWith the new code, since we keep refreshing avail_out (see below),\ngit_inflate() does not even learn how many bytes we are _expecting_\nto see.  Is the error checking in the loop, with the updated code,\ncatch the mismatch between expected and actual size (plausibly\ncaused by a corrupted zstream) the same way as we do in the \nnon dry-run code path?\n\n>  \t\t}\n>  \t\tstream.next_in = fill(1);\n>  \t\tstream.avail_in = len;\n> +\t\tif (dry_run) {\n> +\t\t\t/* reuse the buffer in dry_run mode */\n> +\t\t\tstream.next_out = buf;\n> +\t\t\tstream.avail_out = bufsize;\n> +\t\t}\n>  \t}\n>  \tgit_inflate_end(&stream);\n> +\tif (dry_run)\n> +\t\tFREE_AND_NULL(buf);\n>  \treturn buf;\n>  }\n"},{"id":"456740","messageId":"xmqqilpd7hrh.fsf@gitster.g","threadId":"56672","inReplyTo":"patch-v13-2.7-b3568f0c5c0-20220604T095113Z-avarab@gmail.com","subject":"Re: [PATCH v13 2/7] object-file.c: do fsync() and close() before post-write die()","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-06-06T18:45:38Z","receivedAt":"2022-06-06T18:45:46Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason  <avarab@gmail.com> writes:\n\n> Change write_loose_object() to do an fsync() and close() before the\n> oideq() sanity check at the end. This change re-joins code that was\n> split up by the die() sanity check added in 748af44c63e (sha1_file: be\n> paranoid when creating loose objects, 2010-02-21).\n>\n> I don't think that this change matters in itself, if we called die()\n> it was possible that our data wouldn't fully make it to disk, but in\n> any case we were writing data that we'd consider corrupted. It's\n> possible that a subsequent \"git fsck\" will be less confused now.\n\nwrite_loose_object() \n\n - prepares a temporary file\n - deflates into the temporary file\n - closes and syncs it\n - moves the temporary file to the final locaiton\n\nAnd any die() inserted in between any of these steps will cause the\ncorrupt temporary file not to become the final loose object file.\n\nSo, \"git fsck\" does not need this change at all.\n\n> The real reason to make this change is that in a subsequent commit\n> we'll split this code in write_loose_object() into a utility function,\n> all its callers will want the preceding sanity checks, but not the\n> \"oideq\" check. By moving the close_loose_object() earlier it'll be\n> easier to reason about the introduction of the utility function.\n\nIf a \"split\" relies on the \"close and sync\" step being in any\nparticular place, that smells really fishy.  Is the series loosening\nthe object integrity check?  Are we adding some exploitable hole\ninto our codebase without people knowing, or something?  I am not\nsure if I am following the above logic.\n\n\n\n> Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> ---\n>  object-file.c | 4 ++--\n>  1 file changed, 2 insertions(+), 2 deletions(-)\n>\n> diff --git a/object-file.c b/object-file.c\n> index 79eb8339b60..e4a83012ba4 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -2012,12 +2012,12 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n>  \t\t    ret);\n>  \tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n> +\tclose_loose_object(fd, tmp_file.buf);\n> +\n>  \tif (!oideq(oid, &parano_oid))\n>  \t\tdie(_(\"confused by unstable object source data for %s\"),\n>  \t\t    oid_to_hex(oid));\n>  \n> -\tclose_loose_object(fd, tmp_file.buf);\n> -\n>  \tif (mtime) {\n>  \t\tstruct utimbuf utb;\n>  \t\tutb.actime = mtime;\n"},{"id":"456742","messageId":"xmqqy1y960hq.fsf@gitster.g","threadId":"56672","inReplyTo":"patch-v13-5.7-0b07b29836b-20220604T095113Z-avarab@gmail.com","subject":"Re: [PATCH v13 5/7] object-file.c: add \"stream_loose_object()\" to handle large object","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-06-06T19:44:01Z","receivedAt":"2022-06-06T19:44:16Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason  <avarab@gmail.com> writes:\n\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n>\n> If we want unpack and write a loose object using \"write_loose_object\",\n> we have to feed it with a buffer with the same size of the object, which\n> will consume lots of memory and may cause OOM. This can be improved by\n> feeding data to \"stream_loose_object()\" in a stream.\n>\n> Add a new function \"stream_loose_object()\", which is a stream version of\n> \"write_loose_object()\" but with a low memory footprint. We will use this\n> function to unpack large blob object in later commit.\n\nYay.\n\n> Another difference with \"write_loose_object()\" is that we have no chance\n> to run \"write_object_file_prepare()\" to calculate the oid in advance.\n\nThat is somewhat curious.  Is it fundamentally impossible, or is it\njust that this patch was written in such a way that conflates the\ntwo and it is cumbersome to split the \"we repeat the sequence of\nreading and deflating just a bit until we process all\" and the \"we\ncompute the hash over the data first and then we write out for\nreal\"?\n\n> In \"write_loose_object()\", we know the oid and we can write the\n> temporary file in the same directory as the final object, but for an\n> object with an undetermined oid, we don't know the exact directory for\n> the object.\n>\n> Still, we need to save the temporary file we're preparing\n> somewhere. We'll do that in the top-level \".git/objects/\"\n> directory (or whatever \"GIT_OBJECT_DIRECTORY\" is set to). Once we've\n> streamed it we'll know the OID, and will move it to its canonical\n> path.\n\nThis may have negative implications on some filesystems where cross\ndirectory links do not work atomically, but it is a small price to pay.\n\nI am very tempted to ask why we do not do this to _all_ loose object\nfiles.  Instead of running the machinery twice over the data (once to\ncompute the object name, then to compute the contents and write out),\nif we can produce loose object files of any size with a single pass,\nwouldn't that be an overall win?\n\nIs the fixed overhead, i.e. cost of setting up the streaming interface,\nreasonably large to make it not worth doing for smaller objects?\n\n> \"freshen_packed_object()\" or \"freshen_loose_object()\" will be called\n> inside \"stream_loose_object()\" after obtaining the \"oid\".\n\nThat much we can read from the patch text.  Saying just \"we do X\"\nwithout explaining \"why we do so\" in the proposed log message leaves\nreaders more confused than otherwise.  Why is it worth pointing out\nin the proposed log message?  Is the reason why we need to do so\ninvolve something tricky?\n\n> +int stream_loose_object(struct input_stream *in_stream, size_t len,\n> +\t\t\tstruct object_id *oid)\n> +{\n> +\tint fd, ret, err = 0, flush = 0;\n> +\tunsigned char compressed[4096];\n> +\tgit_zstream stream;\n> +\tgit_hash_ctx c;\n> +\tstruct strbuf tmp_file = STRBUF_INIT;\n> +\tstruct strbuf filename = STRBUF_INIT;\n> +\tint dirlen;\n> +\tchar hdr[MAX_HEADER_LEN];\n> +\tint hdrlen;\n> +\n> +\t/* Since oid is not determined, save tmp file to odb path. */\n> +\tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n> +\thdrlen = format_object_header(hdr, sizeof(hdr), OBJ_BLOB, len);\n> +\n> +\t/*\n> +\t * Common steps for write_loose_object and stream_loose_object to\n> +\t * start writing loose objects:\n> +\t *\n> +\t *  - Create tmpfile for the loose object.\n> +\t *  - Setup zlib stream for compression.\n> +\t *  - Start to feed header to zlib stream.\n> +\t */\n> +\tfd = start_loose_object_common(&tmp_file, filename.buf, 0,\n> +\t\t\t\t       &stream, compressed, sizeof(compressed),\n> +\t\t\t\t       &c, hdr, hdrlen);\n> +\tif (fd < 0) {\n> +\t\terr = -1;\n> +\t\tgoto cleanup;\n> +\t}\n> +\n> +\t/* Then the data itself.. */\n> +\tdo {\n> +\t\tunsigned char *in0 = stream.next_in;\n> +\n> +\t\tif (!stream.avail_in && !in_stream->is_finished) {\n> +\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n> +\t\t\tstream.next_in = (void *)in;\n> +\t\t\tin0 = (unsigned char *)in;\n> +\t\t\t/* All data has been read. */\n> +\t\t\tif (in_stream->is_finished)\n> +\t\t\t\tflush = 1;\n> +\t\t}\n> +\t\tret = write_loose_object_common(&c, &stream, flush, in0, fd,\n> +\t\t\t\t\t\tcompressed, sizeof(compressed));\n> +\t\t/*\n> +\t\t * Unlike write_loose_object(), we do not have the entire\n> +\t\t * buffer. If we get Z_BUF_ERROR due to too few input bytes,\n> +\t\t * then we'll replenish them in the next input_stream->read()\n> +\t\t * call when we loop.\n> +\t\t */\n> +\t} while (ret == Z_OK || ret == Z_BUF_ERROR);\n>\n> +\tif (stream.total_in != len + hdrlen)\n> +\t\tdie(_(\"write stream object %ld != %\"PRIuMAX), stream.total_in,\n> +\t\t    (uintmax_t)len + hdrlen);\n\n> +\t/* Common steps for write_loose_object and stream_loose_object to\n\nStyle.\n\n> +\t * end writing loose oject:\n> +\t *\n> +\t *  - End the compression of zlib stream.\n> +\t *  - Get the calculated oid.\n> +\t */\n> +\tif (ret != Z_STREAM_END)\n> +\t\tdie(_(\"unable to stream deflate new object (%d)\"), ret);\n\nGood to check this, after the loop exits above.  I was expecting to\nsee it immediately after the loop, but here is also OK.\n\n> +\tret = end_loose_object_common(&c, &stream, oid);\n> +\tif (ret != Z_OK)\n> +\t\tdie(_(\"deflateEnd on stream object failed (%d)\"), ret);\n> +\tclose_loose_object(fd, tmp_file.buf);\n> +\n> +\tif (freshen_packed_object(oid) || freshen_loose_object(oid)) {\n> +\t\tunlink_or_warn(tmp_file.buf);\n> +\t\tgoto cleanup;\n\nSo, we were told to write an object, we wrote to a temporary file,\nand we wanted to mark the object to be recent and found that there\nindeed is already the object.  We remove the temporary and do not\nleave the new copy of the object, and the value of err at this point\nis 0 (success) which is what is returned from cleanup: label.\n\nGood.\n\n> +\t}\n> +\n> +\tloose_object_path(the_repository, &filename, oid);\n> +\n> +\t/* We finally know the object path, and create the missing dir. */\n> +\tdirlen = directory_size(filename.buf);\n> +\tif (dirlen) {\n> +\t\tstruct strbuf dir = STRBUF_INIT;\n> +\t\tstrbuf_add(&dir, filename.buf, dirlen);\n> +\n> +\t\tif (mkdir_in_gitdir(dir.buf) && errno != EEXIST) {\n> +\t\t\terr = error_errno(_(\"unable to create directory %s\"), dir.buf);\n> +\t\t\tstrbuf_release(&dir);\n> +\t\t\tgoto cleanup;\n> +\t\t}\n> +\t\tstrbuf_release(&dir);\n> +\t}\n> +\n> +\terr = finalize_object_file(tmp_file.buf, filename.buf);\n> +cleanup:\n> +\tstrbuf_release(&tmp_file);\n> +\tstrbuf_release(&filename);\n> +\treturn err;\n> +}\n> +\n\n"},{"id":"456743","messageId":"xmqqpmjl6065.fsf@gitster.g","threadId":"56672","inReplyTo":"patch-v13-6.7-5ed79c58b18-20220604T095113Z-avarab@gmail.com","subject":"Re: [PATCH v13 6/7] core doc: modernize core.bigFileThreshold documentation","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-06-06T19:50:58Z","receivedAt":"2022-06-06T19:51:07Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason  <avarab@gmail.com> writes:\n\n> So let's attempt to summarize 12 years of changes in behavior, which\n> can be seen with:\n>\n>     git log --oneline -Gbig_file_thre 5eef828bc03.. -- '*.c'\n>\n> To do that turn this into a bullet-point list. The summary Han Xin\n> produced in [1] helped a lot, but is a bit too detailed for\n> documentation aimed at users. Let's instead summarize how\n> user-observable behavior differs, and generally describe how we tend\n> to stream these files in various commands.\n\nNicely studied.  Very much appreciated.\n\n>  core.bigFileThreshold::\n> -\tFiles larger than this size are stored deflated, without\n> -\tattempting delta compression.  Storing large files without\n> -\tdelta compression avoids excessive memory usage, at the\n> -\tslight expense of increased disk usage. Additionally files\n> -\tlarger than this size are always treated as binary.\n> +\tThe size of files considered \"big\", which as discussed below\n> +\tchanges the behavior of numerous git commands, as well as how\n> +\tsuch files are stored within the repository. The default is\n> +\t512 MiB. Common unit suffixes of 'k', 'm', or 'g' are\n> +\tsupported.\n>  +\n> -Default is 512 MiB on all platforms.  This should be reasonable\n> -for most projects as source code and other text files can still\n> -be delta compressed, but larger binary media files won't be.\n> +Files above the configured limit will be:\n>  +\n> -Common unit suffixes of 'k', 'm', or 'g' are supported.\n> +* Stored deflated, without attempting delta compression.\n\n\"even in packfiles\" (with or without \"even\") is better be there in\nthe sentence---loose objects are always stored deflated anyway.\n\n> +The default limit is primarily set with this use-case in mind. With it\n> +most projects will have their source code and other text files delta\n> +compressed, but not larger binary media files.\n> ++\n> +Storing large files without delta compression avoids excessive memory\n> +usage, at the slight expense of increased disk usage.\n\n> +* Will be treated as if though they were labeled \"binary\" (see\n> +  linkgit:gitattributes[5]). This means that e.g. linkgit:git-log[1]\n> +  and linkgit:git-diff[1] will not diffs for files above this limit.\n\nGood.  You can lose three words \"This means that\" and the sentence\nmeans the same thing, so lose them (I always recommend people to\nreread the sentence when they write \"This means that\" with an eye to\nrewrite it better---it often is a sign that either the previous\nsentence is insufficiently clear, in which case it can be discarded\nand description after the three words can be enhanced to a better\nresult).\n\n> +* Will be generally be streamed when written, which avoids excessive\n> +memory usage, at the cost of some fixed overhead. Commands that make\n> +use of this include linkgit:git-archive[1],\n> +linkgit:git-fast-import[1], linkgit:git-index-pack[1] and\n> +linkgit:git-fsck[1].\n\nNice.  And this series adds unpack-objects to the mix.\n\n>  core.excludesFile::\n>  \tSpecifies the pathname to the file that contains patterns to\n\nExcellent.\n\nThanks.\n"},{"id":"456751","messageId":"xmqqk09t5zmm.fsf@gitster.g","threadId":"56672","inReplyTo":"xmqqy1y960hq.fsf@gitster.g","subject":"Re: [PATCH v13 5/7] object-file.c: add \"stream_loose_object()\" to handle large object","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-06-06T20:02:41Z","receivedAt":"2022-06-06T20:03:20Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n>> Another difference with \"write_loose_object()\" is that we have no chance\n>> to run \"write_object_file_prepare()\" to calculate the oid in advance.\n>\n> That is somewhat curious.  Is it fundamentally impossible, or is it\n> just that this patch was written in such a way that conflates the\n> two and it is cumbersome to split the \"we repeat the sequence of\n> reading and deflating just a bit until we process all\" and the \"we\n> compute the hash over the data first and then we write out for\n> real\"?\n\nOK, the answer lies somewhere in between.\n\nThe initial user of this streaming interface reads from an incoming\npackfile and feeds the inflated bytestream to the interface, which\nmeans we cannot seek.  That meaks it \"fundamentally impossible\" for\nthat codepath (i.e. unpack-objects to read from packstream and write\nto on-disk loose objects).\n\nBut if the input source is seekable (e.g. a file in the working\ntree), there is no fundamental reason why the new interface has \"no\nchance to run prepare to calculate the oid in advance\".  It's just\nthat the such a different caller is not added by the series and we\nchose not to allow the \"prepare and then write\" two-step process,\nbecause we currently do not need it when this series lands.\n\n> I am very tempted to ask why we do not do this to _all_ loose object\n> files.  Instead of running the machinery twice over the data (once to\n> compute the object name, then to compute the contents and write out),\n> if we can produce loose object files of any size with a single pass,\n> wouldn't that be an overall win?\n\nThere is a patch later in the series whose proposed log message has\nbenchmarks to show that it is slower in general.  It still is\ncurious where the slowness comes from and if it is something we can\ntune, though.\n\nThanks.\n"},{"id":"456852","messageId":"7ba4858a-d1cc-a4eb-b6d6-4c04a5dd6ce7@gmail.com","threadId":"56672","inReplyTo":"patch-v13-5.7-0b07b29836b-20220604T095113Z-avarab@gmail.com","subject":"Re: [PATCH v13 5/7] object-file.c: add \"stream_loose_object()\" to handle large object","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-06-07T19:53:31Z","receivedAt":"2022-06-08T03:26:44Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On 6/4/2022 3:10 AM, Ævar Arnfjörð Bjarmason wrote:\n> From: Han Xin <hanxin.hx@alibaba-inc.com>\n> \n> If we want unpack and write a loose object using \"write_loose_object\",\n> we have to feed it with a buffer with the same size of the object, which\n> will consume lots of memory and may cause OOM. This can be improved by\n> feeding data to \"stream_loose_object()\" in a stream.\n> \n> Add a new function \"stream_loose_object()\", which is a stream version of\n> \"write_loose_object()\" but with a low memory footprint. We will use this\n> function to unpack large blob object in later commit.\n> \n> Another difference with \"write_loose_object()\" is that we have no chance\n> to run \"write_object_file_prepare()\" to calculate the oid in advance.\n> In \"write_loose_object()\", we know the oid and we can write the\n> temporary file in the same directory as the final object, but for an\n> object with an undetermined oid, we don't know the exact directory for\n> the object.\n> \n> Still, we need to save the temporary file we're preparing\n> somewhere. We'll do that in the top-level \".git/objects/\"\n> directory (or whatever \"GIT_OBJECT_DIRECTORY\" is set to). Once we've\n> streamed it we'll know the OID, and will move it to its canonical\n> path.\n> \n\nI think this new logic doesn't play well with batched-fsync. Even \nthrough we don't know the final OID, we should still call \nprepare_loose_object_bulk_checkin to potentially create the bulk checkin \nobjdir.\n\n\n> diff --git a/object-file.c b/object-file.c\n> index 7946fa5e088..9fd449693c4 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -2119,6 +2119,106 @@ static int freshen_packed_object(const struct object_id *oid)\n>   \treturn 1;\n>   }\n>   \n> +int stream_loose_object(struct input_stream *in_stream, size_t len,\n> +\t\t\tstruct object_id *oid)\n> +{\n> +\tint fd, ret, err = 0, flush = 0;\n> +\tunsigned char compressed[4096];\n> +\tgit_zstream stream;\n> +\tgit_hash_ctx c;\n> +\tstruct strbuf tmp_file = STRBUF_INIT;\n> +\tstruct strbuf filename = STRBUF_INIT;\n> +\tint dirlen;\n> +\tchar hdr[MAX_HEADER_LEN];\n> +\tint hdrlen;\n> +\n> +\t/* Since oid is not determined, save tmp file to odb path. */\n> +\tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n> +\thdrlen = format_object_header(hdr, sizeof(hdr), OBJ_BLOB, len);\n> +\n> +\t/*\n> +\t * Common steps for write_loose_object and stream_loose_object to\n> +\t * start writing loose objects:\n> +\t *\n> +\t *  - Create tmpfile for the loose object.\n> +\t *  - Setup zlib stream for compression.\n> +\t *  - Start to feed header to zlib stream.\n> +\t */\n> +\tfd = start_loose_object_common(&tmp_file, filename.buf, 0,\n> +\t\t\t\t       &stream, compressed, sizeof(compressed),\n> +\t\t\t\t       &c, hdr, hdrlen);\n> +\tif (fd < 0) {\n> +\t\terr = -1;\n> +\t\tgoto cleanup;\n> +\t}\n> +\n> +\t/* Then the data itself.. */\n> +\tdo {\n> +\t\tunsigned char *in0 = stream.next_in;\n> +\n> +\t\tif (!stream.avail_in && !in_stream->is_finished) {\n> +\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n> +\t\t\tstream.next_in = (void *)in;\n> +\t\t\tin0 = (unsigned char *)in;\n> +\t\t\t/* All data has been read. */\n> +\t\t\tif (in_stream->is_finished)\n> +\t\t\t\tflush = 1;\n> +\t\t}\n> +\t\tret = write_loose_object_common(&c, &stream, flush, in0, fd,\n> +\t\t\t\t\t\tcompressed, sizeof(compressed));\n> +\t\t/*\n> +\t\t * Unlike write_loose_object(), we do not have the entire\n> +\t\t * buffer. If we get Z_BUF_ERROR due to too few input bytes,\n> +\t\t * then we'll replenish them in the next input_stream->read()\n> +\t\t * call when we loop.\n> +\t\t */\n> +\t} while (ret == Z_OK || ret == Z_BUF_ERROR);\n> +\n> +\tif (stream.total_in != len + hdrlen)\n> +\t\tdie(_(\"write stream object %ld != %\"PRIuMAX), stream.total_in,\n> +\t\t    (uintmax_t)len + hdrlen);\n> +\n> +\t/* Common steps for write_loose_object and stream_loose_object to\n> +\t * end writing loose oject:\n> +\t *\n> +\t *  - End the compression of zlib stream.\n> +\t *  - Get the calculated oid.\n> +\t */\n> +\tif (ret != Z_STREAM_END)\n> +\t\tdie(_(\"unable to stream deflate new object (%d)\"), ret);\n> +\tret = end_loose_object_common(&c, &stream, oid);\n> +\tif (ret != Z_OK)\n> +\t\tdie(_(\"deflateEnd on stream object failed (%d)\"), ret);\n> +\tclose_loose_object(fd, tmp_file.buf);\n> +\n\nIf batch fsync is enabled, the close_loose_object call will refrain from \nsyncing the tmp file.\n\n> +\tif (freshen_packed_object(oid) || freshen_loose_object(oid)) {\n> +\t\tunlink_or_warn(tmp_file.buf);\n> +\t\tgoto cleanup;\n> +\t}\n> +\n> +\tloose_object_path(the_repository, &filename, oid);\n> +\n\nWe expect this loose_object_path call to return a path in the bulk fsync \nobject directory. It might not do so if we don't call \nprepare_loose_object_bulk_checkin.\n\nIn the new test case introduced in (7/7), we seem to be getting lucky\nin that there are some small objects (commits) earlier in the packfile,\nso we go through write_loose_object first.\n\nThanks for including me on the review!\n\n-Neeraj\n"},{"id":"456867","messageId":"xmqqh74vrwww.fsf@gitster.g","threadId":"56672","inReplyTo":"7ba4858a-d1cc-a4eb-b6d6-4c04a5dd6ce7@gmail.com","subject":"Re: [PATCH v13 5/7] object-file.c: add \"stream_loose_object()\" to handle large object","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-06-08T15:34:55Z","receivedAt":"2022-06-08T15:35:32Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neeraj Singh <nksingh85@gmail.com> writes:\n\n>> Still, we need to save the temporary file we're preparing\n>> somewhere. We'll do that in the top-level \".git/objects/\"\n>> directory (or whatever \"GIT_OBJECT_DIRECTORY\" is set to). Once we've\n>> streamed it we'll know the OID, and will move it to its canonical\n>> path.\n>> \n>\n> I think this new logic doesn't play well with batched-fsync. Even\n> through we don't know the final OID, we should still call \n> prepare_loose_object_bulk_checkin to potentially create the bulk\n> checkin objdir.\n\nGood point.  Careful sanity checks like this are very much\nappreciated.\n\n> Thanks for including me on the review!\n\nYes, indeed.\n\nThanks, both.\n"},{"id":"456913","messageId":"20220609030530.51746-1-chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"7ba4858a-d1cc-a4eb-b6d6-4c04a5dd6ce7@gmail.com","subject":"[RFC PATCH] object-file.c: batched disk flushes for stream_loose_object()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-09T03:05:30Z","receivedAt":"2022-06-09T03:05:42Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"Neeraj Singh[1] pointed out that if batch fsync is enabled, we should still\ncall prepare_loose_object_bulk_checkin() to potentially create the bulk checkin\nobjdir.\n\n1. https://lore.kernel.org/git/7ba4858a-d1cc-a4eb-b6d6-4c04a5dd6ce7@gmail.com/\n\nSigned-off-by: Han Xin <chiyutianyi@gmail.com>\n---\n object-file.c                   |  3 +++\n t/t5351-unpack-large-objects.sh | 15 ++++++++++++++-\n 2 files changed, 17 insertions(+), 1 deletion(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 2dd828b45b..3a1be74775 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2131,6 +2131,9 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen;\n \n+\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n+\t\tprepare_loose_object_bulk_checkin();\n+\n \t/* Since oid is not determined, save tmp file to odb path. */\n \tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n \thdrlen = format_object_header(hdr, sizeof(hdr), OBJ_BLOB, len);\ndiff --git a/t/t5351-unpack-large-objects.sh b/t/t5351-unpack-large-objects.sh\nindex 461ca060b2..a66a51f7df 100755\n--- a/t/t5351-unpack-large-objects.sh\n+++ b/t/t5351-unpack-large-objects.sh\n@@ -18,7 +18,10 @@ test_expect_success \"create large objects (1.5 MB) and PACK\" '\n \ttest_commit --append foo big-blob &&\n \ttest-tool genrandom bar 1500000 >big-blob &&\n \ttest_commit --append bar big-blob &&\n-\tPACK=$(echo HEAD | git pack-objects --revs pack)\n+\tPACK=$(echo HEAD | git pack-objects --revs pack) &&\n+\tgit verify-pack -v pack-$PACK.pack |\n+\t    grep -E \"commit|tree|blob\" |\n+\t\tsed -n -e \"s/^\\([0-9a-f]*\\).*/\\1/p\" >obj-list\n '\n \n test_expect_success 'set memory limitation to 1MB' '\n@@ -45,6 +48,16 @@ test_expect_success 'unpack big object in stream' '\n \ttest_dir_is_empty dest.git/objects/pack\n '\n \n+BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\n+test_expect_success 'unpack big object in stream (core.fsyncmethod=batch)' '\n+\tprepare_dest 1m &&\n+\tgit $BATCH_CONFIGURATION -C dest.git unpack-objects <pack-$PACK.pack &&\n+\ttest_dir_is_empty dest.git/objects/pack &&\n+\tgit -C dest.git cat-file --batch-check=\"%(objectname)\" <obj-list >current &&\n+\tcmp obj-list current\n+'\n+\n test_expect_success 'do not unpack existing large objects' '\n \tprepare_dest 1m &&\n \tgit -C dest.git index-pack --stdin <pack-$PACK.pack &&\n-- \n2.36.1\n\n"},{"id":"456914","messageId":"CAO0brD2s-i2Bp7r2n+TRLs2LckzM-i1-293rr=sgmC2TbLozow@mail.gmail.com","threadId":"56672","inReplyTo":"xmqqpmjl7i7y.fsf@gitster.g","subject":"Re: [PATCH v13 1/7] unpack-objects: low memory footprint for get_data() in dry_run mode","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-09T04:10:09Z","receivedAt":"2022-06-09T04:10:46Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Tue, Jun 7, 2022 at 2:35 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Ævar Arnfjörð Bjarmason  <avarab@gmail.com> writes:\n>\n> > From: Han Xin <hanxin.hx@alibaba-inc.com>\n> >\n> > As the name implies, \"get_data(size)\" will allocate and return a given\n> > amount of memory. Allocating memory for a large blob object may cause the\n> > system to run out of memory. Before preparing to replace calling of\n> > \"get_data()\" to unpack large blob objects in latter commits, refactor\n> > \"get_data()\" to reduce memory footprint for dry_run mode.\n> >\n> > Because in dry_run mode, \"get_data()\" is only used to check the\n> > integrity of data, and the returned buffer is not used at all, we can\n> > allocate a smaller buffer and reuse it as zstream output. Therefore,\n>\n> \"reuse\" -> \"use\"\n>\n> > in dry_run mode, \"get_data()\" will release the allocated buffer and\n> > return NULL instead of returning garbage data.\n>\n> It makes it sound as if we used to return garbage data, but I do not\n> think that is what happened in reality.  Perhaps rewrite the last\n> sentence like\n>\n>         Make the function return NULL in the dry-run mode, as no\n>         callers use the returned buffer.\n>\n> or something?\n>\n> The overall logic sounds quite sensible.\n>\n> > The \"find [...]objects/?? -type f | wc -l\" test idiom being used here\n> > is adapted from the same \"find\" use added to another test in\n> > d9545c7f465 (fast-import: implement unpack limit, 2016-04-25).\n>\n>\n> > +/*\n> > + * Decompress zstream from stdin and return specific size of data.\n>\n> \"specific size\"?  The caller specifies the size of data (because it\n> knows a-priori how many bytes the zstream should inflate to), so\n>\n>     Decompress zstream from the standard input into a newly\n>     allocated buffer of specified size and return the buffer.\n>\n> or something, perhaps.  In any case, it needs to say that the caller\n> is responsible for giving the \"right\" size.\n>\n> > + * The caller is responsible to free the returned buffer.\n> > + *\n> > + * But for dry_run mode, \"get_data()\" is only used to check the\n> > + * integrity of data, and the returned buffer is not used at all.\n> > + * Therefore, in dry_run mode, \"get_data()\" will release the small\n> > + * allocated buffer which is reused to hold temporary zstream output\n> > + * and return NULL instead of returning garbage data.\n> > + */\n> >  static void *get_data(unsigned long size)\n> >  {\n> >       git_zstream stream;\n> > -     void *buf = xmallocz(size);\n> > +     unsigned long bufsize = dry_run && size > 8192 ? 8192 : size;\n> > +     void *buf = xmallocz(bufsize);\n>\n> OK.\n>\n> >       memset(&stream, 0, sizeof(stream));\n> >\n> >       stream.next_out = buf;\n> > -     stream.avail_out = size;\n> > +     stream.avail_out = bufsize;\n> >       stream.next_in = fill(1);\n> >       stream.avail_in = len;\n> >       git_inflate_init(&stream);\n> > @@ -125,8 +136,15 @@ static void *get_data(unsigned long size)\n>\n> What's hidden in the pre-context is this bit:\n>\n>                 int ret = git_inflate(&stream, 0);\n>                 use(len - stream.avail_in);\n>                 if (stream.total_out == size && ret == Z_STREAM_END)\n>                         break;\n>                 if (ret != Z_OK) {\n>                         error(\"inflate returned %d\", ret);\n>                         FREE_AND_NULL(buf);\n>                         if (!recover)\n>                                 exit(1);\n>                         has_errors = 1;\n>                         break;\n>                 }\n>\n> and it is correct to use \"size\", not \"bufsize\", for this check.\n> Unless we receive exactly the caller-specified \"size\" bytes from the\n> inflated zstream with Z_STREAM_END, we want to detect an error and\n> bail out.\n>\n> I am not sure if this is not loosening the error checking in the\n> dry-run case, though.  In the original code, we set the avail_out\n> to the total expected size so\n>\n>  (1) if the caller gives too small a size, git_inflate() would stop\n>      at stream.total_out with ret that is not STREAM_END nor OK,\n>      bypassing the \"break\", and we catch the error.\n>\n>  (2) if the caller gives too large a size, git_inflate() would stop\n>      at the true size of inflated zstream, with STREAM_END and would\n>      not hit this \"break\", and we catch the error.\n>\n> With the new code, since we keep refreshing avail_out (see below),\n> git_inflate() does not even learn how many bytes we are _expecting_\n> to see.  Is the error checking in the loop, with the updated code,\n> catch the mismatch between expected and actual size (plausibly\n> caused by a corrupted zstream) the same way as we do in the\n> non dry-run code path?\n>\n\nUnlike the original implementation, if we get a corrupted zstream, we\nwon't break at Z_BUFFER_ERROR, maybe until we've read all the\ninput. I think it can still catch the mismatch between expected and\nactual size when \"fill(1)\" gets an EOF, if it's not too late.\n\nThanks.\n-Han Xin\n\n> >               }\n> >               stream.next_in = fill(1);\n> >               stream.avail_in = len;\n> > +             if (dry_run) {\n> > +                     /* reuse the buffer in dry_run mode */\n> > +                     stream.next_out = buf;\n> > +                     stream.avail_out = bufsize;\n> > +             }\n> >       }\n> >       git_inflate_end(&stream);\n> > +     if (dry_run)\n> > +             FREE_AND_NULL(buf);\n> >       return buf;\n> >  }\n"},{"id":"456918","messageId":"CAO0brD2bZ=QomcrbbSfPZ+8pgPmYr6=aM5WsqHWdcJQycJMi+g@mail.gmail.com","threadId":"56672","inReplyTo":"xmqqk09t5zmm.fsf@gitster.g","subject":"Re: [PATCH v13 5/7] object-file.c: add \"stream_loose_object()\" to handle large object","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-09T06:04:22Z","receivedAt":"2022-06-09T06:04:40Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Tue, Jun 7, 2022 at 4:03 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n> > I am very tempted to ask why we do not do this to _all_ loose object\n> > files.  Instead of running the machinery twice over the data (once to\n> > compute the object name, then to compute the contents and write out),\n> > if we can produce loose object files of any size with a single pass,\n> > wouldn't that be an overall win?\n>\n> There is a patch later in the series whose proposed log message has\n> benchmarks to show that it is slower in general.  It still is\n> curious where the slowness comes from and if it is something we can\n> tune, though.\n>\n\nCompared with getting the whole object buffer, stream_loose_object() uses\nlimited avail_in buffer and never fill new content until the whole\navail_in has been\ndeflated. It will generate small avail_in fragments due to the limited\navail_out,\nand I think it is precisely because these avail_in fragments generate additional\ngit_deflate() loops.\n\nIn \"unpack-objects\", we use a buffer size of 8192. Increasing the buffer\ncan alleviate this problem, but maybe it's not worth it?\n\n> Thanks.\n"},{"id":"456919","messageId":"CAO0brD3tq-18p8g3P7DX=L=2zzUJnZXN_CRisM4cFRDhpZpU-g@mail.gmail.com","threadId":"56672","inReplyTo":"xmqqy1y960hq.fsf@gitster.g","subject":"Re: [PATCH v13 5/7] object-file.c: add \"stream_loose_object()\" to handle large object","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-09T06:14:53Z","receivedAt":"2022-06-09T06:15:11Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Tue, Jun 7, 2022 at 3:44 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n>\n> > \"freshen_packed_object()\" or \"freshen_loose_object()\" will be called\n> > inside \"stream_loose_object()\" after obtaining the \"oid\".\n>\n> That much we can read from the patch text.  Saying just \"we do X\"\n> without explaining \"why we do so\" in the proposed log message leaves\n> readers more confused than otherwise.  Why is it worth pointing out\n> in the proposed log message?  Is the reason why we need to do so\n> involve something tricky?\n>\n\nYes, it really should be made clear why this is done here.\n\nThanks.\n-Han Xin\n\n> > +     ret = end_loose_object_common(&c, &stream, oid);\n> > +     if (ret != Z_OK)\n> > +             die(_(\"deflateEnd on stream object failed (%d)\"), ret);\n> > +     close_loose_object(fd, tmp_file.buf);\n> > +\n> > +     if (freshen_packed_object(oid) || freshen_loose_object(oid)) {\n> > +             unlink_or_warn(tmp_file.buf);\n> > +             goto cleanup;\n>\n> So, we were told to write an object, we wrote to a temporary file,\n> and we wanted to mark the object to be recent and found that there\n> indeed is already the object.  We remove the temporary and do not\n> leave the new copy of the object, and the value of err at this point\n> is 0 (success) which is what is returned from cleanup: label.\n>\n> Good.\n>\n"},{"id":"456923","messageId":"f4a193ab-8226-fbc9-2276-a4f5d5285843@gmail.com","threadId":"56672","inReplyTo":"20220609030530.51746-1-chiyutianyi@gmail.com","subject":"Re: [RFC PATCH] object-file.c: batched disk flushes for stream_loose_object()","fromName":"Neeraj Singh","fromEmail":"nksingh85@gmail.com","sentAt":"2022-06-09T07:35:08Z","receivedAt":"2022-06-09T07:35:16Z","isPatch":true,"sender":{"key":"nksingh85@gmail.com","avatar":null},"body":"On 6/8/2022 8:05 PM, Han Xin wrote:\n> Neeraj Singh[1] pointed out that if batch fsync is enabled, we should still\n> call prepare_loose_object_bulk_checkin() to potentially create the bulk checkin\n> objdir.\n> \n> 1. https://lore.kernel.org/git/7ba4858a-d1cc-a4eb-b6d6-4c04a5dd6ce7@gmail.com/\n> \n> Signed-off-by: Han Xin <chiyutianyi@gmail.com>\n> ---\n>   object-file.c                   |  3 +++\n>   t/t5351-unpack-large-objects.sh | 15 ++++++++++++++-\n>   2 files changed, 17 insertions(+), 1 deletion(-)\n> \n> diff --git a/object-file.c b/object-file.c\n> index 2dd828b45b..3a1be74775 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -2131,6 +2131,9 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n>   \tchar hdr[MAX_HEADER_LEN];\n>   \tint hdrlen;\n>   \n> +\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n> +\t\tprepare_loose_object_bulk_checkin();\n> +\n>   \t/* Since oid is not determined, save tmp file to odb path. */\n>   \tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n>   \thdrlen = format_object_header(hdr, sizeof(hdr), OBJ_BLOB, len);\n> diff --git a/t/t5351-unpack-large-objects.sh b/t/t5351-unpack-large-objects.sh\n> index 461ca060b2..a66a51f7df 100755\n> --- a/t/t5351-unpack-large-objects.sh\n> +++ b/t/t5351-unpack-large-objects.sh\n> @@ -18,7 +18,10 @@ test_expect_success \"create large objects (1.5 MB) and PACK\" '\n>   \ttest_commit --append foo big-blob &&\n>   \ttest-tool genrandom bar 1500000 >big-blob &&\n>   \ttest_commit --append bar big-blob &&\n> -\tPACK=$(echo HEAD | git pack-objects --revs pack)\n> +\tPACK=$(echo HEAD | git pack-objects --revs pack) &&\n> +\tgit verify-pack -v pack-$PACK.pack |\n> +\t    grep -E \"commit|tree|blob\" |\n> +\t\tsed -n -e \"s/^\\([0-9a-f]*\\).*/\\1/p\" >obj-list\n>   '\n>   \n>   test_expect_success 'set memory limitation to 1MB' '\n> @@ -45,6 +48,16 @@ test_expect_success 'unpack big object in stream' '\n>   \ttest_dir_is_empty dest.git/objects/pack\n>   '\n>   \n> +BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n> +\n> +test_expect_success 'unpack big object in stream (core.fsyncmethod=batch)' '\n> +\tprepare_dest 1m &&\n> +\tgit $BATCH_CONFIGURATION -C dest.git unpack-objects <pack-$PACK.pack &&\n> +\ttest_dir_is_empty dest.git/objects/pack &&\n> +\tgit -C dest.git cat-file --batch-check=\"%(objectname)\" <obj-list >current &&\n> +\tcmp obj-list current\n> +'\n> +\n>   test_expect_success 'do not unpack existing large objects' '\n>   \tprepare_dest 1m &&\n>   \tgit -C dest.git index-pack --stdin <pack-$PACK.pack &&\n\nThis fix looks good to me.\n\nThanks.\n\n-Neeraj\n"},{"id":"456929","messageId":"nycvar.QRO.7.76.6.2206091114070.349@tvgsbejvaqbjf.bet","threadId":"56672","inReplyTo":"20220609030530.51746-1-chiyutianyi@gmail.com","subject":"Re: [RFC PATCH] object-file.c: batched disk flushes for stream_loose_object()","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2022-06-09T09:30:00Z","receivedAt":"2022-06-09T09:31:25Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 9 Jun 2022, Han Xin wrote:\n\n> Neeraj Singh[1] pointed out that if batch fsync is enabled, we should still\n> call prepare_loose_object_bulk_checkin() to potentially create the bulk checkin\n> objdir.\n>\n> 1. https://lore.kernel.org/git/7ba4858a-d1cc-a4eb-b6d6-4c04a5dd6ce7@gmail.com/\n>\n> Signed-off-by: Han Xin <chiyutianyi@gmail.com>\n\nI like a good commit message that is concise and yet has all the necessary\ninformation. Well done!\n\n> ---\n>  object-file.c                   |  3 +++\n>  t/t5351-unpack-large-objects.sh | 15 ++++++++++++++-\n>  2 files changed, 17 insertions(+), 1 deletion(-)\n>\n> diff --git a/object-file.c b/object-file.c\n> index 2dd828b45b..3a1be74775 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -2131,6 +2131,9 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n>  \tchar hdr[MAX_HEADER_LEN];\n>  \tint hdrlen;\n>\n> +\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n> +\t\tprepare_loose_object_bulk_checkin();\n> +\n\nMakes sense.\n\n>  \t/* Since oid is not determined, save tmp file to odb path. */\n>  \tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n>  \thdrlen = format_object_header(hdr, sizeof(hdr), OBJ_BLOB, len);\n> diff --git a/t/t5351-unpack-large-objects.sh b/t/t5351-unpack-large-objects.sh\n> index 461ca060b2..a66a51f7df 100755\n> --- a/t/t5351-unpack-large-objects.sh\n> +++ b/t/t5351-unpack-large-objects.sh\n> @@ -18,7 +18,10 @@ test_expect_success \"create large objects (1.5 MB) and PACK\" '\n>  \ttest_commit --append foo big-blob &&\n>  \ttest-tool genrandom bar 1500000 >big-blob &&\n>  \ttest_commit --append bar big-blob &&\n> -\tPACK=$(echo HEAD | git pack-objects --revs pack)\n> +\tPACK=$(echo HEAD | git pack-objects --revs pack) &&\n> +\tgit verify-pack -v pack-$PACK.pack |\n> +\t    grep -E \"commit|tree|blob\" |\n> +\t\tsed -n -e \"s/^\\([0-9a-f]*\\).*/\\1/p\" >obj-list\n\nHere, I would recommend avoiding the pipe, to ensure that we would catch\nproblems in the `verify-pack` invocation, and I think we can avoid the\n`grep` altogether:\n\n\tgit verify-pack -v pack-$PACK.pack >out &&\n\tsed -n 's/^\\([0-9a-f][0-9a-f]*\\).*\\(commit\\|tree\\|blob\\)/\\1/p' \\\n\t\t<out >obj-list\n\n>  '\n>\n>  test_expect_success 'set memory limitation to 1MB' '\n> @@ -45,6 +48,16 @@ test_expect_success 'unpack big object in stream' '\n>  \ttest_dir_is_empty dest.git/objects/pack\n>  '\n>\n> +BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n> +\n> +test_expect_success 'unpack big object in stream (core.fsyncmethod=batch)' '\n> +\tprepare_dest 1m &&\n> +\tgit $BATCH_CONFIGURATION -C dest.git unpack-objects <pack-$PACK.pack &&\n\nI think the canonical way would be to use `test_config core.fsync ...`,\nbut the presented way works, too.\n\n> +\ttest_dir_is_empty dest.git/objects/pack &&\n> +\tgit -C dest.git cat-file --batch-check=\"%(objectname)\" <obj-list >current &&\n\nGood. The `--batch-check=\"%(objectname)\"` part forces `cat-file` to read\nthe actual object.\n\n> +\tcmp obj-list current\n> +'\n\nMy main question about this test case is whether it _actually_ verifies\nthat the batch-mode `fsync()`ing took place.\n\nI kind of had expected to see Trace2 enabled and a `grep` for\n`fsync/hardware-flush`. Do you think that would still make sense to add?\n\nThank you for working on the `fsync()` aspects of Git!\nDscho\n\n> +\n>  test_expect_success 'do not unpack existing large objects' '\n>  \tprepare_dest 1m &&\n>  \tgit -C dest.git index-pack --stdin <pack-$PACK.pack &&\n> --\n> 2.36.1\n>\n>\n"},{"id":"456957","messageId":"xmqqa6allmjl.fsf@gitster.g","threadId":"56672","inReplyTo":"CAO0brD2s-i2Bp7r2n+TRLs2LckzM-i1-293rr=sgmC2TbLozow@mail.gmail.com","subject":"Re: [PATCH v13 1/7] unpack-objects: low memory footprint for get_data() in dry_run mode","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-06-09T18:27:42Z","receivedAt":"2022-06-09T18:27:56Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Han Xin <chiyutianyi@gmail.com> writes:\n\n>> I am not sure if this is not loosening the error checking in the\n>> dry-run case, though.  In the original code, we set the avail_out\n>> to the total expected size so\n>>\n>>  (1) if the caller gives too small a size, git_inflate() would stop\n>>      at stream.total_out with ret that is not STREAM_END nor OK,\n>>      bypassing the \"break\", and we catch the error.\n>>\n>>  (2) if the caller gives too large a size, git_inflate() would stop\n>>      at the true size of inflated zstream, with STREAM_END and would\n>>      not hit this \"break\", and we catch the error.\n>>\n>> With the new code, since we keep refreshing avail_out (see below),\n>> git_inflate() does not even learn how many bytes we are _expecting_\n>> to see.  Is the error checking in the loop, with the updated code,\n>> catch the mismatch between expected and actual size (plausibly\n>> caused by a corrupted zstream) the same way as we do in the\n>> non dry-run code path?\n>>\n>\n> Unlike the original implementation, if we get a corrupted zstream, we\n> won't break at Z_BUFFER_ERROR, maybe until we've read all the\n> input. I think it can still catch the mismatch between expected and\n> actual size when \"fill(1)\" gets an EOF, if it's not too late.\n\nThat is only one half of the two possible failure cases, i.e. input\nis shorter than the expected size.  If the caller specified size is\nsmaller than what the stream inflates to, I do not see the new code\nto be limiting the .avail_out near the end of the iteration, which\nwould be necessary to catch such an error, even if we are not\ninterested in using the inflated contents, no?\n\n"},{"id":"456983","messageId":"CAO0brD3tU+v_5dS9En_fTpEyYVmgEMVb7iVGsPk6iuYNGdpYBg@mail.gmail.com","threadId":"56672","inReplyTo":"xmqqa6allmjl.fsf@gitster.g","subject":"Re: [PATCH v13 1/7] unpack-objects: low memory footprint for get_data() in dry_run mode","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-10T01:50:35Z","receivedAt":"2022-06-10T01:50:54Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Fri, Jun 10, 2022 at 2:27 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Han Xin <chiyutianyi@gmail.com> writes:\n>\n> >> I am not sure if this is not loosening the error checking in the\n> >> dry-run case, though.  In the original code, we set the avail_out\n> >> to the total expected size so\n> >>\n> >>  (1) if the caller gives too small a size, git_inflate() would stop\n> >>      at stream.total_out with ret that is not STREAM_END nor OK,\n> >>      bypassing the \"break\", and we catch the error.\n> >>\n> >>  (2) if the caller gives too large a size, git_inflate() would stop\n> >>      at the true size of inflated zstream, with STREAM_END and would\n> >>      not hit this \"break\", and we catch the error.\n> >>\n> >> With the new code, since we keep refreshing avail_out (see below),\n> >> git_inflate() does not even learn how many bytes we are _expecting_\n> >> to see.  Is the error checking in the loop, with the updated code,\n> >> catch the mismatch between expected and actual size (plausibly\n> >> caused by a corrupted zstream) the same way as we do in the\n> >> non dry-run code path?\n> >>\n> >\n> > Unlike the original implementation, if we get a corrupted zstream, we\n> > won't break at Z_BUFFER_ERROR, maybe until we've read all the\n> > input. I think it can still catch the mismatch between expected and\n> > actual size when \"fill(1)\" gets an EOF, if it's not too late.\n>\n> That is only one half of the two possible failure cases, i.e. input\n> is shorter than the expected size.  If the caller specified size is\n> smaller than what the stream inflates to, I do not see the new code\n> to be limiting the .avail_out near the end of the iteration, which\n> would be necessary to catch such an error, even if we are not\n> interested in using the inflated contents, no?\n>\n\nYes, you are right.\n\nInstead of always using a fixed \"bufsize\" even if there is not enough\nexpected output remaining, we can get a more accurate one by comparing\n\"total_out\" to \"size\", so we can catch problems early by getting\nZ_BUFFER_ERROR.\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 64abba8dba..5d59144883 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -139,7 +139,8 @@ static void *get_data(unsigned long size)\n                if (dry_run) {\n                        /* reuse the buffer in dry_run mode */\n                        stream.next_out = buf;\n-                       stream.avail_out = bufsize;\n+                       stream.avail_out = bufsize > size - stream.total_out ?\n+                               size - stream.total_out : bufsize;\n                }\n        }\n        git_inflate_end(&stream);\n\nThanks\n-Han Xin\n"},{"id":"457005","messageId":"220610.86mtels249.gmgdl@evledraar.gmail.com","threadId":"56672","inReplyTo":"CAO0brD3tU+v_5dS9En_fTpEyYVmgEMVb7iVGsPk6iuYNGdpYBg@mail.gmail.com","subject":"Re: [PATCH v13 1/7] unpack-objects: low memory footprint for get_data() in dry_run mode","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-06-10T02:05:39Z","receivedAt":"2022-06-10T02:07:08Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, Jun 10 2022, Han Xin wrote:\n\n> On Fri, Jun 10, 2022 at 2:27 AM Junio C Hamano <gitster@pobox.com> wrote:\n>>\n>> Han Xin <chiyutianyi@gmail.com> writes:\n>>\n>> >> I am not sure if this is not loosening the error checking in the\n>> >> dry-run case, though.  In the original code, we set the avail_out\n>> >> to the total expected size so\n>> >>\n>> >>  (1) if the caller gives too small a size, git_inflate() would stop\n>> >>      at stream.total_out with ret that is not STREAM_END nor OK,\n>> >>      bypassing the \"break\", and we catch the error.\n>> >>\n>> >>  (2) if the caller gives too large a size, git_inflate() would stop\n>> >>      at the true size of inflated zstream, with STREAM_END and would\n>> >>      not hit this \"break\", and we catch the error.\n>> >>\n>> >> With the new code, since we keep refreshing avail_out (see below),\n>> >> git_inflate() does not even learn how many bytes we are _expecting_\n>> >> to see.  Is the error checking in the loop, with the updated code,\n>> >> catch the mismatch between expected and actual size (plausibly\n>> >> caused by a corrupted zstream) the same way as we do in the\n>> >> non dry-run code path?\n>> >>\n>> >\n>> > Unlike the original implementation, if we get a corrupted zstream, we\n>> > won't break at Z_BUFFER_ERROR, maybe until we've read all the\n>> > input. I think it can still catch the mismatch between expected and\n>> > actual size when \"fill(1)\" gets an EOF, if it's not too late.\n>>\n>> That is only one half of the two possible failure cases, i.e. input\n>> is shorter than the expected size.  If the caller specified size is\n>> smaller than what the stream inflates to, I do not see the new code\n>> to be limiting the .avail_out near the end of the iteration, which\n>> would be necessary to catch such an error, even if we are not\n>> interested in using the inflated contents, no?\n>>\n>\n> Yes, you are right.\n>\n> Instead of always using a fixed \"bufsize\" even if there is not enough\n> expected output remaining, we can get a more accurate one by comparing\n> \"total_out\" to \"size\", so we can catch problems early by getting\n> Z_BUFFER_ERROR.\n>\n> diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n> index 64abba8dba..5d59144883 100644\n> --- a/builtin/unpack-objects.c\n> +++ b/builtin/unpack-objects.c\n> @@ -139,7 +139,8 @@ static void *get_data(unsigned long size)\n>                 if (dry_run) {\n>                         /* reuse the buffer in dry_run mode */\n>                         stream.next_out = buf;\n> -                       stream.avail_out = bufsize;\n> +                       stream.avail_out = bufsize > size - stream.total_out ?\n> +                               size - stream.total_out : bufsize;\n>                 }\n>         }\n>         git_inflate_end(&stream);\n>\n> Thanks\n> -Han Xin\n\nHan, do you want to pick this up again for a v14? It looks like you're\nvery on top of it already, and I re-sent your patches because I saw that\nyour\nhttps://lore.kernel.org/git/cover.1653015534.git.chiyutianyi@gmail.com/\nwasn't picked up in the interim & you hadn't been active on-list\notherwise.\n\nBut it looks like there's some interest now, and that you have more time\nto test & follow-up on this topic than I do at the moment, so if you\nwanted to do the work of properly rebasing ot in tho recent fsync\nchanges that would be great. Thanks.\n"},{"id":"457016","messageId":"CAO0brD0_CidpW8vKD2xbW=tXmZS-1Nup57Daz6DhcSN=WwPdEg@mail.gmail.com","threadId":"56672","inReplyTo":"220610.86mtels249.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v13 1/7] unpack-objects: low memory footprint for get_data() in dry_run mode","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-10T12:04:34Z","receivedAt":"2022-06-10T12:04:51Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Fri, Jun 10, 2022 at 10:07 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Fri, Jun 10 2022, Han Xin wrote:\n>\n> > On Fri, Jun 10, 2022 at 2:27 AM Junio C Hamano <gitster@pobox.com> wrote:\n> >>\n> >> Han Xin <chiyutianyi@gmail.com> writes:\n> >>\n> >> >> I am not sure if this is not loosening the error checking in the\n> >> >> dry-run case, though.  In the original code, we set the avail_out\n> >> >> to the total expected size so\n> >> >>\n> >> >>  (1) if the caller gives too small a size, git_inflate() would stop\n> >> >>      at stream.total_out with ret that is not STREAM_END nor OK,\n> >> >>      bypassing the \"break\", and we catch the error.\n> >> >>\n> >> >>  (2) if the caller gives too large a size, git_inflate() would stop\n> >> >>      at the true size of inflated zstream, with STREAM_END and would\n> >> >>      not hit this \"break\", and we catch the error.\n> >> >>\n> >> >> With the new code, since we keep refreshing avail_out (see below),\n> >> >> git_inflate() does not even learn how many bytes we are _expecting_\n> >> >> to see.  Is the error checking in the loop, with the updated code,\n> >> >> catch the mismatch between expected and actual size (plausibly\n> >> >> caused by a corrupted zstream) the same way as we do in the\n> >> >> non dry-run code path?\n> >> >>\n> >> >\n> >> > Unlike the original implementation, if we get a corrupted zstream, we\n> >> > won't break at Z_BUFFER_ERROR, maybe until we've read all the\n> >> > input. I think it can still catch the mismatch between expected and\n> >> > actual size when \"fill(1)\" gets an EOF, if it's not too late.\n> >>\n> >> That is only one half of the two possible failure cases, i.e. input\n> >> is shorter than the expected size.  If the caller specified size is\n> >> smaller than what the stream inflates to, I do not see the new code\n> >> to be limiting the .avail_out near the end of the iteration, which\n> >> would be necessary to catch such an error, even if we are not\n> >> interested in using the inflated contents, no?\n> >>\n> >\n> > Yes, you are right.\n> >\n> > Instead of always using a fixed \"bufsize\" even if there is not enough\n> > expected output remaining, we can get a more accurate one by comparing\n> > \"total_out\" to \"size\", so we can catch problems early by getting\n> > Z_BUFFER_ERROR.\n> >\n> > diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n> > index 64abba8dba..5d59144883 100644\n> > --- a/builtin/unpack-objects.c\n> > +++ b/builtin/unpack-objects.c\n> > @@ -139,7 +139,8 @@ static void *get_data(unsigned long size)\n> >                 if (dry_run) {\n> >                         /* reuse the buffer in dry_run mode */\n> >                         stream.next_out = buf;\n> > -                       stream.avail_out = bufsize;\n> > +                       stream.avail_out = bufsize > size - stream.total_out ?\n> > +                               size - stream.total_out : bufsize;\n> >                 }\n> >         }\n> >         git_inflate_end(&stream);\n> >\n> > Thanks\n> > -Han Xin\n>\n> Han, do you want to pick this up again for a v14? It looks like you're\n> very on top of it already, and I re-sent your patches because I saw that\n> your\n> https://lore.kernel.org/git/cover.1653015534.git.chiyutianyi@gmail.com/\n> wasn't picked up in the interim & you hadn't been active on-list\n> otherwise.\n>\n> But it looks like there's some interest now, and that you have more time\n> to test & follow-up on this topic than I do at the moment, so if you\n> wanted to do the work of properly rebasing ot in tho recent fsync\n> changes that would be great. Thanks.\n\nOK, I am glad to do that.\n\nThank you very much.\n\n-Han Xin\n"},{"id":"457017","messageId":"CAO0brD3co2FB+UdhNDShT_uwWhTxuKeW9TuX6T_5yZ4CQkFjPw@mail.gmail.com","threadId":"56672","inReplyTo":"nycvar.QRO.7.76.6.2206091114070.349@tvgsbejvaqbjf.bet","subject":"Re: [RFC PATCH] object-file.c: batched disk flushes for stream_loose_object()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-10T12:55:39Z","receivedAt":"2022-06-10T12:55:57Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Thu, Jun 9, 2022 at 5:30 PM Johannes Schindelin\n<Johannes.Schindelin@gmx.de> wrote:\n>\n> Hi,\n>\n> On Thu, 9 Jun 2022, Han Xin wrote:\n>\n> > Neeraj Singh[1] pointed out that if batch fsync is enabled, we should still\n> > call prepare_loose_object_bulk_checkin() to potentially create the bulk checkin\n> > objdir.\n> >\n> > 1. https://lore.kernel.org/git/7ba4858a-d1cc-a4eb-b6d6-4c04a5dd6ce7@gmail.com/\n> >\n> > Signed-off-by: Han Xin <chiyutianyi@gmail.com>\n>\n> I like a good commit message that is concise and yet has all the necessary\n> information. Well done!\n>\n> > ---\n> >  object-file.c                   |  3 +++\n> >  t/t5351-unpack-large-objects.sh | 15 ++++++++++++++-\n> >  2 files changed, 17 insertions(+), 1 deletion(-)\n> >\n> > diff --git a/object-file.c b/object-file.c\n> > index 2dd828b45b..3a1be74775 100644\n> > --- a/object-file.c\n> > +++ b/object-file.c\n> > @@ -2131,6 +2131,9 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n> >       char hdr[MAX_HEADER_LEN];\n> >       int hdrlen;\n> >\n> > +     if (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n> > +             prepare_loose_object_bulk_checkin();\n> > +\n>\n> Makes sense.\n>\n> >       /* Since oid is not determined, save tmp file to odb path. */\n> >       strbuf_addf(&filename, \"%s/\", get_object_directory());\n> >       hdrlen = format_object_header(hdr, sizeof(hdr), OBJ_BLOB, len);\n> > diff --git a/t/t5351-unpack-large-objects.sh b/t/t5351-unpack-large-objects.sh\n> > index 461ca060b2..a66a51f7df 100755\n> > --- a/t/t5351-unpack-large-objects.sh\n> > +++ b/t/t5351-unpack-large-objects.sh\n> > @@ -18,7 +18,10 @@ test_expect_success \"create large objects (1.5 MB) and PACK\" '\n> >       test_commit --append foo big-blob &&\n> >       test-tool genrandom bar 1500000 >big-blob &&\n> >       test_commit --append bar big-blob &&\n> > -     PACK=$(echo HEAD | git pack-objects --revs pack)\n> > +     PACK=$(echo HEAD | git pack-objects --revs pack) &&\n> > +     git verify-pack -v pack-$PACK.pack |\n> > +         grep -E \"commit|tree|blob\" |\n> > +             sed -n -e \"s/^\\([0-9a-f]*\\).*/\\1/p\" >obj-list\n>\n> Here, I would recommend avoiding the pipe, to ensure that we would catch\n> problems in the `verify-pack` invocation, and I think we can avoid the\n> `grep` altogether:\n>\n>         git verify-pack -v pack-$PACK.pack >out &&\n>         sed -n 's/^\\([0-9a-f][0-9a-f]*\\).*\\(commit\\|tree\\|blob\\)/\\1/p' \\\n>                 <out >obj-list\n>\n\nGood suggestion. I will take it.\n\nThanks.\n-Han Xin\n\n> >  '\n> >\n> >  test_expect_success 'set memory limitation to 1MB' '\n> > @@ -45,6 +48,16 @@ test_expect_success 'unpack big object in stream' '\n> >       test_dir_is_empty dest.git/objects/pack\n> >  '\n> >\n> > +BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n> > +\n> > +test_expect_success 'unpack big object in stream (core.fsyncmethod=batch)' '\n> > +     prepare_dest 1m &&\n> > +     git $BATCH_CONFIGURATION -C dest.git unpack-objects <pack-$PACK.pack &&\n>\n> I think the canonical way would be to use `test_config core.fsync ...`,\n> but the presented way works, too.\n>\n> > +     test_dir_is_empty dest.git/objects/pack &&\n> > +     git -C dest.git cat-file --batch-check=\"%(objectname)\" <obj-list >current &&\n>\n> Good. The `--batch-check=\"%(objectname)\"` part forces `cat-file` to read\n> the actual object.\n>\n> > +     cmp obj-list current\n> > +'\n>\n> My main question about this test case is whether it _actually_ verifies\n> that the batch-mode `fsync()`ing took place.\n>\n> I kind of had expected to see Trace2 enabled and a `grep` for\n> `fsync/hardware-flush`. Do you think that would still make sense to add?\n>\n> Thank you for working on the `fsync()` aspects of Git!\n> Dscho\n>\n\nMore rigorous inspection should be adopted.\n\nThanks.\n-Han Xin\n\n> > +\n> >  test_expect_success 'do not unpack existing large objects' '\n> >       prepare_dest 1m &&\n> >       git -C dest.git index-pack --stdin <pack-$PACK.pack &&\n> > --\n> > 2.36.1\n> >\n> >\n"},{"id":"457020","messageId":"cover.1654871915.git.chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"cover-v13-0.7-00000000000-20220604T095113Z-avarab@gmail.com","subject":"[PATCH v14 0/7] unpack-objects: support streaming blobs to disk","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-10T14:46:00Z","receivedAt":"2022-06-10T14:47:30Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"This series makes \"unpack-objects\" capable of streaming large objects\nto disk.\n\nAs 7/7 shows streaming e.g. a 100MB blob now uses ~5MB of memory\ninstead of ~105MB. This streaming method is slower if you've got\nmemory to handle the blobs in-core, but if you don't it allows you to\nunpack objects at all, as you might otherwise OOM.\n\nChanges since v13:\n\n* Make the error checking in the loop of get_data() the same way as\n  we do in the non dry-run mode.\n\n* Add batched disk flushes for stream_loose_object(). This is pointed\n  out by Neeraj Singh[1].\n\n* Minor typo/grammar/comment etc. fixes throughout.\n\n1. https://lore.kernel.org/git/7ba4858a-d1cc-a4eb-b6d6-4c04a5dd6ce7@gmail.com/\n\nHan Xin (4):\n  unpack-objects: low memory footprint for get_data() in dry_run mode\n  object-file.c: refactor write_loose_object() to several steps\n  object-file.c: add \"stream_loose_object()\" to handle large object\n  unpack-objects: use stream_loose_object() to unpack large objects\n\nÆvar Arnfjörð Bjarmason (3):\n  object-file.c: do fsync() and close() before post-write die()\n  object-file.c: factor out deflate part of write_loose_object()\n  core doc: modernize core.bigFileThreshold documentation\n\n Documentation/config/core.txt   |  33 +++--\n builtin/unpack-objects.c        | 106 ++++++++++++--\n object-file.c                   | 240 +++++++++++++++++++++++++++-----\n object-store.h                  |   8 ++\n t/t5351-unpack-large-objects.sh |  76 ++++++++++\n 5 files changed, 408 insertions(+), 55 deletions(-)\n create mode 100755 t/t5351-unpack-large-objects.sh\n\nRange-diff against v13:\n1:  6703df6350 ! 1:  bf600a2fa8 unpack-objects: low memory footprint for get_data() in dry_run mode\n    @@ Commit message\n     \n         Because in dry_run mode, \"get_data()\" is only used to check the\n         integrity of data, and the returned buffer is not used at all, we can\n    -    allocate a smaller buffer and reuse it as zstream output. Therefore,\n    -    in dry_run mode, \"get_data()\" will release the allocated buffer and\n    -    return NULL instead of returning garbage data.\n    +    allocate a smaller buffer and use it as zstream output. Make the function\n    +    return NULL in the dry-run mode, as no callers use the returned buffer.\n     \n         The \"find [...]objects/?? -type f | wc -l\" test idiom being used here\n         is adapted from the same \"find\" use added to another test in\n    @@ builtin/unpack-objects.c: static void use(int bytes)\n      }\n      \n     +/*\n    -+ * Decompress zstream from stdin and return specific size of data.\n    ++ * Decompress zstream from the standard input into a newly\n    ++ * allocated buffer of specified size and return the buffer.\n     + * The caller is responsible to free the returned buffer.\n     + *\n     + * But for dry_run mode, \"get_data()\" is only used to check the\n    @@ builtin/unpack-objects.c: static void *get_data(unsigned long size)\n     +\t\tif (dry_run) {\n     +\t\t\t/* reuse the buffer in dry_run mode */\n     +\t\t\tstream.next_out = buf;\n    -+\t\t\tstream.avail_out = bufsize;\n    ++\t\t\tstream.avail_out = bufsize > size - stream.total_out ?\n    ++\t\t\t\t\t\t   size - stream.total_out :\n    ++\t\t\t\t\t\t   bufsize;\n     +\t\t}\n      \t}\n      \tgit_inflate_end(&stream);\n2:  6e289d25c1 = 2:  a327f484f7 object-file.c: do fsync() and close() before post-write die()\n3:  46f9def06c ! 3:  9bc8002282 object-file.c: refactor write_loose_object() to several steps\n    @@ object-file.c: static int create_tmpfile(struct strbuf *tmp, const char *filenam\n     + *\n     + * - End the compression of zlib stream.\n     + * - Get the calculated oid to \"oid\".\n    -+ * - fsync() and close() the \"fd\"\n     + */\n     +static int end_loose_object_common(git_hash_ctx *c, git_zstream *stream,\n     +\t\t\t\t   struct object_id *oid)\n4:  5a95ebede6 = 4:  7c73815f18 object-file.c: factor out deflate part of write_loose_object()\n5:  26847541aa ! 5:  28a9588f9c object-file.c: add \"stream_loose_object()\" to handle large object\n    @@ Commit message\n         path.\n     \n         \"freshen_packed_object()\" or \"freshen_loose_object()\" will be called\n    -    inside \"stream_loose_object()\" after obtaining the \"oid\".\n    +    inside \"stream_loose_object()\" after obtaining the \"oid\". After the\n    +    temporary file is written, we wants to mark the object to recent and we\n    +    may find that where indeed is already the object. We should remove the\n    +    temporary and do not leave a new copy of the object.\n     \n         Helped-by: René Scharfe <l.s.r@web.de>\n         Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n    @@ object-file.c: static int freshen_packed_object(const struct object_id *oid)\n     +\tchar hdr[MAX_HEADER_LEN];\n     +\tint hdrlen;\n     +\n    ++\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n    ++\t\tprepare_loose_object_bulk_checkin();\n    ++\n     +\t/* Since oid is not determined, save tmp file to odb path. */\n     +\tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n     +\thdrlen = format_object_header(hdr, sizeof(hdr), OBJ_BLOB, len);\n    @@ object-file.c: static int freshen_packed_object(const struct object_id *oid)\n     +\t\tdie(_(\"write stream object %ld != %\"PRIuMAX), stream.total_in,\n     +\t\t    (uintmax_t)len + hdrlen);\n     +\n    -+\t/* Common steps for write_loose_object and stream_loose_object to\n    ++\t/*\n    ++\t * Common steps for write_loose_object and stream_loose_object to\n     +\t * end writing loose oject:\n     +\t *\n     +\t *  - End the compression of zlib stream.\n6:  eb962b60b9 ! 6:  dea5c4172b core doc: modernize core.bigFileThreshold documentation\n    @@ Documentation/config/core.txt: You probably do not need to adjust this value.\n     +Files above the configured limit will be:\n      +\n     -Common unit suffixes of 'k', 'm', or 'g' are supported.\n    -+* Stored deflated, without attempting delta compression.\n    ++* Stored deflated in packfiles, without attempting delta compression.\n     ++\n     +The default limit is primarily set with this use-case in mind. With it\n     +most projects will have their source code and other text files delta\n    @@ Documentation/config/core.txt: You probably do not need to adjust this value.\n     +usage, at the slight expense of increased disk usage.\n     ++\n     +* Will be treated as if though they were labeled \"binary\" (see\n    -+  linkgit:gitattributes[5]). This means that e.g. linkgit:git-log[1]\n    -+  and linkgit:git-diff[1] will not diffs for files above this limit.\n    ++  linkgit:gitattributes[5]). e.g. linkgit:git-log[1] and\n    ++  linkgit:git-diff[1] will not diffs for files above this limit.\n     ++\n     +* Will be generally be streamed when written, which avoids excessive\n     +memory usage, at the cost of some fixed overhead. Commands that make\n7:  88a2754fcb ! 7:  d236230a4c unpack-objects: use stream_loose_object() to unpack large objects\n    @@ t/t5351-unpack-large-objects.sh: test_description='git unpack-objects with large\n      }\n      \n      test_expect_success \"create large objects (1.5 MB) and PACK\" '\n    +@@ t/t5351-unpack-large-objects.sh: test_expect_success \"create large objects (1.5 MB) and PACK\" '\n    + \ttest_commit --append foo big-blob &&\n    + \ttest-tool genrandom bar 1500000 >big-blob &&\n    + \ttest_commit --append bar big-blob &&\n    +-\tPACK=$(echo HEAD | git pack-objects --revs pack)\n    ++\tPACK=$(echo HEAD | git pack-objects --revs pack) &&\n    ++\tgit verify-pack -v pack-$PACK.pack >out &&\n    ++\tsed -n -e \"s/^\\([0-9a-f][0-9a-f]*\\).*\\(commit\\|tree\\|blob\\).*/\\1/p\" \\\n    ++\t\t<out >obj-list\n    + '\n    + \n    + test_expect_success 'set memory limitation to 1MB' '\n     @@ t/t5351-unpack-large-objects.sh: test_expect_success 'set memory limitation to 1MB' '\n      '\n      \n    @@ t/t5351-unpack-large-objects.sh: test_expect_success 'set memory limitation to 1\n     +\ttest_dir_is_empty dest.git/objects/pack\n     +'\n     +\n    ++BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n    ++\n    ++test_expect_success 'unpack big object in stream (core.fsyncmethod=batch)' '\n    ++\tprepare_dest 1m &&\n    ++\tGIT_TRACE2_EVENT=\"$(pwd)/trace2.txt\" \\\n    ++\t\tgit -C dest.git $BATCH_CONFIGURATION unpack-objects <pack-$PACK.pack &&\n    ++\tgrep fsync/hardware-flush trace2.txt &&\n    ++\ttest_dir_is_empty dest.git/objects/pack &&\n    ++\tgit -C dest.git cat-file --batch-check=\"%(objectname)\" <obj-list >current &&\n    ++\tcmp obj-list current\n    ++'\n    ++\n     +test_expect_success 'do not unpack existing large objects' '\n     +\tprepare_dest 1m &&\n     +\tgit -C dest.git index-pack --stdin <pack-$PACK.pack &&\n-- \n2.36.1\n\n"},{"id":"457021","messageId":"bf600a2fa85883193df22436be52ac8ca809aa33.1654871916.git.chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"cover.1654871915.git.chiyutianyi@gmail.com","subject":"[PATCH v14 1/7] unpack-objects: low memory footprint for get_data() in dry_run mode","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-10T14:46:01Z","receivedAt":"2022-06-10T14:47:44Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nAs the name implies, \"get_data(size)\" will allocate and return a given\namount of memory. Allocating memory for a large blob object may cause the\nsystem to run out of memory. Before preparing to replace calling of\n\"get_data()\" to unpack large blob objects in latter commits, refactor\n\"get_data()\" to reduce memory footprint for dry_run mode.\n\nBecause in dry_run mode, \"get_data()\" is only used to check the\nintegrity of data, and the returned buffer is not used at all, we can\nallocate a smaller buffer and use it as zstream output. Make the function\nreturn NULL in the dry-run mode, as no callers use the returned buffer.\n\nThe \"find [...]objects/?? -type f | wc -l\" test idiom being used here\nis adapted from the same \"find\" use added to another test in\nd9545c7f465 (fast-import: implement unpack limit, 2016-04-25).\n\nSuggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n builtin/unpack-objects.c        | 37 ++++++++++++++++++++---------\n t/t5351-unpack-large-objects.sh | 41 +++++++++++++++++++++++++++++++++\n 2 files changed, 67 insertions(+), 11 deletions(-)\n create mode 100755 t/t5351-unpack-large-objects.sh\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 56d05e2725..32e8b47059 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -97,15 +97,27 @@ static void use(int bytes)\n \tdisplay_throughput(progress, consumed_bytes);\n }\n \n+/*\n+ * Decompress zstream from the standard input into a newly\n+ * allocated buffer of specified size and return the buffer.\n+ * The caller is responsible to free the returned buffer.\n+ *\n+ * But for dry_run mode, \"get_data()\" is only used to check the\n+ * integrity of data, and the returned buffer is not used at all.\n+ * Therefore, in dry_run mode, \"get_data()\" will release the small\n+ * allocated buffer which is reused to hold temporary zstream output\n+ * and return NULL instead of returning garbage data.\n+ */\n static void *get_data(unsigned long size)\n {\n \tgit_zstream stream;\n-\tvoid *buf = xmallocz(size);\n+\tunsigned long bufsize = dry_run && size > 8192 ? 8192 : size;\n+\tvoid *buf = xmallocz(bufsize);\n \n \tmemset(&stream, 0, sizeof(stream));\n \n \tstream.next_out = buf;\n-\tstream.avail_out = size;\n+\tstream.avail_out = bufsize;\n \tstream.next_in = fill(1);\n \tstream.avail_in = len;\n \tgit_inflate_init(&stream);\n@@ -125,8 +137,17 @@ static void *get_data(unsigned long size)\n \t\t}\n \t\tstream.next_in = fill(1);\n \t\tstream.avail_in = len;\n+\t\tif (dry_run) {\n+\t\t\t/* reuse the buffer in dry_run mode */\n+\t\t\tstream.next_out = buf;\n+\t\t\tstream.avail_out = bufsize > size - stream.total_out ?\n+\t\t\t\t\t\t   size - stream.total_out :\n+\t\t\t\t\t\t   bufsize;\n+\t\t}\n \t}\n \tgit_inflate_end(&stream);\n+\tif (dry_run)\n+\t\tFREE_AND_NULL(buf);\n \treturn buf;\n }\n \n@@ -326,10 +347,8 @@ static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n {\n \tvoid *buf = get_data(size);\n \n-\tif (!dry_run && buf)\n+\tif (buf)\n \t\twrite_object(nr, type, buf, size);\n-\telse\n-\t\tfree(buf);\n }\n \n static int resolve_against_held(unsigned nr, const struct object_id *base,\n@@ -359,10 +378,8 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\toidread(&base_oid, fill(the_hash_algo->rawsz));\n \t\tuse(the_hash_algo->rawsz);\n \t\tdelta_data = get_data(delta_size);\n-\t\tif (dry_run || !delta_data) {\n-\t\t\tfree(delta_data);\n+\t\tif (!delta_data)\n \t\t\treturn;\n-\t\t}\n \t\tif (has_object_file(&base_oid))\n \t\t\t; /* Ok we have this one */\n \t\telse if (resolve_against_held(nr, &base_oid,\n@@ -398,10 +415,8 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\t\tdie(\"offset value out of bound for delta base object\");\n \n \t\tdelta_data = get_data(delta_size);\n-\t\tif (dry_run || !delta_data) {\n-\t\t\tfree(delta_data);\n+\t\tif (!delta_data)\n \t\t\treturn;\n-\t\t}\n \t\tlo = 0;\n \t\thi = nr;\n \t\twhile (lo < hi) {\ndiff --git a/t/t5351-unpack-large-objects.sh b/t/t5351-unpack-large-objects.sh\nnew file mode 100755\nindex 0000000000..8d84313221\n--- /dev/null\n+++ b/t/t5351-unpack-large-objects.sh\n@@ -0,0 +1,41 @@\n+#!/bin/sh\n+#\n+# Copyright (c) 2022 Han Xin\n+#\n+\n+test_description='git unpack-objects with large objects'\n+\n+. ./test-lib.sh\n+\n+prepare_dest () {\n+\ttest_when_finished \"rm -rf dest.git\" &&\n+\tgit init --bare dest.git\n+}\n+\n+test_expect_success \"create large objects (1.5 MB) and PACK\" '\n+\ttest-tool genrandom foo 1500000 >big-blob &&\n+\ttest_commit --append foo big-blob &&\n+\ttest-tool genrandom bar 1500000 >big-blob &&\n+\ttest_commit --append bar big-blob &&\n+\tPACK=$(echo HEAD | git pack-objects --revs pack)\n+'\n+\n+test_expect_success 'set memory limitation to 1MB' '\n+\tGIT_ALLOC_LIMIT=1m &&\n+\texport GIT_ALLOC_LIMIT\n+'\n+\n+test_expect_success 'unpack-objects failed under memory limitation' '\n+\tprepare_dest &&\n+\ttest_must_fail git -C dest.git unpack-objects <pack-$PACK.pack 2>err &&\n+\tgrep \"fatal: attempting to allocate\" err\n+'\n+\n+test_expect_success 'unpack-objects works with memory limitation in dry-run mode' '\n+\tprepare_dest &&\n+\tgit -C dest.git unpack-objects -n <pack-$PACK.pack &&\n+\ttest_stdout_line_count = 0 find dest.git/objects -type f &&\n+\ttest_dir_is_empty dest.git/objects/pack\n+'\n+\n+test_done\n-- \n2.36.1\n\n"},{"id":"457022","messageId":"a327f484f7f7466597930e87686e7156beabdc45.1654871916.git.chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"cover.1654871915.git.chiyutianyi@gmail.com","subject":"[PATCH v14 2/7] object-file.c: do fsync() and close() before post-write die()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-10T14:46:02Z","receivedAt":"2022-06-10T14:47:49Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n\nChange write_loose_object() to do an fsync() and close() before the\noideq() sanity check at the end. This change re-joins code that was\nsplit up by the die() sanity check added in 748af44c63e (sha1_file: be\nparanoid when creating loose objects, 2010-02-21).\n\nI don't think that this change matters in itself, if we called die()\nit was possible that our data wouldn't fully make it to disk, but in\nany case we were writing data that we'd consider corrupted. It's\npossible that a subsequent \"git fsck\" will be less confused now.\n\nThe real reason to make this change is that in a subsequent commit\nwe'll split this code in write_loose_object() into a utility function,\nall its callers will want the preceding sanity checks, but not the\n\"oideq\" check. By moving the close_loose_object() earlier it'll be\neasier to reason about the introduction of the utility function.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n object-file.c | 4 ++--\n 1 file changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 79eb8339b6..e4a83012ba 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2012,12 +2012,12 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n \t\t    ret);\n \tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n+\tclose_loose_object(fd, tmp_file.buf);\n+\n \tif (!oideq(oid, &parano_oid))\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n \n-\tclose_loose_object(fd, tmp_file.buf);\n-\n \tif (mtime) {\n \t\tstruct utimbuf utb;\n \t\tutb.actime = mtime;\n-- \n2.36.1\n\n"},{"id":"457023","messageId":"9bc8002282ddbb13b707c303281d88f377ecbdbe.1654871916.git.chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"cover.1654871915.git.chiyutianyi@gmail.com","subject":"[PATCH v14 3/7] object-file.c: refactor write_loose_object() to several steps","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-10T14:46:03Z","receivedAt":"2022-06-10T14:48:02Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWhen writing a large blob using \"write_loose_object()\", we have to pass\na buffer with the whole content of the blob, and this behavior will\nconsume lots of memory and may cause OOM. We will introduce a stream\nversion function (\"stream_loose_object()\") in later commit to resolve\nthis issue.\n\nBefore introducing that streaming function, do some refactoring on\n\"write_loose_object()\" to reuse code for both versions.\n\nRewrite \"write_loose_object()\" as follows:\n\n 1. Figure out a path for the (temp) object file. This step is only\n    used in \"write_loose_object()\".\n\n 2. Move common steps for starting to write loose objects into a new\n    function \"start_loose_object_common()\".\n\n 3. Compress data.\n\n 4. Move common steps for ending zlib stream into a new function\n    \"end_loose_object_common()\".\n\n 5. Close fd and finalize the object file.\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n object-file.c | 101 +++++++++++++++++++++++++++++++++++++-------------\n 1 file changed, 75 insertions(+), 26 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex e4a83012ba..f4d7f8c109 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1951,6 +1951,74 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n \treturn fd;\n }\n \n+/**\n+ * Common steps for loose object writers to start writing loose\n+ * objects:\n+ *\n+ * - Create tmpfile for the loose object.\n+ * - Setup zlib stream for compression.\n+ * - Start to feed header to zlib stream.\n+ *\n+ * Returns a \"fd\", which should later be provided to\n+ * end_loose_object_common().\n+ */\n+static int start_loose_object_common(struct strbuf *tmp_file,\n+\t\t\t\t     const char *filename, unsigned flags,\n+\t\t\t\t     git_zstream *stream,\n+\t\t\t\t     unsigned char *buf, size_t buflen,\n+\t\t\t\t     git_hash_ctx *c,\n+\t\t\t\t     char *hdr, int hdrlen)\n+{\n+\tint fd;\n+\n+\tfd = create_tmpfile(tmp_file, filename);\n+\tif (fd < 0) {\n+\t\tif (flags & HASH_SILENT)\n+\t\t\treturn -1;\n+\t\telse if (errno == EACCES)\n+\t\t\treturn error(_(\"insufficient permission for adding \"\n+\t\t\t\t       \"an object to repository database %s\"),\n+\t\t\t\t     get_object_directory());\n+\t\telse\n+\t\t\treturn error_errno(\n+\t\t\t\t_(\"unable to create temporary file\"));\n+\t}\n+\n+\t/*  Setup zlib stream for compression */\n+\tgit_deflate_init(stream, zlib_compression_level);\n+\tstream->next_out = buf;\n+\tstream->avail_out = buflen;\n+\tthe_hash_algo->init_fn(c);\n+\n+\t/*  Start to feed header to zlib stream */\n+\tstream->next_in = (unsigned char *)hdr;\n+\tstream->avail_in = hdrlen;\n+\twhile (git_deflate(stream, 0) == Z_OK)\n+\t\t; /* nothing */\n+\tthe_hash_algo->update_fn(c, hdr, hdrlen);\n+\n+\treturn fd;\n+}\n+\n+/**\n+ * Common steps for loose object writers to end writing loose objects:\n+ *\n+ * - End the compression of zlib stream.\n+ * - Get the calculated oid to \"oid\".\n+ */\n+static int end_loose_object_common(git_hash_ctx *c, git_zstream *stream,\n+\t\t\t\t   struct object_id *oid)\n+{\n+\tint ret;\n+\n+\tret = git_deflate_end_gently(stream);\n+\tif (ret != Z_OK)\n+\t\treturn ret;\n+\tthe_hash_algo->final_oid_fn(oid, c);\n+\n+\treturn Z_OK;\n+}\n+\n static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\t\t      int hdrlen, const void *buf, unsigned long len,\n \t\t\t      time_t mtime, unsigned flags)\n@@ -1968,28 +2036,11 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \n \tloose_object_path(the_repository, &filename, oid);\n \n-\tfd = create_tmpfile(&tmp_file, filename.buf);\n-\tif (fd < 0) {\n-\t\tif (flags & HASH_SILENT)\n-\t\t\treturn -1;\n-\t\telse if (errno == EACCES)\n-\t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n-\t\telse\n-\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n-\t}\n-\n-\t/* Set it up */\n-\tgit_deflate_init(&stream, zlib_compression_level);\n-\tstream.next_out = compressed;\n-\tstream.avail_out = sizeof(compressed);\n-\tthe_hash_algo->init_fn(&c);\n-\n-\t/* First header.. */\n-\tstream.next_in = (unsigned char *)hdr;\n-\tstream.avail_in = hdrlen;\n-\twhile (git_deflate(&stream, 0) == Z_OK)\n-\t\t; /* nothing */\n-\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n+\tfd = start_loose_object_common(&tmp_file, filename.buf, flags,\n+\t\t\t\t       &stream, compressed, sizeof(compressed),\n+\t\t\t\t       &c, hdr, hdrlen);\n+\tif (fd < 0)\n+\t\treturn -1;\n \n \t/* Then the data itself.. */\n \tstream.next_in = (void *)buf;\n@@ -2007,11 +2058,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tif (ret != Z_STREAM_END)\n \t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n \t\t    ret);\n-\tret = git_deflate_end_gently(&stream);\n+\tret = end_loose_object_common(&c, &stream, &parano_oid);\n \tif (ret != Z_OK)\n-\t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n-\t\t    ret);\n-\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n+\t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid), ret);\n \tclose_loose_object(fd, tmp_file.buf);\n \n \tif (!oideq(oid, &parano_oid))\n-- \n2.36.1\n\n"},{"id":"457024","messageId":"7c73815f188f16bb91c9b4ad981d299330dd3424.1654871916.git.chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"cover.1654871915.git.chiyutianyi@gmail.com","subject":"[PATCH v14 4/7] object-file.c: factor out deflate part of write_loose_object()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-10T14:46:04Z","receivedAt":"2022-06-10T14:48:06Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n\nSplit out the part of write_loose_object() that deals with calling\ngit_deflate() into a utility function, a subsequent commit will\nintroduce another function that'll make use of it.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n object-file.c | 31 +++++++++++++++++++++++++------\n 1 file changed, 25 insertions(+), 6 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex f4d7f8c109..cfae54762e 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2000,6 +2000,28 @@ static int start_loose_object_common(struct strbuf *tmp_file,\n \treturn fd;\n }\n \n+/**\n+ * Common steps for the inner git_deflate() loop for writing loose\n+ * objects. Returns what git_deflate() returns.\n+ */\n+static int write_loose_object_common(git_hash_ctx *c,\n+\t\t\t\t     git_zstream *stream, const int flush,\n+\t\t\t\t     unsigned char *in0, const int fd,\n+\t\t\t\t     unsigned char *compressed,\n+\t\t\t\t     const size_t compressed_len)\n+{\n+\tint ret;\n+\n+\tret = git_deflate(stream, flush ? Z_FINISH : 0);\n+\tthe_hash_algo->update_fn(c, in0, stream->next_in - in0);\n+\tif (write_buffer(fd, compressed, stream->next_out - compressed) < 0)\n+\t\tdie(_(\"unable to write loose object file\"));\n+\tstream->next_out = compressed;\n+\tstream->avail_out = compressed_len;\n+\n+\treturn ret;\n+}\n+\n /**\n  * Common steps for loose object writers to end writing loose objects:\n  *\n@@ -2047,12 +2069,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tstream.avail_in = len;\n \tdo {\n \t\tunsigned char *in0 = stream.next_in;\n-\t\tret = git_deflate(&stream, Z_FINISH);\n-\t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n-\t\tif (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n-\t\t\tdie(_(\"unable to write loose object file\"));\n-\t\tstream.next_out = compressed;\n-\t\tstream.avail_out = sizeof(compressed);\n+\n+\t\tret = write_loose_object_common(&c, &stream, 1, in0, fd,\n+\t\t\t\t\t\tcompressed, sizeof(compressed));\n \t} while (ret == Z_OK);\n \n \tif (ret != Z_STREAM_END)\n-- \n2.36.1\n\n"},{"id":"457025","messageId":"28a9588f9ceda2252d8ca9c4b3912177c45cb95c.1654871916.git.chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"cover.1654871915.git.chiyutianyi@gmail.com","subject":"[PATCH v14 5/7] object-file.c: add \"stream_loose_object()\" to handle large object","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-10T14:46:05Z","receivedAt":"2022-06-10T14:48:30Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nIf we want unpack and write a loose object using \"write_loose_object\",\nwe have to feed it with a buffer with the same size of the object, which\nwill consume lots of memory and may cause OOM. This can be improved by\nfeeding data to \"stream_loose_object()\" in a stream.\n\nAdd a new function \"stream_loose_object()\", which is a stream version of\n\"write_loose_object()\" but with a low memory footprint. We will use this\nfunction to unpack large blob object in later commit.\n\nAnother difference with \"write_loose_object()\" is that we have no chance\nto run \"write_object_file_prepare()\" to calculate the oid in advance.\nIn \"write_loose_object()\", we know the oid and we can write the\ntemporary file in the same directory as the final object, but for an\nobject with an undetermined oid, we don't know the exact directory for\nthe object.\n\nStill, we need to save the temporary file we're preparing\nsomewhere. We'll do that in the top-level \".git/objects/\"\ndirectory (or whatever \"GIT_OBJECT_DIRECTORY\" is set to). Once we've\nstreamed it we'll know the OID, and will move it to its canonical\npath.\n\n\"freshen_packed_object()\" or \"freshen_loose_object()\" will be called\ninside \"stream_loose_object()\" after obtaining the \"oid\". After the\ntemporary file is written, we wants to mark the object to recent and we\nmay find that where indeed is already the object. We should remove the\ntemporary and do not leave a new copy of the object.\n\nHelped-by: René Scharfe <l.s.r@web.de>\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n object-file.c  | 104 +++++++++++++++++++++++++++++++++++++++++++++++++\n object-store.h |   8 ++++\n 2 files changed, 112 insertions(+)\n\ndiff --git a/object-file.c b/object-file.c\nindex cfae54762e..0b8383ad47 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2118,6 +2118,110 @@ static int freshen_packed_object(const struct object_id *oid)\n \treturn 1;\n }\n \n+int stream_loose_object(struct input_stream *in_stream, size_t len,\n+\t\t\tstruct object_id *oid)\n+{\n+\tint fd, ret, err = 0, flush = 0;\n+\tunsigned char compressed[4096];\n+\tgit_zstream stream;\n+\tgit_hash_ctx c;\n+\tstruct strbuf tmp_file = STRBUF_INIT;\n+\tstruct strbuf filename = STRBUF_INIT;\n+\tint dirlen;\n+\tchar hdr[MAX_HEADER_LEN];\n+\tint hdrlen;\n+\n+\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n+\t\tprepare_loose_object_bulk_checkin();\n+\n+\t/* Since oid is not determined, save tmp file to odb path. */\n+\tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n+\thdrlen = format_object_header(hdr, sizeof(hdr), OBJ_BLOB, len);\n+\n+\t/*\n+\t * Common steps for write_loose_object and stream_loose_object to\n+\t * start writing loose objects:\n+\t *\n+\t *  - Create tmpfile for the loose object.\n+\t *  - Setup zlib stream for compression.\n+\t *  - Start to feed header to zlib stream.\n+\t */\n+\tfd = start_loose_object_common(&tmp_file, filename.buf, 0,\n+\t\t\t\t       &stream, compressed, sizeof(compressed),\n+\t\t\t\t       &c, hdr, hdrlen);\n+\tif (fd < 0) {\n+\t\terr = -1;\n+\t\tgoto cleanup;\n+\t}\n+\n+\t/* Then the data itself.. */\n+\tdo {\n+\t\tunsigned char *in0 = stream.next_in;\n+\n+\t\tif (!stream.avail_in && !in_stream->is_finished) {\n+\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n+\t\t\tstream.next_in = (void *)in;\n+\t\t\tin0 = (unsigned char *)in;\n+\t\t\t/* All data has been read. */\n+\t\t\tif (in_stream->is_finished)\n+\t\t\t\tflush = 1;\n+\t\t}\n+\t\tret = write_loose_object_common(&c, &stream, flush, in0, fd,\n+\t\t\t\t\t\tcompressed, sizeof(compressed));\n+\t\t/*\n+\t\t * Unlike write_loose_object(), we do not have the entire\n+\t\t * buffer. If we get Z_BUF_ERROR due to too few input bytes,\n+\t\t * then we'll replenish them in the next input_stream->read()\n+\t\t * call when we loop.\n+\t\t */\n+\t} while (ret == Z_OK || ret == Z_BUF_ERROR);\n+\n+\tif (stream.total_in != len + hdrlen)\n+\t\tdie(_(\"write stream object %ld != %\"PRIuMAX), stream.total_in,\n+\t\t    (uintmax_t)len + hdrlen);\n+\n+\t/*\n+\t * Common steps for write_loose_object and stream_loose_object to\n+\t * end writing loose oject:\n+\t *\n+\t *  - End the compression of zlib stream.\n+\t *  - Get the calculated oid.\n+\t */\n+\tif (ret != Z_STREAM_END)\n+\t\tdie(_(\"unable to stream deflate new object (%d)\"), ret);\n+\tret = end_loose_object_common(&c, &stream, oid);\n+\tif (ret != Z_OK)\n+\t\tdie(_(\"deflateEnd on stream object failed (%d)\"), ret);\n+\tclose_loose_object(fd, tmp_file.buf);\n+\n+\tif (freshen_packed_object(oid) || freshen_loose_object(oid)) {\n+\t\tunlink_or_warn(tmp_file.buf);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tloose_object_path(the_repository, &filename, oid);\n+\n+\t/* We finally know the object path, and create the missing dir. */\n+\tdirlen = directory_size(filename.buf);\n+\tif (dirlen) {\n+\t\tstruct strbuf dir = STRBUF_INIT;\n+\t\tstrbuf_add(&dir, filename.buf, dirlen);\n+\n+\t\tif (mkdir_in_gitdir(dir.buf) && errno != EEXIST) {\n+\t\t\terr = error_errno(_(\"unable to create directory %s\"), dir.buf);\n+\t\t\tstrbuf_release(&dir);\n+\t\t\tgoto cleanup;\n+\t\t}\n+\t\tstrbuf_release(&dir);\n+\t}\n+\n+\terr = finalize_object_file(tmp_file.buf, filename.buf);\n+cleanup:\n+\tstrbuf_release(&tmp_file);\n+\tstrbuf_release(&filename);\n+\treturn err;\n+}\n+\n int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t    enum object_type type, struct object_id *oid,\n \t\t\t    unsigned flags)\ndiff --git a/object-store.h b/object-store.h\nindex 539ea43904..5222ee5460 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -46,6 +46,12 @@ struct object_directory {\n \tchar *path;\n };\n \n+struct input_stream {\n+\tconst void *(*read)(struct input_stream *, unsigned long *len);\n+\tvoid *data;\n+\tint is_finished;\n+};\n+\n KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n \tstruct object_directory *, 1, fspathhash, fspatheq)\n \n@@ -269,6 +275,8 @@ static inline int write_object_file(const void *buf, unsigned long len,\n int write_object_file_literally(const void *buf, unsigned long len,\n \t\t\t\tconst char *type, struct object_id *oid,\n \t\t\t\tunsigned flags);\n+int stream_loose_object(struct input_stream *in_stream, size_t len,\n+\t\t\tstruct object_id *oid);\n \n /*\n  * Add an object file to the in-memory object store, without writing it\n-- \n2.36.1\n\n"},{"id":"457026","messageId":"dea5c4172b3908104d703dd09b644796ae52d873.1654871916.git.chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"cover.1654871915.git.chiyutianyi@gmail.com","subject":"[PATCH v14 6/7] core doc: modernize core.bigFileThreshold documentation","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-10T14:46:06Z","receivedAt":"2022-06-10T14:48:49Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n\nThe core.bigFileThreshold documentation has been largely unchanged\nsince 5eef828bc03 (fast-import: Stream very large blobs directly to\npack, 2010-02-01).\n\nBut since then this setting has been expanded to affect a lot more\nthan that description indicated. Most notably in how \"git diff\" treats\nthem, see 6bf3b813486 (diff --stat: mark any file larger than\ncore.bigfilethreshold binary, 2014-08-16).\n\nIn addition to that, numerous commands and APIs make use of a\nstreaming mode for files above this threshold.\n\nSo let's attempt to summarize 12 years of changes in behavior, which\ncan be seen with:\n\n    git log --oneline -Gbig_file_thre 5eef828bc03.. -- '*.c'\n\nTo do that turn this into a bullet-point list. The summary Han Xin\nproduced in [1] helped a lot, but is a bit too detailed for\ndocumentation aimed at users. Let's instead summarize how\nuser-observable behavior differs, and generally describe how we tend\nto stream these files in various commands.\n\n1. https://lore.kernel.org/git/20220120112114.47618-5-chiyutianyi@gmail.com/\n\nHelped-by: Han Xin <chiyutianyi@gmail.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n Documentation/config/core.txt | 33 ++++++++++++++++++++++++---------\n 1 file changed, 24 insertions(+), 9 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex 41e330f306..f2e75dd824 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -444,17 +444,32 @@ You probably do not need to adjust this value.\n Common unit suffixes of 'k', 'm', or 'g' are supported.\n \n core.bigFileThreshold::\n-\tFiles larger than this size are stored deflated, without\n-\tattempting delta compression.  Storing large files without\n-\tdelta compression avoids excessive memory usage, at the\n-\tslight expense of increased disk usage. Additionally files\n-\tlarger than this size are always treated as binary.\n+\tThe size of files considered \"big\", which as discussed below\n+\tchanges the behavior of numerous git commands, as well as how\n+\tsuch files are stored within the repository. The default is\n+\t512 MiB. Common unit suffixes of 'k', 'm', or 'g' are\n+\tsupported.\n +\n-Default is 512 MiB on all platforms.  This should be reasonable\n-for most projects as source code and other text files can still\n-be delta compressed, but larger binary media files won't be.\n+Files above the configured limit will be:\n +\n-Common unit suffixes of 'k', 'm', or 'g' are supported.\n+* Stored deflated in packfiles, without attempting delta compression.\n++\n+The default limit is primarily set with this use-case in mind. With it\n+most projects will have their source code and other text files delta\n+compressed, but not larger binary media files.\n++\n+Storing large files without delta compression avoids excessive memory\n+usage, at the slight expense of increased disk usage.\n++\n+* Will be treated as if though they were labeled \"binary\" (see\n+  linkgit:gitattributes[5]). e.g. linkgit:git-log[1] and\n+  linkgit:git-diff[1] will not diffs for files above this limit.\n++\n+* Will be generally be streamed when written, which avoids excessive\n+memory usage, at the cost of some fixed overhead. Commands that make\n+use of this include linkgit:git-archive[1],\n+linkgit:git-fast-import[1], linkgit:git-index-pack[1] and\n+linkgit:git-fsck[1].\n \n core.excludesFile::\n \tSpecifies the pathname to the file that contains patterns to\n-- \n2.36.1\n\n"},{"id":"457027","messageId":"d236230a4c5edf5e1c2685468c8ec0441743066e.1654871916.git.chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"cover.1654871915.git.chiyutianyi@gmail.com","subject":"[PATCH v14 7/7] unpack-objects: use stream_loose_object() to unpack large objects","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-10T14:46:07Z","receivedAt":"2022-06-10T14:49:07Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nMake use of the stream_loose_object() function introduced in the\npreceding commit to unpack large objects. Before this we'd need to\nmalloc() the size of the blob before unpacking it, which could cause\nOOM with very large blobs.\n\nWe could use the new streaming interface to unpack all blobs, but\ndoing so would be much slower, as demonstrated e.g. with this\nbenchmark using git-hyperfine[0]:\n\n\trm -rf /tmp/scalar.git &&\n\tgit clone --bare https://github.com/Microsoft/scalar.git /tmp/scalar.git &&\n\tmv /tmp/scalar.git/objects/pack/*.pack /tmp/scalar.git/my.pack &&\n\tgit hyperfine \\\n\t\t-r 2 --warmup 1 \\\n\t\t-L rev origin/master,HEAD -L v \"10,512,1k,1m\" \\\n\t\t-s 'make' \\\n\t\t-p 'git init --bare dest.git' \\\n\t\t-c 'rm -rf dest.git' \\\n\t\t'./git -C dest.git -c core.bigFileThreshold={v} unpack-objects </tmp/scalar.git/my.pack'\n\nHere we'll perform worse with lower core.bigFileThreshold settings\nwith this change in terms of speed, but we're getting lower memory use\nin return:\n\n\tSummary\n\t  './git -C dest.git -c core.bigFileThreshold=10 unpack-objects </tmp/scalar.git/my.pack' in 'origin/master' ran\n\t    1.01 ± 0.01 times faster than './git -C dest.git -c core.bigFileThreshold=1k unpack-objects </tmp/scalar.git/my.pack' in 'origin/master'\n\t    1.01 ± 0.01 times faster than './git -C dest.git -c core.bigFileThreshold=1m unpack-objects </tmp/scalar.git/my.pack' in 'origin/master'\n\t    1.01 ± 0.02 times faster than './git -C dest.git -c core.bigFileThreshold=1m unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n\t    1.02 ± 0.00 times faster than './git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/scalar.git/my.pack' in 'origin/master'\n\t    1.09 ± 0.01 times faster than './git -C dest.git -c core.bigFileThreshold=1k unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n\t    1.10 ± 0.00 times faster than './git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n\t    1.11 ± 0.00 times faster than './git -C dest.git -c core.bigFileThreshold=10 unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n\nA better benchmark to demonstrate the benefits of that this one, which\ncreates an artificial repo with a 1, 25, 50, 75 and 100MB blob:\n\n\trm -rf /tmp/repo &&\n\tgit init /tmp/repo &&\n\t(\n\t\tcd /tmp/repo &&\n\t\tfor i in 1 25 50 75 100\n\t\tdo\n\t\t\tdd if=/dev/urandom of=blob.$i count=$(($i*1024)) bs=1024\n\t\tdone &&\n\t\tgit add blob.* &&\n\t\tgit commit -mblobs &&\n\t\tgit gc &&\n\t\tPACK=$(echo .git/objects/pack/pack-*.pack) &&\n\t\tcp \"$PACK\" my.pack\n\t) &&\n\tgit hyperfine \\\n\t\t--show-output \\\n\t\t-L rev origin/master,HEAD -L v \"512,50m,100m\" \\\n\t\t-s 'make' \\\n\t\t-p 'git init --bare dest.git' \\\n\t\t-c 'rm -rf dest.git' \\\n\t\t'/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold={v} unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum'\n\nUsing this test we'll always use >100MB of memory on\norigin/master (around ~105MB), but max out at e.g. ~55MB if we set\ncore.bigFileThreshold=50m.\n\nThe relevant \"Maximum resident set size\" lines were manually added\nbelow the relevant benchmark:\n\n  '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=50m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'origin/master' ran\n        Maximum resident set size (kbytes): 107080\n    1.02 ± 0.78 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'origin/master'\n        Maximum resident set size (kbytes): 106968\n    1.09 ± 0.79 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=100m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'origin/master'\n        Maximum resident set size (kbytes): 107032\n    1.42 ± 1.07 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=100m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'HEAD'\n        Maximum resident set size (kbytes): 107072\n    1.83 ± 1.02 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=50m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'HEAD'\n        Maximum resident set size (kbytes): 55704\n    2.16 ± 1.19 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'HEAD'\n        Maximum resident set size (kbytes): 4564\n\nThis shows that if you have enough memory this new streaming method is\nslower the lower you set the streaming threshold, but the benefit is\nmore bounded memory use.\n\nAn earlier version of this patch introduced a new\n\"core.bigFileStreamingThreshold\" instead of re-using the existing\n\"core.bigFileThreshold\" variable[1]. As noted in a detailed overview\nof its users in [2] using it has several different meanings.\n\nStill, we consider it good enough to simply re-use it. While it's\npossible that someone might want to e.g. consider objects \"small\" for\nthe purposes of diffing but \"big\" for the purposes of writing them\nsuch use-cases are probably too obscure to worry about. We can always\nsplit up \"core.bigFileThreshold\" in the future if there's a need for\nthat.\n\n0. https://github.com/avar/git-hyperfine/\n1. https://lore.kernel.org/git/20211210103435.83656-1-chiyutianyi@gmail.com/\n2. https://lore.kernel.org/git/20220120112114.47618-5-chiyutianyi@gmail.com/\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Derrick Stolee <stolee@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n Documentation/config/core.txt   |  4 +-\n builtin/unpack-objects.c        | 69 ++++++++++++++++++++++++++++++++-\n t/t5351-unpack-large-objects.sh | 43 ++++++++++++++++++--\n 3 files changed, 109 insertions(+), 7 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex f2e75dd824..a599dcb96b 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -468,8 +468,8 @@ usage, at the slight expense of increased disk usage.\n * Will be generally be streamed when written, which avoids excessive\n memory usage, at the cost of some fixed overhead. Commands that make\n use of this include linkgit:git-archive[1],\n-linkgit:git-fast-import[1], linkgit:git-index-pack[1] and\n-linkgit:git-fsck[1].\n+linkgit:git-fast-import[1], linkgit:git-index-pack[1],\n+linkgit:git-unpack-objects[1] and linkgit:git-fsck[1].\n \n core.excludesFile::\n \tSpecifies the pathname to the file that contains patterns to\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 32e8b47059..43789b8ef2 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -351,6 +351,68 @@ static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n \t\twrite_object(nr, type, buf, size);\n }\n \n+struct input_zstream_data {\n+\tgit_zstream *zstream;\n+\tunsigned char buf[8192];\n+\tint status;\n+};\n+\n+static const void *feed_input_zstream(struct input_stream *in_stream,\n+\t\t\t\t      unsigned long *readlen)\n+{\n+\tstruct input_zstream_data *data = in_stream->data;\n+\tgit_zstream *zstream = data->zstream;\n+\tvoid *in = fill(1);\n+\n+\tif (in_stream->is_finished) {\n+\t\t*readlen = 0;\n+\t\treturn NULL;\n+\t}\n+\n+\tzstream->next_out = data->buf;\n+\tzstream->avail_out = sizeof(data->buf);\n+\tzstream->next_in = in;\n+\tzstream->avail_in = len;\n+\n+\tdata->status = git_inflate(zstream, 0);\n+\n+\tin_stream->is_finished = data->status != Z_OK;\n+\tuse(len - zstream->avail_in);\n+\t*readlen = sizeof(data->buf) - zstream->avail_out;\n+\n+\treturn data->buf;\n+}\n+\n+static void stream_blob(unsigned long size, unsigned nr)\n+{\n+\tgit_zstream zstream = { 0 };\n+\tstruct input_zstream_data data = { 0 };\n+\tstruct input_stream in_stream = {\n+\t\t.read = feed_input_zstream,\n+\t\t.data = &data,\n+\t};\n+\tstruct obj_info *info = &obj_list[nr];\n+\n+\tdata.zstream = &zstream;\n+\tgit_inflate_init(&zstream);\n+\n+\tif (stream_loose_object(&in_stream, size, &info->oid))\n+\t\tdie(_(\"failed to write object in stream\"));\n+\n+\tif (data.status != Z_STREAM_END)\n+\t\tdie(_(\"inflate returned (%d)\"), data.status);\n+\tgit_inflate_end(&zstream);\n+\n+\tif (strict) {\n+\t\tstruct blob *blob = lookup_blob(the_repository, &info->oid);\n+\n+\t\tif (!blob)\n+\t\t\tdie(_(\"invalid blob object from stream\"));\n+\t\tblob->object.flags |= FLAG_WRITTEN;\n+\t}\n+\tinfo->obj = NULL;\n+}\n+\n static int resolve_against_held(unsigned nr, const struct object_id *base,\n \t\t\t\tvoid *delta_data, unsigned long delta_size)\n {\n@@ -483,9 +545,14 @@ static void unpack_one(unsigned nr)\n \t}\n \n \tswitch (type) {\n+\tcase OBJ_BLOB:\n+\t\tif (!dry_run && size > big_file_threshold) {\n+\t\t\tstream_blob(size, nr);\n+\t\t\treturn;\n+\t\t}\n+\t\t/* fallthrough */\n \tcase OBJ_COMMIT:\n \tcase OBJ_TREE:\n-\tcase OBJ_BLOB:\n \tcase OBJ_TAG:\n \t\tunpack_non_delta_entry(type, size, nr);\n \t\treturn;\ndiff --git a/t/t5351-unpack-large-objects.sh b/t/t5351-unpack-large-objects.sh\nindex 8d84313221..8ce8aa3b14 100755\n--- a/t/t5351-unpack-large-objects.sh\n+++ b/t/t5351-unpack-large-objects.sh\n@@ -9,7 +9,8 @@ test_description='git unpack-objects with large objects'\n \n prepare_dest () {\n \ttest_when_finished \"rm -rf dest.git\" &&\n-\tgit init --bare dest.git\n+\tgit init --bare dest.git &&\n+\tgit -C dest.git config core.bigFileThreshold \"$1\"\n }\n \n test_expect_success \"create large objects (1.5 MB) and PACK\" '\n@@ -17,7 +18,10 @@ test_expect_success \"create large objects (1.5 MB) and PACK\" '\n \ttest_commit --append foo big-blob &&\n \ttest-tool genrandom bar 1500000 >big-blob &&\n \ttest_commit --append bar big-blob &&\n-\tPACK=$(echo HEAD | git pack-objects --revs pack)\n+\tPACK=$(echo HEAD | git pack-objects --revs pack) &&\n+\tgit verify-pack -v pack-$PACK.pack >out &&\n+\tsed -n -e \"s/^\\([0-9a-f][0-9a-f]*\\).*\\(commit\\|tree\\|blob\\).*/\\1/p\" \\\n+\t\t<out >obj-list\n '\n \n test_expect_success 'set memory limitation to 1MB' '\n@@ -26,16 +30,47 @@ test_expect_success 'set memory limitation to 1MB' '\n '\n \n test_expect_success 'unpack-objects failed under memory limitation' '\n-\tprepare_dest &&\n+\tprepare_dest 2m &&\n \ttest_must_fail git -C dest.git unpack-objects <pack-$PACK.pack 2>err &&\n \tgrep \"fatal: attempting to allocate\" err\n '\n \n test_expect_success 'unpack-objects works with memory limitation in dry-run mode' '\n-\tprepare_dest &&\n+\tprepare_dest 2m &&\n \tgit -C dest.git unpack-objects -n <pack-$PACK.pack &&\n \ttest_stdout_line_count = 0 find dest.git/objects -type f &&\n \ttest_dir_is_empty dest.git/objects/pack\n '\n \n+test_expect_success 'unpack big object in stream' '\n+\tprepare_dest 1m &&\n+\tgit -C dest.git unpack-objects <pack-$PACK.pack &&\n+\ttest_dir_is_empty dest.git/objects/pack\n+'\n+\n+BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\n+test_expect_success 'unpack big object in stream (core.fsyncmethod=batch)' '\n+\tprepare_dest 1m &&\n+\tGIT_TRACE2_EVENT=\"$(pwd)/trace2.txt\" \\\n+\t\tgit -C dest.git $BATCH_CONFIGURATION unpack-objects <pack-$PACK.pack &&\n+\tgrep fsync/hardware-flush trace2.txt &&\n+\ttest_dir_is_empty dest.git/objects/pack &&\n+\tgit -C dest.git cat-file --batch-check=\"%(objectname)\" <obj-list >current &&\n+\tcmp obj-list current\n+'\n+\n+test_expect_success 'do not unpack existing large objects' '\n+\tprepare_dest 1m &&\n+\tgit -C dest.git index-pack --stdin <pack-$PACK.pack &&\n+\tgit -C dest.git unpack-objects <pack-$PACK.pack &&\n+\n+\t# The destination came up with the exact same pack...\n+\tDEST_PACK=$(echo dest.git/objects/pack/pack-*.pack) &&\n+\ttest_cmp pack-$PACK.pack $DEST_PACK &&\n+\n+\t# ...and wrote no loose objects\n+\ttest_stdout_line_count = 0 find dest.git/objects -type f ! -name \"pack-*\"\n+'\n+\n test_done\n-- \n2.36.1\n\n"},{"id":"457043","messageId":"xmqqh74scjxl.fsf@gitster.g","threadId":"56672","inReplyTo":"dea5c4172b3908104d703dd09b644796ae52d873.1654871916.git.chiyutianyi@gmail.com","subject":"Re: [PATCH v14 6/7] core doc: modernize core.bigFileThreshold documentation","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-06-10T21:01:10Z","receivedAt":"2022-06-10T21:01:18Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Han Xin <chiyutianyi@gmail.com> writes:\n\n>  core.bigFileThreshold::\n\n> +\tThe size of files considered \"big\", which as discussed below\n> +\tchanges the behavior of numerous git commands, as well as how\n> +\tsuch files are stored within the repository. The default is\n> +\t512 MiB. Common unit suffixes of 'k', 'm', or 'g' are\n> +\tsupported.\n>  +\n> +Files above the configured limit will be:\n>  +\n> +* Stored deflated in packfiles, without attempting delta compression.\n> ++\n> +The default limit is primarily set with this use-case in mind. With it\n\n\"With it\" -> \"With it,\"\n\n> +most projects will have their source code and other text files delta\n> +compressed, but not larger binary media files.\n> +\n> +Storing large files without delta compression avoids excessive memory\n> +usage, at the slight expense of increased disk usage.\n\nMakes sense.\n\n> +* Will be treated as if though they were labeled \"binary\" (see\n\n\"as if though\" -> \"as if\"\n\n> +  linkgit:gitattributes[5]). e.g. linkgit:git-log[1] and\n> +  linkgit:git-diff[1] will not diffs for files above this limit.\n\n\"will not diffs\" -ECANNOTPARSE.  \"will not compute diffs\", probably?\n\n> ++\n> +* Will be generally be streamed when written, which avoids excessive\n\n\"be generally be\" -> \"generally be\"\n\n> +memory usage, at the cost of some fixed overhead. Commands that make\n> +use of this include linkgit:git-archive[1],\n> +linkgit:git-fast-import[1], linkgit:git-index-pack[1] and\n> +linkgit:git-fsck[1].\n>  \n>  core.excludesFile::\n>  \tSpecifies the pathname to the file that contains patterns to\n"},{"id":"457044","messageId":"0b9bc499-18c7-e8ab-5c89-f9e1a98685bc@web.de","threadId":"56672","inReplyTo":"a327f484f7f7466597930e87686e7156beabdc45.1654871916.git.chiyutianyi@gmail.com","subject":"Re: [PATCH v14 2/7] object-file.c: do fsync() and close() before post-write die()","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2022-06-10T21:10:17Z","receivedAt":"2022-06-10T21:10:40Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 10.06.22 um 16:46 schrieb Han Xin:\n> From: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n>\n> Change write_loose_object() to do an fsync() and close() before the\n> oideq() sanity check at the end. This change re-joins code that was\n> split up by the die() sanity check added in 748af44c63e (sha1_file: be\n> paranoid when creating loose objects, 2010-02-21).\n>\n> I don't think that this change matters in itself, if we called die()\n> it was possible that our data wouldn't fully make it to disk, but in\n> any case we were writing data that we'd consider corrupted. It's\n> possible that a subsequent \"git fsck\" will be less confused now.\n\nThis is done before renaming the file, so git fsck is going to see (at\nmost) a tmp_obj_?????? file, which it ignores in either case, right?\n\n> The real reason to make this change is that in a subsequent commit\n> we'll split this code in write_loose_object() into a utility function,\n> all its callers will want the preceding sanity checks, but not the\n> \"oideq\" check. By moving the close_loose_object() earlier it'll be\n> easier to reason about the introduction of the utility function.\n\nThis sounds like a patch would move the close_loose_object() call to\nsome other place, but that's not the case.  The sequence below (starting\nfrom the close_loose_object() call) is still present after applying the\nwhole series, so it seems this patch is not necessary.\n\n>\n> Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> ---\n>  object-file.c | 4 ++--\n>  1 file changed, 2 insertions(+), 2 deletions(-)\n>\n> diff --git a/object-file.c b/object-file.c\n> index 79eb8339b6..e4a83012ba 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -2012,12 +2012,12 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n>  \t\t    ret);\n>  \tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n> +\tclose_loose_object(fd, tmp_file.buf);\n> +\n>  \tif (!oideq(oid, &parano_oid))\n>  \t\tdie(_(\"confused by unstable object source data for %s\"),\n>  \t\t    oid_to_hex(oid));\n>\n> -\tclose_loose_object(fd, tmp_file.buf);\n> -\n>  \tif (mtime) {\n>  \t\tstruct utimbuf utb;\n>  \t\tutb.actime = mtime;\n"},{"id":"457045","messageId":"xmqqczfgcigc.fsf@gitster.g","threadId":"56672","inReplyTo":"0b9bc499-18c7-e8ab-5c89-f9e1a98685bc@web.de","subject":"Re: [PATCH v14 2/7] object-file.c: do fsync() and close() before post-write die()","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-06-10T21:33:07Z","receivedAt":"2022-06-10T21:33:19Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"René Scharfe <l.s.r@web.de> writes:\n\n> Am 10.06.22 um 16:46 schrieb Han Xin:\n>> From: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n>>\n>> Change write_loose_object() to do an fsync() and close() before the\n>> oideq() sanity check at the end. This change re-joins code that was\n>> split up by the die() sanity check added in 748af44c63e (sha1_file: be\n>> paranoid when creating loose objects, 2010-02-21).\n>>\n>> I don't think that this change matters in itself, if we called die()\n>> it was possible that our data wouldn't fully make it to disk, but in\n>> any case we were writing data that we'd consider corrupted. It's\n>> possible that a subsequent \"git fsck\" will be less confused now.\n>\n> This is done before renaming the file, so git fsck is going to see (at\n> most) a tmp_obj_?????? file, which it ignores in either case, right?\n\nYes, I thought I pointed that out in my review on the previous\nround, but I missed that it was still here in this round X-<.\n\nThanks for noticing.\n\n"},{"id":"457050","messageId":"CAO0brD3ZmZUnBSKGcsXHJacZkOgbgow8-2ykBoF1k6V=5stfeg@mail.gmail.com","threadId":"56672","inReplyTo":"xmqqczfgcigc.fsf@gitster.g","subject":"Re: [PATCH v14 2/7] object-file.c: do fsync() and close() before post-write die()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-11T01:50:41Z","receivedAt":"2022-06-11T01:50:59Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"On Sat, Jun 11, 2022 at 5:33 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> René Scharfe <l.s.r@web.de> writes:\n>\n> > Am 10.06.22 um 16:46 schrieb Han Xin:\n> >> From: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> >>\n> >> Change write_loose_object() to do an fsync() and close() before the\n> >> oideq() sanity check at the end. This change re-joins code that was\n> >> split up by the die() sanity check added in 748af44c63e (sha1_file: be\n> >> paranoid when creating loose objects, 2010-02-21).\n> >>\n> >> I don't think that this change matters in itself, if we called die()\n> >> it was possible that our data wouldn't fully make it to disk, but in\n> >> any case we were writing data that we'd consider corrupted. It's\n> >> possible that a subsequent \"git fsck\" will be less confused now.\n> >\n> > This is done before renaming the file, so git fsck is going to see (at\n> > most) a tmp_obj_?????? file, which it ignores in either case, right?\n>\n> Yes, I thought I pointed that out in my review on the previous\n> round, but I missed that it was still here in this round X-<.\n>\n> Thanks for noticing.\n>\n\nYes, agree with both of you, I'll be removing this patch in the next series.\n\nThis patch was first introduced in v10[1], close_loose_object() was moved\nto end_loose_object_common(), but it was put back in v12[2]. It is indeed\nno longer necessary now.\n\n1. https://lore.kernel.org/git/patch-v10-3.6-0e33d2a6e35-20220204T135538Z-avarab@gmail.com/\n2. https://lore.kernel.org/git/patch-v12-2.8-54060eb8c6b-20220329T135446Z-avarab@gmail.com/\n\nThank you both.\n-Han Xin\n\n-Han Xin\n"},{"id":"457051","messageId":"cover.1654914555.git.chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"cover.1654871915.git.chiyutianyi@gmail.com","subject":"[PATCH v15 0/6] unpack-objects: support streaming blobs to disk","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-11T02:44:15Z","receivedAt":"2022-06-11T02:44:39Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"This series makes \"unpack-objects\" capable of streaming large objects\nto disk.\n\nAs 6/6 shows streaming e.g. a 100MB blob now uses ~5MB of memory\ninstead of ~105MB. This streaming method is slower if you've got\nmemory to handle the blobs in-core, but if you don't it allows you to\nunpack objects at all, as you might otherwise OOM.\n\nChanges since v14:\n\n* Remove \"object-file.c: do fsync() and close() before post-write die()\"\n  as it's not necessary anymore. It was first introduced in v10 and was\n  no longer in the utility function end_loose_object_common() since v12.\n  We can see the discussion[1].\n\n* Minor grammar/comment etc. fixes throughout.\n\n1. https://lore.kernel.org/git/0b9bc499-18c7-e8ab-5c89-f9e1a98685bc@web.de/\n\nHan Xin (4):\n  unpack-objects: low memory footprint for get_data() in dry_run mode\n  object-file.c: refactor write_loose_object() to several steps\n  object-file.c: add \"stream_loose_object()\" to handle large object\n  unpack-objects: use stream_loose_object() to unpack large objects\n\nÆvar Arnfjörð Bjarmason (2):\n  object-file.c: factor out deflate part of write_loose_object()\n  core doc: modernize core.bigFileThreshold documentation\n\n Documentation/config/core.txt   |  33 +++--\n builtin/unpack-objects.c        | 106 +++++++++++++--\n object-file.c                   | 233 ++++++++++++++++++++++++++++----\n object-store.h                  |   8 ++\n t/t5351-unpack-large-objects.sh |  76 +++++++++++\n 5 files changed, 405 insertions(+), 51 deletions(-)\n create mode 100755 t/t5351-unpack-large-objects.sh\n\nRange-diff against v14:\n1:  bf600a2fa8 ! 1:  9a776f717d unpack-objects: low memory footprint for get_data() in dry_run mode\n    @@ Commit message\n         d9545c7f465 (fast-import: implement unpack limit, 2016-04-25).\n     \n         Suggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n    -    Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n    +    Signed-off-by: Han Xin <chiyutianyi@gmail.com>\n         Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n      ## builtin/unpack-objects.c ##\n2:  a327f484f7 < -:  ---------- object-file.c: do fsync() and close() before post-write die()\n3:  9bc8002282 ! 2:  a1e090d338 object-file.c: refactor write_loose_object() to several steps\n    @@ Commit message\n     \n         Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n         Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n    -    Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n    +    Signed-off-by: Han Xin <chiyutianyi@gmail.com>\n         Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n      ## object-file.c ##\n    @@ object-file.c: static int write_loose_object(const struct object_id *oid, char *\n     -\tret = git_deflate_end_gently(&stream);\n     +\tret = end_loose_object_common(&c, &stream, &parano_oid);\n      \tif (ret != Z_OK)\n    --\t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n    --\t\t    ret);\n    + \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n    + \t\t    ret);\n     -\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n    -+\t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid), ret);\n    - \tclose_loose_object(fd, tmp_file.buf);\n    - \n      \tif (!oideq(oid, &parano_oid))\n    + \t\tdie(_(\"confused by unstable object source data for %s\"),\n    + \t\t    oid_to_hex(oid));\n4:  7c73815f18 = 3:  0ddf912d47 object-file.c: factor out deflate part of write_loose_object()\n5:  28a9588f9c ! 4:  f9e51d3c68 object-file.c: add \"stream_loose_object()\" to handle large object\n    @@ Commit message\n         Helped-by: René Scharfe <l.s.r@web.de>\n         Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n         Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n    -    Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n    +    Signed-off-by: Han Xin <chiyutianyi@gmail.com>\n         Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n      ## object-file.c ##\n6:  dea5c4172b ! 5:  61ae1c1632 core doc: modernize core.bigFileThreshold documentation\n    @@ Documentation/config/core.txt: You probably do not need to adjust this value.\n     -Common unit suffixes of 'k', 'm', or 'g' are supported.\n     +* Stored deflated in packfiles, without attempting delta compression.\n     ++\n    -+The default limit is primarily set with this use-case in mind. With it\n    ++The default limit is primarily set with this use-case in mind. With it,\n     +most projects will have their source code and other text files delta\n     +compressed, but not larger binary media files.\n     ++\n     +Storing large files without delta compression avoids excessive memory\n     +usage, at the slight expense of increased disk usage.\n     ++\n    -+* Will be treated as if though they were labeled \"binary\" (see\n    ++* Will be treated as if they were labeled \"binary\" (see\n     +  linkgit:gitattributes[5]). e.g. linkgit:git-log[1] and\n    -+  linkgit:git-diff[1] will not diffs for files above this limit.\n    ++  linkgit:git-diff[1] will not compute diffs for files above this limit.\n     ++\n    -+* Will be generally be streamed when written, which avoids excessive\n    ++* Will generally be streamed when written, which avoids excessive\n     +memory usage, at the cost of some fixed overhead. Commands that make\n     +use of this include linkgit:git-archive[1],\n     +linkgit:git-fast-import[1], linkgit:git-index-pack[1] and\n7:  d236230a4c ! 6:  5a4782d746 unpack-objects: use stream_loose_object() to unpack large objects\n    @@ Commit message\n         Helped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n         Helped-by: Derrick Stolee <stolee@gmail.com>\n         Helped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\n    -    Signed-off-by: Han Xin <hanxin.hx@alibaba-inc.com>\n    +    Signed-off-by: Han Xin <chiyutianyi@gmail.com>\n         Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n      ## Documentation/config/core.txt ##\n     @@ Documentation/config/core.txt: usage, at the slight expense of increased disk usage.\n    - * Will be generally be streamed when written, which avoids excessive\n    + * Will generally be streamed when written, which avoids excessive\n      memory usage, at the cost of some fixed overhead. Commands that make\n      use of this include linkgit:git-archive[1],\n     -linkgit:git-fast-import[1], linkgit:git-index-pack[1] and\n-- \n2.36.1\n\n"},{"id":"457052","messageId":"9a776f717d512dc63888a9334074bbf1728395c5.1654914555.git.chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"cover.1654914555.git.chiyutianyi@gmail.com","subject":"[PATCH v15 1/6] unpack-objects: low memory footprint for get_data() in dry_run mode","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-11T02:44:16Z","receivedAt":"2022-06-11T02:44:48Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nAs the name implies, \"get_data(size)\" will allocate and return a given\namount of memory. Allocating memory for a large blob object may cause the\nsystem to run out of memory. Before preparing to replace calling of\n\"get_data()\" to unpack large blob objects in latter commits, refactor\n\"get_data()\" to reduce memory footprint for dry_run mode.\n\nBecause in dry_run mode, \"get_data()\" is only used to check the\nintegrity of data, and the returned buffer is not used at all, we can\nallocate a smaller buffer and use it as zstream output. Make the function\nreturn NULL in the dry-run mode, as no callers use the returned buffer.\n\nThe \"find [...]objects/?? -type f | wc -l\" test idiom being used here\nis adapted from the same \"find\" use added to another test in\nd9545c7f465 (fast-import: implement unpack limit, 2016-04-25).\n\nSuggested-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <chiyutianyi@gmail.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n builtin/unpack-objects.c        | 37 ++++++++++++++++++++---------\n t/t5351-unpack-large-objects.sh | 41 +++++++++++++++++++++++++++++++++\n 2 files changed, 67 insertions(+), 11 deletions(-)\n create mode 100755 t/t5351-unpack-large-objects.sh\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 56d05e2725..32e8b47059 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -97,15 +97,27 @@ static void use(int bytes)\n \tdisplay_throughput(progress, consumed_bytes);\n }\n \n+/*\n+ * Decompress zstream from the standard input into a newly\n+ * allocated buffer of specified size and return the buffer.\n+ * The caller is responsible to free the returned buffer.\n+ *\n+ * But for dry_run mode, \"get_data()\" is only used to check the\n+ * integrity of data, and the returned buffer is not used at all.\n+ * Therefore, in dry_run mode, \"get_data()\" will release the small\n+ * allocated buffer which is reused to hold temporary zstream output\n+ * and return NULL instead of returning garbage data.\n+ */\n static void *get_data(unsigned long size)\n {\n \tgit_zstream stream;\n-\tvoid *buf = xmallocz(size);\n+\tunsigned long bufsize = dry_run && size > 8192 ? 8192 : size;\n+\tvoid *buf = xmallocz(bufsize);\n \n \tmemset(&stream, 0, sizeof(stream));\n \n \tstream.next_out = buf;\n-\tstream.avail_out = size;\n+\tstream.avail_out = bufsize;\n \tstream.next_in = fill(1);\n \tstream.avail_in = len;\n \tgit_inflate_init(&stream);\n@@ -125,8 +137,17 @@ static void *get_data(unsigned long size)\n \t\t}\n \t\tstream.next_in = fill(1);\n \t\tstream.avail_in = len;\n+\t\tif (dry_run) {\n+\t\t\t/* reuse the buffer in dry_run mode */\n+\t\t\tstream.next_out = buf;\n+\t\t\tstream.avail_out = bufsize > size - stream.total_out ?\n+\t\t\t\t\t\t   size - stream.total_out :\n+\t\t\t\t\t\t   bufsize;\n+\t\t}\n \t}\n \tgit_inflate_end(&stream);\n+\tif (dry_run)\n+\t\tFREE_AND_NULL(buf);\n \treturn buf;\n }\n \n@@ -326,10 +347,8 @@ static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n {\n \tvoid *buf = get_data(size);\n \n-\tif (!dry_run && buf)\n+\tif (buf)\n \t\twrite_object(nr, type, buf, size);\n-\telse\n-\t\tfree(buf);\n }\n \n static int resolve_against_held(unsigned nr, const struct object_id *base,\n@@ -359,10 +378,8 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\toidread(&base_oid, fill(the_hash_algo->rawsz));\n \t\tuse(the_hash_algo->rawsz);\n \t\tdelta_data = get_data(delta_size);\n-\t\tif (dry_run || !delta_data) {\n-\t\t\tfree(delta_data);\n+\t\tif (!delta_data)\n \t\t\treturn;\n-\t\t}\n \t\tif (has_object_file(&base_oid))\n \t\t\t; /* Ok we have this one */\n \t\telse if (resolve_against_held(nr, &base_oid,\n@@ -398,10 +415,8 @@ static void unpack_delta_entry(enum object_type type, unsigned long delta_size,\n \t\t\tdie(\"offset value out of bound for delta base object\");\n \n \t\tdelta_data = get_data(delta_size);\n-\t\tif (dry_run || !delta_data) {\n-\t\t\tfree(delta_data);\n+\t\tif (!delta_data)\n \t\t\treturn;\n-\t\t}\n \t\tlo = 0;\n \t\thi = nr;\n \t\twhile (lo < hi) {\ndiff --git a/t/t5351-unpack-large-objects.sh b/t/t5351-unpack-large-objects.sh\nnew file mode 100755\nindex 0000000000..8d84313221\n--- /dev/null\n+++ b/t/t5351-unpack-large-objects.sh\n@@ -0,0 +1,41 @@\n+#!/bin/sh\n+#\n+# Copyright (c) 2022 Han Xin\n+#\n+\n+test_description='git unpack-objects with large objects'\n+\n+. ./test-lib.sh\n+\n+prepare_dest () {\n+\ttest_when_finished \"rm -rf dest.git\" &&\n+\tgit init --bare dest.git\n+}\n+\n+test_expect_success \"create large objects (1.5 MB) and PACK\" '\n+\ttest-tool genrandom foo 1500000 >big-blob &&\n+\ttest_commit --append foo big-blob &&\n+\ttest-tool genrandom bar 1500000 >big-blob &&\n+\ttest_commit --append bar big-blob &&\n+\tPACK=$(echo HEAD | git pack-objects --revs pack)\n+'\n+\n+test_expect_success 'set memory limitation to 1MB' '\n+\tGIT_ALLOC_LIMIT=1m &&\n+\texport GIT_ALLOC_LIMIT\n+'\n+\n+test_expect_success 'unpack-objects failed under memory limitation' '\n+\tprepare_dest &&\n+\ttest_must_fail git -C dest.git unpack-objects <pack-$PACK.pack 2>err &&\n+\tgrep \"fatal: attempting to allocate\" err\n+'\n+\n+test_expect_success 'unpack-objects works with memory limitation in dry-run mode' '\n+\tprepare_dest &&\n+\tgit -C dest.git unpack-objects -n <pack-$PACK.pack &&\n+\ttest_stdout_line_count = 0 find dest.git/objects -type f &&\n+\ttest_dir_is_empty dest.git/objects/pack\n+'\n+\n+test_done\n-- \n2.36.1\n\n"},{"id":"457053","messageId":"a1e090d338ed9761336d78104bd7df71953138b4.1654914555.git.chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"cover.1654914555.git.chiyutianyi@gmail.com","subject":"[PATCH v15 2/6] object-file.c: refactor write_loose_object() to several steps","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-11T02:44:17Z","receivedAt":"2022-06-11T02:44:50Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nWhen writing a large blob using \"write_loose_object()\", we have to pass\na buffer with the whole content of the blob, and this behavior will\nconsume lots of memory and may cause OOM. We will introduce a stream\nversion function (\"stream_loose_object()\") in later commit to resolve\nthis issue.\n\nBefore introducing that streaming function, do some refactoring on\n\"write_loose_object()\" to reuse code for both versions.\n\nRewrite \"write_loose_object()\" as follows:\n\n 1. Figure out a path for the (temp) object file. This step is only\n    used in \"write_loose_object()\".\n\n 2. Move common steps for starting to write loose objects into a new\n    function \"start_loose_object_common()\".\n\n 3. Compress data.\n\n 4. Move common steps for ending zlib stream into a new function\n    \"end_loose_object_common()\".\n\n 5. Close fd and finalize the object file.\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <chiyutianyi@gmail.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n object-file.c | 98 ++++++++++++++++++++++++++++++++++++++-------------\n 1 file changed, 74 insertions(+), 24 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 79eb8339b6..b5bce03274 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1951,6 +1951,74 @@ static int create_tmpfile(struct strbuf *tmp, const char *filename)\n \treturn fd;\n }\n \n+/**\n+ * Common steps for loose object writers to start writing loose\n+ * objects:\n+ *\n+ * - Create tmpfile for the loose object.\n+ * - Setup zlib stream for compression.\n+ * - Start to feed header to zlib stream.\n+ *\n+ * Returns a \"fd\", which should later be provided to\n+ * end_loose_object_common().\n+ */\n+static int start_loose_object_common(struct strbuf *tmp_file,\n+\t\t\t\t     const char *filename, unsigned flags,\n+\t\t\t\t     git_zstream *stream,\n+\t\t\t\t     unsigned char *buf, size_t buflen,\n+\t\t\t\t     git_hash_ctx *c,\n+\t\t\t\t     char *hdr, int hdrlen)\n+{\n+\tint fd;\n+\n+\tfd = create_tmpfile(tmp_file, filename);\n+\tif (fd < 0) {\n+\t\tif (flags & HASH_SILENT)\n+\t\t\treturn -1;\n+\t\telse if (errno == EACCES)\n+\t\t\treturn error(_(\"insufficient permission for adding \"\n+\t\t\t\t       \"an object to repository database %s\"),\n+\t\t\t\t     get_object_directory());\n+\t\telse\n+\t\t\treturn error_errno(\n+\t\t\t\t_(\"unable to create temporary file\"));\n+\t}\n+\n+\t/*  Setup zlib stream for compression */\n+\tgit_deflate_init(stream, zlib_compression_level);\n+\tstream->next_out = buf;\n+\tstream->avail_out = buflen;\n+\tthe_hash_algo->init_fn(c);\n+\n+\t/*  Start to feed header to zlib stream */\n+\tstream->next_in = (unsigned char *)hdr;\n+\tstream->avail_in = hdrlen;\n+\twhile (git_deflate(stream, 0) == Z_OK)\n+\t\t; /* nothing */\n+\tthe_hash_algo->update_fn(c, hdr, hdrlen);\n+\n+\treturn fd;\n+}\n+\n+/**\n+ * Common steps for loose object writers to end writing loose objects:\n+ *\n+ * - End the compression of zlib stream.\n+ * - Get the calculated oid to \"oid\".\n+ */\n+static int end_loose_object_common(git_hash_ctx *c, git_zstream *stream,\n+\t\t\t\t   struct object_id *oid)\n+{\n+\tint ret;\n+\n+\tret = git_deflate_end_gently(stream);\n+\tif (ret != Z_OK)\n+\t\treturn ret;\n+\tthe_hash_algo->final_oid_fn(oid, c);\n+\n+\treturn Z_OK;\n+}\n+\n static int write_loose_object(const struct object_id *oid, char *hdr,\n \t\t\t      int hdrlen, const void *buf, unsigned long len,\n \t\t\t      time_t mtime, unsigned flags)\n@@ -1968,28 +2036,11 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \n \tloose_object_path(the_repository, &filename, oid);\n \n-\tfd = create_tmpfile(&tmp_file, filename.buf);\n-\tif (fd < 0) {\n-\t\tif (flags & HASH_SILENT)\n-\t\t\treturn -1;\n-\t\telse if (errno == EACCES)\n-\t\t\treturn error(_(\"insufficient permission for adding an object to repository database %s\"), get_object_directory());\n-\t\telse\n-\t\t\treturn error_errno(_(\"unable to create temporary file\"));\n-\t}\n-\n-\t/* Set it up */\n-\tgit_deflate_init(&stream, zlib_compression_level);\n-\tstream.next_out = compressed;\n-\tstream.avail_out = sizeof(compressed);\n-\tthe_hash_algo->init_fn(&c);\n-\n-\t/* First header.. */\n-\tstream.next_in = (unsigned char *)hdr;\n-\tstream.avail_in = hdrlen;\n-\twhile (git_deflate(&stream, 0) == Z_OK)\n-\t\t; /* nothing */\n-\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n+\tfd = start_loose_object_common(&tmp_file, filename.buf, flags,\n+\t\t\t\t       &stream, compressed, sizeof(compressed),\n+\t\t\t\t       &c, hdr, hdrlen);\n+\tif (fd < 0)\n+\t\treturn -1;\n \n \t/* Then the data itself.. */\n \tstream.next_in = (void *)buf;\n@@ -2007,11 +2058,10 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tif (ret != Z_STREAM_END)\n \t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n \t\t    ret);\n-\tret = git_deflate_end_gently(&stream);\n+\tret = end_loose_object_common(&c, &stream, &parano_oid);\n \tif (ret != Z_OK)\n \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n \t\t    ret);\n-\tthe_hash_algo->final_oid_fn(&parano_oid, &c);\n \tif (!oideq(oid, &parano_oid))\n \t\tdie(_(\"confused by unstable object source data for %s\"),\n \t\t    oid_to_hex(oid));\n-- \n2.36.1\n\n"},{"id":"457054","messageId":"0ddf912d479eeda47c47e6b770816831aed4ebdb.1654914555.git.chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"cover.1654914555.git.chiyutianyi@gmail.com","subject":"[PATCH v15 3/6] object-file.c: factor out deflate part of write_loose_object()","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-11T02:44:18Z","receivedAt":"2022-06-11T02:44:56Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n\nSplit out the part of write_loose_object() that deals with calling\ngit_deflate() into a utility function, a subsequent commit will\nintroduce another function that'll make use of it.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n object-file.c | 31 +++++++++++++++++++++++++------\n 1 file changed, 25 insertions(+), 6 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex b5bce03274..18dbf2a4e4 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2000,6 +2000,28 @@ static int start_loose_object_common(struct strbuf *tmp_file,\n \treturn fd;\n }\n \n+/**\n+ * Common steps for the inner git_deflate() loop for writing loose\n+ * objects. Returns what git_deflate() returns.\n+ */\n+static int write_loose_object_common(git_hash_ctx *c,\n+\t\t\t\t     git_zstream *stream, const int flush,\n+\t\t\t\t     unsigned char *in0, const int fd,\n+\t\t\t\t     unsigned char *compressed,\n+\t\t\t\t     const size_t compressed_len)\n+{\n+\tint ret;\n+\n+\tret = git_deflate(stream, flush ? Z_FINISH : 0);\n+\tthe_hash_algo->update_fn(c, in0, stream->next_in - in0);\n+\tif (write_buffer(fd, compressed, stream->next_out - compressed) < 0)\n+\t\tdie(_(\"unable to write loose object file\"));\n+\tstream->next_out = compressed;\n+\tstream->avail_out = compressed_len;\n+\n+\treturn ret;\n+}\n+\n /**\n  * Common steps for loose object writers to end writing loose objects:\n  *\n@@ -2047,12 +2069,9 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tstream.avail_in = len;\n \tdo {\n \t\tunsigned char *in0 = stream.next_in;\n-\t\tret = git_deflate(&stream, Z_FINISH);\n-\t\tthe_hash_algo->update_fn(&c, in0, stream.next_in - in0);\n-\t\tif (write_buffer(fd, compressed, stream.next_out - compressed) < 0)\n-\t\t\tdie(_(\"unable to write loose object file\"));\n-\t\tstream.next_out = compressed;\n-\t\tstream.avail_out = sizeof(compressed);\n+\n+\t\tret = write_loose_object_common(&c, &stream, 1, in0, fd,\n+\t\t\t\t\t\tcompressed, sizeof(compressed));\n \t} while (ret == Z_OK);\n \n \tif (ret != Z_STREAM_END)\n-- \n2.36.1\n\n"},{"id":"457055","messageId":"f9e51d3c680fcd2bfcd19a069d0500f73e3a3bac.1654914555.git.chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"cover.1654914555.git.chiyutianyi@gmail.com","subject":"[PATCH v15 4/6] object-file.c: add \"stream_loose_object()\" to handle large object","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-11T02:44:19Z","receivedAt":"2022-06-11T02:45:11Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nIf we want unpack and write a loose object using \"write_loose_object\",\nwe have to feed it with a buffer with the same size of the object, which\nwill consume lots of memory and may cause OOM. This can be improved by\nfeeding data to \"stream_loose_object()\" in a stream.\n\nAdd a new function \"stream_loose_object()\", which is a stream version of\n\"write_loose_object()\" but with a low memory footprint. We will use this\nfunction to unpack large blob object in later commit.\n\nAnother difference with \"write_loose_object()\" is that we have no chance\nto run \"write_object_file_prepare()\" to calculate the oid in advance.\nIn \"write_loose_object()\", we know the oid and we can write the\ntemporary file in the same directory as the final object, but for an\nobject with an undetermined oid, we don't know the exact directory for\nthe object.\n\nStill, we need to save the temporary file we're preparing\nsomewhere. We'll do that in the top-level \".git/objects/\"\ndirectory (or whatever \"GIT_OBJECT_DIRECTORY\" is set to). Once we've\nstreamed it we'll know the OID, and will move it to its canonical\npath.\n\n\"freshen_packed_object()\" or \"freshen_loose_object()\" will be called\ninside \"stream_loose_object()\" after obtaining the \"oid\". After the\ntemporary file is written, we wants to mark the object to recent and we\nmay find that where indeed is already the object. We should remove the\ntemporary and do not leave a new copy of the object.\n\nHelped-by: René Scharfe <l.s.r@web.de>\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <chiyutianyi@gmail.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n object-file.c  | 104 +++++++++++++++++++++++++++++++++++++++++++++++++\n object-store.h |   8 ++++\n 2 files changed, 112 insertions(+)\n\ndiff --git a/object-file.c b/object-file.c\nindex 18dbf2a4e4..2ca2576ab1 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2119,6 +2119,110 @@ static int freshen_packed_object(const struct object_id *oid)\n \treturn 1;\n }\n \n+int stream_loose_object(struct input_stream *in_stream, size_t len,\n+\t\t\tstruct object_id *oid)\n+{\n+\tint fd, ret, err = 0, flush = 0;\n+\tunsigned char compressed[4096];\n+\tgit_zstream stream;\n+\tgit_hash_ctx c;\n+\tstruct strbuf tmp_file = STRBUF_INIT;\n+\tstruct strbuf filename = STRBUF_INIT;\n+\tint dirlen;\n+\tchar hdr[MAX_HEADER_LEN];\n+\tint hdrlen;\n+\n+\tif (batch_fsync_enabled(FSYNC_COMPONENT_LOOSE_OBJECT))\n+\t\tprepare_loose_object_bulk_checkin();\n+\n+\t/* Since oid is not determined, save tmp file to odb path. */\n+\tstrbuf_addf(&filename, \"%s/\", get_object_directory());\n+\thdrlen = format_object_header(hdr, sizeof(hdr), OBJ_BLOB, len);\n+\n+\t/*\n+\t * Common steps for write_loose_object and stream_loose_object to\n+\t * start writing loose objects:\n+\t *\n+\t *  - Create tmpfile for the loose object.\n+\t *  - Setup zlib stream for compression.\n+\t *  - Start to feed header to zlib stream.\n+\t */\n+\tfd = start_loose_object_common(&tmp_file, filename.buf, 0,\n+\t\t\t\t       &stream, compressed, sizeof(compressed),\n+\t\t\t\t       &c, hdr, hdrlen);\n+\tif (fd < 0) {\n+\t\terr = -1;\n+\t\tgoto cleanup;\n+\t}\n+\n+\t/* Then the data itself.. */\n+\tdo {\n+\t\tunsigned char *in0 = stream.next_in;\n+\n+\t\tif (!stream.avail_in && !in_stream->is_finished) {\n+\t\t\tconst void *in = in_stream->read(in_stream, &stream.avail_in);\n+\t\t\tstream.next_in = (void *)in;\n+\t\t\tin0 = (unsigned char *)in;\n+\t\t\t/* All data has been read. */\n+\t\t\tif (in_stream->is_finished)\n+\t\t\t\tflush = 1;\n+\t\t}\n+\t\tret = write_loose_object_common(&c, &stream, flush, in0, fd,\n+\t\t\t\t\t\tcompressed, sizeof(compressed));\n+\t\t/*\n+\t\t * Unlike write_loose_object(), we do not have the entire\n+\t\t * buffer. If we get Z_BUF_ERROR due to too few input bytes,\n+\t\t * then we'll replenish them in the next input_stream->read()\n+\t\t * call when we loop.\n+\t\t */\n+\t} while (ret == Z_OK || ret == Z_BUF_ERROR);\n+\n+\tif (stream.total_in != len + hdrlen)\n+\t\tdie(_(\"write stream object %ld != %\"PRIuMAX), stream.total_in,\n+\t\t    (uintmax_t)len + hdrlen);\n+\n+\t/*\n+\t * Common steps for write_loose_object and stream_loose_object to\n+\t * end writing loose oject:\n+\t *\n+\t *  - End the compression of zlib stream.\n+\t *  - Get the calculated oid.\n+\t */\n+\tif (ret != Z_STREAM_END)\n+\t\tdie(_(\"unable to stream deflate new object (%d)\"), ret);\n+\tret = end_loose_object_common(&c, &stream, oid);\n+\tif (ret != Z_OK)\n+\t\tdie(_(\"deflateEnd on stream object failed (%d)\"), ret);\n+\tclose_loose_object(fd, tmp_file.buf);\n+\n+\tif (freshen_packed_object(oid) || freshen_loose_object(oid)) {\n+\t\tunlink_or_warn(tmp_file.buf);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tloose_object_path(the_repository, &filename, oid);\n+\n+\t/* We finally know the object path, and create the missing dir. */\n+\tdirlen = directory_size(filename.buf);\n+\tif (dirlen) {\n+\t\tstruct strbuf dir = STRBUF_INIT;\n+\t\tstrbuf_add(&dir, filename.buf, dirlen);\n+\n+\t\tif (mkdir_in_gitdir(dir.buf) && errno != EEXIST) {\n+\t\t\terr = error_errno(_(\"unable to create directory %s\"), dir.buf);\n+\t\t\tstrbuf_release(&dir);\n+\t\t\tgoto cleanup;\n+\t\t}\n+\t\tstrbuf_release(&dir);\n+\t}\n+\n+\terr = finalize_object_file(tmp_file.buf, filename.buf);\n+cleanup:\n+\tstrbuf_release(&tmp_file);\n+\tstrbuf_release(&filename);\n+\treturn err;\n+}\n+\n int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t    enum object_type type, struct object_id *oid,\n \t\t\t    unsigned flags)\ndiff --git a/object-store.h b/object-store.h\nindex 539ea43904..5222ee5460 100644\n--- a/object-store.h\n+++ b/object-store.h\n@@ -46,6 +46,12 @@ struct object_directory {\n \tchar *path;\n };\n \n+struct input_stream {\n+\tconst void *(*read)(struct input_stream *, unsigned long *len);\n+\tvoid *data;\n+\tint is_finished;\n+};\n+\n KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n \tstruct object_directory *, 1, fspathhash, fspatheq)\n \n@@ -269,6 +275,8 @@ static inline int write_object_file(const void *buf, unsigned long len,\n int write_object_file_literally(const void *buf, unsigned long len,\n \t\t\t\tconst char *type, struct object_id *oid,\n \t\t\t\tunsigned flags);\n+int stream_loose_object(struct input_stream *in_stream, size_t len,\n+\t\t\tstruct object_id *oid);\n \n /*\n  * Add an object file to the in-memory object store, without writing it\n-- \n2.36.1\n\n"},{"id":"457056","messageId":"61ae1c1632582ba1cfd9e15e375c57fdb3f559af.1654914555.git.chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"cover.1654914555.git.chiyutianyi@gmail.com","subject":"[PATCH v15 5/6] core doc: modernize core.bigFileThreshold documentation","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-11T02:44:20Z","receivedAt":"2022-06-11T02:45:14Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n\nThe core.bigFileThreshold documentation has been largely unchanged\nsince 5eef828bc03 (fast-import: Stream very large blobs directly to\npack, 2010-02-01).\n\nBut since then this setting has been expanded to affect a lot more\nthan that description indicated. Most notably in how \"git diff\" treats\nthem, see 6bf3b813486 (diff --stat: mark any file larger than\ncore.bigfilethreshold binary, 2014-08-16).\n\nIn addition to that, numerous commands and APIs make use of a\nstreaming mode for files above this threshold.\n\nSo let's attempt to summarize 12 years of changes in behavior, which\ncan be seen with:\n\n    git log --oneline -Gbig_file_thre 5eef828bc03.. -- '*.c'\n\nTo do that turn this into a bullet-point list. The summary Han Xin\nproduced in [1] helped a lot, but is a bit too detailed for\ndocumentation aimed at users. Let's instead summarize how\nuser-observable behavior differs, and generally describe how we tend\nto stream these files in various commands.\n\n1. https://lore.kernel.org/git/20220120112114.47618-5-chiyutianyi@gmail.com/\n\nHelped-by: Han Xin <chiyutianyi@gmail.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n Documentation/config/core.txt | 33 ++++++++++++++++++++++++---------\n 1 file changed, 24 insertions(+), 9 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex 41e330f306..87e4c04836 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -444,17 +444,32 @@ You probably do not need to adjust this value.\n Common unit suffixes of 'k', 'm', or 'g' are supported.\n \n core.bigFileThreshold::\n-\tFiles larger than this size are stored deflated, without\n-\tattempting delta compression.  Storing large files without\n-\tdelta compression avoids excessive memory usage, at the\n-\tslight expense of increased disk usage. Additionally files\n-\tlarger than this size are always treated as binary.\n+\tThe size of files considered \"big\", which as discussed below\n+\tchanges the behavior of numerous git commands, as well as how\n+\tsuch files are stored within the repository. The default is\n+\t512 MiB. Common unit suffixes of 'k', 'm', or 'g' are\n+\tsupported.\n +\n-Default is 512 MiB on all platforms.  This should be reasonable\n-for most projects as source code and other text files can still\n-be delta compressed, but larger binary media files won't be.\n+Files above the configured limit will be:\n +\n-Common unit suffixes of 'k', 'm', or 'g' are supported.\n+* Stored deflated in packfiles, without attempting delta compression.\n++\n+The default limit is primarily set with this use-case in mind. With it,\n+most projects will have their source code and other text files delta\n+compressed, but not larger binary media files.\n++\n+Storing large files without delta compression avoids excessive memory\n+usage, at the slight expense of increased disk usage.\n++\n+* Will be treated as if they were labeled \"binary\" (see\n+  linkgit:gitattributes[5]). e.g. linkgit:git-log[1] and\n+  linkgit:git-diff[1] will not compute diffs for files above this limit.\n++\n+* Will generally be streamed when written, which avoids excessive\n+memory usage, at the cost of some fixed overhead. Commands that make\n+use of this include linkgit:git-archive[1],\n+linkgit:git-fast-import[1], linkgit:git-index-pack[1] and\n+linkgit:git-fsck[1].\n \n core.excludesFile::\n \tSpecifies the pathname to the file that contains patterns to\n-- \n2.36.1\n\n"},{"id":"457057","messageId":"5a4782d746a496e8edd1654296bac390d8e1c9d3.1654914555.git.chiyutianyi@gmail.com","threadId":"56672","inReplyTo":"cover.1654914555.git.chiyutianyi@gmail.com","subject":"[PATCH v15 6/6] unpack-objects: use stream_loose_object() to unpack large objects","fromName":"Han Xin","fromEmail":"chiyutianyi@gmail.com","sentAt":"2022-06-11T02:44:21Z","receivedAt":"2022-06-11T02:45:15Z","isPatch":true,"sender":{"key":"chiyutianyi@gmail.com","avatar":null},"body":"From: Han Xin <hanxin.hx@alibaba-inc.com>\n\nMake use of the stream_loose_object() function introduced in the\npreceding commit to unpack large objects. Before this we'd need to\nmalloc() the size of the blob before unpacking it, which could cause\nOOM with very large blobs.\n\nWe could use the new streaming interface to unpack all blobs, but\ndoing so would be much slower, as demonstrated e.g. with this\nbenchmark using git-hyperfine[0]:\n\n\trm -rf /tmp/scalar.git &&\n\tgit clone --bare https://github.com/Microsoft/scalar.git /tmp/scalar.git &&\n\tmv /tmp/scalar.git/objects/pack/*.pack /tmp/scalar.git/my.pack &&\n\tgit hyperfine \\\n\t\t-r 2 --warmup 1 \\\n\t\t-L rev origin/master,HEAD -L v \"10,512,1k,1m\" \\\n\t\t-s 'make' \\\n\t\t-p 'git init --bare dest.git' \\\n\t\t-c 'rm -rf dest.git' \\\n\t\t'./git -C dest.git -c core.bigFileThreshold={v} unpack-objects </tmp/scalar.git/my.pack'\n\nHere we'll perform worse with lower core.bigFileThreshold settings\nwith this change in terms of speed, but we're getting lower memory use\nin return:\n\n\tSummary\n\t  './git -C dest.git -c core.bigFileThreshold=10 unpack-objects </tmp/scalar.git/my.pack' in 'origin/master' ran\n\t    1.01 ± 0.01 times faster than './git -C dest.git -c core.bigFileThreshold=1k unpack-objects </tmp/scalar.git/my.pack' in 'origin/master'\n\t    1.01 ± 0.01 times faster than './git -C dest.git -c core.bigFileThreshold=1m unpack-objects </tmp/scalar.git/my.pack' in 'origin/master'\n\t    1.01 ± 0.02 times faster than './git -C dest.git -c core.bigFileThreshold=1m unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n\t    1.02 ± 0.00 times faster than './git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/scalar.git/my.pack' in 'origin/master'\n\t    1.09 ± 0.01 times faster than './git -C dest.git -c core.bigFileThreshold=1k unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n\t    1.10 ± 0.00 times faster than './git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n\t    1.11 ± 0.00 times faster than './git -C dest.git -c core.bigFileThreshold=10 unpack-objects </tmp/scalar.git/my.pack' in 'HEAD'\n\nA better benchmark to demonstrate the benefits of that this one, which\ncreates an artificial repo with a 1, 25, 50, 75 and 100MB blob:\n\n\trm -rf /tmp/repo &&\n\tgit init /tmp/repo &&\n\t(\n\t\tcd /tmp/repo &&\n\t\tfor i in 1 25 50 75 100\n\t\tdo\n\t\t\tdd if=/dev/urandom of=blob.$i count=$(($i*1024)) bs=1024\n\t\tdone &&\n\t\tgit add blob.* &&\n\t\tgit commit -mblobs &&\n\t\tgit gc &&\n\t\tPACK=$(echo .git/objects/pack/pack-*.pack) &&\n\t\tcp \"$PACK\" my.pack\n\t) &&\n\tgit hyperfine \\\n\t\t--show-output \\\n\t\t-L rev origin/master,HEAD -L v \"512,50m,100m\" \\\n\t\t-s 'make' \\\n\t\t-p 'git init --bare dest.git' \\\n\t\t-c 'rm -rf dest.git' \\\n\t\t'/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold={v} unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum'\n\nUsing this test we'll always use >100MB of memory on\norigin/master (around ~105MB), but max out at e.g. ~55MB if we set\ncore.bigFileThreshold=50m.\n\nThe relevant \"Maximum resident set size\" lines were manually added\nbelow the relevant benchmark:\n\n  '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=50m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'origin/master' ran\n        Maximum resident set size (kbytes): 107080\n    1.02 ± 0.78 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'origin/master'\n        Maximum resident set size (kbytes): 106968\n    1.09 ± 0.79 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=100m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'origin/master'\n        Maximum resident set size (kbytes): 107032\n    1.42 ± 1.07 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=100m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'HEAD'\n        Maximum resident set size (kbytes): 107072\n    1.83 ± 1.02 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=50m unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'HEAD'\n        Maximum resident set size (kbytes): 55704\n    2.16 ± 1.19 times faster than '/usr/bin/time -v ./git -C dest.git -c core.bigFileThreshold=512 unpack-objects </tmp/repo/my.pack 2>&1 | grep Maximum' in 'HEAD'\n        Maximum resident set size (kbytes): 4564\n\nThis shows that if you have enough memory this new streaming method is\nslower the lower you set the streaming threshold, but the benefit is\nmore bounded memory use.\n\nAn earlier version of this patch introduced a new\n\"core.bigFileStreamingThreshold\" instead of re-using the existing\n\"core.bigFileThreshold\" variable[1]. As noted in a detailed overview\nof its users in [2] using it has several different meanings.\n\nStill, we consider it good enough to simply re-use it. While it's\npossible that someone might want to e.g. consider objects \"small\" for\nthe purposes of diffing but \"big\" for the purposes of writing them\nsuch use-cases are probably too obscure to worry about. We can always\nsplit up \"core.bigFileThreshold\" in the future if there's a need for\nthat.\n\n0. https://github.com/avar/git-hyperfine/\n1. https://lore.kernel.org/git/20211210103435.83656-1-chiyutianyi@gmail.com/\n2. https://lore.kernel.org/git/20220120112114.47618-5-chiyutianyi@gmail.com/\n\nHelped-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\nHelped-by: Derrick Stolee <stolee@gmail.com>\nHelped-by: Jiang Xin <zhiyou.jx@alibaba-inc.com>\nSigned-off-by: Han Xin <chiyutianyi@gmail.com>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n Documentation/config/core.txt   |  4 +-\n builtin/unpack-objects.c        | 69 ++++++++++++++++++++++++++++++++-\n t/t5351-unpack-large-objects.sh | 43 ++++++++++++++++++--\n 3 files changed, 109 insertions(+), 7 deletions(-)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex 87e4c04836..3ea3124f7f 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -468,8 +468,8 @@ usage, at the slight expense of increased disk usage.\n * Will generally be streamed when written, which avoids excessive\n memory usage, at the cost of some fixed overhead. Commands that make\n use of this include linkgit:git-archive[1],\n-linkgit:git-fast-import[1], linkgit:git-index-pack[1] and\n-linkgit:git-fsck[1].\n+linkgit:git-fast-import[1], linkgit:git-index-pack[1],\n+linkgit:git-unpack-objects[1] and linkgit:git-fsck[1].\n \n core.excludesFile::\n \tSpecifies the pathname to the file that contains patterns to\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 32e8b47059..43789b8ef2 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -351,6 +351,68 @@ static void unpack_non_delta_entry(enum object_type type, unsigned long size,\n \t\twrite_object(nr, type, buf, size);\n }\n \n+struct input_zstream_data {\n+\tgit_zstream *zstream;\n+\tunsigned char buf[8192];\n+\tint status;\n+};\n+\n+static const void *feed_input_zstream(struct input_stream *in_stream,\n+\t\t\t\t      unsigned long *readlen)\n+{\n+\tstruct input_zstream_data *data = in_stream->data;\n+\tgit_zstream *zstream = data->zstream;\n+\tvoid *in = fill(1);\n+\n+\tif (in_stream->is_finished) {\n+\t\t*readlen = 0;\n+\t\treturn NULL;\n+\t}\n+\n+\tzstream->next_out = data->buf;\n+\tzstream->avail_out = sizeof(data->buf);\n+\tzstream->next_in = in;\n+\tzstream->avail_in = len;\n+\n+\tdata->status = git_inflate(zstream, 0);\n+\n+\tin_stream->is_finished = data->status != Z_OK;\n+\tuse(len - zstream->avail_in);\n+\t*readlen = sizeof(data->buf) - zstream->avail_out;\n+\n+\treturn data->buf;\n+}\n+\n+static void stream_blob(unsigned long size, unsigned nr)\n+{\n+\tgit_zstream zstream = { 0 };\n+\tstruct input_zstream_data data = { 0 };\n+\tstruct input_stream in_stream = {\n+\t\t.read = feed_input_zstream,\n+\t\t.data = &data,\n+\t};\n+\tstruct obj_info *info = &obj_list[nr];\n+\n+\tdata.zstream = &zstream;\n+\tgit_inflate_init(&zstream);\n+\n+\tif (stream_loose_object(&in_stream, size, &info->oid))\n+\t\tdie(_(\"failed to write object in stream\"));\n+\n+\tif (data.status != Z_STREAM_END)\n+\t\tdie(_(\"inflate returned (%d)\"), data.status);\n+\tgit_inflate_end(&zstream);\n+\n+\tif (strict) {\n+\t\tstruct blob *blob = lookup_blob(the_repository, &info->oid);\n+\n+\t\tif (!blob)\n+\t\t\tdie(_(\"invalid blob object from stream\"));\n+\t\tblob->object.flags |= FLAG_WRITTEN;\n+\t}\n+\tinfo->obj = NULL;\n+}\n+\n static int resolve_against_held(unsigned nr, const struct object_id *base,\n \t\t\t\tvoid *delta_data, unsigned long delta_size)\n {\n@@ -483,9 +545,14 @@ static void unpack_one(unsigned nr)\n \t}\n \n \tswitch (type) {\n+\tcase OBJ_BLOB:\n+\t\tif (!dry_run && size > big_file_threshold) {\n+\t\t\tstream_blob(size, nr);\n+\t\t\treturn;\n+\t\t}\n+\t\t/* fallthrough */\n \tcase OBJ_COMMIT:\n \tcase OBJ_TREE:\n-\tcase OBJ_BLOB:\n \tcase OBJ_TAG:\n \t\tunpack_non_delta_entry(type, size, nr);\n \t\treturn;\ndiff --git a/t/t5351-unpack-large-objects.sh b/t/t5351-unpack-large-objects.sh\nindex 8d84313221..8ce8aa3b14 100755\n--- a/t/t5351-unpack-large-objects.sh\n+++ b/t/t5351-unpack-large-objects.sh\n@@ -9,7 +9,8 @@ test_description='git unpack-objects with large objects'\n \n prepare_dest () {\n \ttest_when_finished \"rm -rf dest.git\" &&\n-\tgit init --bare dest.git\n+\tgit init --bare dest.git &&\n+\tgit -C dest.git config core.bigFileThreshold \"$1\"\n }\n \n test_expect_success \"create large objects (1.5 MB) and PACK\" '\n@@ -17,7 +18,10 @@ test_expect_success \"create large objects (1.5 MB) and PACK\" '\n \ttest_commit --append foo big-blob &&\n \ttest-tool genrandom bar 1500000 >big-blob &&\n \ttest_commit --append bar big-blob &&\n-\tPACK=$(echo HEAD | git pack-objects --revs pack)\n+\tPACK=$(echo HEAD | git pack-objects --revs pack) &&\n+\tgit verify-pack -v pack-$PACK.pack >out &&\n+\tsed -n -e \"s/^\\([0-9a-f][0-9a-f]*\\).*\\(commit\\|tree\\|blob\\).*/\\1/p\" \\\n+\t\t<out >obj-list\n '\n \n test_expect_success 'set memory limitation to 1MB' '\n@@ -26,16 +30,47 @@ test_expect_success 'set memory limitation to 1MB' '\n '\n \n test_expect_success 'unpack-objects failed under memory limitation' '\n-\tprepare_dest &&\n+\tprepare_dest 2m &&\n \ttest_must_fail git -C dest.git unpack-objects <pack-$PACK.pack 2>err &&\n \tgrep \"fatal: attempting to allocate\" err\n '\n \n test_expect_success 'unpack-objects works with memory limitation in dry-run mode' '\n-\tprepare_dest &&\n+\tprepare_dest 2m &&\n \tgit -C dest.git unpack-objects -n <pack-$PACK.pack &&\n \ttest_stdout_line_count = 0 find dest.git/objects -type f &&\n \ttest_dir_is_empty dest.git/objects/pack\n '\n \n+test_expect_success 'unpack big object in stream' '\n+\tprepare_dest 1m &&\n+\tgit -C dest.git unpack-objects <pack-$PACK.pack &&\n+\ttest_dir_is_empty dest.git/objects/pack\n+'\n+\n+BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n+\n+test_expect_success 'unpack big object in stream (core.fsyncmethod=batch)' '\n+\tprepare_dest 1m &&\n+\tGIT_TRACE2_EVENT=\"$(pwd)/trace2.txt\" \\\n+\t\tgit -C dest.git $BATCH_CONFIGURATION unpack-objects <pack-$PACK.pack &&\n+\tgrep fsync/hardware-flush trace2.txt &&\n+\ttest_dir_is_empty dest.git/objects/pack &&\n+\tgit -C dest.git cat-file --batch-check=\"%(objectname)\" <obj-list >current &&\n+\tcmp obj-list current\n+'\n+\n+test_expect_success 'do not unpack existing large objects' '\n+\tprepare_dest 1m &&\n+\tgit -C dest.git index-pack --stdin <pack-$PACK.pack &&\n+\tgit -C dest.git unpack-objects <pack-$PACK.pack &&\n+\n+\t# The destination came up with the exact same pack...\n+\tDEST_PACK=$(echo dest.git/objects/pack/pack-*.pack) &&\n+\ttest_cmp pack-$PACK.pack $DEST_PACK &&\n+\n+\t# ...and wrote no loose objects\n+\ttest_stdout_line_count = 0 find dest.git/objects -type f ! -name \"pack-*\"\n+'\n+\n test_done\n-- \n2.36.1\n\n"},{"id":"458316","messageId":"xmqqk08xa8ww.fsf@gitster.g","threadId":"56672","inReplyTo":"5a4782d746a496e8edd1654296bac390d8e1c9d3.1654914555.git.chiyutianyi@gmail.com","subject":"Re: [PATCH v15 6/6] unpack-objects: use stream_loose_object() to unpack large objects","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-07-01T02:01:03Z","receivedAt":"2022-07-01T02:01:11Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Han Xin <chiyutianyi@gmail.com> writes:\n\n> +BATCH_CONFIGURATION='-c core.fsync=loose-object -c core.fsyncmethod=batch'\n> +\n> +test_expect_success 'unpack big object in stream (core.fsyncmethod=batch)' '\n> +\tprepare_dest 1m &&\n> +\tGIT_TRACE2_EVENT=\"$(pwd)/trace2.txt\" \\\n> +\t\tgit -C dest.git $BATCH_CONFIGURATION unpack-objects <pack-$PACK.pack &&\n> +\tgrep fsync/hardware-flush trace2.txt &&\n> +\ttest_dir_is_empty dest.git/objects/pack &&\n> +\tgit -C dest.git cat-file --batch-check=\"%(objectname)\" <obj-list >current &&\n> +\tcmp obj-list current\n> +'\n\nThis test without any prerequisite expects that \"hardware-flush\"\nwill always appear in the trace, but is that reasonable?  Don't\nwe need either \n\n (1) some sort of prerequisite to make sure this test piece runs\n     only on platforms that will use hardware-flush, or\n\n (2) loosen grep pattern to look for just \"fsync/\", or\n\n (3) something else?\n\nIt will become even worse when we queue Ævar's \"trace2 squelch\"\npatch on top, as we will stop emitting trace entries for that did\nnot trigger.\n"}]}